Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Enhancing Building Change Detection with UVT-BCD: A UNet-Vision Transformer Fusion Approach

View through CrossRef
Abstract Building change detection (BCD) is particularly important for comprehending ground changes and activities carried out by humans. Since its introduction, deep learning has emerged as the dominant method for BCD. Despite this, the detection accuracy continues to be inadequate because of the constraints imposed by feature extraction requirements. Consequently, the purpose of this study is to present a feature enhancement network that combines a UNet encoder and a vision transformer (UVT) structure in order to identify BCD (UVT-BCD). A deep convolutional network and a section of the vision transformer structure are combined in this model. The result is a strong feature extraction capability that can be used for a wide variety of building types. To improve the ability of small-scale structures to be detected, you should design an attention mechanism that takes into consideration both the spatial and channel dimensions. A cross-channel context semantic aggregation module is used to carry out information aggregation in the channel dimension. Experiments have been conducted in numerous cases using two different BCD datasets to evaluate the performance of the previously suggested model. The findings reveal that UVT-BCD outperforms existing approaches, achieving improvements of 5.95% in overall accuracy, 5.33% in per-class accuracy, and 8.28% in the Cohen's Kappa statistic for the LEVIR-CD dataset. Furthermore, it demonstrates enhancements of 6.05% and 6.4% in overall accuracy, 6.56% and 5.89% in per-class accuracy, and 6.71% and 6.23% in the Cohen's Kappa statistic for the WHU-CD dataset.
Research Square Platform LLC
Title: Enhancing Building Change Detection with UVT-BCD: A UNet-Vision Transformer Fusion Approach
Description:
Abstract Building change detection (BCD) is particularly important for comprehending ground changes and activities carried out by humans.
Since its introduction, deep learning has emerged as the dominant method for BCD.
Despite this, the detection accuracy continues to be inadequate because of the constraints imposed by feature extraction requirements.
Consequently, the purpose of this study is to present a feature enhancement network that combines a UNet encoder and a vision transformer (UVT) structure in order to identify BCD (UVT-BCD).
A deep convolutional network and a section of the vision transformer structure are combined in this model.
The result is a strong feature extraction capability that can be used for a wide variety of building types.
To improve the ability of small-scale structures to be detected, you should design an attention mechanism that takes into consideration both the spatial and channel dimensions.
A cross-channel context semantic aggregation module is used to carry out information aggregation in the channel dimension.
Experiments have been conducted in numerous cases using two different BCD datasets to evaluate the performance of the previously suggested model.
The findings reveal that UVT-BCD outperforms existing approaches, achieving improvements of 5.
95% in overall accuracy, 5.
33% in per-class accuracy, and 8.
28% in the Cohen's Kappa statistic for the LEVIR-CD dataset.
Furthermore, it demonstrates enhancements of 6.
05% and 6.
4% in overall accuracy, 6.
56% and 5.
89% in per-class accuracy, and 6.
71% and 6.
23% in the Cohen's Kappa statistic for the WHU-CD dataset.

Related Results

The Nuclear Fusion Award
The Nuclear Fusion Award
The Nuclear Fusion Award ceremony for 2009 and 2010 award winners was held during the 23rd IAEA Fusion Energy Conference in Daejeon. This time, both 2009 and 2010 award winners w...
Automatic Load Sharing of Transformer
Automatic Load Sharing of Transformer
Transformer plays a major role in the power system. It works 24 hours a day and provides power to the load. The transformer is excessive full, its windings are overheated which lea...
Machining simulation of titanium alloy under different cutting methods
Machining simulation of titanium alloy under different cutting methods
Abstract Titanium alloy is difficult to machine due to its low thermal conductivity, high strength, and low machining efficiency, posing a significant challenge to t...
Enhancing Brain Tumor Segmentation with Transformer-Based Models: A Study on the BraTS 2020 Dataset
Enhancing Brain Tumor Segmentation with Transformer-Based Models: A Study on the BraTS 2020 Dataset
Abstract Accurate segmentation of brain tumors from medical images is crucial for clin- ical diagnosis, treatment planning, and patient outcome prediction. While the UNet a...
Depth-aware salient object segmentation
Depth-aware salient object segmentation
Object segmentation is an important task which is widely employed in many computer vision applications such as object detection, tracking, recognition, and ret...
VM-UNet++ research on crack image segmentation based on improved VM-UNet
VM-UNet++ research on crack image segmentation based on improved VM-UNet
Abstract Cracks are common defects in physical structures, and if not detected and addressed in a timely manner, they can pose a severe threat to the overall safety of th...
Performance Comparison of 8-Digit BCD Adders using CLA and Brent–Kung Architectures
Performance Comparison of 8-Digit BCD Adders using CLA and Brent–Kung Architectures
Abstract - Binary Coded Decimal (BCD) adders play a crucial role in digital systems requiring precise decimal arithmetic, particularly in financial computing, commercial applicatio...

Back to Top