Dabing Yu

dblp:245/9891 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0003-1500-0276ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 VolFormer: Explore More Comprehensive Cube Interaction for Hyperspectral Image Restoration and Beyond
abstract
Capitalizing on the talent of self-attention in capturing non-local features, Transformer architectures have exhibited remarkable performance in single hyperspectral image restoration. For hyperspectral images, each pixel is located in the hyperspectral image cubes with a large spectral dimension and two spatial dimensions. Although uni-dimensional self-attention, like channel self-attention or spatial self-attention, builds long-range dependencies in spectral or spatial dimensions, they lack more comprehensive interactions across dimensions. To tackle the above drawback, we propose a VolFormer, a volumetric self-attention embedded Transformer network for single hyperspectral image restoration. Specifically, we propose volumetric self-attention (VolSA), which extends the interaction from 2D flat to 3D cube. VolSA can simultaneously model token interaction in the 3D cube, mining the potential correlations between the hyperspectral image cube. An attention decomposition form is proposed to reduce the computational burden of modeling volumetric information. In practical terms, VolSA adapts double similarity matrixes in spatial and channel dimensions to implicitly model 3D context information while transforming the complexity from cubic to quadratic. Additionally, we introduce the explicit spectral location prior to enhance the proposed self-attention. This property allows the target token to perceive global spectral information while simultaneously assigning different levels of attention to tokens at varying wavelength bands. Extensive experiments demonstrate that VolFormer achieves record-high performance on hyperspectral image super-resolution, denoise and classification benchmarks. The source code is available at https://github.com/yudadabing/VolFormer.
Dabing Yu
CVPR1
2023 DSTrans: Dual-Stream Transformer for Hyperspectral Image Restoration
abstract
Most CNN models exhibit two major flaws in hyper-spectral image (HSI) restoration tasks. First, limited high-dimensional HSI training examples exacerbate the difficulty of deep learning methods in learning effective spatial and spectral representations. Second, the existing CNN-based methods model local relations and present limitations in capturing long-range dependencies. In this paper, we customize a novel dual-stream Transformer (DSTrans) for HSI restoration, which mainly consists of the dual-stream attention and the dual-stream feed-forward network. Specifically, we develop the dual-stream attention consisting of Multi-Dconv-head spectral attention (MDSA) and Multi-head Spatial self-attention (MSSA). MDSA and MSSA respectively calculate self-attention along the spectral and spatial dimensions in local windows to capture long-range spectrum dependencies and model global spatial interactions. Meanwhile, the dual-stream feed-forward network is developed to extract global signals and local details in parallel branches. In addition, we exploit a multi-tasking network to train the auxiliary RGB image (RGBI) task and HSI task jointly so that both numerous RGBI samples and limited HSI samples are exploited to learn parameter distribution for DSTrans. Extensive experimental results demonstrate that our method achieves state-of-the-art results on HSI restoration tasks, including HSI super-resolution and denoising. The source code can be obtained at: https://github.com/yudadabing/Dual-Stream-Transformer-for-Hyperspectral-Image-Restoration.
Dabing Yu, Qingwu Li, Yixi Qian
WACV1
2023 ROV-based binocular vision system for underwater structure crack detection and width measurement
Qingwu Li, Yaqin Zhou, Dabing Yu
Multim. Tools Appl.5
2023 Visual saliency detection via invariant feature constrained stacked denoising autoencoder
Zhihong Yu, Yaqin Zhou, Dabing Yu
Multim. Tools Appl.5
2023 Dual-Space Graph-Based Interaction Network for RGB-Thermal Semantic Segmentation in Electric Power Scene
abstract
Real-time scene comprehension is the basis for automatic electric power inspection. However, existing RGB-based scene comprehension methods may achieve unsatisfied performance when dealing with complex scenarios, insufficient illumination or occluded appearances. To solve this problem, by cooperating visual and thermal images, the Dual-Space Graph-based Interaction Network (DSGBINet) is proposed to achieve all-day time semantic segmentation of power equipment in high-voltage power transmission line and electric transformer substation scenes. Specifically, modality-specific features are first extracted via two separate backbone networks with the same architecture. Multi-modality high-level features are first fused via long-range relationship in coordinate space. Then, multi-modality features from regular grids are further clustered and assigned to vertices in feature space. Cross-graph and inner-graph regional relations are utilized for reasoning and enhancement, which could exploit the mutual benefits and extract rich contextual information in a semantic view. Furthermore, to overcome the huge scale difference and the inherent characteristic of thermal images, the idea of multi-task learning is integrated into the decoding process. The edge detection and semantic segmentation are achieved collaboratively, which could segment the different power equipment more accurately and completely. The comparative and ablation experiments on the proposed two RGB-T semantic segmentation datasets evaluate the effectiveness and robustness of the proposed network compared with existing state-of-the-art methods. The extended experiments on the public datasets further demonstrate the superiority of the proposed method. Our dataset and code will be released at:https://github.com/hhujiang/DSGBINet.
Chang Xu 0022, Qingwu Li, Xiongbiao Jiang, Dabing Yu, Yaqin Zhou
IEEE Trans. Circuits Syst. Video Technol.4
2023 Enhanced visual perception for underwater images based on multistage generative adversarial network
Dabing Yu, Yaqin Zhou
Vis. Comput.2
2022 Asymmetric cross-modal activation network for RGB-T salient object detection
Chang Xu 0022, Qingwu Li, Qingkai Zhou, Xiongbiao Jiang, Dabing Yu, Yaqin Zhou
Knowl. Based Syst.5
2022 A Cross-Level Spectral-Spatial Joint Encode Learning Framework for Imbalanced Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have dominated the research of hyperspectral image (HSI) classification, attributing to the superior feature representation capacity. Patch-free global learning (FPGA) as a fast learning framework for HSI classification has received wide interest. Despite their promising results from the perspective of fast inference, recent works have difficulty modeling spectral-spatial relationships with imbalanced samples. In this paper, we revisit the encoder–decoder-based fully convolutional network (FCN) and propose a cross-level spectral-spatial joint encoding framework (CLSJE) for Imbalanced HSI classification. First, a multi-scale input encoder and multiple-to-one multi-scale features connection are introduced to obtain abundant features and facilitate multi-scale contextual information flow between encoder and decoder. Second, in the encoder layer, we propose the spectral-spatial joint attention (SSJA) mechanism consisting of the high-frequency spatial attention (HFSA) and spectral-transform channel attention (STCA). HFSA and STCA encode spectral-spatial features jointly to improve the learning of the discriminative spectral-spatial features. Powered by these two components, CLSJE enjoys a high capability to capture both spatial and spectral dependencies for HSI classification. Besides, a class-proportion sampling strategy is developed to increase the attention to insufficiency samples. Extensive experiments demonstrate the superiority of our proposed CLSJE both at classification accuracy and inference speed, and show the state-of-the-art results on four benchmark datasets. Code can be obtained at: https://github.com/yudadabing/CLSJE.
Dabing Yu, Qingwu Li, Chang Xu 0022, Yaqin Zhou
IEEE Trans. Geosci. Remote. Sens.1