EDBT 2026 Demo / reviewers in the wild / expert
Yuting Yang 0008
dblp:25/3635-8
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-6720-4134ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compositional Attribute Imbalance in Vision DatasetsabstractVisual attribute imbalance is a common yet underexplored issue in image classification, significantly impacting model performance and generalization. In this work, we first define the first-level and second-level attributes of images and then introduce a CLIP-based framework to construct a visual attribute dictionary, enabling automatic evaluation of image attributes. By systematically analyzing both single-attribute imbalance and compositional attribute imbalance, we reveal how the rarity of attributes affects model performance. To tackle these challenges, we propose adjusting the sampling probability of samples based on the rarity of their compositional attributes. This strategy is further integrated with various data augmentation techniques (such as CutMix, Fmix, and SaliencyMix) to enhance the model's ability to represent rare attributes. Extensive experiments on benchmark datasets demonstrate that our method effectively mitigates attribute imbalance, thereby improving the robustness and fairness of deep neural networks. Our research highlights the importance of modeling visual attribute distributions and provides a scalable solution for long-tail image classification tasks. Yanbiao Ma, Wei Dai 0015, Yuting Yang 0008, Bowei Liu, Jiaxuan Zhao, Andi Zhang 0004 |
AAAI | 6 |
| 2026 | ACA-Net: Adaptive cloud-aware network for remote sensing image thick cloud removal
Baopu Hou, Xin Dang, Jinguang Wang, Quankai Zhao, Hongjia Qu, Yuting Yang 0008, Xiaoxuan Chen, Bo Jiang 0014 |
Expert Syst. Appl. | 7 |
| 2026 | Direction-aware frequency domain network for fine-grained cloud removal in remote sensing images
Junhao Jia, Hongjia Qu, Shengmei Chen, Yuting Yang 0008, Xiaoxuan Chen, Bo Jiang 0014 |
Expert Syst. Appl. | 6 |
| 2025 | Local-Global Spectral Feature-Aware Learning for Hyperspectral Imagery ClassificationabstractEffective modeling of the relationship between local spectral details and global contextual information remains a core challenge in hyperspectral image (HSI) classification. In this paper, a local-global spectral feature-aware network (LGSFA-Net) is proposed, which achieves local and global spectral feature learning through synergistic integration of local convolutional inductive biases and global state-space models (SSMs). The architecture of LGSFA-Net comprises three sequential components, including an embedding stage using convolutions for fundamental feature extraction, an encoding stage that cascades standard Mamba blocks with specialized interactive Mamba (IMamba) blocks and an enhanced spatial-spectral feature fusion (ESSFF) module. The proposed IMamba blocks employ separable convolutions and feature interaction learning for explicitly modeling the cross-channel spectral correlations learning, which can be effective in awareness of the spectral feature. And then, the ESSFF module utilizes self-attention mechanisms to dynamically balance local and global spatial-spectral feature contributions. The final prediction stage incorporates a lightweight classification head for efficient inference. Experimental results validate the effectiveness of the proposed methods for HSI classification on four benchmark datasets, including the PaviaU, Houston, Honghu, and Hanchuan datasets. The proposed LGSFA-Net achieves approximately 1.48%-2.51% increased overall accuracy (OA), 1.34%-2.06% increased average accuracy(AA), and 1.37%-3.75% increased Kappa on the aforementioned four datasets, respectively, outperforming the contrasting methods. The code implementation will be available at https://github.com/yutinyang/LGSFA-Net. Yuting Yang 0008, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuo Li 0010, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Multimodal Segformer for Flood Rapid Mapping with Sentinel-2 DataabstractFlood rapid mapping products play an important role in informing flood emergency response and management. To this end, the 2024 IEEE GRSS Data Fusion Contest Track 2 (DFC24-T2) establishes a multimodal benchmark for the segmentation of flood areas from Sentinel-2 multispectral images. However, the problems of imbalanced data distribution, data scarsity, and inter-modal differences severely inhibit the performance of deep-learning-based segmentation networks. In this work, we propose an end-to-end Multimodal Transformer-based Segmentation Network (MTSN) for accurate flood rapid mapping. MTSN first employs two Siamese encoders with shared parameters to accept multimodal inputs and output their respective hierarchical multiscale features, which are then enriched by several channel attention blocks. Subsequently, a Cross-modal Feature Fusion Module (CFFM) based on a gated mechanism is proposed to efficiently integrate the benefits of multimodal features, and generate informative representations. Finally, the fused features are decoded by a lightweight pure multilayer perception decoder to quickly generate mapping results of flood areas. Moreover, we introduce offline data augmentation, semi-supervised learning, test-time augmentation, and multimodal post-process to further boost the performance and generalization of our MTSN. Experimental results and extensive ablations show the effectiveness of our method. Code is available at https://github.com/xiaoqiang-lu/MMSegFormer. Xiaoqiang Lu, Tong Gou, Zhongjian Huang, Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001 |
IGARSS | 4 |
| 2024 | Efficient LWPooling: Rethinking the Wavelet Pooling for Scene ParsingabstractExisting wavelet pooling methods discard the high-frequency sub-bands, which can improve the noise-robustness of convolutional neural networks (CNNs) but lose the essential detailed features. Besides, most of them depend on different wavelets, which is not adaptive. In this paper, a novel efficient lifting-based wavelet pooling (LWPooling) is proposed to alleviate the problems above. Firstly, wavelet pooling is rethought based on the equivalence of 2D discrete wavelet transform (DWT) and standard average pooling (SAP), which suggests the lack of detailed information on traditional wavelet pooling. Secondly, the efficient LWPooling module is proposed to adaptively capture and preserve the critical high-frequency features via lifting-based wavelets. It can constrain the features linear independence, which efficiently makes important features salient. Thirdly, the lifting-based wavelet collaborative network (LWCNet) is constructed for classification and segmentation tasks based on the efficient LWPooling module. Experiments are validated on Cifar10, Cifar100, and ADE20K datasets. It suggests that the efficient LWPooling can enhance CNN’s representation and achieve a particular performance advantage compared to average, maximum, and original wavelet pooling. Besides, the proposed LWCNet shows the potential for scene parsing. The code implementation will be available at https://github.com/yutinyang/LWCNet. Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | LGLFormer: Local-Global Lifting Transformer for Remote Sensing Scene ParsingabstractIn deep learning, convolutional neural networks (CNNs) and transformers have gained excellent achievements in remote sensing scene parsing. Strong feature representation ability is still a challenge for them. Besides, the complex scenes are still essential challenges for deep learning in remote sensing scene parsing. In this article, an efficient local–global lifting transformer (LGLFormer) framework is proposed to ease the challenges above. It effectively combines CNNs, transformer, and wavelet transform to build a strong local–global (LG) feature representation network. Besides, global feature learning driven by LG adaptive features is proposed based on the 2-D LG adaptive feature extractor (LGAFE) and refined global feature attention module. The 2-D LG lifting feature extractor is inspired by the lifting scheme, which introduces local and global dependency. Furthermore, two LG lifting schemes are proposed, including the series and parallel modes, which can effectively learn LG relations between pixels. Finally, experiments are validated on three remote sensing benchmark datasets. The proposed LGLFormer achieves the state-of-the-art with 99.02%, 99.2%, and 99.48% overall accuracy (OA) on AID, WHU-RS19, and UCM datasets, respectively. In addition, LGLFormer shows good convergence with competitive parameters. The experimental code will be available athttps://github.com/yutinyang/LGLFormer. Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Strong Vision Transformer Adapter with Adaptive Thresholding for fine-Grained Building ClassificationabstractFine-grained building classification provides a solid basis for the comparison of city morphologies and the investigation of urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for the classification of building roof types. However, the problems of long-tailed distribution, data insufficient, inter-class similarity, and intra-class difference severely inhibit the performance of the detector. In this work, we build a strong vision transformer adapter fine-tuned on the cropped building instances to enhance the capacity of feature extraction and design a cross-modal fusion (CMF) module to effectively aggregate features from RGB and SAR data. When transferring to building instance segmentation, we construct a robust training pipeline and a two-stage test-time results ensemble scheme. Furthermore, we introduce self-training with two key denoising techniques, global average filtering (GAF) and intra-class adaptive thresholding (IAT), to boost the generalization of the model. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008 |
IGARSS | 7 |
| 2023 | Trident Cooperation Network for Building Extraction and Height EstimationabstractBuilding extraction and height estimation provide solid fundamentals for reconstructing city morphologies and investigating urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for multi-task learning of building reconstruction. However, the problems of data limitation and fore-background confusion severely inhibit the performance of the model. In this work, we propose a novel trident cooperation network (TCNet) to perform end-to-end building extraction and height estimation using RGB and SAR data. Specifically, to enrich the feature representation and generalization of the shared backbone, we introduce a vision transformer adapter to inject vision-specific inductive biases and design a cross-modal fusion (CMF) module to effectively aggregate features from multi-modal data. For downstream visual tasks, we construct trident decoders including a detector, a lightweight MLP segmentation head, and a pixel-wise regression head. Moreover, to highlight the foreground object, we use the binary mask predicted by the MLP head to cooperate with the height estimation map predicted by the estimator. And the weighted sub-task losses are gathered to optimize our TCNet. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008 |
IGARSS | 7 |
| 2023 | Dense Cross-Scale Transformer with Channel Learning for Remote Sensing Scene ClassificationabstractThe recent explosive on Transformer has suggested its potential to become the mainstream model for feature representation and classification of remote sensing. Although Transformer has excellent global modeling capabilities, it lacks inductive bias. In contrast, CNNs show excellent performance in computer vision due to their strong inductive bias. To solve the above issue, and benefit from combining the Transformer model with CNNs, a dense cross-scale Transformer is proposed for remote sensing scene classification. Firstly, the attention aggregation based feature pyramid network (A2-FPN) is adopted to obtain the multi-scale features. And then, the multi-scale features are input into the multi-scale global features learning with channel attention (MSCA) module to obtain the global features. Besides, the multi-scale features are input into the dense cross-scale attention (DCSA) module to learn the multi-level cross-scale features. Finally, the outputs of these two modules are considered for computing the final class score. Experimental results obtained from the three public datasets indicate that the proposed method surpasses other remote sensing classification methods. Yuting Yang 0008, Xu Liu 0006, Wenping Ma 0001, Licheng Jiao |
IGARSS | 2 |
| 2023 | Dual Wavelet Attention Networks for Image ClassificationabstractGlobal average pooling (GAP) plays an important role in traditional channel attention. However, there is the disadvantage of insufficient information to use the result of GAP as the channel scalar. At the same time, the existing spatial attention models focus on the areas of interest using average pooling or convolutional networks, but there is a loss of feature information and neglect of the structural feature. In this paper, dual wavelet attention is proposed, which can effectively alleviate the aforementioned problems and enhance the representation ability of CNNs. Firstly, the equivalence between the sum of the low-frequency subband coefficients of 2D DWT (Haar) and GAP is proved. On this basis, the statistical characteristics of low-frequency and high-frequency subbands are effectively combined to obtain the channel scalars, which can better measure the importance of each channel. In addition, 2D DWT can effectively capture the approximate and detailed structural features. Thus, wavelet spatial attention is proposed, which can effectively focus on the key spatial structural features. Different from traditional spatial attention, it can better curve the structural and spatial attention for different channels. The experiments are verified on four natural image data sets and three remote sensing scene classification data sets, which shows the effectiveness and versatility of the proposed methods. The code of this paper will be available athttps://github.com/yutinyang/DWAN. Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Lingling Li 0002, Puhua Chen, Xiufang Li, Zhongjian Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | An Explainable Spatial-Frequency Multiscale Transformer for Remote Sensing Scene ClassificationabstractDeep convolutional neural networks (CNNs) are significant in remote sensing. Due to the strong local representation learning ability, CNNs have excellent performance in remote sensing scene classification. However, CNNs focus on location-sensitive representations in the spatial domain and lack contextual information mining capabilities. Meanwhile, remote sensing scene classification still faces challenges, such as complex scenes and significant differences in target sizes. To address the problems and challenges above, more robust feature representation learning networks are necessary. In this paper, a novel and explainable spatial-frequency multi-scale Transformer framework, SF-MSFormer, is proposed for remote sensing scene classification. It mainly comprises spatial-domain and frequency-domain multi-scale Transformer branches, which consider the spatial-frequency global multi-scale representation features. Besides, the texture-enhanced encoder is designed in the frequency-domain multi-scale Transformer branch, which is adaptive to capture the global texture features. In addition, an adaptive feature aggregation module is designed to integrate the spatial-frequency multi-scale feature for final recognition. The experimental results verify the effectiveness of SF-MSFormer and show better convergence. It achieves state-of-the-art results (98.72%, 98.6%, 99.72%, and 94.83% overall accuracies, respectively) on the AID, UCM, WHU-RS19, and NWPU-RESISC45 datasets. Besides, the feature visualizations evaluate the explainability of the texture-enhanced encoder. The code implementation of this article will be available at https://github.com/yutinyang/SF-MSFormer. Yuting Yang 0008, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | MRTA: Multi-Resolution Training Algorithm for Multitemporal Semantic Change DetectionabstractThe multitemporal semantic change detection challenge track (Track MSD) in the 2021 Data Fusion Contest is to extract the land cover changes of the US state of Maryland from 2013 to 2017, but only low-resolution label data is provided. We present a multi-resolution training algorithm (MRTA) to alleviate the overfitting of the model on the coarse labels. First, using low-resolution coarse labels to train FCN, the average IoU can reach 0.5253. Generating pseudo-labels using this network, they are combined with coarse labels to form a multi-resolution label combination, and perform iterative fine-tuning and retrain. After that, by analyzing the loss and gain indicators of specific categories, it was found that the discrimination effect of the water area was poor, so a strong classifier was trained for the water area. We also implemented strategies such as model voting and weighted training to improve model performance. Finally, our method achieves 0.6445 mIoU on the test set, ranking 3rd in Track 2 of the IEEE Data Fusion Contest of 2021. Qianyue Bao, Yang Liu 0349, Zixiao Zhang, Dafan Chen, Yuting Yang 0008, Licheng Jiao, Fang Liu 0001 |
IGARSS | 5 |
| 2021 | Multisource Data Fusion for the Detection of Settlements Without ElectricityabstractThe international charity SolarAid aims to provide access to lights in areas without electricity, and it is a challenge to accurately and efficiently transmit the lights to the areas in need. Multisource, multitemporal, and multimodal remote sensing images can provide rich information about the target area, so using multisource remote sensing images for accurate detection of human settlements without electricity is a feasible solution. In this paper two separate detection tasks are formulated: building two attention SENet for settlements detection and light detection using the Sentinel-2 dataset and the Suomi Visible Infrared Imaging Radiometer Suite (VIIRS) night time dataset, respectively. In addition, we study a new outlier removal method based on the pixel distribution characteristics of the VIIRS dataset for data pre-processing, and propose a post-processing method based on region continuity for further correction of the results. Experiments show that our method can maximize the use of multisource data information and rank first in the detection of settlements without electricity challenge track (Track DSE) of the 2021 IEEE GRSS Data Fusion Contest. Yanbiao Ma, Kexin Feng, Xueli Geng, Licheng Jiao, Fang Liu 0001, Yuting Yang 0008 |
IGARSS | 7 |