Yang Liu 0352

dblp:51/3710-352 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-0010-5200ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge Integration
abstract
Multi-task model merging aims to consolidate knowledge from multiple fine-tuned task-specific experts into a unified model while minimizing performance degradation. Existing methods primarily approach this by minimizing differences between task-specific experts and the unified model, either from a parameter-level or a task-loss perspective. However, parameter-level methods exhibit a significant performance gap compared to the upper bound, while task-loss approaches entail costly secondary training procedures. In contrast, we observe that performance degradation closely correlates with feature drift, i.e., differences in feature representations of the same sample caused by model merging. Motivated by this observation, we propose Layer-wise Optimal Task Vector Merging (LOT Merging), a technique that explicitly minimizes feature drift between task-specific experts and the unified model in a layer-by-layer manner. LOT Merging can be formulated as a convex quadratic optimization problem, enabling us to analytically derive closed-form solutions for the parameters of linear and normalization layers. Consequently, LOT Merging achieves efficient model consolidation through basic matrix operations. Extensive experiments across vision and vision-language benchmarks demonstrate that LOT Merging significantly outperforms baseline methods, achieving improvements of up to 4.4% (ViT-B/32) over state-of-the-art approaches. The source code is available at https://github.com/SunWenJu123/model-merging.
Wenju Sun, Qingyong Li, Wen Wang 0019, Yang Liu 0352, Boyang Li 0001
NeurIPS4
2024 CLDiff: Weakly Supervised Cloud Detection With Denoising Diffusion Probabilistic Models
abstract
Cloud detection is an essential step in remote sensing (RS) image processing, contributing to various applications. However, existing fully supervised cloud detection methods rely on massive pixel-wise annotations, which are expensive and time-consuming. To alleviate the annotation burden, weakly supervised cloud detection (WSCD) has received extensive attention recently. One standard approach performs cloud detection within a classification paradigm, which inevitably faces category ambiguity when detecting semitransparent clouds. To tackle this problem, we propose a novel WSCD framework based on the diffusion model, termed CLDiff. Specifically, a multiscale feature rectification (MFR) module is introduced to extract multiscale semantic features in the encoder, enabling a definite identification of clouds and mitigating interference from bright objects in the background. Considering that clouds exhibit varying optical thicknesses, a diffusion decoder is developed to model the intraclass variations of clouds in a generative strategy, improving thin cloud detection. Initially, it devises a Gaussian modulation function to recalibrate ambiguous cloud activations and emphasize semitransparent clouds. Subsequently, these modulated activations serve as semantic guidance to optimize the diffusion process. This approach enables CLDiff to activate cloud contours under definite semantic conditions and avoids the additional branches for semantic learning as found in previous methods. Experimental results demonstrate that CLDiff achieves state-of-the-art performance in WSCD. A public reference implementation of this work in PyTorch is available athttps://github.com/YLiu-creator/CLDiff.
Yang Liu 0352, Qingyong Li, Zhigang Yao, Tony Z. Qiu, Wen Wang 0019
IEEE Trans. Geosci. Remote. Sens.1
2023 Confidence-adapted meta-interaction for unsupervised person re-identification
Xiaobao Li, Qingyong Li, Wenyuan Xue, Yang Liu 0352, Fengjiao Liang, Wen Wang 0019
Appl. Intell.4
2023 A General Dual-Branch Framework for Land Cover Mapping Models With Multispectral Data
abstract
Land cover mapping based on multispectral images can, in principle, be considered an application of semantic segmentation, but land cover mapping inputs include near-infrared (NIR) data in addition to RGB data. It has been experimentally found that LULC mapping performance based on RGB data alone is better than that based on RGB and NIR data when using some established single-branch encoder–decoder models. To address this issue, we propose a dual-branch encoder–decoder (DBED) framework that can be applied to existing encoder–decoder models for semantic segmentation. First, the multispectral input data is divided into two parts: RGB and NIR and fed to the respective branch for encoding. The dual-branch structure facilitates cross-modal complementary information encoding without deteriorating the original RGB modality-specific feature extraction. Second, an attention module named the multispectral attention module (MSAM) is proposed to mine the contextual correlation between the multispectral feature maps, leading to further performance boosting. We apply this framework to three mainstream semantic segmentation models and validate it on the Gaofen Image Dataset (GID). Experimental results show that this structure brings performance improvements. The source code of DBED is publicly available athttps://github.com/MalignusCN/DBED.
Qingyong Li, Yang Liu 0352, Wen Wang 0019
IEEE Geosci. Remote. Sens. Lett.3
2023 Leveraging Physical Rules for Weakly Supervised Cloud Detection in Remote Sensing Images
abstract
Cloud detection plays a significant role in remote sensing image applications. Existing deep learning-based cloud detection methods rely on massive precise pixel-wise annotations, which are time-consuming and expensive. To alleviate this problem, we propose a weakly supervised cloud detection framework that leverages physical rules to generate weak supervision for cloud detection in remote sensing images. Specifically, a rule-based adaptive pseudo labeling (RAPL) algorithm is devised to adaptively annotate potential cloud pixels based on cloud spectral properties without manual intervention. Unlike existing physical annotations using fixed thresholds, RAPL employs the bidirectional threshold segmentation and adaptive gating mechanism to annotate cloud and boundary masks with more explicit semantic categories and spatial structures separately. Subsequently, these pseudo masks are treated as weak supervision to optimize the heuristic cloud detection network for pixel-wise segmentation. Considering that clouds appear as complex geometric structures and nonuniform spectral reflectance, a deformable boundary refining module is designed to enhance the modeling ability of spatial transformation and activate sharp boundaries from translucent cloud regions. Moreover, a harmonic loss is employed to recognize clouds with nonuniform spectral reflectance and suppress the interference of bright backgrounds. Extensive experiments on the GF-1, L8 Biome, and WDCD datasets demonstrate that the proposed method achieves state-of-the-art results. A public reference implementation of this work in PyTorch is available at https://github.com/NiAn-creator/HeuristicCloudDetection.
Yang Liu 0352, Qingyong Li, Xiaobao Li, Shuyi He, Fengjiao Liang, Zhigang Yao, Wen Wang 0019
IEEE Trans. Geosci. Remote. Sens.1
2022 Semantic Segmentation of Remote Sensing Images With Self-Supervised Semantic-Aware Inpainting
abstract
Semantic segmentation of remote sensing imageries plays a crucial role in resource exploration, urban planning, weather forecasting, etc. For this task, deep learning-based methods have shown significant achievement, typically trained with large-scale labeled data. However, these methods often suffer the performance deterioration facing limited labeled data in real-world applications. To address this problem, a novel self-supervised semantic segmentation framework is proposed for remote sensing imageries with limited labeled data. Specifically, image inpainting is acted as pixel-level pretext task for learning dense feature representations suitable for semantic segmentation. Further, rather than trivially leveraging the conventional random inpainting strategy, a novel adversarial training scheme is proposed to drive the pretext task to adaptively mask and restore salient local regions. The adversarial training scheme consists of instructor network and inpainting network, the instructor network increasingly predicts meaningful salient regions as erased regions, and meanwhile the inpainting network seeks for restoring the corrupted image as pretext task to learn its intrinsic representation. Moreover, the structural similarity (SSIM) is applied as a patch-level loss function for semantic segmentation considering that remote sensing images are highly structured. The experimental results on the ISPRS Potsdam dataset demonstrate that our method outperforms state-of-the-art self-supervised methods and the ImageNet pre-training methods. The source code is available at https://github.com/JasmineBJTU/self-supervised_RSSS.
Shuyi He, Qingyong Li, Yang Liu 0352, Wen Wang 0019
IEEE Geosci. Remote. Sens. Lett.3
2022 DCNet: A Deformable Convolutional Cloud Detection Network for Remote Sensing Imagery
abstract
Recently, deep convolutional neural networks (CNNs) have made important progress in cloud detection with powerful representation learning capability and yield significant performance. However, most existing CNN-based cloud detection methods still face serious challenges because of the variable geometry of clouds and the complexity of underlying surfaces. It is attributed that they only use the fixed grid to extract contextual information, which lacks internal mechanisms to handle the geometric transformations of clouds. To tackle this problem, we propose a deformable convolutional cloud detection network with an encoder-decoder architecture, named DCNet, which can enhance the adaptability of a model to cloud variations. Specifically, we introduce deformable convolution blocks at the encoder to capture saliency spatial contexts adaptively based on the morphological characteristics of clouds and generate high-level semantic representations. After this, we incorporate skip-connection mechanisms into the decoder that integrate low-level spatial contexts as guidance to recover high-level semantic pixel localization and export precise cloud-detection results. Extensive experiments on the GF-1 wide field-of-view (WFV) Satellite Imagery demonstrate that DCNet outperforms several state-of-the-art methods. A public reference implementation of our proposed model in PyTorch is available athttps://github.com/NiAn-creator/deformableCloudDetection.git.
Yang Liu 0352, Wen Wang 0019, Qingyong Li, Min Min, Zhigang Yao
IEEE Geosci. Remote. Sens. Lett.1