Yuan Wang 0064

dblp:41/3241-64 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-8553-7901ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021
YearPublicationVenuePosition
2026 DA2-LiDAR: A Generic Density-Adaptive Framework for Unsupervised Domain Adaptation in LiDAR Segmentation
abstract
This paper addresses the critical challenge of domain adaptation for LiDAR-based semantic segmentation, particularly the significant density disparities that emerge when transferring models from synthetic to real-world environments. We present DA2-LiDAR, a novel density-adaptive domain adaptation framework that bridges domain gaps through the construction of intermediate domains with density-varying point distributions. Our approach employs a simple yet effective masking strategy that systematically reduces density discrepancies between domains while extracting more effective supervisory signals, as well as preserving critical semantic information. The framework consists of three key components: (1) a Density Adaptation Module that establishes a continuous spectrum of intermediate domains through dataset-agnostic masking operations; (2) a Contextual Consistency Module that enforces relational coherence across differently masked variants of the same scan at varying degrees, providing additional supervision signals, enhancing the model's ability to extract features; and (3) a Semantic Preservation Module that mitigates information loss in heavily masked scans by reconstructing domain-specific data distributions. Extensive experiments on synthetic-to-real and other benchmarks demonstrate that DA2-LiDAR consistently outperforms state-of-the-art methods, achieving significant improvements in cross-domain generalization without requiring dataset-specific prior knowledge or introducing computational overhead.
Rui Sun 0006, Wangkai Li, Naisong Luo, Yuan Wang 0064, Tianzhu Zhang 0001, Feng Wu 0005
IEEE Trans. Image Process.5
2025 Exploring the Better Multimodal Synergy Strategy for Vision-Language Models
abstract
Vision-Language models (VLMs) have shown great potential in enhancing open-world visual concept comprehension. Recent researches focus on an optimum multimodal collaboration strategy that significantly advances CLIP-based few-shot tasks. However, existing prompt-based solutions suffer from unidirectional information flow and increased parameters since they explicitly condition the vision prompts on textual prompts across different transformer layers using non-shareable coupling functions. To address this issue, we propose a Dual-shared mechanism based on LoRA (DsRA) that addresses VLM adaptation in low-data regimes. The proposed DsRA enjoys several merits. First, we design an inter-modal shared coefficient that focuses on capturing visual and textual shared patterns, ensuring effective mutual synergy between image and text features. Second, an intra-modal shared matrix is proposed to achieve efficient parameter fine-tuning by combining the different coefficients to generate layer-wise adapters placed in encoder layers. Our extensive experiments demonstrate that DsRA improves the generalizability under few-shot classification, base-to-new generalization, and domain generalization settings. Our code will be released soon.
Xiaotian Yin, Xin Liu 0089, Yuan Wang 0064, Yuwen Pan, Tianzhu Zhang 0001
AAAI4
2025 Dual-Agent Optimization framework for Cross-Domain Few-Shot Segmentation
abstract
Cross-Domain Few-Shot Segmentation (CD-FSS) extends the generalization ability of Few-Shot Segmentation (FSS) beyond a single domain, enabling more practical applications. However, directly employing conventional FSS methods suffers from severe performance degradation in cross-domain settings, primarily due to feature sensitivity and support-to-query matching process sensitivity across domains. Existing methods for CD-FSS either focus on domain adaptation of features or delve into designing matching strategies for enhanced cross-domain robustness. Nonetheless, they overlook the fact that these two issues are interdependent and should be addressed jointly. In this work, we tackle these two issues within a unified framework by optimizing features in the frequency domain and enhancing the matching process in the spatial domain, working jointly to handle the deviations introduced by the domain gap. To this end, we propose a coherent Dual-Agent Optimization (DATO) framework, including a consistent mutual aggregation (CMA) and a correlation rectification strategy (CRS). In the consistent mutual aggregation module, we employ a set of agents to learn domain-invariant features across domains, and then use these features to enhance the original representations for feature adaptation. In the correlation rectification strategy, the agent-aggregated domain-invariant features serve as a bridge, transforming the support-to-query matching process into a referable feature space and reducing its domain sensitivity. Extensive experiments demonstrate the efficacy of our approach.
Zhaoyang Li 0010, Yuan Wang 0064, Wangkai Li, Tianzhu Zhang 0001, Xiang Liu 0020
CVPR2
2025 Generalized Few-Shot Point Cloud Segmentation via LLM-Assisted Hyper-Relation Matching
Zhaoyang Li 0010, Yuan Wang 0064, Guoxin Xiong, Wangkai Li, Yuwen Pan, Tianzhu Zhang 0001
ICCV2
2025 Two Losses, One Goal: Balancing Conflict Gradients for Semi-Supervised Semantic Segmentation
Rui Sun 0006, Huayu Mai, Wangkai Li, Yuan Wang 0064
ICCV5
2025 Beyond Confidence: Exploiting Homogeneous Pattern for Semi-Supervised Semantic Segmentation
abstract
The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model’s generalization performance for robust segmentation. Existing methods mainly rely on confidence-based scoring functions in the prediction space to filter pseudo labels, which suffer from the inherent trade-off between true and false positive rates. In this paper, we carefully design an agent construction strategy to build clean sets of correct (positive) and incorrect (negative) pseudo labels, and propose the Agent Score function (AgScore) to measure the consensus between candidate pixels and these sets. In this way, AgScore takes a step further to capture homogeneous patterns in the embedding space, conditioned on clean positive/negative agents stemming from the prediction space, without sacrificing the merits of confidence score, yielding better trad-off. We provide theoretical analysis to understand the mechanism of AgScore, and demonstrate its effectiveness by integrating it into three semi-supervised segmentation frameworks on Pascal VOC, Cityscapes, and COCO datasets, showing consistent improvements across all data partitions.
Rui Sun 0006, Huayu Mai, Wangkai Li, Naisong Luo, Yuan Wang 0064, Tianzhu Zhang 0001
ICML6
2025 Focus on the Object: Gradient-based Feature Modulation for Camouflaged Object Segmentation
abstract
Camouflaged Object Segmentation (COS) seeks to accurately identify and segment objects that are intricately blended with their surroundings, making them challenging to distinguish at the pixel level. Existing COS methods often struggle to capture the subtle distinctions between targets and backgrounds, despite their improved adaptability to camouflaged objects. To address this challenge, we propose a novel Adaptive Camouflage Discrimination Network (ACDNet), to focus more attention on object-relevant features while suppressing attached camouflage features. The proposed ACDNet enjoys several merits. First, we design a gradient-based feature modulator that injects gradient information into channel-wise attention layers, thereby enhancing the discriminability between camouflaged objects and background features. Second, a hierarchical prompting strategy is introduced to endow the prototype-based classifier with target awareness and multi-level perception, mitigating the impact of camouflage diversity. Extensive experimental results on four benchmarks demonstrate that our ACDNet performs favorably against state-of-the-art COS methods.
Naisong Luo, Yuan Wang 0064, Yuwen Pan, Rui Sun 0006
ACM Multimedia2
2025 Exploring the Better Correlation for Few-Shot Video Object Segmentation
abstract
Few-shot video object segmentation (FSVOS) aims to achieve accurate segmentation of novel objects in given video sequences, where the target objects are specified by limited annotated images as support. Most previous top-performing methods adopt the support-query semantic correlation learning paradigm or the intra-query temporal correlation learning paradigm. Nevertheless, they either fail to model temporal consistency across frames, resulting in inconsecutive segmentation, or lose diverse support object information, leading to incomplete segmentation. Therefore, we argue that it is more desirable to achieve both correlations in a collaborative manner. In this work, we delve into the issues present in the combination of few-shot image segmentation methods and video object segmentation methods and propose a dedicated Collaborative Correlation Network (CoCoNet) to address these problems, including a pixel correlation calibration module and a temporal correlation mining module. The proposed CoCoNet enjoys several merits. First, the pixel correlation calibration module aims to mitigate the noise issue in support-query correlation by integrating the affinity learning strategy and the prototype learning strategy. Specifically, we employ Optimal Transport to enrich pixel correlation with contextual information, thereby reducing intra-class differences between support and query. Second, the temporal correlation mining module is responsible for alleviating the issue of uncertainty in the initial frame and establishing reliable guidance for subsequent frames of the query video. With the collaboration of these two modules, our CoCoNet can effectively establish support-query and temporal correlation simultaneously and achieve accurate FSVOS. Extensive experimental results on two challenging benchmarks demonstrate that our method performs favorably against state-of-the-art FSVOS methods.
Naisong Luo, Yuan Wang 0064, Rui Sun 0006, Guoxin Xiong, Tianzhu Zhang 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Purify Then Guide: A Bi-Directional Bridge Network for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation (OVSS) aims to segment an image into regions of corresponding semantic vocabularies, without being limited to a predefined set of object categories. Existing works mainly utilize large-scale vision-language models (e.g., CLIP) to leverage their superior open-vocabulary classification abilities in a two-stage manner. However, their heavy reliance on the first-stage segmentation network leaves the full potential of CLIP untapped, creating an unresolved gap between the rich pre-training knowledge and the challenging per-pixel classification task. Although the recent one-stage paradigm has further leveraged pre-trained vision knowledge from CLIP, it fails to effectively utilize text information due to the inclusion of numerous unrelated semantics in the vocabulary list. How to avoid noise interference in text information and utilize language guidance remains a Gordian knot. In this paper, we propose a bi-directional bridge network (BBN) to bridge the gap between upstream pre-trained models and downstream segmentation tasks. It first purifies the noisy text embedding and then guides semantics-vision aggregation with the purified information in a purification-then-guidance manner, thereby facilitating effective semantic utilization. Specifically, we design an optimal purification modulator to purify noisy text information via the optimal transport algorithm, and a reliable guidance modulator to integrate proper textual information into vision embedding via the designed reliable attention in an adaptive manner. Extensive experimental results on five challenging benchmarks demonstrate that our BBN performs favorably against state-of-the-art open-vocabulary semantic segmentation methods.
Yuwen Pan, Rui Sun 0006, Yuan Wang 0064, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Pay Attention to Target: Relation-Aware Temporal Consistency for Domain Adaptive Video Semantic Segmentation
abstract
Video semantic segmentation has achieved conspicuous achievements attributed to the development of deep learning, but suffers from labor-intensive annotated training data gathering. To alleviate the data-hunger issue, domain adaptation approaches are developed in the hope of adapting the model trained on the labeled synthetic videos to the real videos in the absence of annotations. By analyzing the dominant paradigm consistency regularization in the domain adaptation task, we find that the bottlenecks exist in previous methods from the perspective of pseudo-labels. To take full advantage of the information contained in the pseudo-labels and empower more effective supervision signals, we propose a coherent PAT network including a target domain focalizer and relation-aware temporal consistency. The proposed PAT network enjoys several merits. First, the target domain focalizer is responsible for paying attention to the target domain, and increasing the accessibility of pseudo-labels in consistency training. Second, the relation-aware temporal consistency aims at modeling the inter-class consistent relationship across frames to equip the model with effective supervision signals. Extensive experimental results on two challenging benchmarks demonstrate that our method performs favorably against state-of-the-art domain adaptive video semantic segmentation methods.
Huayu Mai, Rui Sun 0006, Yuan Wang 0064, Tianzhu Zhang 0001, Feng Wu 0001
AAAI3
2024 Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation (OVS) aims to segment images of arbitrary categories specified by class labels or captions. However, most previous best-performing methods, whether pixel grouping methods or region recognition methods, suffer from false matches between image features and category labels. We attribute this to the natural gap between the textual features and visual features. In this work, we rethink how to mitigate false matches from the perspective of image-to-image matching and propose a novel relation-aware intra-modal matching (RIM) framework for OVS based on visual foundation models. RIM achieves robust region classification by firstly constructing diverse image-modal reference features and then matching them with region features based on relation-aware ranking distribution. The proposed RIM enjoys several merits. First, the intra-modal reference features are better aligned, circumventing potential ambiguities that may arise in cross-modal matching. Second, the ranking-based matching process harnesses the structure information implicit in the inter-class relationships, making it more robust than comparing individually. Extensive experiments on three benchmarks demonstrate that RIM outperforms previous state-of-the-art methods by large margins, obtaining a lead of more than 10% in mIoU on PASCAL VOC benchmark.
Yuan Wang 0064, Rui Sun 0006, Naisong Luo, Yuwen Pan, Tianzhu Zhang 0001
CVPR1
2024 Localization and Expansion: A Decoupled Framework for Point Cloud Few-Shot Semantic Segmentation
Zhaoyang Li 0010, Yuan Wang 0064, Wangkai Li, Rui Sun 0006, Tianzhu Zhang 0001
ECCV (73)2
2024 Aggregation and Purification: Dual Enhancement Network for Point Cloud Few-shot Segmentation
Guoxin Xiong, Yuan Wang 0064, Zhaoyang Li 0010, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001
IJCAI2
2024 Rethinking the Implicit Optimization Paradigm with Dual Alignments for Referring Remote Sensing Image Segmentation
abstract
Referring Remote Sensing Image Segmentation (RRSIS) is a challenging task that aims to identify specific regions in aerial images that are relevant to given textual conditions. Existing methods tend to adopt the paradigm of implicit optimization, utilizing a framework consisting of early cross-modal feature fusion and a fixed convolutional kernel-based predictor, neglecting the inherent inter-domain gap and conducting class-agnostic predictions. In this paper, we rethink the issues with the implicit optimization paradigm and address the RRSIS task from a dual-alignment perspective. Specifically, we prepend the dedicated Dual Alignment Network (DANet), including an explicit alignment strategy and a reliable agent alignment module. The explicit alignment strategy effectively reduces domain discrepancies by narrowing the inter-domain affinity distribution. Meanwhile, the reliable agent alignment module aims to enhance the predictor's multi-modality awareness and alleviate the impact of deceptive noise interference. Extensive experiments on two remote sensing datasets demonstrate the effectiveness of our proposed DANet in achieving superior segmentation performance without introducing additional learnable parameters compared to state-of-the-art methods.
Yuwen Pan, Rui Sun 0006, Yuan Wang 0064, Tianzhu Zhang 0001, Yongdong Zhang 0001
ACM Multimedia3
2023 Rethinking the Correlation in Few-Shot Segmentation: A Buoys View
abstract
Few-shot segmentation (FSS) aims to segment novel ob-jects in a given query image with only a few annotated support images. However, most previous best-performing methods, whether prototypical learning methods or affinity learning methods, neglect to alleviate false matches caused by their own pixel-level correlation. In this work, we rethink how to mitigate the false matches from the perspective of representative reference features (referred to as buoys), and propose a novel adaptive buoys correlation (ABC) network to rectify direct pairwise pixel-level correlation, including a buoys mining module and an adaptive correlation module. The proposed ABC enjoys several merits. First, to learn the buoys well without any correspondence supervision, we customize the buoys mining module according to the three characteristics of representativeness, task awareness and re-silience. Second, the proposed adaptive correlation module is responsible for further endowing buoy-correlation-based pixel matching with an adaptive ability. Extensive experimen-tal results with two different backbones on two challenging benchmarks demonstrate that our ABC, as a general plu-gin, achieves consistent improvements over several leading methods on both I-shot and 5-shot settings.
Yuan Wang 0064, Rui Sun 0006, Tianzhu Zhang 0001
CVPR1
2023 Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object Segmentation
abstract
Memory-based methods in semi-supervised video object segmentation task achieve competitive performance by performing dense matching between query and memory frames. However, most of the existing methods neglect the fact that videos carry rich temporal information yet redundant spatial information. In this case, direct pixel-level global matching will lead to ambiguous correspondences. In this work, we reconcile the inherent tension of spatial and temporal information to retrieve memory frame information along the object trajectory, and propose a novel and coherent Trajectory Memory Retrieval Network (TMRN) to equip with the trajectory information, including a spatial alignment module and a temporal aggregation module. The proposed TMRN enjoys several merits. First, TMRN is empowered to characterize the temporal correspondence which is in line with the nature of video in a data-driven manner. Second, we elegantly customize the spatial alignment module by coupling SVD initialization with agent-level correlation for representative agent construction and rectifying false matches caused by direct pairwise pixel-level correlation, respectively. Extensive experimental results on challenging benchmarks including DAVIS 2017 validation / test and Youtube-VOS 2018/2019 demonstrate that our TMRN, as a general plugin module, achieves consistent improvements over several leading methods.
Rui Sun 0006, Yuan Wang 0064, Huayu Mai, Tianzhu Zhang 0001, Feng Wu 0001
ICCV2
2023 Focus on Query: Adversarial Mining Transformer for Few-Shot Segmentation
abstract
Few-shot segmentation (FSS) aims to segment objects of new categories given only a handful of annotated samples. Previous works focus their efforts on exploring the support information while paying less attention to the mining of the critical query branch. In this paper, we rethink the importance of support information and propose a new query-centric FSS model Adversarial Mining Transformer (AMFormer), which achieves accurate query image segmentation with only rough support guidance or even weak support labels. The proposed AMFormer enjoys several merits. First, we design an object mining transformer (G) that can achieve the expansion of incomplete region activated by support clue, and a detail mining transformer (D) to discriminate the detailed local difference between the expanded mask and the ground truth. Second, we propose to train G and D via an adversarial process, where G is optimized to generate more accurate masks approaching ground truth to fool D. We conduct extensive experiments on commonly used Pascal-5i and COCO-20i benchmarks and achieve state-of-the-art results across all settings. In addition, the decent performance with weak support labels in our query-centric paradigm may inspire the development of more general FSS models.
Yuan Wang 0064, Naisong Luo, Tianzhu Zhang 0001
NeurIPS1
2022 Adaptive Agent Transformer for Few-Shot Segmentation
Yuan Wang 0064, Rui Sun 0006, Tianzhu Zhang 0001
ECCV (29)1