VLDB 2026 Research / reviewers in the wild / expert
Rui Sun 0006
dblp:01/3595-6
· DBLP profile ↗
34ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0002-8009-4240ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DA2-LiDAR: A Generic Density-Adaptive Framework for Unsupervised Domain Adaptation in LiDAR SegmentationabstractThis paper addresses the critical challenge of domain adaptation for LiDAR-based semantic segmentation, particularly the significant density disparities that emerge when transferring models from synthetic to real-world environments. We present DA2-LiDAR, a novel density-adaptive domain adaptation framework that bridges domain gaps through the construction of intermediate domains with density-varying point distributions. Our approach employs a simple yet effective masking strategy that systematically reduces density discrepancies between domains while extracting more effective supervisory signals, as well as preserving critical semantic information. The framework consists of three key components: (1) a Density Adaptation Module that establishes a continuous spectrum of intermediate domains through dataset-agnostic masking operations; (2) a Contextual Consistency Module that enforces relational coherence across differently masked variants of the same scan at varying degrees, providing additional supervision signals, enhancing the model's ability to extract features; and (3) a Semantic Preservation Module that mitigates information loss in heavily masked scans by reconstructing domain-specific data distributions. Extensive experiments on synthetic-to-real and other benchmarks demonstrate that DA2-LiDAR consistently outperforms state-of-the-art methods, achieving significant improvements in cross-domain generalization without requiring dataset-specific prior knowledge or introducing computational overhead. Rui Sun 0006, Wangkai Li, Naisong Luo, Yuan Wang 0064, Tianzhu Zhang 0001, Feng Wu 0005 |
IEEE Trans. Image Process. | 2 |
| 2026 | DA-Cal: Toward Cross-Domain Calibration in Semantic SegmentationabstractWhile existing unsupervised domain adaptation (UDA) methods greatly enhance target domain performance in semantic segmentation, they often neglect network calibration quality, resulting in misalignment between prediction confidence and actual accuracy-a significant risk in safety-critical applications. Our key insight emerges from observing that performance degrades substantially when soft pseudo-labels replace hard pseudo-labels in cross-domain scenarios due to poor calibration, despite the theoretical equivalence of perfectly calibrated soft pseudo-labels to hard pseudo-labels. Based on this finding, we propose DA-Cal, a dedicated cross-domain calibration framework that transforms target domain calibration into soft pseudo-label optimization. DA-Cal introduces a Meta Temperature Network to generate pixel-level calibration parameters and employs bi-level optimization to establish the relationship between soft pseudo-labels and UDA supervision, while utilizing complementary domain-mixing strategies to prevent overfitting and reduce domain discrepancies. Experiments demonstrate that DA-Cal seamlessly integrates with existing self-training frameworks across multiple UDA segmentation benchmarks, significantly improving target domain calibration while delivering performance gains without inference overhead. The code will be released. Wangkai Li, Rui Sun 0006, Zhaoyang Li 0010, Tianzhu Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Alleviate and Mining: Rethinking Unsupervised Domain Adaptation for Mitochondria Segmentation from Pseudo-Label PerspectiveabstractMitochondria segmentation from electron microscopy (EM) images plays a crucial role in biological and medical research. However, models trained on source domains often suffer from performance degradation when applied to target domains due to domain shift. Unsupervised domain adaptation (UDA) methods have been proposed to address this issue, but they often overlook the reliability of pseudo-labels and the effectiveness of supervision signals. In this paper, we propose R4MITO, a novel UDA framework for robust mitochondria segmentation. First, we introduce Reliable Prototype Pseudo-labels to mitigate the inconsistency of class-level features between across domains by leveraging source prototypes to model target prototypes. Second, we devise Correlation-wise Consistency Regularization to exploit inter-pixel correlations, aligning agent-level correlations under various perturbations. Third, we propose Rank-aware Relationship Consistency Regularization to fully utilize the rich information encoded in inter-agent relationships by imposing rank-aware constraints on agent-ranking probability distributions. Extensive experiments on multiple EM datasets demonstrate the superiority of our R4MITO over existing state-of-the-art UDA methods for mitochondria segmentation. Rui Sun 0006, Wangkai Li, Huayu Mai, Naisong Luo, Yuwen Pan, Tianzhu Zhang 0001 |
AAAI | 2 |
| 2025 | Relaxed Class-consensus Consistency for Semi-supervised Semantic SegmentationabstractThe key to semi-supervised semantic segmentation lies in how to fully exploit a large amount of unlabeled data to improve the model’s generalization performance. Most methods are lured into the trap of taking each class independently (i.e., class-independent consistency) and neglecting the fact that there exist semantic dependencies among classes. In this paper, we analyze the bottlenecks of class-independent consistency inherent in previous methods and offer a fresh perspective of cooperative game theory to explicitly encourage class-consensus alignment (i.e., class-consensus consistency between the teacher (weak augmented view) and student network (strong augmented view). We formulate classes as players in an cooperative game to model their interpretable consensus and shed light on the possibility of closer collaboration between consensus themselves and consistency regularization, yielding more comprehensive and effective supervision signals. To this end, we carefully design the class-consensus consistency without introducing any external knowledge to model class structure information which renders better interpretability, and further, prepend relaxed class-consensus consistency (RCC) to unlock the potential of modeling class consensus by relaxing the strict alignment of direct class consensus values to ranking alignment. Extensive experimental results on multiple benchmarks demonstrate that RCC performs favorably against state-of-the-art methods. Particularly in the low-data regimes, RCC achieves significant improvements. Huayu Mai, Rui Sun 0006, Feng Wu 0001 |
AAAI | 2 |
| 2025 | Exploring Weather-aware Aggregation and Adaptation for Semantic Segmentation under Adverse Conditions
Yuwen Pan, Rui Sun 0006, Wangkai Li, Tianzhu Zhang 0001 |
ICCV | 2 |
| 2025 | Two Losses, One Goal: Balancing Conflict Gradients for Semi-Supervised Semantic Segmentation
Rui Sun 0006, Huayu Mai, Wangkai Li, Yuan Wang 0064 |
ICCV | 1 |
| 2025 | Towards Unbiased Learning in Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation aims to learn from a limited amount of labeled data and a large volume of unlabeled data, which has witnessed impressive progress with the recent advancement of deep neural networks. However, existing methods tend to neglect the fact of class imbalance issues, leading to the Matthew effect, that is, the poorly calibrated model’s predictions can be biased towards the majority classes and away from minority classes with fewer samples. In this work, we analyze the Matthew effect present in previous methods that hinder model learning from a discriminative perspective. In light of this background, we integrate generative models into semi-supervised learning, taking advantage of their better class-imbalance tolerance. To this end, we propose DiffMatch to formulate the semi-supervised semantic segmentation task as a conditional discrete data generation problem to alleviate the Matthew effect of discriminative solutions from a generative perspective. Plus, to further reduce the risk of overfitting to the head classes and to increase coverage of the tail class distribution, we mathematically derive a debiased adjustment to adjust the conditional reverse probability towards unbiased predictions during each sampling step. Extensive experimental results across multiple benchmarks, especially in the most limited label scenarios with the most serious class imbalance issues, demonstrate that DiffMatch performs favorably against state-of-the-art methods. Rui Sun 0006, Huayu Mai, Wangkai Li, Tianzhu Zhang 0001 |
ICLR | 1 |
| 2025 | Beyond Confidence: Exploiting Homogeneous Pattern for Semi-Supervised Semantic SegmentationabstractThe critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model’s generalization performance for robust segmentation. Existing methods mainly rely on confidence-based scoring functions in the prediction space to filter pseudo labels, which suffer from the inherent trade-off between true and false positive rates. In this paper, we carefully design an agent construction strategy to build clean sets of correct (positive) and incorrect (negative) pseudo labels, and propose the Agent Score function (AgScore) to measure the consensus between candidate pixels and these sets. In this way, AgScore takes a step further to capture homogeneous patterns in the embedding space, conditioned on clean positive/negative agents stemming from the prediction space, without sacrificing the merits of confidence score, yielding better trad-off. We provide theoretical analysis to understand the mechanism of AgScore, and demonstrate its effectiveness by integrating it into three semi-supervised segmentation frameworks on Pascal VOC, Cityscapes, and COCO datasets, showing consistent improvements across all data partitions. Rui Sun 0006, Huayu Mai, Wangkai Li, Naisong Luo, Yuan Wang 0064, Tianzhu Zhang 0001 |
ICML | 1 |
| 2025 | Balanced Learning for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation (UDA) for semantic segmentation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Despite the effectiveness of self-training techniques in UDA, they struggle to learn each class in a balanced manner due to inherent class imbalance and distribution shift in both data and label space between domains. To address this issue, we propose Balanced Learning for Domain Adaptation (BLDA), a novel approach to directly assess and alleviate class bias without requiring prior knowledge about the distribution shift. First, we identify over-predicted and under-predicted classes by analyzing the distribution of predicted logits. Subsequently, we introduce a post-hoc approach to align the logits distributions across different classes using shared anchor distributions. To further consider the network’s need to generate unbiased pseudo-labels during self-training, we estimate logits distributions online and incorporate logits correction terms into the loss function. Moreover, we leverage the resulting cumulative density as domain-shared structural knowledge to connect the source and target domains. Extensive experiments on two standard UDA semantic segmentation benchmarks demonstrate that BLDA consistently improves performance, especially for under-predicted classes, when integrated into various existing methods. Wangkai Li, Rui Sun 0006, Bohao Liao, Zhaoyang Li 0010, Tianzhu Zhang 0001 |
ICML | 2 |
| 2025 | Focus on the Object: Gradient-based Feature Modulation for Camouflaged Object SegmentationabstractCamouflaged Object Segmentation (COS) seeks to accurately identify and segment objects that are intricately blended with their surroundings, making them challenging to distinguish at the pixel level. Existing COS methods often struggle to capture the subtle distinctions between targets and backgrounds, despite their improved adaptability to camouflaged objects. To address this challenge, we propose a novel Adaptive Camouflage Discrimination Network (ACDNet), to focus more attention on object-relevant features while suppressing attached camouflage features. The proposed ACDNet enjoys several merits. First, we design a gradient-based feature modulator that injects gradient information into channel-wise attention layers, thereby enhancing the discriminability between camouflaged objects and background features. Second, a hierarchical prompting strategy is introduced to endow the prototype-based classifier with target awareness and multi-level perception, mitigating the impact of camouflage diversity. Extensive experimental results on four benchmarks demonstrate that our ACDNet performs favorably against state-of-the-art COS methods. Naisong Luo, Yuan Wang 0064, Yuwen Pan, Rui Sun 0006 |
ACM Multimedia | 4 |
| 2025 | BeyondMix: Leveraging Structural Priors and Long-Range Dependencies for Domain-Invariant LiDAR SegmentationabstractDomain adaptation for LiDAR semantic segmentation remains challenging due to the complex structural properties of point cloud data. While mix-based paradigms have shown promise, they often fail to fully leverage the rich structural priors inherent in 3D LiDAR point clouds. In this paper, we identify three critical yet underexploited structural priors: permutation invariance, local consistency, and geometric consistency. We introduce BeyondMix, a novel framework that harnesses the capabilities of State Space Models (specifically Mamba) to construct and exploit these structural priors while modeling long-range dependencies that transcend the limited receptive fields of conventional voxel-based approaches. By employing space-filling curves to impose sequential ordering on point cloud data and implementing strategic spatial partitioning schemes, BeyondMix effectively captures domain-invariant representations. Extensive experiments on challenging LiDAR semantic segmentation benchmarks demonstrate that our approach consistently outperforms existing state-of-the-art methods, establishing a new paradigm for unsupervised domain adaptation in 3D point cloud understanding. Rui Sun 0006, Wangkai Li, Huayu Mai, Zhixin Cheng, Tianzhu Zhang 0001 |
NeurIPS | 2 |
| 2025 | Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding PerspectiveabstractPseudo-label learning is widely used in semantic segmentation, particularly in label-scarce scenarios such as unsupervised domain adaptation (UDA) and semi-supervised learning (SSL). Despite its success, this paradigm can generate erroneous pseudo-labels, which are further amplified during training due to utilization of one-hot encoding. To address this issue, we propose ECOCSeg, a novel perspective for segmentation models that utilizes error-correcting output codes (ECOC) to create a fine-grained encoding for each class. ECOCSeg offers several advantages. First, an ECOC-based classifier is introduced, enabling model to disentangle classes into attributes and handle partial inaccurate bits, improving stability and generalization in pseudo-label learning. Second, a bit-level label denoising mechanism is developed to generate higher-quality pseudo-labels, providing adequate and robust supervision for unlabeled images. ECOCSeg can be easily integrated with existing methods and consistently demonstrates significant improvements on multiple UDA and SSL benchmarks across different segmentation architectures. Code is available at https://github.com/Woof6/ECOCSeg. Wangkai Li, Rui Sun 0006, Zhaoyang Li 0010, Tianzhu Zhang 0001 |
NeurIPS | 2 |
| 2025 | Towards Unsupervised Domain Bridging via Image Degradation in Semantic SegmentationabstractSemantic segmentation suffers from significant performance degradation when the trained network is applied to a different domain. To address this issue, unsupervised domain adaptation (UDA) has been extensively studied.
Despite the effectiveness of selftraining techniques in UDA, they still overlook the explicit modeling
of domain-shared feature extraction.
In this paper, we propose DiDA, an unsupervised domain bridging approach for semantic segmentation. DiDA consists of two key modules: (1) Degradation-based Intermediate Domain Construction, which creates continuous intermediate domains through simple image degradation operations to encourage learning domain-invariant features as domain differences gradually diminish; (2) Semantic Shift Compensation, which leverages a diffusion encoder to disentangle and compensate for semantic shift information with degraded time-steps, preserving discriminative representations in the intermediate domains.
As a plug-and-play solution, DiDA supports various degradation operations and seamlessly integrates with existing UDA methods. Extensive experiments on multiple domain adaptive semantic segmentation benchmarks demonstrate that DiDA consistently achieves significant performance improvements across all settings.
Code is available at https://github.com/Woof6/DiDA. Wangkai Li, Rui Sun 0006, Huayu Mai, Tianzhu Zhang 0001 |
NeurIPS | 2 |
| 2025 | Agent-Based Control Prompt Tuning for Video-Text RetrievalabstractLarge-scale image-text pre-trained models have shown promising transferability to various downstream tasks. Video-text retrieval benefits from it by transferring pre-trained CLIP to video-text domain. Although these pre-trained models have shown impressive performance, full fine-tuning becomes prohibitively expensive as the size of these pre-trained models grows rapidly. To solve this, parameter-efficient tuning methods have been proposed, and prompt tuning is one of the most promising directions. However, existing prompt tuning methods do not have sufficient performance due to the lack of cross-modal interaction and prompt reliability assurance. To address these issues, we propose an effective and efficient Agent-based Control Prompt Tuning method (AbC-PT) for parameter-efficient video-text retrieval. The proposed AbC-PT enjoys several merits. Firstly, we design a parameter-efficient agent decoder with a carefully designed consistent attention mechanism to effectively capture video temporal information, mine contextual texts and perform cross-modal interaction between them. Secondly, we introduce two different sets of prompts, i.e., the vanilla prompt prepended to the input tokens and the concept prompt as the agent of the agent decoder. In addition, to ensure cross-modal semantic consistency of the concept prompt, we design a semantic consistency constraint loss. Thirdly, we devise a parameter-free prompt controller for adaptively calibrating each vanilla prompt based on its semantic in a data-driven way. Extensive experiments on five challenging benchmarks demonstrate that our method not only outperforms state-of-the-art parameter-efficient tuning methods, but even surpasses the full fine-tuning with 0.46% parameter overhead. Huakai Lai, Rui Sun 0006, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Exploring the Better Correlation for Few-Shot Video Object SegmentationabstractFew-shot video object segmentation (FSVOS) aims to achieve accurate segmentation of novel objects in given video sequences, where the target objects are specified by limited annotated images as support. Most previous top-performing methods adopt the support-query semantic correlation learning paradigm or the intra-query temporal correlation learning paradigm. Nevertheless, they either fail to model temporal consistency across frames, resulting in inconsecutive segmentation, or lose diverse support object information, leading to incomplete segmentation. Therefore, we argue that it is more desirable to achieve both correlations in a collaborative manner. In this work, we delve into the issues present in the combination of few-shot image segmentation methods and video object segmentation methods and propose a dedicated Collaborative Correlation Network (CoCoNet) to address these problems, including a pixel correlation calibration module and a temporal correlation mining module. The proposed CoCoNet enjoys several merits. First, the pixel correlation calibration module aims to mitigate the noise issue in support-query correlation by integrating the affinity learning strategy and the prototype learning strategy. Specifically, we employ Optimal Transport to enrich pixel correlation with contextual information, thereby reducing intra-class differences between support and query. Second, the temporal correlation mining module is responsible for alleviating the issue of uncertainty in the initial frame and establishing reliable guidance for subsequent frames of the query video. With the collaboration of these two modules, our CoCoNet can effectively establish support-query and temporal correlation simultaneously and achieve accurate FSVOS. Extensive experimental results on two challenging benchmarks demonstrate that our method performs favorably against state-of-the-art FSVOS methods. Naisong Luo, Yuan Wang 0064, Rui Sun 0006, Guoxin Xiong, Tianzhu Zhang 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Purify Then Guide: A Bi-Directional Bridge Network for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation (OVSS) aims to segment an image into regions of corresponding semantic vocabularies, without being limited to a predefined set of object categories. Existing works mainly utilize large-scale vision-language models (e.g., CLIP) to leverage their superior open-vocabulary classification abilities in a two-stage manner. However, their heavy reliance on the first-stage segmentation network leaves the full potential of CLIP untapped, creating an unresolved gap between the rich pre-training knowledge and the challenging per-pixel classification task. Although the recent one-stage paradigm has further leveraged pre-trained vision knowledge from CLIP, it fails to effectively utilize text information due to the inclusion of numerous unrelated semantics in the vocabulary list. How to avoid noise interference in text information and utilize language guidance remains a Gordian knot. In this paper, we propose a bi-directional bridge network (BBN) to bridge the gap between upstream pre-trained models and downstream segmentation tasks. It first purifies the noisy text embedding and then guides semantics-vision aggregation with the purified information in a purification-then-guidance manner, thereby facilitating effective semantic utilization. Specifically, we design an optimal purification modulator to purify noisy text information via the optimal transport algorithm, and a reliable guidance modulator to integrate proper textual information into vision embedding via the designed reliable attention in an adaptive manner. Extensive experimental results on five challenging benchmarks demonstrate that our BBN performs favorably against state-of-the-art open-vocabulary semantic segmentation methods. Yuwen Pan, Rui Sun 0006, Yuan Wang 0064, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Electron Microscopy Images as Set of Fragments for Mitochondrial SegmentationabstractAutomatic mitochondrial segmentation enjoys great popularity with the development of deep learning. However, the coarse prediction raised by the presence of regular 3D grids in previous methods regardless of 3D CNN or the vision transformers suggest a possibly sub-optimal feature arrangement. To mitigate this limitation, we attempt to interpret the 3D EM image stacks as a set of interrelated 3D fragments for a better solution. However, it is non-trivial to model the 3D fragments without introducing excessive computational overhead. In this paper, we design a coherent fragment vision transformer (FragViT) combined with affinity learning to manipulate features on 3D fragments yet explore mutual relationships to model fragment-wise context, enjoying locality prior without sacrificing global reception. The proposed FragViT includes a fragment encoder and a hierarchical fragment aggregation module. The fragment encoder is equipped with affinity heads to transform the tokens into fragments with homogeneous semantics, and the multi-layer self-attention is used to explicitly learn inter-fragment relations with long-range dependencies. The hierarchical fragment aggregation module is responsible for hierarchically aggregating fragment-wise prediction back to the final voxel-wise prediction in a progressive manner. Extensive experimental results on the challenging MitoEM, Lucchi, and AC3/AC4 benchmarks demonstrate the effectiveness of the proposed method. Naisong Luo, Rui Sun 0006, Yuwen Pan, Tianzhu Zhang 0001, Feng Wu 0001 |
AAAI | 2 |
| 2024 | Pay Attention to Target: Relation-Aware Temporal Consistency for Domain Adaptive Video Semantic SegmentationabstractVideo semantic segmentation has achieved conspicuous achievements attributed to the development of deep learning, but suffers from labor-intensive annotated training data gathering. To alleviate the data-hunger issue, domain adaptation approaches are developed in the hope of adapting the model trained on the labeled synthetic videos to the real videos in the absence of annotations. By analyzing the dominant paradigm consistency regularization in the domain adaptation task, we find that the bottlenecks exist in previous methods from the perspective of pseudo-labels. To take full advantage of the information contained in the pseudo-labels and empower more effective supervision signals, we propose a coherent PAT network including a target domain focalizer and relation-aware temporal consistency. The proposed PAT network enjoys several merits. First, the target domain focalizer is responsible for paying attention to the target domain, and increasing the accessibility of pseudo-labels in consistency training. Second, the relation-aware temporal consistency aims at modeling the inter-class consistent relationship across frames to equip the model with effective supervision signals. Extensive experimental results on two challenging benchmarks demonstrate that our method performs favorably against state-of-the-art domain adaptive video semantic segmentation methods. Huayu Mai, Rui Sun 0006, Yuan Wang 0064, Tianzhu Zhang 0001, Feng Wu 0001 |
AAAI | 2 |
| 2024 | RankMatch: Exploring the Better Consistency Regularization for Semi-Supervised Semantic SegmentationabstractThe key lie in semi-supervised semantic segmentation is how to fully exploit substantial unlabeled data to im-prove the model's generalization performance by resorting to constructing effective supervision signals. Most methods tend to directly apply contrastive learning to seek additional supervision to complement independent regular pixel-wise consistency regularization. However, these methods tend not to be preferred ascribed to their complicated designs, heavy memory footprints and susceptibility to confirmation bias. In this paper, we analyze the bottlenecks exist in con-trastive learning-based methods and offer a fresh perspective on inter-pixel correlations to construct more safe and effective supervision signals, which is in line with the nature of semantic segmentation. To this end, we develop a coherent RankMatch network, including the construction of representative agents to model inter-pixel correlation beyond regular individual pixel-wise consistency, and fur-ther unlock the potential of agents by modeling inter-agent relationships in pursuit of rank-aware correlation consis-tency. Extensive experimental results on multiple bench-marks, including mitochondria segmentation, demonstrate that RankMatch performs favorably against state-of-the-art methods. Particularly in the low-data regimes, RankMatch achieves significant improvements. Huayu Mai, Rui Sun 0006, Tianzhu Zhang 0001, Feng Wu 0001 |
CVPR | 2 |
| 2024 | Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation (OVS) aims to segment images of arbitrary categories specified by class labels or captions. However, most previous best-performing methods, whether pixel grouping methods or region recognition methods, suffer from false matches between image features and category labels. We attribute this to the natural gap between the textual features and visual features. In this work, we rethink how to mitigate false matches from the perspective of image-to-image matching and propose a novel relation-aware intra-modal matching (RIM) framework for OVS based on visual foundation models. RIM achieves robust region classification by firstly constructing diverse image-modal reference features and then matching them with region features based on relation-aware ranking distribution. The proposed RIM enjoys several merits. First, the intra-modal reference features are better aligned, circumventing potential ambiguities that may arise in cross-modal matching. Second, the ranking-based matching process harnesses the structure information implicit in the inter-class relationships, making it more robust than comparing individually. Extensive experiments on three benchmarks demonstrate that RIM outperforms previous state-of-the-art methods by large margins, obtaining a lead of more than 10% in mIoU on PASCAL VOC benchmark. Yuan Wang 0064, Rui Sun 0006, Naisong Luo, Yuwen Pan, Tianzhu Zhang 0001 |
CVPR | 2 |
| 2024 | Localization and Expansion: A Decoupled Framework for Point Cloud Few-Shot Semantic Segmentation
Zhaoyang Li 0010, Yuan Wang 0064, Wangkai Li, Rui Sun 0006, Tianzhu Zhang 0001 |
ECCV (73) | 4 |
| 2024 | Exploring Reliable Matching with Phase Enhancement for Night-Time Semantic Segmentation
Yuwen Pan, Rui Sun 0006, Naisong Luo, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
ECCV (53) | 2 |
| 2024 | Rethinking the Implicit Optimization Paradigm with Dual Alignments for Referring Remote Sensing Image SegmentationabstractReferring Remote Sensing Image Segmentation (RRSIS) is a challenging task that aims to identify specific regions in aerial images that are relevant to given textual conditions. Existing methods tend to adopt the paradigm of implicit optimization, utilizing a framework consisting of early cross-modal feature fusion and a fixed convolutional kernel-based predictor, neglecting the inherent inter-domain gap and conducting class-agnostic predictions. In this paper, we rethink the issues with the implicit optimization paradigm and address the RRSIS task from a dual-alignment perspective. Specifically, we prepend the dedicated Dual Alignment Network (DANet), including an explicit alignment strategy and a reliable agent alignment module. The explicit alignment strategy effectively reduces domain discrepancies by narrowing the inter-domain affinity distribution. Meanwhile, the reliable agent alignment module aims to enhance the predictor's multi-modality awareness and alleviate the impact of deceptive noise interference. Extensive experiments on two remote sensing datasets demonstrate the effectiveness of our proposed DANet in achieving superior segmentation performance without introducing additional learnable parameters compared to state-of-the-art methods. Yuwen Pan, Rui Sun 0006, Yuan Wang 0064, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | Camouflaged Instance Segmentation via Explicit De-CamouflagingabstractCamouflaged Instance Segmentation (CIS) aims at predicting the instance-level masks of camouflaged objects, which are usually the animals in the wild adapting their appearance to match the surroundings. Previous instance segmentation methods perform poorly on this task as they are easily disturbed by the deceptive camouflage. To address these challenges, we propose a novel De-camouflaging Network (DCNet) including a pixel-level camouflage decoupling module and an instance-level camouflage suppression module. The proposed DCNet enjoys several merits. First, the pixel-level camouflage decoupling module can extract camouflage characteristics based on the Fourier transformation. Then a difference attention mechanism is proposed to eliminate the camouflage characteristics while reserving target object characteristics in the pixel feature. Second, the instance-level camouflage suppression module can aggregate rich instance information from pixels by use of instance prototypes. To mitigate the effect of background noise during segmentation, we introduce some reliable reference points to build a more robust similarity measurement. With the aid of these two modules, our DCNet can effectively model de-camouflaging and achieve accurate segmentation for camouflaged instances. Extensive experimental results on two benchmarks demonstrate that our DCNet performs favorably against state-of-the-art CIS methods, e.g., with more than 5% performance gains on COD10K and NC4K datasets in average precision. Naisong Luo, Yuwen Pan, Rui Sun 0006, Tianzhu Zhang 0001, Zhiwei Xiong, Feng Wu 0001 |
CVPR | 3 |
| 2023 | DualRel: Semi-Supervised Mitochondria Segmentation from A Prototype PerspectiveabstractAutomatic mitochondria segmentation enjoys great popularity with the development of deep learning. However, existing methods rely heavily on the labor-intensive manual gathering by experienced domain experts. And naively applying semi-supervised segmentation methods in the natural image field to mitigate the labeling cost is undesirable. In this work, we analyze the gap between mitochondrial images and natural images and rethink how to achieve effective semi-supervised mitochondria segmentation, from the perspective of reliable prototype-level supervision. We propose a novel end-to-end dual-reliable (DualRel) network, including a reliable pixel aggregation module and a reliable prototype selection module. The proposed DualRel enjoys several merits. First, to learn the prototypes well without any explicit supervision, we carefully design the referential correlation to rectify the direct pairwise correlation. Second, the reliable prototype selection module is responsible for further evaluating the reliability of prototypes in constructing prototype-level consistency regularization. Extensive experimental results on three challenging benchmarks demonstrate that our method performs favorably against state-of-the-art semi-supervised segmentation methods. Importantly, with extremely few samples used for training, DualRel is also on par with current state-of-the-art fully supervised methods. Huayu Mai, Rui Sun 0006, Tianzhu Zhang 0001, Zhiwei Xiong, Feng Wu 0001 |
CVPR | 2 |
| 2023 | Rethinking the Correlation in Few-Shot Segmentation: A Buoys ViewabstractFew-shot segmentation (FSS) aims to segment novel ob-jects in a given query image with only a few annotated support images. However, most previous best-performing methods, whether prototypical learning methods or affinity learning methods, neglect to alleviate false matches caused by their own pixel-level correlation. In this work, we rethink how to mitigate the false matches from the perspective of representative reference features (referred to as buoys), and propose a novel adaptive buoys correlation (ABC) network to rectify direct pairwise pixel-level correlation, including a buoys mining module and an adaptive correlation module. The proposed ABC enjoys several merits. First, to learn the buoys well without any correspondence supervision, we customize the buoys mining module according to the three characteristics of representativeness, task awareness and re-silience. Second, the proposed adaptive correlation module is responsible for further endowing buoy-correlation-based pixel matching with an adaptive ability. Extensive experimen-tal results with two different backbones on two challenging benchmarks demonstrate that our ABC, as a general plu-gin, achieves consistent improvements over several leading methods on both I-shot and 5-shot settings. Yuan Wang 0064, Rui Sun 0006, Tianzhu Zhang 0001 |
CVPR | 2 |
| 2023 | Adaptive Template Transformer for Mitochondria Segmentation in Electron Microscopy ImagesabstractMitochondria, as tiny structures within the cell, are of significant importance in studying cell functions for biological and clinical analysis. And exploring how to automatically segment mitochondria in electron microscopy (EM) images has attracted increasing attention. However, most of existing methods struggle to adapt to different scales and appearances of the input due to the inherent limitations of the traditional CNN architecture. To mitigate these limitations, we propose a novel adaptive template transformer (ATFormer) for mitochondria segmentation. The proposed ATFormer model enjoys several merits. First, the designed structural template learning module can acquire appearance-adaptive templates of background, foreground and contour to sense the characteristics of different shapes of mitochondria. And we further adopt an optimal transport algorithm to enlarge the discrepancy among diverse templates to activate corresponding regions fully. Second, we introduce a hierarchical attention learning mechanism to absorb multi-level information for templates to be adaptive scale-aware classifiers for dense prediction. Extensive experimental results on three challenging benchmarks including MitoEM, Lucchi and NucMM-Z datasets demonstrate that our ATFormer performs favorably against state-of-the-art mitochondria segmentation methods. Yuwen Pan, Naisong Luo, Rui Sun 0006, Tianzhu Zhang 0001, Zhiwei Xiong, Yongdong Zhang 0001 |
ICCV | 3 |
| 2023 | Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object SegmentationabstractMemory-based methods in semi-supervised video object segmentation task achieve competitive performance by performing dense matching between query and memory frames. However, most of the existing methods neglect the fact that videos carry rich temporal information yet redundant spatial information. In this case, direct pixel-level global matching will lead to ambiguous correspondences. In this work, we reconcile the inherent tension of spatial and temporal information to retrieve memory frame information along the object trajectory, and propose a novel and coherent Trajectory Memory Retrieval Network (TMRN) to equip with the trajectory information, including a spatial alignment module and a temporal aggregation module. The proposed TMRN enjoys several merits. First, TMRN is empowered to characterize the temporal correspondence which is in line with the nature of video in a data-driven manner. Second, we elegantly customize the spatial alignment module by coupling SVD initialization with agent-level correlation for representative agent construction and rectifying false matches caused by direct pairwise pixel-level correlation, respectively. Extensive experimental results on challenging benchmarks including DAVIS 2017 validation / test and Youtube-VOS 2018/2019 demonstrate that our TMRN, as a general plugin module, achieves consistent improvements over several leading methods. Rui Sun 0006, Yuan Wang 0064, Huayu Mai, Tianzhu Zhang 0001, Feng Wu 0001 |
ICCV | 1 |
| 2023 | Appearance Prompt Vision Transformer for Connectome ReconstructionabstractNeural connectivity reconstruction aims to understand the function of biological reconstruction and promote basic scientific research. The intricate morphology and densely intertwined branches make it an extremely challenging task. Most previous best-performing methods adopt affinity learning or metric learning. Nevertheless, they either neglect to model explicit voxel semantics caused by implicit optimization or are hysteresis to spatial information. Furthermore, the inherent locality of 3D CNNs limits modeling long-range dependencies, leading to sub-optimal results. In this work, we propose a coherent and unified Appearance Prompt Vision Transformer (APViT) to integrate affinity and metric learning to exploit the complementarity by learning long-range spatial dependencies. The proposed APViT enjoys several merits. First, the extension continuity-aware attention module aims at constructing hierarchical attention customized for neuron extensibility and slice continuity to learn instance voxel semantic context from a global perspective and utilize continuity priors to enhance voxel spatial awareness. Second, the appearance prompt modulator is responsible for leveraging voxel-adaptive appearance knowledge conditioned on affinity rich in spatial information to instruct instance voxel semantics, exploiting the potential of affinity learning to complement metric learning. Extensive experimental results on multiple challenging benchmarks demonstrate that our APViT achieves consistent improvements with huge flexibility under the same post-processing strategy. Rui Sun 0006, Naisong Luo, Yuwen Pan, Huayu Mai, Tianzhu Zhang 0001, Zhiwei Xiong, Feng Wu 0001 |
IJCAI | 1 |
| 2023 | Structure-Decoupled Adaptive Part Alignment Network for Domain Adaptive Mitochondria Segmentation
Rui Sun 0006, Huayu Mai, Naisong Luo, Tianzhu Zhang 0001, Zhiwei Xiong, Feng Wu 0001 |
MICCAI (4) | 1 |
| 2023 | DAW: Exploring the Better Weighting Function for Semi-supervised Semantic SegmentationabstractThe critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model’s generalization performance for robust segmentation. Existing methods tend to employ certain criteria (weighting function) to select pixel-level pseudo labels. However, the trade-off exists between inaccurate yet utilized pseudo-labels, and correct yet discarded pseudo-labels in these methods when handling pseudo-labels without thoughtful consideration of the weighting function, hindering the generalization ability of the model. In this paper, we systematically analyze the trade-off in previous methods when dealing with pseudo-labels. We formally define the trade-off between inaccurate yet utilized pseudo-labels, and correct yet discarded pseudo-labels by explicitly modeling the confidence distribution of correct and inaccurate pseudo-labels, equipped with a unified weighting function. To this end, we propose Distribution-Aware Weighting (DAW) to strive to minimize the negative equivalence impact raised by the trade-off. We find an interesting fact that the optimal solution for the weighting function is a hard step function, with the jump point located at the intersection of the two confidence distributions. Besides, we devise distribution alignment to mitigate the issue of the discrepancy between the prediction distributions of labeled and unlabeled data. Extensive experimental results on multiple benchmarks including mitochondria segmentation demonstrate that DAW performs favorably against state-of-the-art methods. Rui Sun 0006, Huayu Mai, Tianzhu Zhang 0001, Feng Wu 0001 |
NeurIPS | 1 |
| 2022 | Adaptive Agent Transformer for Few-Shot Segmentation
Yuan Wang 0064, Rui Sun 0006, Tianzhu Zhang 0001 |
ECCV (29) | 2 |
| 2022 | Electron Microscopy Image Registration with Transformers
Fuyu Feng, Tianzhu Zhang 0001, Rui Sun 0006, Zhiwei Xiong, Feng Wu 0001 |
ICONIP (3) | 3 |
| 2021 | Lesion-Aware Transformers for Diabetic Retinopathy GradingabstractDiabetic retinopathy (DR) is the leading cause of permanent blindness in the working-age population. And automatic DR diagnosis can assist ophthalmologists to design tailored treatments for patients, including DR grading and lesion discovery. However, most of existing methods treat DR grading and lesion discovery as two independent tasks, which require lesion annotations as a learning guidance and limits the actual deployment. To alleviate this problem, we propose a novel lesion-aware transformer (LAT) for DR grading and lesion discovery jointly in a unified deep model via an encoder-decoder structure including a pixel relation based encoder and a lesion filter based decoder. The proposed LAT enjoys several merits. First, to the best of our knowledge, this is the first work to formulate lesion discovery as a weakly supervised lesion localization problem via a transformer decoder. Second, to learn lesion filters well with only image-level labels, we design two effective mechanisms including lesion region importance and lesion region diversity for identifying diverse lesion regions. Extensive experimental results on three challenging benchmarks including Messidor-1, Messidor-2 and EyePACS demonstrate that the proposed LAT performs favorably against state-of-the-art DR grading and lesion discovery methods. Rui Sun 0006, Tianzhu Zhang 0001, Zhendong Mao 0001, Feng Wu 0001, Yongdong Zhang 0001 |
CVPR | 1 |