Tao Chen 0012

dblp:69/510-12 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0001-8239-1698ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Quadratic: Linear-Time Change Detection with RWKV
abstract
Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this conflict. By building upon the Receptance Weighted Key Value (RWKV) framework, our ChangeRWKV uniquely combines the parallelizable training of Transformers with the linear-time inference of RNNs. Our approach core features two key innovations: a hierarchical RWKV encoder that builds multi-resolution feature representation, and a novel Spatial-Temporal Fusion Module (STFM) engineered to resolve spatial misalignments across scales while distilling fine-grained temporal discrepancies. ChangeRWKV not only achieves state-of-the-art performance on the LEVIR-CD benchmark, with an 85.46% IoU and 92.16% F1 score, but does so while drastically reducing parameters and FLOPs compared to previous leading methods. This work demonstrates a new, efficient, and powerful paradigm for operational-scale change detection.
Gensheng Pei, Tao Chen 0012, Xia Yuan, Haofeng Zhang 0001, Xiangbo Shu, Yazhou Yao
AAAI3
2026 DepMatch: Boosting Semi-Supervised Semantic Segmentation by Exploring Depth Difference Knowledge
abstract
Existing semi-supervised semantic segmentation (SSS) methods fail to explore the potential of depth information in unlabeled data, as they suffer from 1) inter-class depth similarity, and 2) intra-class depth discrepancy. To address these challenges, this paper proposes DepMatch, a simple yet effective approach that leverages depth difference knowledge to guide consistency learning. Specifically, a Class-wise Depth Disparity Perception (CDDP) module is designed to exploit depth difference information, driven by class prediction priors, facilitating robust feature learning. Depth-feature discrepancy set is first constructed and then reliable pixel pairs are selected for inter-class depth disparity knowledge distillation. Simultaneously, exponential normalization is applied to intra-category depth disparity for suppressing large outlier variations, and an entropy-based adaptive weight is derived to prioritize feature learning of high entropy areas. Moreover, we propose the Uncertain Logit Disparity Regulation (ULDR) module, which leverages the depth variations at class boundaries to promote the mutual regulation of uncertain pixel logit information, enhancing the model's spatial understanding. Experiments on five public benchmarks show that DepMatch can be seamlessly incorporated as a plug-and-play plugin into popular SSS frameworks, achieving significant performance improvements across various visual encoders. The source code and models are made available at https://github.com/NUST-Machine-Intelligence-Laboratory/DepMatch.
Jianjian Yin, Xiruo Jiang, Tao Chen 0012, Gensheng Pei, Yazhou Yao, Fumin Shen, Heng Tao Shen
IEEE Trans. Image Process.3
2025 Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
abstract
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot classification and retrieval capabilities on various domains. However, CLIP’s training remains computationally intensive, with high demands on both data processing and memory. To address these challenges, recent masking strategies have emerged, focusing on the selective removal of image patches to improve training efficiency. Although effective, these methods often compromise key semantic information, resulting in suboptimal alignment between visual features and text descriptions. In this work, we present a concise yet effective approach called Patch Generation-to-Selection (CLIP-PGS) to enhance CLIP’s training efficiency while preserving critical semantic con tent. Our method introduces a gradual masking process in which a small set of candidate patches is first pre-selected as potential mask regions. Then, we apply Sobel edge detection across the entire image to generate an edge mask that prioritizes the retention of the primary object areas. Finally, similarity scores between the candidate mask patches and their neighboring patches are computed, with optimal transport normalization refining the selection process to ensure a balanced similarity matrix. Our approach, CLIP-PGS, sets new state-of-the-art results in zero-shot classification and retrieval tasks, achieving superior performance in robustness evaluation and language compositionality benchmarks.
Gensheng Pei, Tao Chen 0012, Xinhao Cai, Xiangbo Shu, Tianfei Zhou, Yazhou Yao
CVPR2
2025 DynaPlane-Lane: Dynamic Multi-Plane Geometry Learning for Robust Monocular 3D Lane Detection
abstract
Accurate monocular 3D lane detection remains a fundamental challenge in autonomous driving, due to the need to infer spatial geometry from single-view images under diverse road and environmental conditions. Existing approaches often struggle with limited geometric adaptability and inconsistent structural predictions, particularly in the presence of sloped or uneven road surfaces. In this paper, we propose DynaPlane-Lane, a geometry-aware framework that explicitly models road topology through a dynamic multi-plane representation. By partitioning the scene into near, transition, and far regions with learnable geometric transformations, our method effectively captures height changes and surface irregularities. To enhance spatial reasoning, we design a Multi-Scale Adaptive Feature Aggregation module with channel and spatial attention to fuse encoder features across multiple resolutions. Additionally, a Topology-Aware Lane Optimization objective is formulated to enforce continuity and smoothness of predicted lane curves under partial visibility. Comprehensive experiments on the OpenLane and Apollo 3D Synthetic benchmarks demonstrate that our approach achieves robust performance across diverse scenarios, validating the effectiveness of incorporating adaptive geometry and structural constraints into monocular 3D lane detection.
Chunying Song, Huafeng Liu 0004, Tao Chen 0012, Qiong Wang 0003
MMAsia3
2025 Semi-Supervised Semantic Segmentation With Multi-Constraint Consistency Learning
abstract
Consistency regularization has prevailed in semi-supervised semantic segmentation and achieved promising performance. However, existing methods typically concentrate on enhancing the Image-augmentation based Prediction consistency and optimizing the segmentation network as a whole, resulting in insufficient utilization of potential supervisory information. In this paper, we propose a Multi-Constraint Consistency Learning (MCCL) approach to facilitate the staged enhancement of the encoder and decoder. Specifically, we first design a feature knowledge alignment (FKA) strategy to promote the feature consistency learning of the encoder from image-augmentation. Our FKA encourages the encoder to derive consistent features for strongly and weakly augmented views from the perspectives of point-to-point alignment and prototype-based intra-class compactness. Moreover, we propose a self-adaptive intervention (SAI) module to increase the discrepancy of aligned intermediate feature representations, promoting Feature-perturbation based Prediction consistency learning. Self-adaptive feature masking and noise injection are designed in an instance-specific manner to perturb the features for robust learning of the decoder. Experimental results on Pascal VOC2012 and Cityscapes datasets demonstrate that our proposed MCCL achieves new state-of-the-art performance. The source code and models are made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/MCCL.
Jianjian Yin, Tao Chen 0012, Gensheng Pei, Huafeng Liu 0004, Yazhou Yao, Liqiang Nie, Xian-Sheng Hua 0001
IEEE Trans. Multim.2
2024 Adaptive Integration of Partial Label Learning and Negative Learning for Enhanced Noisy Label Learning
abstract
There has been significant attention devoted to the effectiveness of various domains, such as semi-supervised learning, contrastive learning, and meta-learning, in enhancing the performance of methods for noisy label learning (NLL) tasks. However, most existing methods still depend on prior assumptions regarding clean samples amidst different sources of noise (e.g., a pre-defined drop rate or a small subset of clean samples). In this paper, we propose a simple yet powerful idea called NPN, which revolutionizes Noisy label learning by integrating Partial label learning (PLL) and Negative learning (NL). Toward this goal, we initially decompose the given label space adaptively into the candidate and complementary labels, thereby establishing the conditions for PLL and NL. We propose two adaptive data-driven paradigms of label disambiguation for PLL: hard disambiguation and soft disambiguation. Furthermore, we generate reliable complementary labels using all non-candidate labels for NL to enhance model robustness through indirect supervision. To maintain label reliability during the later stage of model training, we introduce a consistency regularization term that encourages agreement between the outputs of multiple augmentations. Experiments conducted on both synthetically corrupted and real-world noisy datasets demonstrate the superiority of NPN compared to other state-of-the-art (SOTA) methods. The source code has been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/NPN.
Mengmeng Sheng, Zeren Sun, Zhenhuang Cai, Tao Chen 0012, Yazhou Yao
AAAI4
2024 VideoMAC: Video Masked Autoencoders Meet ConvNets
abstract
Recently, the advancement of self-supervised learning techniques, like masked autoencoders (MAE), has greatly influenced visual representation learning for images and videos. Nevertheless, it is worth noting that the predomi-nant approaches in existing masked image / video modeling rely excessively on resource-intensive vision transformers (ViTs) as the feature encoder. In this paper, we propose a new approach termed as VideoMAC, which combines video masked autoencoders with resource-friendly Con-vNets. Specifically, VideoMAC employs symmetric masking on randomly sampled pairs of video frames. To prevent the issue of mask pattern dissipation, we utilize ConvNets which are implemented with sparse convolutional operators as en-coders. Simultaneously, we present a simple yet effective masked video modeling (MVM) approach, a dual encoder architecture comprising an online encoder and an exponential moving average target encoder, aimed to facilitate inter-frame reconstruction consistency in videos. Additionally, we demonstrate that VideoMAC, empowering classical (ResNet) / modern (ConvNeXt) convolutional encoders to harness the benefits of MVM, outperforms ViT-based approaches on downstream tasks, including video object segmentation (+5.2% /6.4% J&F), body part propagation (+6.3% /3.1% mIoU), and human pose tracking (+10.2% / 11.1% [email protected]).
Gensheng Pei, Tao Chen 0012, Xiruo Jiang, Huafeng Liu 0004, Zeren Sun, Yazhou Yao
CVPR2
2024 Knowledge Transfer with Simulated Inter-image Erasing for Weakly Supervised Semantic Segmentation
Tao Chen 0012, Xiruo Jiang, Gensheng Pei, Zeren Sun, Yucheng Wang 0013, Yazhou Yao
ECCV (42)1
2024 Foster Adaptivity and Balance in Learning with Noisy Labels
Mengmeng Sheng, Zeren Sun, Tao Chen 0012, Shuchao Pang, Yucheng Wang 0013, Yazhou Yao
ECCV (27)3
2024 Relating CNN-Transformer Fusion Network for Remote Sensing Change Detection
abstract
While deep learning, particularly convolutional neural networks (CNNs), has revolutionized remote sensing (RS) change detection (CD), existing approaches often miss crucial features due to neglecting global context and incomplete change learning. Additionally, transformer networks struggle with low-level details. RCTNet addresses these limitations by introducing (1) an early fusion backbone to exploit both spatial and temporal features early on, (2) a Cross-Stage Aggregation (CSA) module for enhanced temporal representation, (3) a Multi-Scale Feature Fusion (MSF) module for enriched feature extraction in the decoder, and (4) an Efficient Self-deciphering Attention (ESA) module utilizing transformers to capture global information and fine-grained details for accurate change detection. Extensive experiments demonstrate RCTNet’s clear superiority over traditional RS image CD methods, showing significant improvement and an optimal balance between accuracy and computational cost. Our source codes and pre-trained models are available at: https://github.com/NUST-Machine-Intelligence-Laboratory/RCTNet.
Yuhao Gao, Gensheng Pei, Mengmeng Sheng, Zeren Sun, Tao Chen 0012, Yazhou Yao
ICME5
2024 Enhancing Robustness in Learning with Noisy Labels: An Asymmetric Co-Training Approach
abstract
Label noise, an inevitable issue in various real-world datasets, tends to impair the performance of deep neural networks. A large body of literature focuses on symmetric co-training, aiming to enhance model robustness by exploiting interactions between models with distinct capabilities. However, the symmetric training processes employed in existing methods often culminate in model consensus, diminishing their efficacy in handling noisy labels. To this end, we propose an Asymmetric Co-Training (ACT) method to mitigate the detrimental effects of label noise. Specifically, we introduce an asymmetric training framework in which one model (i.e., RTM) is robustly trained with a selected subset of clean samples while the other (i.e., NTM) is conventionally trained using the entire training set. We propose two novel criteria based on agreement and discrepancy between models, establishing asymmetric sample selection and mining. Moreover, a metric, derived from the divergence between models, is devised to quantify label memorization, guiding our method in determining the optimal stopping point for sample mining. Finally, we propose to dynamically re-weight identified clean samples according to their reliability inferred from historical information. We additionally employ consistency regularization to achieve further performance improvement. Extensive experimental results on synthetic and real-world datasets demonstrate the effectiveness and superiority of our method.
Mengmeng Sheng, Zeren Sun, Gensheng Pei, Tao Chen 0012, Haonan Luo 0002, Yazhou Yao
ACM Multimedia4
2024 Class Probability Space Regularization for semi-supervised semantic segmentation
Jianjian Yin, Tao Chen 0012, Yi Chen 0023, Yazhou Yao
Comput. Vis. Image Underst.3
2024 Holistic Prototype Attention Network for Few-Shot Video Object Segmentation
abstract
Few-shot video object segmentation (FSVOS) aims to segment dynamic objects of unseen classes by resorting to a small set of support images that contain pixel-level object annotations. Existing methods have demonstrated that the domain agent-based attention mechanism is effective in FSVOS by learning the correlation between support images and query frames. However, the agent frame contains redundant pixel information and background noise, resulting in inferior segmentation performance. Moreover, existing methods tend to ignore inter-frame correlations in query videos. To alleviate the above dilemma, we propose a holistic prototype attention network (HPAN) for advancing FSVOS. Specifically, HPAN introduces a prototype graph attention module (PGAM) and a bidirectional prototype attention module (BPAM), transferring informative knowledge from seen to unseen classes. PGAM generates local prototypes from all foreground features and then utilizes their internal correlations to enhance the representation of the holistic prototypes. BPAM exploits the holistic information from support images and video frames by fusing co-attention and self-attention to achieve support-query semantic consistency and inner-frame temporal consistency. Extensive experiments on YouTube-FSVOS have been provided to demonstrate the effectiveness and superiority of our proposed HPAN method. Our source code and models are available anonymously at https://github.com/NUST-Machine-Intelligence-Laboratory/HPAN.
Tao Chen 0012, Xiruo Jiang, Yazhou Yao, Guosen Xie, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.2
2024 Spatial Structure Constraints for Weakly Supervised Semantic Segmentation
abstract
The image-level label has prevailed in weakly supervised semantic segmentation tasks due to its easy availability. Since image-level labels can only indicate the existence or absence of specific categories of objects, visualization-based techniques have been widely adopted to provide object location clues. Considering class activation maps (CAMs) can only locate the most discriminative part of objects, recent approaches usually adopt an expansion strategy to enlarge the activation area for more integral object localization. However, without proper constraints, the expanded activation will easily intrude into the background region. In this paper, we propose spatial structure constraints (SSC) for weakly supervised semantic segmentation to alleviate the unwanted object over-activation of attention expansion. Specifically, we propose a CAM-driven reconstruction module to directly reconstruct the input image from deep CAM features, which constrains the diffusion of last-layer object attention by preserving the coarse spatial structure of the image content. Moreover, we propose an activation self-modulation module to refine CAMs with finer spatial structure details by enhancing regional consistency. Without external saliency models to provide background clues, our approach achieves 72.7% and 47.0% mIoU on the PASCAL VOC 2012 and COCO datasets, respectively, demonstrating the superiority of our proposed approach. The source codes and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/SSC.
Tao Chen 0012, Yazhou Yao, Xingguo Huang, Zechao Li, Liqiang Nie, Jinhui Tang 0001
IEEE Trans. Image Process.1
2023 Semi-Supervised Semantic Segmentation With Region Relevance
abstract
Semi-supervised semantic segmentation aims to learn from a small amount of labeled data and plenty of unlabeled ones for the segmentation task. The most common approach is to generate pseudo-labels for unlabeled images to augment the training data. However, the noisy pseudo-labels will lead to cumulative classification errors and aggravate the local inconsistency in prediction. This paper proposes a Region Relevance Network (RRN) to alleviate the problem mentioned above. Specifically, we first introduce a local pseudo-label filtering module that leverages discriminator networks to assess the accuracy of the pseudo-label at the region level. A local selection loss is proposed to mitigate the negative impact of wrong pseudo-labels in consistency regularization training. In addition, we propose a dynamic region-loss correction module, which takes the merit of network diversity to further rate the reliability of pseudo-labels and correct the convergence direction of the segmentation network with a dynamic region loss. Extensive experiments are conducted on PASCAL VOC 2012 and Cityscapes datasets with varying amounts of labeled data, demonstrating that our proposed approach achieves state-of-the-art performance compared to current counterparts. Our code is available at: https://github.com/NUST-Machine-Intelligence-Laboratory/TorchSemiSeg2.
Tao Chen 0012, Qiong Wang 0003, Yazhou Yao
ICME2
2023 Multi-Granularity Denoising and Bidirectional Alignment for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) models relying on class activation maps (CAMs) have achieved desirable performance comparing to the non-CAMs-based counterparts. However, to guarantee WSSS task feasible, we need to generate pseudo labels by expanding the seeds from CAMs which is complex and time-consuming, thus hindering the design of efficient end-to-end (single-stage) WSSS approaches. To tackle the above dilemma, we resort to the off-the-shelf and readily accessible saliency maps for directly obtaining pseudo labels given the image-level class labels. Nevertheless, the salient regions may contain noisy labels and cannot seamlessly fit the target objects, and saliency maps can only be approximated as pseudo labels for simple images containing single-class objects. As such, the achieved segmentation model with these simple images cannot generalize well to the complex images containing multi-class objects. To this end, we propose an end-to-end multi-granularity denoising and bidirectional alignment (MDBA) model, to alleviate the noisy label and multi-class generalization issues. Specifically, we propose the online noise filtering and progressive noise detection modules to tackle image-level and pixel-level noise, respectively. Moreover, a bidirectional alignment mechanism is proposed to reduce the data distribution gap at both input and output space with simple-to-complex image synthesis and complex-to-simple adversarial learning. MDBA can reach the mIoU of 69.5% and 70.2% on validation and test sets for the PASCAL VOC 2012 dataset. The source codes and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/MDBA.
Tao Chen 0012, Yazhou Yao, Jinhui Tang 0001
IEEE Trans. Image Process.1
2023 Hierarchical Graph Pattern Understanding for Zero-Shot Video Object Segmentation
abstract
The optical flow guidance strategy is ideal for obtaining motion information of objects in the video. It is widely utilized in video segmentation tasks. However, existing optical flow-based methods have a significant dependency on optical flow, which results in poor performance when the optical flow estimation fails for a particular scene. The temporal consistency provided by the optical flow could be effectively supplemented by modeling in a structural form. This paper proposes a new hierarchical graph neural network (GNN) architecture, dubbed hierarchical graph pattern understanding (HGPU), for zero-shot video object segmentation (ZS-VOS). Inspired by the strong ability of GNNs in capturing structural relations, HGPU innovatively leverages motion cues (i.e., optical flow) to enhance the high-order representations from the neighbors of target frames. Specifically, a hierarchical graph pattern encoder with message aggregation is introduced to acquire different levels of motion and appearance features in a sequential manner. Furthermore, a decoder is designed for hierarchically parsing and understanding the transformed multi-modal contexts to achieve more accurate and robust results. HGPU achieves state-of-the-art performance on four publicly available benchmarks (DAVIS-16, YouTube-Objects, Long-Videos and DAVIS-17). Code and pre-trained model can be found at https://github.com/NUST-Machine-Intelligence-Laboratory/HGPU.
Gensheng Pei, Fumin Shen, Yazhou Yao, Tao Chen 0012, Xian-Sheng Hua 0001, Heng Tao Shen
IEEE Trans. Image Process.4
2023 Saliency Guided Inter- and Intra-Class Relation Constraints for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation with only image-level labels aims to reduce annotation costs for the segmentation task. Existing approaches generally leverage class activation maps (CAMs) to locate the object regions for pseudo label generation. However, CAMs can only discover the most discriminative parts of objects, thus leading to inferior pixel-level pseudo labels. To address this issue, we propose a saliency guidedInter- andIntra-ClassRelationConstrained (I$^{2}$CRC) framework to assist the expansion of the activated object regions in CAMs. Specifically, we propose a saliency guided class-agnostic distance module to pull the intra-category features closer by aligning features to their class prototypes. Further, we propose a class-specific distance module to push the inter-class features apart and encourage the object region to have a higher activation than the background. Besides strengthening the capability of the classification network to activate more integral object regions in CAMs, we also introduce an object guided label refinement module to take a full use of both the segmentation prediction and the initial labels for obtaining superior pseudo-labels. Extensive experiments on PASCAL VOC 2012 and COCO datasets demonstrate well the effectiveness of I$^{2}$CRC over other state-of-the-art counterparts.
Tao Chen 0012, Yazhou Yao, Lei Zhang 0054, Qiong Wang 0003, Guosen Xie, Fumin Shen
IEEE Trans. Multim.1
2023 FECANet: Boosting Few-Shot Semantic Segmentation With Feature-Enhanced Context-Aware Network
abstract
Few-shot semantic segmentation is the task of learning to locate each pixel of the novel class in the query image with only a few annotated support images. The current correlation-based methods construct pair-wise feature correlations to establish the many-to-many matching because the typical prototype-based approaches cannot learn fine-grained correspondence relations. However, the existing methods still suffer from the noise contained in naive correlations and the lack of context semantic information in correlations. To alleviate these problems mentioned above, we propose a Feature-Enhanced Context-Aware Network (FECANet). Specifically, a feature enhancement module is proposed to suppress the matching noise caused by inter-class local similarity and enhance the intra-class relevance in the naive correlation. In addition, we propose a novel correlation reconstruction module that encodes extra correspondence relations between foreground and background and multi-scale context semantic features, significantly boosting the encoder to capture a reliable matching pattern. Experiments on PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that our proposed FECANet leads to remarkable improvement compared to previous state-of-the-arts, demonstrating its effectiveness. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/FECANET.
Huafeng Liu 0004, Tao Chen 0012, Qiong Wang 0003, Yazhou Yao, Xian-Sheng Hua 0001
IEEE Trans. Multim.3
2022 A Novel Multi-Sample Data Augmentation Method for Oriented Object Detection in Remote Sensing Images
abstract
Data augmentation is widely used in computer vision tasks for enhancing the diversity of training data. However, due to sample redundancy and lack of object background, it is challenging to apply traditional data augmentation techniques to oriented object detection in remote sensing images. In this work, we propose SSMup (specifically synthetic mineral oversampling with mosaic and mixup), a multi-sample data augmentation method, to improve object detection performance in remote sensing images. Our method integrates Mosaic, Mixup, and SSMOTE to enable even distribution of target objects in augmented samples. Moreover, it equips the augmented samples with rich background information. Compared to existing state-of-the-art methods, our proposed method can remarkably improve the detection and generalization performance in remote sensing images. Comprehensive experiments are provided to demonstrate the effectiveness of the proposed method.
Guhua Chen, Gensheng Pei, Tao Chen 0012, Zhenmin Tang
MMSP4
2022 Feature Difference Enhancement Fusion for Remote Sensing Image Change Detection
Gensheng Pei, Tao Chen 0012, Yazhou Yao
PRCV (3)4
2022 Few-Shot Object Detection via Understanding Convolution and Attention
Jiaxing Tong, Tao Chen 0012, Qiong Wang 0003, Yazhou Yao
PRCV (1)2
2022 Enhanced Feature Alignment for Unsupervised Domain Adaptation of Semantic Segmentation
abstract
Unsupervised domain adaptation for semantic segmentation aims to transfer knowledge from a labeled source domain to another unlabeled target domain. However, due to the label noise and domain mismatch, learning directly from source domain data tends to have poor performance. Though adversarial learning methods strive to reduce domain discrepancies by aligning feature distributions, traditional methods suffer from the training imbalance and feature distortion problems. Besides, due to the absence of target domain labels, the classifier is blind to features from the target domain during training. Consequently, the final classifier overfits the source domain features and usually fails to predict the structured outputs of the target domain. To alleviate these problems, we focus on enhancing the adversarial learning based feature alignment from three perspectives. First, a classification constrained discriminator is proposed to balance the adversarial training and alleviate the feature distortion problem. Next, to alleviate the classifier overfitting problem, self-training is collaboratively used to learn a domain robust classifier with target domain pseudo labels. Moreover, an efficient class centroid calculation module is proposed and the domain discrepancy is further reduced by aligning the feature centroids of the same class from different domains. Experimental evaluations on GTA5$\rightarrow$Cityscapes and SYNTHIA$\rightarrow$Cityscapes demonstrate state-of-the-art results compared to other counterpart methods. The source code and models have been made available at.11[Online]. Available:https://github.com/NUST-Machine-Intelligence-Laboratory/EFA.
Tao Chen 0012, Shuihua Wang, Qiong Wang 0003, Zheng Zhang 0006, Guosen Xie, Zhenmin Tang
IEEE Trans. Multim.1
2022 Semantically Meaningful Class Prototype Learning for One-Shot Image Segmentation
abstract
One-shot semantic image segmentation aims to segment the object regions for the novel class with only one annotated image. Recent works adopt the episodic training strategy to mimic the expected situation at testing time. However, these existing approaches simulate the test conditions too strictly during the training process, and thus cannot make full use of the given label information. Besides, these approaches mainly focus on the foreground-background target class segmentation setting. They only utilize binary mask labels for training. In this paper, we propose to leverage the multi-class label information during the episodic training. It will encourage the network to generate more semantically meaningful features for each category. After integrating the target class cues into the query features, we then propose a pyramid feature fusion module to mine the fused features for the final classifier. Furthermore, to take more advantage of the support image-mask pair, we propose a self-prototype guidance branch to support image segmentation. It can constrain the network for generating more compact features and a robust prototype for each semantic class. For inference, we propose a fused prototype guidance branch for the segmentation of the query image. Specifically, we leverage the prediction of the query image to extract the pseudo-prototype and combine it with the initial prototype. Then we utilize the fused prototype to guide the final segmentation of the query image. Extensive experiments demonstrate the superiority of our proposed approach. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/SMCP.
Tao Chen 0012, Guosen Xie, Yazhou Yao, Qiong Wang 0003, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.1
2021 Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of pseudo labels within the image’s salient region. In this work, we propose a non-salient region object mining approach for weakly supervised semantic segmentation. We introduce a graph-based global reasoning unit to strengthen the classification network’s ability to capture global relations among disjoint and distant regions. This helps the network activate the object features outside the salient area. To further mine the non-salient region objects, we propose to exert the segmentation network’s self-correction ability. Specifically, a potential object mining module is proposed to reduce the false-negative rate in pseudo labels. Moreover, we propose a non-salient region masking module for complex images to generate masked pseudo labels. Our non-salient region masking module helps further discover the objects in the non-salient region. Extensive experiments on the PASCAL VOC dataset demonstrate state-of-the-art results compared to current methods. The source codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/nsrom.
Yazhou Yao, Tao Chen 0012, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Zhenmin Tang, Jian Zhang 0002
CVPR2
2020 Classification Constrained Discriminator For Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptation for semantic segmentation aims to transfer knowledge from label-rich synthetic datasets to real-world images without any annotation. The traditional adversarial learning methods for domain adaptation learn to extract domain-invariant feature representations by aligning the feature distributions of both domains. However, these methods suffer from an imbalance in adversarial training and feature distortion. In this work, we propose a classification constrained discriminator to alleviate these problems. Specifically, we first propose to balance the adversarial training by eliminating any pooling layers or strided convolutions in the discriminator. Then, we propose to constrain the discriminator with an auxiliary classification loss to help the feature generator extract the domain-invariant features that are useful for segmentation rather than just ambiguous features to fool the domain discriminator. Extensive experiments demonstrate the superiority of our proposed approach. The source code and models have been made available at https://github.com/NUSTMachine-Intelligence-Laboratory/ccd.
Tao Chen 0012, Jian Zhang 0002, Guosen Xie, Yazhou Yao, Xiaoshui Huang, Zhenmin Tang
ICME1