EDBT 2026 Demo / reviewers in the wild / expert
Zhe Chen 0013
dblp:06/4240-13
· DBLP profile ↗
36ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0001-5004-8975ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 14 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CERA: Conflict-Explicit Reflective Agent for Multimodal Emotion ReasoningabstractMultimodal Emotion Recognition (MER) aims to understand complex human emotions by jointly analyzing visual and textual data. However, in real-world scenarios, emotional cues from different modalities often contain conflict information, such as a smiling face paired with negative text, which poses great challenges for existing multimodal language models (MLLMs). Existing emotion MLLMs and multimodal emotion benchmarks often overlook or even intentionally avoid scenarios involving multimodal emotion conflicts, limiting their ability to reason about complex and contradictory affective cues. By addressing this, we propose Conflict-Explicit Reflective Agent (CERA), a training-free, conflict-aware, and language-driven agentic framework for MER. The concept of CERA is to treat modality emotion conflicts as meaningful signals and resolve them via a three-stage perception–evaluation–reflection reasoning loop. Firstly, the agent’s conflict-perceptive emotion graph construction module builds emotion graphs from fine-grained cues to reveal conflicts, and progressively refines them through iterative updates. Secondly, a reward model evaluates these graphs and produces natural language feedback that identifies unresolved conflicts. Lastly, the language-driven conflict refinement module generates graph editing signals from the feedback without any parameter tuning, enabling the overall CERA to refine its reasoning without training. Extensive experiments on two multimodal emotion datasets, MAFW and CH-SIMS, demonstrate that CERA significantly outperforms state-of-the-art training-free methods in both recognition accuracy and conflict interpretability, providing an effective training-free solution for complex emotional reasoning. Kejun Liu, Chang Tang, Zhe Chen 0013, Yibing Zhan |
ICMR | 7 |
| 2025 | ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry AreaabstractLarge Language Models (LLMs) have achieved remarkable success and have been applied across various scientific fields, including chemistry. However, many chemical tasks require the processing of visual information, which cannot be successfully handled by existing chemical LLMs. This brings a growing need for models capable of integrating multimodal information in the chemical domain. In this paper, we introduce ChemVLM, an open-source chemical multimodal large language model specifically designed for chemical applications. ChemVLM is trained on a carefully curated bilingual multimodal dataset that enhances its ability to understand both textual and visual chemical information, including molecular structures, reactions, and chemistry examination questions. We develop three datasets for comprehensive evaluation, tailored to Chemical Optical Character Recognition (OCR), Multimodal Chemical Reasoning (MMCR), and Multimodal Molecule Understanding tasks. We benchmark ChemVLM against a range of open-source and proprietary multimodal large language models on various tasks. Experimental results demonstrate that ChemVLM achieves competitive performance across all evaluated tasks. Junxian Li 0001, Di Zhang 0026, Xunzhi Wang, Zeying Hao, Jingdi Lei, Cai Zhou, Wei Liu 0123, Yaotian Yang, Xinrui Xiong, Weiyun Wang, Zhe Chen 0013, Wenhai Wang, Wei Li 0076, Mao Su, Shufei Zhang, Wanli Ouyang, Dongzhan Zhou |
AAAI | 12 |
| 2025 | On Geometry-Enhanced Parameter-Efficient Fine-Tuning for 3D Scene SegmentationabstractThe emergence of large-scale pre-trained point cloud models has significantly advanced 3D scene understanding, but adapting these models to specific downstream tasks typically demands full fine-tuning, incurring high computational and storage costs. Parameter-efficient fine-tuning (PEFT) techniques, successful in natural language processing and 2D vision tasks, would underperform when naively applied to 3D point cloud models due to significant geometric and spatial distribution shifts.
Existing PEFT methods commonly treat points as orderless tokens, neglecting important local spatial structures and global geometric contexts in 3D modeling.
To bridge this gap, we introduce the Geometric Encoding Mixer (GEM), a novel geometry-aware PEFT module specifically designed for 3D point cloud transformers. GEM explicitly integrates fine-grained local positional encodings with a lightweight latent attention mechanism to capture comprehensive global context, thereby effectively addressing the spatial and geometric distribution mismatch.
Extensive experiments demonstrate that GEM achieves performance comparable to or sometimes even exceeding full fine-tuning, while only updating 1.6% of the model's parameters, fewer than other PEFT methods.
With significantly reduced training time and memory requirements, our approach thus sets a new benchmark for efficient, scalable, and geometry-aware fine-tuning of large-scale 3D point cloud models.
Code is available at https://github.com/LiyaoTang/GEM. Liyao Tang, Zhe Chen 0013, Dacheng Tao |
NeurIPS | 2 |
| 2025 | Sample-Cohesive Pose-Aware Contrastive Facial Representation LearningabstractAbstract Self-supervised facial representation learning (SFRL) methods, especially contrastive learning (CL) methods, have been increasingly popular due to their ability to perform face understanding without heavily relying on large-scale well-annotated datasets. However, analytically, current CL-based SFRL methods still perform unsatisfactorily in learning facial representations due to their tendency to learn pose-insensitive features, resulting in the loss of some useful pose details. This could be due to the inappropriate positive/negative pair selection within CL. To conquer this challenge, we propose a Pose-disentangled Contrastive Facial Representation Learning (PCFRL) framework to enhance pose awareness for SFRL. We achieve this by explicitly disentangling the pose-aware features from non-pose face-aware features and introducing appropriate sample calibration schemes for better CL with the disentangled features. In PCFRL, we first devise a pose-disentangled decoder with a delicately designed orthogonalizing regulation to perform the disentanglement; therefore, the learning on the pose-aware and non-pose face-aware features would not affect each other. Then, we introduce a false-negative pair calibration module to overcome the issue that the two types of disentangled features may not share the same negative pairs for CL. Our calibration employs a novel neighborhood-cohesive pair alignment method to identify pose and face false-negative pairs, respectively, and further help calibrate them to appropriate positive pairs. Lastly, we devise two calibrated CL losses, namely calibrated pose-aware and face-aware CL losses, for adaptively learning the calibrated pairs more effectively, ultimately enhancing the learning with the disentangled features and providing robust facial representations for various downstream tasks. In the experiments, we perform linear evaluations on four challenging downstream facial tasks with SFRL using our method, including facial expression recognition, face recognition, facial action unit detection, and head pose estimation. Experimental results show that PCFRL outperforms existing state-of-the-art methods by a substantial margin, demonstrating the importance of improving pose awareness for SFRL. Our evaluation code and model will be available at https://github.com/fulaoze/CV/tree/main . Yuanyuan Liu 0004, Shaoze Feng, Yibing Zhan, Dapeng Tao, Zijing Chen, Zhe Chen 0013 |
Int. J. Comput. Vis. | 7 |
| 2025 | Noise-Resistant Multimodal Transformer for Emotion Recognition
Yuanyuan Liu 0004, Haoyu Zhang 0001, Yibing Zhan, Zijing Chen, Guanghao Yin, Zhe Chen 0013 |
Int. J. Comput. Vis. | 7 |
| 2025 | Learning General and Specific Embedding with Transformer for Few-Shot Object DetectionabstractAbstract Few-shot object detection (FSOD) studies how to detect novel objects with few annotated examples effectively. Recently, it has been demonstrated that decent feature embeddings, including the general feature embeddings that are more invariant to visual changes and the specific feature embeddings that are more discriminative for different object classes, are both important for FSOD. However, current methods lack appropriate mechanisms to sensibly cooperate both types of feature embeddings based on their importance to detecting objects of novel classes, which may result in sub-optimal performance. In this paper, to achieve more effective FSOD, we attempt to explicitly encode both general and specific feature embeddings using learnable tensors and apply a Transformer to help better incorporate them in FSOD according to their relations to the input object features. We thus propose a Transformer-based general and specific embedding learning (T-GSEL) method for FSOD. In T-GSEL, learnable tensors are employed in a three-stage pipeline, encoding feature embeddings in general level, intermediate level, and specific level, respectively. In each stage, we apply a Transformer to first model the relations of the corresponding embedding to input object features and then apply the estimated relations to refine the input features. Meanwhile, we further introduce cross-stage connections between embeddings of different stages to make them complement and cooperate with each other, delivering general, intermediate, and specific feature embeddings stage by stage and utilizing them together for feature refinement in FSOD. In practice, a T-GSEL module is easy to inject. Extensive empirical results further show that our proposed T-GSEL method achieves compelling FSOD performance on both PASCAL VOC and MS COCO datasets compared with other state-of-the-art approaches. Zhe Chen 0013, Jing Zhang 0037, Tongliang Liu, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2025 | Degradation-adaptive attack-robust self-supervised facial representation learning
Yuanyuan Liu 0004, Chang Tang, Kun Sun 0002, Yibing Zhan, Zhe Chen 0013 |
Neurocomputing | 6 |
| 2025 | Leveraging Eye Movement for Instructing Robust Video-Based Facial Expression RecognitionabstractVideo-based facial expression recognition (VFER) is challenging due to variations caused by cultural background and expression camouflage. To tackle these problems, researchers introduced eye movement signals to complement visual information. However, existing methods either require expensive devices to capture high-quality eye movements or can only extract low-quality eye movements visually, making them ineffective in the real world. To address this, we propose an eye movement-instructed VFER (EM-VFER) that leverages high-quality eye movements to instruct the visual learning, obtaining robust performance without requiring costly devices during inference. Specifically, our EM-VFER operates in two stages: the high-quality eye movement pre-training stage and the eye movement-instructed video fine-tuning stage. In the pre-training, we compile an Eye-behavior-aided Multimodal Emotion Recognition (EMER) dataset and use it to train a multimodal Transformer. During the fine-tuning, we propose a novel progressive eye movement-instructed learning to take better advantage of the prior knowledge about high-quality eye movement signals from EMER. The instructed fine-tuning model could then make more robust predictions on downstream facial expression datasets. We evaluate our approach on three macroexpression datasets (DFEW, MAFW and Aff-wild2) and two micro-expression datasets (CASME III and CASME II). The results demonstrate that EM-VFER significantly outperforms existing methods. The code will be available. Yuanyuan Liu 0004, Kejun Liu, Zijing Chen, Zhe Chen 0013, Chang Tang, Jingying Chen 0001, Shiguang Shan |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | MASDG: Multiview Augmented Single-Source Domain Generalization Method for Robust Remote Sensing Building ExtractionabstractDespite advances in deep learning for remote sensing building extraction (RSBE), Multi-target Domain RSBE (MD-RSBE) remains challenging, as it requires transferring knowledge from a labeled source domain to multiple unlabeled target domains, with domain shifts in texture, style, and semantics. Existing domain adaptation (DA) and generalization (DG) methods face significant limitations: DA requires target-domain training, while DG needs multi-source training, leading to high training costs and low generalization in practical MD-RSBE scenarios. To address this, we propose a Multi-view Augmented Single-source Domain Generalization (MASDG) method, which effectively mitigates domain shifts across RS source and target domains for robust MD-RSBE performance by enriching the diversity of the source domain through multi-view augmentation and enforcing semantic consistency. Specifically, MASDG consists of three key components: Texture-level Domain Augmentation (TDA) module, Style-level Domain Augmentation (SDA) module and Semantic-invariant Representation Learning (SRL). To mitigate texture-level domain shift, TDA first introduces parameter-optimized multi-layer random convolution to modify the texture of source image, generating texture-augmented image pairs for simulating real-world texture diversity across various RS domains. Then, with each image pair from TDA, SDA employs two paralleled encoders, namely the general feature encoder and the batch-guided style encoder, to formulate multi-view building features, further mitigating style-level domain shift. Finally, SRL ensures semantic-invariant representation learning via a dual mechanism, including multi-view segmentation loss and semantic consistency loss. The former generates predictions from diverse feature views (original, texture-augmented, style-augmented, etc.), while the latter performs semantic alignment by minimizing distribution discrepancies among predictions, bridging semantic inconsistency to enable robust segmentation. Extensive experiments across three different MD-RSBE settings with 7 different target domains demonstrate that our MASDG outperforms existing state-of-the-art methods by a significant margin. Yunjiao Liu, Yuanyuan Liu 0004, Kejun Liu, Chang Tang, Wujie Zhou, Zhe Chen 0013, Wei Xiang 0001, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Emotion-Oriented Cross-Modal Prompting and Alignment for Human-Centric Emotional Video CaptioningabstractHuman-centric Emotional Video Captioning (H-EVC) aims to generate fine-grained, emotion-related sentences for human-based videos, enhancing the understanding of human emotions and facilitating human-computer emotional interaction. However, existing video captioning methods often overlook subtle emotional clues and interactions in videos. As a result, the generated captions frequently lack emotional information. To address this, we proposeEmotion-orientedCross-modalPrompting andAlignment (ECPA), which improves HEVC accuracy by modeling fine-grained visual-textual emotion clues. Using large foundation models, ECPA introduces two learnable prompting strategies: visual emotion prompting (VEP) and textual emotion prompting (TEP), along with an emotion-oriented cross-modal alignment (ECA) module. VEP uses two levels of visual prompts,i.e., emotion recognition (ER) and action unit (AU), to focus on both coarse and fine visual emotional features. TEP devise two-level learnable textual prompts,i.e., sentence-level emotional tokens and word-level masked tokens to capture global and local textual emotion representations. ECA introduces another two levels of emotion-oriented prompt alignment learning mechanisms: the ER-sentence level and the AU-word level alignment losses. Both enhance the model's ability to capture and integrate both global and local cross-modal emotion semantics, thereby enabling the generation of fine-grained emotional linguistic descriptions in video captioning. Experiments show ECPA significantly outperforms state-of-the-art methods on various H-EVC datasets (relative improvements of 9.98%, 5.72%, 4.46%, 24.52% on MAFW, and 12.82%, 20.27%, 4.22%, 5.01% on EmVidCap across four evaluation metrics) and supports zero-shot tasks on MSVD and MSRVTT, demonstrating strong applicability and generalization. Yu Wang 0246, Yuanyuan Liu 0004, Shunping Zhou, Chang Tang, Wujie Zhou, Zhe Chen 0013 |
IEEE Trans. Multim. | 7 |
| 2024 | AVSegFormer: Audio-Visual Segmentation with TransformerabstractAudio-visual segmentation (AVS) aims to locate and segment the sounding objects in a given video, which demands audio-driven pixel-level scene understanding. The existing methods cannot fully process the fine-grained correlations between audio and visual cues across various situations dynamically. They also face challenges in adapting to complex scenarios, such as evolving audio, the coexistence of multiple objects, and more. In this paper, we propose AVSegFormer, a novel framework for AVS that leverages the transformer architecture. Specifically, It comprises a dense audio-visual mixer, which can dynamically adjust interested visual features, and a sparse audio-visual decoder, which implicitly separates audio sources and automatically matches optimal visual features. Combining both components provides a more robust bidirectional conditional multi-modal representation, improving the segmentation performance in different scenarios. Extensive experiments demonstrate that AVSegFormer achieves state-of-the-art results on the AVS benchmark. The code is available at https://github.com/vvvb-github/AVSegFormer. Shengyi Gao, Zhe Chen 0013, Wenhai Wang |
AAAI | 2 |
| 2024 | Structural Information Guided Multimodal Pre-training for Vehicle-Centric PerceptionabstractUnderstanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neglect the specific characteristics of vehicle perception in different tasks and might thus lead to sub-optimal performance. To address this issue, we propose a novel vehicle-centric pre-training framework called VehicleMAE, which incorporates the structural information including the spatial structure from vehicle profile information and the semantic structure from informative high-level natural language descriptions for effective masked vehicle appearance reconstruction. To be specific, we explicitly extract the sketch lines of vehicles as a form of the spatial structure to guide vehicle reconstruction. The more comprehensive knowledge distilled from the CLIP big model based on the similarity between the paired/unpaired vehicle image-text sample is further taken into consideration to help achieve a better understanding of vehicles. A large-scale dataset is built to pre-train our model, termed Autobot1M, which contains about 1M vehicle images and 12693 text information. Extensive experiments on four vehicle-based downstream tasks fully validated the effectiveness of our VehicleMAE. The source code and pre-trained models will be released at https://github.com/Event-AHU/VehicleMAE. Xiao Wang 0014, Chenglong Li 0002, Zhicheng Zhao 0001, Zhe Chen 0013, Yukai Shi, Jin Tang 0001 |
AAAI | 5 |
| 2024 | SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionabstractMulti-view camera-based 3D object detection has become popular due to its low cost, but accurately inferring 3D geometry solely from camera data remains challenging and may lead to inferior performance. Although distilling precise 3D geometry knowledge from LiDAR data could help tackle this challenge, the benefits of LiDAR information could be greatly hindered by the significant modality gap between different sensory modalities. To address this issue, we propose a Simulated multi-modal Distillation (SimDistill) method by carefully crafting the model architecture and distillation strategy. Specifically, we devise multi-modal architectures for both teacher and student models, including a LiDAR-camera fusion-based teacher and a simulated fusion-based student. Owing to the ``identical'' architecture design, the student can mimic the teacher to generate multi-modal features with merely multi-view images as input, where a geometry compensation module is introduced to bridge the modality gap. Furthermore, we propose a comprehensive multi-modal distillation scheme that supports intra-modal, cross-modal, and multi-modal fusion distillation simultaneously in the Bird's-eye-view space. Incorporating them together, our SimDistill can learn better feature representations for 3D object detection while maintaining a cost-effective camera-only deployment. Extensive experiments validate the effectiveness and superiority of SimDistill over state-of-the-art methods, achieving an improvement of 4.8% mAP and 4.1% NDS over the baseline detector. The source code will be released at https://github.com/ViTAE-Transformer/SimDistill. Haimei Zhao, Qiming Zhang 0001, Shanshan Zhao 0001, Zhe Chen 0013, Jing Zhang 0037, Dacheng Tao |
AAAI | 4 |
| 2024 | Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive PromptingabstractIn Video-based Facial Expression Recognition (V-FER), models are typically trained on closed-set datasets with a fixed number of known classes. However, these models struggle with unknown classes common in real-world scenarios. In this paper, we introduce a challenging Open-set Video-based Facial Expression Recognition (OV-FER) task, aiming to identify both known and new, unseen facial expressions. While existing approaches use large-scale vision-language models like CLIP to identify unseen classes, we argue that these methods may not adequately capture the subtle human expressions needed for OV-FER. To address this limitation, we propose a novel Human Expression-Sensitive Prompting (HESP) mechanism to significantly enhance CLIP's ability to model video-based facial expression details effectively. Our proposed HESP comprises three components: 1) a textual prompting module with learnable prompts to enhance CLIP's textual representation of both known and unknown emotions, 2) a visual prompting module that encodes temporal emotional information from video frames using expression-sensitive attention, equipping CLIP with a new visual modeling ability to extract emotion-rich information, and 3) an open-set multi-task learning scheme that promotes interaction between the textual and visual modules, improving the understanding of novel human emotions in video sequences. Extensive experiments conducted on four OV-FER task settings demonstrate that HESP can significantly boost CLIP's performance (a relative improvement of 17.93% on AUROC and 106.18% on OSCR) and outperform other state-of-the-art open-set video understanding methods by a large margin. Code is available at https://github.com/cosinehuang/HESP. Yuanyuan Liu 0004, Yibing Zhan, Zijing Chen, Zhe Chen 0013 |
ACM Multimedia | 6 |
| 2024 | VisEvent: Reliable Object Tracking via Collaboration of Frame and Event FlowsabstractDifferent from visible cameras which record intensity images frame by frame, the biologically inspired event camera produces a stream of asynchronous and sparse events with much lower latency. In practice, visible cameras can better perceive texture details and slow motion, while event cameras can be free from motion blurs and have a larger dynamic range which enables them to work well under fast motion and low illumination (LI). Therefore, the two sensors can cooperate with each other to achieve more reliable object tracking. In this work, we propose a large-scale Visible-Event benchmark (termed VisEvent) due to the lack of a realistic and scaled dataset for this task. Our dataset consists of 820 video pairs captured under LI, high speed, and background clutter scenarios, and it is divided into a training and a testing subset, each of which contains 500 and 320 videos, respectively. Based on VisEvent, we transform the event flows into event images and construct more than 30 baseline methods by extending current single-modality trackers into dual-modality versions. More importantly, we further build a simple but effective tracking algorithm by proposing a cross-modality transformer, to achieve more effective feature fusion between visible and event data. Extensive experiments on the proposed VisEvent dataset, FE108, COESOT, and two simulated datasets (i.e., OTB-DVS and VOT-DVS), validated the effectiveness of our model. The dataset and source code have been released on: https://github.com/wangxiao5791509/VisEvent_SOT_Benchmark. Xiao Wang 0014, Jianing Li 0001, Lin Zhu 0012, Zhe Chen 0013, Xin Li 0034, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Cybern. | 5 |
| 2023 | Pose-disentangled Contrastive Learning for Self-supervised Facial RepresentationabstractSelf-supervised facial representation has recently attracted increasing attention due to its ability to perform face understanding without relying on large-scale annotated datasets heavily. However, analytically, current contrastive-based self-supervised learning (SSL) still performs unsatisfactorily for learning facial representation. More specifically, existing contrastive learning (CL) tends to learn pose-invariant features that cannot depict the pose details of faces, compromising the learning performance. To conquer the above limitation of CL, we propose a novel Pose-disentangled Contrastive Learning (PCL) method for general self-supervised facial representation. Our PCL first devises a pose-disentangled decoder (PDD) with a delicately designed orthogonalizing regulation, which disentangles the pose-related features from the face-aware features; therefore, pose-related and other pose-unrelated facial information could be performed in individual subnetworks and do not affect each other's training. Furthermore, we introduce a pose-related contrastive learning scheme that learns pose-related information based on data augmentation of the same image, which would deliver more effective face-aware representation for various downstream tasks. We conducted linear evaluation on four challenging downstream facial understanding tasks, i.e., facial expression recognition, face recognition, AU detection and head pose estimation. Experimental results demonstrate that PCL significantly outperforms cuttingedge SSL methods. Our Code is available at https://github.com/DreamMr/PCL. Yuanyuan Liu 0004, Wenbin Wang 0001, Yibing Zhan, Shaoze Feng, Kejun Liu, Zhe Chen 0013 |
CVPR | 6 |
| 2023 | CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal PoseabstractAnimal pose estimation is challenging for existing image-based methods because of limited training data and large intra- and inter-species variances. Motivated by the progress of visual-language research, we propose that pre-trained language models (e.g., CLIP) can facilitate animal pose estimation by providing rich prior knowledge for describing animal keypoints in text. However, we found that building effective connections between pre-trained language models and visual animal keypoints is non-trivial since the gap between text-based descriptions and keypoint-based visual features about animal pose can be significant. To address this issue, we introduce a novel prompt-based Contrastive learning scheme for connecting Language and AniMal Pose (CLAMP) effectively. The CLAMP attempts to bridge the gap by adapting the text prompts to the animal keypoints during network training. The adaptation is decomposed into spatialaware and feature-aware processes, and two novel contrastive losses are devised correspondingly. In practice, the CLAMP enables the first cross-modal animal pose estimation paradigm. Experimental results show that our method achieves state-of-the-art performance under the supervised, few-shot, and zero-shot settings, outperforming image-based methods by a large margin. The code is available at https://github.com/xuzhang1199/CLAMP. Wen Wang 0009, Zhe Chen 0013, Yufei Xu, Jing Zhang 0037, Dacheng Tao |
CVPR | 3 |
| 2023 | All Points Matter: Entropy-Regularized Distribution Alignment for Weakly-supervised 3D SegmentationabstractPseudo-labels are widely employed in weakly supervised 3D segmentation tasks where only sparse ground-truth labels are available for learning.
Existing methods often rely on empirical label selection strategies, such as confidence thresholding, to generate beneficial pseudo-labels for model training.
This approach may, however, hinder the comprehensive exploitation of unlabeled data points.
We hypothesize that this selective usage arises from the noise in pseudo-labels generated on unlabeled data. The noise in pseudo-labels may result in significant discrepancies between pseudo-labels and model predictions, thus confusing and affecting the model training greatly.
To address this issue, we propose a novel learning strategy to regularize the generated pseudo-labels and effectively narrow the gaps between pseudo-labels and model predictions.
More specifically, our method introduces an Entropy Regularization loss and a Distribution Alignment loss for weakly supervised learning in 3D segmentation tasks, resulting in an ERDA learning strategy.
Interestingly, by using KL distance to formulate the distribution alignment loss, it reduces to a deceptively simple cross-entropy-based loss which optimizes both the pseudo-label generation network and the 3D segmentation network simultaneously.
Despite the simplicity, our method promisingly improves the performance.
We validate the effectiveness through extensive experiments on various baselines and large-scale datasets.
Results show that ERDA effectively enables the effective usage of all unlabeled data points for learning and achieves state-of-the-art performance under different settings.
Remarkably, our method can outperform fully-supervised baselines using only 1\% of true annotations.
Code and model will be made publicly available at https://github.com/LiyaoTang/ERDA. Liyao Tang, Zhe Chen 0013, Shanshan Zhao 0001, Dacheng Tao |
NeurIPS | 2 |
| 2023 | Transformer-Based Context Condensation for Boosting Feature Pyramids in Object DetectionabstractAbstract Current object detectors typically have a feature pyramid (FP) module for multi-level feature fusion (MFF) which aims to mitigate the gap between features from different levels and form a comprehensive object representation to achieve better detection performance. However, they usually require heavy cross-level connections or iterative refinement to obtain better MFF result, making them complicated in structure and inefficient in computation. To address these issues, we propose a novel and efficient context modeling mechanism that can help existing FPs deliver better MFF results while reducing the computational costs effectively. In particular, we introduce a novel insight that comprehensive contexts can be decomposed and condensed into two types of representations for higher efficiency. The two representations include a locally concentrated representation and a globally summarized representation, where the former focuses on extracting context cues from nearby areas while the latter extracts general contextual representations of the whole image scene as global context cues. By collecting the condensed contexts, we employ a Transformer decoder to investigate the relations between them and each local feature from the FP and then refine the MFF results accordingly. As a result, we obtain a simple and light-weight Transformer-based Context Condensation (TCC) module, which can boost various FPs and lower their computational costs simultaneously. Extensive experimental results on the challenging MS COCO dataset show that TCC is compatible to four representative FPs and consistently improves their detection accuracy by up to 7.8% in terms of average precision and reduce their complexities by up to around 20% in terms of GFLOPs, helping them achieve state-of-the-art performance more efficiently. Code will be released at https://github.com/zhechen/TCC . Zhe Chen 0013, Jing Zhang 0037, Yufei Xu, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2023 | Expression snippet transformer for robust video-based facial expression recognition
Yuanyuan Liu 0004, Wenbin Wang 0001, Chuanxu Feng, Haoyu Zhang 0001, Zhe Chen 0013, Yibing Zhan |
Pattern Recognit. | 5 |
| 2022 | SASA: Semantics-Augmented Set Abstraction for Point-Based 3D Object DetectionabstractAlthough point-based networks are demonstrated to be accurate for 3D point cloud modeling, they are still falling behind their voxel-based competitors in 3D detection. We observe that the prevailing set abstraction design for down-sampling points may maintain too much unimportant background information that can affect feature learning for detecting objects. To tackle this issue, we propose a novel set abstraction method named Semantics-Augmented Set Abstraction (SASA). Technically, we first add a binary segmentation module as the side output to help identify foreground points. Based on the estimated point-wise foreground scores, we then propose a semantics-guided point sampling algorithm to help retain more important foreground points during down-sampling. In practice, SASA shows to be effective in identifying valuable points related to foreground objects and improving feature learning for point-based 3D detection. Additionally, it is an easy-to-plug-in module and able to boost various point-based detectors, including single-stage and two-stage ones. Extensive experiments on the popular KITTI and nuScenes datasets validate the superiority of SASA, lifting point-based detection models to reach comparable performance to state-of-the-art voxel-based methods. Code is available at https://github.com/blakechen97/SASA. Zhe Chen 0013, Jing Zhang 0037, Dacheng Tao |
AAAI | 2 |
| 2022 | Recurrent Glimpse-based Decoder for Detection with TransformerabstractAlthough detection with Transformer (DETR) is increasingly popular, its global attention modeling requires an extremely long training period to optimize and achieve promising detection performance. Alternative to existing studies that mainly develop advanced feature or embedding designs to tackle the training issue, we point out that the Region-of-Interest (RoI) based detection refinement can easily help mitigate the difficulty of training for DETR methods. Based on this, we introduce a novel REcurrent Glimpse-based decOder (REGO) in this paper. In particular, the REGO employs a multi-stage recurrent processing structure to help the attention of DETR gradually focus on foreground objects more accurately. In each processing stage, visual features are extracted as glimpse features from RoIs with enlarged bounding box areas of detection results from the previous stage. Then, a glimpse-based decoder is introduced to provide refined detection results based on both the glimpse features and the attention modeling outputs of the previous stage. In practice, REGO can be easily embedded in representative DETR variants while maintaining their fully end-to-end training and inference pipelines. In particular, REGO helps Deformable DETR achieve 44.8 AP on the MSCOCO dataset with only 36 training epochs, compared with the first DETR and the Deformable DETR that require 500 and 50 epochs to achieve comparable performance, respectively. Experiments also show that REGO consistently boosts the performance of different DETR detectors by up to 7% relative gain at the same setting of 50 training epochs. Code is available via https://github.com/zhechen/Deformable-DETR-REGO. Zhe Chen 0013, Jing Zhang 0037, Dacheng Tao |
CVPR | 1 |
| 2022 | Contrastive Boundary Learning for Point Cloud SegmentationabstractPoint cloud segmentation is fundamental in understanding 3D environments. However, current 3D point cloud segmentation methods usually perform poorly on scene boundaries, which degenerates the overall segmentation performance. In this paper, we focus on the segmentation of scene boundaries. Accordingly, we first explore metrics to evaluate the segmentation performance on scene boundaries. To address the unsatisfactory performance on boundaries, we then propose a novel contrastive boundary learning (CBL) framework for point cloud segmentation. Specifically, the proposed CBL enhances feature discrimination between points across boundaries by contrasting their representations with the assistance of scene contexts at multiple scales. By applying CBL on three different baseline methods, we experimentally show that CBL consistently improves different baselines and assists them to achieve compelling performance on boundaries, as well as the overall performance, e.g. in mIoU. The experimental results demonstrate the effectiveness of our method and the importance of boundaries for 3D point cloud segmentation. Code and model will be made publicly available at https://github.com/LiyaoTang/contrastBoundary. Liyao Tang, Yibing Zhan, Zhe Chen 0013, Baosheng Yu, Dacheng Tao |
CVPR | 3 |
| 2022 | Pedestrian attribute recognition: A survey
Xiao Wang 0014, Shaofei Zheng, Aihua Zheng, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 5 |
| 2022 | Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-Based Beam SearchabstractTo track the target in a video, current visual trackers usually adopt greedy search for target object localization in each frame, that is, the candidate region with the maximum response score will be selected as the tracking result of each frame. However, we found that this may be not an optimal choice, especially when encountering challenging tracking scenarios such as heavy occlusion and fast motion. In particular, if a tracker drifts, errors will be accumulated and would further make response scores estimated by the tracker unreliable in future frames. To address this issue, we propose to maintain multiple tracking trajectories and apply beam search strategy for visual tracking, so that the trajectory with fewer accumulated errors can be identified. Accordingly, this paper introduces a novel multi-agent reinforcement learning based beam search tracking strategy, termed BeamTracking. It is mainly inspired by the image captioning task, which takes an image as input and generates diverse descriptions using beam search algorithm. Accordingly, we formulate the tracking as a sample selection problem fulfilled by multiple parallel decision-making processes, each of which aims at picking out one sample as their tracking result in each frame. Each maintained trajectory is associated with an agent to perform the decision-making and determine what actions should be taken to update related information. More specifically, using the classification-based tracker as the baseline, we first adopt bi-GRU to encode the target feature, proposal feature, and its response score into a unified state representation. The state feature and greedy search result are then fed into the first agent for independent action selection. Afterwards, the output action and state features are fed into the subsequent agent for diverse results prediction. When all the frames are processed, we select the trajectory with the maximum accumulated score as the tracking result. Extensive experiments on seven popular tracking benchmark datasets validated the effectiveness of the proposed algorithm. Xiao Wang 0014, Zhe Chen 0013, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2021 | A Shape Transformation-based Dataset Augmentation Framework for Pedestrian Detection
Zhe Chen 0013, Wanli Ouyang, Tongliang Liu, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2021 | Recursive Context Routing for Object Detection
Zhe Chen 0013, Jing Zhang 0037, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2021 | Towards High Performance Human Keypoint Detection
Jing Zhang 0037, Zhe Chen 0013, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2021 | Terra: A Smart and Sensible Digital Twin Framework for Robust Robot Deployment in Challenging EnvironmentsabstractDigital twin (DT) systems that replicate the physical world digitally are powerful tools for monitoring physical systems and evaluating algorithms, but current DT systems are commonly not applicable for robotic deployment and investigation. Meanwhile, current 3-D simulation-based robotic platforms do not model the dynamics of the physical world on-the-fly as done in DT systems, limiting their potential for the development of robotics in challenging environments. To tackle this issue, we propose the first robot-centered smart DT framework, namely, Terra, to facilitate the deployment of robots in challenging environments. The proposed Terra framework introduces a comprehensive DT representation to encode the useful real-time dynamics of both the physical world and the robot agent deployed therein. A multiview multimodality perception module is further devised for Terra to obtain high-level semantics and deliver a precise description of the current status of the environment and the robot agent. By mapping the perceived results to the virtual replica of the physical environment, Terra actively updates the action policy and sends it back to the agent, forming an integral and real-time information feedback loop. In practice, to help demonstrate the effectiveness and feasibility of the proposed framework, we deliberately set up a challenging unordered physical environment with many obstacles and a very simple robot aiming to fulfill a navigation task. Empirical results show that the proposed Terra framework successfully facilitates the robot to accomplish the task without causing hazards. Yamin Mo, Sihan Ma, Zhe Chen 0013, Jing Zhang 0037, Dacheng Tao |
IEEE Internet Things J. | 4 |
| 2021 | Dynamic Attention Guided Multi-Trajectory Analysis for Single Object TrackingabstractMost of the existing single object trackers track the target in a unitary local search window, making them particularly vulnerable to challenging factors such as heavy occlusions and out-of-view movements. Despite the attempts to further incorporate global search, prevailing mechanisms that cooperate local and global search are relatively static, thus are still sub-optimal for improving tracking performance. By further studying the local and global search results, we raise a question: can we allow more dynamics for cooperating both results? In this paper, we propose to introduce more dynamics by devising a dynamic attention-guided multi-trajectory tracking strategy. In particular, we construct dynamic appearance model that contains multiple target templates, each of which provides its own attention for locating the target in the new frame. Guided by different attention, we maintain diversified tracking results for the target to build multi-trajectory tracking history, allowing more candidates to represent the true target trajectory. After spanning the whole sequence, we introduce a multi-trajectory selection network to find the best trajectory that deliver improved tracking performance. Extensive experimental results show that our proposed tracking strategy achieves compelling performance on various large-scale tracking benchmarks. The project page of this paper can be found athttps://sites.google.com/view/mt-track/. Xiao Wang 0014, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | TextFuseNet: Scene Text Detection with Richer Fused FeaturesabstractArbitrary shape text detection in natural scenes is an extremely challenging task. Unlike existing text detection approaches that only perceive texts based on limited feature representations, we propose a novel framework, namely TextFuseNet, to exploit the use of richer features fused for text detection. More specifically, we propose to perceive texts from three levels of feature representations, i.e., character-, word- and global-level, and then introduce a novel text representation fusion technique to help achieve robust arbitrary text detection. The multi-level feature representation can adequately describe texts by dissecting them into individual characters while still maintaining their general semantics. TextFuseNet then collects and merges the texts’ features from different levels using a multi-path fusion architecture which can effectively align and fuse different representations. In practice, our proposed TextFuseNet can learn a more adequate description of arbitrary shapes texts, suppressing false positives and producing more accurate detection results. Our proposed framework can also be trained with weak supervision for those datasets that lack character-level annotations. Experiments on several datasets show that the proposed TextFuseNet achieves state-of-the-art performance. Specifically, we achieve an F-measure of 94.3% on ICDAR2013, 92.1% on ICDAR2015, 87.1% on Total-Text and 86.6% on CTW-1500, respectively. Zhe Chen 0013, Juhua Liu, Bo Du 0001 |
IJCAI | 2 |
| 2020 | ASTS: A Unified Framework for Arbitrary Shape Text SpottingabstractArbitrary shape text spotting remains a challenging computer vision task. In this paper, we propose an end-to-end trainable unified framework for arbitrary shape text spotting to overcome the limitations inherent in the existing methods. Specifically, we propose to perceive and understand text based on different levels of semantics, i.e ., holistic-, pixel- and sequence-level semantics, and then unify the recognized semantics for robust text spotting. To implement the framework, we customize the detection and mask branches of Mask R-CNN to explore both holistic- and pixel-level semantics for text recognition. According to the recognition results, the text spotting task can then be formulated in the two-dimensional feature space. Then, by feeding the two-dimensional feature maps into an additional text recognition branch, our framework further delivers one-dimensional sequence-level semantics for text recognition based on an attention-based sequence-to-sequence network. Finally, the results from all the three levels of semantics are merged as the final result. Therefore, our framework is capable of simultaneously recognizing texts from both the one- and two-dimensional perspectives, achieving highly comprehensive text recognition. In addition, because some existing datasets lack character-level annotations, the extensive descriptions of texts from our framework further allow us to use only word-level annotations as weak supervision for training a robust text spotting model. Experiments on ICDAR 2013, ICDAR 2015, and Total-Text show that our framework achieves state-of-the-art performance for both detection and recognition. Juhua Liu, Zhe Chen 0013, Bo Du 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2018 | Context Refinement for Object Detection
Zhe Chen 0013, Shaoli Huang, Dacheng Tao |
ECCV (8) | 1 |
| 2017 | RBNet: A Deep Neural Network for Unified Road and Road Boundary Detection
Zhe Chen 0013, Zijing Chen |
ICONIP (1) | 1 |
| 2017 | Generic Pixel Level Object Tracker Using Bi-Channel Fully Convolutional Network
Zijing Chen, Jun Li 0010, Zhe Chen 0013, Xinge You |
ICONIP (1) | 3 |
| 2015 | MUlti-Store Tracker (MUSTer): A cognitive psychology inspired approach to object trackingabstractVariations in the appearance of a tracked object, such as changes in geometry/photometry, camera viewpoint, illumination, or partial occlusion, pose a major challenge to object tracking. Here, we adopt cognitive psychology principles to design a flexible representation that can adapt to changes in object appearance during tracking. Inspired by the well-known Atkinson-Shiffrin Memory Model, we propose MUlti-Store Tracker (MUSTer), a dual-component approach consisting of short- and long-term memory stores to process target appearance memories. A powerful and efficient Integrated Correlation Filter (ICF) is employed in the short-term store for short-term tracking. The integrated long-term component, which is based on keypoint matching-tracking and RANSAC estimation, can interact with the long-term memory and provide additional information for output control. MUSTer was extensively evaluated on the CVPR2013 Online Object Tracking Benchmark (OOTB) and ALOV++ datasets. The experimental results demonstrated the superior performance of MUSTer in comparison with other state-of-art trackers. Zhibin Hong, Zhe Chen 0013, Chaohui Wang, Xue Mei, Danil V. Prokhorov, Dacheng Tao |
CVPR | 2 |