EDBT 2026 Demo / reviewers in the wild / expert
Zhe Cao 0001
dblp:93/5097-1
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0009-0005-1556-8058ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | It Takes Two: Multi-Frequency Perception With Complementary Fusion Network for Complex Scene SegmentationabstractComplex scene segmentation aims to segment objects with intricate details or those concealed within the background. Despite significant advancements, a persistent challenge remains: accurately identifying object edges in backgrounds with high inherent similarity and complex structures. To address this, we identify the prevalent spectral bias in image segmentation, where networks preferentially learn low-frequency information, as a key impediment to recognizing and learning object edges, which are rich in high-frequency details. To mitigate this bias, we propose MCNet, a segmentation framework designed to promote balanced frequency learning. MCNet comprises two primary components: multi-frequency perception (MP), which independently captures high-frequency details and low-frequency structural components of objects, and complementary fusion (CF), which intelligently fuses these distinct frequency features through learnable, adaptive mechanisms. Crucially, MCNet employs a novel frequency-aware consistency adversarial loss to explicitly guide the learning across different frequency bands. MCNet effectively integrates MP and CF, enhancing the detection of high-frequency details and low-frequency structures, thereby alleviating challenges posed by spectral bias. We evaluate the proposed method on complex scene segmentation tasks, including camouflaged object detection and dichotomous image segmentation. Through extensive comparisons with 31 existing methods across 8 benchmark datasets, we demonstrate the superiority of the proposed method. Jin Zhang 0021, Ruiheng Zhang 0001, Zhe Cao 0001, Lixin Xu 0001, Xi Chen 0090, Min Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | RADCI: A Synchronized Radar-RGBT Object Detecting-Tracking Dataset And A BenchmarkabstractHigh-quality perception is crucial in autonomous driving and monitoring systems, where millimeter-wave radar and infrared cameras play important roles due to their robustness and reliability under harsh conditions. Both technologies can serve as low-cost supplements to optical image detection, improving overall system robustness. However, there is currently a lack of widely applicable feature-level fusion methods and multimodal datasets to effectively integrate visible light with these two heterogeneous data types for multiple tasks. In this work, we collect a new multimodal dataset, RADCI8, which synchronizes data from a camera, an infrared camera, and a radar for target detection and tracking. The dataset includes 2D image annotations, radar RAD tensor data with distance, angle, and Doppler information, as well as target ID annotations in both data formats. In addition, to address the incomplete use of radar data in previous fusion algorithms, we propose a detection method that fuses image and radar features using feature concatenation and an attention mechanism. Our proposed algorithm achieves 51.5% AP with an IOU of 50:95 on 2D bounding box prediction, significantly improving average detection accuracy over vision-based methods and maintaining robustness even when a single sensor degrades. Ruiheng Zhang 0001, Zhe Cao 0001, Biwen Yang, Jin Zhang 0021, Guanyu Liu |
ICASSP | 4 |
| 2025 | IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale BenchmarkabstractReal-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infrared-text. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models. Zhe Cao 0001, Jin Zhang 0021, Ruiheng Zhang 0001 |
ICCV | 1 |
| 2025 | MMCSBench: A Fine-Grained Benchmark for Large Vision-Language Models in Camouflage ScenesabstractCurrent camouflaged object detection methods predominantly follow discriminative segmentation paradigms and heavily rely on predefined categories present in the training data, limiting their generalization to unseen or emerging camouflage objects. This limitation is further compounded by the labor-intensive and time-consuming nature of collecting camouflage imagery. Although Large Vision-Language Models (LVLMs) show potential to improve such issues with their powerful generative capabilities, their understanding of camouflage scenes is still insufficient. To bridge this gap, we introduce MMCSBench, the first comprehensive multimodal benchmark designed to evaluate and advance LVLM capabilities in camouflage scenes. MMCSBench comprises 22,537 images and 76,843 corresponding image-text pairs across five fine-grained camouflage tasks. Additionally, we propose a new task, Camouflage Efficacy Assessment (CEA), aimed at quantitatively evaluating the camouflage effectiveness of objects in images and enabling automated collection of camouflage images from large-scale databases. Extensive experiments on 26 LVLMs reveal significant shortcomings in models' ability to perceive and interpret camouflage scenes. These findings highlight the fundamental differences between natural and camouflaged visual inputs, offering insights for future research in advancing LVLM capabilities within this challenging domain. Jing Zhang 0037, Ruiheng Zhang 0001, Zhe Cao 0001, Kaizheng Chen |
NeurIPS | 3 |
| 2025 | Visible-Infrared Person Re-Identification With Real-World Label NoiseabstractIn recent years, growing needs for advanced security and traffic management have significantly heightened the prominence of the visible-infrared person re-identification community (VI-ReID), garnering considerable attention. A critical challenge in VI-ReID is the performance degradation attributable to label noise, an issue that becomes even more pronounced in cross-modal scenarios due to an increased likelihood of data confusion. While previous methods have achieved notable successes, they often overlook the complexities of instance-dependent and real-world noise, creating a disconnect from the practical applications of person re-identification. To bridge this gap, our research analyzes the primary sources of label noise in real-world settings, which include a) instantiated identities, b) blurry infrared images, and c) annotators’ errors. In response to these challenges, we develop a Robust Hybrid Loss function (RHL) that enables targeted recognition and retrieval optimization through a more fine-grained division of the noisy dataset. The proposed method categorises data into three sets: clean, obviously noisy, and indistinguishably noisy, with bespoke loss calculations for each category. The identification loss is structured to address the varied nature of these sets specifically. For the retrieval sub-task, we utilize an enhanced triplet loss, adept at handling noisy correspondences. Furthermore, to empirically validate our method, we have re-annotated a real-world dataset, SYSU-Real. Our experiments on SYSU-MM01 and RegDB, conducted under various noise ratios of random and instance-dependent label noise, demonstrate the generalized robustness and effectiveness of our proposed approach. Ruiheng Zhang 0001, Zhe Cao 0001, Yan Huang 0023, Shuo Yang 0006, Lixin Xu 0001, Min Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Cognition-Driven Structural Prior for Instance-Dependent Label Transition Matrix EstimationabstractThe label transition matrix has emerged as a widely accepted method for mitigating label noise in machine learning. In recent years, numerous studies have centered on leveraging deep neural networks to estimate the label transition matrix for individual instances within the context of instance-dependent noise. However, these methods suffer from low search efficiency due to the large space of feasible solutions. Behind this drawback, we have explored that the real murderer lies in the invalid class transitions, that is, the actual transition probability between certain classes is zero but is estimated to have a certain value. To mask the invalid class transitions, we introduced a human-cognition-assisted method with structural information from human cognition. Specifically, we introduce a structured transition matrix network (STMN) designed with an adversarial learning process to balance instance features and prior information from human cognition. The proposed method offers two advantages: 1) better estimation effectiveness is obtained by sparing the transition matrix and 2) better estimation accuracy is obtained with the assistance of human cognition. By exploiting these two advantages, our method parametrically estimates a sparse label transition matrix, effectively converting noisy labels into true labels. The efficiency and superiority of our proposed method are substantiated through comprehensive comparisons with state-of-the-art methods on three synthetic datasets and a real-world dataset. Our code will be available at https://github.com/WheatCao/STMN-Pytorch. Ruiheng Zhang 0001, Zhe Cao 0001, Shuo Yang 0006, Lingyu Si, Lixin Xu 0001, Fuchun Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Learning Camouflaged Object Detection from Noisy Pseudo Label
Jin Zhang 0021, Ruiheng Zhang 0001, Yanjiao Shi, Zhe Cao 0001, Nian Liu 0002, Fahad Shahbaz Khan |
ECCV (1) | 4 |
| 2024 | Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryabstractExisting paradigms of pushing the state of the art require exponentially more training data in many fields. Coreset selection seeks to mitigate this growing demand by identifying the most efficient subset of training data. In this paper, we delve into geometry-based coreset methods and preliminarily link the geometry of data distribution with models’ generalization capability in theoretics. Leveraging these theoretical insights, we propose a novel coreset construction method by selecting training samples to reconstruct the decision boundary of a deep neural network learned on the full dataset. Extensive experiments across various popular benchmarks demonstrate the superiority of our method over multiple competitors. For the first time, our method achieves a 50% data pruning rate on the ImageNet-1K dataset while sacrificing less than 1% in accuracy. Additionally, we showcase and analyze the remarkable cross-architecture transferability of the coresets derived from our approach. Shuo Yang 0006, Zhe Cao 0001, Ruiheng Zhang 0001, Ping Luo 0002, Shengping Zhang, Liqiang Nie |
ICML | 2 |
| 2024 | Concentrating Estimation Attention: Human Prior Constrained Methods for Robust Classification
Zhe Cao 0001, Shuo Yang 0006, Hongbin Pei, Yan Huang 0023, Yushu Yu, Ruiheng Zhang 0001 |
PRCV (15) | 1 |
| 2024 | Part-Aware Correlation Networks for Few-Shot LearningabstractFew-shot learning brings the machine close to human thinking which enables fast learning with limited samples. Recent work considers local features to achieve contextual semantic complementation, while they are merely coarsened feature observations that can only extract insignificant label correlations. On the contrary, partial properties of few-shot examples significantly draw the implicit feature observations that can reveal the underlying label correlation of rare label classification. To fully explore the correlation between labels and partial features, this paper proposes a Part-Aware Correlation Network (PACNet) based on Partial Representation (PR) and Semantic Covariance Matrix (SCM). Specifically, we develop a partial representing module of an object that eliminates object-independent information and allows the model to focus on more distinctive parts. Furthermore, a semantic covariance measure function is redefined as a way to learn the semantic relationships of partial representations and to compute the partial similarity between the query sample and the support set. Experiments on three benchmark datasets consistently show that the proposed method outperforms the state-of-the-art counterparts,e.g., on the PartImageNet dataset, the performance gains of up to 12% and 5.9% are observed for the 5-way 1-shot and 5-way 5-shot settings, respectively. Ruiheng Zhang 0001, Jinyu Tan, Zhe Cao 0001, Lixin Xu 0001, Lingyu Si, Fuchun Sun 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Few-Shot Infrared Image Classification with Partial Concept Feature
Jinyu Tan, Ruiheng Zhang 0001, Qi Zhang 0004, Zhe Cao 0001, Lixin Xu 0001 |
PRCV (4) | 4 |