VLDB 2026 Research / reviewers in the wild / expert
Xinzi Cao
dblp:330/7405
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0001-7966-9724ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Segmentation and scene understanding · 29% Image recognition and object detection · 21% 3D vision · 14% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding › category discovery
generalized category discovery |
1.6 | 2 | 2025 | ALLGCD: Leveraging All Unlabeled Data for Generalized Category Discovery · ICCV 2025 Solving the Catastrophic Forgetting Problem in Generalized Category Discovery · CVPR 2024 |
Machine learning › Representation and self-supervised learning › pre-training › visual pre-training
event camera data pre-training |
0.9 | 1 | 2025 | Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion · ICCV 2025 |
Computer vision › 3D vision › event-based vision
event camera data processing |
0.9 | 1 | 2025 | Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion · ICCV 2025 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.8 | 1 | 2024 | Solving the Catastrophic Forgetting Problem in Generalized Category Discovery · CVPR 2024 |
Computer vision › Image recognition and object detection
image classification |
0.7 | 1 | 2023 | LocLoc: Low-level Cues and Local-area Guides for Weakly Supervised Object Localization · ACM Multimedia 2023 |
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization |
0.7 | 1 | 2023 | LocLoc: Low-level Cues and Local-area Guides for Weakly Supervised Object Localization · ACM Multimedia 2023 |
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification |
0.7 | 1 | 2023 | LocLoc: Low-level Cues and Local-area Guides for Weakly Supervised Object Localization · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
self-supervised pretraining · 0.9prompt fusion · 0.9local entropy regularization · 0.8kullback-leibler divergence · 0.8contrastive learning · 0.8transformer · 0.7graph cuts · 0.7grabcut · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Margin-Aware Prototype Debiasing for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is a challenging task that aims to identify both seen and novel categories in unlabeled data. We argue that a clear margin between seen and novel class representations is essential for accurate recognition. However, existing methods often ignore this margin, mapping representations to prototypes without enforcing separation between seen and novel classes. This leads to a bias where seen samples are misclassified as novel. To address this issue, we propose DebiasGCD, a debiasing framework that enhances prototype separation through margin-aware learning. Unlike prior work that relies on static prototype learning and overlooks fine-grained representations, our method introduces Dynamic Prototype Debiasing (DPD) and Spatial-Aware Representation Distillation (SARD) to mitigate this bias. First, DPD dynamically enforces inter-prototype margins, improving class-specific feature learning and prototype discrimination. Meanwhile, SARD promotes local representation of spatial learning, supporting DPD to capture subtle details that further refine class-specific features. By synergizing these components, DebiasGCD significantly improves prototype discriminability, generating more reliable predictions for seen classes. Extensive experiments demonstrate that our approach effectively mitigates pseudo-labeling bias across datasets, especially on fine-grained ones, achieving +8.3% and +9.6% improvements on the ‘All’ classes in CUB and Stanford Cars, respectively. Xinzi Cao, Feidiao Yang, Xiawu Zheng, Quanmin Liang, Yutong Lu, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | ALLGCD: Leveraging All Unlabeled Data for Generalized Category Discovery
Xinzi Cao, Ke Chen 0004, Feidiao Yang, Xiawu Zheng, Yonghong Tian 0001, Yutong Lu |
ICCV | 1 |
| 2025 | Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion
Quanmin Liang, Shuai Liu 0009, Xinzi Cao, Jinyi Lu, Feidiao Yang, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001 |
ICCV | 4 |
| 2024 | Solving the Catastrophic Forgetting Problem in Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) aims to identify a mix of known and novel categories within unlabeled data sets, providing a more realistic setting for image recognition. Essentially, GCD needs to remember existing patterns thoroughly to recognize novel categories. Recent state-of-the-art method SimGCD transfers the knowledge from known-class data to the learning of novel classes through debiased learning. However, some patterns are catastrophically forgot during adaptation and thus lead to poor performance in novel categories classification. To address this issue, we propose a novel learning approach, LegoGCD, which is seamlessly integrated into previous methods to enhance the discrimination of novel classes while maintaining performance on previously encountered known classes. Specifically, we design two types of techniques termed as Local Entropy Regularization (LER) and Dual-views Kullback-Leibler divergence constraint (DKL). The LER optimizes the distribution of potential known class samples in unlabeled data, thus ensuring the preservation of knowledge related to known categories while learning novel classes. Meanwhile, DKL introduces Kullback-Leibler divergence to encourage the model to produce a similar prediction distribution of two view samples from the same image. In this way, it successfully avoids mismatched prediction and generates more reliable potential known class samples simultaneously. Extensive experiments validate that the proposed LegoGCD effectively addresses the known category forgetting issue across all datasets, e.g., delivering a 7.74% and 2.51% accuracy boost on known and novel classes in CUB, respectively. Our code is available at: https://github.com/Cliffia123/LegoGCD. Xinzi Cao, Xiawu Zheng, Guanhong Wang, Weijiang Yu, Yunhang Shen, Ke Li 0015, Yutong Lu, Yonghong Tian 0001 |
CVPR | 1 |
| 2023 | LocLoc: Low-level Cues and Local-area Guides for Weakly Supervised Object LocalizationabstractWeakly Supervised Object Localization (WSOL) aims to localize objects using only image-level labels while ensuring competitive classification performance. However, previous efforts have prioritized localization over classification accuracy in discriminative features, in which low-level information is neglected. We argue that low-level image representations, such as edges, color, texture, and motions are crucial for accurate detection. That is, using such information further achieves more refined localization, which can be used to promote classification accuracy. In this paper, we propose a unified framework that simultaneously improves localization and classification accuracy, termed as LocLoc (Low-level Cues and Local-area Guides). It leverages low-level image cues to explore global and local representations for accurate localization and classification. Specifically, we introduce a GrabCut-Enhanced Generator (GEG) to learn global semantic representations for localization based on graph cuts to enhance low-level information based on long-range dependencies captured by the transformer. We further design a Local Feature Digging Module (LFDM) that utilizes low-level cues to guide the learning route of local feature representations for accurate classification. Extensive experiments demonstrate the effectiveness of LocLoc with 84.4%(↑5.2%) Top-1 Loc., 85.8% Top-1 Cls. on CUB-200-2011 and 57.6% (↑1.5%) Top-1 Loc., 78.6% Top-1Cls. on ILSVRC 2012, indicating that our method achieves competitive performance with a large margin compared to previous approaches. Code and models are available at https://github.com/Cliffia123/LocLoc. Xinzi Cao, Xiawu Zheng, Yunhang Shen, Ke Li 0015, Jie Chen 0001, Yutong Lu, Yonghong Tian 0001 |
ACM Multimedia | 1 |
| 2022 | Exploring Pixel Alignment on Shallow Feature for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to cover the entire target object only under the image-level supervision. Most WSOL methods are stuck in mining the CAMs (class activation maps) of deep semantic features for they only focus on limited discriminative regions playing key role in classification. Recently, a new paradigm has emerged by localizing objects using the low-level feature through two stages. Existing two-stages methods usually train a classification network first to yield CAMs as pseudo labels to guide the learning of segment network, yet it does not consider the activations with more background noise or less discriminative area. In this paper, we propose a Pixel Alignment strategy to refine the object localization by improving the shallow-feature based CAMs generator with the joint supervision of pseudo-label mask, classification evaluation, and absolution size constraint on the activation map. More specifically, we utilize the class-specific pixel gradient to achieve a robust activation pseudo mask to background noise, which further supervises the activation generator with confident foreground and background regions. We also adapt a post-processing to excavate the target region in the conflict area (i.e., the non-overlap area of CAMs and the activations). Extensive experiments on CUB-2002011 and ILSVRC datasets indicate that our method outperforms the state-of-the-art among the two-stage works. Xinzi Cao, Meng Yang 0001, Guoying Sun |
IJCNN | 1 |