VLDB 2026 Research / reviewers in the wild / expert
Chenxing Gao
dblp:342/2723
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion ModelabstractA recent endeavor in one class of video anomaly detection is to leverage diffusion models and posit the task as a generation problem, where the diffusion model is trained to recover normal patterns exclusively, thus reporting abnormal patterns as outliers. Yet, existing attempts neglect the various formations of anomaly and predict normal samples at the feature level regardless that abnormal objects in surveillance videos are often relatively small. To address this, a novel patch-based diffusion model is proposed, specifically engineered to capture fine-grained local information. We further observe that anomalies in videos manifest themselves as deviations in both appearance and motion. Therefore, we argue that a comprehensive solution must consider both of these aspects simultaneously to achieve accurate frame prediction. To address this, we introduce innovative motion and appearance conditions that are seamlessly integrated into our patch diffusion model. These conditions are designed to guide the model in generating coherent and contextually appropriate predictions for both semantic content and motion relations. Experimental results on four challenging video anomaly detection datasets empirically substantiate the efficacy of our proposed approach, demonstrating that it consistently outperforms most existing methods in detecting abnormal behaviors. Hang Zhou 0010, Jiale Cai, Yuteng Ye, Yonghui Feng, Chenxing Gao, Junqing Yu, Zikai Song, Wei Yang 0034 |
AAAI | 5 |
| 2025 | Causal Feature Supervision Decoupling: A Novel Method for Clothes-Changing Person Re-identification AlgorithmabstractClothes-Changing person re-identification algorithm (Re-ID) is the task of retrieving the query person in the case of the change of pedestrian clothing. Changes in pedestrian clothing lead to an offset in clothing features, resulting in a decrease in identification performance. Simply removing clothing may lead to the loss of contour information. Furthermore, algorithms based on feature decoupling cannot guarantee the accuracy of the positional relationship between clothing features and other features due to the lack of groundtruth. To address these issues, we propose a novel clothes-changing person re-identification algorithm based on causal feature supervision decoupling. Utilizing a multi-scale feature fusion module to extract fine-grained features of clothing and add supervised information labels. This enables the dual-branch network to separately approach overall and clothing features, promoting the extraction of effective identity information by the causal decoupling module, and obtaining unbiased estimations of pedestrians. Experimental results show that the proposed algorithm achieves the highest mAP and Top-1 accuracy on the LTCC and PRCC datasets. The source code is available at https://github.com/zhihu250/CISupNet. Wenxin Hu, Caidan Zhao, Chenxing Gao, Zhiqiang Wu 0001 |
ICASSP | 3 |
| 2024 | Attacking Transformers with Feature Diversity Adversarial PerturbationabstractUnderstanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturbations, is crucial for addressing challenges in its real-world applications. Existing ViT adversarial attackers rely on labels to calculate the gradient for perturbation, and exhibit low transferability to other structures and tasks. In this paper, we present a label-free white-box attack approach for ViT-based models that exhibits strong transferability to various black-box models, including most ViT variants, CNNs, and MLPs, even for models developed for other modalities. Our inspiration comes from the feature collapse phenomenon in ViTs, where the critical attention mechanism overly depends on the low-frequency component of features, causing the features in middle-to-end layers to become increasingly similar and eventually collapse. We propose the feature diversity attacker to naturally accelerate this process and achieve remarkable performance and transferability. Chenxing Gao, Hang Zhou 0010, Junqing Yu, Yuteng Ye, Jiale Cai, Junle Wang, Wei Yang 0034 |
AAAI | 1 |
| 2024 | Progressive Text-to-Image Diffusion with Soft Latent DirectionabstractIn spite of the rapidly evolving landscape of text-to-image generation, the synthesis and manipulation of multiple entities while adhering to specific relational constraints pose enduring challenges. This paper introduces an innovative progressive synthesis and editing operation that systematically incorporates entities into the target image, ensuring their adherence to spatial and relational constraints at each sequential step. Our key insight stems from the observation that while a pre-trained text-to-image diffusion model adeptly handles one or two entities, it often falters when dealing with a greater number. To address this limitation, we propose harnessing the capabilities of a Large Language Model (LLM) to decompose intricate and protracted text descriptions into coherent directives adhering to stringent formats. To facilitate the execution of directives involving distinct semantic operations—namely insertion, editing, and erasing—we formulate the Stimulus, Response, and Fusion (SRF) framework. Within this framework, latent regions are gently stimulated in alignment with each operation, followed by the fusion of the responsive latent components to achieve cohesive entity manipulation. Our proposed framework yields notable advancements in object synthesis, particularly when confronted with intricate and lengthy textual inputs. Consequently, it establishes a new benchmark for text-to-image generation tasks, further elevating the field's performance standards. Yuteng Ye, Jiale Cai, Hang Zhou 0010, Guanwen Li, Youjia Zhang, Zikai Song, Chenxing Gao, Junqing Yu, Wei Yang 0034 |
AAAI | 7 |
| 2024 | Dynamic Feature Pruning and Consolidation for Occluded Person Re-identificationabstractOccluded person re-identification (ReID) is a challenging problem due to contamination from occluders. Existing approaches address the issue with prior knowledge cues, such as human body key points and semantic segmentations, which easily fail in the presence of heavy occlusion and other humans as occluders. In this paper, we propose a feature pruning and consolidation (FPC) framework to circumvent explicit human structure parsing. The framework mainly consists of a sparse encoder, a multi-view feature mathcing module, and a feature consolidation decoder. Specifically, the sparse encoder drops less important image tokens, mostly related to background noise and occluders, solely based on correlation within the class token attention. Subsequently, the matching stage relies on the preserved tokens produced by the sparse encoder to identify k-nearest neighbors in the gallery by measuring the image and patch-level combined similarity. Finally, we use the feature consolidation module to compensate pruned features using identified neighbors for recovering essential information while disregarding disturbance from noise and occlusion. Experimental results demonstrate the effectiveness of our proposed framework on occluded, partial, and holistic Re-ID datasets. In particular, our method outperforms state-of-the-art results by at least 8.6% mAP and 6.0% Rank-1 accuracy on the challenging Occluded-Duke dataset. Yuteng Ye, Hang Zhou 0010, Jiale Cai, Chenxing Gao, Youjia Zhang, Junle Wang, Qiang Hu 0003, Junqing Yu, Wei Yang 0034 |
AAAI | 4 |
| 2024 | Video Anomaly Detection Framework Based on Motion ConsistencyabstractMost methods rely on unsupervised learning due to the limited availability of anomaly data. However, most of the current unsupervised learning methods are based on deep self-encoders, which do not pay enough attention to the consistency of the motion process. Therefore, we propose a video anomaly detection framework based on motion consistency (VADMC).The framework uses a CVAE network as a generator to generate predicted frames. In order to increase the reconstruction error of the CVAE network, we embed a memory module in the optical flow coding features, which is used to memorize the feature distribution of the normal patterns. The reconstruction error is increased by perturbing the a priori distribution, thus increasing the reconstruction error. A discriminator is used to discriminate the generated optical flow maps to ensure the consistency of the forward and backward motions of the normal samples. We conducted experiments on three public datasets to demonstrate the effectiveness of the VADMC framework. The accuracy on the UCSD PED2, CHUK Avenue, and Shanghai Tech datasets reached 97.2%, 76.3%, and 76.2%, respectively. Compared with previous state-of-the-art methods, our method shows competitive results. Caidan Zhao, Chenxing Gao |
CSCWD | 3 |
| 2023 | Lightweight Image Dehazing Algorithm Based on Detail Feature EnhancementabstractHaze can reduce the visibility of the captured image, making it hard to accurately distinguish the details of each object in the captured image scene. Aiming at the problem of detail loss in existing dehazing models, this paper proposes a lightweight end-to-end image dehazing framework called DFE-GAN (Detail Feature Enhancement-GAN). The missing detail contours in the haze image can be predicted by employing a densely connected detail feature prediction network. Supplemented with a patch discriminator and an improved loss function, the restoration of details in the dehazing image is enhanced to improve image quality. We apply inverse residual modules to extract and fuse multi-scale features from images, which can ensure the real-time processing capability of the model. Compared with previous state-of-the-art approaches, solid experimental results on various benchmark datasets validate the robustness and effectiveness of our model. Chenxing Gao, Lingjun Chen, Caidan Zhao, Xiangyu Huang, Zhiqiang Wu 0001 |
CSCWD | 1 |
| 2023 | Synthetic Pseudo Anomalies for Unsupervised Video Anomaly Detection: A Simple Yet Efficient Framework Based on Masked AutoencoderabstractDue to the limited availability of anomalous samples for training, video anomaly detection is commonly viewed as a one-class classification problem. Many prevalent methods investigate the reconstruction difference produced by AutoEncoders (AEs) under the assumption that the AEs would reconstruct the normal data well while reconstructing anomalies poorly. However, even with only normal data training, the AEs often reconstruct anomalies well, which depletes their anomaly detection performance. To alleviate this issue, we propose a simple yet efficient framework for video anomaly detection. The pseudo anomaly samples are introduced, which are synthesized from only normal data by embedding random mask tokens without extra data processing. We also propose a normalcy consistency training strategy that encourages the AEs to better learn the regular knowledge from normal and corresponding pseudo anomaly data. This way, the AEs learn more distinct reconstruction boundaries between normal and abnormal data, resulting in superior anomaly discrimination capability. Experimental results demonstrate the effectiveness of the proposed method. Xiangyu Huang, Caidan Zhao, Chenxing Gao, Lvdong Chen, Zhiqiang Wu 0001 |
ICASSP | 3 |
| 2023 | Multi-Level Memory-Augmented Appearance-Motion Correspondence Framework for Video Anomaly DetectionabstractFrame prediction based on AutoEncoder plays a significant role in unsupervised video anomaly detection. Ideally, the models trained on the normal data could generate larger prediction errors of anomalies. However, the correlation between appearance and motion information is underutilized, which makes the models lack an understanding of normal patterns. Moreover, the models do not work well due to the uncontrollable generalizability of deep AutoEncoder. To tackle these problems, we propose a multi-level memory-augmented appearance-motion correspondence framework. The latent correspondence between appearance and motion is explored via appearance-motion semantics alignment and semantics replacement training. Besides, we also introduce a Memory-Guided Suppression Module, which utilizes the difference from normal prototype features to suppress the reconstruction capacity caused by skip-connection, achieving the tradeoff between the good reconstruction of normal data and the poor reconstruction of abnormal data. Experimental results show that our framework outperforms the state-of-the-art methods, achieving AUCs of 99.6%, 93.8%, and 76.3% on UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. Xiangyu Huang, Caidan Zhao, Chenxing Gao, Zhiqiang Wu 0001 |
ICME | 4 |