EDBT 2026 Demo / reviewers in the wild / expert
Haoxing Chen
dblp:168/5619
· DBLP profile ↗
22ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0001-6637-8741ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
Yi Xin 0003, Jianjiang Yang, Yuntao Du 0001, Haoxing Chen, Kangrui Cen, Yangfan He, Yuewen Cao, Junjun He, Xiaokang Yang 0001, Guangtao Zhai, Ming-Hsuan Yang 0001, Xiaohong Liu 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images DetectionabstractThe development of text-to-image generative models has enabled the creation of images so realistic that distinguishing between AI-generated images and real photos is becoming a challenge. This progress offers new possibilities but also raises concerns over privacy, authenticity, and security. Detecting AI-generated images is crucial to prevent misuse. To assess the generalizability and robustness of AI-generated image detection, we present a large-scale dataset, referred to as WildFake. This dataset features cutting-edge image generators, a wide variety of generator categories, and generators for various applications, organized in a hierarchical framework. WildFake collects fake images from the open-source community, enriching its diversity with a broad range of image classes and image styles. Its design significantly improves the effectiveness of detection algorithms, making it a valuable resource for enhancing AI-generated image detection in practical applications. Our evaluations offer insights into the performance of generative models at various levels, showcasing WildFake's unique hierarchical structure's benefits. Yan Hong 0001, Jianming Feng, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003 |
AAAI | 3 |
| 2025 | Efficient Transfer Learning for Video-language Foundation ModelsabstractPre-trained vision-language models provide a robust foundation for efficient transfer learning across various downstream tasks. In the field of video action recognition, mainstream approaches often introduce additional modules to capture temporal information. Although the additional modules increase the capacity of model, enabling it to better capture video-specific inductive biases, existing methods typically introduce a substantial number of new parameters and are prone to catastrophic forgetting of previously acquired generalizable knowledge. In this paper, we propose a parameter-efficient Multi-modal Spatio-Temporal Adapter (MSTA) to enhance the alignment between textual and visual representations, achieving a balance between generalizable knowledge and task-specific adaptation. Furthermore, to mitigate over-fitting and enhance generalizability, we introduce a spatio-temporal description-guided consistency constraint. This constraint involves providing template inputs (e.g., "a video of {cls}") to the trainable language branch and LLM-generated spatio-temporal descriptions to the pre-trained language branch, enforcing output consistency between the branches. This approach reduces overfitting to downstream tasks and enhances the distinguishability of the trainable branch within the spatio-temporal semantic space. We evaluate the effectiveness of our approach across four tasks: zero-shot transfer, few-shot learning, base-to-novel generalization, and fully-supervised learning. Compared to many state-of-the-art methods, our MSTA achieves outstanding performance across all evaluations, while using only 2-7% of the trainable parameters in the original model. Haoxing Chen, Zizheng Huang, Yan Hong 0001, Yanshuo Wang, Zhongcai Lyu, Zhuoer Xu, Jun Lan 0001, Zhangxuan Gu |
CVPR | 1 |
| 2025 | Dynamic Model-Bank Test-Time Adaptation for Automatic Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) based on deep learning has achieved impressive progress in recent years.However, the performance of ASR foundation model often degrades significantly on out-of-domain data due to real-world domain shifts.Test-Time Adaptation (TTA) methods aim to mitigate this issue by adapting models during inference without access to source data.Despite recent progress, existing ASR TTA methods often struggle with instability under continual and long-term distribution shifts.To alleviate the risk of performance collapse due to error accumulation, we propose Dynamic Model-bank Single-Utterance Test-time Adaptation (DM-SUTA), a sustainable continual TTA framework based on adaptive ASR model ensembling.DMSUTA maintains a dynamic model bank, from which a subset of checkpoints is selected for each test sample based on confidence and uncertainty criteria.To preserve both model plasticity and long-term stability, DMSUTA actively manages the bank by filtering out potentially collapsed models.This design allows DMSUTA to continually adapt to evolving domain shifts in ASR test-time scenarios.Experiments on diverse, continuously shifting ASR TTA benchmarks show that DM-SUTA consistently outperforms existing continual TTA baselines, demonstrating superior robustness to domain shifts in ASR. Yanshuo Wang, Yanghao Zhou, Yukang Lin, Haoxing Chen |
EMNLP | 4 |
| 2025 | Stochastic Layer-Wise Shuffle for Improving Vision Mamba TrainingabstractRecent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies and their potential are still not sufficiently explored. In this paper, we investigate strategies for Vim and propose Stochastic Layer-Wise Shuffle (SLWS), a novel regularization method that can effectively improve the Vim training. Without architectural modifications, this approach enables the non-hierarchical Vim to get leading performance on ImageNet-1K compared with the similar type counterparts. Our method operates through four simple steps per layer: probability allocation to assign layer-dependent shuffle rates, operation sampling via Bernoulli trials, sequence shuffling of input tokens, and order restoration of outputs. SLWS distinguishes itself through three principles: \textit{(1) Plug-and-play:} No architectural modifications are needed, and it is deactivated during inference. \textit{(2) Simple but effective:} The four-step process introduces only random permutations and negligible overhead. \textit{(3) Intuitive design:} Shuffling probabilities grow linearly with layer depth, aligning with the hierarchical semantic abstraction in vision models. Our work underscores the importance of tailored training strategies for Vim models and provides a helpful way to explore their scalability. Code and models are available at https://github.com/huangzizheng01/ShuffleMamba Zizheng Huang, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Limin Wang 0002 |
ICML | 2 |
| 2025 | Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsabstractProgress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake. Yikun Ji, Yan Hong 0001, Jiahui Zhan, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Liqing Zhang 0001, Jianfu Zhang 0003 |
ACM Multimedia | 4 |
| 2025 | InterAnimate: Taming Region-Aware Diffusion Model for Realistic Human Interaction Animation
Yukang Lin, Yan Hong 0001, Zunnan Xu, Xindi Li, Chuanbiao Song, Ronghui Li, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Xiu Li 0001 |
ACM Multimedia | 8 |
| 2025 | Conditional Prototype Rectification Prompt LearningabstractPre-trained large-scale vision-language models (VLMs) have acquired profound understanding of general visual concepts. Recent advancements in efficient transfer learning (ETL) have shown remarkable success in fine-tuning VLMs within the scenario of limited data, introducing only a few parameters to harness task-specific insights from VLMs. Despite significant progress, current leading ETL methods tend to overfit the narrow distributions of base classes seen during training and encounter two primary challenges: (i) only utilizing uni-modal information to modeling task-specific knowledge; and (ii) using costly and time-consuming methods to supplement knowledge. To address these issues, we propose a Conditional Prototype Rectification Prompt Learning (CPR) method to correct the bias of the base examples and augment limited data in an effective way. Specifically, we alleviate over-fitting on base classes from two aspects. First, each input image acquires knowledge from both textual and visual prototypes and then generates sample-conditional text tokens. Second, we extract utilizable knowledge from unlabeled data to further refine the prototypes. These two strategies mitigate biases that stem from base classes, yielding a more effective classifier. Extensive experiments on 11 benchmark datasets show that our CPR achieves state-of-the-art performance on few-shot classification, base-to-new generalization, and cross-dataset generalization tasks. Our code is available at https://github.com/chenhaoxing/CPR. Haoxing Chen, Zizheng Huang, Yan Hong 0001, Zhuoer Xu, Zhangxuan Gu, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | ComFusion: Enhancing Personalized Generation by Instance-Scene Compositing and Fusion
Yan Hong 0001, Yuxuan Duan, Bo Zhang 0075, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003 |
ECCV (44) | 4 |
| 2024 | Segment Anything Model Meets Image HarmonizationabstractImage harmonization is a crucial technique in image composition that aims to seamlessly match the background by adjusting the foreground of composite images. Current methods adopt either global-level or pixel-level feature matching. Global-level feature matching ignores the proximity prior, treating foreground and background as separate entities. On the other hand, pixel-level feature matching loses contextual information. Therefore, it is necessary to use the information from semantic maps that describe different objects to guide harmonization. In this paper, we propose Semantic-guided Region-aware Instance Normalization (SRIN) that can utilize the semantic segmentation maps output by a pre-trained Segment Anything Model (SAM) to guide the visual consistency learning of foreground and background features. Abundant experiments demonstrate the superiority of our method for image harmonization over state-of-the-art methods. Haoxing Chen, Zhangxuan Gu, Zhuoer Xu, Jun Lan 0001, Huaxiong Li |
ICASSP | 1 |
| 2024 | Diffusioninst: Diffusion Model for Instance SegmentationabstractDiffusion frameworks have achieved comparable performance with previous state-of-the-art image generation models. This paper proposes DiffusionInst, a novel framework representing instances as vectors and formulates instance segmentation as a noise-to-vector denoising process. The model is trained to reverse the noisy groundtruth mask without any inductive bias from RPN. It takes a randomly generated vector as input and outputs mask with multi-step denoising during inference. Extensive experimental results on COCO and LVIS show that DiffusionInst achieves competitive performance. Our code is available at https://github.com/chenhaoxing/DiffusionInst. Zhangxuan Gu, Haoxing Chen, Zhuoer Xu |
ICASSP | 2 |
| 2024 | Learning latent disentangled embeddings and graphs for multi-view clustering
Chao Zhang 0078, Haoxing Chen, Huaxiong Li, Chunlin Chen 0001 |
Pattern Recognit. | 2 |
| 2023 | Mobile User Interface Element Detection Via Adaptively Prompt TuningabstractRecent object detection approaches rely on pretrained vision-language models for image-text alignment. However, they fail to detect the Mobile User Interface (MUI) element since it contains additional OCR information, which describes its content and function but is often ignored. In this paper, we develop a new MUI element detection dataset named MUI-zh and propose an Adaptively Prompt Tuning (APT) module to take advantage of discriminating OCR information. APT is a lightweight and effective module to jointly optimize category prompts across different modalities. For every element, APT uniformly encodes its visual features and OCR descriptions to dynamically adjust the representation of frozen category prompts. We evaluate the effectiveness of our plug-and-play APT upon several existing CLIP-based detectors for both standard and open-vocabulary MUI element detection. Extensive experiments show that our method achieves considerable improvements on two datasets. The datasets is available at github.com/antmachineintelligence/MUI-zh. Zhangxuan Gu, Zhuoer Xu, Haoxing Chen, Jun Lan 0001, Changhua Meng, Weiqiang Wang 0002 |
CVPR | 3 |
| 2023 | Model-Aware Contrastive Learning: Towards Escaping the DilemmasabstractContrastive learning (CL) continuously achieves significant breakthroughs across multiple domains. However, the most common InfoNCE-based methods suffer from some dilemmas, such as uniformity-tolerance dilemma (UTD) and gradient reduction, both of which are related to a $\mathcal{P}_{ij}$ term. It has been identified that UTD can lead to unexpected performance degradation. We argue that the fixity of temperature is to blame for UTD. To tackle this challenge, we enrich the CL loss family by presenting a Model-Aware Contrastive Learning (MACL) strategy, whose temperature is adaptive to the magnitude of alignment that reflects the basic confidence of the instance discrimination task, then enables CL loss to adjust the penalty strength for hard negatives adaptively. Regarding another dilemma, the gradient reduction issue, we derive the limits of an involved gradient scaling factor, which allows us to explain from a unified perspective why some recent approaches are effective with fewer negative samples, and summarily present a gradient reweighting to escape this dilemma. Extensive remarkable empirical results in vision, sentence, and graph modality validate our approach’s general improvement for representation learning and downstream tasks. Zizheng Huang, Haoxing Chen, Ziqi Wen, Chao Zhang 0078, Huaxiong Li, Bo Wang 0027, Chunlin Chen 0001 |
ICML | 2 |
| 2023 | Hierarchical Dynamic Image HarmonizationabstractImage harmonization is a critical task in computer vision, which aims to adjust the foreground to make it compatible with the background. Recent works mainly focus on using global transformations (i.e., normalization and color curve rendering) to achieve visual consistency. However, these models ignore local visual consistency and their huge model sizes limit their harmonization ability on edge devices. In this paper, we propose a hierarchical dynamic network (HDNet) to adapt features from local to global view for better feature transformation in efficient image harmonization. Inspired by the success of various dynamic models, local dynamic (LD) module and mask-aware global dynamic (MGD) module are proposed in this paper. Specifically, LD matches local representations between the foreground and background regions based on semantic similarities, then adaptively adjust every foreground local representation according to the appearance of its K-nearest neighbor background regions. In this way, LD can produce more realistic images at a more fine-grained level, and simultaneously enjoy the characteristic of semantic alignment. The MGD effectively applies distinct convolution to the foreground and background, learning the representations of foreground and background regions as well as their correlations to the global harmonization, facilitating local visual consistency for the images much more efficiently. Experimental results demonstrate that the proposed HDNet significantly reduces the total model parameters by more than 80% compared to previous methods, while still attaining state-of-the-art performance on the popular iHarmony4 dataset. Additionally, we introduced a lightweight version of HDNet, i.e., HDNet-lite, which has only 0.65MB parameters, yet it still achieve competitive performance. Our code is avaliable at https://github.com/chenhaoxing/HDNet. Haoxing Chen, Zhangxuan Gu, Jun Lan 0001, Changhua Meng, Weiqiang Wang 0002, Huaxiong Li |
ACM Multimedia | 1 |
| 2023 | DiffUTE: Universal Text Editing Diffusion ModelabstractDiffusion model based language-guided image editing has achieved great success recently. However, existing state-of-the-art diffusion models struggle with rendering correct text and text style during generation. To tackle this problem, we propose a universal self-supervised text editing diffusion model (DiffUTE), which aims to replace or modify words in the source image with another one while maintaining its realistic appearance. Specifically, we build our model on a diffusion model and carefully modify the network structure to enable the model for drawing multilingual characters with the help of glyph and position information. Moreover, we design a self-supervised learning framework to leverage large amounts of web data to improve the representation ability of the model. Experimental results show that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. Our code will be avaliable in \url{https://github.com/chenhaoxing/DiffUTE}. Haoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan 0001, Xing Zheng, Changhua Meng, Huijia Zhu, Weiqiang Wang 0002 |
NeurIPS | 1 |
| 2023 | Sparse spatial transformers for few-shot learning
Haoxing Chen, Huaxiong Li, Chunlin Chen 0001 |
Sci. China Inf. Sci. | 1 |
| 2022 | Multi-level Metric Learning for Few-Shot Image Recognition
Haoxing Chen, Huaxiong Li, Chunlin Chen 0001 |
ICANN (1) | 1 |
| 2022 | Multi-Scale Adaptive Task Attention Network for Few-Shot LearningabstractFew-shot learning has aroused considerable interest in recent years, which aims to recognize unseen categories by using a few labeled samples. In various few-shot methods, pixel-level metric-learning based methods have achieved promising performance. However, most of these methods deal with each category in the support set independently, which may be insufficient to measure the relations among features, especially in a specific task. Besides, the coexistence of dominant objects at different scales may degrade the performance of these methods. To address these issues, a novel Multi-Scale Adaptive Task Attention Network, MATANet for short, is proposed for few-shot learning. In MATANet, a multi-scale feature generator is first constructed to extract the image features at different scales. Then, an adaptive task attention module is built to select the most important local representations among the entire task. Finally, a similarity-to-class module is adapted to measure the similarities between query and support set. Extensive experiments on popular benchmarks show the effectiveness of the proposed MATANet compared with state-of-the-art methods. Our source code is available at: https://github.com/chenhaoxing/MATANet. Haoxing Chen, Huaxiong Li, Chunlin Chen 0001 |
ICPR | 1 |
| 2022 | Transductive Aesthetic Preference Propagation for Personalized Image Aesthetics AssessmentabstractPersonalized image aesthetics assessment (PIAA) aims at capturing individual aesthetic preference. Fine-tuning on personalized data has been proven to be effective in PIAA task. However, a fixed fine-tuning strategy may cause under/over-fitting on limited personal data and it also brings additional training cost. To alleviate these issues, we employ a meta learning-based Transductive Aesthetic Preference Propagation (TAPP-PIAA) algorithm under regression manner to substitute the fine-tuning strategy. Specifically, each user's data is regarded as a meta-task and spilt into support and query set. Then, we extract deep aesthetic features with a pre-trained generic image aesthetic assessment (GIAA) model. Next, we treat image features as graph nodes and their similarities as edge weights to construct an undirected nearest neighbor graph for inference. Instead of fine-tuning on support set, TAPP-PIAA propagates aesthetic preference from support to query set with a predefined propagation formula. Finally, to learn a generalizable aesthetic representation for various users, we optimize our TAPP-PIAA across different users with meta-learning framework. Experimental results indicate that our TAPP-PIAA can surpass the state-of-the-art methods on benchmark databases. Yuzhe Yang 0001, Huaxiong Li, Haoxing Chen, Liwu Xu, Leida Li, Yandong Guo |
ACM Multimedia | 4 |
| 2022 | Shaping Visual Representations With Attributes for Few-Shot RecognitionabstractFew-shot recognition aims to recognize novel categories under low-data regimes. Some recent few-shot recognition methods introduce auxiliary semantic modality, i.e., category attribute information, into representation learning, which enhances the feature discrimination and improves the recognition performance. Most of these existing methods only consider the attribute information of support set while ignoring the query set, resulting in a potential loss of performance. In this letter, we propose a novel attribute-shaped learning (ASL) framework, which can jointly perform query attributes generation and discriminative visual representation learning for few-shot recognition. Specifically, a visual-attribute predictor (VAP) is constructed to predict the attributes of queries. By leveraging the attributes information, an attribute-visual attention module (AVAM) is designed, which can adaptively utilize attributes and visual representations to learn more discriminative features. Under the guidance of attribute modality, our method can learn enhanced semantic-aware representation for classification. Experiments demonstrate that our method can achieve competitive results on CUB and SUN benchmarks. Our source code is available at:https://github.com/chenhaoxing/ASL. Haoxing Chen, Huaxiong Li, Chunlin Chen 0001 |
IEEE Signal Process. Lett. | 1 |
| 2021 | Local Mutual Metric Network for Few-Shot Image Classification
Huaxiong Li, Haoxing Chen, Chunlin Chen 0001 |
PRCV (1) | 3 |