EDBT 2026 Demo / reviewers in the wild / expert
Qidong Huang
dblp:289/8148
· DBLP profile ↗
18ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-2702-8516ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring a novel data- and parameter-efficient fine-tuning method for robust face recognition
Yin Lin, Qidong Huang, Yang Cao 0010, Zengfu Wang |
Pattern Recognit. | 3 |
| 2025 | MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World ConversationabstractRecent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, their abilities to memorize, recall, and reason in sustained interactions within real-world scenarios remain underexplored. This paper introduces MMRC, a Multi-Modal Real-world Conversation benchmark for evaluating six core open-ended abilities of MLLMs: information extraction, multi-turn reasoning, information update, image management, memory recall, and answer refusal. With data collected from real-world scenarios, MMRC comprises 5,120 conversations and 28,720 corresponding manually labeled questions, posing a significant challenge to existing MLLMs. Evaluations on 20 MLLMs in MMRC indicate an accuracy drop during open-ended interactions. We identify four common failure patterns: long-term memory degradation, inadequacies in updating factual knowledge, accumulated assumption of error propagation, and reluctance to “say no.” To mitigate these issues, we propose a simple yet effective NOTE-TAKING strategy, which can record key information from the conversation and remind the model during its responses, enhancing conversational capabilities. Experiments across six MLLMs demonstrate significant performance improvements. Haochen Xue, Yexin Liu, Qidong Huang, Yulong Li 0002, Zhongxing Xu, Chong Zhang 0006, Yutong Xie 0001, Muhammad Imran Razzak, ZongYuan Ge, Jionglong Su, Junjun He, Yu Qiao 0001 |
ACL (1) | 5 |
| 2025 | Conical Visual Concentration for Efficient Large Vision-Language ModelsabstractIn large vision-language models (LVLMs), images serve as inputs that carry a wealth of information. As the idiom "A picture is worth a thousand words" implies, representing a single image in current LVLMs can require hundreds or even thousands of tokens, which thereby severely impacting the efficiency. Previous approaches have attempted to reduce the number of image tokens either before or within the early layers of LVLMs. However, these strategies inevitably result in the loss of crucial image information. To address this challenge, we conduct an empirical study revealing that all visual tokens are necessary for LVLMs in the shallow layers, and token redundancy progressively increases in the deeper layers. To this end, we propose PyramidDrop, a visual redundancy reduction strategy for LVLMs to boost their efficiency in both training and inference with neglectable performance loss. Specifically, we partition the LVLM into several stages and drop part of the image tokens at the end of each stage with a pre-defined ratio. The dropping is based on a lightweight similarity calculation with a negligible time overhead. Extensive experiments demonstrate that PyramidDrop can achieve over 40% training time reduction and 55% inference FLOPs acceleration on leading LVLMs like LLaVA-NeXT, maintaining comparable multi-modal performance. Besides, PyramidDrop can also serve as a plug-and-play strategy to accelerate inference in a free way, with better performance and lower inference cost than counterparts. Our code is available at https://github.com/Cooperx521/PyramidDrop. Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu, Pan Zhang 0001, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang 0003, Dahua Lin |
CVPR | 2 |
| 2025 | Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging ScenariosabstractFace recognition plays a crucial role in human life, prompting numerous excellent research efforts. However, face recognition in real-world applications presents various scenarios such as occluded, overexposed and near-infrared face recognition. Due to domain discrepancy and a lack of large-scale training data, effectively transferring pre-trained face recognition models to these scenarios has become a challenge. Recently, Parameter-Efficient Fine-Tuning (PEFT) methods have shown great potential in natural language processing tasks, but their effectiveness in computer vision tasks, especially in face recognition tasks, remains under-explored. In this paper, we propose a Data-Parameter-Efficient Fine-Tuning (DPEFT) approach for the face recognition tasks, encompassing two kinds of fine-tuning strategies. With these strategies, the DPEFT method requires only an additional 2.7% learnable parameters and 20% of the training data during the training phase to achieve competitive results. Moreover, by further integrating the concept of structural re-parameterization, our approach maintains the same model architecture and parameters as the pre-trained model during inference. Extensive experimental results on both holistic and occluded face datasets demonstrate that our approach achieves performance comparable to or better than the fully fine-tuning methods, and significantly lower training costs. Our DPEFT enables the pre-trained face recognition model to adapt efficiently and effectively to a variety of scenarios, indicating its potential in practical applications. Yin Lin, Ziyang Wu, Qidong Huang, Jinshui Hu, Zengfu Wang |
ICASSP | 3 |
| 2025 | Deciphering Cross-Modal Alignment in Large Vision-Language Models Via Modality Integration Rate
Qidong Huang, Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Yuhang Cao, Jiaqi Wang 0003, Weiming Zhang 0001, Nenghai Yu |
ICCV | 1 |
| 2025 | Light-a-Video: Training-Free Video Relighting via Progressive Light FusionabstractRecent advancements in image relighting models, driven by large-scale datasets and pre-trained diffusion models, have enabled the imposition of consistent lighting. However, video relighting still lags, primarily due to the excessive training costs and the scarcity of diverse, high-quality video relighting datasets. A simple application of image relighting models on a frame-by-frame basis leads to several issues: lighting source inconsistency and relighted appearance inconsistency, resulting in flickers in the generated videos. In this work, we propose Light-A-Video, a training-free approach to achieve temporally smooth video relighting. Adapted from image relighting models, Light-A-Video introduces two key techniques to enhance lighting consistency. First, we design a Consistent Light Attention (CLA) module, which enhances cross-frame interactions within the self-attention layers of the image relight model to stabilize the generation of the background lighting source. Second, leveraging the physical principle of light transport independence, we apply linear blending between the source video's appearance and the relighted appearance, using a Progressive Light Fusion (PLF) strategy to ensure smooth temporal transitions in illumination. Experiments show that Light-A-Video improves the temporal consistency of relighted video while maintaining the relighted image quality, ensuring coherent lighting transitions across frames. Project page: https://bujiazi.github.io/light-a-video.github.io/. Jiazi Bu, Pengyang Ling, Pan Zhang 0001, Qidong Huang, Jinsong Li 0001, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Anyi Rao, Jiaqi Wang 0003, Li Niu 0002 |
ICCV | 6 |
| 2025 | Multi-scale count-task guided feature enhancement face detection
Ziyang Wu, Yin Lin, Qidong Huang, Wengang Zhou 0001, Houqiang Li |
Multim. Syst. | 3 |
| 2024 | OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-AllocationabstractHallucination, posed as a pervasive challenge of multi-modal large language models (MLLMs), has significantly impeded their real-world usage that demands precise judgment. Existing methods mitigate this issue with either training with specific designed data or inferencing with external knowledge from other sources, incurring inevitable additional costs. In this paper, we present OPERA, a novel MLLM decoding method grounded in an Over-trust Penalty and a Retrospection-Allocation strategy, serving as a nearly free lunch to alleviate the hallucination issue without additional data, knowledge, or training. Our approach begins with an interesting observation that, most hallucinations are closely tied to the knowledge aggregation patterns manifested in the self-attention matrix, i.e., MLLMs tend to generate new tokens by focusing on a few summary tokens, but not all the previous tokens. Such partial overtrust inclination results in the neglecting of image tokens and describes the image content with hallucination. Based on the observation, OPERA introduces a penalty term on the model logits during the beam-search decoding to mitigate the over-trust issue, along with a rollback strategy that retrospects the presence of summary tokens in the previously generated tokens, and re-allocate the token selection if necessary. With extensive experiments, OPERA shows significant hallucination-mitigating performance on different MLLMs and metrics, proving its effectiveness and generality. Our code is at: https://github.com/shikiw/OPERA. Qidong Huang, Xiaoyi Dong, Pan Zhang 0001, Bin Wang 0065, Conghui He, Jiaqi Wang 0003, Dahua Lin, Weiming Zhang 0001, Nenghai Yu |
CVPR | 1 |
| 2024 | SimAC: A Simple Anti-Customization Method for Protecting Face Privacy Against Text-to-Image Synthesis of Diffusion ModelsabstractDespite the success of diffusion-based customization methods on visual content creation, increasing concerns have been raised about such techniques from both privacy and political perspectives. To tackle this issue, several anti-customization methods have been proposed in very recent months, predominantly grounded in adversarial attacks. Unfortunately, most of these methods adopt straight-forward designs, such as end-to-end optimization with afocus on adversarially maximizing the original training loss, thereby neglecting nuanced internal properties intrinsic to the diffusion model, and even leading to ineffective optimization in some diffusion time steps. In this paper, we strive to bridge this gap by undertaking a comprehensive exploration of these inherent properties, to boost the performance of current anti-customization approaches. Two aspects of properties are investigated: 1) We examine the relationship between time step selection and the model's perception in the frequency domain of images and find that lower time steps can give much more contributions to adversarial noises. This inspires us to propose an adaptive greedy search for optimal time steps that seamlessly integrates with existing anti-customization methods. 2) We scrutinize the roles of features at different layers during denoising and devise a sophisticated feature-based optimization framework for anti-customization. Experiments on facial benchmarks demonstrate that our approach significantly increases identity disruption, thereby protecting user privacy and copyright. Our code is available at:: http//github.com/somuchtome/SimAC. Zhentao Tan, Tianyi Wei, Qidong Huang |
CVPR | 5 |
| 2024 | PointCAT: Contrastive Adversarial Training for Robust Point Cloud RecognitionabstractNotwithstanding the prominent performance shown in various applications, point cloud recognition models have often suffered from natural corruptions and adversarial perturbations. In this paper, we delve into boosting the general robustness of point cloud recognition, proposing Point-Cloud Contrastive Adversarial Training (PointCAT). The main intuition of PointCAT is encouraging the target recognition model to narrow the decision gap between clean point clouds and corrupted point clouds by devising feature-level constraints rather than logit-level constraints. Specifically, we leverage a supervised contrastive loss to facilitate the alignment and the uniformity of hypersphere representations, and design a pair of centralizing losses with dynamic prototype guidance to prevent features from deviating outside their belonging category clusters. To generate more challenging corrupted point clouds, we adversarially train a noise generator concurrently with the recognition model from the scratch. This differs from previous adversarial training methods that utilized gradient-based attacks as the inner loop. Comprehensive experiments show that the proposed PointCAT outperforms the baseline methods, significantly enhancing the robustness of diverse point cloud recognition models under various corruptions, including isotropic point noises, the LiDAR simulated noises, random point dropping, and adversarial perturbations. Our code is available at: https://github.com/shikiw/PointCAT. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Hang Zhou 0007, Weiming Zhang 0001, Gang Hua 0001, Yueqiang Cheng, Nenghai Yu |
IEEE Trans. Image Process. | 1 |
| 2023 | Diversity-Aware Meta Visual PromptingabstractWe present Diversity-Aware Meta Visual Prompting (DAM-VP), an efficient and effective prompting method for transferring pre-trained models to downstream tasks with frozen backbone. A challenging issue in visual prompting is that image datasets sometimes have a large data diversity whereas a per-dataset generic prompt can hardly handle the complex distribution shift toward the original pretraining data distribution properly. To address this issue, we propose a dataset Diversity-Aware prompting strategy whose initialization is realized by a Meta-prompt. Specifically, we cluster the downstream dataset into small homogeneity subsets in a diversity-adaptive way, with each subset has its own prompt optimized separately. Such a divide-and-conquer design reduces the optimization difficulty greatly and significantly boosts the prompting performance. Furthermore, all the prompts are initialized with a meta-prompt, which is learned across several datasets. It is a bootstrapped paradigm, with the key observation that the prompting knowledge learned from previous datasets could help the prompt to converge faster and perform better on a new dataset. During inference, we dynamically select a proper prompt for each input, based on the feature distance between the input and each subset. Through extensive experiments, our DAM-VP demonstrates superior efficiency and effectiveness, clearly surpassing previous prompting methods in a series of downstream datasets for different pretraining models. Our code is available at: https://github.com/shikiw/DAM-VP. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
CVPR | 1 |
| 2023 | Improving Adversarial Robustness of Masked Autoencoders via Test-time Frequency-domain PromptingabstractIn this paper, we investigate the adversarial robustness of vision transformers that are equipped with BERT pretraining (e.g., BEiT, MAE). A surprising observation is that MAE has significantly worse adversarial robustness than other BERT pretraining methods. This observation drives us to rethink the basic differences between these BERT pretraining methods and how these differences affect the robustness against adversarial perturbations. Our empirical analysis reveals that the adversarial robustness of BERT pretraining is highly related to the reconstruction target, i.e., predicting the raw pixels of masked image patches will degrade more adversarial robustness of the model than predicting the semantic context, since it guides the model to concentrate more on medium-/high-frequency components of images. Based on our analysis, we provide a simple yet effective way to boost the adversarial robustness of MAE. The basic idea is using the dataset-extracted domain knowledge to occupy the medium-/high-frequency of images, thus narrowing the optimization space of adversarial perturbations. Specifically, we group the distribution of pretraining data and optimize a set of cluster-specific visual prompts on frequency domain. These prompts are incorporated with input images through prototype-based prompt selection during test period. Extensive evaluation shows that our method clearly boost MAE’s adversarial robustness while maintaining its clean performance on ImageNet-1k classification. Our code is available at: https://github.com/shikiw/RobustMAE. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Yinpeng Chen, Lu Yuan 0001, Gang Hua 0001, Weiming Zhang 0001, Nenghai Yu |
ICCV | 1 |
| 2023 | Hierarchical Terrain Attention and Multi-Scale Rainfall Guidance for Flood Image PredictionabstractWith the deterioration of climate, the phenomenon of rain-induced flooding has become frequent. To mitigate its impact, recent works adopt convolutional neural network or its variants to predict the floods. However, these methods directly force the model to reconstruct the raw pixels of flood images through a global constraint, overlooking the underlying information contained in terrain features and rainfall patterns. To address this, we present a novel framework for precise flood map prediction, which incorporates hierarchical terrain spatial attention to help the model focus on spatially-salient areas of terrain features and constructs multi-scale rainfall embedding to extensively integrate rainfall pattern information into generation. To better adapt the model in various rainfall conditions, we leverage a rainfall regression loss for both the generator and the discriminator as additional supervision. Extensive evaluations on real catchment datasets demonstrate the superior performance of our method, which greatly surpasses the previous arts under different rainfall conditions. Qidong Huang, Shaoqing Chen |
ICIP | 4 |
| 2023 | Ada3Diff: Defending against 3D Adversarial Point Clouds via Adaptive DiffusionabstractDeep 3D point cloud models are sensitive to adversarial attacks, which poses threats to safety-critical applications such as autonomous driving. Robust training and defend-by-denoising are typical strategies for defending adversarial perturbations. However, they either induce massive computational overhead or rely heavily upon specified priors, limiting generalized robustness against attacks of all kinds. To remedy it, this paper introduces a novel distortion-aware defense framework that can rebuild the pristine data distribution with a tailored intensity estimator and a diffusion model. To perform distortion-aware forward diffusion, we design a distortion estimation algorithm that is obtained by summing the distance of each point to the best-fitting plane of its local neighboring points, which is based on the observation of the local spatial properties of the adversarial point cloud. By iterative diffusion and reverse denoising, the perturbed point cloud under various distortions can be restored back to a clean distribution. This approach enables effective defense against adaptive attacks with varying noise budgets, enhancing the robustness of existing 3D deep recognition models. Hang Zhou 0007, Jie Zhang 0073, Qidong Huang, Weiming Zhang 0001, Nenghai Yu |
ACM Multimedia | 4 |
| 2022 | Shape-invariant 3D Adversarial Point CloudsabstractAdversary and invisibility are two fundamental but conflict characters of adversarial perturbations. Previous adversarial attacks on 3D point cloud recognition have often been criticized for their noticeable point outliers, since they just involve an “implicit constrain” like global distance loss in the time-consuming optimization to limit the generated noise. While point cloud is a highly structured data format, it is hard to constrain its perturbation with a simple loss or metric properly. In this paper, we propose a novel Point-Cloud Sensitivity Map to boost both the efficiency and imperceptibility of point perturbations. This map reveals the vulnerability of point cloud recognition models when encountering shape-invariant adversarial noises. These noises are designed along the shape surface with an “explicit constrain” instead of extra distance loss. Specifically, we first apply a reversible coordinate transformation on each point of the point cloud input, to reduce one degree of point freedom and limit its movement on the tangent plane. Then we calculate the best attacking direction with the gradients of the transformed point cloud obtained on the white-box model. Finally we assign each point with a non-negative score to construct the sensitivity map, which benefits both white-box adversarial invisibility and black-box query-efficiency extended in our work. Extensive evaluations prove that our method can achieve the superior performance on various point cloud recognition models, with its satisfying adversarial imperceptibility and strong resistance to different point cloud defense settings. Our code is available at: https://github.com/shikiw/SI-Adv. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Hang Zhou 0007, Weiming Zhang 0001, Nenghai Yu |
CVPR | 1 |
| 2022 | Poison Ink: Robust and Invisible Backdoor AttackabstractRecent research shows deep neural networks are vulnerable to different types of attacks, such as adversarial attacks, data poisoning attacks, and backdoor attacks. Among them, backdoor attacks are the most cunning and can occur in almost every stage of the deep learning pipeline. Backdoor attacks have attracted lots of interest from both academia and industry. However, most existing backdoor attack methods are visible or fragile to some effortless pre-processing such as common data transformations. To address these limitations, we propose a robust and invisible backdoor attack called "Poison Ink". Concretely, we first leverage the image structures as target poisoning areas and fill them with poison ink (information) to generate the trigger pattern. As the image structure can keep its semantic meaning during the data transformation, such a trigger pattern is inherently robust to data transformations. Then we leverage a deep injection network to embed such input-aware trigger pattern into the cover image to achieve stealthiness. Compared to existing popular backdoor attack methods, Poison Ink outperforms both in stealthiness and robustness. Through extensive experiments, we demonstrate that Poison Ink is not only general to different datasets and network architectures but also flexible for different attack scenarios. Besides, it also has very strong resistance against many state-of-the-art defense techniques. Jie Zhang 0073, Dongdong Chen 0001, Qidong Huang, Jing Liao 0001, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Image Process. | 3 |
| 2021 | Initiative Defense against Facial ManipulationabstractBenefiting from the development of generative adversarial networks (GAN), facial manipulation has achieved significant progress in both academia and industry recently. It inspires an increasing number of entertainment applications but also incurs severe threats to individual privacy and even political security meanwhile. To mitigate such risks, many countermeasures have been proposed. However, the great majority methods are designed in a passive manner, which is to detect whether the facial images or videos are tampered after their wide propagation. These detection-based methods have a fatal limitation, that is, they only work for ex-post forensics but can not prevent the engendering of malicious behavior. To address the limitation, in this paper, we propose a novel framework of initiative defense to degrade the performance of facial manipulation models controlled by malicious users. The basic idea is to actively inject imperceptible venom into target facial data before manipulation. To this end, we first imitate the target manipulation model with a surrogate model, and then devise a poison perturbation generator to obtain the desired venom. An alternating training strategy are further leveraged to train both the surrogate model and the perturbation generator. Two typical facial manipulation tasks: face attribute editing and face reenactment, are considered in our initiative defense framework. Extensive experiments demonstrate the effectiveness and robustness of our framework in different settings. Finally, we hope this work can shed some light on initiative countermeasures against more adversarial scenarios. Qidong Huang, Jie Zhang 0073, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
AAAI | 1 |
| 2021 | Deep Template-Based WatermarkingabstractTraditional watermarking algorithms have been extensively studied. As an important type of watermarking schemes, template-based approaches maintain a very high embedding rate. In such scheme, the message is often represented by some dedicatedly designed templates, and then the message embedding process is carried out by additive operation with the templates and the host image. To resist potential distortions, these templates often need to contain some special statistical features so that they can be successfully recovered at the extracting side. But in existing methods, most of these features are handcrafted and too simple, thus making them not robust enough to resist serious distortions unless very strong and obvious templates are used. Inspired by the powerful feature learning capacity of deep neural network, we propose the first deep template-based watermarking algorithm in this paper. Specifically, at the embedding side, we first design two new templates for message embedding and locating, which is achieved by leveraging the special properties of human visual system, i.e., insensitivity to specific chrominance components, the proximity principle and the oblique effect. At the extracting side, we propose a novel two-stage deep neural network, which consists of an auxiliary enhancing sub-network and a classification sub-network. Thanks to the power of deep neural networks, our method achieves both digital editing resilience and camera shooting resilience based on typical application scenarios. Through extensive experiments, we demonstrate that the proposed method can achieve much better robustness than existing methods while guaranteeing the original visual quality. Han Fang 0004, Dongdong Chen 0001, Qidong Huang, Jie Zhang 0073, Zehua Ma, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |