Qinlei Huang

dblp:409/8573 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0009-0318-7599ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are correctly matched. However, in real-world scenarios, due to high triplet annotation costs, CIR datasets inevitably contain annotation errors, resulting in incorrectly matched triplets. To address this issue, the problem of Noisy Triplet Correspondence (NTC) has attracted growing attention. We argue that noise in CIR can be categorized into two types: cross-modal correspondence noise and modality-inherent noise. The former arises from mismatches across modalities, whereas the latter originates from intra-modal background interference or visual factors irrelevant to the coarse-grained modification annotations. However, modality-inherent noise is often overlooked, and research on cross-modal correspondence noise remains nascent. To tackle above issues, we propose the Invariance and discrimiNaTion-awarE Noise neTwork (INTENT), comprising two components: Visual Invariant Composition and Bi-Objective Discriminative Learning, specifically designed to handle the two-aspect noise. The former applies causal intervention on the visual side via Fast Fourier Transform (FFT) to generate intervened composed features, enforcing visual invariance and enabling the model to ignore modality-inherent noise during composition. The latter adopts collaborative optimization with both positive and negative samples, and constructs a scalable decision boundary that dynamically adjusts decisions based on the loyalty degree, enabling robust correspondence discrimination. Extensive experiments on two widely used benchmark datasets demonstrate the superiority and robustness of INTENT.
Zhiwei Chen 0003, Yupeng Hu 0003, Zhiheng Fu, Zixu Li 0001, Qinlei Huang, Yinwei Wei
AAAI6
2026 ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval
abstract
With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a multi-modal query consisting of a reference video and a piece of modification text as input. The modification text conveys the user's intended alterations to the reference video. Based on this input, the model aims to retrieve the most relevant target video. In the CVR task, there exists a substantial discrepancy in information density between video and text modalities. Traditional composition methods tend to bias the composed feature toward the reference video, which leads to suboptimal retrieval performance. This limitation is significant due to the presence of three core challenges: (1) modal contribution entanglement, (2) explicit optimization of composed features, and (3) retrieval uncertainty. To address these challenges, we propose the evidence-dRivEn dual-sTream diRectionAl anChor calibration networK (ReTrack). ReTrack is the first CVR framework that improves multi-modal query understanding by calibrating directional bias in composed features. It consists of three key modules: Semantic Contribution Disentanglement, Composition Geometry Calibration, and Reliable Evidence-driven Alignment. Specifically, ReTrack estimates the semantic contribution of each modality to calibrate the directional bias of the composed feature. It then uses the calibrated directional anchors to compute bidirectional evidence that drives reliable composed-to-target similarity estimation. Moreover, ReTrack exhibits strong generalization to the Composed Image Retrieval (CIR) task, achieving SOTA performance across three benchmark datasets in both CVR and CIR scenarios.
Zixu Li 0001, Yupeng Hu 0003, Zhiwei Chen 0003, Qinlei Huang, Guozhi Qiu, Zhiheng Fu, Meng Liu 0006
AAAI4
2026 HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recommendation systems, it encounters a severe challenge in practical scenarios known as the Noise Triplet Correspondence (NTC) problem. This issue primarily arises from the high cost and subjectivity involved in annotating triplet data. To address this problem, we identify two central challenges: the precise estimation of composed semantic discrepancy and the insufficient progressive adaptation to modification discrepancy. To tackle these challenges, we propose a cHrono-synergiA roBust progressIve learning framework for composed image reTrieval (HABIT), which consists of two core modules. First, the Mutual Knowledge Estimation Module quantifies sample cleanliness by calculating the Transition Rate of mutual information between the composed feature and the target image, thereby effectively identifying clean samples that align with the intended modification semantics. Second, the Dual-consistency Progressive Learning Module introduces a collaborative mechanism between the historical and current models, simulating human habit formation to retain good habits and calibrate bad habits, ultimately enabling robust learning under the presence of NTC. Extensive experiments conducted on two standard CIR datasets demonstrate that HABIT significantly outperforms most methods under various noise ratios, exhibiting superior robustness and retrieval performance.
Zixu Li 0001, Yupeng Hu 0003, Zhiwei Chen 0003, Qinlei Huang, Zhiheng Fu, Yinwei Wei
AAAI5
2026 RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) constitutes a pivotal paradigm requiring models to perform joint reasoning on reference images and modification texts. However, the prevalence of Noisy Triplet Correspondence (NTC) in large-scale datasets severely constrains model performance. Existing denoising methods either target binary mismatches or rely on scalar-based point-wise estimation, neglecting rich global structural correlations among sample populations and dynamic value variations during training, thereby yielding suboptimal results. This paper identifies two critical unresolved challenges: Global Structural Inconsistency of Semantic Correlations and Hard Sample Discrimination Uncertainty. To address these, we propose RankVR, a framework designed to construct a robust CIR model via global structure consistency and dynamic value perception. Specifically, we introduce the Global Structure Consistency Perception (GSCP) module, which utilizes the Effective Rank of the Correlation Matrix to decouple clean samples from structural noise. By measuring rank difference, GSCP identifies samples disrupting macroscopic semantic symmetry. Furthermore, we develop the Adaptive Semantic Value Calibration (ASVC) module to distinguish high-value hard clean samples. By integrating training potential and reliability, it dynamically quantifies the semantic value of each triplet, ensuring effective utilization of hard samples while suppressing noise characterized by logical conflicts. Extensive experiments on the FashionIQ and CIRR benchmark datasets demonstrate that RankVR significantly outperforms existing state-of-the-art methods, validating its superior robustness in noisy environments.
Zixu Li 0001, Zhiheng Fu, Zhiwei Chen 0003, Qinlei Huang, Yupeng Hu 0003
ICMR5
2026 REFINE: Composed Video Retrieval via Shared and Differential Semantics Enhancement
abstract
Composed Video Retrieval (CVR) is a novel video retrieval paradigm. Unlike traditional single-modal video retrieval paradigms (e.g., text to video or video to video), CVR employs multi-modal queries (including both a reference video and a natural language modification) to retrieve the target video that best matches the modified reference video. Existing CVR methods primarily rely on generalized knowledge from vision-language pretrained models or utilize caption expansions to enhance video comprehension. However, these approaches overlook the benefits offered by the shareability and variability of videos for multi-modal query understanding. To overcome this limitation, we introduce a novel CVR framework named shaREd and diFferential semantIcs eNhancement nEtwork ( REFINE ). REFINE is the first framework to exploit the shareability and variability of videos to improve multi-modal query comprehension. Specifically, REFINE leverages learnable tokens to achieve enhanced shared feature representation. Moreover, it introduces a carefully designed Differential Block to disentangle differential semantics between frames and employs modification associations to guide multi-modal query feature fusion. Additionally, REFINE has been extended to the Composed Image Retrieval task, making it effectively generalize across existing composed multi-modal retrieval scenarios and outperform existing methods. Extensive qualitative and quantitative evaluations on four benchmark datasets validate the superiority of the proposed REFINE framework.
Yupeng Hu 0003, Zixu Li 0001, Zhiwei Chen 0003, Qinlei Huang, Zhiheng Fu, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.4
2025 MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image Retrieval
abstract
The Composed Image Retrieval (CIR) task aims to retrieve a target image that meets the requirements based on a given multimodal query (includes a reference image and modification text). Most existing works align multimodal semantics at both local and global granularity. However, they have failed to consider the mining of semantic correspondences at the intermediate-grained level, which has resulted in sub-optimal model performance. In this paper, we propose an adaptive interMEDiate-graIned Aggregation Network (MEDIAN). Compared with the conventional CIR models, MEDIAN is capable of generating intermediate-grained feature aggregation supervised signals and constructing graph attention networks to extract intermediate-grained features. Concurrently, MEDIAN also devises cross-modal semantic correspondence aligning guided by the target image, which in turn enables accurate multi-grained feature composition. The superiority of MEDIAN is demonstrated by extensive experiments on three benchmark datasets. Our code is available at https://windlikeo.github.io/MEDIAN.github.io/.
Qinlei Huang, Zhiwei Chen 0003, Zixu Li 0001, Xuemeng Song, Yupeng Hu 0003, Liqiang Nie
ICASSP1