EDBT 2026 Demo / reviewers in the wild / expert
Zixu Li 0001
dblp:336/5965-1
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2026
0009-0001-5136-159XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are correctly matched. However, in real-world scenarios, due to high triplet annotation costs, CIR datasets inevitably contain annotation errors, resulting in incorrectly matched triplets. To address this issue, the problem of Noisy Triplet Correspondence (NTC) has attracted growing attention. We argue that noise in CIR can be categorized into two types: cross-modal correspondence noise and modality-inherent noise. The former arises from mismatches across modalities, whereas the latter originates from intra-modal background interference or visual factors irrelevant to the coarse-grained modification annotations. However, modality-inherent noise is often overlooked, and research on cross-modal correspondence noise remains nascent. To tackle above issues, we propose the Invariance and discrimiNaTion-awarE Noise neTwork (INTENT), comprising two components: Visual Invariant Composition and Bi-Objective Discriminative Learning, specifically designed to handle the two-aspect noise. The former applies causal intervention on the visual side via Fast Fourier Transform (FFT) to generate intervened composed features, enforcing visual invariance and enabling the model to ignore modality-inherent noise during composition. The latter adopts collaborative optimization with both positive and negative samples, and constructs a scalable decision boundary that dynamically adjusts decisions based on the loyalty degree, enabling robust correspondence discrimination. Extensive experiments on two widely used benchmark datasets demonstrate the superiority and robustness of INTENT. Zhiwei Chen 0003, Yupeng Hu 0003, Zhiheng Fu, Zixu Li 0001, Qinlei Huang, Yinwei Wei |
AAAI | 4 |
| 2026 | ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalabstractWith the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a multi-modal query consisting of a reference video and a piece of modification text as input. The modification text conveys the user's intended alterations to the reference video. Based on this input, the model aims to retrieve the most relevant target video. In the CVR task, there exists a substantial discrepancy in information density between video and text modalities. Traditional composition methods tend to bias the composed feature toward the reference video, which leads to suboptimal retrieval performance. This limitation is significant due to the presence of three core challenges: (1) modal contribution entanglement, (2) explicit optimization of composed features, and (3) retrieval uncertainty. To address these challenges, we propose the evidence-dRivEn dual-sTream diRectionAl anChor calibration networK (ReTrack). ReTrack is the first CVR framework that improves multi-modal query understanding by calibrating directional bias in composed features. It consists of three key modules: Semantic Contribution Disentanglement, Composition Geometry Calibration, and Reliable Evidence-driven Alignment. Specifically, ReTrack estimates the semantic contribution of each modality to calibrate the directional bias of the composed feature. It then uses the calibrated directional anchors to compute bidirectional evidence that drives reliable composed-to-target similarity estimation. Moreover, ReTrack exhibits strong generalization to the Composed Image Retrieval (CIR) task, achieving SOTA performance across three benchmark datasets in both CVR and CIR scenarios. Zixu Li 0001, Yupeng Hu 0003, Zhiwei Chen 0003, Qinlei Huang, Guozhi Qiu, Zhiheng Fu, Meng Liu 0006 |
AAAI | 1 |
| 2026 | HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recommendation systems, it encounters a severe challenge in practical scenarios known as the Noise Triplet Correspondence (NTC) problem. This issue primarily arises from the high cost and subjectivity involved in annotating triplet data. To address this problem, we identify two central challenges: the precise estimation of composed semantic discrepancy and the insufficient progressive adaptation to modification discrepancy. To tackle these challenges, we propose a cHrono-synergiA roBust progressIve learning framework for composed image reTrieval (HABIT), which consists of two core modules. First, the Mutual Knowledge Estimation Module quantifies sample cleanliness by calculating the Transition Rate of mutual information between the composed feature and the target image, thereby effectively identifying clean samples that align with the intended modification semantics. Second, the Dual-consistency Progressive Learning Module introduces a collaborative mechanism between the historical and current models, simulating human habit formation to retain good habits and calibrate bad habits, ultimately enabling robust learning under the presence of NTC. Extensive experiments conducted on two standard CIR datasets demonstrate that HABIT significantly outperforms most methods under various noise ratios, exhibiting superior robustness and retrieval performance. Zixu Li 0001, Yupeng Hu 0003, Zhiwei Chen 0003, Qinlei Huang, Zhiheng Fu, Yinwei Wei |
AAAI | 1 |
| 2026 | TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image RetrievalabstractComposed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that consists of a reference image and modification text. Although research on CIR has made significant progress, prevailing setups still rely simple modification texts that typically cover only a limited range of salient changes, which induces two limitations highly relevant to practical applications, namely Insufficient Entity Coverage and Clause-Entity Misalignment. In order to address these issues and bring CIR closer to real-world use, we construct two instruction-rich multi-modification datasets, M-FashionIQ and M-CIRR. In addition, we propose TEMA, the Text-oriented Entity Mapping Architecture, which is the first CIR framework designed for multi-modification while also accommodating simple modifications. Extensive experiments on four benchmark datasets demonstrate that TEMA’s superiority in both original and multi-modification scenarios, while maintaining an optimal balance between retrieval accuracy and computational efficiency. Our codes and constructed multi-modification dataset (M-FashionIQ and M-CIRR) are available at https://github.com/lee-zixu/ACL26-TEMA/ Zixu Li 0001, Yupeng Hu 0003, Zhiheng Fu, Zhiwei Chen 0003, Yongqi Li 0001, Liqiang Nie |
ACL (1) | 1 |
| 2026 | IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video RetrievalabstractComposed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal correspondences, they often assume modified objects appear directly in videos. However, modification texts frequently describe concepts not explicitly presented but implicitly expressed through semantically related visual cues (e.g., “cake” implying “birthday party”). Current approaches typically rely on aligning explicit feature representations within the concrete space, neglecting critical latent associations. To address this, we propose an adaptIve scheMa-ImAGery enhanced composItional NEtwork (IMAGINE). Unlike standard explicit matching, IMAGINE materializes implicit semantics (termed schema imagery) via dynamic multimodal prototypes. These prototypes capture shared latent concepts to adaptively modulate visual features, effectively injecting implicit guidance into the retrieval process. By bridging the gap between explicit visual contents and implicit retrieval intentions, IMAGINE achieves state-of-the-art performance in both CVR and Composed Image Retrieval (CIR) across three widely used benchmarks. Zixu Li 0001, Zhiwei Chen 0003, Zhiheng Fu, Yupeng Hu 0003 |
ICMR | 2 |
| 2026 | RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image RetrievalabstractComposed Image Retrieval (CIR) constitutes a pivotal paradigm requiring models to perform joint reasoning on reference images and modification texts. However, the prevalence of Noisy Triplet Correspondence (NTC) in large-scale datasets severely constrains model performance. Existing denoising methods either target binary mismatches or rely on scalar-based point-wise estimation, neglecting rich global structural correlations among sample populations and dynamic value variations during training, thereby yielding suboptimal results. This paper identifies two critical unresolved challenges: Global Structural Inconsistency of Semantic Correlations and Hard Sample Discrimination Uncertainty. To address these, we propose RankVR, a framework designed to construct a robust CIR model via global structure consistency and dynamic value perception. Specifically, we introduce the Global Structure Consistency Perception (GSCP) module, which utilizes the Effective Rank of the Correlation Matrix to decouple clean samples from structural noise. By measuring rank difference, GSCP identifies samples disrupting macroscopic semantic symmetry. Furthermore, we develop the Adaptive Semantic Value Calibration (ASVC) module to distinguish high-value hard clean samples. By integrating training potential and reliability, it dynamically quantifies the semantic value of each triplet, ensuring effective utilization of hard samples while suppressing noise characterized by logical conflicts. Extensive experiments on the FashionIQ and CIRR benchmark datasets demonstrate that RankVR significantly outperforms existing state-of-the-art methods, validating its superior robustness in noisy environments. Zixu Li 0001, Zhiheng Fu, Zhiwei Chen 0003, Qinlei Huang, Yupeng Hu 0003 |
ICMR | 2 |
| 2026 | ERASE: Bypassing Collaborative Detection of AI Counterfeit via Comprehensive Artifacts EliminationabstractThe rapid advancement of AI-Generated Images (AIGI) has amplified concerns about increasingly undetectable deepfakes. Recent adversarial techniques further worsen this problem by enhancing the imperceptibility of synthetic forgeries to both human viewers and automated detection systems. To simulate realistic adversaries and expose detection vulnerabilities, AI-Generated Image Stealth (AIGI-S) methods specifically aim to make synthetic images harder to detect. However, existing AIGI-S approaches often lack universality and transferability across diverse detection models—especially in collaborative detection settings—and tend to prioritize machine deception over human perceptual fidelity, resulting in visible artifacts. Inspired by real-world antique painting forgery, we propose ERASE (comprehensivE counteRfeit ArtifactS Elimination), a stealth-oriented optimization framework designed for multi-detector environments. ERASE comprehensively suppresses generative artifacts and incorporates a perceptual optimization objective to improve deception against both detection algorithms and human examiners. Extensive evaluations across eight distinct generative subsets from the GenImage benchmark and fifteen detection models demonstrate that ERASE delivers substantially improved attack performance—improving single-detector evasion by +10.5% and collaborative detection evasion by +17.9%—while preserving high image quality. Qianyun Yang, Peizhuo Lv, Yingjiu Li, Shengzhi Zhang, Zhiwei Chen 0003, Zixu Li 0001, Yupeng Hu 0003 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2026 | COMBINER: Composed Image Retrieval Guided by Attribute-Based Neighbor RelationsabstractComposed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases where images appear visually alike yet differ in attributes, potentially undermining both multimodal feature fusion and similarity modeling. To mitigate this limitation, we design a unified representation of cross-modal features based on attribute prototypes. Nevertheless, the task is far from straightforward, owing to three core issues: (1) entanglement in attribute-level semantics, (2) inconsistency across modalities, and (3) supervised signal missing. To tackle the above obstacles, we introduce a COMposed image retrieval network guided By attrIbute-based NEighbor Relations (COMBINER). Specifically, we first design an Adaptive Semantic Disentanglement module, which is capable of disentangling attribute features based on multimodal primitive features. Secondly, we propose a Unified Prototype-based Composition module, which can construct cross-modal unified prototypes (CUP) and facilitate multimodal feature composition. Finally, we introduce a Dual Relations Modeling module, which can mine pairwise and neighbor relations based on attribute similarity. Compared to traditional neighbor relations modeling CIR methods, COMBINER represents the first study addressing the phenomenon of visually similar but attribute-unrelated samples. It achieves a more accurate understanding of the semantic relations among samples by employing an attribute prototype-based similarity metric. Comprehensive experiments conducted on three benchmark datasets confirm the effectiveness of our proposed COMBINER. The implementation of our method will be accessed at https://github.com/Lee-zixu/COMBINER. Zixu Li 0001, Yupeng Hu 0003, Zhiwei Chen 0003, Haokun Wen, Xuemeng Song, Liqiang Nie |
IEEE Trans. Image Process. | 1 |
| 2026 | STABLE: Efficient Hybrid Nearest Neighbor Search via Magnitude-Uniformity and Cardinality-RobustnessabstractHybrid Approximate Nearest Neighbor Search (Hybrid ANNS) is a foundational search technology for large-scale heterogeneous data and has gained significant attention in both academia and industry. However, current approaches overlook the heterogeneity in data distribution, thus ignoring two major challenges: the Compatibility Barrier for Similarity Magnitude Heterogeneity and the Tolerance Bottleneck to Attribute Cardinality. To overcome these issues, we propose the robuSt he Terogeneity-Aware hyBrid retrievaL framEwork, STABLE, designed for accurate, efficient, and robust hybrid ANNS under datasets with various distributions. Specifically, we introduce an enhAnced heterogeneoUs semanTic perceptiOn (AUTO) metric to achieve a joint measurement of feature similarity and attribute consistency, addressing similarity magnitude heterogeneity and improving robustness to datasets with various attribute cardinalities. Thereafter, we construct our Heterogeneous Emanticre Lation graPh (HELP) index based on AUTO to organize heterogeneous semantic relations. Finally, we employ a novel Dynamic Heterogeneity Routing method to ensure an efficient search. Extensive experiments on five feature vector benchmarks with various attribute cardinalities demonstrate the superior performance of STABLE. Qianyun Yang, Zhiwei Chen 0003, Yupeng Hu 0003, Zixu Li 0001, Zhiheng Fu, Liqiang Nie |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2026 | REFINE: Composed Video Retrieval via Shared and Differential Semantics EnhancementabstractComposed Video Retrieval (CVR) is a novel video retrieval paradigm. Unlike traditional single-modal video retrieval paradigms (e.g., text to video or video to video), CVR employs multi-modal queries (including both a reference video and a natural language modification) to retrieve the target video that best matches the modified reference video. Existing CVR methods primarily rely on generalized knowledge from vision-language pretrained models or utilize caption expansions to enhance video comprehension. However, these approaches overlook the benefits offered by the shareability and variability of videos for multi-modal query understanding. To overcome this limitation, we introduce a novel CVR framework named shaREd and diFferential semantIcs eNhancement nEtwork ( REFINE ). REFINE is the first framework to exploit the shareability and variability of videos to improve multi-modal query comprehension. Specifically, REFINE leverages learnable tokens to achieve enhanced shared feature representation. Moreover, it introduces a carefully designed Differential Block to disentangle differential semantics between frames and employs modification associations to guide multi-modal query feature fusion. Additionally, REFINE has been extended to the Composed Image Retrieval task, making it effectively generalize across existing composed multi-modal retrieval scenarios and outperform existing methods. Extensive qualitative and quantitative evaluations on four benchmark datasets validate the superiority of the proposed REFINE framework. Yupeng Hu 0003, Zixu Li 0001, Zhiwei Chen 0003, Qinlei Huang, Zhiheng Fu, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalabstractThe objective of Composed Image Retrieval (CIR) is to identify a target image that meets the requirement based on a multimodal query (including the reference image and the modification text) provided by the user. Despite the notable success of existing approaches, they fail to adequately address the modification relation between visual entities and modification actions. This limitation is non-trivial due to three challenges: 1) irrelevant factor perturbation, 2) vague semantic boundaries, and 3) implicit modification relations. To address the above challenges, we propose an Entity miNing and modifiCation relatiOn binDing nEtwoRk (ENCODER), which has been designed to mine visual entities and modification actions, and then bind modification relations. Among the various components of the proposed ENCODER, we have initially designed the Latent Factor Filter (LFF) module to filter visual and textual latent factors related to modification semantics based on a threshold gating mechanism. Secondly, we propose Entity-Action Binding (EAB), which comprises modality-shared Learnable Relation Queries (LRQ) that are capable of mining visual entities and modification actions, as well as learning implicit modification relations for entity-action binding. Finally, the Multi-scale Composition module is introduced to achieve multi-scale feature composition, with guidance provided by entity-action binding. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed method. Zixu Li 0001, Zhiwei Chen 0003, Haokun Wen, Zhiheng Fu, Yupeng Hu 0003, Weili Guan |
AAAI | 1 |
| 2025 | PAIR: Complementarity-guided Disentanglement for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a novel image retrieval paradigm that aims at searching for the target images via the multimodal query including a reference image and a modification text. Although existing works have made significant progress, they overlook the inter-modal coherence and incoherence relations modeling, hindering the retrieval accuracy of CIR models. This limitation is non-trivial due to the following two challenges: 1) inter-modal incoherence and 2) intra-modal entanglement. To address the above challenges, we propose a comPlementArity-guided dIsentanglement netwoRk (PAIR), which can disentangle the features of multimodal queries from a semantic coherence perspective, thereby facilitating the identification of both complementary coherent and incoherent features. Furthermore, based on disentangled features, PAIR develops an asymmetric feature composition module, which is designed to enhance the retrieval performance of the model. Extensive experiments on three benchmark datasets demonstrate the superiority of PAIR. The code is available at https://zhihfu.github.io/PAIR.github.io/. Zhiheng Fu, Zixu Li 0001, Zhiwei Chen 0003, Xuemeng Song, Yupeng Hu 0003, Liqiang Nie |
ICASSP | 2 |
| 2025 | MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image RetrievalabstractThe Composed Image Retrieval (CIR) task aims to retrieve a target image that meets the requirements based on a given multimodal query (includes a reference image and modification text). Most existing works align multimodal semantics at both local and global granularity. However, they have failed to consider the mining of semantic correspondences at the intermediate-grained level, which has resulted in sub-optimal model performance. In this paper, we propose an adaptive interMEDiate-graIned Aggregation Network (MEDIAN). Compared with the conventional CIR models, MEDIAN is capable of generating intermediate-grained feature aggregation supervised signals and constructing graph attention networks to extract intermediate-grained features. Concurrently, MEDIAN also devises cross-modal semantic correspondence aligning guided by the target image, which in turn enables accurate multi-grained feature composition. The superiority of MEDIAN is demonstrated by extensive experiments on three benchmark datasets. Our code is available at https://windlikeo.github.io/MEDIAN.github.io/. Qinlei Huang, Zhiwei Chen 0003, Zixu Li 0001, Xuemeng Song, Yupeng Hu 0003, Liqiang Nie |
ICASSP | 3 |
| 2025 | OFFSET: Segmentation-based Focus Shift Revision for Composed Image RetrievalabstractComposed Image Retrieval (CIR) represents a novel retrieval paradigm that is capable of expressing users' intricate retrieval requirements flexibly. It enables the user to give a multimodal query, comprising a reference image and a modification text, and subsequently retrieve the target image. Notwithstanding the considerable advances made by prevailing methodologies, CIR remains in its nascent stages due to two limitations: 1) inhomogeneity between dominant and noisy portions in visual data is ignored, leading to query feature degradation, and 2) the priority of textual data in the image modification process is overlooked, which leads to a visual focus bias. To address these two limitations, this work presents a focus mapping-based feature extractor, which consists of two modules: dominant portion segmentation and dual focus mapping. It is designed to identify significant dominant portions in images and guide the extraction of visual and textual data features, thereby reducing the impact of noise interference. Subsequently, we propose a textually guided focus revision module, which can utilize the modification requirements implied in the text to perform adaptive focus revision on the reference image, thereby enhancing the perception of the modification focus on the composed features. The aforementioned modules collectively constitute the segmentatiOn-based Focus shiFt reviSion nETwork (OFFSET), and comprehensive experiments on four benchmark datasets substantiate the superiority of our proposed method. The codes and data are available on https://zivchen-ty.github.io/OFFSET.github.io/. Zhiwei Chen 0003, Yupeng Hu 0003, Zixu Li 0001, Zhiheng Fu, Xuemeng Song, Liqiang Nie |
ACM Multimedia | 3 |
| 2025 | HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video RetrievalabstractComposed Video Retrieval (CVR) is a challenging video retrieval task that utilizes multi-modal queries, consisting of a reference video and modification text, to retrieve the desired target video. The core of this task lies in understanding the multi-modal composed query and achieving accurate composed feature learning. Within multi-modal queries, the video modality typically carries richer semantic content compared to the textual modality. However, previous works have largely overlooked the disparity in information density between these two modalities. This limitation can lead to two critical issues: 1) modification subject referring ambiguity and 2) limited detailed semantic focus, both of which degrade the performance of CVR models. To address the aforementioned issues, we propose a novel CVR framework, namely the Hierarchical Uncertainty-aware Disambiguation network (HUD). HUD is the first framework that leverages the disparity in information density between video and text to enhance multi-modal query understanding. It comprises three key components: (a) Holistic Pronoun Disambiguation, (b) Atomistic Uncertainty Modeling, and (c) Holistic-to-Atomistic Alignment. By exploiting overlapping semantics through holistic cross-modal interaction and fine-grained semantic alignment via atomistic-level cross-modal interaction, HUD enables effective object disambiguation and enhances the focus on detailed semantics, thereby achieving precise composed feature learning. Moreover, our proposed HUD is also applicable to the Composed Image Retrieval (CIR) task and achieves state-of-the-art performance across three benchmark datasets for both CVR and CIR tasks. The codes are available on https://zivchen-ty.github.io/HUD.github.io/. Zhiwei Chen 0003, Yupeng Hu 0003, Zixu Li 0001, Zhiheng Fu, Haokun Wen, Weili Guan |
ACM Multimedia | 3 |
| 2024 | Explicit Granularity and Implicit Scale Correspondence Learning for Point-Supervised Video Moment LocalizationabstractVideo moment localization (VML) aims to identify the temporal boundary semantically matching the given query. Point-supervised VML balances localization accuracy and annotation cost but is still immature due to granularity alignment and scale perception issues. To this end, we propose a Semantic Granularity and Scale Correspondence Integration (SG-SCI) framework aimed at leveraging limited single-frame annotation for correspondence learning. It explicitly models semantic relations of different feature granularities and adaptively mines the implicit semantic scale, thereby enhancing feature representations of varying granularities and scales. SG-SCI uses granularity correspondence alignment to align semantics via latent prior knowledge and a scale correspondence learning to identify and address semantic scale differences. Extensive experiments on benchmark datasets have demonstrated the promising performance of our model over several state-of-the-art competitors. Kun Wang 0039, Hao Liu 0072, Lirong Jie, Zixu Li 0001, Yupeng Hu 0003, Liqiang Nie |
ACM Multimedia | 4 |