EDBT 2026 Demo / reviewers in the wild / expert
Fengling Li 0001
dblp:07/2900-1
· DBLP profile ↗
34ranked-venue papers
6as first author
34since 2021 · last 2026
0009-0000-2942-4868ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalizing Vision-Language Models with Dedicated Prompt GuidanceabstractFine-tuning large pretrained vision-language models (VLMs) has emerged as a prevalent paradigm for downstream adaptation, yet it faces a critical trade-off between domain specificity and domain generalization (DG) ability. Current methods typically fine-tune a universal model on the entire dataset, which potentially compromises the ability to generalize to unseen domains. To fill this gap, we provide a theoretical understanding of the generalization ability for VLM fine-tuning, which reveals that training multiple parameter-efficient expert models on partitioned source domains leads to better generalization than fine-tuning a universal model. Inspired by this finding, we propose a two-step domain-expert-Guided DG (GuiDG) framework. GuiDG first employs prompt tuning to obtain source domain experts, then introduces a Cross-Modal Attention module to guide the fine-tuning of the vision encoder via adaptive expert integration. To better evaluate few-shot DG, we construct ImageNet-DG from ImageNet and its variants. Extensive experiments on standard DG benchmarks and ImageNet-DG demonstrate that GuiDG improves upon state-of-the-art fine-tuning methods while maintaining efficiency. Yinjie Min, Zhekai Du, Fengling Li 0001, Jingjing Li 0001 |
AAAI | 5 |
| 2026 | Same Frequency Begets Shared Interests: Popular-Niche Wavelet Graph Learning for Multimodal Recommendation
Jingxi Xie, Fengling Li 0001, Jingjing Li 0001 |
SIGIR | 4 |
| 2026 | Depression Detection from Social Media: A Mutual Guidance Multi-modal Network with Complementary Graph LearningabstractDepression has become a critical global public health challenge, creating an urgent need for automated and scalable screening solutions. Social media platforms, which capture rich and spontaneous multi-modal behavioral data, offer a promising avenue for detecting early signs of mental distress. However, existing depression detection methods predominantly rely on static multi-modal fusion strategies and frequently fail to effectively tackle cross-modal semantic gaps. To address these limitations, we propose a Mutual Guidance Multi-modal Network with Complementary Graph Learning (MGMN) for depression detection by observing individuals' behavioral performance on social media. Specifically, a cross-modal mutual guidance mechanism is designed to dynamically construct a complementary graph by using mutual similarities within and across visual and acoustic modalities common in social media. More specifically, based on this complementary graph, a modality-specific adaptive residual learning module is applied to each modality to stabilize deep feature learning and preserve modality-specific and complementary information via graph-conditioned adaptive residual fusion. Furthermore, the refined uni-modal features are subsequently fed into a joint-modal fusion and prediction module to output the final disease prediction probability. Extensive experiments on the MUD3, LMVD, and D-vlog datasets demonstrate our proposed method's superiority over state-of-the-art methods, confirming that the proposed framework provides a robust and effective solution for mental health monitoring. Codes are available at https://github.com/Petofi-romance/MGMN Guocheng Hu, Chaoqun Zheng, Ruifan Zuo, Fengling Li 0001, Dan Shi 0003, Xiaofeng Qu, Wenpeng Lu |
SIGIR | 4 |
| 2026 | GANPOI: A multi-factor generative adversarial framework for POI recommendationabstractThe Point-of-Interest (POI) recommender systems utilize users’ historical check-in sequences to predict future visits and are widely used in location-based social service platforms. However, users’ check-in behaviors often exhibit strong spatial constraints and temporal dependencies, as users typically visit only a limited set of POIs. This leads to a long-tail distribution and severe data sparsity, which obscures the contribution of individual decision factors such as specific POI characteristics, temporal preferences, and geographic proximity, and limits model generalization due to insufficient fine-grained supervision. To address these issues, we propose a multi-factor generative adversarial framework, GANPOI, that integrates a multi-head discriminator within a generative adversarial network architecture to mitigate data sparsity and improve recommendation accuracy. Specifically, the proposed model consists of two primary components: (1) the generator module models user trajectories using a Transformer-based architecture and employs adversarial training to capture self-supervised signals from check-in contexts; and (2) a multi-head discriminator module composed of several sub-discriminators, each performing task-specific discrimination from different semantic perspectives. This structure enhances the model’s ability to capture fine-grained contextual factors influencing user decisions. Extensive experiments on three real-world datasets demonstrate that GANPOI consistently outperforms existing state-of-the-art POI recommendation methods. To facilitate reproducibility, we have released both the code and the POI recommendation datasets we used. The source codes and datasets are available at https://anonymous.4open.science/r/GANPOI-YYR . Yarong Yu, Yang Xu 0025, Lei Zhu 0002, Fengling Li 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Dynamic Bit-Wise Semantic Transformer Hashing for Multi-Modal RetrievalabstractMulti-modal hashing aims to succinctly encode heterogeneous modalities into binary hash codes, facilitating efficient multimedia retrieval characterized by low storage demands and high retrieval speed. Despite the commendable achievements of existing methods, they still face three crucial challenges: 1) Inadequate bridging of the heterogeneous modality gap through coarse, global feature-level alignment and fusion. 2) The erosion of bit independence and consequent limitations on the semantic representation capacity of hash codes during feature-level hash code learning. 3) The insufficiency of binary label-based pairwise semantic preservation strategies in capturing intricate fine-grained semantic correlations within multi-modal data. To address these challenges, this paper introduces the Dynamic Bit-wise Semantic Transformer Hashing (DBSTH) framework. Remarkably, it treats each hash bit as a unique semantic concept, facilitating concept-level alignment of heterogeneous modalities. This safeguards bit independence and augments representation capabilities. Specifically, we devise a dynamic unit fusion strategy for the adaptive combination of local multi-modal information units, facilitating the acquisition of bit-wise semantic concepts. Subsequently, we incorporate a transformer encoder to refine these concepts by uncovering latent correlations among distinct concepts. Finally, we perform the multi-modal alignment and fusion on the fine-grained concept-level, independently encoding each concept to its corresponding hash bit. To provide enhanced guidance for concept learning, a label prototype learning mechanism is introduced, which learns prototype embeddings for all categories through the consideration of co-occurrence priors. This mechanism effectively captures fine-grained explicit semantic correlations and generates supervising hash codes. Additionally, to improve the robustness of the hashing model in handling noisy multi-modal data, a masked concept learning strategy is introduced, facilitating the acquisition of resilient semantic concepts. Extensive experiments conducted on three widely tested multi-modal retrieval datasets demonstrate the superiority of our method in conventional, noisy, and open-set retrieval scenarios. Wentao Tan, Fengling Li 0001, Lei Zhu 0002, Weili Guan, Jingjing Li 0001, Zhiyong Cheng 0001, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Unified stable and generalizable online hashing for cross-modal retrieval
Tianshi Wang 0001, Fengling Li 0001, Guohua Dong, Lei Zhu 0002 |
Pattern Recognit. | 3 |
| 2026 | Prompt-Driven Bit Extension Hashing for Continual Cross-Modal RetrievalabstractContinual cross-modal hashing is critical for efficient retrieval across heterogeneous modalities in dynamic environments. Yet, existing approaches primarily focus on mitigating catastrophic forgetting, while overlooking two key challenges: 1) the hash collision arising from the excessive utilization of the Hamming space across tasks, and 2) the absence of consistency modeling for cross-modal dynamic associations. To address these challenges, we introduce Prompt-driven Bit Extension Hashing (PBEH), a novel framework that dynamically extends hash codes to prevent hash collisions and capture evolving modality-aligned semantics in continuously expanding multi-modal data. Specifically, PBEH first adaptively initializes a set of modality-shared prompts for each task, which are jointly optimized with the hashing functions to enhance model plasticity and retain task-specific knowledge, enabling continual cross-modal semantic alignment. In parallel, a dynamic Hamming space extension mechanism allocates dedicated capacity per task, alleviating bottlenecks and inter-task collisions. During retrieval, queries are encoded via the extended hash functions and matched to stored codes using a truncated strategy for compatibility. To ensure efficiency and semantic stability, only the prompts and hashing functions are updated while the pre-trained backbone remains frozen. Extensive experiments demonstrate that PBEH achieves superior and stable performance in continual cross-modal retrieval with substantially reduced computational overhead. The source codes and datasets are available at https://github.com/Liuwwhh/PBEH. Tianshi Wang 0001, Fengling Li 0001, Jingjing Li 0001, Lei Zhu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Training-Free Open-Set Domain Adaptation With Vision-Language ModelsabstractWith the prevalence of pre-trained vision-language models like CLIP, leveraging the generic knowledge embedded in CLIP for domain adaptation has proved to be a promising direction. However, most existing CLIP-based methods are limited to closed-set settings. This is primarily because CLIP needs the semantic labels of unknown classes for inference, thus making it not applicable to Open-Set Domain Adaptation (OSDA). To utilize the complementary roles of CLIP and the source model, our paper proposes a novel Semantic-guided Target Adaptation (SemTA) framework for OSDA in a training-free manner. Specifically, we introduce an unknown semantic discovery module. It uses the cluster centroids of the target data to obtain the semantic labels of unknown classes from the worldwide corpus. Then, the semantic-based inference can be performed with CLIP. Additionally, the dual sample attention mechanism is implemented to output sample-based inference. Representative features from both the source model and CLIP serve as the key to improve task specificity. Compared to previous OSDA methods which reject unknown data by confidence threshold, the proposed approach is more practical and offers better interpretability. Comprehensive evaluations on four benchmarks reveal our method sets a new state-of-the-art even without training. Our code will be publicly available soon. Zhiqi Yu, Ke Lu 0001, Kangkai Wu, Fengling Li 0001, Jingjing Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | PIC-CMH: Efficient Prompt-Infused Continual Cross-Modal HashingabstractCross-modal hashing models face significant challenges in handling continuous data growth, particularly in balancing the plasticity to learn new knowledge and the stability to retain prior cross-modal knowledge. Existing studies partially address this by maintaining previous mappings or extending hash codes, but struggle to reconcile plasticity and stability while requiring heavy parameter optimization. To tackle this, we propose an efficient Prompt-Infused Continual Cross-Modal Hashing (PIC-CMH) approach designed for hash learning with the continuous growth of multi-modal data and emerging knowledge. Specifically, PIC-CMH introduces a finite set of learnable multi-modal prompts, including global and task-specific expert prompts, which work in synergy with multi-modal representations. All prompts are optimized with the hash functions via backpropagation after Gaussian initialization. Global prompts stay learnable throughout, linking tasks, while expert prompts are updated only within their tasks, facilitating knowledge acquisition and mitigating catastrophic forgetting in continual learning. By freezing the pre-trained models used for multi-modal representations, continual learning is confined to the lightweight multi-modal prompts and hash functions, significantly reducing computational overhead. Extensive experiments demonstrate that PIC-CMH effectively addresses the stability-plasticity trade-off in cross-modal hash learning, delivering high retrieval accuracy with low computational cost and a simple yet efficient architecture. The source codes and datasets are available athttps://github.com/styx29-0/PIC-CMH. Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Xiaojun Chang |
IEEE Trans. Multim. | 1 |
| 2026 | EffiPOI: A Product Quantization Framework Based on Knowledge Distillation for Efficient POI RecommendationsabstractIn large-scale Point-of-Interest (POI) recommendation, the conflict between accuracy and computational efficiency intensifies as POI catalogs grow. Traditional deep models struggle to balance quality with efficiency. To address this challenge, we propose a knowledge-distilled product quantization framework EffiPOI for efficient POI recommendation. EffiPOI jointly optimizes accuracy and efficiency by integrating product quantization with multi-modal knowledge distillation. Specifically, we first construct service-oriented multi-modal POI representations, which comprehensively capture each POI’s spatial coverage, temporal activity patterns, and semantic attributes. Based on these representations, we design a teacher-student distillation paradigm. The teacher model adopts a Mixture-of-Experts architecture to generate discriminative and semantically expressive POI representations, which serve as high-quality supervision signals for guiding the student model through knowledge distillation. The student model leverages product quantization to encode POIs into compact and computation-friendly representations, achieving a favorable tradeoff between representational compactness and predictive accuracy. To alleviate the performance degradation due to quantization, we develop a hybrid knowledge distillation strategy that transfers both response-aware and feature-aware knowledge from the teacher model to the student model. Experimental results on three real-world datasets show that the proposed method achieves 4.6%–12.1% improvements in accuracy and over 10× speedup in inference efficiency, outperforming existing POI recommendation models. Code is available at: https://github.com/pcm1217/EffiPOI . Chengmei Peng, Yang Xu 0025, Lei Zhu 0002, Fengling Li 0001, Huaxiang Zhang 0001, Zhigang Ma |
ACM Trans. Inf. Syst. | 4 |
| 2026 | Noise-Robust Generative Hashing for Cross-Modal RetrievalabstractDeep hashing has proven remarkable effectiveness for large-scale cross-modal retrieval, yet its performance is highly vulnerable to supervisory noise, such as mismatched cross-modal correspondences and incorrect category labels. Such noise is prevalent in real-world scenarios, where correspondence mismatches and label inaccuracies often coexist, posing significant challenges for learning accurate multimodal representations. Existing methods typically address only a single type of noise in isolation and neglect the potential value of noisy data, resulting in limited performance gains. To address these challenges, we propose Noise-Robust Generative Hashing (NRGH), a unified framework designed to accommodate various forms of noise inherent in cross-modal retrieval. Specifically, NRGH introduces a hash-driven noise estimation module that computes the confidence score for each multimodal sample by combining frozen auxiliary hash functions with a Gaussian mixture model. Guided by these confidence scores, NRGH performs data correction through two stages: generative text refinement and multi-label probability calibration. The former leverages a pre-trained vision-language model to generate descriptive captions that refine noisy textual information, while the latter corrects noisy labels using confidence-aware soft labels. Furthermore, a dynamic margin contrastive loss adaptively modulates the data contribution of each sample based on its confidence, enabling sample-level adaptive learning. Extensive experiments on benchmark datasets demonstrate that NRGH significantly exceeds state-of-the-art baselines in various noisy scenarios, delivering superior robustness and accuracy. Our source codes and datasets are available at https://github.com/xiaolaohuuu/NRGH . Tianshi Wang 0001, Fengling Li 0001, Jingjing Li 0001, Lei Zhu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Source-free Domain Adaptation with Multiple Alignment for Efficient Image RetrievalabstractDomain adaptation techniques help models generalize to target domains by addressing domain discrepancies between the source and target domain data distributions. These techniques are particularly valuable for cross-domain hashing retrieval, as they reduce training costs while maintaining high retrieval efficiency. However, existing unsupervised domain adaptative hashing methods often require access to both source and target domain data, which may raise privacy concerns regarding source domain data. To address these concerns, restricting access to source domain data is crucial, but this restriction also makes the alignment process between domains more challenging. In this paper, we propose a Source-free Domain Adaptive Hashing with Multiple Alignment (SFDAH-MA) approach for image retrieval. SFDAH-MA integrates model structure alignment, class moment alignment, and semantic relationship alignment to maximize inter-domain knowledge transfer. This comprehensive alignment strategy not only enhances retrieval performance across domains but also minimizes reliance on source domain data, thereby supporting data privacy protection. Extensive experiments show that SFDAH-MA achieves state-of-the-art performance in source-free settings and achieves comparable results to existing unsupervised domain adaptive hashing methods under conventional settings. The source codes of our method are available at: https://github.com/Nikphyc/SFDAH-MA. Jinghui Ni, Hui Cui 0004, Lihai Zhao, Fengling Li 0001, Xiaohui Han, Lijuan Xu 0001 |
ICASSP | 4 |
| 2025 | Flip is Better than Noise: Unbiased Interest Generation for Multimedia RecommendationabstractGenerative diffusion model approaches have achieved remarkable success in multimodal recommendation by generating latent user interest interaction graphs. However, current diffusion methods based on Gaussian noise introduce uncertain interest bias noise. This noise not only disrupts the original user-item interaction bipartite graph structure but also undermines the model's ability to accurately capture user interest preferences. To address these challenges, we propose Unbiased Interest Generation for Multimodal Recommendation (GenRec). Our approach aims to generate valid latent user interests while non-invasively preserving the original interest graph structure. We innovatively introduce the Multi-Modal Interest Generaction Module. During the forward process, we simulate user interest state transitions using ''forward flipping.'' In the reverse stage, we generate binary interaction graphs following a Bernoulli distribution. To further mitigate the random uncertainty during the generation process, we design a Multi-Modal Interest Debiase Module. By constructing a multimodal interest clustering space and using user interest hashing, we correct and enhance the generated interest graphs. Finally, the Multi-modal High-Order Graph Learning Optimization is employed to capture high-order interaction information between users and items. A contrastive learning loss function is used for model optimization. We conduct extensive experiments on different real-world open datasets from the industrial sector. Compared with the state-of-the-art DiffMM, our method significantly improves NDCG@20 by 7.3% on the TikTok dataset and boosts Recall@20 by 4.2% on the Sports dataset. The experimental results validate the effectiveness of GenRec. The code is publicly available at https://github.com/orangeheyue/GenRec-V1. Jingxi Xie, Fengling Li 0001, Lei Zhu 0002, Jingjing Li 0001 |
ACM Multimedia | 3 |
| 2025 | Attack as Defense: Proactive Adversarial Multi-Modal Learning to Evade RetrievalabstractWith growing concerns about information security, protecting the privacy of user-sensitive data has become crucial. The rapid development of multi-modal retrieval technologies poses new threats, making sensitive data more vulnerable to leakage and malicious mining. To address this, we introduce a Proactive Adversarial Multi-modal Learning (PAML) approach that transforms sensitive data into adversarial counterparts, evading malicious multi-modal retrieval and ensuring privacy. Our method starts by sending queries to a knowledge-agnostic retrieval system and analyzing the results to understand the retrieval feedback mechanism. Using a U-Net-based diffusion model, we create a semantic perturbation network that subtly alters the implicit semantics of sensitive data. This, combined with multi-modal retrieved results and random noise, shifts the data's semantics towards outliers, preventing retrieval as neighbors to relevant queries. Additionally, a discriminator and pre-trained model enhance the visual realism and outlier generalization of protected data. Extensive experiments show that PAML outperforms potential baselines in data privacy protection. Ablation analysis validates each component's effectiveness, and our approach's variants are applicable to diverse retrieval systems. Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Jingjing Li 0001, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Generative Augmentation Hashing for Few-Shot Cross-Modal RetrievalabstractDeep cross-modal hashing has demonstrated strong performance in large-scale retrieval but remains challenging in few-shot scenarios due to limited data and weak cross-modal alignment. We propose Generative Augmentation Hashing (GAH), a new framework that synergizes Visual-Language Models (VLMs) and generation-driven hashing to address these limitations. GAH first introduces a cycle generative augmentation mechanism: VLMs generate descriptive textual captions for images, which, combined with label semantics, guide diffusion models to synthesize semantically aligned images via inconsistency filtering. These images then regenerate coherent textual descriptions through VLMs, forming a self-reinforcing cycle that iteratively expands cross-modal data. To resolve the diversity-alignment trade-off in augmentation, we design cross-modal perturbation enhancement, injecting synchronized perturbations with controlled noise to preserve inter-modal semantic relationships while enhancing robustness. Finally, GAH employs dual-level adversarial hash learning, where adversarial alignment of modality-specific and shared latent spaces optimizes both cross-modal consistency and discriminative hash code generation, effectively bridging heterogeneous gaps. Extensive experiments on benchmark datasets show that GAH outperforms state-of-the-art methods in few-shot cross-modal retrieval, achieving significant improvements in retrieval accuracy. Our source codes and datasets are available at https://github.com/xiaolaohuuu/GAH. Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Xiaojun Chang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Online Adaptive Fault Diagnosis With Test-Time Domain AdaptationabstractCross-domain bearing fault diagnosis algorithms have garnered considerable attention in recent years due to their robust ability to address domain bias. However, prevailing methods often grapple with two key challenges: the absence of privacy preservation (necessitating access to source domain data) and the inability to facilitate real-time predictions (requiring iterative training on complete target domain data). In response to these issues, this article introduces an algorithm designed to adapt a pretrained model to the target domain in an online fashion. Notably, data augmentation is employed for pretraining the source domain model, enhancing the generalization capabilities. Subsequently, self-supervised learning is integrated through weight average updating. Furthermore, a memory bank-based approach is introduced to augment the compactness of features within the same class. Evaluation on several public datasets demonstrates that our model not only effectively enhances the diagnostic accuracy of the source model, but also achieves state-of-the-art results compared to other test-time adaptation methods. Kangkai Wu, Jingjing Li 0001, Lichao Meng, Fengling Li 0001, Ke Lu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Fast Partial-Modal Online Cross-Modal HashingabstractCross-Modal Hashing (CMH) has become a powerful technique for large-scale cross-modal retrieval, offering benefits like fast computation and efficient storage. However, most CMH models struggle to adapt to streaming multimodal data in real-time once deployed. Although recent online CMH studies have made progress in this area, they often overlook two key challenges: 1) learning effectively from streaming partial-modal multimodal data, and 2) avoiding the high costs associated with frequent hash function re-training and large-scale updates to database hash codes. To address these issues, we propose Fast Partial-modal Online Cross-Modal Hashing (FPO-CMH), the first approach to tackle online cross-modal hash learning with partial-modal data. This marks a significant shift from previous methods that rely on fully-available multimodal data. Specifically, our approach introduces a multimodal dual-tier anchor bank, initialized using offline training data, which allows offline-trained CMH models to adapt seamlessly to partial-modal data while progressively updating the anchor bank. By leveraging gradient accumulation and asynchronous optimization, FPO-CMH facilitates efficient online cross-modal hash learning. Additionally, an initial-anchor rehearsal strategy is employed to prevent model catastrophic forgetting during online optimization, ensuring the code invariance of database hash codes and eliminating the need for frequent hash function re-training. Extensive experiments validate the superiority of FPO-CMH, especially in handling streaming partial-modal multimodal data, a more realistic scenario. The source codes and datasets are available at https://github.com/DandelionWow/FPO-CMH. Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Xiaojun Chang |
IEEE Trans. Image Process. | 1 |
| 2025 | SecureDA: Privacy-Preserving Source-Free Domain Adaptation for Person Re-IdentificationabstractConventional domain adaptation (DA) for person re-identification (ReID) aims to bridge the domain gap but often requires direct use of fully labeled source and target domains, raising significant data privacy concerns due to the inclusion of personal identity information (PII) in raw data. Source-free domain adaptation (SFDA) for person ReID effectively preserves PII within the authorized source model. Nevertheless, these methods are vulnerable to data privacy (e.g., portrait rights) of the target domain during retrieval, where attackers can exploit pedestrian images for malicious generation, leading to damage to an individual’s reputation. Beyond these limitations, we propose a novel framework called SecureDA to address privacy-preserving SFDA for person ReID, which can generate a privacy key to defend against potential attacks on PII. Technically, we introduce domain-specific adversarial attacks into DA, where the protected query and gallery images are encrypted to ensure secure image retrieval. Furthermore, we employ two simultaneous processes: 1) The global–local adversarial pathway (GLAP) leverages encrypted and original images as adversarial pairs, thereby fostering the development of robust ReID models; 2) The global–local collaborative pathway (GLCP) is mastered through positive pairs collected from the same domain, effectively mitigating the pernicious catastrophic forgetting phenomenon. Extensive experiments show that SecureDA achieves state-of-the-art performance on multiple DA benchmarks and even outperforms the conventional DA and SFDA methods, which inherently compromise data privacy. Xiaofeng Qu, Li Liu 0031, Huaxiang Zhang 0001, Lei Zhu 0002, Liqiang Nie, Xiaojun Chang, Fengling Li 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Plug-In Open-Set Cross-Modal HashingabstractUnsupervised Cross-Modal Hashing (UCMH) models the intrinsic semantic correlations across different modalities to generate binary hash codes, facilitating efficient cross-modal retrieval. This technology offers notable advantages, such as independence from labeled data and superior generalization capabilities compared to supervised methods. However, most UCMH methods are designed for closed-set retrieval scenarios and have difficulty generalizing to open multi-modal data, which is common in real-world retrieval settings. This limitation hampers their performance in open retrieval tasks, particularly when these tasks involve novel categories. To address the above issue, we propose an Open-set Cross-Modal Hashing (OCMH) method, which enhances the generalization capability of trained UCMH models in an efficient plug-in manner for open cross-modal retrieval. Our method enables the model to learn from novel categories in open-set scenarios by increasing the pre-defined hash code length, while simultaneously preventing the catastrophic forgetting of trained knowledge from the closed-set domain using basic hash codes. Additionally, we introduce a historical-category detection module and an asymmetric optimization strategy to support the joint learning of basic and increased hash codes by replaying detected samples related to historical categories. By plugging our proposed method into several representative UCMH methods on three widely used datasets, experimental results show that the enhanced UCMH methods achieve superior retrieval performance in both open-set and closed-set scenarios. The source code is available at:https://github.com/WangBowen7/OCMH/. Lei Zhu 0002, Fengling Li 0001, Hui Cui 0004, Jingjing Li 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Effective Comparative Prototype Hashing for Unsupervised Domain AdaptationabstractUnsupervised domain adaptive hashing is a highly promising research direction within the field of retrieval. It aims to transfer valuable insights from the source domain to the target domain while maintaining high storage and retrieval efficiency. Despite its potential, this field remains relatively unexplored. Previous methods usually lead to unsatisfactory retrieval performance, as they frequently directly apply slightly modified domain adaptation algorithms to hash learning framework, or pursue domain alignment within the Hamming space characterized by limited semantic information. In this paper, we propose a simple yet effective approach named Comparative Prototype Hashing (CPH) for unsupervised domain adaptive image retrieval. We establish a domain-shared unit hypersphere space through prototype contrastive learning and then obtain the Hamming hypersphere space via mapping from the shared hypersphere. This strategy achieves a cohesive synergy between learning uniformly distributed and category conflict-averse feature representations, eliminating domain discrepancies, and facilitating hash code learning. Moreover, by leveraging dual-domain information to supervise the entire hashing model training process, we can generate hash codes that retain inter-sample similarity relationships within both domains. Experimental results validate that our CPH significantly outperforms the state-of-the-art counterparts across multiple cross-domain and single-domain retrieval tasks. Notably, on Office-Home and Office-31 datasets, CPH achieves an average performance improvement of 19.29% and 13.85% on cross-domain retrieval tasks compared to the second-best results, respectively. The source codes of our method are available at: https://github.com/christinecui/CPH. Hui Cui 0004, Lihai Zhao, Fengling Li 0001, Lei Zhu 0002, Xiaohui Han, Jingjing Li 0001 |
AAAI | 3 |
| 2024 | Agile Multi-Source-Free Domain AdaptationabstractEfficiently utilizing rich knowledge in pretrained models has become a critical topic in the era of large models. This work focuses on adaptively utilize knowledge from multiple source-pretrained models to an unlabeled target domain without accessing the source data. Despite being a practically useful setting, existing methods require extensive parameter tuning over each source model, which is computationally expensive when facing abundant source domains or larger source models. To address this challenge, we propose a novel approach which is free of the parameter tuning over source backbones. Our technical contribution lies in the Bi-level ATtention ENsemble (Bi-ATEN) module, which learns both intra-domain weights and inter-domain ensemble weights to achieve a fine balance between instance specificity and domain consistency. By slightly tuning source bottlenecks, we achieve comparable or even superior performance on a challenging benchmark DomainNet with less than 3% trained parameters and 8 times of throughput compared with SOTA method. Furthermore, with minor modifications, the proposed module can be easily equipped to existing methods and gain more than 4% performance boost. Code is available at https://github.com/TL-UESTC/Bi-ATEN. Jingjing Li 0001, Fengling Li 0001, Lei Zhu 0002, Ke Lu 0001 |
AAAI | 3 |
| 2024 | Domain-Agnostic Mutual Prompting for Unsupervised Domain AdaptationabstractConventional Unsupervised Domain Adaptation (UDA) strives to minimize distribution discrepancy between do-mains, which neglects to harness rich semantics from data and struggles to handle complex domain shifts. A promising technique is to leverage the knowledge of large-scale pretrained vision-language models for more guided adaptation. Despite some endeavors, current methods often learn textual prompts to embed domain semantics for source and target domains separately and perform classification within each domain, limiting cross-domain knowledge transfer. Moreover, prompting only the language branch lacks flex-ibility to adapt both modalities dynamically. To bridge this gap, we propose Domain-Agnostic Mutual Prompting (DAMP) to exploit domain-invariant semantics by mutually aligning visual and textual embeddings. Specifically, the image contextual information is utilized to prompt the language branch in a domain-agnostic and instance-conditioned way. Meanwhile, visual prompts are im-posed based on the domain-agnostic textual prompt to elicit domain-invariant visual embeddings. These two branches of prompts are learned mutually with a cross-attention module and regularized with a semantic-consistency loss and an instance-discrimination contrastive loss. Experiments on three UDA benchmarks demonstrate the superiority of DAMP over state-of-the-art approaches1. Zhekai Du, Fengling Li 0001, Ke Lu 0001, Lei Zhu 0002, Jingjing Li 0001 |
CVPR | 3 |
| 2024 | Split to Merge: Unifying Separated Modalities for Unsupervised Domain AdaptationabstractLarge vision-language models (VLMs) like CLIP have demonstrated good zero-shot learning performance in the unsupervised domain adaptation task. Yet, most transfer approaches for VLMs focus on either the language or visual branches, overlooking the nuanced interplay between both modalities. In this work, we introduce a Unified Modality Separation (UniMoS) framework for unsupervised domain adaptation. Leveraging insights from modality gap studies, we craft a nimble modality separation network that distinctly disentangles CLIP's features into language-associated and vision-associated components. Our proposed Modality-Ensemble Training (MET) method fosters the exchange of modality-agnostic information while maintaining modality-specific nuances. We align features across domains using a modality discriminator. Comprehensive evaluations on three benchmarks reveal our approach sets a new state-of-the-art with minimal computational costs. Code: https://github.com/TL-UESTC/UniMoS. Zhekai Du, Fengling Li 0001, Ke Lu 0001, Jingjing Li 0001 |
CVPR | 4 |
| 2024 | SOIL: Contrastive Second-Order Interest Learning for Multimodal RecommendationabstractMainstream multimodal recommender systems are designed to learn user interest by analyzing user-item interaction graphs. However, what they learn about user interest needs to be completed because historical interactions only record items that best match user interest (i.e., the first-order interest), while suboptimal items are absent. To fully exploit user interest, we propose a Second-Order Interest Learning (SOIL) framework to retrieve second-order interest from unrecorded suboptimal items. In this framework, we build a user-item interaction graph augmented by second-order interest, an interest-aware item-item graph for the visual modality, and a similar graph for the textual modality. In our work, all three graphs are constructed from user-item interaction records and multimodal feature similarity. Similarly to other graph-based approaches, we apply graph convolutional networks to each of the three graphs to learn representations of users and items. To improve the exploitation of both first-order and second-order interest, we optimize the model by implementing contrastive learning modules for user and item representations at both the user-item and item-item levels. The proposed framework is evaluated on three real-world public datasets in online shopping scenarios. Experimental results verify that our method is able to significantly improve prediction performance. For instance, our method outperforms the previous state-of-the-art method MGCN by an average of 8.1% in terms of Recall@10. Code: https://github.com/TL-UESTC/SOIL. Hongzu Su, Jingjing Li 0001, Fengling Li 0001, Ke Lu 0001, Lei Zhu 0002 |
ACM Multimedia | 3 |
| 2024 | Cross-Modal Retrieval: A Systematic Review of Methods and Future DirectionsabstractWith the exponential surge in diverse multimodal data, traditional unimodal retrieval methods struggle to meet the needs of users seeking access to data across various modalities. To address this, cross-modal retrieval has emerged, enabling interaction across modalities, facilitating semantic matching, and leveraging complementarity and consistency between heterogeneous data. Although prior literature has reviewed the field of cross-modal retrieval, it suffers from numerous deficiencies in terms of timeliness, taxonomy, and comprehensiveness. This article conducts a comprehensive review of cross-modal retrieval’s evolution, spanning from shallow statistical analysis techniques to vision-language pretraining (VLP) models. Commencing with a comprehensive taxonomy grounded in machine learning paradigms, mechanisms, and models, this article delves deeply into the principles and architectures underpinning existing cross-modal retrieval methods. Furthermore, it offers an overview of widely used benchmarks, metrics, and performances. Lastly, this article probes the prospects and challenges that confront contemporary cross-modal retrieval, while engaging in a discourse on potential directions for further progress in the field. To facilitate the ongoing research on cross-modal retrieval, we develop a user-friendly toolbox and an open-source repository athttps://cross-modal-retrieval.github.io. Tianshi Wang 0001, Fengling Li 0001, Lei Zhu 0002, Jingjing Li 0001, Zheng Zhang 0006, Heng Tao Shen |
Proc. IEEE | 2 |
| 2024 | Online Query Expansion Hashing for Efficient Image RetrievalabstractUnsupervised hashing has the desirable advantages of label independence, high storage, and retrieval efficiency, which is suitable for scalable image retrieval. Most existing methods focus on enhancing the image hashing model training process at the offline stage. However, little attention has been paid to the query content analysis by them at the online retrieval stage. They still suffer from important query semantic shortages, and thus limit the online retrieval performance, which is the ultimate objective of the image retrieval system. In this paper, we propose an Online Query Expansion Hashing (OQEH) for efficient image retrieval, by adaptively enhancing the discriminative capability of query hash codes in an expansion manner at the online retrieval stage. Specifically, we first design a self-expansion network to learn semantically invariant feature representations from images and their visual augmentations. Then, we conduct neighborhood-expansion to search similar samples for each image from a query expansion set with the semantically invariant features and design a Transformer architecture to adaptively transfer the semantics of neighbor samples to their corresponding images. With the support of semantically invariant features, query expansion set, and adaptive semantic transfer, the representation capability of query hash codes can be enhanced at the online retrieval stage. Experimental results demonstrate that the proposed OQEH method achieves superior retrieval accuracy and comparable retrieval efficiency compared with the state-of-the-art methods. Particularly, on MS COCO dataset, OQEH can obtain about 6% performance improvement compared with the state-of-the-art results. The source codes of our method are available at:https://github.com/christinecui/OQEH. Hui Cui 0004, Fengling Li 0001, Lei Zhu 0002, Jingjing Li 0001, Zheng Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Cross-Domain Transfer Hashing for Efficient Cross-Modal RetrievalabstractUnsupervised cross-modal hashing presents significant advantages in heterogeneous modality retrieval, offering label scalability, high retrieval efficiency, and low storage costs. However, the lack of explicit semantic supervision in this process results in a noticeable semantic deficit, impacting retrieval performance. In this paper, we address this challenge with a dual-pronged approach: Cross-Domain Transfer Hashing (CDTH), a lightweight weakly-supervised cross-modal hashing model. Our method leverages a semantically rich auxiliary domain to augment the target unsupervised cross-modal hash learning process. Simultaneously, we design a lightweight target cross-modal hashing network to reduce semantic requirements, lessening the burden of parameter optimization. Within the auxiliary domain, we perform direct semantic transfer with hashing network parameter transfer and indirect correlation semantic transfer by constructing an auxiliary semantic correlation graph with the identified cross-domain semantic consistent samples. In the target domain, we generate pseudo-labels using CLIP and establish a target weak semantic correlation graph. These two graphs collaborate to bolster the target cross-modal hashing training process. Extensive experiments on three publicly available datasets affirm the superiority of our approach in both retrieval accuracy and training efficiency. The source code for our method is accessible at: https://github.com/WangBowen7/CDTH. Fengling Li 0001, Lei Zhu 0002, Jingjing Li 0001, Zheng Zhang 0006, Xiaojun Chang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Temporal Social Graph Network Hashing for Efficient RecommendationabstractHashing-based recommender systems that represent users and items as binary hash codes are recently proposed to significantly improve time and space efficiency. However, the highly developed social media presents two major challenges to hashing-based recommendation algorithms. Firstly, the boundary between information producers and consumers becomes blurred, resulting in the rapid emergence of massive online content. Meanwhile, users' limited information consumption capacity inevitably causes further interaction sparsity. The inherent high sparsity of data leads to insufficient hash learning. Secondly, a considerable amount of online content becomes fast-moving consumer goods, such as short videos and news commentary, causing frequent changes in user interests and item popularity. To address the above problems, we propose a Temporal Social Graph Network Hashing (TSGNH) method for efficient recommendation, which generates binary hash codes of users and items through dynamic-adaptive aggregation on a constructed temporal social graph network. Specifically, we build a temporal social graph network to fully capture the social information widely existing in practical recommendation scenarios and propose a dynamic-adaptive aggregation method to capture long-term and short-term characters of users and items. Furthermore, different from the discrete optimization approaches used by existing hashing-based recommendation methods, we devise an end-to-end hashing learning approach that incorporates balanced and de-correlated constraints to learn compact and informative binary hash codes tailored for recommendation scenarios. Extensive experiments on three widely evaluated recommendation datasets demonstrate the superiority of the proposed method. Yang Xu 0025, Lei Zhu 0002, Jingjing Li 0001, Fengling Li 0001, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Invisible Black-Box Backdoor Attack against Deep Cross-Modal Hashing RetrievalabstractDeep cross-modal hashing has promoted the field of multi-modal retrieval due to its excellent efficiency and storage, but its vulnerability to backdoor attacks is rarely studied. Notably, current deep cross-modal hashing methods inevitably require large-scale training data, resulting in poisoned samples with imperceptible triggers that can easily be camouflaged into the training data to bury backdoors in the victim model. Nevertheless, existing backdoor attacks focus on the uni-modal vision domain, while the multi-modal gap and hash quantization weaken their attack performance. In addressing the aforementioned challenges, we undertake an invisible black-box backdoor attack against deep cross-modal hashing retrieval in this article. To the best of our knowledge, this is the first attempt in this research field. Specifically, we develop a flexible trigger generator to generate the attacker’s specified triggers, which learns the sample semantics of the non-poisoned modality to bridge the cross-modal attack gap. Then, we devise an input-aware injection network, which embeds the generated triggers into benign samples in the form of sample-specific stealth and realizes cross-modal semantic interaction between triggers and poisoned samples. Owing to the knowledge-agnostic of victim models, we enable any cross-modal hashing knockoff to facilitate the black-box backdoor attack and alleviate the attack weakening of hash quantization. Moreover, we propose a confusing perturbation and mask strategy to induce the high-performance victim models to focus on imperceptible triggers in poisoned samples. Extensive experiments on benchmark datasets demonstrate that our method has a state-of-the-art attack performance against deep cross-modal hashing retrieval. Besides, we investigate the influences of transferable attacks, few-shot poisoning, multi-modal poisoning, perceptibility, and potential defenses on backdoor attacks. Our codes and datasets are available at https://github.com/tswang0116/IB3A. Tianshi Wang 0001, Fengling Li 0001, Lei Zhu 0002, Jingjing Li 0001, Zheng Zhang 0006, Heng Tao Shen |
ACM Trans. Inf. Syst. | 2 |
| 2023 | Prototype-guided Knowledge Transfer for Federated Unsupervised Cross-modal HashingabstractAlthough deep cross-modal hashing methods have shown superiorities for cross-modal retrieval recently, there is a concern about potential data privacy leakage when training the models. Federated learning adopts a distributed machine learning strategy, which can collaboratively train models without leaking local private data. It is a promising technique to support privacy-preserving cross-modal hashing. However, existing federated learning-based cross-modal retrieval methods usually rely on a large number of semantic annotations, which limits the scalability of the retrieval models. Furthermore, they mostly update the global models by aggregating local model parameters, ignoring the differences in the quantity and category of multi-modal data from multiple clients. To address these issues, we propose a Prototype Transfer-based Federated Unsupervised Cross-modal Hashing(PT-FUCH) method for solving the privacy leakage problem in cross-modal retrieval model learning. PT-FUCH protects local private data by exploring unified global prototypes for different clients, without relying on any semantic annotations. Global prototypes are used to guide the local cross-modal hash learning and promote the alignment of the feature space, thereby alleviating the model bias caused by the difference in the distribution of local multi-modal data and improving the retrieval accuracy. Additionally, we design an adaptive cross-modal knowledge distillation to transfer valuable semantic knowledge from modal-specific global models to local prototype learning processes, reducing the risk of overfitting. Experimental results on three benchmark cross-modal retrieval datasets validate that our PT-FUCH method can achieve outstanding retrieval performance when trained under distributed privacy-preserving mode. The source codes of our method are available at https://github.com/exquisite1210/PT-FUCH_P. Jingzhi Li 0005, Fengling Li 0001, Lei Zhu 0002, Hui Cui 0004, Jingjing Li 0001 |
ACM Multimedia | 2 |
| 2023 | Task-Adversarial Adaptation for Multi-modal RecommendationabstractAn ideal multi-modal recommendation system is supposed to be timely updated with the latest modality information and interaction data because the distribution discrepancy between new data and historical data will lead to severe recommendation performance deterioration. However, upgrading a recommendation system with numerous new data consumes much time and computing resources. To mitigate this problem, we propose a Task-Adversarial Adaptation (TAA) framework, which is able to align data distributions and reduce resource consumption at the same time. This framework is specifically designed to align distributions of embedded features for different recommendation tasks between the source domain (i.e., historical data) and the target domain (i.e., new data). Technically, we design a domain feature discriminator for each task to distinguish which domain a feature comes from. By the two-player min-max game between the feature discriminator and the feature embedding network, the feature embedding network is able to align the source and target data distributions. With the ability to align source and target distributions, we are able to reduce the number of training samples by random sampling. In addition, we formulate the proposed approach as a plug-and-play module to accelerate the model training and improve the performance of mainstream multi-modal multi-task recommendation systems. We evaluate our method by predicting the Click-Through Rate (CTR) in e-commerce scenarios. Extensive experiments verify that our method is able to significantly improve prediction performance and accelerate model training on the target domain. For instance, our method is able to surpass the previous state-of-the-art method by 2.45% in terms of Area Under Curve (AUC) on AliExpress_US dataset while only utilizing one percent of the target data in training. Code: https://github.com/TL-UESTC/TAA. Hongzu Su, Jingjing Li 0001, Fengling Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002 |
ACM Multimedia | 3 |
| 2023 | Noise-Robust Continual Test-Time Domain AdaptationabstractContinual test-time domain adaptation (TTA) is a challenging topic in the field of source-free domain adaptation, which focuses on addressing cross-domain multimedia information during inference with a continuously changing data distribution. Previous methods have been found to lack noise robustness, leading to a significant increase in errors under strong noise. In this paper, we address the noise-robustness problem in continual TTA by offering three effective recipes to mitigate it. At the category level, we employ the Taylor cross-entropy loss to alleviate the low confidence category bias commonly associated with cross-entropy. At the sample level, we reweight the target samples based on uncertainty to prevent the model from overfitting on noisy samples. Finally, to reduce pseudo-label noise, we propose a soft ensemble negative learning mechanism to guide the model optimization using ensemble complementary pseudo labels. Our method achieves state-of-the-art performance on three widely used continual TTA datasets, particularly in the strong noise setting that we introduced. Zhiqi Yu, Jingjing Li 0001, Zhekai Du, Fengling Li 0001, Lei Zhu 0002, Yang Yang 0002 |
ACM Multimedia | 4 |
| 2023 | One for more: Structured Multi-Modal Hashing for multiple multimedia retrieval tasks
Chaoqun Zheng, Fengling Li 0001, Lei Zhu 0002, Zheng Zhang 0006, Wenpeng Lu |
Expert Syst. Appl. | 2 |
| 2021 | Task-adaptive Asymmetric Deep Cross-modal Hashing
Fengling Li 0001, Lei Zhu 0002, Zheng Zhang 0006, Xinhua Wang 0003 |
Knowl. Based Syst. | 1 |