EDBT 2026 Demo / reviewers in the wild / expert
Lei Zhu 0005
dblp:99/549-5
· DBLP profile ↗
27ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-4569-1429ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HAViG: Hierarchical adaptive visual grounding framework for video question answering
Lei Zhu 0005, Lingmin Pan, Siqiao Tan, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Farid Boussaïd, Mohammed Bennamoun |
Pattern Recognit. | 1 |
| 2026 | Multi-Modal Refined Prompting for Advancing Knowledge-Based Visual Question AnsweringabstractKnowledge-based Visual Question Answering (KB-VQA) has surfaced as a critical task in advancing AI capabilities. Despite significant progress enabled by large language models (LLMs), there are still three major challenges: (1) flawed image captions cause unreliable reasoning; (2) noisy explicit knowledge can disrupts answering; and (3) massive LLMs scale is irreplaceable to robustness. To overcome these challenges, we develop a novel approach, Multi-Modal Refined Prompting (MMRP), which generates high-quality prompts tailored for LLMs. To tackle the first challenge, a multi-faceted image captioning strategy is employed to generate detailed, contextually relevant visual descriptions. In addition, we introduce a complementary knowledge retrieval and refinement strategy to deliver concise, contextually relevant knowledge, effectively overcoming the second challenge. These enhanced image captions and explicit knowledge are then integrated into a knowledge-infused in-context prompt, effectively activating the reasoning capabilities of LLMs. Importantly, MMRP eliminates reliance on massive LLMs and avoids the need for model fine-tuning, while achieving significant improvements in answer accuracy. Extensive evaluations on the widely-used OK-VQA benchmark against 22 baselines prove the superiority of MMRP, establishing a new state-of-the-art in KB-VQA. Lei Zhu 0005, Mengxi Ying, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Shichao Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Textual semantics enhancement adversarial hashing for cross-modal retrieval
Lei Zhu 0005, Runbing Wu, Deyin Liu, Chengyuan Zhang 0001, Lin Wu 0001, Ying Zhang 0001, Shichao Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2025 | Information bottleneck-guided KNN contrastive hashing for unsupervised cross-modal retrievalabstractUnsupervised cross-modal hashing (UCMH) has emerged as a promising solution for scalable multi-modal retrieval without costly annotations. However, existing methods often rely on rigid pairwise contrastive learning and fixed-size neighborhood selection, which suffer from false negatives and semantic noise, respectively—limiting their ability to model complex semantic structures in open-world scenarios. In this paper, we propose a novel framework, I nformation B ottleneck-guided K NN C ontrastive H ashing ( IBKCH ), which introduces a flexible and semantically adaptive contrastive paradigm for UCMH. Specifically, we design an information-aware neighbor sampling strategy that integrates: (1) a Hard-negative and Soft-positive (HN-SP) mechanism to adaptively distinguish informative negatives and softly aggregate latent positives; (2) an information bottleneck loss to retain task-relevant semantics while suppressing redundancy; and (3) an entropy sparsity regularizer to mitigate noisy neighbor interference. Furthermore, we develop an adaptive KNN contrastive learning scheme that unifies intra-modal and inter-modal alignment, enabling robust and discriminative hash code learning. Extensive experiments on three benchmark datasets demonstrate that IBKCH consistently outperforms state-of-the-art methods, especially under noisy or semantically diverse conditions—highlighting its effectiveness and generalizability in real-world UCMH applications. Lei Zhu 0005, Zhengchang Yuan, Zeqian Yi, Chengyuan Zhang 0001, Lin Wu 0001, Ying Zhang 0001, Farid Boussaïd, Mohammed Bennamoun, Shichao Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2025 | Bi-Direction Label-Guided Semantic Enhancement for Cross-Modal HashingabstractSupervised cross-modal hashing has gained significant attention due to its efficiency in reducing storage and computation costs while maintaining rich semantic information. Despite substantial progress in generating compact binary codes, two key challenges remain: (1) insufficient utilization of labels to mine and fuse multi-grained semantic information, and (2) unreliable cross-modal interaction, which does not fully leverage multi-grained semantics or accurately capture sample relationships. To address these limitations, we propose a novel method called Bi-direction Label-Guided Semantic Enhancement for cross-modal Hashing (BiLGSEH). To tackle the first challenge, we introduce a label-guided semantic fusion strategy that extracts and integrates multi-grained semantic features guided by multi-labels. For the second challenge, we propose a semantic-enhanced relation aggregation strategy that constructs and aggregates multi-modal relational information through bi-directional similarity. Additionally, we incorporate CLIP features to improve the alignment between multi-modal content and complex semantics. In summary, BiLGSEH generates discriminative hash codes by effectively aligning semantic distribution and relational structure across modalities. Extensive performance evaluations against 18 competitive methods demonstrate the superiority of our approach. The source code for our method is publicly available at:https://github.com/yileicc/BiLGSEH. Lei Zhu 0005, Runbing Wu, Chengyuan Zhang 0001, Lin Wu 0001, Shichao Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | MvHAAN: multi-view hierarchical attention adversarial network for person re-identification
Lei Zhu 0005, Weiren Yu, Chengyuan Zhang 0001, Yangding Li, Shichao Zhang 0001 |
World Wide Web (WWW) | 1 |
| 2022 | MSSPQ: Multiple Semantic Structure-Preserving Quantization for Cross-Modal RetrievalabstractCross-modal hashing is a hot issue in the multimedia community, which is to generate compact hash code from multimedia content for efficient cross-modal search. Two challenges, i.e., (1) How to efficiently enhance cross-modal semantic mining is essential for cross-modal hash code learning, and (2) How to combine multiple semantic correlations learning to improve the semantic similarity preserving, cannot be ignored. To this end, this paper proposed a novel end-to-end cross-modal hashing approach, named Multiple Semantic Structure-Preserving Quantization (MSSPQ) that is to integrate deep hashing model with multiple semantic correlation learning to boost hash learning performance. The multiple semantic correlation learning consists of inter-modal and intra-modal pairwise correlation learning and Cosine correlation learning, which can comprehensively capture cross-modal consistent semantics and realize semantic similarity preserving. Extensive experiments are conducted on three multimedia datasets, which confirms that the proposed method outperforms the baselines. Lei Zhu 0005, Liewu Cai, Jiayu Song, Chengyuan Zhang 0001, Shichao Zhang 0001 |
ICMR | 1 |
| 2022 | Two-step learning for crowdsourcing data classification
Jiaye Li 0001, Zhaojiang Wu, Lei Zhu 0005 |
Multim. Tools Appl. | 5 |
| 2022 | CAESAR: concept augmentation based semantic representation for cross-modal retrieval
Lei Zhu 0005, Jiayu Song, Xiangxiang Wei |
Multim. Tools Appl. | 1 |
| 2022 | PPIS-JOIN: A Novel Privacy-Preserving Image Similarity Join Method
Chengyuan Zhang 0001, Fangxin Xie, Lei Zhu 0005, Yangding Li |
Neural Process. Lett. | 5 |
| 2022 | DAP2CMH: Deep Adversarial Privacy-Preserving Cross-Modal Hashing
Lei Zhu 0005, Jiayu Song, Zhan Yang 0001, Wenti Huang, Chengyuan Zhang 0001, Weiren Yu |
Neural Process. Lett. | 1 |
| 2022 | Multi-Graph Heterogeneous Interaction Fusion for Social RecommendationabstractWith the rapid development of online social recommendation system, substantial methods have been proposed. Unlike traditional recommendation system, social recommendation performs by integrating social relationship features, where there are two major challenges, i.e., early summarization and data sparsity. Thus far, they have not been solved effectively. In this article, we propose a novel social recommendation approach, namely Multi-Graph Heterogeneous Interaction Fusion (MG-HIF), to solve these two problems. Our basic idea is to fuse heterogeneous interaction features from multi-graphs, i.e., user–item bipartite graph and social relation network, to improve the vertex representation learning. A meta-path cross-fusion model is proposed to fuse multi-hop heterogeneous interaction features via discrete cross-correlations. Based on that, a social relation GAN is developed to explore latent friendships of each user. We further fuse representations from two graphs by a novel multi-graph information fusion strategy with attention mechanism. To the best of our knowledge, this is the first work to combine meta-path with social relation representation. To evaluate the performance of MG-HIF, we compare MG-HIF with seven states of the art over four benchmark datasets. The experimental results show that MG-HIF achieves better performance. Chengyuan Zhang 0001, Yang Wang 0023, Lei Zhu 0005, Jiayu Song, Hongzhi Yin |
ACM Trans. Inf. Syst. | 3 |
| 2021 | Multi-Graph Based Hierarchical Semantic Fusion for Cross-Modal RepresentationabstractThe main challenge of cross-modal retrieval is how to efficiently realize semantic alignment and reduce the heterogeneity gap. However, existing approaches ignore the multi-grained semantic knowledge learning from different modalities. To this end, this paper proposes a novel end-to-end cross-modal representation method, termed as Multi-Graph based Hierarchical Semantic Fusion (MG-HSF). This method is an integration of multi-graph hierarchical semantic fusion with cross-modal adversarial learning, which captures fine-grained and coarse-grained semantic knowledge from cross-modal samples, and generate modalities-invariant representations in a common subspace. To evaluate the performance, extensive experiments are conducted on three benchmarks. The experimental results show that our method is superior than the state-of-the-arts. Lei Zhu 0005, Chengyuan Zhang 0001, Jiayu Song, Shichao Zhang 0001, Yangding Li |
ICME | 1 |
| 2021 | M2GUDA: Multi-Metrics Graph-Based Unsupervised Domain Adaptation for Cross-Modal HashingabstractCross-modal hashing is a critical but very challenging task that is to retrieve similar samples of one modality via queries of other modalities. To improve the unsupervised cross-modal hashing, domain adaptation techniques can be used to support unsupervised hashing learning by transferring semantic knowledge from labeled source domain to unlabeled target domain. However, there are two problems that cannot be ignored: (1) most of domain adaptation based researches mainly focused on unimodal hashing or cross-modal real value-based retrieval but the study for cross-modal hashing is limited; (2) most existing studies only consider one or two consistency constraints during the domain adaptation learning. To this end, this paper propose a novel end-to-end framework to realize unsupervised domain adaptation for cross-modal hashing. This method, dubbed M$^2$GUDA, including four different consistency constraints: structure consistency, domain consistency, semantic consistency and modality consistency for domain adaptation learning. Besides, to enhance the structure consistency learning, we develop a multi-metrics graph modeling method to capture structure information comprehensively. Extensive experiments are performed on three common used benchmarks to evaluate the effectivity of our method. The results show that our method outperforms several state-of-the-art cross-modal hashing methods. Chengyuan Zhang 0001, Lei Zhu 0005, Shichao Zhang 0001, Da Cao |
ICMR | 3 |
| 2021 | NSDH: A Nonlinear Supervised Discrete Hashing framework for large-scale cross-modal retrieval
Zhan Yang 0001, Liu Yang 0015, Osolo Ian Raymond, Lei Zhu 0005, Wenti Huang, Zhifang Liao |
Knowl. Based Syst. | 4 |
| 2021 | HCMSL: Hybrid Cross-modal Similarity Learning for Cross-modal RetrievalabstractThe purpose of cross-modal retrieval is to find the relationship between different modal samples and to retrieve other modal samples with similar semantics by using a certain modal sample. As the data of different modalities presents heterogeneous low-level feature and semantic-related high-level features, the main problem of cross-modal retrieval is how to measure the similarity between different modalities. In this article, we present a novel cross-modal retrieval method, named Hybrid Cross-Modal Similarity Learning model (HCMSL for short). It aims to capture sufficient semantic information from both labeled and unlabeled cross-modal pairs and intra-modal pairs with same classification label. Specifically, a coupled deep fully connected networks are used to map cross-modal feature representations into a common subspace. Weight-sharing strategy is utilized between two branches of networks to diminish cross-modal heterogeneity. Furthermore, two Siamese CNN models are employed to learn intra-modal similarity from samples of same modality. Comprehensive experiments on real datasets clearly demonstrate that our proposed technique achieves substantial improvements over the state-of-the-art cross-modal retrieval techniques. Chengyuan Zhang 0001, Jiayu Song, Xiaofeng Zhu 0001, Lei Zhu 0005, Shichao Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Nonlinear Robust Discrete Hashing for Cross-Modal RetrievalabstractHashing techniques have recently been successfully applied to solve similarity search problems in the information retrieval field because of their significantly reduced storage and high-speed search capabilities. However, the hash codes learned from most recent cross-modal hashing methods lack the ability to comprehensively preserve adequate information, resulting in a less than desirable performance. To solve this limitation, we propose a novel method termed Nonlinear Robust Discrete Hashing (NRDH), for cross-modal retrieval. The main idea behind NRDH is motivated by the success of neural networks, i.e., nonlinear descriptors, in the field of representation learning, and the use of nonlinear descriptors instead of simple linear transformations is more in line with the complex relationships that exist between common latent representation and heterogeneous multimedia data in the real world. In NRDH, we first learn a common latent representation through nonlinear descriptors to encode complementary and consistent information from the features of the heterogeneous multimedia data. Moreover, an asymmetric learning scheme is proposed to correlate the learned hash codes with the common latent representation. Empirically, we demonstrate that NRDH is able to successfully generate a comprehensive common latent representation that significantly improves the quality of the learned hash codes. Then, NRDH adopts a linear learning strategy to fast learn the hash function with the learned hash codes. Extensive experiments performed on two benchmark datasets highlight the superiority of NRDH over several state-of-the-art methods. Zhan Yang 0001, Lei Zhu 0005, Wenti Huang |
SIGIR | 3 |
| 2020 | Scalable deep asymmetric hashing via unequal-dimensional embeddings for image similarity search
Zhan Yang 0001, Osolo Ian Raymond, Wenti Huang, Zhifang Liao, Lei Zhu 0005 |
Neurocomputing | 5 |
| 2020 | PAC-GAN: An effective pose augmentation scheme for unsupervised cross-view person re-identificationabstractPerson re-identification (person Re-Id) aims to retrieve the pedestrian images of the same person that captured by disjoint and non-overlapping cameras. Lots of researchers recently focused on this hot issue and proposed deep learning based methods to enhance the recognition rate in a supervised or unsupervised manner . However,there are two limitations that cannot be ignored: firstly, compared with other image retrieval benchmarks, the size of existing person Re-Id datasets is far from meeting the requirement, which cannot provide sufficient pedestrian samples for the training of deep model; secondly, the samples in existing datasets do not have sufficient human motions or postures coverage to provide more priori knowledges for learning. In this paper, we introduce a novel unsupervised pose augmentation cross-view person Re-Id scheme called PAC-GAN to overcome these limitations. We firstly present the formal definition of cross-view pose augmentation and then propose the framework of PAC-GAN that is a novel conditional generative adversarial network (CGAN) based approach to improve the performance of unsupervised corss-view person Re-Id. Specifically, the pose generation model in PAC-GAN called CPG-Net is to generate enough quantity of pose-rich samples from original image and skeleton samples. The pose augmentation dataset is produced by combining the synthesized pose-rich samples with the original samples, which is fed into the corss-view person Re-Id model named Cross-GAN. Besides, we use weight-sharing strategy in the CPG-Net to improve the quality of new generated samples. To the best of our knowledge, we are the first to enhance the unsupervised cross-view person Re-Id by pose augmentation, and the results of extensive experiments show that the proposed scheme can combat the state-of-the-arts with recognition rate. Chengyuan Zhang 0001, Lei Zhu 0005, Shichao Zhang 0001, Weiren Yu |
Neurocomputing | 2 |
| 2020 | TDHPPIR: An Efficient Deep Hashing Based Privacy-Preserving Image Retrieval Method
Chengyuan Zhang 0001, Lei Zhu 0005, Shichao Zhang 0001, Weiren Yu |
Neurocomputing | 2 |
| 2020 | Relation classification via knowledge graph enhanced transformer encoder
Wenti Huang, Yiyu Mao, Zhan Yang 0001, Lei Zhu 0005 |
Knowl. Based Syst. | 4 |
| 2019 | Efficient interactive search for geo-tagged multimedia data
Lei Zhu 0005, Chengyuan Zhang 0001, Zhan Yang 0001, Yunwu Lin, Ruipeng Chen |
Multim. Tools Appl. | 2 |
| 2019 | Efficient continuous top-k geo-image search on road network
Chengyuan Zhang 0001, Kesheng Cheng, Lei Zhu 0005, Ruipeng Chen, Zuping Zhang 0001, Fang Huang 0004 |
Multim. Tools Appl. | 3 |
| 2019 | Hierarchical information quadtree: efficient spatial temporal image search for multimedia stream
Chengyuan Zhang 0001, Ruipeng Chen, Lei Zhu 0005, Anfeng Liu, Yunwu Lin, Fang Huang 0004 |
Multim. Tools Appl. | 3 |
| 2019 | Hierarchical one permutation hashing: efficient multimedia near duplicate detection
Chengyuan Zhang 0001, Yunwu Lin, Lei Zhu 0005, Xinpan Yuan, Fang Huang 0004 |
Multim. Tools Appl. | 3 |
| 2019 | Efficient region of visual interests search for geo-multimedia data
Chengyuan Zhang 0001, Yunwu Lin, Lei Zhu 0005, Zuping Zhang 0001, Fang Huang 0004 |
Multim. Tools Appl. | 3 |
| 2019 | CNN-VWII: An efficient approach for large-scale video retrieval by image queries
Chengyuan Zhang 0001, Yunwu Lin, Lei Zhu 0005, Anfeng Liu, Zuping Zhang 0001, Fang Huang 0004 |
Pattern Recognit. Lett. | 3 |