Xu Lu 0004

dblp:53/6317-4 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-8459-3186ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Explicit semantic guided bi-incomplete multi-modal hashing with label co-occurrence and label graph constraints
Xu Lu 0004, Li Liu 0031, Huaxiang Zhang 0001
Neural Networks2
2025 Informative Scene Graph Generation via Debiasing
Lianli Gao, Xinyu Lyu, Yuyu Guo 0001, Yuan-Fang Li, Xu Lu 0004, Heng Tao Shen, Jingkuan Song
Int. J. Comput. Vis.6
2025 Incomplete Multi-Modal Weakly-Supervised Hashing With Consensus Bipartite Graph
abstract
Due to its excellent query and storage efficiency to facilitate large-scale multimedia retrieval, multi-modal hashing (MMH) has garnered a lot of attention from researchers. Nevertheless, existing MMH methods still suffer from several challenges: 1) Existing MMH methods often rely on graphs to represent complex correlation, but are constrained by the quality of graph construction and the storage overhead. 2) Existing MMH methods only deal with complete multi-modal data where all modalities of each instance are available, but cannot work with incomplete multi-modal data which encounter the problem of missing modalities. 3) Existing MMH methods often ignore the inevitable weak-supervision issue. To address these challenges, this paper proposes an Incomplete Multi-modal wEakly-supervised Hashing with Consensus Bipartite Graph (IMEH-CBG) method, which learns consensus bipartite graph for incomplete multi-modal fusion and corrects weak labels for discriminant hash learning. As far as we know, this is the first MMH method to work with incomplete and weakly-supervised multi-modal data in an unified framework. IMEH-CBG selects unified anchor set and builds consensus bipartite graph jointly for incomplete multi-modal fusion to tackle the first and the second challenges. Then, the semantic labels are predicted and utilized to learn hash code in an asymmetric way to tackle the third challenge. Extensive experiments demonstrate the superiority of IMEH-CBG.
Xu Lu 0004, Li Liu 0031, Huaxiang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Primary Code Guided Targeted Attack against Cross-modal Hashing Retrieval
abstract
Deep hashing algorithms have demonstrated considerable success in recent years, particularly in cross-modal retrieval tasks. Although hash-based cross-modal retrieval methods have demonstrated considerable efficacy, the vulnerability of deep networks to adversarial examples represents a significant challenge for the hash retrieval. In the absence of target semantics, previous non-targeted attack methods attempt to attack depth models by adding disturbance to the input data, yielding some positive outcomes. Nevertheless, they still lack specific instance-level hash codes and fail to consider the diversity and semantic association of different modalities, which is insufficient to meet the attacker's expectations. In response, we present a novel Primary code Guided Targeted Attack (PGTA) against cross-modal hashing retrieval. Specifically, we integrate cross-modal instances and labels to obtain well-fused target semantics, thereby enhancing cross-modal interaction. Secondly, the primary code is designed to generate discriminable information with fine-grained semantics for target labels. Benign samples and target semantics collectively generate adversarial examples under the guidance of primary codes, thereby enhancing the efficacy of targeted attacks. Extensive experiments demonstrate that our PGTA outperforms the most advanced methods on three datasets, achieving State-of-the-Art targeted attack performance.
Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007, Xu Lu 0004
IEEE Trans. Multim.5
2024 Joint-Modal Graph Convolutional Hashing for unsupervised cross-modal retrieval
Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007, Xu Lu 0004
Neurocomputing5
2024 Hypergraph clustering based multi-label cross-modal retrieval
Shengtang Guo, Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007, Xu Lu 0004, Liujian Li
J. Vis. Commun. Image Represent.5
2024 Multi-Facet Weighted Asymmetric Multi-Modal Hashing Based on Latent Semantic Distribution
abstract
With the advent of multi-modal data, multi-modal hashing has received increasing attention for it can configure complementary multi-modal fusion and support fast multimedia retrieval. Nevertheless, the “coarse-grained” modality weighting strategy widely used in existing methods always ignores the distinctive contributions of different features and is troubled by parameter adjustment. Besides, traditional supervised methods usually adopt “hard semantic” that reflects the logical relationship between data and labels, but fails to poring on the description degree of categories to data. To solve these problems, we propose amulti-Facet weIghting aSymmetric Multi-modal Hashing based on latent semantic distribution (FISMH)approach, which is divided into supervised paradigm SFISMH and unsupervised paradigm UFISMH. First, we design aMulti-facet Weighted Multi-modal Fusion modulethat utilizes both modality- and feature- wise weights to achieve multi-modal fusion, where the weight learning requires no additional parameter adjustment. Then, we design aLatent Semantic Distribution based Asymmetric Hash Learning module, which utilizes the pair- wise similarity and semantic distribution to guide hash learning, and avoids the challenging pair- wise factorization through asymmetric form. The semantic distribution is learned from the inherent information of feature space, which can further preserve the intra-class relationships. Finally, a discrete hash optimization is developed to reduce quantization and directly learn hash codes. The main difference between SFISMH and UFISMH is that the former utilizes category information while the latter explores the underlying data structure when constructing the pair- wise similarity. Extensive experiments demonstrate that both SFIMH and UFISMH outperform existing supervised and unsupervised multi-modal hashing methods, showcasing their exceptional performance.
Xu Lu 0004, Li Liu 0031, Lixin Ning, Shaomin Mu, Huaxiang Zhang 0001
IEEE Trans. Multim.1
2022 Efficient Multi-modal Hashing with Online Query Adaption for Multimedia Retrieval
abstract
Multi-modal hashing supports efficient multimedia retrieval well. However, existing methods still suffer from two problems: (1) Fixed multi-modal fusion. They collaborate the multi-modal features with fixed weights for hash learning, which cannot adaptively capture the variations of online streaming multimedia contents. (2) Binary optimization challenge. To generate binary hash codes, existing methods adopt either two-step relaxed optimization that causes significant quantization errors or direct discrete optimization that consumes considerable computation and storage cost. To address these problems, we first propose a Supervised Multi-modal Hashing with Online Query-adaption method. A self-weighted fusion strategy is designed to adaptively preserve the multi-modal features into hash codes by exploiting their complementarity. Besides, the hash codes are efficiently learned with the supervision of pair-wise semantic labels to enhance their discriminative capability while avoiding the challenging symmetric similarity matrix factorization. Further, we propose an efficient Unsupervised Multi-modal Hashing with Online Query-adaption method with an adaptive multi-modal quantization strategy. The hash codes are directly learned without the reliance on the specific objective formulations. Finally, in both methods, we design a parameter-free online hashing module to adaptively capture query variations at the online retrieval stage. Experiments validate the superiority of our proposed methods.
Lei Zhu 0002, Chaoqun Zheng, Xu Lu 0004, Zhiyong Cheng 0001, Liqiang Nie, Huaxiang Zhang 0001
ACM Trans. Inf. Syst.3
2021 From General to Specific: Informative Scene Graph Generation via Balance Adjustment
abstract
The scene graph generation (SGG) task aims to detect visual relationship triplets, i.e., subject, predicate, object, in an image, providing a structural vision layout for scene understanding. However, current models are stuck in common predicates, e.g., "on" and "at", rather than informative ones, e.g., "standing on" and "looking at", resulting in the loss of precise information and overall performance. If a model only uses "stone on road" rather than "blocking" to describe an image, it is easy to misunderstand the scene. We argue that this phenomenon is caused by two key imbalances between informative predicates and common ones, i.e., semantic space level imbalance and training sample level imbalance. To tackle this problem, we propose BA-SGG, a simple yet effective SGG framework based on balance adjustment but not the conventional distribution fitting. It integrates two components: Semantic Adjustment (SA) and Balanced Predicate Learning (BPL), respectively for adjusting these imbalances. Benefited from the model-agnostic process, our method is easily applied to the state-of-the-art SGG models and significantly improves the SGG performance. Our method achieves 14.3%, 8.0%, and 6.1% higher Mean Recall (mR) than that of the Transformer model at three scene graph generation sub-tasks on Visual Genome, respectively. Codes are publicly available1.
Yuyu Guo 0001, Lianli Gao, Xuanhan Wang, Xing Xu 0001, Xu Lu 0004, Heng Tao Shen, Jingkuan Song
ICCV6
2021 Graph Convolutional Multi-modal Hashing for Flexible Multimedia Retrieval
abstract
Multi-modal hashing makes an important contribution to multimedia retrieval, where a key challenge is to encode heterogeneous modalities into compact hash codes. To solve this dilemma, graph-based multi-modal hashing methods generally define individual affinity matrix of each independent modality and apply linear algorithm for heterogeneous modalities fusion and compact hash learning. Several other methods construct graph Laplacian matrix based on semantic information to help learn discriminative hash code. However, these conventional methods roughly ignore the structural similarity of training set and the complex relations among multi-modal samples, which leads to unsatisfactory complementarity of fused hash codes. More notably, they are faced with two other important problems: huge computing and storage costs caused by graph construction and partial modality feature lost problem when incomplete query sample comes. In this paper, we propose a Flexible Graph Convolutional Multi-modal Hashing (FGCMH) method that adopts GCNs with linear complexity to preserve both the modality-individual and modality-fused structural similarity for discriminative hash learning. Necessarily, accurate multimedia retrieval can be performed on complete and incomplete datasets with our method. Specifically, multiple modality-individual GCNs under semantic guidance are proposed to act on each individual modality independently for intra-modality similarity preserving, then the output representations are fused into a fusion graph with adaptive weighting scheme. Hash GCN and semantic GCN, which share parameters in the first two layers, propagate fusion information and generate hash codes under high-level label space supervision. In the query stage, our method adaptively captures various multi-modal contents in a flexible and robust way, even if partial modality features are lost. Experimental results on three publicly datasets show the flexibility and effectiveness of our proposed method.
Xu Lu 0004, Lei Zhu 0002, Li Liu 0031, Liqiang Nie, Huaxiang Zhang 0001
ACM Multimedia1
2021 Iterative graph attention memory network for cross-modal retrieval
Huaxiang Zhang 0001, Xu Lu 0004
Knowl. Based Syst.4
2021 Semantic-Driven Interpretable Deep Multi-Modal Hashing for Large-Scale Multimedia Retrieval
abstract
Multi-modal hashing focuses on fusing different modalities and exploring the complementarity of heterogeneous multi-modal data for compact hash learning. However, existing multi-modal hashing methods still suffer from several problems, including: 1) Almost all existing methods generate unexplainable hash codes. They roughly assume that the contribution of each hash code bit to the retrieval results is the same, ignoring the discriminative information embedded in hash learning and semantic similarity in hash retrieval. Moreover, the length of hash code is empirically set, which will cause bit redundancy and affect retrieval accuracy. 2) Most existing methods exploit shallow models which fail to fully capture higher-level correlation of multi-modal data. 3) Most existing methods adopt online hashing strategy based on immutable direct projection, which generates query codes for new samples without considering the differences of semantic categories. In this paper, we propose a Semantic-driven Interpretable Deep Multi-modal Hashing (SIDMH) method to generate interpretable hash codes driven by semantic categories within a deep hashing architecture, which can solve all these three problems in an integrated model. The main contributions are: 1) A novel deep multi-modal hashing network is developed to progressively extract hidden representations of heterogeneous modality features and deeply exploit the complementarity of multi-modal data. 2) Learning interpretable hash codes, with discriminant information of different categories distinctively embedded into hash codes and their different impacts on hash retrieval intuitively explained. Besides, the code length depends on the number of categories in the dataset, which can reduce the bit redundancy and improve the retrieval accuracy. 3) The semantic-driven online hashing strategy encodes the significant branches and discards the negligible branches of each query sample according to the semantics contained in it, therefore it could capture different semantics in dynamic queries. Finally, we consider both the nearest neighbor similarity and semantic similarity of hash codes. Experiments on several public multimedia retrieval datasets validate the superiority of the proposed method.
Xu Lu 0004, Li Liu 0031, Liqiang Nie, Xiaojun Chang, Huaxiang Zhang 0001
IEEE Trans. Multim.1
2020 Deep Collaborative Multi-View Hashing for Large-Scale Image Search
abstract
Hashing could significantly accelerate large-scale image search by transforming the high-dimensional features into binary Hamming space, where efficient similarity search can be achieved with very fast Hamming distance computation and extremely low storage cost. As an important branch of hashing methods, multi-view hashing takes advantages of multiple features from different views for binary hash learning. However, existing multi-view hashing methods are either based on shallow models which fail to fully capture the intrinsic correlations of heterogeneous views, or unsupervised deep models which suffer from insufficient semantics and cannot effectively exploit the complementarity of view features. In this paper, we propose a novel Deep Collaborative Multi-view Hashing (DCMVH) method to deeply fuse multi-view features and learn multi-view hash codes collaboratively under a deep architecture. DCMVH is a new deep multi-view hash learning framework. It mainly consists of 1) multiple view-specific networks to extract hidden representations of different views, and 2) a fusion network to learn multi-view fused hash code. DCMVH associates different layers with instance-wise and pair-wise semantic labels respectively. In this way, the discriminative capability of representation layers can be progressively enhanced and meanwhile the complementarity of different view features can be exploited effectively. Finally, we develop a fast discrete hash optimization method based on augmented Lagrangian multiplier to efficiently solve the binary hash codes. Experiments on public multi-view image search datasets demonstrate our approach achieves substantial performance improvement over state-of-the-art methods.
Lei Zhu 0002, Xu Lu 0004, Zhiyong Cheng 0001, Jingjing Li 0001, Huaxiang Zhang 0001
IEEE Trans. Image Process.2
2020 Flexible Multi-modal Hashing for Scalable Multimedia Retrieval
abstract
Multi-modal hashing methods could support efficient multimedia retrieval by combining multi-modal features for binary hash learning at the both offline training and online query stages. However, existing multi-modal methods cannot binarize the queries, when only one or part of modalities are provided. In this article, we propose a novel Flexible Multi-modal Hashing (FMH) method to address this problem. FMH learns multiple modality-specific hash codes and multi-modal collaborative hash codes simultaneously within a single model. The hash codes are flexibly generated according to the newly coming queries, which provide any one or combination of modality features. Besides, the hashing learning procedure is efficiently supervised by the pair-wise semantic matrix to enhance the discriminative capability. It could successfully avoid the challenging symmetric semantic matrix factorization and O ( n 2 ) storage cost of semantic matrix. Finally, we design a fast discrete optimization to learn hash codes directly with simple operations. Experiments validate the superiority of the proposed approach.
Lei Zhu 0002, Xu Lu 0004, Zhiyong Cheng 0001, Jingjing Li 0001, Huaxiang Zhang 0001
ACM Trans. Intell. Syst. Technol.2
2020 Fast Discrete Collaborative Multi-Modal Hashing for Large-Scale Multimedia Retrieval
abstract
Many achievements have been made on learning to hash for uni-modal and cross-modal retrieval. However, it is still an unsolved problem that how to directly and efficiently learn discriminative discrete hash codes for the multimedia retrieval, where both query and database samples are represented with heterogeneous multi-modal features. With this motivation, we propose a Fast Discrete Collaborative Multi-modal Hashing (FDCMH) method in this paper. We first propose an efficient collaborative multi-modal mapping that first transforms heterogeneous multi-modal features into the unified factors to exploit the complementarity of multi-modal features and preserve the semantic correlations in multiple modalities with linear computation and space complexity. Such shared factors also bridge the heterogeneous modality gap and remove the inter-modality redundancy. Further, we develop an asymmetric hashing learning module to simultaneously correlate the learned hash codes with low-level data distribution and high-level semantics. In particular, this design could avoid the challenging symmetric semantic matrix factorization and O(n2) memory cost (n is the number of training samples). It can support both computation and memory efficient discrete hash optimization. Experiments on several public multimedia retrieval datasets demonstrate the superiority of the proposed approach compared with state-of-the-art hashing techniques, in terms of both model learning efficiency and retrieval accuracy.
Chaoqun Zheng, Lei Zhu 0002, Xu Lu 0004, Jingjing Li 0001, Zhiyong Cheng 0001, Hanwang Zhang
IEEE Trans. Knowl. Data Eng.3
2020 Efficient Supervised Discrete Multi-View Hashing for Large-Scale Multimedia Search
abstract
Hashing has recently received substantial attention in large-scale multimedia search for its extremely low-cost storage cost and high retrieval efficiency. However, most existing hashing techniques focus on learning hash codes for single-view or cross-view retrieval. It is still an unsolved problem that how to efficiently learn discriminative binary codes for multi-view data that is common in real world multimedia search. In this paper, we propose an efficient Supervised Discrete Multi-view Hashing (SDMH) to solve the problem. SDMH first properly detects the shared binary hash codes, with an integrated multi-view feature mapping and latent hash coding, by exploiting the complementarity of different view-specific features and removing the involved inter-view redundancy. To further enhance the discriminative capability of hash codes, SDMH directly represses the explicit semantic labels of data samples with their corresponding binary codes. Different from most existing multi-view hashing methods that adopt “relaxing+rounding” hash optimization strategy or the discrete optimization method based on discrete cyclic coordinate descent, an efficient augmented Lagrangian multiplier (ALM) based discrete hash optimization method is developed in this paper to optimize the hash codes within a single step. Experimental results on four benchmark datasets demonstrate the superior performance of the proposed approach over state-of-the-art hashing techniques, in terms of both learning efficiency and retrieval accuracy.
Xu Lu 0004, Lei Zhu 0002, Jingjing Li 0001, Huaxiang Zhang 0001, Heng Tao Shen
IEEE Trans. Multim.1
2019 Flexible Online Multi-modal Hashing for Large-scale Multimedia Retrieval
abstract
Multi-modal hashing fuses multi-modal features at both offline training and online query stage for compact binary hash learning. It has aroused extensive attention in research filed of efficient large-scale multimedia retrieval. However, existing methods adopt batch-based learning scheme or unsupervised learning paradigm. They cannot efficiently handle the very common online streaming multi-modal data (for batch-learning methods), or learn the hash codes suffering from limited discriminative capability and less flexibility for varied streaming data (for existing online multi-modal hashing methods). In this paper, we develop a supervised Flexible Online Multi-modal Hashing (FOMH) method to adaptively fuse heterogeneous modalities and flexibly learn the discriminative hash code for the newly coming data, even if part of the modalities is missing. Specifically, instead of adopting the fixed weights, the modalities weights in FOMH are automatically learned with the proposed flexible multi-modal binary projection to timely capture the variations of streaming samples. Further, we design an efficient asymmetric online supervised hashing strategy to enhance the discriminative capability of the hash codes, while avoiding the challenging symmetric semantic matrix decomposition and storage cost. Moreover, to support fast hash updating and avoid the propagation of binary quantization errors in online learning process, we propose to directly update the hash codes with an efficient discrete online optimization. Experiments on several public multimedia retrieval datasets validate the superiority of the proposed method from various aspects.
Xu Lu 0004, Lei Zhu 0002, Zhiyong Cheng 0001, Jingjing Li 0001, Xiushan Nie, Huaxiang Zhang 0001
ACM Multimedia1
2019 Online Multi-modal Hashing with Dynamic Query-adaption
abstract
Multi-modal hashing is an effective technique to support large-scale multimedia retrieval, due to its capability of encoding heterogeneous multi-modal features into compact and similarity-preserving binary codes. Although great progress has been achieved so far, existing methods still suffer from several problems, including: 1) All existing methods simply adopt fixed modality combination weights in online hashing process to generate the query hash codes. This strategy cannot adaptively capture the variations of different queries. 2) They either suffer from insufficient semantics (for unsupervised methods) or require high computation and storage cost (for the supervised methods, which rely on pair-wise semantic matrix). 3) They solve the hash codes with relaxed optimization strategy or bit-by-bit discrete optimization, which results in significant quantization loss or consumes considerable computation time. To address the above limitations, in this paper, we propose an Online Multi-modal Hashing with Dynamic Query-adaption (OMH-DQ) method in a novel fashion. Specifically, a self-weighted fusion strategy is designed to adaptively preserve the multi-modal feature information into hash codes by exploiting their complementarity. The hash codes are learned with the supervision of pair-wise semantic labels to enhance their discriminative capability, while avoiding the challenging symmetric similarity matrix factorization. Under such learning framework, the binary hash codes can be directly obtained with efficient operations and without quantization errors. Accordingly, our method can benefit from the semantic labels, and simultaneously, avoid the high computation complexity. Moreover, to accurately capture the query variations, at the online retrieval stage, we design a parameter-free online hashing module which can adaptively learn the query hash codes according to the dynamic query contents. Extensive experiments demonstrate the state-of-the-art performance of the proposed approach from various aspects.
Xu Lu 0004, Lei Zhu 0002, Zhiyong Cheng 0001, Liqiang Nie, Huaxiang Zhang 0001
SIGIR1
2019 Efficient discrete latent semantic hashing for scalable cross-modal retrieval
Xu Lu 0004, Lei Zhu 0002, Zhiyong Cheng 0001, Xuemeng Song, Huaxiang Zhang 0001
Signal Process.1
2018 Discriminative correlation hashing for supervised cross-modal retrieval
Xu Lu 0004, Huaxiang Zhang 0001, Jiande Sun 0001, Zhenhua Wang 0004, Peilian Guo, Wenbo Wan
Signal Process. Image Commun.1