Li Liu 0031

dblp:33/4528-31 · DBLP profile ↗
← Back
70ranked-venue papers
2as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 2 first-author · 38 since 2021Artificial intelligence and machine learning · 26 · 19 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal Alignment
abstract
Cross-modal alignment is a promising yet challenging task in multimodal learning. Existing methods typically assess it by measuring the cross-modal semantic similarity from both global and local perspectives. However, these methods often neglect their potential interdependence. Specifically, global matching methods suffer from the over-compression of local features, while local matching methods rarely consider the inherent spatial topology of image patches. To address these limitations, we propose MG-Net, a unified framework with two collaborative modules: Multi-View Differential Mixer (MDM) and Graph-Guided Structural Region Selector (GSRS). The MDM is designed to capture discriminative global representations. It generates a series of views by decomposing feature vectors through multi-order differential operations, and adaptively fuses them via a lightweight Mixture-of-Experts (MoE) network. Meanwhile, the GSRS organizes image patches as a spatial graph and employs text-guided contextual reasoning to select spatially coherent and semantically complete structural regions. Extensive experiments on the Flickr30K and MS-COCO benchmarks demonstrate that the proposed MG-Net outperforms state-of-the-art methods in most cases.
Linlin Ji, Li Liu 0031
AAAI2
2026 Unsupervised multidimensional sample dynamic optimization for cross-modal hashing
Li Liu 0031, Huaxiang Zhang 0001, Dongmei Liu 0007
Appl. Intell.2
2026 CLIP-based knowledge projector for image-text matching
Dingwen Zhang, Longfei Han, Huaxiang Zhang 0001, Li Liu 0031, Junwei Han 0001
Inf. Process. Manag.5
2026 Lightweight target network with cross-domain multi-source knowledge fusion for cross-modal hashing
Li Liu 0031, Huaxiang Zhang 0001, Dongmei Liu 0007, Xiaochang Fang
Knowl. Based Syst.2
2026 Explicit semantic guided bi-incomplete multi-modal hashing with label co-occurrence and label graph constraints
Xu Lu 0004, Li Liu 0031, Huaxiang Zhang 0001
Neural Networks4
2026 DEANet : Adaptive RGB-T salient object detection with two-dimensional entropy-guided dual-domain feature interaction
Zerui Zhu, Dongmei Liu 0007, Huaxiang Zhang 0001, Li Liu 0031, Feng-Fei Jin
Signal Process.4
2026 Video and text semantic center alignment for text-video cross-modal retrieval
Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031
Signal Process. Image Commun.5
2026 VisualRAG: Knowledge-Guided Retrieval Augmentation for Image-Text Matching
abstract
Image-text matching as a fundamental cross-modal understanding task presents unique challenges in weakly-aligned scenarios. Such data typically feature highly abstract textual captions with sparse entity references, creating a significant semantic gap with visual content. Current mainstream methods, primarily designed for strongly aligned data pairs, employ dynamic modeling or multi-dimensional similarity computation to achieve feature space mapping. However, they struggle with information asymmetry and modal heterogeneity in weakly aligned cases. To address this, we propose a Visual Perception Knowledge Enhancement (VPKE) framework. Unlike existing methods based on strong alignment assumptions, this framework mines latent image semantics through vision-language models and generates auxiliary captions, overcoming the information bottleneck of traditional text modalities. Its core innovation lies in an adaptive knowledge distillation mechanism that combines retrieval-augmented generation (RAG) with key entity extraction. This mechanism effectively filters noise when introducing external knowledge while optimizing cross-modal feature integration. The framework employs multi-level similarity evaluation to dynamically adjust fusion weights among original text, key entities, and auxiliary captions, enabling adaptive integration of diverse semantic features and significantly improving model flexibility. Additionally, multi-scale feature extraction further enhances cross-modal representation capabilities. Experimental results show that the proposed method performs excellently in image-text retrieval tasks on the MSCOCO and Flickr30K datasets, validating its effectiveness.
Hengchang Wang, Li Liu 0031, Huaxiang Zhang 0001, Lei Zhu 0002, Xiaojun Chang
IEEE Trans. Circuits Syst. Video Technol.2
2026 SPENav: Dynamic Object Filtering With Spatial Perception Enhancement for Vision-Language Navigation
abstract
The Vision-language navigation task requires agents to efficiently interpret visual cues in the environment and accurately follow long-range instructions, posing significant challenges to their scene memory and spatial reasoning capabilities. Existing methods typically construct memory systems directly from raw visual observations. However, task-irrelevant cues commonly present in the environment can continuously introduce localization errors during navigation, severely limiting the agent’s performance in complex scenes. Meanwhile, due to the lack of transferable general knowledge priors, existing agents exhibit notable limitations in spatial perception, which undermines the reliability of their decision-making in unseen environments. To address these issues, this paper proposes the dynamic object filtering with Spatial Perception Enhancement for Vision-Language Navigation (SPENav), which aggregates open-vocabulary perception with multi-level information modeling. At the local level, the Hierarchical Semantic Prior Extractor and Room-Information-Guided Filtering construct task-oriented semantic priors to capture critical objects and suppress irrelevant features. At the global level, the Spatial-Instructional Guided Dual Attention module leverages spatial information and instruction guidance to enable the agent to develop selective memory that is goal- and task-oriented. On the unseen test split of R2R, SPENav achieves a 76% Success Rate (SR) and a 65% Success weighted by Path Length (SPL). These results demonstrate the effectiveness of task-oriented feature selection and multi-level semantic modeling in enhancing cross-modal understanding and adaptive navigation performance.
Huaxiang Zhang 0001, Li Liu 0031, Lei Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2025 Video Frame Enhancement based Text Semantic Fusion for Cross-modal Text-video Retrieval
Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007
ICMR3
2025 ViSSNet: RGB-T Salient Object Detection via Vision State-Space Network
Zerui Zhu, Dongmei Liu 0007, Huaxiang Zhang 0001, Li Liu 0031, Feng-Fei Jin
PRCV (8)4
2025 External information-augmented contrastive learning framework for fake news detection
Xiaochang Fang, Huaxiang Zhang 0001, Hongchen Wu, Li Liu 0031, Hongzhu Yu, Zhaorong Jing
Appl. Intell.4
2025 MSPL: Multi-granularity Semantic Prototype Learning for occluded person re-identification
Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031
Neurocomputing5
2025 MFCQA: Multi-Range Feature Cross-Attention Mechanism for no-reference image quality assessment
Nu Sun, Lili Meng, Weisi Lin, Li Liu 0031, Huaxiang Zhang 0001
Knowl. Based Syst.6
2025 Multi-Scale Feature Fusion Based on Piecewise Polynomial Activation Function for Image-Text Matching
abstract
Image-text matching remains challenging in big data processing. Matching accuracy is influenced by various factors, including the correlation between images and texts, feature extraction and fusion. Although activation functions play a crucial role in image-text matching, their design has received limited attention. This paper proposes an image-text matching model that utilizes multi-scale feature fusion based on a piecewise polynomial activation function. On one hand, a feature correlation optimization method is proposed to minimize the distance between paired images and texts. This method introduces large-scale downsampling odd-even feature embeddings and mean downsampling feature embeddings. After feature enhancement using a self-attention module, the odd-even feature embeddings are corrected with large-scale mean downsampling features to improve their representational ability. Additionally, a new multi-scale feature fusion method is utilized to enhance the robustness of the feature enhancement algorithm. On the other hand, we propose PCPAF (Piecewise Cubic Polynomial Activation Function), which offers advantages such as low computational cost,C1continuity, and superior generalizability. The PCPAF significantly improves model accuracy compared to existing activation functions. By adjusting the parameters of PCPAF, different activation functions can be derived, thereby improving matching accuracy in other image-text matching models. Experimental results on the Flickr30k and MS-COCO datasets demonstrate that the proposed model outperforms state-of-the-art models in terms of overall performance.
Linlin Ji, Li Liu 0031
IEEE Trans. Circuits Syst. Video Technol.2
2025 Heterogeneous Generative Tokens and Distance-Aware Recovery Network for Occluded Person Re-Identification
abstract
In real-world surveillance scenarios, person re-identification tasks are often seriously affected by occlusion problems, which requires the model to be able to not only extract powerful features, but also effectively recover features when they are occluded. Although existing methods disentangle visible human bodies by clustering semantic information, they often damage discriminative appearance due to the introduction of background noises. To solve this problem, we propose Heterogeneous Generative Tokens and Distance-aware Recovery (HGTDR) network, which aims to effectively extract discriminative appearance and recover the occluded body regions. HGTDR mainly contains two branches: a holistic stream and a part stream. The holistic stream utilizes ViT to capture the global context information and provide stable global features by establishing long-range relationships. In the part stream, we propose a Semantic Patch Generator (SPG), which combines the local attention mechanism to capture rich local semantics and further generate semantic patches. Further, considering the discrimination score and relevance score of semantic patches, we feed them into the proposed Adaptive Heterogeneous Semantic Token Generator (AHSTG) to gradually generate strong-response foreground and weak-response background features. In addition, to complete the features of occluded regions, the Distance-based Feature Recovery (DFR) module is designed. The module calculates the planar Euclidean distance of heterogeneous tokens and adaptively allocates the corresponding weights to dynamically recover the invisible bodies. Finally, we obtain discriminative and robust person descriptors. Extensive experiments on several challenging occluded, partial and holistic Re-ID datasets demonstrate that our proposed HGTDR network achieves superior performance and outperforms various state-of-the-art methods.
Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031
IEEE Trans. Circuits Syst. Video Technol.5
2025 Incomplete Multi-Modal Weakly-Supervised Hashing With Consensus Bipartite Graph
abstract
Due to its excellent query and storage efficiency to facilitate large-scale multimedia retrieval, multi-modal hashing (MMH) has garnered a lot of attention from researchers. Nevertheless, existing MMH methods still suffer from several challenges: 1) Existing MMH methods often rely on graphs to represent complex correlation, but are constrained by the quality of graph construction and the storage overhead. 2) Existing MMH methods only deal with complete multi-modal data where all modalities of each instance are available, but cannot work with incomplete multi-modal data which encounter the problem of missing modalities. 3) Existing MMH methods often ignore the inevitable weak-supervision issue. To address these challenges, this paper proposes an Incomplete Multi-modal wEakly-supervised Hashing with Consensus Bipartite Graph (IMEH-CBG) method, which learns consensus bipartite graph for incomplete multi-modal fusion and corrects weak labels for discriminant hash learning. As far as we know, this is the first MMH method to work with incomplete and weakly-supervised multi-modal data in an unified framework. IMEH-CBG selects unified anchor set and builds consensus bipartite graph jointly for incomplete multi-modal fusion to tackle the first and the second challenges. Then, the semantic labels are predicted and utilized to learn hash code in an asymmetric way to tackle the third challenge. Extensive experiments demonstrate the superiority of IMEH-CBG.
Xu Lu 0004, Li Liu 0031, Huaxiang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Primary Code Guided Targeted Attack against Cross-modal Hashing Retrieval
abstract
Deep hashing algorithms have demonstrated considerable success in recent years, particularly in cross-modal retrieval tasks. Although hash-based cross-modal retrieval methods have demonstrated considerable efficacy, the vulnerability of deep networks to adversarial examples represents a significant challenge for the hash retrieval. In the absence of target semantics, previous non-targeted attack methods attempt to attack depth models by adding disturbance to the input data, yielding some positive outcomes. Nevertheless, they still lack specific instance-level hash codes and fail to consider the diversity and semantic association of different modalities, which is insufficient to meet the attacker's expectations. In response, we present a novel Primary code Guided Targeted Attack (PGTA) against cross-modal hashing retrieval. Specifically, we integrate cross-modal instances and labels to obtain well-fused target semantics, thereby enhancing cross-modal interaction. Secondly, the primary code is designed to generate discriminable information with fine-grained semantics for target labels. Benign samples and target semantics collectively generate adversarial examples under the guidance of primary codes, thereby enhancing the efficacy of targeted attacks. Extensive experiments demonstrate that our PGTA outperforms the most advanced methods on three datasets, achieving State-of-the-Art targeted attack performance.
Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007, Xu Lu 0004
IEEE Trans. Multim.3
2025 SecureDA: Privacy-Preserving Source-Free Domain Adaptation for Person Re-Identification
abstract
Conventional domain adaptation (DA) for person re-identification (ReID) aims to bridge the domain gap but often requires direct use of fully labeled source and target domains, raising significant data privacy concerns due to the inclusion of personal identity information (PII) in raw data. Source-free domain adaptation (SFDA) for person ReID effectively preserves PII within the authorized source model. Nevertheless, these methods are vulnerable to data privacy (e.g., portrait rights) of the target domain during retrieval, where attackers can exploit pedestrian images for malicious generation, leading to damage to an individual’s reputation. Beyond these limitations, we propose a novel framework called SecureDA to address privacy-preserving SFDA for person ReID, which can generate a privacy key to defend against potential attacks on PII. Technically, we introduce domain-specific adversarial attacks into DA, where the protected query and gallery images are encrypted to ensure secure image retrieval. Furthermore, we employ two simultaneous processes: 1) The global–local adversarial pathway (GLAP) leverages encrypted and original images as adversarial pairs, thereby fostering the development of robust ReID models; 2) The global–local collaborative pathway (GLCP) is mastered through positive pairs collected from the same domain, effectively mitigating the pernicious catastrophic forgetting phenomenon. Extensive experiments show that SecureDA achieves state-of-the-art performance on multiple DA benchmarks and even outperforms the conventional DA and SFDA methods, which inherently compromise data privacy.
Xiaofeng Qu, Li Liu 0031, Huaxiang Zhang 0001, Lei Zhu 0002, Liqiang Nie, Xiaojun Chang, Fengling Li 0001
IEEE Trans. Multim.2
2024 A Unified Contrastive Framework with Multi-Granularity Fusion for Text-to-Image Generation
Yachao He, Li Liu 0031, Huaxiang Zhang 0001, Dongmei Liu 0007, Hongzhen Li
MMAsia2
2024 Self-similarity guided probabilistic embedding matching based on transformer for occluded person re-identification
Yunxiao Pang, Huaxiang Zhang 0001, Lei Zhu 0002, Dongmei Liu 0007, Li Liu 0031
Expert Syst. Appl.5
2024 Joint-Modal Graph Convolutional Hashing for unsupervised cross-modal retrieval
Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007, Xu Lu 0004
Neurocomputing3
2024 Accurate multi-view clustering to seek the cross-viewed yet uniform sample assignment via tensor feature matching
Yue Zhang 0045, Wuxiu Quan, Tatsuya Akutsu, Li Liu 0031, Hongmin Cai, Bin Zhang 0050
Inf. Sci.4
2024 Hypergraph clustering based multi-label cross-modal retrieval
Shengtang Guo, Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007, Xu Lu 0004, Liujian Li
J. Vis. Commun. Image Represent.3
2024 Source-free Style-diversity Adversarial Domain Adaptation with Privacy-preservation for person re-identification
Xiaofeng Qu, Li Liu 0031, Lei Zhu 0002, Liqiang Nie, Huaxiang Zhang 0001
Knowl. Based Syst.2
2024 Locally controllable network based on visual-linguistic relation alignment for text-to-image generation
Zaike Li, Li Liu 0031, Huaxiang Zhang 0001, Dongmei Liu 0007, Boqun Li
Multim. Syst.2
2024 Multi-Facet Weighted Asymmetric Multi-Modal Hashing Based on Latent Semantic Distribution
abstract
With the advent of multi-modal data, multi-modal hashing has received increasing attention for it can configure complementary multi-modal fusion and support fast multimedia retrieval. Nevertheless, the “coarse-grained” modality weighting strategy widely used in existing methods always ignores the distinctive contributions of different features and is troubled by parameter adjustment. Besides, traditional supervised methods usually adopt “hard semantic” that reflects the logical relationship between data and labels, but fails to poring on the description degree of categories to data. To solve these problems, we propose amulti-Facet weIghting aSymmetric Multi-modal Hashing based on latent semantic distribution (FISMH)approach, which is divided into supervised paradigm SFISMH and unsupervised paradigm UFISMH. First, we design aMulti-facet Weighted Multi-modal Fusion modulethat utilizes both modality- and feature- wise weights to achieve multi-modal fusion, where the weight learning requires no additional parameter adjustment. Then, we design aLatent Semantic Distribution based Asymmetric Hash Learning module, which utilizes the pair- wise similarity and semantic distribution to guide hash learning, and avoids the challenging pair- wise factorization through asymmetric form. The semantic distribution is learned from the inherent information of feature space, which can further preserve the intra-class relationships. Finally, a discrete hash optimization is developed to reduce quantization and directly learn hash codes. The main difference between SFISMH and UFISMH is that the former utilizes category information while the latter explores the underlying data structure when constructing the pair- wise similarity. Extensive experiments demonstrate that both SFIMH and UFISMH outperform existing supervised and unsupervised multi-modal hashing methods, showcasing their exceptional performance.
Xu Lu 0004, Li Liu 0031, Lixin Ning, Shaomin Mu, Huaxiang Zhang 0001
IEEE Trans. Multim.2
2024 AAMT: Adversarial Attack-Driven Mutual Teaching for Source-Free Domain-Adaptive Person Reidentification
abstract
Conventional domain adaptive (DA) methods for person re-identification (ReID) face knowledge transfer challenges when labeled data from the source domain cannot be accessed due to privacy constraints. Although the methods operating under source-absent DA settings attempt to address this challenge by using models pretrained on the source domain in their mutual teaching frameworks, failing to capture domain divergence in scenarios in which the source data are completely inaccessible can simultaneously introduce issues related to mutual convergence. In response, we introduce an adversarial attacks-driven mutual teaching (AAMT) framework as an innovative and applicable source-free DA person ReID scheme. Specifically, we first carefully develop a perturbation generator to generate source-style adversarial examples by leveraging a pretrained source model. Then, these diverse adversarial examples are employed to attack the mutual teaching model, implicitly measuring the domain divergence. Accordingly, we design a contrastive learning loss to enlarge the differences between the training pairs and further mitigate the mutual convergence issue. Extensive experiments demonstrate that AAMT outperforms the existing methods under both conventional and source-absent DA settings, achieving state-of-the-art performance.
Xiaofeng Qu, Huaxiang Zhang 0001, Lei Zhu 0002, Liqiang Nie, Li Liu 0031
IEEE Trans. Multim.5
2024 Instance-level Adversarial Source-free Domain Adaptive Person Re-identification
abstract
Domain adaption (DA) for person re-identification (ReID) has attained considerable progress by transferring knowledge from a source domain with labels to a target domain without labels. Nonetheless, most of the existing methods require access to source data, which raises privacy concerns. Source-free DA has recently emerged as a response to these privacy challenges, yet its direct application to open-set pedestrian re-identification tasks is hindered by the reliance on a shared category space in existing methods. Current source-free DA approaches for person ReID still encounter several obstacles, particularly the divergence-agnostic problem and the notable domain divergence due to the absent source data. In this article, we introduce an Instance-level Adversarial Mutual Teaching (IAMT) framework, which utilizes adversarial views to tackle the challenges mentioned above. Technically, we first elaborately develop a variance-based division (VBD) module to segregate the target data into instance-level subsets based on their similarity and dissimilarity to the source using the source-trained model, implicitly tackling the divergence-agnostic problem. To mitigate domain divergence, we additionally introduce a dynamic adversarial alignment (DAA) strategy, aiming to enhance the consistence of feature distribution across domains by employing adversarial instances from the target data to confuse the discriminators. Experiments reveal the superiority of the IAMT over state-of-the-art methods for DA person ReID tasks, while preserving the privacy of the source data.
Xiaofeng Qu, Li Liu 0031, Lei Zhu 0002, Liqiang Nie, Huaxiang Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Effective Occlusion Suppression Network via Grouped Pose Estimation for Occluded Person Re-Identification
abstract
The occluded target person often lead to incorrect matching results, so eliminating the interference of background clutter is a key matter. To deal with the challenging occluded person re-identification (Re-ID) tasks, this paper proposes a new grouped pose estimation occlusion generation (GPEOG) network to fully use the person grouped keypoint information and the stable non-occluded features, and make the extracted features more robust and discriminative. The local branch uses the proposed Dynamic Adaptive module (DAM) to automatically adjust the patch size during training process and provides adaptive local features. The occlusion-suppression generation branch mainly includes the proposed effective occlusion suppression (EOS) module. It judges whether the person’s keypoint groups are occluded, and then generates the masks of the corresponding parts to suppress occlusion. Extensive experiments show that the proposed method has obvious advantages compared with several state-of-the-art methods.
Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031
ICME5
2023 Giving Text More Imagination Space for Image-text Matching
abstract
Image-text matching is a hot topic in multi-modal analysis. The existing image-text matching algorithms focus on bridging the heterogeneity gap and mapping the feature into a common space under strong alignment assumption. However, these methods have unsatisfactory performance under the weak alignment scenario, which assumes that the text contains more abstract information, and the number of entities in the text is always fewer than objects in image. This is the first time, from our knowledge, to solve the image-text matching problem from the perspective of information difference with weak alignment. In order to both narrow the cross-modal heterogeneity gap and balance the information discrepancy, we proposed an imagination network to enrich the text modality based on pre-trained framework, which is helpful for image-text matching. The imagination network utilizes reinforcement learning to enhance the semantic information for text modality, and an action refinement strategy is designed to constrain the freedom and divergence of imagination. The experiment results show the superiority and generality of the proposed framework based on two pre-trained models, CLIP and BLIP on two most frequently-used datasets MSCOCO and Flickr30K.
Longfei Han, Dingwen Zhang, Li Liu 0031, Junwei Han 0001, Huaxiang Zhang 0001
ACM Multimedia4
2023 Generative adversarial network based on semantic consistency for text-to-image generation
Li Liu 0031, Huaxiang Zhang 0001, Chunjing Wang
Appl. Intell.2
2023 Weight grouping operators selection strategy for a multiobjective evolutionary algorithm based on decomposition
Yanyan Tan, Zeyuan Yan, Lili Meng, Li Liu 0031
Appl. Intell.5
2023 Feature generation based on relation learning and image partition for occluded person re-identification
Yunxiao Pang, Huaxiang Zhang 0001, Lei Zhu 0002, Dongmei Liu 0007, Li Liu 0031
J. Vis. Commun. Image Represent.5
2023 Semantic-embedding Guided Graph Network for cross-modal retrieval
Mengru Yuan, Huaxiang Zhang 0001, Dongmei Liu 0007, Lin Wang 0112, Li Liu 0031
J. Vis. Commun. Image Represent.5
2023 Attribute-aware style adaptation for person re-identification
Xiaofeng Qu, Li Liu 0031, Lei Zhu 0002, Huaxiang Zhang 0001
Multim. Syst.2
2023 Generative adversarial text-to-image generation with style image constraint
Li Liu 0031, Huaxiang Zhang 0001, Dongmei Liu 0007
Multim. Syst.2
2023 Multi-view unsupervised feature selection with tensor robust principal component analysis and consensus graph learning
abstract
Recently, multi-view unsupervised feature selection has attracted much attention due to its efficiency and better interpretability in processing high-dimensional multi-view datasets. Most existing methods rely on the constructed similarity matrices to obtain reliable pseudo labels to guide the feature selection. However, the considerable adverse noise in the raw data inevitably impedes the exploration of true underlying similarity structures. Besides, the inter-view correlations are often ignored during the common representation learning , which limits the effective fusion of the essential information from multiple views. To solve these issues, we design a novel robust multi-view unsupervised feature selection framework. Specifically, our method seeks a set of noise-free view-specific similarity matrices by leveraging tensor robust principal component analysis , where the high-order connections among different views are well exploited through the constructed low-rank tensor. Meanwhile, a high-quality consensus similarity matrix is adaptively learned from the view-specific representations within the same unified framework to capture the shared local structures. To enhance the discriminative ability of the feature selection matrix, we further impose a rank constraint on the consensus similarity matrix to obtain reliable pseudo cluster indicators. We present an efficient optimization algorithm ground on the alternating direction method of multipliers to solve the proposed model. Experimental results on six multi-view datasets confirm the superiority of our method.
Cheng Liang 0001, Lianzhi Wang, Li Liu 0031, Huaxiang Zhang 0001, Fei Guo 0001
Pattern Recognit.3
2023 Multi-label adversarial fine-grained cross-modal retrieval
Chunpu Sun, Huaxiang Zhang 0001, Li Liu 0031, Dongmei Liu 0007, Lin Wang 0112
Signal Process. Image Commun.3
2023 Multi-level adversarial attention cross-modal hashing
Benhui Wang, Huaxiang Zhang 0001, Lei Zhu 0002, Liqiang Nie, Li Liu 0031
Signal Process. Image Commun.5
2023 Video Sampled Frame Category Aggregation and Consistent Representation for Cross-Modal Retrieval
abstract
Many current video and text cross-modal retrieval research works focus on narrowing the semantic gap between video and text, but ignore the semantic difference between different sampled frames in the same video and the correlation of feature distribution of objects contained in different sampled frames in the same video, as a result, the features of the sampled frames in the final learned video cannot well represent the semantic features of the whole video. To overcome the shortcomings of existing studies, we first use a pre-trained video frame classification-aggregation network to make the object categories contained in different sampled frames in the same video be more close to the important object categories contained in the whole video, so as to promote the feature distribution of different sampled frames in the same video to be consistent, and increase the relevance of object features in different frames. Then we propose a video internal frame aggregation loss module to solve the problem of inconsistent feature distribution between different frame features encoded by video encoder in the same video and the aggregation feature of the sampled frame, thus enhancing the ability of video sampled frame aggregation feature representation. Experiments conducted on three common datasets MSVD, MSR-VTT and DiDeMo demonstrate the validity of the proposed approach.
Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031
IEEE Trans. Circuits Syst. Video Technol.5
2022 Group-pair deep feature learning for multi-view 3d model retrieval
Xiuxiu Chen, Li Liu 0031, Huaxiang Zhang 0001, Lili Meng, Dongmei Liu 0007
Appl. Intell.2
2022 Visible-infrared person re-identification based on key-point feature extraction and optimization
Li Liu 0031, Lei Zhu 0002, Huaxiang Zhang 0001
J. Vis. Commun. Image Represent.2
2022 HCFN: Hierarchical cross-modal shared feature network for visible-infrared person re-identification
Huaxiang Zhang 0001, Li Liu 0031
J. Vis. Commun. Image Represent.3
2022 Dual-path image pair joint discrimination for visible-infrared person re-identification
Li Liu 0031, Huaxiang Zhang 0001
J. Vis. Commun. Image Represent.2
2022 Coarse-to-fine dual-level attention for video-text cross modal retrieval
Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031
Knowl. Based Syst.5
2022 A dual-operator strategy for a multiobjective evolutionary algorithm based on decomposition
Zeyuan Yan, Yanyan Tan, Li Liu 0031, Huaxiang Zhang 0001
Knowl. Based Syst.4
2022 Adversarial Graph Convolutional Network for Cross-Modal Retrieval
abstract
The completeness of semantic expression plays an important role in cross-modal retrieval tasks, which contributes to align the cross-modal data and thus narrow the modality gap. But due to the abstractness of semantics, the same topic may have different aspects to be well described so it may be incomplete to express semantics with only one sample. In order to obtain semantic complementary information and strengthen similar information for samples with the same semantics, we utilize a graph convolutional network (GCN) to reconstruct the sample representation based on the adjacency relationship between the sample itself and its neighborhoods. We construct a local graph for each instance, and propose a novel Graph Feature Generator based on GCN and a fully-connected network to reconstruct node features based on local graph and map the features of two modalities into a common space. The Graph Feature Generator and Graph Feature Discriminator adopt a minimax game strategy to generate modality-invariant graph feature representations. Experiments on three benchmark datasets demonstrate the superiority of our proposed model compared with several state-of-the-art methods.
Li Liu 0031, Lei Zhu 0002, Liqiang Nie, Huaxiang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Hierarchical Feature Aggregation Based on Transformer for Image-Text Matching
abstract
In order to carry out more accurate retrieval across image-text modalities, some scholars use fine-grained feature to align image and text. Most of them directly use attention mechanism to align image regions and words in the sentence, and ignore the fact that semantics related to an object is abstract and cannot be accurately expressed by object information alone. To overcome this weakness, we propose a hierarchical feature aggregation algorithm based on graph convolutional networks (GCN) to facilitate object semantic integrity by integrating attributes of an object and relations between objects hierarchically in both image and text modalities. In order to eliminate the semantic gap between modalities, we propose a cross-modal feature fusion method based on transformer to generate modal-specific feature representations by integrating both the object feature and global feature from the other modality. Then we map the fusion feature into a common space. Experiment results on the most frequently-used datasets MSCOCO and Flickr30K show the effectiveness of the proposed model compared with the latest methods.
Huaxiang Zhang 0001, Lei Zhu 0002, Liqiang Nie, Li Liu 0031
IEEE Trans. Circuits Syst. Video Technol.5
2021 Graph Convolutional Multi-modal Hashing for Flexible Multimedia Retrieval
abstract
Multi-modal hashing makes an important contribution to multimedia retrieval, where a key challenge is to encode heterogeneous modalities into compact hash codes. To solve this dilemma, graph-based multi-modal hashing methods generally define individual affinity matrix of each independent modality and apply linear algorithm for heterogeneous modalities fusion and compact hash learning. Several other methods construct graph Laplacian matrix based on semantic information to help learn discriminative hash code. However, these conventional methods roughly ignore the structural similarity of training set and the complex relations among multi-modal samples, which leads to unsatisfactory complementarity of fused hash codes. More notably, they are faced with two other important problems: huge computing and storage costs caused by graph construction and partial modality feature lost problem when incomplete query sample comes. In this paper, we propose a Flexible Graph Convolutional Multi-modal Hashing (FGCMH) method that adopts GCNs with linear complexity to preserve both the modality-individual and modality-fused structural similarity for discriminative hash learning. Necessarily, accurate multimedia retrieval can be performed on complete and incomplete datasets with our method. Specifically, multiple modality-individual GCNs under semantic guidance are proposed to act on each individual modality independently for intra-modality similarity preserving, then the output representations are fused into a fusion graph with adaptive weighting scheme. Hash GCN and semantic GCN, which share parameters in the first two layers, propagate fusion information and generate hash codes under high-level label space supervision. In the query stage, our method adaptively captures various multi-modal contents in a flexible and robust way, even if partial modality features are lost. Experimental results on three publicly datasets show the flexibility and effectiveness of our proposed method.
Xu Lu 0004, Lei Zhu 0002, Li Liu 0031, Liqiang Nie, Huaxiang Zhang 0001
ACM Multimedia3
2021 PBNet: Position-specific Text-to-image Generation by Boundary
abstract
Most existing methods focus on improving the clarity and semantic consistency of the image with a given text, but do not pay attention to the multiple control of generated image content, such as the position of the object in generated image. In this paper, we introduce a novel position-based generative network (PBNet) which can generate fine-grained images with the object at the specified location. PBNet combines iterative structure with generative adversarial network (GAN). A location information embedding module (LIEM) is proposed to combine the location information extracted from the boundary block image with the semantic information extracted from the text. In addition, a silhouette generation module (SGM) is proposed to train the generator to generate object based on location information. The experimental results on CUB dataset demonstrate that PBNet effectively controls the location of the object in the generated image.
Li Liu 0031, Huaxiang Zhang 0001, Dongmei Liu 0007
MMAsia2
2021 Level set method with Retinex-corrected saliency embedded for image segmentation
abstract
Abstract It can be a very challenging task when using level set method segmenting natural images with high intensity inhomogeneity and complex background scenes. A new synthesis level set method for robust image segmentation based on the combination of Retinex‐corrected saliency region information and edge information is proposed in this work. First, the Retinex theory is introduced to correct the saliency information extraction. Second, the Retinex‐corrected saliency information is embedded into the level set method due to its advantageous quality which makes a foreground object stand out relative to the backgrounds. Combined with the edge information, the boundary of segmentation will be more precise and smooth. Experiments indicate that the proposed segmentation algorithm is efficient, fast, reliable, and robust.
Dongmei Liu 0007, Faliang Chang, Huaxiang Zhang 0001, Li Liu 0031
IET Image Process.4
2021 Label projection online hashing for balanced similarity
Yuzhi Fang, Huaxiang Zhang 0001, Li Liu 0031
J. Vis. Commun. Image Represent.3
2021 Dual-modality hard mining triplet-center loss for visible infrared person re-identification
Li Liu 0031, Lei Zhu 0002, Huaxiang Zhang 0001
Knowl. Based Syst.2
2021 Person re-identification based on multi-scale feature learning
Li Liu 0031, Lei Zhu 0002, Huaxiang Zhang 0001
Knowl. Based Syst.2
2021 CGNet: A Cascaded Generative Network for dense point cloud reconstruction from a single image
Li Liu 0031, Huaxiang Zhang 0001, Tianshi Wang 0001
Knowl. Based Syst.2
2021 Unsupervised Deep K-Means Hashing for Efficient Image Retrieval and Clustering
abstract
Recent studies show that hashing technology can achieve efficient similarity searching and many works have been done on supervised deep hash learning. However, under unsupervised scenarios, there are several issues to be solved when learning hashing codes based on visual features for image retrieval and clustering. In this article, we propose a simple but effective Unsupervised Deep K-means Hashing (UDKH) method to simultaneously alleviate the problems of image retrieval and clustering within a single learning framework. UDKH progressively improves the quality of cluster labels and binary hash codes by minimizing pair-wise supervision loss and optimizing the binary K-means to generate discriminative hash codes under the supervision of the learned cluster labels for effective image retrieval. Since the learned hash codes are discriminative, UDKH also improves the image clustering accuracy. Experiments on test datasets demonstrate its effectiveness for image retrieval and clustering.
Li Liu 0031, Lei Zhu 0002, Zhiyong Cheng 0001, Huaxiang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Semantic-Driven Interpretable Deep Multi-Modal Hashing for Large-Scale Multimedia Retrieval
abstract
Multi-modal hashing focuses on fusing different modalities and exploring the complementarity of heterogeneous multi-modal data for compact hash learning. However, existing multi-modal hashing methods still suffer from several problems, including: 1) Almost all existing methods generate unexplainable hash codes. They roughly assume that the contribution of each hash code bit to the retrieval results is the same, ignoring the discriminative information embedded in hash learning and semantic similarity in hash retrieval. Moreover, the length of hash code is empirically set, which will cause bit redundancy and affect retrieval accuracy. 2) Most existing methods exploit shallow models which fail to fully capture higher-level correlation of multi-modal data. 3) Most existing methods adopt online hashing strategy based on immutable direct projection, which generates query codes for new samples without considering the differences of semantic categories. In this paper, we propose a Semantic-driven Interpretable Deep Multi-modal Hashing (SIDMH) method to generate interpretable hash codes driven by semantic categories within a deep hashing architecture, which can solve all these three problems in an integrated model. The main contributions are: 1) A novel deep multi-modal hashing network is developed to progressively extract hidden representations of heterogeneous modality features and deeply exploit the complementarity of multi-modal data. 2) Learning interpretable hash codes, with discriminant information of different categories distinctively embedded into hash codes and their different impacts on hash retrieval intuitively explained. Besides, the code length depends on the number of categories in the dataset, which can reduce the bit redundancy and improve the retrieval accuracy. 3) The semantic-driven online hashing strategy encodes the significant branches and discards the negligible branches of each query sample according to the semantics contained in it, therefore it could capture different semantics in dynamic queries. Finally, we consider both the nearest neighbor similarity and semantic similarity of hash codes. Experiments on several public multimedia retrieval datasets validate the superiority of the proposed method.
Xu Lu 0004, Li Liu 0031, Liqiang Nie, Xiaojun Chang, Huaxiang Zhang 0001
IEEE Trans. Multim.2
2020 Consensus Achievement of Multi-agent Systems under Delayed State Information
abstract
In this paper, the consensus achievement for first-order multi-agent systems is researched over undirected graph. Suppose that the system is unstable and exactly knows the communication delay that affects the actually transmitted information, consensus gain in the protocol is designed with the delay information. Then, conditions are obtained to guarantee consensus for any large yet bounded communication delay. At the end of the article, an example is given to demonstrate the validity of the conclusion.
Yanli Zhu, Zhenhua Wang 0004, Li Liu 0031, Huaxiang Zhang 0001
ICARCV3
2020 A background-induced generative network with multi-level discriminator for text-to-image generation
abstract
Most existing text-to-image generation methods focus on synthesizing images using only text descriptions, but this cannot meet the requirement of generating desired objects with given backgrounds. In this paper, we propose a Background-induced Generative Network (BGNet) that combines attention mechanisms, background synthesis, and multi-level discriminator to generate realistic images with given backgrounds according to text descriptions. BGNet takes a multi-stage generation as the basic framework to generate fine-grained images and introduces a hybrid attention mechanism to capture the local semantic correlation between texts and images. To adjust the impact of the given backgrounds on the synthesized images, synthesis blocks are added at each stage of image generation, which appropriately combines the foreground objects generated by the text descriptions with the given background images. Besides, a multi-level discriminator and its corresponding loss function are proposed to optimize the synthesized images. The experimental results on the CUB bird dataset demonstrate the superiority of our method and its ability to generate realistic images with given backgrounds.
Li Liu 0031, Huaxiang Zhang 0001, Tianshi Wang 0001
MMAsia2
2020 A multi-label text classification method via dynamic semantic representation model and deep neural network
Tianshi Wang 0001, Li Liu 0031, Naiwen Liu, Huaxiang Zhang 0001
Appl. Intell.2
2020 Deep semantic cross modal hashing with correlation alignment
Meijia Zhang, Junzheng Li, Huaxiang Zhang 0001, Li Liu 0031
Neurocomputing4
2020 Cross-modal dual subspace learning with adversarial network
Fei Shang, Huaxiang Zhang 0001, Jiande Sun 0001, Liqiang Nie, Li Liu 0031
Neural Networks5
2020 Fuzzy weighted sparse reconstruction error-steered semi-supervised learning for face recognition
Li Liu 0031, Xiuxiu Chen, Tianshi Wang 0001
Vis. Comput.1
2019 Semantic consistency cross-modal dictionary learning with rank constraint
Fei Shang, Huaxiang Zhang 0001, Jiande Sun 0001, Li Liu 0031
J. Vis. Commun. Image Represent.4
2019 Detecting the latent associations hidden in multi-source information for better group recommendation
Huaxiang Zhang 0001, Lei Wang 0078, Li Liu 0031, Yuchang Xu
Knowl. Based Syst.4
2019 Hierarchical prediction based on two-level Gaussian mixture model clustering for bike-sharing system
Wenzhen Jia, Yanyan Tan, Li Liu 0031, Jing Li 0046, Huaxiang Zhang 0001, Kai Zhao 0011
Knowl. Based Syst.3
2019 Graph steered discriminative projections based on collaborative representation for Image recognition
Li Liu 0031, Bin Zhang 0050, Huaxiang Zhang 0001
Multim. Tools Appl.1
2018 Joint feature selection and graph regularization for modality-dependent cross-modal retrieval
Li Wang 0148, Lei Zhu 0002, Li Liu 0031, Jiande Sun 0001, Huaxiang Zhang 0001
J. Vis. Commun. Image Represent.4
2017 Discriminative Dictionary Learning With Ranking Metric Embedded for Person Re-Identification
abstract
The goal of person re-identification (Re-Id) is to match pedestrians captured from multiple non-overlapping cameras. In this paper, we propose a novel dictionary learning based method with the ranking metric embedded, for person Re-Id. A new and essential ranking graph Laplacian term is introduced, which minimizes the intra-personal compactness and maximizes the inter-personal dispersion in the objective. Different from the traditional dictionary learning based approaches and their extensions, which just use the same or not information, our proposed method can explore the ranking relationship among the person images, which is essential for such retrieval related tasks. Simultaneously, one distance measurement has been explicitly learned in the model to further improve the performance. Since we have reformulated these ranking constraints into the graph Laplacian form, the proposed method is easy-to-implement but effective. We conduct extensive experiments on three widely used person Re-Id benchmark datasets, and achieve state-of-the-art performances.
De Cheng, Xiaojun Chang, Li Liu 0031, Alex Hauptmann 0001, Yihong Gong, Nanning Zheng 0001
IJCAI3