Xingbo Liu

dblp:10/8084 · DBLP profile ↗
← Back
42ranked-venue papers
13as first author
31since 2021 · last 2026
0000-0002-7236-8715ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 10 first-author · 20 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 PEOCH: Online Cross-Modal Hashing with Semi-Supervised Streaming Data Driving Prototype Evolution
abstract
The exponential growth of streaming multi-modal data presents critical challenges for cross-modal retrieval: distribution shifts, modality gap, and scarce labels. Semi-supervised online cross-modal hashing has gained increasing interest due to its ability to encode complex streaming data and update hash functions simultaneously. Nevertheless, existing methods can hardly generate high-quality unsupervised hash codes, which fundamentally limits diversity and flexibility during the retrieval process. To this end, we propose a novel method named Prototype Evolution Online Cross-modal Hashing (PEOCH). By driving prototype evolution with semi-supervised streaming data, precise and stable hash codes are generated for both labeled and unlabeled data. Specifically, two prototype updates with stability guarantee are conducted: labeled samples push semantic knowledge into the supervised prototypes, while unlabeled samples perform clustering to generate unsupervised prototypes. Simultaneously, a co-optimization mechanism is designed to ensure the prototypes continuously evolve and preserve the consistency of the entire streaming data. Besides, an elasticity regularizer integrates discriminability and smoothness constraints, improving the reliability of prototypes. Extensive experiments on three benchmark datasets demonstrate that PEOCH outperforms state-of-the-art methods, achieving an average improvement of 6.7% in mAP@all across various retrieval tasks.
Xiao Kang, Xingbo Liu, Shuo Pan, Xuening Zhang, Xiushan Nie, Yilong Yin
AAAI2
2026 MedCTM: A CNN-Transformer-Mamba Hybrid Network for Medical Image Classification
Weichao Pan, Jiaju Kang, Chengze Lv, Xuening Zhang, Xingbo Liu
Inf. Process. Manag.6
2026 SOR-BDNet: Semantic-Optical Representation for Boundary-Aware Video Anomaly Detection with GPT-4o
abstract
In recent years, Video Anomaly Detection (VAD) has shifted from conventional appearance-based modeling to semantically driven frameworks empowered by LLMs. Traditional reconstruction- and prediction-based methods, relying on motion or appearance patterns learned from normal data, often misclassify previously unseen yet semantically normal events as anomalies. To address this limitation, we propose SOR-BDNet (Semantic-Optical Representation with Boundary Detection Network), an annotation-free multimodal VAD framework that jointly leverages visual appearance and motion dynamics to generate interpretable semantic representations at the frame level. Specifically, we employ RAFT to estimate dense motion fields and concatenate the resulting flow maps with RGB images to form unified spatiotemporal inputs. These fused representations are fed into a GPT-4o-based module that generates semantic captions capturing object semantics and motion cues. Anomalies are detected by measuring semantic deviations from a memory bank constructed from normal captions. To further refine temporal boundaries, we design a boundary refinement module that integrates visual continuity constraints with contrastive feature learning based on a Swin Transformer backbone. Extensive experiments on four challenging benchmarks—UCSD-Ped2, Avenue, ShanghaiTech, and UCF-Crime—demonstrate that SOR-BDNet achieves frame-level accuracies of 97.96%, 82.86%, 87.36%, and 85.64%, respectively. These results highlight the robustness and scalability of the proposed framework, while significantly improving interpretability and generalization across diverse real-world surveillance scenarios. The source code and pretrained models are available at https://github.com/syi-coder/SOR-BDNet-Semantic-Optical-Representation-for-Boundary-Aware-Video-Anomaly-Detection-with-GPT-4o .
Bryan W. Scotney, Xiushan Nie, Xingbo Liu, Shuai Zhang 0001, Lanting Qiu
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Semi-Supervised Online Cross-Modal Hashing
abstract
Online cross-modal hashing has gained increasing interest due to its ability to encode streaming data and update hash functions simultaneously. Existing online methods often assume either fully supervised or completely unsupervised settings. However, they overlook the prevalent and challenging scenario of semi-supervised cross-modal streaming data, where diverse data types, including labeled/unlabeled, paired/unpaired, and multi-modal, are intertwined. To address this issue, we propose Semi-Supervised Online Cross-modal Hashing (SSOCH). It presents an alignment-free pseudo-labeling strategy that extracts semantic information from unlabeled streaming data without relying on pairing relations. Furthermore, we design an online tri-consistent preserving scheme, integrating pseudo-labeled data regularization, discriminative label embedding, and fine-grained similarity preservation. This scheme fully explores consistency across data annotation, modalities, and streaming chunks, improving the model's adaptiveness in these challenging scenarios. Extensive experiments on benchmark datasets demonstrate the superiority of SSOCH under various scenarios, highlighting the importance of semi-supervised learning for online cross-modal hashing.
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
AAAI2
2025 Generalized Debiased Semi-Supervised Hashing for Large-Scale Image Retrieval
abstract
Semi-supervised hashing has shown promising efficacy in large-scale image retrieval, which learns similarity-preserving codes from both labeled and unlabeled data. To enable the use of advanced supervised hashing techniques, pseudo labels are widely applied. However, existing methods typically suffer from a biased learning issue due to pseudo label noise, which can be further aggravated during optimization. Although such a bias can adversely affect hashing accuracy, it has not been investigated sufficiently. In view of this, we present a comprehensive discussion on potential causes of biases, involving processes of pseudo-labeling, hash learning and optimization. Accordingly, a novel Generalized Debiased Semi-supervised Hashing (GDSH) method is proposed as a unified solution to mitigate the biases. Specifically, reliable pseudo labels are first predicted via a robust label completion strategy. Secondly, a debiased hash learning module is designed by combining label denoising and similarity updating. This can not only refine the supervision, but also obtain hash codes that are semantically debiased in both category and sample levels. Finally, a discrete semi-supervised hashing algorithm is proposed to alleviate the bias arising from optimization. Experimental results on three single-label and three multi-label image benchmarks demonstrate that GDSH remarkably outperforms the state-of-the-arts in different semi-supervised settings.
Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
AAAI1
2025 Binary Continual Stream-View Clustering
abstract
Multi-view clustering is valued for uncovering latent common semantics lying in multi-view data, which has been a hot topic in unsupervised learning. However, when dealing with incremental streaming views, existing approaches typically require reconstructing the view data and aggregating streaming representations, leading to misalignment between representation and clusters. More importantly, conducting the clustering process frequently results in significant time consumption. To address these issues, we propose a novel method called Binary Continual Stream-View Clustering (BCSVC). Specifically, we design a continual clustering method that seamlessly unifies streaming representation learning and cluster assignment within a single framework. We also introduce a variance-weighted center updating mechanism to smooth the frequent clustering operation and absorb the semantics of previous views. In addition, to reduce the time and space expenditure on computation and storage, binary code for clustering representations is introduced, which can also significantly improve the computational efficiency of continuous updates in streaming scenarios. Last but not least, comprehensive theoretical analysis and extensive experimental results demonstrate its superior performance under various scenarios.
Xingbo Liu, Kang Xiao, Xuening Zhang, Xiushan Nie
ECAI2
2025 Supervised Discriminative Transformer Hashing for Large-Scale Remote Sensing Image Retrieval
abstract
With the advancement of remote sensing technology and the exponential growth of remote sensing visual data, efficiently retrieving remote sensing images from extensive databases has become increasingly important. Deep hashing, which combines the advantages of deep learning and hashing techniques, has emerged as a significant research direction in remote sensing image retrieval (RSIR). However, remote sensing images often contain substantial amounts of complex background information that are unrelated to the target. This noise can obscure or interfere with the target features, making it challenging for the model to effectively distinguish the target objects. To address these challenges, we propose a novel method called Supervised Discriminative Transformer Hashing (SDTH) for large-scale RSIR task, designed to enhance the retrieval of remote sensing images by extracting more distinguishable features. Our approach utilizes the Swin Transformer V2 architecture to improve feature extraction capabilities, thereby acquiring richer global contextual information and multi-scale features. To mitigate the impact of image noise and generate more discriminative hash codes, we propose to integrate batch-hard triplet loss with symmetric cross entropy loss. Experimental results on three benchmark datasets for remote sensing demonstrate the effectiveness and superiority of the proposed method.
Xingbo Liu, Xuening Zhang, Xiushan Nie
IEEE Geosci. Remote. Sens. Lett.2
2025 CTPT: Continual Test-time Prompt Tuning for vision-language models
Zhongyi Han, Xingbo Liu, Yilong Yin, Xin Gao 0001
Pattern Recognit.3
2025 Online Hashing with Discriminative Attribute Embedding
abstract
Online hashing has emerged as a powerful tool for efficiently processing large-scale and streaming data. However, existing approaches often struggle with scalability limitations in similarity relations and inadequate discrimination provided by one-hot labels. To address these challenges, we propose Online Hashing with Discriminative Attribute Embedding (OHDAE). This novel method leverages a triple-matrix decomposition framework to dynamically decompose features into a dictionary, attributes, and category representations, effectively capturing semantic consistency without relying on accumulated data. To enhance the consistency and discriminability of attributes, we introduce an attribute construction strategy that integrates dictionary constraints with an online optimization strategy. Additionally, fine-grained semantic labels are embedded to improve the discriminability of hash codes by incorporating both semantic and similarity relationships. Experiments conducted on three benchmark datasets validate the superior performance, scalability, and robustness of OHDAE compared to existing state-of-the-art methods.
Xingbo Liu, Zhijie Zhao, Xuening Zhang, Xiao Kang, Xiushan Nie
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Unsupervised Online Cross-modal Hashing With Multiple Association Exploitation
abstract
Unsupervised online cross-modal hashing has gained increasing attention for its effectiveness in streaming data retrieval. However, existing methods primarily focus on exploiting shared properties, overlooking semantic shifts among chunks and specific properties of each modality. To address these challenges, we propose a novel method called Unsupervised Online Cross-Modal Hashing with multiple association exploitation, UOCMH in short. Specifically, we design a hierarchical matrix factorization framework. It skillfully constructs robust orthogonal bases, multi-modality specific representations, and unified common representations, thereby capturing semantic associations among multi-modality streaming data more sufficiently. Additionally, we present a semantic auto-encoder scheme as hash functions. It builds the association between features and hash codes, facilitating the stability of the hashing process. Extensive experiments on the widely-used benchmark datasets demonstrate the superiority of the proposed UOCMH.
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
ICME2
2024 Fast Multi-view Clustering With Binary Anchor Graph
abstract
Multi-view clustering has achieved remarkable efficacy in integrating multi-view information, and received much research interest. Although anchor-based clustering algorithms have been well-investigated in past years, the separation of graph construction and category partitioning, can lead to suboptimal clustering performance and learning efficiency. To address these challenges, we propose a novel fast clustering algorithm named FAST-BAG. The proposed method can integrate the anchor graph construction and clustering partitioning seamlessly, breaking the separation between data fusion and task processes. Specifically, the multi-view data is unified into a consistent binary anchor graph with linear time complexity. Additionally, we leverage the high efficiency of binary distance computation to expedite the category partitioning process. Experiments conducted on five benchmark datasets validate the effectiveness and efficiency of the proposed method.
Xingbo Liu, Xiao Kang, Xuening Zhang, Xiushan Nie, Yilong Yin
ICME2
2024 Completely Unpaired Cross-Modal Hashing Based on Coupled Subspace
abstract
Unpaired cross-modal hashing which requires no supervision is a promising candidate to support large-scale retrieval across heterogeneous data. However, existing works focus on recovering pairwise relationships, which are usually time-consuming and sensitive to outliers. To tackle this issue, we propose a novel method termed Completely Unpaired Crossmodal Hashing (CUCH), which is applicable to scenarios where neither pairwise correspondence nor label information is available. The proposed CUCH creatively combines the merits of subspace recovery and cross-modal hashing, producing an effective subspace with both robustness and high efficiency. It first discovers robust subspace from each modality by excluding outliers. Then latent space translation is elaborated to obtain coupled subspace, based on which intermodal similarities can be captured. Moreover, the similarity-preservation property for CUCH is guaranteed. By manipulating subspaces rather than pairwise relations, CUCH reduces computational cost significantly. Experimental results demonstrate its advantages in various settings.
Xuening Zhang, Xingbo Liu, Xiao Kang, Xiushan Nie, Yilong Yin
ICME2
2024 Multi-Scale Temporal Relations and Segmented Channel Attention for Video Anomaly Detection
abstract
In recent years, the rapid advancement in video surveillance technology has significantly enhanced public safety and security. In conventional video anomaly detection approaches, there is often an exclusive focus on local information, with key temporal dynamics being overlooked. This oversight could potentially lead to a failure in recognizing dynamic anomalies, such as the sudden running of a person or the rapid movement of objects. Therefore, this study proposes a model framework structure called MTR-SCA. By utilizing widerresnet38 and Multi-Scale Temporal Relations (MTR) to capture the multi-scale temporal relationships in video time series, the framework achieves an understanding of spatial and temporal information. It introduces the Segmented Channel Attention (SCA) to enhance key information in the input feature maps and suppress less important channels for refined feature selection. We conducted experiments with the MTR-SCA network on three datasets: Avenue, ped2, and ShanghaiTech, achieving results of 97.8%, 86.8%, and 74.1% respectively.
Xiushan Nie, Bryan W. Scotney, Shuai Zhang 0001, Xingbo Liu
IJCNN5
2024 Generalized Universal Domain Adaptation
Wan Su, Zhongyi Han, Xingbo Liu, Yilong Yin
Knowl. Based Syst.3
2024 Discrete online cross-modal hashing with consistency preservation
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
Pattern Recognit.2
2024 Scalable Unsupervised Hashing via Exploiting Robust Cross-Modal Consistency
abstract
Unsupervised cross-modal hashing has received increasing attention because of its efficiency and scalability for large-scale data retrieval and analysis. However, existing unsupervised cross-modal hashing methods primarily focus on learning shared feature embedding, ignoring robustness and consistency across different modalities. To this end, this study proposes a novel method called scalable unsupervised hashing (SUH) for large-scale cross-modal retrieval. In the proposed method, latent semantic information and common semantic embedding within heterogeneous data are simultaneously exploited using multimodal clustering and collective matrix factorization, respectively. Furthermore, the robust norm is seamlessly integrated into the two processes, making SUH insensitive to outliers. Based on the robust consistency exploited from the latent semantic information and feature embedding, hash codes can be learned discretely to avoid cumulative quantitation loss. The experimental results on five benchmark datasets demonstrate the effectiveness of the proposed method under various scenarios.
Xingbo Liu, Jiamin Li 0003, Xiushan Nie, Xuening Zhang, Yilong Yin
IEEE Trans. Big Data1
2024 Online Discriminative Cross-Modal Hashing
abstract
Online cross-modal hashing has received increasing research attention due to its capability of encoding streaming data and updating hash functions simultaneously. Despite significant progress, there is still room for further improving accuracy from two aspects,i.e., 1) enhancing discrimination of hash codes with an efficient training process; 2) elevating generalization performance by harmonizing the training and retrieval process. Inspired by this, we propose an Online Discriminative Cross-modal Hashing method, called ODCH. To enlarge the inter-class margin and magnify the intra-class similarity, ODCH skillfully constructs a discriminative semantic space and seamlessly integrates bit balance and uncorrelation constraints, discrete optimization, and asymmetric strategy for embedding the discriminative semantic information into hamming space. Furthermore, ODCH attempts to boost the generalization process by bridging the gap between learning and generalization. It develops adaptive bit-wise weights to reflect different learning conditions among bits and transmits them into the generalization process. Besides, the proposed discriminative embedding and adaptive weighting can be adopted by existing supervised cross-modal hashing methods, achieving more precise performance than the original versions. Extensive experiments on three benchmarked datasets show that ODCH achieves up to an average of 4.17% mAP score gains compared to state-of-the-art online cross-modal hashing methods, indicating its superiority.
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.2
2024 Semi-Supervised Semi-Paired Cross-Modal Hashing
abstract
Large-scale cross-modal hashing has drawn extensive attention due to its attractive efficiency in both storage and retrieval. Existing methods exhibit poor performance when exploiting the semantic correlations implied in unsupervised and unpaired data during training process. To deal with this issue, we propose a novel hashing method, named Semi-supervised Semi-paired Cross-modal Hashing (SSCH). By leveraging a general and flexible two-step scheme, the proposed method can handle the complex training data effectively and efficiently, where both the common semantics and the modality-specific optimal pseudo semantics are well captured. Specifically, the proposed SSCH performs an alignment-free pseudo-labeling process to get strengthened semantic information. Furthermore, hash representations for various data are learned via a label-enhanced strategy, through which the cross-modal correlations are strengthened and preserved with considering efficiency. The semantic-preserving proof of SSCH is given based on statistical analysis. Also, we prove the stability of the proposed time-saving algorithm using properties of Bregman divergence. Experimental results on three benchmark datasets show that SSCH can obtain satisfactory precision and scalability in various scenarios.
Xuening Zhang, Xingbo Liu, Xiushan Nie, Xiao Kang, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.2
2024 Online Cross-modal Hashing With Dynamic Prototype
abstract
Online cross-modal hashing has received increasing attention due to its efficiency and effectiveness in handling cross-modal streaming data retrieval. Despite the promising performance, these methods mainly focus on the supervised learning paradigm, demanding expensive and laborious work to obtain clean annotated data. Existing unsupervised online hashing methods mostly struggle to construct instructive semantic correlations among data chunks, resulting in the forgetting of accumulated data distribution. To this end, we propose a Dynamic Prototype-based Online Cross-modal Hashing method, called DPOCH. Based on the pre-learned reliable common representations, DPOCH generates prototypes incrementally as sketches of accumulated data and updates them dynamically for adapting streaming data. Thereafter, the prototype-based semantic embedding and similarity graphs are designed to promote stability and generalization of the hashing process, thereby obtaining globally adaptive hash codes and hash functions. Experimental results on benchmarked datasets demonstrate that the proposed DPOCH outperforms state-of-the-art unsupervised online cross-modal hashing methods.
Xiao Kang, Xingbo Liu, Xiushan Nie, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Fast Unsupervised Cross-Modal Hashing with Robust Factorization and Dual Projection
abstract
Unsupervised hashing has attracted extensive attention in effectively and efficiently tackling large-scale cross-modal retrieval task. Existing methods typically try to mine the latent common subspace across multimodal data without any category annotation. Despite the exciting progress, there are still three challenges that need to be further addressed: (1) efficiently improving the robustness during latent common subspace learning; (2) harmoniously embedding the intra-modal inherence and inter-modal relevance of multimodal data into Hamming space; and (3) effectively reducing the training time complexity and making the model scalable for large-scale datasets. To well address the above challenges, this study proposes a method named Fast Unsupervised Cross-Modal Hashing (FUCH). Specifically, FUCH proposes a semantic-aware collective matrix factorization to learn robust representation via exploiting latent category-specific attributes, and introduces Cauchy loss to measure the factorization process. Accordingly, the above process can effectively embed potential discriminative information into common space, while making the model insensitive for outliers. Moreover, FUCH designs a dual projection learning scheme, which not only learns modality-unique hash functions to excavate individual properties but also learns modality-mutual hash functions to multimodal correlational properties. Experimental results on three benchmark datasets verify the effectiveness of FUCH under various scenarios.
Xingbo Liu, Jiamin Li 0003, Xiushan Nie, Xuening Zhang, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Supervised Discrete Multiple-Length Hashing for Image Retrieval
abstract
Hashing can facilitate efficient retrieval and storage for large-scale images due to the binary representation. In the real applications, the trade-off between retrieval accuracy and speed is essential for designing a hashing framework, which is reflected by variable hash code lengths. In light of this, the existing hashing methods need to train different models for different lengths of hash codes, leading to considerable training time cost and hashing flexibility reduction. Given that a sample can be represented by various hash codes with different lengths, there are some helpful relationships that can boost the performance of hashing methods. However, the existing hashing methods do not fully utilize these relationships. To address the aforementioned issues, we propose a new model, known as supervised discrete multiple-length hashing (SDMLH), to simultaneously learn hash codes with multiple lengths. In this proposed SDMLH method, three types of information are respectively derived, from the hash codes with different lengths. The original features of the samples, and the label, are applied for hash learning. Unlike the existing hashing methods, SDMLH can fully employ the assistance among hash codes with different lengths and learn them in one step. Furthermore, given a hash length meeting the demand of users, we propose a hash fusion strategy to obtain the hash code with this desirable length by fusing the multiple-length hash codes. This obtained hash code outperforms the one learned directly. In addition, SDMLH can generate the hash code of any length that is shorter than the sum length of given multiple hash codes with the fusion strategy. To the best of our knowledge, SDMLH is one of the first attempts for learning multiple-length hash codes simultaneously. We conduct extensive experiments based on three benchmark datasets, demonstrating the superiority of this proposed method.
Xiushan Nie, Xingbo Liu, Jie Guo 0012, Yilong Yin
IEEE Trans. Big Data2
2023 Zero-Shot Hashing via Asymmetric Ratio Similarity Matrix
abstract
Zero-shot hashing targets to learn the hash codes of images in unseen classes based on the limited training data provided by seen classes. In zero-shot hashing, transferring the supervised knowledge, such as attributes and semantic relations, from seen classes to unseen ones is a widely employed method, where the performance is always subject to the ability to capture these supervised knowledge (which is always difficult to obtain). Therefore, in this study, we propose a new methodology for zero-shot hashing via an asymmetric ratio similarity matrix (ASZH), which only needs to calculate the semantic similarity among seen classes for hash learning. Specifically, we use an asymmetric ratio matrix in the similarity calculation to further explore the influence of similarity, where the values of positive weights for similar samples are not equivalent to those of negative ones for dissimilar samples. Additionally, a theoretical analysis regarding the utilization of an asymmetric ratio matrix is provided in this study. The experiments on three large benchmark datasets indicate that the proposed method achieves excellent performance than several state-of-the-art hashing methods.
Xiushan Nie, Xingbo Liu, Lu Yang 0005, Yilong Yin
IEEE Trans. Knowl. Data Eng.3
2022 Supervised discrete hashing for hamming space retrieval
Xiao Kang, Fasheng Liu, Xiushan Nie, Xingbo Liu
Pattern Recognit. Lett.5
2022 Supervised Adaptive Similarity Matrix Hashing
abstract
Compact hash codes can facilitate large-scale multimedia retrieval, significantly reducing storage and computation. Most hashing methods learn hash functions based on the data similarity matrix, which is predefined by supervised labels or a distance metric type. However, this predefined similarity matrix cannot accurately reflect the real similarity relationship among images, which results in poor retrieval performance of hashing methods, especially in multi-label datasets and zero-shot datasets that are highly dependent on similarity relationships. Toward this end, this study proposes a new supervised hashing method called supervised adaptive similarity matrix hashing (SASH) via feature-label space consistency. SASH not only learns the similarity matrix adaptively, but also extracts the label correlations by maintaining consistency between the feature and the label space. This correlation information is then used to optimize the similarity matrix. The experiments on three large normal benchmark datasets (including two multi-label datasets) and three large zero-shot benchmark datasets show that SASH has an excellent performance compared with several state-of-the-art techniques.
Xiushan Nie, Xingbo Liu, Yilong Yin
IEEE Trans. Image Process.3
2022 Learning Binary Semantic Embedding for Large-Scale Breast Histology Image Analysis
abstract
With the progress of clinical imaging innovation and machine learning, the computer-assisted diagnosis of breast histology images has attracted broad attention. Nonetheless, the use of computer-assisted diagnoses has been blocked due to the incomprehensibility of customary classification models. In view of this question, we propose a novel method for Learning Binary Semantic Embedding (LBSE). In this study, bit balance and uncorrela-tion constraints, double supervision, discrete optimization and asymmetric pairwise similarity are seamlessly integrated for learning binary semantic-preserving embedding. Moreover, a fusion-based strategy is carefully designed to handle the intractable problem of parameter setting, saving huge amounts of time for boundary tuning. Based on the above-mentioned proficient and effective embedding, classification and retrieval are simultaneously performed to give interpretable image-based deduction and model helped conclusions for breast histology images. Extensive experiments are conducted on three benchmark datasets to approve the predominance of LBSE in different situations.
Xingbo Liu, Xiao Kang, Xiushan Nie, Jie Guo 0012, Yilong Yin
IEEE J. Biomed. Health Informatics1
2021 Learning Binary Semantic Embedding for Breast Histology Image Classification and Retrieval
abstract
With the development of medical imaging technology and machine learning, the computer-assisted diagnosis has attracted extensive research attention, which can provide beneficial reference to pathologists. However, the exponential growth of medical images and uninterpretability of traditional classification models have hindered the applications of the computer-assisted diagnosis. To address this issues, we propose a novel method for Learning Binary Semantic Embedding (LBSE). Based on this efficient and effective embedding, classification and retrieval are performed to provide interpretable computer-assisted diagnosis for histology images. Furthermore, double supervision, bit uncorrelation and balance constraint, asymmetric strategy and discrete optimization are seamlessly integrated in the proposed method for learning binary embedding. Experiments conducted on three benchmark datasets validate the superiority of LBSE under various scenarios.
Xiao Kang, Xingbo Liu, Xiushan Nie, Yilong Yin
ICASSP2
2021 Deep Multiple Length Hashing via Multi-task Learning
abstract
Hashing can compress heterogeneous high-dimensional data into compact binary codes. For most existing hash methods, they first predetermine a fixed length for the hash code and then train the model based on this fixed length. However, when the task requirements change, these methods need to retrain the model for a new length of hash codes, which increases time cost. To address this issue, we propose a deep supervised hashing method, called deep multiple length hashing(DMLH), which can learn multiple length hash codes simultaneously based on a multi-task learning network. This proposed DMLH can well utilize the relationships with a hard parameter sharing-based multi-task network. Specifically, in DMLH, the multiple hash codes with different lengths are regarded as different views of the same sample. Furthermore, we introduce a type of mutual information loss to mine the association among hash codes of different lengths. Extensive experiments have indicated that DMLH outperforms most existing models, verifying its effectiveness.
Xiushan Nie, Xingbo Liu
MMAsia5
2021 Discrete hashing with triple supervision learning
Xiao Kang, Fasheng Liu, Xiushan Nie, Xingbo Liu
J. Vis. Commun. Image Represent.5
2021 Supervised discrete hashing through similarity learning
Xingbo Liu, Xiushan Nie
Multim. Tools Appl.2
2021 Reinforced Short-Length Hashing
abstract
Given that retrieval and storage have compelling efficiency, similarity-preserving hashing has been extensively employed to approximate nearest neighbor search in large-scale image retrieval. Hash codes that are extremely compact not only can further lower the storage cost, but also accelerate the retrieval speed. However, existing methods perform poorly in retrieval based on an extremely short-length hash code, which attributes to the weak ability of classification and poor distribution of hash bit. To tackle this issue, in this study, we propose a novel reinforced short-length hashing (RSLH). In particular, this proposed method applies the mutual reconstruction between the hash representation and semantic label to retain the semantic information. Furthermore, to enhance the accuracy of hash representation, a pairwise similarity matrix is designed to make a balance between accuracy and training expenditure on memory. Besides, we integrate a parameter boosting strategy to strengthen the precision with the consideration of bit balance and uncorrelation constraints. Extensive experiments on three large-scale image benchmarks demonstrate the superior performance of RSLH under various short-length hashing scenarios.
Xingbo Liu, Xiushan Nie, Qi Dai 0001, Yupan Huang, Li Lian, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.1
2021 Fast Unmediated Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing is for the purpose of compressing heterogeneous multi-modal data into compact binary codes for the cross-modal retrieval, where accuracy and efficiency are two primary issues. To achieve high accuracy and efficiency, we put forward a novel method named Fast Unmediated Hashing (FUH) for cross-modal retrieval. For this method, motivated by the fact that label vector is a natural binary representation of samples for retrieval, we directly learn the cross-modal hash codes from semantic labels without any intermediate representation. This will capture more relations among different modalities, and reduce the number of variables. However, directly learning hash codes from labels would weaken the discrimination of hash codes. To address this issue, double supervision involving label information and pairwise similarity is proposed to enhance the discrimination. In addition, to decrease the training time, we present a strategy to bypass the similarity matrix-related operation in each iteration of optimization, thus some other related terms can also be computed offline to lower training complexity. Compared to several state-of-the-art techniques on three public datasets, the experimental results have manifested the superiority of FUH concerning efficiency and accuracy.
Xiushan Nie, Xingbo Liu, Xiaoming Xi, Chenglong Li 0004, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.2
2020 Focusing on Detail: Deep Hashing Based on Multiple Region Details (Student Abstract)
abstract
Fast retrieval efficiency and high performance hashing, which aims to convert multimedia data into a set of short binary codes while preserving the similarity of the original data, has been widely studied in recent years. Majority of the existing deep supervised hashing methods only utilize the semantics of a whole image in learning hash codes, but ignore the local image details, which are important in hash learning. To fully utilize the detailed information, we propose a novel deep multi-region hashing (DMRH), which learns hash codes from local regions, and in which the final hash codes of the image are obtained by fusing the local hash codes corresponding to local regions. In addition, we propose a self-similarity loss term to address the imbalance problem (i.e., the number of dissimilar pairs is significantly more than that of the similar ones) of methods based on pairwise similarity.
Xiushan Nie, Xingbo Liu, Yilong Yin
AAAI4
2020 Deep Multi-Region Hashing
abstract
Hashing has been widely used for large-scale approximate nearest neighbors retrieval own to its high efficiency. In the existing hashing methods, deep supervised hashing methods have achieved the best performance by utilizing the semantic labels on data with deep learning. However, most of these methods only consider the semantics of whole image but ignore the local information which contains much more semantic details. Evidently, the semantic details are beneficial for hash learning. To address this issue, in this paper, we proposed a novel Deep Multi-Region Hashing (DMRH) method to fully utilize the semantic details, which uses overlapping N × N regions of an image to learn N2hash codes for getting a final hash code. Extensive experimental results with three datasets show that DMRH can achieve state-of-the-art performance.
Xiushan Nie, Xingbo Liu, Yilong Yin
ICASSP4
2020 Modality correlation-based video summarization
Xingrun Wang, Xiushan Nie, Xingbo Liu, Binze Wang, Yilong Yin
Multim. Tools Appl.3
2020 Model Optimization Boosting Framework for Linear Model Hash Learning
abstract
Efficient hashing techniques have attracted extensive research interests in both storage and retrieval of highdimensional data, such as images and videos. In existing hashing methods, a linear model is commonly utilized owing to its efficiency. To obtain better accuracy, linear-based hashing methods focus on designing a generalized linear objective function with different constraints or penalty terms that consider the inherent characteristics and neighborhood information of samples. Differing from existing hashing methods, in this study, we propose a self-improvement framework called Model Boost (MoBoost) to improve model parameter optimization for linear-based hashing methods without adding new constraints or penalty terms. In the proposed MoBoost, for a linear-based hashing method, we first repeatedly execute the hashing method to obtain several hash codes to training samples. Then, utilizing two novel fusion strategies, these codes are fused into a single set. We also propose two new criteria to evaluate the goodness of hash bits during the fusion process. Based on the fused set of hash codes, we learn new parameters for the linear hash function that can significantly improve the accuracy. In general, the proposed MoBoost can be adopted by existing linear-based hashing methods, achieving more precise and stable performance compared to the original methods, and adopting the proposed MoBoost will incur negligible time and space costs. To evaluate the proposed MoBoost, we performed extensive experiments on four benchmark datasets, and the results demonstrate superior performance.
Xingbo Liu, Xiushan Nie, Liqiang Nie, Yilong Yin
IEEE Trans. Image Process.1
2019 Jointly Multiple Hash Learning
abstract
Hashing can compress heterogeneous high-dimensional data into compact binary codes while preserving the similarity to facilitate efficient retrieval and storage, and thus hashing has recently received much attention from information retrieval researchers. Most of the existing hashing methods first predefine a fixed length (e.g., 32, 64, or 128 bit) for the hash codes before learning them with this fixed length. However, one sample can be represented by various hash codes with different lengths, and thus there must be some associations and relationships among these different hash codes because they represent the same sample. Therefore, harnessing these relationships will boost the performance of hashing methods. Inspired by this possibility, in this study, we propose a new model jointly multiple hash learning (JMH), which can learn hash codes with multiple lengths simultaneously. In the proposed JMH method, three types of information are used for hash learning, which come from hash codes with different lengths, the original features of the samples and label. In contrast to the existing hashing methods, JMH can learn hash codes with different lengths in one step. Users can select appropriate hash codes for their retrieval tasks according to the requirements in terms of accuracy and complexity. To the best of our knowledge, JMH is one of the first attempts to learn multi-length hash codes simultaneously. In addition, in the proposed model, discrete and closed-form solutions for variables can be obtained by cyclic coordinate descent, thereby making the proposed model much faster during training. Extensive experiments were performed based on three benchmark datasets and the results demonstrated the superior performance of the proposed method.
Xingbo Liu, Xiushan Nie, Yingxin Wang, Yilong Yin
AAAI1
2019 MoBoost: A Self-improvement Framework for Linear-based Hashing
abstract
The linear model is commonly utilized in hashing methods owing to its efficiency. To obtain better accuracy, linear-based hashing methods focus on designing a generalized linear objective function with different constraints or penalty terms that consider neighborhood information. In this study, we propose a novel generalized framework called Model Boost (MoBoost), which can achieve the self-improvement of the linear-based hashing. The proposed MoBoost is used to improve model parameter optimization for linear-based hashing methods without adding new constraints or penalty terms. In the proposed MoBoost, given a linear-based hashing method, we first execute the method several times to get several different hash codes for training samples, and then combine these different hash codes into one set utilizing one novel fusion strategy. Based on this set of hash codes, we learn some new parameters for the linear hash function that can significantly improve accuracy. The proposed MoBoost can be generally adopted in existing linear-based hashing methods, achieving more precise and stable performance compared to the original methods while imposing negligible added expenditure in terms of time and space. Extensive experiments are performed based on three benchmark datasets, and the results demonstrate the superior performance of the proposed framework.
Xingbo Liu, Xiushan Nie, Xiaoming Xi, Lei Zhu 0002, Yilong Yin
CIKM1
2019 Supervised Short-Length Hashing
abstract
Hashing can compress high-dimensional data into compact binary codes, while preserving the similarity, to facilitate efficient retrieval and storage. However, when retrieving using an extremely short length hash code learned by the existing methods, the performance cannot be guaranteed because of severe information loss. To address this issue, in this study, we propose a novel supervised short-length hashing (SSLH). In this proposed SSLH, mutual reconstruction between the short-length hash codes and original features are performed to reduce semantic loss. Furthermore, to enhance the robustness and accuracy of the hash representation, a robust estimator term is added to fully utilize the label information. Extensive experiments conducted on four image benchmarks demonstrate the superior performance of the proposed SSLH with short-length hash codes. In addition, the proposed SSLH outperforms the existing methods, with long-length hash codes. To the best of our knowledge, this is the first linear-based hashing method that focuses on both short and long-length hash codes for maintaining high precision.
Xingbo Liu, Xiushan Nie, Xiaoming Xi, Lei Zhu 0002, Yilong Yin
IJCAI1
2019 Supervised Discrete Hashing With Mutual Linear Regression
abstract
Supervised linear hashing can compress high-dimensional data into compact binary codes owing to its efficiency. Generally, the relation between label and hash codes is widely used in the existing hashing methods because of its effectiveness of improving the accuracy. The existing hashing methods always use two different projections to represent the mutual regression between hash codes and class labels. In contrast to the existing methods, we propose a novel learning-based hashing method termed supervised discrete hashing with mutual linear regression (SDHMLR) in this study, where only one stable projection is used to describe the linear correlation between hash codes and corresponding labels. To the best of our knowledge, this strategy has not been used for hashing previously. In addition, we further use a boosting strategy to improve the final performance of the proposed method without adding extra constraints and with little extra expenditure in terms of time and space. Extensive experiments conducted on three image benchmarks demonstrate the superior performance of the proposed method.
Xingbo Liu, Xiushan Nie, Yilong Yin
ACM Multimedia1
2018 Modality-Specific Structure Preserving Hashing for Cross-Modal Retrieval
abstract
Hashing-based methods have made great advancements in cross-modal retrieval in both computational efficiency and storage. Learning a common space from different modalities is the common strategy of hashing-based methods, however, relational and structural information between samples in each modality, namely, a modality-specific structure, is always discarded during learning. In addition, cross-modality samples sometimes suffer from inter-class ambiguity and intra-class variability because of the uncertainty of manual labeling. To address these issues, we propose a novel method named Modality-specific structure Preserving Hashing (MsPH), which learns hashes by preserving the local structure and relations between samples in each modality. Moreover, label enhancement is utilized in MsPH to address label ambiguity and variability. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of MsPH under various cross-modal scenarios.
Xingbo Liu, Haoliang Sun, Xiushan Nie, Chaoran Cui, Yilong Yin
ICASSP1
2018 Fast Discrete Cross-modal Hashing With Regressing From Semantic Labels
abstract
Hashing has recently received great attention in cross-modal retrieval. Cross-modal retrieval aims at retrieving information across heterogeneous modalities (e.g., texts vs. images). Cross-modal hashing compresses heterogeneous high-dimensional data into compact binary codes with similarity preserving, which provides efficiency and facility in both retrieval and storage. In this study, we propose a novel fast discrete cross-modal hashing (FDCH) method with regressing from semantic labels to take advantage of supervised labels to improve retrieval performance. In contrast to existing methods that learn the projection from hash codes to semantic labels, the proposed FDCH regresses the semantic labels of training examples to the corresponding hash codes with a drift. It not only accelerates the hash learning process, but also helps generate stable hash codes. Furthermore, the drift can adjust the regression and enhance the discriminative capability of hash codes. Especially in the case of training efficiency, FDCH is much faster than existing methods. Comparisons with several state-of-the-art techniques on three benchmark datasets have demonstrated the superiority of FDCH under various cross-modal retrieval scenarios.
Xingbo Liu, Xiushan Nie, Wenjun Zeng 0001, Chaoran Cui, Lei Zhu 0002, Yilong Yin
ACM Multimedia1
2018 Cross-modal hashing based on category structure preserving
Xiushan Nie, Xingbo Liu, Leilei Geng
J. Vis. Commun. Image Represent.3