Yongxin Wang 0001

dblp:95/9248-1 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
21since 2021 · last 2025
0000-0002-0172-9085ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Attraction Diminishing and Distributing for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel classes with limited samples after pre-training on a set of base classes. To avoid catastrophic forgetting and overfitting, most FSCIL methods first train the model on the base classes and then freeze the feature extractor in the incremental sessions. However, the reliance on nearest neighbor classification makes FSCIL prone to the hubness phenomenon, which negatively impacts performance in this dynamic and open scenario. While recent methods attempt to adapt to the dynamic and open nature of FSCIL, they are often limited to biased optimizations to the feature space. In this paper, we pioneer the theoretical analysis of the inherent hubness in FSCIL. To mitigate the negative effects of hubness, we propose a novel Attraction Diminishing and Distributing (D2A) method from the essential perspectives of distance metric and feature space. Extensive experimental results demonstrate that our method can broadly and significantly improve the performance of existing methods.
Li-Jun Zhao 0005, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
CVPR3
2025 Tag-Aware Weakly-Supervised Online Hashing with Enhanced Joint Representation
abstract
Weakly-supervised online hashing has garnered significant attention recently, yet several challenges remain unresolved, such as how to effectively denoise tags, and how to efficiently learn hash functions in dynamic online scenarios. To tackle these challenges, we propose a novel method named Tag-Aware Weakly-supervised Online Hashing with enhanced joint representation (TA-WOH). Our method creates an enhanced joint representation with CLIP-based features in order to reduce the tag noise. Additionally, we introduce a tag association and noise model for improved similarity matrix and hash code learning. A novel mapping mechanism is developed to align joint representations with the optimal tag space, enhancing both the accuracy and robustness of the model. The computational complexity of TA-WOH is dependent solely on the size of the incoming data, ensuring scalability and efficiency for large-scale datasets. Extensive experiments on two datasets demonstrate that our method surpasses several state-of-the-art methods in both accuracy and efficiency.
Yu-Wei Zhan, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
ICASSP4
2025 Evolving and Regularizing Meta-Environment Learner for Fine-Grained Few-Shot Class-Incremental Learning
abstract
Recently proposed Fine-Grained Few-Shot Class-Incremental Learning (FG-FSCIL) offers a practical and efficient solution for enabling models to incrementally learn new fine-grained categories under limited data conditions. However, existing methods still settle for the fine-grained feature extraction capabilities learned from the base classes. Unlike conventional datasets, fine-grained categories exhibit subtle inter-class variations, naturally fostering latent synergy among sub-categories. Meanwhile, the incremental learning framework offers an opportunity to progressively strengthen this synergy by incorporating new sub-category data over time. Motivated by this, we theoretically formulate the FSCIL problem and derive a generalization error bound within a shared fine-grained meta-category environment. Guided by our theoretical insights, we design a novel Meta-Environment Learner (MEL) for FG-FSCIL, which evolves fine-grained feature extraction to enhance meta-environment understanding and simultaneously regularizes hypothesis space complexity. Extensive experiments demonstrate that our method consistently and significantly outperforms existing approaches.
Li-Jun Zhao 0005, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
NeurIPS3
2025 OH-CMH: Towards cross-modal hashing for streaming data with hierarchical labels and label increment scenario
Chong-Yu Zhang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
Knowl. Based Syst.6
2025 BITS: Bit-Extendable Incremental Hashing in Open Environments
abstract
Hashing is an effective technique for large-scale image retrieval. However, traditional hashing models typically follow a closed-set assumption, which fails to satisfy the practicality of real-world tasks. In this paper, we explore a meaningful yet overlooked question: is there a hashing paradigm that not only supports rehearsal-free online incremental coding for single-pass data streams but also adapts to potentially expanding concept spaces in open environments? Instead of presetting fixed bit lengths, we suggest adjusting the bit length dynamically based on the number of encountered categories, meanwhile enabling bit extension of existing hash codes to match the adaptive code lengths without knowledge forgetting. Therefore, we propose a Bit-extendable IncremenTal haShing (BITS) method for image retrieval in open environments. Specifically, we identify a blurry incremental setup to better simulate realistic scenarios, revisiting the widely-used data-incremental and class-incremental settings. With this challenging setup, a three-phase framework is designed to efficiently perform incremental hashing, which jointly solves online continual coding and bit extension with adaptive code lengths. Through the well-designed hashing paradigm, BITS achieves comparable performance to offline hashing methods while significantly saving computational resources. Comprehensive experiments on six benchmarks demonstrate the superiority of our BITS in dynamic scenarios. The source code is available at https://github.com/yxinwang/BITS.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Image Process.1
2025 Domain-Aware Semantic Alignment Hashing for Large-Scale Zero-Shot Image Retrieval
abstract
Hashing has been proven to be effective in the field of large-scale image retrieval. However, traditional hashing is stuck in performance dilemmas under zero-shot scenarios due to the concept shift problem. Although some zero-shot hashing methods exploit category attributes to facilitate knowledge transfer across domains, they usually struggle to generate domain-adaptive hash codes, making it hard to distinguish samples between unknown and known classes. With this motivation, we propose a novel approach called Domain-Aware Semantic Alignment Zero-Shot Hashing (DSAZH), which reveals three issues that suppress performance: semantic misalignment, biased optimization, and ambiguous Hamming distance. To address these challenges, multiple initiatives are innovatively integrated into a unified framework: First, it generates semantic-aligned hash codes through class-level and instance-level semantic alignment ; then it learns unbiased hash codes and domain-adaptive hash function through unbiased optimization equipped with asymmetric processing and class-prompting regression; finally, it distinguishes seen instances from unseen using domain-aware thresholding . Extensive experiments show that DSAZH achieves up to 15.82% MAP improvement (e.g., 69.71% vs. 53.89% on large-scale ImageNet with 256-bit codes) while reducing training time by two orders of magnitude (e.g., 3.07 s vs. 202.84 s), demonstrating its superior accuracy and efficiency compared to state-of-the-art ZSH methods. The source code is available at https://github.com/yxinwang/DSAZH .
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2024 FedCAFE: Federated Cross-Modal Hashing with Adaptive Feature Enhancement
abstract
Deep Cross-Modal Hashing (CMH) has become one of the most popular solutions for cross-modal retrieval. Existing methods need to first collect data and then be trained with these accumulated data. However, in real world, data may be generated and possessed by different owners. Considering the concerns about privacy, data may not be shared or transmitted, leading to the failure of sufficient training of CMH. To solve the problem, we propose a new framework called Federated Cross-modal Hashing with Adaptive Feature Enhancement (FedCAFE). FedCAFE is a federated method which could use distributed data to train existing CMH methods under the privacy protection. To overcome the data heterogeneity challenge of distributed data and improve the generalization ability of global model, FedCAFE is endowed with a novel adaptive feature enhancement module and a new weighted aggregation strategy. Besides, it could fully utilize the rich global information carried in the global model to constrain the model during the local training process. We have conducted extensive experiments on four widely-used datasets in CMH domain with both IID and non-IID settings. The reported results demonstrate that the proposed FedCAFE achieves better performance than several state-of-the-art baselines.
Yu-Wei Zhan, Chong-Yu Zhang, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xun Yang 0001, Xin-Shun Xu
ACM Multimedia6
2024 POLISH: Adaptive Online Cross-Modal Hashing for Class Incremental Data
abstract
In recent years, hashing-based online cross-modal retrieval has garnered growing attention. This trend is motivated by the fact that web data is increasingly delivered in a streaming manner as opposed to batch processing. Simultaneously, the sheer scale of web data sometimes makes it impractical to fully load for the training of hashing models. Despite the evolution of online cross-modal hashing techniques, several challenges remain: 1) Most existing methods learn hash codes by considering the relevance among newly arriving data or between new data and the existing data, often disregarding valuable global semantic information. 2) A common but limiting assumption in many methods is that the label space remains constant, implying that all class labels should be provided within the first data chunk. This assumption does not hold in real-world scenarios, and the presence of new labels in incoming data chunks can severely degrade or even break these methods.
Yu-Wei Zhan, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Yinwei Wei, Xin-Shun Xu
WWW4
2024 Weighted cross-modal hashing with label enhancement
Yongxin Wang 0001, Kuikui Wang, Xiushan Nie, Zhen-Duo Chen 0001
Knowl. Based Syst.1
2024 A vision transformer for fine-grained classification by reducing noise and enhancing discriminative information
Zi-Chao Zhang 0002, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.3
2024 Multiple Information Embedded Hashing for Large-Scale Cross-Modal Retrieval
abstract
Recently, many efforts have been devoted to improving the retrieval performance of supervised cross-modal hashing; however, current methods are gradually reaching a performance bottleneck, especially when dealing with real-world multimedia data. This is mainly due to their application of coarse-grained semantics, unrobust hash functions, and inflexible workflows. Therefore, discovering refined semantics hidden in data, designing robust hash functions, and creating a non-interfering but facilitative learning workflow are much more significant. With this motivation, in this paper, we propose a novel supervised cross-modal hashing method, i.e., Multiple Information Embedded Hashing, MIEH for short. It consists of a three-step working flow that flexibly handles multiple information mining, hash code learning, and hash function learning. First, it explores the multimedia data from multiple perspectives such as modal-level consistency, class-level discriminability, and instance-level similarity to mine comprehensive semantic information, which not only contributes to the generation of discriminative hash codes, but also accelerates convergence. Subsequently, MIEH is committed to embed the refined semantics into targeted hash codes with an efficient discrete optimization algorithm. Finally, it improves the learning ability of linear hash function by noisy example erasing and deviation correcting. Considering this, MIEH is able to garner more robust hash function. Extensive experiments conducted on three popular benchmark datasets highlight the superiority of our MIEH on large-scale cross-modal retrieval tasks and demonstrate its competitive performance against state-of-the-art approaches. The source code is available1.
Yongxin Wang 0001, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.1
2023 Prototype-Based Layered Federated Cross-Modal Hashing
abstract
Recently, deep cross-modal hashing has gained increasing attention. However, in many practical cases, data are distributed and cannot be collected due to privacy concerns, which greatly reduces the cross-modal hashing performance on each client. And due to the problems of statistical heterogeneity, model heterogeneity, and forcing each client to accept the same parameters, applying federated learning to cross-modal hash learning becomes very tricky. In this paper, we propose a novel method called prototype-based layered federated cross-modal hashing. Specifically, the prototype is introduced to learn the similarity between instances and classes on server, reducing the impact of statistical heterogeneity (non-IID) on different clients. And we monitor the distance between local and global prototypes to further improve the performance. To realize personalized federated learning, a hypernetwork is deployed on server to dynamically update different layers’ weights of local model. Experimental results on benchmark datasets show that our method outperforms state-of-the-art methods.
Yu-Wei Zhan, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin-Shun Xu
ICASSP5
2023 Self-Distillation Dual-Memory Online Hashing with Hash Centers for Streaming Data Retrieval
abstract
With the continuous generation of massive amounts of multimedia data nowadays, hashing has demonstrated significant potentials for large-scale search. To handle the emerging needs for streaming data retrieval, online hashing is drawing more and more attention. For online scenario, data distribution may change and concept drifts may occur as new data is continuously added to the database. Inevitably, hashing models may lose or disrupt the previously obtained knowledge when learning from new information, which is called the problem of catastrophic forgetting. In this paper, we propose a new online hashing method called Self-distillation Dual-memory Online Hashing with Hash Centers, which is abbreviated to SDOH-HC, to overcome this challenge. Specifically, SDOH-HC contains replay and distillation modules. For replay, a dual-memory mechanism is proposed which involves hash centers and exemplars. For knowledge distillation, we let hash centers distill information from themselves, i.e., the version of last round. Additionally, a new objective function is further built on above modules and is solved discretely to learn hash codes. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our method.
Chong-Yu Zhang, Xin Luo 0006, Yu-Wei Zhan, Peng-Fei Zhang 0001, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xun Yang 0001, Xin-Shun Xu
ACM Multimedia6
2023 Diagnose Like Doctors: Weakly Supervised Fine-Grained Classification of Breast Cancer
abstract
Breast cancer is the most common type of cancers in women. Therefore, how to accurately and timely diagnose it becomes very important. Some computer-aided diagnosis models based on pathological images have been proposed for this task. However, there are still some issues that need to be further addressed. For example, most deep learning based models suffer from a lack of interpretability. In addition, some of them cannot fully exploit the information in medical data, e.g., hierarchical label structure and scattered distribution of target objects. To address these issues, we propose a weakly supervised fine-grained medical image classification method for breast cancer diagnosis, i.e., DLD-Net for short. It simulates the diagnostic procedures of pathologists by multiple attention-guided cropping and dropping operations, making it have good clinical interpretability. Moreover, it cannot only exploit the global information of a whole image, but also further mine the critical local information by generating and selecting critical regions from the image. In light of this, those subtle discriminating information hidden in scattered regions can be exploited. In addition, we also design a novel hierarchical cross-entropy loss to utilize the hierarchical label information in medical images, making the classification results more discriminative. Furthermore, DLD-Net is a weakly supervised network, which can be trained end-to-end without any additional region annotations. Extensive experimental results on three benchmark datasets demonstrate that DLD-Net is able to achieve good results and outperforms some state-of-the-art methods.
Jieru Tian, Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
ACM Trans. Intell. Syst. Technol.2
2022 Discrete online cross-modal hashing
Yu-Wei Zhan, Yongxin Wang 0001, Xiao-Ming Wu 0002, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.2
2022 A High-Dimensional Sparse Hashing Framework for Cross-Modal Retrieval
abstract
In recent years, many achievements have been made in improving the performance of supervised cross-modal hashing. However, it remains an open issue on how to fully explore the data information to achieve fine-grained retrieval performance. Most methods employ logical labels or a binary similarity matrix to supervise the hash learning, losing a lot of useful information. From another point of view, the low expressiveness of dense hash code severely limits its preservation of fine-grained data information. With this motivation, in this paper, we propose a high-dimensional sparse hashing framework for cross-modal retrieval, i.e., High-dimensional Sparse Cross-modal Hashing, HSCH for short. It leverages not only high-level semantic labels but also low-level multi-modal features to construct a fine-grained similarity. In particular, based on two well-designed rules, i.e., multi-level and prioritized, it is able to avoid semantic conflicts. Additionally, it leverages the strong power of high-dimensional sparse hash codes to preserve the fine-grained similarity. Then, it efficiently solves the sparse and discrete constraints of sparse hash codes through an efficient discrete optimization algorithm. In light of this, it is much more efficient and scalable to large-scale datasets. More importantly, the computational complexity of HSCH in the retrieval phase is as efficient as those naive hashing methods that use dense hash codes. Moreover, to support online learning scenarios, this paper also extends HSCH into an online version, i.e., HSCH_on. Extensive experiments on three benchmark datasets demonstrate the superiority of our framework compared with some state-of-the-art cross-modal hashing approaches in terms of both accuracy and efficiency.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.1
2022 Fast Cross-Modal Hashing With Global and Local Similarity Embedding
abstract
Recently, supervised cross-modal hashing has attracted much attention and achieved promising performance. To learn hash functions and binary codes, most methods globally exploit the supervised information, for example, preserving an at-least-one pairwise similarity into hash codes or reconstructing the label matrix with binary codes. However, due to the hardness of the discrete optimization problem, they are usually time consuming on large-scale datasets. In addition, they neglect the class correlation in supervised information. From another point of view, they only explore the global similarity of data but overlook the local similarity hidden in the data distribution. To address these issues, we present an efficient supervised cross-modal hashing method, that is, fast cross-modal hashing (FCMH). It leverages not only global similarity information but also the local similarity in a group. Specifically, training samples are partitioned into groups; thereafter, the local similarity in each group is extracted. Moreover, the class correlation in labels is also exploited and embedded into the learning of binary codes. In addition, to solve the discrete optimization problem, we further propose an efficient discrete optimization algorithm with a well-designed group updating scheme, making its computational complexity linear to the size of the training set. In light of this, it is more efficient and scalable to large-scale datasets. Extensive experiments on three benchmark datasets demonstrate that FCMH outperforms some state-of-the-art cross-modal hashing approaches in terms of both retrieval accuracy and learning efficiency.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Rui Li 0090, Xin-Shun Xu
IEEE Trans. Cybern.1
2022 Fine-Grained Hashing With Double Filtering
abstract
Fine-grained hashing is a new topic in the field of hashing-based retrieval and has not been well explored up to now. In this paper, we raise three key issues that fine-grained hashing should address simultaneously, i.e., fine-grained feature extraction, feature refinement as well as a well-designed loss function. In order to address these issues, we propose a novel Fine-graIned haSHing method with a double-filtering mechanism and a proxy-based loss function, FISH for short. Specifically, the double-filtering mechanism consists of two modules, i.e., Space Filtering module and Feature Filtering module, which address the fine-grained feature extraction and feature refinement issues, respectively. Thereinto, the Space Filtering module is designed to highlight the critical regions in images and help the model to capture more subtle and discriminative details; the Feature Filtering module is the key of FISH and aims to further refine extracted features by supervised re- weighting and enhancing. Moreover, the proxy-based loss is adopted to train the model by preserving similarity relationships between data instances and proxy-vectors of each class rather than other data instances, further making FISH much efficient and effective. Experimental results demonstrate that FISH achieves much better retrieval performance compared with state-of-the-art fine-grained hashing methods, and converges very fast. The source code is publicly available: https://github.com/chenzhenduo/FISH.
Zhen-Duo Chen 0001, Xin Luo 0006, Yongxin Wang 0001, Shanqing Guo, Xin-Shun Xu
IEEE Trans. Image Process.3
2021 Weakly-Supervised Online Hashing
abstract
With the rapid development of social websites, recent years have witnessed an explosive growth of social images with user-provided tags. Most existing hashing methods for social image retrieval are batch-based which may violate the nature of social images, i.e., social images are usually generated periodically or collected in a stream fashion. Although there exist many online hashing methods, they either adopt unsupervised learning which ignore the relevant tags, or are designed in the supervised manner which needs high-quality labels. In this paper, to overcome the above limitations, we propose a new method named Weakly-supervised Online Hashing (WOH). In order to learn high-quality hash codes, WOH exploits the weak supervision, i.e., tags, by considering the semantics of tags and removing the noise. Besides, we develop a discrete online optimization algorithm, which is efficient and scalable. Extensive experiments conducted on two real-world datasets demonstrate the superiority of WOH.
Yu-Wei Zhan, Xin Luo 0006, Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin-Shun Xu
ICME4
2021 High-Dimensional Sparse Cross-Modal Hashing with Fine-Grained Similarity Embedding
abstract
Recently, with the discoveries in neurobiology, high-dimensional sparse hashing has attracted increasing attention. In contrast with general hashing that generates low-dimensional hash codes, the high-dimensional sparse hashing maps inputs into a higher dimensional space and generates sparse hash codes, achieving superior performance. However, the sparse hashing has not been fully studied in hashing literature yet. For example, how to fully explore the power of sparse coding in cross-modal retrieval tasks; how to discretely solve the binary and sparse constraints so as to avoid the quantization error problem. Motivated by these issues, in this paper, we present an efficient sparse hashing method, i.e., High-dimensional Sparse Cross-modal Hashing, HSCH for short. It not only takes the high-level semantic similarity of data into consideration, but also properly exploits the low-level feature similarity. In specific, we theoretically design a fine-grained similarity with two critical fusion rules. Then we take advantage of sparse codes to embed the fine-grained similarity into the to-be-learnt hash codes. Moreover, an efficient discrete optimization algorithm is proposed to solve the binary and sparse constraints, reducing the quantization error. In light of this, it becomes much more trainable, and the learnt hash codes are more discriminative. More importantly, the retrieval complexity of HSCH is as efficient as general hash methods. Extensive experiments on three widely-used datasets demonstrate the superior performance of HSCH compared with several state-of-the-art cross-modal hashing approaches.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
WWW1
2021 BATCH: A Scalable Asymmetric Discrete Cross-Modal Hashing
abstract
Supervised cross-modal hashing has attracted much attention. However, there are still some challenges, e.g., how to effectively embed the label information into binary codes, how to avoid using a large similarity matrix and make a model scalable to large-scale datasets, how to efficiently solve the binary optimization problem. To address these challenges, in this paper, we present a novel supervised cross-modal hashing method, i.e., scalaBle Asymmetric discreTe Cross-modal Hashing, BATCH for short. It leverages collective matrix factorization to learn a common latent space for the labels and different modalities, and embeds the labels into binary codes by minimizing a distance-distance difference problem. Furthermore, it builds a connection between the common latent space and the hash codes by an asymmetric strategy. In the light of this, it can perform cross-modal retrieval and embed more similarity information into the binary codes. In addition, it introduces a quantization minimization term and orthogonal constraints into the optimization problem, and generates the binary codes discretely. Therefore, the quantization error and redundancy may be much reduced. Moreover, it is a two-step method, making the optimization simple and scalable to large-scale datasets. Extensive experimental results on three benchmark datasets demonstrate that BATCH outperforms some state-of-the-art cross-modal hashing methods in terms of accuracy and efficiency.
Yongxin Wang 0001, Xin Luo 0006, Liqiang Nie, Jingkuan Song, Wei Zhang 0021, Xin-Shun Xu
IEEE Trans. Knowl. Data Eng.1
2020 Label Embedding Online Hashing for Cross-Modal Retrieval
abstract
Supervised cross-modal hashing has gained a lot of attention recently. However, most existing methods learn binary codes or hash functions in a batch-based scheme, which is inefficient in an online scenario, i.e., data points come in a streaming fashion. Online hashing is a promising solution; however, there still exist several challenges, e.g., how to effectively exploit semantic information, how to discretely solve the binary optimization problem, how to efficiently update hash codes and hash functions. To address these issues, in this paper, we propose a novel supervised online cross-modal hashing method, i.e., Label EMbedding ONline hashing, LEMON for short. It builds a label embedding framework including label similarity preserving and label reconstructing, which may generate discriminative binary codes and reduce the computational complexity. Furthermore, it not only preserves the pairwise similarity of incoming data, but also establishes a connection between newly coming data and existing data by the inner product minimization on a block similarity matrix. In the light of this, it can exploit more similarity information and make the optimization less sensitive to incoming data, leading to effective binary codes. In addition, we design a discrete optimization algorithm to solve the binary optimization problem without relaxation. Therefore, the quantization error can be reduced. Moreover, its computational complexity is only relevant to the size of incoming data, making it very efficient and scalable to large-scale datasets. Extensive experimental results on three benchmark datasets demonstrate that LEMON outperforms some state-of-the-art offline and online cross-modal hashing methods in terms of accuracy and efficiency.
Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
ACM Multimedia1
2020 Supervised Hierarchical Deep Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing has attracted much attention in the large-scale multimedia search area. In many real applications, labels of samples have hierarchical structure which also contains much useful information for learning. However, most existing methods are originally designed for non-hierarchical labeled data and thus fail to exploit the rich information of the label hierarchy. In this paper, we propose an effective cross-modal hashing method, named Supervised Hierarchical Deep Cross-modal Hashing, SHDCH for short, to learn hash codes by explicitly delving into the hierarchical labels. Specifically, both the similarity at each layer of the label hierarchy and the relatedness across different layers are implanted into the hash-code learning. Besides, an iterative optimization algorithm is proposed to directly learn the discrete hash codes instead of relaxing the binary constraints. We conducted extensive experiments on two real-world datasets and the experimental results show the superior performance of SHDCH over several state-of-the-art methods.
Yu-Wei Zhan, Xin Luo 0006, Yongxin Wang 0001, Xin-Shun Xu
ACM Multimedia3
2019 A Two-Step Cross-Modal Hashing by Exploiting Label Correlations and Preserving Similarity in Both Steps
abstract
In this paper, we present a novel Two-stEp Cross-modal Hashing method, TECH for short, for cross-modal retrieval tasks. As a two-step method, it first learns hash codes based on semantic labels, while preserving the similarity in the original space and exploiting the label correlations in the label space. In the light of this, it is able to make better use of label information and generate better binary codes. In addition, different from other two-step methods that mainly focus on the hash codes learning, TECH adopts a new hash function learning strategy in the second step, which also preserves the similarity in the original space. Moreover, with the help of well designed objective function and optimization scheme, it is able to generate hash codes discretely and scalable for large scale data. To the best of our knowledge, it is the first cross-modal hashing method exploiting label correlations, and also the first two-step hashing model preserving the similarity while leaning hash function. Extensive experiments demonstrate that the proposed approach outperforms some state-of-the-art cross-modal hashing methods.
Zhen-Duo Chen 0001, Yongxin Wang 0001, Huiqiong Li, Xin Luo 0006, Liqiang Nie, Xin-Shun Xu
ACM Multimedia2
2018 SDMCH: Supervised Discrete Manifold-Embedded Cross-Modal Hashing
abstract
Cross-modal hashing methods have attracted considerable attention. Most pioneer approaches only preserve the neighborhood relationship by constructing the correlations among heterogeneous modalities. However, they neglect the fact that the high-dimensional data often exists on a low-dimensional manifold embedded in the ambient space and the relative proximity between the neighbors is also important. Although some methods leverage the manifold learning to generate the hash codes, most of them fail to explicitly explore the discriminative information in the class labels and discard the binary constraints during optimization, generating large quantization errors. To address these issues, in this paper, we present a novel cross-modal hashing method, named Supervised Discrete Manifold-Embedded Cross-Modal Hashing (SDMCH). It can not only exploit the non-linear manifold structure of data and construct the correlation among heterogeneous multiple modalities, but also fully utilize the semantic information. Moreover, the hash codes can be generated discretely by an iterative optimization algorithm, which can avoid the large quantization errors. Extensive experimental results on three benchmark datasets demonstrate that SDMCH outperforms ten state-of-the-art cross-modal hashing methods.
Xin Luo 0006, Xiao-Ya Yin, Liqiang Nie, Xuemeng Song, Yongxin Wang 0001, Xin-Shun Xu
IJCAI5