Zhen-Duo Chen 0001

dblp:312/8153-1 · also Zhenduo Chen 0001 · DBLP profile ↗
← Back
53ranked-venue papers
5as first author
45since 2021 · last 2026
0000-0002-3481-4892ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 29 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 18 since 2021Computer networks · 6 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Federated class-incremental learning with prompting
Xin Luo 0006, Fang-Yi Liang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin-Shun Xu
Expert Syst. Appl.5
2026 FusionHash: Empowering cross-modal hashing with fine-grained knowledge fusion
Bo-Lin Zhang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
Knowl. Based Syst.6
2026 RISEN: Incremental multilingual text recognition by sharing and fusing cross-language knowledge
Zi-Xin Li, Tai Zheng, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.4
2026 Dynamic Clustering-driven weakly-supervised online hashing with enhanced similarity
Chong-Yu Zhang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.4
2026 Fine-Grained Augmentation and Progressive Feature Integration for Unsupervised Fine-Grained Hashing
abstract
Unsupervised fine-grained image retrieval aims to retrieve specific subcategory images from large-scale unlabeled databases. The small inter-class and large intra-class variances inherent in fine-grained images present significant challenges for unsupervised model training and feature recognition. Without the guidance of supervised information, existing methods often fail to focus on fine-grained details, and multi-region features struggle to embed effectively into hash codes. In this article, we propose Fine-Grained Augmentation and Progressive Feature Integration for unsupervised fine-grained hashing, named FAPI. Specifically, from the perspective of unsupervised contrastive learning, we design fine-grained feature augmentation and cross-contrastive learning modules to enhance the capture of critical discriminative details. Additionally, from a feature extraction standpoint, we propose a progressive granularity feature integration module to extract and fuse multi-layer, multi-granularity features, ensuring effective fine-grained feature extraction and hash code embedding. Extensive experiments on five widely recognized fine-grained datasets demonstrate that FAPI significantly outperforms existing unsupervised methods, achieving state-of-the-art performance.
Yun-Cong Liu, Zhen-Duo Chen 0001, Qingze Bai, Xiao-Dong Xie, Xin Luo 0006, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Gleaning Wisdom from the Past: Towards Label Incremental Learning for Online Hashing with a Plug-and-Play Framework
abstract
Benefiting from low storage requirements, high search efficiency, and promising performance, online hashing has attracted considerable attention for large-scale streaming data retrieval. However, most existing online hashing methods implicitly assume that the label space is fixed and fail to effectively handle the label incremental problem. With this motivation, in this article, we introduce a novel supervised online hashing framework named Incremental Plug-and-play Online Hashing (IPOH), which endows prevailing methods with the ability to solve the above issues. Specifically, IPOH could extract the patterns of old classes from the learned hash codes at the last round and then generate the class-wise representations for unseen new labels. With these representations, we could learn the hash codes for streaming data with new labels. Experimental results on two widely used datasets demonstrate the effectiveness of our method in handling class incremental data. We hope this article could convince more researchers to look into this interesting problem, and the source code for the implementation of our IPOH framework has been released at https://github.com/ZCyueternal/IPOH.git .
Chong-Yu Zhang, Xin Luo 0006, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Few-Shot Fine-Grained Image Classification with Progressively Feature Refinement and Continuous Relationship Modeling
abstract
Recently, a number of effective methods have been proposed to tackle the challenging task of Few-Shot Fine-Grained Image Classification (FS-FGIC). However, how to fully leverage the backbone network to discover and extract detailed features to generate more discriminative class prototypes, as well as how to accurately model the similarity relationship between query samples and the class prototypes, are still issues to be further considered. Therefore, we propose a novel progreSsively featUre refInement and conTinuous rElationship moDeling method, SUITED for short, to address these two issues existing in the State-of-the-Art FS-FGIC methods. Specifically, we design the Progressive Feature Refinement Module (PFRM) to fully exploit the backbone network's progressive feature extraction capabilities, forming multi-scale feature representations to further enhance discriminative features. Then, the Continuous Relationship Modeling Module (CRMM) is proposed to capture the dependencies between query samples and the corresponding class prototypes, achieving precise optimization of the distances among corresponding sample points in the feature space. We conducted extensive experiments on five fine-grained benchmark datasets, and the experimental results demonstrate that the proposed method is comprehensively ahead of the existing State-of-the-Art methods.
Zhen-Xiang Ma, Zhen-Duo Chen 0001, Tai Zheng, Xin Luo 0006, Zixia Jia, Xin-Shun Xu
AAAI2
2025 Attraction Diminishing and Distributing for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel classes with limited samples after pre-training on a set of base classes. To avoid catastrophic forgetting and overfitting, most FSCIL methods first train the model on the base classes and then freeze the feature extractor in the incremental sessions. However, the reliance on nearest neighbor classification makes FSCIL prone to the hubness phenomenon, which negatively impacts performance in this dynamic and open scenario. While recent methods attempt to adapt to the dynamic and open nature of FSCIL, they are often limited to biased optimizations to the feature space. In this paper, we pioneer the theoretical analysis of the inherent hubness in FSCIL. To mitigate the negative effects of hubness, we propose a novel Attraction Diminishing and Distributing (D2A) method from the essential perspectives of distance metric and feature space. Extensive experimental results demonstrate that our method can broadly and significantly improve the performance of existing methods.
Li-Jun Zhao 0005, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
CVPR2
2025 Tag-Aware Weakly-Supervised Online Hashing with Enhanced Joint Representation
abstract
Weakly-supervised online hashing has garnered significant attention recently, yet several challenges remain unresolved, such as how to effectively denoise tags, and how to efficiently learn hash functions in dynamic online scenarios. To tackle these challenges, we propose a novel method named Tag-Aware Weakly-supervised Online Hashing with enhanced joint representation (TA-WOH). Our method creates an enhanced joint representation with CLIP-based features in order to reduce the tag noise. Additionally, we introduce a tag association and noise model for improved similarity matrix and hash code learning. A novel mapping mechanism is developed to align joint representations with the optimal tag space, enhancing both the accuracy and robustness of the model. The computational complexity of TA-WOH is dependent solely on the size of the incoming data, ensuring scalability and efficiency for large-scale datasets. Extensive experiments on two datasets demonstrate that our method surpasses several state-of-the-art methods in both accuracy and efficiency.
Yu-Wei Zhan, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
ICASSP3
2025 Evolving and Regularizing Meta-Environment Learner for Fine-Grained Few-Shot Class-Incremental Learning
abstract
Recently proposed Fine-Grained Few-Shot Class-Incremental Learning (FG-FSCIL) offers a practical and efficient solution for enabling models to incrementally learn new fine-grained categories under limited data conditions. However, existing methods still settle for the fine-grained feature extraction capabilities learned from the base classes. Unlike conventional datasets, fine-grained categories exhibit subtle inter-class variations, naturally fostering latent synergy among sub-categories. Meanwhile, the incremental learning framework offers an opportunity to progressively strengthen this synergy by incorporating new sub-category data over time. Motivated by this, we theoretically formulate the FSCIL problem and derive a generalization error bound within a shared fine-grained meta-category environment. Guided by our theoretical insights, we design a novel Meta-Environment Learner (MEL) for FG-FSCIL, which evolves fine-grained feature extraction to enhance meta-environment understanding and simultaneously regularizes hypothesis space complexity. Extensive experiments demonstrate that our method consistently and significantly outperforms existing approaches.
Li-Jun Zhao 0005, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
NeurIPS2
2025 OH-CMH: Towards cross-modal hashing for streaming data with hierarchical labels and label increment scenario
Chong-Yu Zhang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
Knowl. Based Syst.5
2025 Hypergraph-based CLIP hashing for unsupervised cross-modal retrieval
Jia-Rui Zhao, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
Knowl. Based Syst.5
2025 DGPrompt: Dual-guidance prompts generation for vision-language models
Tai Zheng, Zhen-Duo Chen 0001, Zi-Chao Zhang 0002, Zhen-Xiang Ma, Li-Jun Zhao 0005, Chong-Yu Zhang, Xin Luo 0006, Xin-Shun Xu
Neural Networks2
2025 Class-Aware Prompting for Federated Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn new classes from limited samples while preventing catastrophic forgetting. With the increasing distribution of learning data across different clients and privacy concerns, FSCIL faces a more realistic scenario where few learning samples are distributed across different clients, thereby necessitating a Federated Few-Shot Class-Incremental Learning (FedFSCIL) scenario. However, this integration faces challenges from non-IID problem, which affects model generalization and training efficiency. The communication overhead in federated settings also presents a significant challenge. To address these issues, we propose Class-Aware Prompting for Federated Few-Shot Class-Incremental Learning (FedCAP). Our framework leverages pre-trained models enhanced by a class-wise prompt pool, where shared class-wise keys enable clients to utilize global class information during training. This unifies the understanding of base class features across clients and enhances model consistency. We further incorporate a class-level information fusion module to improve class representation and model generalization. Our approach requires very few parameter transmission during model aggregation, ensuring communication efficiency. To our knowledge, this is the first study to explore the scenario of FedFSCIL. Consequently, we designed comprehensive experimental setups and made the code publicly available.
Fang-Yi Liang, Yu-Wei Zhan, Chong-Yu Zhang, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.5
2025 BTG-Net++: Enhanced Bi-Directional Task-Guided Network for Few-Shot Fine-Grained Image Classification
abstract
In recent years, a number of effective Few-Shot Fine-Grained Image Classification (FS-FGIC) methods have been proposed, which mainly focus on extracting discriminative information within high-level features in a single episode/task. However, this is insufficient for addressing the cross-task challenges of FS-FGIC, which is represented in two aspects. On the one hand, from the perspective of the Fine-Grained Image Classification (FGIC) task, there is a need to supplement the model with mid-level features containing rich fine-grained information. On the other hand, from the perspective of the Few-Shot Learning (FSL) task, explicit modeling of cross-task general knowledge is required. In this paper, we propose a novel Enhanced Bi-directional Task-Guided Network (BTG-Net++) to tackle these issues. Specifically, from the FGIC task perspective, we design the Semantic-Guided Noise Filtering (SGNF) module to filter noise on mid-level features rich in detailed information with the assistance of high-level features. Further, from the FSL task perspective, the General Knowledge Prompt Modeling (GKPM) module is proposed to retain the cross-task general knowledge by utilizing the prompting mechanism, thereby enhancing the model’s generalization performance on unseen novel classes. We have conducted extensive experiments on five fine-grained benchmark datasets, and the results demonstrate that BTG-Net++ shows considerable improvements compared with state-of-the-art methods.
Zhen-Xiang Ma, Zhen-Duo Chen 0001, Tai Zheng, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Distributed Learning for Privacy-Preserving Semi-Supervised Video Anomaly Detection
abstract
Semi-supervised video anomaly detection (SS-VAD) is essential for intelligent monitoring. However, collecting large-scale surveillance videos from various organizations raises significant privacy concerns regarding sensitive information. Federated learning offers a promising solution by enabling distributed learning among multiple participants while safeguarding privacy. Despite its potential, research on applying federated learning to SS-VAD remains unexplored due to the inherent challenges of this task. In this paper, we solve this task via proposing DLPP, a novel distributed learning framework for privacy-preserving SS-VAD. It addresses the issue of statistical heterogeneity among data from different participants in real-world federated SS-VAD applications, particularly focusing on non-independent and identically distributed (non-IID) data and imbalanced data volumes. In specific, it addresses these challenges in two key innovations: 1) For the non-IID data challenge, it dynamically updates the client model based on the overall gradient at the client of the previous training round and the degree of divergence between the server model and the client model. In this way, it can better adapt the server model to each client and promote convergence. 2) For the imbalanced data volumes challenge, it adaptively allocates client aggregation weights by comprehensively considering the data volumes, model quality, and learning efficiency of clients. This means a more robust server model can be obtained, and model bias reduced. We conduct extensive experiments to evaluate the performance of DLPP on benchmark datasets by partitioning data to simulate various degrees of non-IID environments. The results show that DLPP significantly outperforms both Baseline and SOTA methods, achieving up to a 3.89% improvement, and its communication efficiency is 3x better than FedAvg.
Xiao-Dong Xie, Yu-Wei Zhan, Zhen-Xiang Ma, Hong-Mei Liu, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Self-Supervised Discovery of Cross-Lingual Shared Knowledge for Continual Text Recognition
abstract
Incremental multilingual text recognition (IMLTR) aims to advance continual learning by retaining knowledge from previously learned languages while adapting to new ones. Existing methods typically perform under a constrained assumption that each text instance originates from a specific single-language domain. However, this assumption is inaccurate in multilingual scenarios, as it overlooks the inherent cross-lingual knowledge, i.e., the incremental sharing problem. To address this issue, we propose a novel self-supervised cross-lingual knowledge discovery framework, CrossKnow, tailored for IMLTR tasks. Specifically, an innovative shared knowledge discovery strategy is developed to identify potential shared knowledge by leveraging prediction consistency across multiple recognizers, thus eliminating the reliance on language labels of all characters. Building upon this shared knowledge, we further design a multi-granularity, multi-task language domain discriminator to capture dependency relationships among incremental languages, which could adequately guide the hierarchical sequence decoding. By mining shared knowledge, CrossKnow can not only mitigate the forgetting of old knowledge but also efficiently achieve cross-lingual knowledge transfer, thereby promoting the continual learning of incremental multilingual text recognition models. Experiments on two widely used datasets, MLT17 and MLT19, demonstrate the superiority of CrossKnow. Compared to methods that leverage additional language supervision of characters, CrossKnow achieves competitive performance while eliminating storage overhead and improving computation efficiency.
Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Image Process.2
2025 BITS: Bit-Extendable Incremental Hashing in Open Environments
abstract
Hashing is an effective technique for large-scale image retrieval. However, traditional hashing models typically follow a closed-set assumption, which fails to satisfy the practicality of real-world tasks. In this paper, we explore a meaningful yet overlooked question: is there a hashing paradigm that not only supports rehearsal-free online incremental coding for single-pass data streams but also adapts to potentially expanding concept spaces in open environments? Instead of presetting fixed bit lengths, we suggest adjusting the bit length dynamically based on the number of encountered categories, meanwhile enabling bit extension of existing hash codes to match the adaptive code lengths without knowledge forgetting. Therefore, we propose a Bit-extendable IncremenTal haShing (BITS) method for image retrieval in open environments. Specifically, we identify a blurry incremental setup to better simulate realistic scenarios, revisiting the widely-used data-incremental and class-incremental settings. With this challenging setup, a three-phase framework is designed to efficiently perform incremental hashing, which jointly solves online continual coding and bit extension with adaptive code lengths. Through the well-designed hashing paradigm, BITS achieves comparable performance to offline hashing methods while significantly saving computational resources. Comprehensive experiments on six benchmarks demonstrate the superiority of our BITS in dynamic scenarios. The source code is available at https://github.com/yxinwang/BITS.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Image Process.2
2025 Domain-Aware Semantic Alignment Hashing for Large-Scale Zero-Shot Image Retrieval
abstract
Hashing has been proven to be effective in the field of large-scale image retrieval. However, traditional hashing is stuck in performance dilemmas under zero-shot scenarios due to the concept shift problem. Although some zero-shot hashing methods exploit category attributes to facilitate knowledge transfer across domains, they usually struggle to generate domain-adaptive hash codes, making it hard to distinguish samples between unknown and known classes. With this motivation, we propose a novel approach called Domain-Aware Semantic Alignment Zero-Shot Hashing (DSAZH), which reveals three issues that suppress performance: semantic misalignment, biased optimization, and ambiguous Hamming distance. To address these challenges, multiple initiatives are innovatively integrated into a unified framework: First, it generates semantic-aligned hash codes through class-level and instance-level semantic alignment ; then it learns unbiased hash codes and domain-adaptive hash function through unbiased optimization equipped with asymmetric processing and class-prompting regression; finally, it distinguishes seen instances from unseen using domain-aware thresholding . Extensive experiments show that DSAZH achieves up to 15.82% MAP improvement (e.g., 69.71% vs. 53.89% on large-scale ImageNet with 256-bit codes) while reducing training time by two orders of magnitude (e.g., 3.07 s vs. 202.84 s), demonstrating its superior accuracy and efficiency compared to state-of-the-art ZSH methods. The source code is available at https://github.com/yxinwang/DSAZH .
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Cross-Layer and Cross-Sample Feature Optimization Network for Few-Shot Fine-Grained Image Classification
abstract
Recently, a number of Few-Shot Fine-Grained Image Classification (FS-FGIC) methods have been proposed, but they primarily focus on better fine-grained feature extraction while overlooking two important issues. The first one is how to extract discriminative features for Fine-Grained Image Classification tasks while reducing trivial and non-generalizable sample level noise introduced in this procedure, to overcome the over-fitting problem under the setting of Few-Shot Learning. The second one is how to achieve satisfying feature matching between limited support and query samples with variable spatial positions and angles. To address these issues, we propose a novel Cross-layer and Cross-sample feature optimization Network for FS-FGIC, C2-Net for short. The proposed method consists of two main modules: Cross-Layer Feature Refinement (CLFR) module and Cross-Sample Feature Adjustment (CSFA) module. The CLFR module further refines the extracted features while integrating outputs from multiple layers to suppress sample-level feature noise interference. Additionally, the CSFA module addresses the feature mismatch between query and support samples through both channel activation and position matching operations. Extensive experiments have been conducted on five fine-grained benchmark datasets, and the results show that the C2-Net outperforms other state-of-the-art methods by a significant margin in most cases. Our code is available at: https://github.com/zenith0923/C2-Net.
Zhen-Xiang Ma, Zhen-Duo Chen 0001, Li-Jun Zhao 0005, Zi-Chao Zhang 0002, Xin Luo 0006, Xin-Shun Xu
AAAI2
2024 Characteristics Matching Based Hash Codes Generation for Efficient Fine-Grained Image Retrieval
abstract
The rapidly growing scale of data in practice poses demands on the efficiency of retrieval models. However, for fine-grained image retrieval task, there are inherent contradictions in the design of hashing based efficient models. Firstly, the limited information embedding capacity of low-dimensional binary hash codes, coupled with the detailed information required to describe fine-grained categories, results in a contradiction in feature learning. Secondly, there is also a contradiction between the complexity of fine-grained feature extraction models and retrieval efficiency. To address these issues, in this paper, we propose the characteristics matching based hash codes generation method. Coupled with the cross-layer semantic information transfer module and the multi-region feature embedding module, the proposed method can generate hash codes that effectively capture fine-grained differences among samples while ensuring efficient inference. Extensive experiments on widely used datasets demonstrate that our method can significantly outperform state-of-the-art methods.
Zhen-Duo Chen 0001, Li-Jun Zhao 0005, Zi-Chao Zhang 0002, Xin Luo 0006, Xin-Shun Xu
CVPR1
2024 FedCAFE: Federated Cross-Modal Hashing with Adaptive Feature Enhancement
abstract
Deep Cross-Modal Hashing (CMH) has become one of the most popular solutions for cross-modal retrieval. Existing methods need to first collect data and then be trained with these accumulated data. However, in real world, data may be generated and possessed by different owners. Considering the concerns about privacy, data may not be shared or transmitted, leading to the failure of sufficient training of CMH. To solve the problem, we propose a new framework called Federated Cross-modal Hashing with Adaptive Feature Enhancement (FedCAFE). FedCAFE is a federated method which could use distributed data to train existing CMH methods under the privacy protection. To overcome the data heterogeneity challenge of distributed data and improve the generalization ability of global model, FedCAFE is endowed with a novel adaptive feature enhancement module and a new weighted aggregation strategy. Besides, it could fully utilize the rich global information carried in the global model to constrain the model during the local training process. We have conducted extensive experiments on four widely-used datasets in CMH domain with both IID and non-IID settings. The reported results demonstrate that the proposed FedCAFE achieves better performance than several state-of-the-art baselines.
Yu-Wei Zhan, Chong-Yu Zhang, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xun Yang 0001, Xin-Shun Xu
ACM Multimedia5
2024 Hierarchical Multi-label Learning for Incremental Multilingual Text Recognition
abstract
Multilingual text recognition (MLTR) is increasingly essential for facilitating cultural communication. However, existing methods often struggle with retaining previous knowledge when learning new languages. A straightforward solution is performing incremental learning (IL) on MLTR tasks. However, it ignores the shared words and characters across incremental languages, which we first term as an incremental sharing problem. Motivated by this observation, we propose HierArchicalMulti-label learning framework for Multilingual tExt Recognition, termed HAMMER. An online knowledge analysis is designed to identify shared knowledge and provide corresponding multi-label language supervision. Specifically, only words and characters appearing simultaneously in multiple languages are considered shared knowledge. Additionally, to further capture language dependencies, we introduce a hierarchical language evaluation mechanism to predict language scores at word and character levels. These scores, supervised by the knowledge analysis, guide the specific recognizers to effectively utilize both old and new language knowledge, thereby mitigating catastrophic forgetting caused by imbalanced rehearsal sets. Extensive experiments conducted on benchmark datasets, MLT17 and MLT19, show that HAMMER exhibits remarkable results and outperforms other state-of-the-art approaches.
Minghui Liu 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
ACM Multimedia3
2024 Bi-directional Task-Guided Network for Few-Shot Fine-Grained Image Classification
abstract
In recent years, the Few-Shot Fine-Grained Image Classification (FS-FGIC) problem has gained widespread attention. A number of effective methods have been proposed that focus on extracting discriminative information within high-level features in a single episode/task. However, this is insufficient for addressing the cross-task challenges of FS-FGIC, which is represented in two aspects. On the one hand, from the perspective of the Fine-Grained Image Classification (FGIC) task, there is a need to supplement the model with mid-level features containing rich fine-grained information. On the other hand, from the perspective of the Few-Shot Learning (FSL) task, explicit modeling of cross-task general knowledge is required. In this paper, we propose a novel Bi-directional Task-Guided Network (BTG-Net) to tackle these issues. Specifically, from the FGIC task perspective, we design the Semantic-Guided Noise Filtering (SGNF) module to filter noise on mid-level features rich in detailed information. Further, from the FSL task perspective, the General Knowledge Prompt Modeling (GKPM) module is proposed to retain the cross-task general knowledge by utilizing the prompting mechanism, thereby enhancing the model's generalization performance on novel classes. We have conducted extensive experiments on five fine-grained benchmark datasets, and the results demonstrate that BTG-Net outperforms state-of-the-art methods comprehensively.
Zhen-Xiang Ma, Zhen-Duo Chen 0001, Li-Jun Zhao 0005, Zi-Chao Zhang 0002, Tai Zheng, Xin Luo 0006, Xin-Shun Xu
ACM Multimedia2
2024 POLISH: Adaptive Online Cross-Modal Hashing for Class Incremental Data
abstract
In recent years, hashing-based online cross-modal retrieval has garnered growing attention. This trend is motivated by the fact that web data is increasingly delivered in a streaming manner as opposed to batch processing. Simultaneously, the sheer scale of web data sometimes makes it impractical to fully load for the training of hashing models. Despite the evolution of online cross-modal hashing techniques, several challenges remain: 1) Most existing methods learn hash codes by considering the relevance among newly arriving data or between new data and the existing data, often disregarding valuable global semantic information. 2) A common but limiting assumption in many methods is that the label space remains constant, implying that all class labels should be provided within the first data chunk. This assumption does not hold in real-world scenarios, and the presence of new labels in incoming data chunks can severely degrade or even break these methods.
Yu-Wei Zhan, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Yinwei Wei, Xin-Shun Xu
WWW3
2024 Weighted cross-modal hashing with label enhancement
Yongxin Wang 0001, Kuikui Wang, Xiushan Nie, Zhen-Duo Chen 0001
Knowl. Based Syst.5
2024 A vision transformer for fine-grained classification by reducing noise and enhancing discriminative information
Zi-Chao Zhang 0002, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.2
2024 Multiple Information Embedded Hashing for Large-Scale Cross-Modal Retrieval
abstract
Recently, many efforts have been devoted to improving the retrieval performance of supervised cross-modal hashing; however, current methods are gradually reaching a performance bottleneck, especially when dealing with real-world multimedia data. This is mainly due to their application of coarse-grained semantics, unrobust hash functions, and inflexible workflows. Therefore, discovering refined semantics hidden in data, designing robust hash functions, and creating a non-interfering but facilitative learning workflow are much more significant. With this motivation, in this paper, we propose a novel supervised cross-modal hashing method, i.e., Multiple Information Embedded Hashing, MIEH for short. It consists of a three-step working flow that flexibly handles multiple information mining, hash code learning, and hash function learning. First, it explores the multimedia data from multiple perspectives such as modal-level consistency, class-level discriminability, and instance-level similarity to mine comprehensive semantic information, which not only contributes to the generation of discriminative hash codes, but also accelerates convergence. Subsequently, MIEH is committed to embed the refined semantics into targeted hash codes with an efficient discrete optimization algorithm. Finally, it improves the learning ability of linear hash function by noisy example erasing and deviation correcting. Considering this, MIEH is able to garner more robust hash function. Extensive experiments conducted on three popular benchmark datasets highlight the superiority of our MIEH on large-scale cross-modal retrieval tasks and demonstrate its competitive performance against state-of-the-art approaches. The source code is available1.
Yongxin Wang 0001, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Angular Isotonic Loss Guided Multi-Layer Integration for Few-Shot Fine-Grained Image Classification
abstract
Recent research on few-shot fine-grained image classification (FSFG) has predominantly focused on extracting discriminative features. The limited attention paid to the role of loss functions has resulted in weaker preservation of similarity relationships between query and support instances, thereby potentially limiting the performance of FSFG. In this regard, we analyze the limitations of widely adopted cross-entropy loss and introduce a novel Angular ISotonic (AIS) loss. The AIS loss introduces an angular margin to constrain the prototypes to maintain a certain distance from a pre-set threshold. It guides the model to converge more stably, learn clearer boundaries among highly similar classes, and achieve higher accuracy faster with limited instances. Moreover, to better accommodate the feature requirements of the AIS loss and fully exploit its potential in FSFG, we propose a Multi-Layer Integration (MLI) network that captures object features from multiple perspectives to provide more comprehensive and informative representations of the input images. Extensive experiments demonstrate the effectiveness of our proposed method on four standard fine-grained benchmarks. Codes are available at: https://github.com/Legenddddd/AIS-MLI.
Li-Jun Zhao 0005, Zhen-Duo Chen 0001, Zhen-Xiang Ma, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Image Process.2
2024 S3Mix: Same Category Same Semantics Mixing for Augmenting Fine-grained Images
abstract
Data augmentation is a common technique to improve the generalization performance of models for image classification. Although methods such as Mixup and CutMix that mix images randomly are indeed instrumental in general image classification, randomly swapping or masking regions is not friendly to fine-grained images, since the key to fine-grained image classification precisely lies in discriminative and informative regions, and it is unreasonable to generate labels solely consistent with the proportion of synthesis. Some erasing methods like Cutout even endanger fine-grained image classification because of erasing the discriminative regions by chance. In this article, we propose the Same Category Same Semantics Mixing method (S3Mix) corresponding to the characteristics of fine-grained images. Specifically, we limit the mixture to regions of the same category and semantics. The core of the method is two constraints. The exchange with the semantic region ensures the discrimination and semantics integrity of the generated image, and the exchange in the same class avoids the problem of unreasonable label generation. At the same time, we propose a homology loss to promote the semantic relationship between the generated positive image pairs. Experiments have been conducted on four fine-grained datasets, and the results show the proposed method is superior to the traditional image augmentation methods as well as some fine-grained data augmentation methods.
Zi-Chao Zhang 0002, Zhen-Duo Chen 0001, Zhen-Yu Xie, Xin Luo 0006, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Prototype-Based Layered Federated Cross-Modal Hashing
abstract
Recently, deep cross-modal hashing has gained increasing attention. However, in many practical cases, data are distributed and cannot be collected due to privacy concerns, which greatly reduces the cross-modal hashing performance on each client. And due to the problems of statistical heterogeneity, model heterogeneity, and forcing each client to accept the same parameters, applying federated learning to cross-modal hash learning becomes very tricky. In this paper, we propose a novel method called prototype-based layered federated cross-modal hashing. Specifically, the prototype is introduced to learn the similarity between instances and classes on server, reducing the impact of statistical heterogeneity (non-IID) on different clients. And we monitor the distance between local and global prototypes to further improve the performance. To realize personalized federated learning, a hypernetwork is deployed on server to dynamically update different layers’ weights of local model. Experimental results on benchmark datasets show that our method outperforms state-of-the-art methods.
Yu-Wei Zhan, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin-Shun Xu
ICASSP4
2023 FedVMR: A New Federated Learning Method for Video Moment Retrieval
abstract
Despite the great success achieved, existing video moment retrieval (VMR) methods are developed under the assumption that data are centralizedly stored. However, in real-world applications, due to the inherent nature of data generation and privacy concerns, data are often distributed on different silos, bringing huge challenges to effective large-scale training. In this work, we try to overcome above limitation by leveraging the recent success of federated learning. As the first that is explored in VMR field, the new task is defined as video moment retrieval with distributed data. Then, a novel federated learning method named FedVMR is proposed to facilitate large-scale and secure training of VMR models in decentralized environment. Experiments on benchmark datasets demonstrate its effectiveness. This work is the very first attempt to enable safe and efficient VMR training in decentralized scene, which is hoped to pave the way for further study in the related research field.
Xin Luo 0006, Zhen-Duo Chen 0001, Peng-Fei Zhang 0001, Meng Liu 0006, Xin-Shun Xu
ICASSP3
2023 Self-Distillation Dual-Memory Online Hashing with Hash Centers for Streaming Data Retrieval
abstract
With the continuous generation of massive amounts of multimedia data nowadays, hashing has demonstrated significant potentials for large-scale search. To handle the emerging needs for streaming data retrieval, online hashing is drawing more and more attention. For online scenario, data distribution may change and concept drifts may occur as new data is continuously added to the database. Inevitably, hashing models may lose or disrupt the previously obtained knowledge when learning from new information, which is called the problem of catastrophic forgetting. In this paper, we propose a new online hashing method called Self-distillation Dual-memory Online Hashing with Hash Centers, which is abbreviated to SDOH-HC, to overcome this challenge. Specifically, SDOH-HC contains replay and distillation modules. For replay, a dual-memory mechanism is proposed which involves hash centers and exemplars. For knowledge distillation, we let hash centers distill information from themselves, i.e., the version of last round. Additionally, a new objective function is further built on above modules and is solved discretely to learn hash codes. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our method.
Chong-Yu Zhang, Xin Luo 0006, Yu-Wei Zhan, Peng-Fei Zhang 0001, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xun Yang 0001, Xin-Shun Xu
ACM Multimedia5
2023 A multi-layer memory sharing network for video captioning
Tian-Zi Niu, Zhen-Duo Chen 0001, Xin Luo 0006, Zi Huang, Shanqing Guo, Xin-Shun Xu
Pattern Recognit.3
2023 Expansion window local alignment weighted network for fine-grained sketch-based image retrieval
Zi-Chao Zhang 0002, Zhen-Yu Xie, Zhen-Duo Chen 0001, Yu-Wei Zhan, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.3
2023 Diagnose Like Doctors: Weakly Supervised Fine-Grained Classification of Breast Cancer
abstract
Breast cancer is the most common type of cancers in women. Therefore, how to accurately and timely diagnose it becomes very important. Some computer-aided diagnosis models based on pathological images have been proposed for this task. However, there are still some issues that need to be further addressed. For example, most deep learning based models suffer from a lack of interpretability. In addition, some of them cannot fully exploit the information in medical data, e.g., hierarchical label structure and scattered distribution of target objects. To address these issues, we propose a weakly supervised fine-grained medical image classification method for breast cancer diagnosis, i.e., DLD-Net for short. It simulates the diagnostic procedures of pathologists by multiple attention-guided cropping and dropping operations, making it have good clinical interpretability. Moreover, it cannot only exploit the global information of a whole image, but also further mine the critical local information by generating and selecting critical regions from the image. In light of this, those subtle discriminating information hidden in scattered regions can be exploited. In addition, we also design a novel hierarchical cross-entropy loss to utilize the hierarchical label information in medical images, making the classification results more discriminative. Furthermore, DLD-Net is a weakly supervised network, which can be trained end-to-end without any additional region annotations. Extensive experimental results on three benchmark datasets demonstrate that DLD-Net is able to achieve good results and outperforms some state-of-the-art methods.
Jieru Tian, Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
ACM Trans. Intell. Syst. Technol.3
2023 Video Captioning by Learning from Global Sentence and Looking Ahead
abstract
Video captioning aims to automatically generate natural language sentences describing the content of a video. Although encoder-decoder-based models have achieved promising progress, it is still very challenging to effectively model the linguistic behavior of humans in generating video captions. In this paper, we propose a novel video captioning model by learning from gLobal sEntence and looking AheaD, LEAD for short. Specifically, LEAD consists of two modules: a Vision Module (VM) and a Language Module (LM) . Thereinto, VM is a novel attention network, which can map visual features to high-level language space and model entire sentences explicitly. LM can not only effectively make use of the information of the previous sequence when generating the current word, but also have a look at the future word. Therefore, based on VM and LM, LEAD can obtain global sentence information and future word information to make video captioning more like a fill-in-the-blank task than a word-by-word sentence generation. In addition, we also propose an autonomous strategy and a multi-stage training scheme to optimize the model, which can mitigate the problem of information leakage. Extensive experiments show that LEAD outperforms some state-of-the-art methods on MSR-VTT, MSVD, and VATEX, demonstrating the effectiveness of the proposed approach in video captioning. In addition, we release the code of our proposed model to be publicly available. 1
Tian-Zi Niu, Zhen-Duo Chen 0001, Xin Luo 0006, Peng-Fei Zhang 0001, Zi Huang, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Semantic Enhanced Video Captioning with Multi-feature Fusion
abstract
Video captioning aims to automatically describe a video clip with informative sentences. At present, deep learning-based models have become the mainstream for this task and achieved competitive results on public datasets. Usually, these methods leverage different types of features to generate sentences, e.g., semantic information, 2D or 3D features. However, some methods only treat semantic information as a complement of visual representations and cannot fully exploit it; some of them ignore the relationship between different types of features. In addition, most of them select multiple frames of a video with an equally spaced sampling scheme, resulting in much redundant information. To address these issues, we present a novel video-captioning framework, Semantic Enhanced video captioning with Multi-feature Fusion, SEMF for short. It optimizes the use of different types of features from three aspects. First, a semantic encoder is designed to enhance meaningful semantic features through a semantic dictionary to boost performance. Second, a discrete selection module pays attention to important features and obtains different contexts at different steps to reduce feature redundancy. Finally, a multi-feature fusion module uses a novel relation-aware attention mechanism to separate the common and complementary components of different features to provide more effective visual features for the next step. Moreover, the entire framework can be trained in an end-to-end manner. Extensive experiments are conducted on Microsoft Research Video Description Corpus (MSVD) and MSR-Video to Text (MSR-VTT) datasets. The results demonstrate that SEMF is able to achieve state-of-the-art results.
Tian-Zi Niu, Zhen-Duo Chen 0001, Xin Luo 0006, Shanqing Guo, Zi Huang, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Online Enhanced Semantic Hashing: Towards Effective and Efficient Retrieval for Streaming Multi-Modal Data
abstract
With the vigorous development of multimedia equipments and applications, efficient retrieval of large-scale multi-modal data has become a trendy research topic. Thereinto, hashing has become a prevalent choice due to its retrieval efficiency and low storage cost. Although multi-modal hashing has drawn lots of attention in recent years, there still remain some problems. The first point is that existing methods are mainly designed in batch mode and not able to efficiently handle streaming multi-modal data. The second point is that all existing online multi-modal hashing methods fail to effectively handle unseen new classes which come continuously with streaming data chunks. In this paper, we propose a new model, termed Online enhAnced SemantIc haShing (OASIS). We design novel semantic-enhanced representation for data, which could help handle the new coming classes, and thereby construct the enhanced semantic objective function. An efficient and effective discrete online optimization algorithm is further proposed for OASIS. Extensive experiments show that our method can exceed the state-of-the-art models. For good reproducibility and benefiting the community, our code and data are already publicly available.
Xiao-Ming Wu 0002, Xin Luo 0006, Yu-Wei Zhan, Chenlu Ding, Zhen-Duo Chen 0001, Xin-Shun Xu
AAAI5
2022 A High-Dimensional Sparse Hashing Framework for Cross-Modal Retrieval
abstract
In recent years, many achievements have been made in improving the performance of supervised cross-modal hashing. However, it remains an open issue on how to fully explore the data information to achieve fine-grained retrieval performance. Most methods employ logical labels or a binary similarity matrix to supervise the hash learning, losing a lot of useful information. From another point of view, the low expressiveness of dense hash code severely limits its preservation of fine-grained data information. With this motivation, in this paper, we propose a high-dimensional sparse hashing framework for cross-modal retrieval, i.e., High-dimensional Sparse Cross-modal Hashing, HSCH for short. It leverages not only high-level semantic labels but also low-level multi-modal features to construct a fine-grained similarity. In particular, based on two well-designed rules, i.e., multi-level and prioritized, it is able to avoid semantic conflicts. Additionally, it leverages the strong power of high-dimensional sparse hash codes to preserve the fine-grained similarity. Then, it efficiently solves the sparse and discrete constraints of sparse hash codes through an efficient discrete optimization algorithm. In light of this, it is much more efficient and scalable to large-scale datasets. More importantly, the computational complexity of HSCH in the retrieval phase is as efficient as those naive hashing methods that use dense hash codes. Moreover, to support online learning scenarios, this paper also extends HSCH into an online version, i.e., HSCH_on. Extensive experiments on three benchmark datasets demonstrate the superiority of our framework compared with some state-of-the-art cross-modal hashing approaches in terms of both accuracy and efficiency.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.2
2022 Fast Cross-Modal Hashing With Global and Local Similarity Embedding
abstract
Recently, supervised cross-modal hashing has attracted much attention and achieved promising performance. To learn hash functions and binary codes, most methods globally exploit the supervised information, for example, preserving an at-least-one pairwise similarity into hash codes or reconstructing the label matrix with binary codes. However, due to the hardness of the discrete optimization problem, they are usually time consuming on large-scale datasets. In addition, they neglect the class correlation in supervised information. From another point of view, they only explore the global similarity of data but overlook the local similarity hidden in the data distribution. To address these issues, we present an efficient supervised cross-modal hashing method, that is, fast cross-modal hashing (FCMH). It leverages not only global similarity information but also the local similarity in a group. Specifically, training samples are partitioned into groups; thereafter, the local similarity in each group is extracted. Moreover, the class correlation in labels is also exploited and embedded into the learning of binary codes. In addition, to solve the discrete optimization problem, we further propose an efficient discrete optimization algorithm with a well-designed group updating scheme, making its computational complexity linear to the size of the training set. In light of this, it is more efficient and scalable to large-scale datasets. Extensive experiments on three benchmark datasets demonstrate that FCMH outperforms some state-of-the-art cross-modal hashing approaches in terms of both retrieval accuracy and learning efficiency.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Rui Li 0090, Xin-Shun Xu
IEEE Trans. Cybern.2
2022 Fine-Grained Hashing With Double Filtering
abstract
Fine-grained hashing is a new topic in the field of hashing-based retrieval and has not been well explored up to now. In this paper, we raise three key issues that fine-grained hashing should address simultaneously, i.e., fine-grained feature extraction, feature refinement as well as a well-designed loss function. In order to address these issues, we propose a novel Fine-graIned haSHing method with a double-filtering mechanism and a proxy-based loss function, FISH for short. Specifically, the double-filtering mechanism consists of two modules, i.e., Space Filtering module and Feature Filtering module, which address the fine-grained feature extraction and feature refinement issues, respectively. Thereinto, the Space Filtering module is designed to highlight the critical regions in images and help the model to capture more subtle and discriminative details; the Feature Filtering module is the key of FISH and aims to further refine extracted features by supervised re- weighting and enhancing. Moreover, the proxy-based loss is adopted to train the model by preserving similarity relationships between data instances and proxy-vectors of each class rather than other data instances, further making FISH much efficient and effective. Experimental results demonstrate that FISH achieves much better retrieval performance compared with state-of-the-art fine-grained hashing methods, and converges very fast. The source code is publicly available: https://github.com/chenzhenduo/FISH.
Zhen-Duo Chen 0001, Xin Luo 0006, Yongxin Wang 0001, Shanqing Guo, Xin-Shun Xu
IEEE Trans. Image Process.1
2021 Weakly-Supervised Online Hashing
abstract
With the rapid development of social websites, recent years have witnessed an explosive growth of social images with user-provided tags. Most existing hashing methods for social image retrieval are batch-based which may violate the nature of social images, i.e., social images are usually generated periodically or collected in a stream fashion. Although there exist many online hashing methods, they either adopt unsupervised learning which ignore the relevant tags, or are designed in the supervised manner which needs high-quality labels. In this paper, to overcome the above limitations, we propose a new method named Weakly-supervised Online Hashing (WOH). In order to learn high-quality hash codes, WOH exploits the weak supervision, i.e., tags, by considering the semantics of tags and removing the noise. Besides, we develop a discrete online optimization algorithm, which is efficient and scalable. Extensive experiments conducted on two real-world datasets demonstrate the superiority of WOH.
Yu-Wei Zhan, Xin Luo 0006, Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin-Shun Xu
ICME5
2021 TEACH: Attention-Aware Deep Cross-Modal Hashing
abstract
Hashing methods for cross-modal retrieval have recently been widely investigated due to the explosive growth of multimedia data. Generally, real-world data is imperfect and has more or less redundancy, making cross-modal retrieval task challenging. However, most existing cross-modal hashing methods fail to deal with the redundancy, leading to unsatisfactory performance on such data. In this paper, to address this issue, we propose a novel cross-modal hashing method, namely aTtEntion-Aware deep Cross-modal Hashing (TEACH). It could perform feature learning and hash-code learning simultaneously. Besides, with designed attention modules for different modalities, one for each, TEACH can effectively highlight the useful information of data while suppressing the redundant information. Extensive experiments on benchmark datasets demonstrate that our method outperforms some state-of-the-art hashing methods in cross-modal retrieval tasks.
Honglei Yao, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
ICMR3
2021 High-Dimensional Sparse Cross-Modal Hashing with Fine-Grained Similarity Embedding
abstract
Recently, with the discoveries in neurobiology, high-dimensional sparse hashing has attracted increasing attention. In contrast with general hashing that generates low-dimensional hash codes, the high-dimensional sparse hashing maps inputs into a higher dimensional space and generates sparse hash codes, achieving superior performance. However, the sparse hashing has not been fully studied in hashing literature yet. For example, how to fully explore the power of sparse coding in cross-modal retrieval tasks; how to discretely solve the binary and sparse constraints so as to avoid the quantization error problem. Motivated by these issues, in this paper, we present an efficient sparse hashing method, i.e., High-dimensional Sparse Cross-modal Hashing, HSCH for short. It not only takes the high-level semantic similarity of data into consideration, but also properly exploits the low-level feature similarity. In specific, we theoretically design a fine-grained similarity with two critical fusion rules. Then we take advantage of sparse codes to embed the fine-grained similarity into the to-be-learnt hash codes. Moreover, an efficient discrete optimization algorithm is proposed to solve the binary and sparse constraints, reducing the quantization error. In light of this, it becomes much more trainable, and the learnt hash codes are more discriminative. More importantly, the retrieval complexity of HSCH is as efficient as general hash methods. Extensive experiments on three widely-used datasets demonstrate the superior performance of HSCH compared with several state-of-the-art cross-modal hashing approaches.
Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
WWW2
2020 SCRATCH: A Scalable Discrete Matrix Factorization Hashing Framework for Cross-Modal Retrieval
abstract
In this paper, we present a novel supervised cross-modal hashing framework, namely Scalable disCRete mATrix faCtorization Hashing (SCRATCH). First, it utilizes collective matrix factorization on original features together with label semantic embedding, to learn the latent representations in a shared latent space. Thereafter, it generates binary hash codes based on the latent representations. During optimization, it avoids using a large n × n similarity matrix and generates hash codes discretely. Besides, based on different objective functions, learning strategy, and features, we further present three models in this framework, i.e., SCRATCH-o, SCRATCH-t, and SCRATCH-d. The first one is a one-step method, learning the hash functions and the binary codes in the same optimization problem. The second is a two-step method, which first generates the binary codes and then learns the hash functions based on the learned hash codes. The third one is a deep version of SCRATCH-t, which utilizes deep neural networks as hash functions. The extensive experiments on two widely used benchmark datasets demonstrate that SCRATCH-o and SCRATCH-t outperform some state-of-the-art shallow hashing methods for cross-modal retrieval. The SCRATCH-d also outperforms some state-of-the-art deep hashing models.
Zhen-Duo Chen 0001, Chuan-Xiang Li, Xin Luo 0006, Liqiang Nie, Wei Zhang 0021, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.1
2019 A Two-Step Cross-Modal Hashing by Exploiting Label Correlations and Preserving Similarity in Both Steps
abstract
In this paper, we present a novel Two-stEp Cross-modal Hashing method, TECH for short, for cross-modal retrieval tasks. As a two-step method, it first learns hash codes based on semantic labels, while preserving the similarity in the original space and exploiting the label correlations in the label space. In the light of this, it is able to make better use of label information and generate better binary codes. In addition, different from other two-step methods that mainly focus on the hash codes learning, TECH adopts a new hash function learning strategy in the second step, which also preserves the similarity in the original space. Moreover, with the help of well designed objective function and optimization scheme, it is able to generate hash codes discretely and scalable for large scale data. To the best of our knowledge, it is the first cross-modal hashing method exploiting label correlations, and also the first two-step hashing model preserving the similarity while leaning hash function. Extensive experiments demonstrate that the proposed approach outperforms some state-of-the-art cross-modal hashing methods.
Zhen-Duo Chen 0001, Yongxin Wang 0001, Huiqiong Li, Xin Luo 0006, Liqiang Nie, Xin-Shun Xu
ACM Multimedia1
2019 DELTA: A deep dual-stream network for multi-label image classification
Wan-Jin Yu, Zhen-Duo Chen 0001, Xin Luo 0006, Wu Liu 0005, Xin-Shun Xu
Pattern Recognit.2
2018 Dual Deep Neural Networks Cross-Modal Hashing
abstract
Recently, deep hashing methods have attracted much attention in multimedia retrieval task. Some of them can even perform cross-modal retrieval. However, almost all existing deep cross-modal hashing methods are pairwise optimizing methods, which means that they become time-consuming if they are extended to large scale datasets. In this paper, we propose a novel tri-stage deep cross-modal hashing method – Dual Deep Neural Networks Cross-Modal Hashing, i.e., DDCMH, which employs two deep networks to generate hash codes for different modalities. Specifically, in Stage 1, it leverages a single-modal hashing method to generate the initial binary codes of textual modality of training samples; in Stage 2, these binary codes are treated as supervised information to train an image network, which maps visual modality to a binary representation; in Stage 3, the visual modality codes are reconstructed according to a reconstruction procedure, and used as supervised information to train a text network, which generates the binary codes for textual modality. By doing this, DDCMH can make full use of inter-modal information to obtain high quality binary codes, and avoid the problem of pairwise optimization by optimizing different modalities independently. The proposed method can be treated as a framework which can extend any single-modal hashing method to perform cross-modal search task. DDCMH is tested on several benchmark datasets. The results demonstrate that it outperforms both deep and shallow state-of-the-art hashing methods.
Zhen-Duo Chen 0001, Wan-Jin Yu, Chuan-Xiang Li, Liqiang Nie, Xin-Shun Xu
AAAI1
2018 Asymmetric Discrete Cross-Modal Hashing
abstract
Recently, cross-modal hashing (CMH) methods have attracted much attention. Many methods have been explored; however, there are still some issues that need to be further considered. 1) How to efficiently construct the correlations among heterogeneous modalities. 2) How to solve the NP-hard optimization problem and avoid the large quantization errors generated by relaxation. 3) How to handle the complex and difficult problem in most CMH methods that simultaneously learning the hash codes and hash functions. To address these challenges, we present a novel cross-modal hashing algorithm, named Asymmetric Discrete Cross-Modal Hashing (ADCH). Specifically, it leverages the collective matrix factorization technique to learn the common latent representations while preserving not only the cross-correlation from different modalities but also the semantic similarity. Instead of relaxing the binary constraints, it generates the hash codes directly using an iterative optimization algorithm proposed in this work. Based the learnt hash codes, ADCH further learns a series of binary classifiers as hash functions, which is flexible and effective. Extensive experiments are conducted on three real-world datasets. The results demonstrate that ADCH outperforms several state-of-the-art cross-modal hashing baselines.
Xin Luo 0006, Peng-Fei Zhang 0001, Zhen-Duo Chen 0001, Hua-Junjie Huang, Xin-Shun Xu
ICMR4
2018 SCRATCH: A Scalable Discrete Matrix Factorization Hashing for Cross-Modal Retrieval
abstract
In recent years, many hashing methods have been proposed for the cross-modal retrieval task. However, there are still some issues that need to be further explored. For example, some of them relax the binary constraints to generate the hash codes, which may generate large quantization error. Although some discrete schemes have been proposed, most of them are time-consuming. In addition, most of the existing supervised hashing methods use an n x n similarity matrix during the optimization, making them unscalable. To address these issues, in this paper, we present a novel supervised cross-modal hashing method---Scalable disCRete mATrix faCtorization Hashing, SCRATCH for short. It leverages the collective matrix factorization on the kernelized features and the semantic embedding with labels to find a latent semantic space to preserve the intra- and inter-modality similarities. In addition, it incorporates the label matrix instead of the similarity matrix into the loss function. Based on the proposed loss function and the iterative optimization algorithm, it can learn the hash functions and binary codes simultaneously. Moreover, the binary codes can be generated discretely, reducing the quantization error generated by the relaxation scheme. Its time complexity is linear to the size of the dataset, making it scalable to large-scale datasets. Extensive experiments on three benchmark datasets, namely, Wiki, MIRFlickr-25K, and NUS-WIDE, have verified that our proposed SCRATCH model outperforms several state-of-the-art unsupervised and supervised hashing methods for cross-modal retrieval.
Chuan-Xiang Li, Zhen-Duo Chen 0001, Peng-Fei Zhang 0001, Xin Luo 0006, Liqiang Nie, Wei Zhang 0021, Xin-Shun Xu
ACM Multimedia2
2018 Fast Scalable Supervised Hashing
abstract
Despite significant progress in supervised hashing, there are three common limitations of existing methods. First, most pioneer methods discretely learn hash codes bit by bit, making the learning procedure rather time-consuming. Second, to reduce the large complexity of the n by n pairwise similarity matrix, most methods apply sampling strategies during training, which inevitably results in information loss and suboptimal performance; some recent methods try to replace the large matrix with a smaller one, but the size is still large. Third, among the methods that leverage the pairwise similarity matrix, most of them only encode the semantic label information in learning the hash codes, failing to fully capture the characteristics of data. In this paper, we present a novel supervised hashing method, called Fast Scalable Supervised Hashing (FSSH), which circumvents the use of the large similarity matrix by introducing a pre-computed intermediate term whose size is independent with the size of training data. Moreover, FSSH can learn the hash codes with not only the semantic information but also the features of data. Extensive experiments on three widely used datasets demonstrate its superiority over several state-of-the-art methods in both accuracy and scalability. Our experiment codes are available at: https://lcbwlx.wixsite.com/fssh.
Xin Luo 0006, Liqiang Nie, Xiangnan He 0001, Zhen-Duo Chen 0001, Xin-Shun Xu
SIGIR5
2017 Improving Hashing by Leveraging Multiple Layers of Deep Networks
Xin Luo 0006, Zhen-Duo Chen 0001, Gao-Yuan Du, Xin-Shun Xu
ICONIP (1)2