EDBT 2026 Demo / reviewers in the wild / expert
Chao Zhang 0078
dblp:94/3019-78
· DBLP profile ↗
28ranked-venue papers
13as first author
28since 2021 · last 2026
0000-0002-5929-0853ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Online Cross-Modal Hashing with Expanding Label SpaceabstractDue to the continuous increase of multimedia data on the internet, online hashing has garnered considerable attention for handling multi-modal data streams. However, most existing online hashing approaches focus solely on data growth of samples, overlooking the dynamics of classes. In this paper, we simultaneously address the challenges of both sample-level and class-level growth, and propose a novel Online Hashing method with Expanding Label Space (OH-ELS) for cross-modal retrieval. In OH-ELS, multi-modal data arrives continuously, and incoming data may introduce new classes. To avoid catastrophic forgetting, we transfer the historical knowledge at both the sample and class levels. At the sample-level, a small subset of anchor codes from old data are replayed to preserve the similarities between new data and old data. At the class-level, a consistency regularizer is applied to new classifiers to leverage the priors of historical classes. To ensure both efficiency and accuracy, a discrete optimization algorithm is proposed to solve the binary-constrained optimization problem without relaxation. Experimental results illustrate the effectiveness and superiority of OH-ELS in class-incremental cross-modal retrieval compared with the state-of-the-art methods. Wentao Fan 0003, Chao Zhang 0078, Chunlin Chen 0001, Huaxiong Li |
AAAI | 2 |
| 2026 | Semantic-Aware Feature Enhancement for Partial Label LearningabstractPartial label learning (PLL) aims to learn from the data where each instance is associated with a candidate label set, with only one being valid. Most existing approaches are designed to eliminate noisy labels and use the remaining reliable ones for model training, following a label-centric learning paradigm. In this paper, we propose a new PLL method called Semantic-Aware Feature Enhancement (SAFE), which tackles the problem through a novel feature-centric learning paradigm. SAFE presumes that the candidate labels are correct while the observed features are partial, and thus seeks to recover the underlying missing features. In this manner, a desired predictive model is constructed by integrating the observed and recovered features, which are responsible for predicting the true label and the remaining candidate labels, respectively. To ensure the quality of recovered features, SAFE jointly explores the intrinsic topological structures via dynamic graphs in both feature and label spaces as guidance for semantic-aware feature enhancement. Extensive experimental results on some popular datasets demonstrate the effectiveness and superiority of the proposed method over state-of-the-art PLL approaches. Haowei Mei, Chao Zhang 0078, Wentao Fan 0003, Xiuyi Jia, Chunlin Chen 0001, Huaxiong Li |
AAAI | 2 |
| 2026 | Semantic-Augmented Image Clustering via Adaptive Multi-Modal CollaborationabstractImage clustering is a fundamental task in unsupervised visual learning. While recent self-supervised methods have explored various pretext tasks to generate supervision signals for clustering, they typically depend exclusively on raw images, resulting in insufficient supervision signals that are inherently constrained by limited visual semantics. In this paper, we propose a novel Semantic-Augmented image Clustering (SAC) method, which transcends the inherent limitations of purely visual representations through the integration of external knowledge. Specifically, SAC utilizes Vision-Language pre-trained Models (VLMs) to flexibly generate textual descriptions for each image, providing external semantic cues to supplement the visual information. By integrating both visual and textual information, SAC achieves image clustering through a multi-modal learning framework. To mitigate the negative impact of inaccurate textual information, SAC designs an uncertainty-driven adaptive weighting mechanism that explores both intra-modal and inter-modal neighborhood structures, and incorporates the adaptive weights into intra-modal and inter-modal contrastive learning, which improves the robustness against noisy image-text correspondences. Experiments on several popular datasets demonstrate the superiority of SAC compared to state-of-the-art methods. Chao Zhang 0078, Deng Xu, Hong Yu 0007, Chunlin Chen 0001, Huaxiong Li |
AAAI | 2 |
| 2025 | Fast Incomplete Multi-view Clustering with Adaptive Similarity Completion and ReconstructionabstractRecently, anchor-based incomplete multi-view clustering (IMVC) has been widely adopted for fast clustering, but most existing approaches still encounter some issues: (1) They generally rely on the observed samples to construct anchor graphs, ignoring the potentially useful information of missing instances. (2) Most methods attempt to learn a consensus anchor graph, failing to fully excavate the complementary information and high-order correlations across views. (3) They generally apply post-processing on learned anchor graph to seek latent embeddings, making them not globally-optimal. To address these issues, this paper proposes a novel fast IMVC approach with Adaptive Similarity Completion and Reconstruction (ASCR), which unifies anchor learning, anchor-sample similarity construction and completion, and latent multi-view embedding learning in a joint framework. Specifically, ASCR learns an anchor-sample similarity graph for each view, and the missing values are fulfilled to mitigate the adverse effects. To explore the consistent and complementary information across views, ASCR simultaneously seeks the view-specific anchor embeddings and sample embeddings in a latent subspace by similarity reconstruction, which not only preserves the semantic information into latent embeddings but also enhances the low-rank property of similarity graphs, achieving a reliable graph completion process. Furthermore, the high-order cross-view correlations are explored with tensor-based regularization. Extensive experimental results demonstrate the superiority and efficiency of ASCR compared with SOTA approaches. Deng Xu, Chao Zhang 0078, Cong Guo 0008, Chunlin Chen 0001, Huaxiong Li |
AAAI | 2 |
| 2025 | Semi-supervised Multi-view Clustering with Active ConstraintsabstractMulti-view clustering has attracted increasing attention in recent years. However, most existing multi-view clustering approaches are performed in a purely unsupervised manner, while ignoring the valuable weak supervision information that can be obtained (e.g., active query) in many real applications. This paper considers the weak pairwise constraints among samples to enhance the clustering performance, and proposes a Semi-supervised Multi-view Clustering method with Active Constraints, SMCAC for short. SMCAC consists of two stages, clustering (C-stage) and active query (A-stage). In the C-stage, we design a tensor based multi-view graph learning model equipped with sample pairwise constraints regularization to facilitate the discriminative graph learning and fusion. An effective optimization algorithm based on alternating direction minimization is devised to solve the clustering model. In the A-stage, the most uncertain or difficult sample pairs are actively selected to query the constraints, based on the divergence of multi-view similarities learned in the C-stage. The two processes alternate iteratively until the maximum number of queries is reached. Extensive experiments on several popular datasets well validate the effectiveness of the proposed method. Chao Zhang 0078, Deng Xu, Chunlin Chen 0001, Huaxiong Li |
KDD (1) | 1 |
| 2025 | Online Cross-Modal Hashing with Multi-Level MemoryabstractOnline cross-modal hashing has recently gained significant attention due to its remarkable capability to handle cross-modal streaming data retrieval. Despite promising progress, existing methods still face challenges in fully exploiting the intricate relations across heterogeneous modalities and streaming data chunks, limiting the retrieval performance. In this paper, a novel Online Cross-modal Hashing method with Multi-level Memory (OCH-MM) is proposed. OCH-MM captures the cross-modal consistency and sample semantic correlations for discrete hash learning with latent feature disentanglement, and designs a multi-level memory framework for effective knowledge transfer. Specifically, for discriminative hash learning, OCH-MM maps the multi-modal data into a latent feature space that is further disentangled into a common Hamming space and a modality-specific feature space. The semantic correlations among samples are also preserved into discrete hash codes without relaxation in a nonlinear manner. For effectively learning from streaming data, OCH-MM designs an intra-space feature association memory, an inter-space feature association memory, and a hash codes memory, which encode the historical feature correlations within original multi-modal spaces, the feature correlations between original and latent space, and a subset of hash codes, respectively. By dynamically updating and utilizing the multi-level memory, the data correlations between different chunks are well explored and the historical knowledge is effectively reused to guide future learning. The proposed model is solved by an efficient discrete optimization algorithm. Experimental results on three benchmark datasets demonstrate that our proposed method achieves better retrieval accuracy over the state-of-the-art baselines. Wentao Fan 0003, Chao Zhang 0078, Chunlin Chen 0001, Huaxiong Li |
ACM Multimedia | 2 |
| 2025 | Multi-View Clustering With Incremental Instances and ViewsabstractMulti-view clustering (MVC) has attracted increasing attention with the emergence of various data collected from multiple sources. In real-world dynamic environment, instances are continually gathered, and the number of views expands as new data sources become available. Learning for such simultaneous increment of instances and views, particularly in unsupervised scenarios, is crucial yet underexplored. In this paper, we address this problem by proposing a novel MVC method with Incremental Instances and Views, MVC-IIV for short. MVC-IIV contains two stages, an initial stage and an incremental stage. In the initial stage, a basic latent multi-view subspace clustering model is constructed to handle existing data, which can be viewed as traditional static MVC. In the incremental stage, the previously trained model is reused to guide learning for newly arriving instances with new views, transferring historical knowledge while avoiding redundant computations. In specific, we design and reuse two modules, i.e., multi-view embedding module for low-dimensional representation learning, and consensus centroids module for cluster probability learning. By adding consistency regularization on the two modules, the knowledge acquired from previous data is used, which not only enhances the exploration within current data batch, but also extracts the between-batch data correlations. The proposed model can be efficiently solved with linear space and time complexity. Extensive experiments demonstrate the effectiveness and efficiency of our method compared with the state-of-the-art approaches. Chao Zhang 0078, Zhi Wang 0001, Xiuyi Jia, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
IEEE Trans. Image Process. | 1 |
| 2025 | Label Distribution Guided Hashing for Cross-Modal RetrievalabstractHashing methods have recently attracted extensive attention in cross-modal retrieval. Most supervised hashing methods attempt to preserve the semantic information into hash codes by leveraging the original logical label matrix. However, they generally treat all labels equally, and ignore the relative significance of different labels due to the variety of data features. In this article, we argue that exploring the relative importance of labels benefits the enhancement of semantic information, and we propose a novel LAbel Distribution Guided Hashing (LADH) method for cross-modal retrieval. In particular, LADH first learns a feature-induced label distribution for each sample to weigh different labels, which leverages the multi-modal feature information to enrich the semantic label information. By jointly using the learned label distributions and multi-modal features, the latent representation and hash codes are obtained with multi-modal feature selection and enhanced semantic similarities embedded. An efficient algorithm is designed to solve the proposed method whose time complexity is linear to the number of the training instances. Experimental results on several public benchmark datasets verify the effectiveness and efficiency of our method compared with the state-of-the-art methods. Fatang Lei, Chao Zhang 0078, Huaxiong Li, Yang Gao 0001, Chunlin Chen 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2025 | Fast Disentangled Slim Tensor Learning for Multi-View ClusteringabstractTensor-based multi-view clustering has recently received significant attention due to its exceptional ability to explore cross-view high-order correlations. However, most existing methods still encounter some limitations. (1) Most of them explore the correlations among different affinity matrices, making them unscalable to large-scale data. (2) Although some methods address it by introducing bipartite graphs, they may result in sub-optimal solutions caused by an unstable anchor selection process. (3) They generally ignore the negative impact of latent semantic-unrelated information in each view. To tackle these issues, we propose a new approach termed fast Disentangled Slim Tensor Learning (DSTL) for multi-view clustering. Instead of focusing on the multi-view graph structures, DSTL directly explores the high-order correlations among multi-view latent semantic representations based on matrix factorization. To alleviate the negative influence of feature redundancy, inspired by robust PCA, DSTL disentangles the latent low-dimensional representation into a semantic-unrelated part and a semantic-related part for each view. Subsequently, two slim tensors are constructed with tensor-based regularization. To further enhance the quality of feature disentanglement, the semantic-related representations are aligned across views through a consensus alignment indicator. Our proposed model is computationally efficient and can be solved effectively. Extensive experiments demonstrate the superiority and efficiency of DSTL over state-of-the-art approaches. Deng Xu, Chao Zhang 0078, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
IEEE Trans. Multim. | 2 |
| 2025 | Three-Stage Semisupervised Cross-Modal Hashing With Pairwise Relations ExploitationabstractHashing methods have sparked a great revolution in cross-modal retrieval due to the low cost of storage and computation. Benefiting from the sufficient semantic information of labeled data, supervised hashing methods have shown better performance compared with unsupervised ones. Nevertheless, it is expensive and labor intensive to annotate the training samples, which restricts the feasibility of supervised methods in real applications. To deal with this limitation, a novel semisupervised hashing method, i.e., three-stage semisupervised hashing (TS3H) is proposed in this article, where both labeled and unlabeled data are seamlessly handled. Different from other semisupervised approaches that learn the pseudolabels, hash codes, and hash functions simultaneously, the new approach is decomposed into three stages as the name implies, in which all of the stages are conducted individually to make the optimization cost-effective and precise. Specifically, the classifiers of different modalities are learned via the provided supervised information to predict the labels of unlabeled data at first. Then, hash code learning is achieved with a simple but efficient scheme by unifying the provided and the newly predicted labels. To capture the discriminative information and preserve the semantic similarities, we leverage pairwise relations to supervise both classifier learning and hash code learning. Finally, the modality-specific hash functions are obtained by transforming the training samples to the generated hash codes. The new approach is compared with the state-of-the-art shallow and deep cross-modal hashing (DCMH) methods on several widely used benchmark databases, and the experiment results verify its efficiency and superiority. Wentao Fan 0003, Chao Zhang 0078, Huaxiong Li, Xiuyi Jia, Guoyin Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Learning Cluster-Wise Anchors for Multi-View ClusteringabstractDue to its effectiveness and efficiency, anchor based multi-view clustering (MVC) has recently attracted much attention. Most existing approaches try to adaptively learn anchors to construct an anchor graph for clustering. However, they generally focus on improving the diversity among anchors by using orthogonal constraint and ignore the underlying semantic relations, which may make the anchors not representative and discriminative enough. To address this problem, we propose an adaptive Cluster-wise Anchor learning based MVC method, CAMVC for short. We first make an anchor cluster assumption that supposes the prior cluster structure of target anchors by pre-defining a consensus cluster indicator matrix. Based on the prior knowledge, an explicit cluster structure of latent anchors is enforced by learning diverse cluster centroids, which can explore both inter-cluster diversity and intra-cluster consistency of anchors, and improve the subspace representation discrimination. Extensive results demonstrate the effectiveness and superiority of our proposed method compared with some state-of-the-art MVC approaches. Chao Zhang 0078, Xiuyi Jia, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
AAAI | 1 |
| 2024 | Continual Multi-View Clustering with Consistent Anchor Guidance
Chao Zhang 0078, Deng Xu, Xiuyi Jia, Chunlin Chen 0001, Huaxiong Li |
IJCAI | 1 |
| 2024 | Transparent Projection Networks for Interpretable Image RecognitionabstractThe absence of interpretability in deep convolutional neural networks (CNNs) leaves us with the dilemma of how to explain their decision mechanism in terms of human-understandable semantics due to the black-box nature. In this paper, we present a novel network architecture called Transparent Projection Networks (ProNets) for interpretable image recognition, by constructing a meaningful latent space through transparent projection learning. Specifically, we exploit the expressiveness of NNs to learn a group of input-dependent pixel-level weights, which project input images into the latent space, thus allowing for a transparent weighted connection between raw pixel-level features and latent high-level representations. We organize our layers in light of ResNets with valid modifications, where we propose to develop new block designs to learn projection weights and use shortcuts to enable linear projection flexibly every few layers. Our model inherits the properties of linear models, decomposing the output into a linear combination of contributions from each input feature, which can deliver transparent interpretations of the reasoning process along with high visual quality. Further, as an interpretable model for image classification, through experimental studies on several benchmark datasets, ProNets achieve competitive accuracy results with classic CNNs like VGG and ResNets, while demonstrating a remarkably high level of interpretability. Chao Zhang 0078, Chunlin Chen 0001, Huaxiong Li |
IJCNN | 2 |
| 2024 | Fast Dynamic Multi-view Clustering with semantic-consistency inheritance
Shuyao Lu, Deng Xu, Chao Zhang 0078, Zhangqing Zhu |
Knowl. Based Syst. | 3 |
| 2024 | Multi-level graph regularized robust multi-modal feature selection for Alzheimer's disease classification
Chao Zhang 0078, Wentao Fan 0003, Huaxiong Li, Chunlin Chen 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Learning latent disentangled embeddings and graphs for multi-view clustering
Chao Zhang 0078, Haoxing Chen, Huaxiong Li, Chunlin Chen 0001 |
Pattern Recognit. | 1 |
| 2024 | Joint Projection Learning and Tensor Decomposition-Based Incomplete Multiview ClusteringabstractIncomplete multiview clustering (IMVC) has received increasing attention since it is often that some views of samples are incomplete in reality. Most existing methods learn similarity subgraphs from original incomplete multiview data and seek complete graphs by exploring the incomplete subgraphs of each view for spectral clustering. However, the graphs constructed on the original high-dimensional data may be suboptimal due to feature redundancy and noise. Besides, previous methods generally ignored the graph noise caused by the interclass and intraclass structure variation during the transformation of incomplete graphs and complete graphs. To address these problems, we propose a novel joint projection learning and tensor decomposition (JPLTD)-based method for IMVC. Specifically, to alleviate the influence of redundant features and noise in high-dimensional data, JPLTD introduces an orthogonal projection matrix to project the high-dimensional features into a lower-dimensional space for compact feature learning. Meanwhile, based on the lower-dimensional space, the similarity graphs corresponding to instances of different views are learned, and JPLTD stacks these graphs into a third-order low-rank tensor to explore the high-order correlations across different views. We further consider the graph noise of projected data caused by missing samples and use a tensor-decomposition-based graph filter for robust clustering. JPLTD decomposes the original tensor into an intrinsic tensor and a sparse tensor. The intrinsic tensor models the true data similarities. An effective optimization algorithm is adopted to solve the JPLTD model. Comprehensive experiments on several benchmark datasets demonstrate that JPLTD outperforms the state-of-the-art methods. The code of JPLTD is available at https://github.com/weilvNJU/JPLTD. Chao Zhang 0078, Huaxiong Li, Xiuyi Jia, Chunlin Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Low-Rank Tensor Regularized Views Recovery for Incomplete Multiview ClusteringabstractIn real applications, it is often that the collected multiview data contain missing views. Most existing incomplete multiview clustering (IMVC) methods cannot fully utilize the underlying information of missing data or sufficiently explore the consistent and complementary characteristics. In this article, we propose a novel Low-rAnk Tensor regularized viEws Recovery (LATER) method for IMVC, which jointly reconstructs and utilizes the missing views and learns multilevel graphs for comprehensive similarity discovery in a unified model. The missing views are recovered from a common latent representation, and the recovered views conversely improve the learning of shared patterns. Based on the shared subspace representations and recovered complete multiview data, the multilevel graphs are learned by self-representation to fully exploit the consistent and complementary information among views. Besides, a tensor nuclear norm regularizer is introduced to pursue the global low-rank property and explore the interview correlations. An alternating direction minimization algorithm is presented to optimize the proposed model. Moreover, a new initialization method is proposed to promote the effectiveness of our method for latent representation learning and missing data recovery. Extensive experiments demonstrate that our method outperforms the state-of-the-art approaches. Chao Zhang 0078, Huaxiong Li, Caihua Chen, Xiuyi Jia, Chunlin Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Enhanced Tensor Low-Rank and Sparse Representation Recovery for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC) has attracted remarkable attention due to the emergence of multi-view data with missing views in real applications. Recent methods attempt to recover the missing information to address the IMVC problem. However, they generally cannot fully explore the underlying properties and correlations of data similarities across views. This paper proposes a novel Enhanced Tensor Low-rank and Sparse Representation Recovery (ETLSRR) method, which reformulates the IMVC problem as a joint incomplete similarity graphs learning and complete tensor representation recovery problem. Specifically, ETLSRR learns the intra-view similarity graphs and constructs a 3-way tensor by stacking the graphs to explore the inter-view correlations. To alleviate the negative influence of missing views and data noise, ETLSRR decomposes the tensor into two parts: a sparse tensor and an intrinsic tensor, which models the noise and underlying true data similarities, respectively. Both global low-rank and local structured sparse characteristics of the intrinsic tensor are considered, which enhances the discrimination of similarity matrix. Moreover, instead of using the convex tensor nuclear norm, ETLSRR introduces a generalized non-convex tensor low-rank regularization to alleviate the biased approximation. Experiments on several datasets demonstrate the effectiveness of our method compared with the state-of-the-art methods. Chao Zhang 0078, Huaxiong Li, Zizheng Huang, Yang Gao 0001, Chunlin Chen 0001 |
AAAI | 1 |
| 2023 | Model-Aware Contrastive Learning: Towards Escaping the DilemmasabstractContrastive learning (CL) continuously achieves significant breakthroughs across multiple domains. However, the most common InfoNCE-based methods suffer from some dilemmas, such as uniformity-tolerance dilemma (UTD) and gradient reduction, both of which are related to a $\mathcal{P}_{ij}$ term. It has been identified that UTD can lead to unexpected performance degradation. We argue that the fixity of temperature is to blame for UTD. To tackle this challenge, we enrich the CL loss family by presenting a Model-Aware Contrastive Learning (MACL) strategy, whose temperature is adaptive to the magnitude of alignment that reflects the basic confidence of the instance discrimination task, then enables CL loss to adjust the penalty strength for hard negatives adaptively. Regarding another dilemma, the gradient reduction issue, we derive the limits of an involved gradient scaling factor, which allows us to explain from a unified perspective why some recent approaches are effective with fewer negative samples, and summarily present a gradient reweighting to escape this dilemma. Extensive remarkable empirical results in vision, sentence, and graph modality validate our approach’s general improvement for representation learning and downstream tasks. Zizheng Huang, Haoxing Chen, Ziqi Wen, Chao Zhang 0078, Huaxiong Li, Bo Wang 0027, Chunlin Chen 0001 |
ICML | 4 |
| 2023 | Robust Spectral Embedding Completion Based Incomplete Multi-view ClusteringabstractGraph based methods have been widely used in incomplete multi-view clustering (IMVC). Most recent methods try to fill the original missing samples or incomplete affinity matrices to obtain a complete similarity graph for the subsequent spectral clustering. However, recovering the original high-dimensional data or complete n X n similarity matrix is usually time-consuming and noise-sensitive. Besides, they generally separate the cluster indicator learning into an individual step, which may result in sub-optimal graphs or spectral embeddings for clustering. To address these problems, this paper proposes a robust Spectral Embedding Completion based IMVC (SEC-IMVC) method, which incorporates spectral embedding completion and discrete cluster indicator learning into a unified framework. SEC-IMVC performs completion on spectral embeddings, and the embedding noise is eliminated to reduce the negative influence of original data noise. The discrete cluster indicator matrix is seamlessly learned by using spectral rotation, and it can explore the first-order feature consistency among different views. To further improve the completion robustness, the second-order correlation consistency is also captured by pairwise relations alignment. We compare our method with some state-of-the-art approaches on several datasets, and the experimental results show the effectiveness and advantages of our method. Chao Zhang 0078, Jingwen Wei, Bo Wang 0027, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
ACM Multimedia | 1 |
| 2023 | A robust mixed error coding method based on nonconvex sparse representation
Chao Zhang 0078, Huaxiong Li, Bo Wang 0027, Chunlin Chen 0001 |
Inf. Sci. | 2 |
| 2023 | Adaptive Label Correlation Based Asymmetric Discrete Hashing for Cross-Modal RetrievalabstractHashing methods have captured much attention for cross-modal retrieval in recent years. Most existing approaches mainly focus on preserving the semantic similarity across heterogeneous modalities in a shared Hamming subspace, while the label information and potential correlations of multi-label semantics are not fully excavated. In this article, a novel Adaptive Label correlation based asymmEtric Cross-modal Hashing method, i.e., ALECH, is proposed for cross-modal retrieval. ALECH decomposes hash learning into two steps, hash codes learning and hash functions learning. For hash codes learning, the high-order semantic label correlations are adaptively exploited to guide the latent feature learning, while simultaneously generating the binary codes in a discrete manner. The asymmetric strategy is utilized to connect the latent feature space and Hamming space, and preserve the pairwise semantic similarity. Different from other two-step methods that directly adopt simple least-squares regression to learn hash functions based on binary codes, ALECH leverages both hash codes and semantic labels for hash functions learning which further preserves the similarity. Experiments on several benchmark datasets demonstrate that the proposed ALECH method outperforms the state-of-the-art cross-hashing methods. Huaxiong Li, Chao Zhang 0078, Xiuyi Jia, Yang Gao 0001, Chunlin Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Weakly-Supervised Enhanced Semantic-Aware Hashing for Cross-Modal RetrievalabstractOwing to its query and storage efficiency, hash learning has sparked much interest for Cross-Modal Retrieval (CMR) task. Previous literatures have proved the superiority of supervised Cross-Modal Hashing (CMH) methods over unsupervised ones. Nevertheless, most existing supervised CMH methods still suffer from some limitations: 1) it is assumed that the observed labels of training data are complete and accurate, which may be impractical due to the missing and wrong class assignments in real applications, and 2) the semantic information is not fully excavated, especially for the semantic correlations among labels. To address these issues, this paper proposes a Weakly-supervised enhAnced Semantic-aware Hashing (WASH) method which simultaneously estimates the label noises and performs enhanced semantic-aware hash learning. WASH employs the low-rank and sparse decomposition to alleviate the label noises, and a high-level semantic factor as well as a semantic correlation matrix is obtained by low-rank factorization on the noise-reduced labels. The low-rank semantic factors and multi-modal features are jointly factorized into a common subspace to reduce the heterogeneity gaps, so as to enhance the semantic awareness of shared representation. In this way, the hash codes can be obtained by binarizing the shared representation with pairwise semantic similarity preserved. Experiments on several benchmark datasets verify the effectiveness of the proposed method in comparison with the state-of-the-art CMH approaches. Chao Zhang 0078, Huaxiong Li, Yang Gao 0001, Chunlin Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Adaptive Marginalized Semantic Hashing for Unpaired Cross-Modal RetrievalabstractIn recent years, Cross-Modal Hashing (CMH) has attracted much attention due to its fast query speed and efficient storage. Previous studies have achieved promising results for Cross-Modal Retrieval (CMR) by discovering discriminative hash codes and modality-specific hash functions. Nonetheless, most existing CMR works are subjected to some restrictions: 1) It is assumed that data of different modalities are fully paired, which is impractical in real applications due to sample missing and false data alignment, and 2) binary regression targets including the label matrix and binary codes are too rigid to effectively learn semantic-preserving hash codes and hash functions. To address these problems, this paper proposes an Adaptive Marginalized Semantic Hashing (AMSH) method which not only enhances the discrimination of latent representations and hash codes by adaptive margins, but can also be used for both paired and unpaired CMR. As a two-step method, in the first step, AMSH generates semantic-aware modality-specific latent representations with adaptively marginalized labels, thereby enlarging the distances between different classes, and exploiting the labels to preserve the inter-modal and intra-modal semantic similarities into latent representations and hash codes. In the second step, adaptive margin matrices are embedded into the hash codes, and enlarge the gaps between positive and negative bits, which improves the discrimination and robustness of hash functions. On this basis, AMSH generates similarity-preserving hash codes and robust hash functions without the strict one-to-one data correspondence requirement. Experiments are conducted on several benchmark datasets to demonstrate the superiority and flexibility of AMSH over some state-of-the-art CMR methods. The source code is available athttps://github.com/LKYLKYZ/AMSH. Kaiyi Luo, Chao Zhang 0078, Huaxiong Li, Xiuyi Jia, Chunlin Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Enhanced Group Sparse Regularized Nonconvex Regression for Face RecognitionabstractRegression analysis based methods have shown strong robustness and achieved great success in face recognition. In these methods, convex$l_1$-norm and nuclear norm are usually utilized to approximate the$l_0$-norm and rank function. However, such convex relaxations may introduce a bias and lead to a suboptimal solution. In this paper, we propose a novel Enhanced Group Sparse regularized Nonconvex Regression (EGSNR) method for robust face recognition. An upper bounded nonconvex function is introduced to replace$l_1$-norm for sparsity, which alleviates the bias problem and adverse effects caused by outliers. To capture the characteristics of complex errors, we propose a mixed model by combining$\gamma$-norm and matrix$\gamma$-norm induced from the nonconvex function. Furthermore, an$l_{2,\gamma }$-norm based regularizer is designed to directly seek the interclass sparsity or group sparsity instead of traditional$l_{2,1}$-norm. The locality of data, i.e., the distance between the query sample and multi-subspaces, is also taken into consideration. This enhanced group sparse regularizer enables EGSNR to learn more discriminative representation coefficients. Comprehensive experiments on several popular face datasets demonstrate that the proposed EGSNR outperforms the state-of-the-art regression based methods for robust face recognition. Chao Zhang 0078, Huaxiong Li, Chunlin Chen 0001, Xianzhong Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Locality-Constrained Discriminative Matrix Regression for Robust Face IdentificationabstractRegression-based methods have been widely applied in face identification, which attempts to approximately represent a query sample as a linear combination of all training samples. Recently, a matrix regression model based on nuclear norm has been proposed and shown strong robustness to structural noises. However, it may ignore two important issues: the label information and local relationship of data. In this article, a novel robust representation method called locality-constrained discriminative matrix regression (LDMR) is proposed, which takes label information and locality structure into account. Instead of focusing on the representation coefficients, LDMR directly imposes constraints on representation components by fully considering the label information, which has a closer connection to identification process. The locality structure characterized by subspace distances is used to learn class weights, and the correct class is forced to make more contribution to representation. Furthermore, the class weights are also incorporated into a competitive constraint on the representation components, which reduces the pairwise correlations between different classes and enhances the competitive relationships among all classes. An iterative optimization algorithm is presented to solve LDMR. Experiments on several benchmark data sets demonstrate that LDMR outperforms some state-of-the-art regression-based methods. Chao Zhang 0078, Huaxiong Li, Chunlin Chen 0001, Xianzhong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Pairwise Relations Oriented Discriminative RegressionabstractLinear Regression (LR) is a popular and effective technique in pattern recognition area, which aims to find a transform matrix between source data and target data (usually label matrix). However, a binary zero-one label matrix may be too strict and inappropriate for regression. Besides, directly projecting source data to target data by one transform matrix may lose some intrinsic data information. To address these issues, this paper proposes a novel Pairwise Relations oriented Discriminative Regression (PRDR) method. In PRDR, the source data is regressed into a latent space instead of label space. To supervise the discriminative projection learning, the pairwise relations in source data space and label space are exploited in the latent space simultaneously. The pairwise label relations are transferred into the latent subspace by solving a distance-distance difference minimization problem, and the intraclass instance relations are also preserved in latent space. These two constraints ensure the pairwise similarity of data points after transformation which is beneficial for classification. By further enlarging the margins between true and false classes, PRDR is extended to a robust version, i.e., R-PRDR. An efficient algorithm is presented to solve the PRDR model. Extensive experiments on several popular image datasets demonstrate the effectiveness and efficiency of the proposed method compared with some state-of-the-art regression approaches. Chao Zhang 0078, Huaxiong Li, Chunlin Chen 0001, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |