EDBT 2026 Demo / reviewers in the wild / expert
Shizhe Hu
dblp:208/4268
· DBLP profile ↗
43ranked-venue papers
15as first author
36since 2021 · last 2026
0000-0003-1301-2396ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 8 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interest-driven Deep Multi-modal ClusteringabstractDeep multi-modal clustering fully learns semantically consistent and discriminative cluster representations between multiple modalities in an unlabeled manner. However, existing methods treat all samples equally, ignoring varying sample quality, which limits clustering performance. Inspired by the concept of interest in the recommendation system, we propose a novel interest-driven deep multi-modal clustering (IDMC) framework. It designs a new paradigm to quantify the importance of each sample base on the attention it receives from other samples, which called interest value. This value jointly captures the local geometric structure through the Euclidean distance in feature space and the consistency of pseudo-labels. Then, we design a novel adaptive Bayesian fusion mechanism to dynamically balance the prior features and self-supervisory signals to ensure confidence-based sample importance estimation. Furthermore, we introduce a median normalization constraint and a label consistency constraint to further refine the construction of the interest value. By embedding this interest-guided value into representation learning and cluster optimization, IDMC focuses on the samples with the most information and the most stable semantics, thereby enhancing the performance of multi-modal representation learning. Extensive experiments verify that IDMC is superior to existing state-of-the-art methods in multiple evaluation metrics. Guoliang Zou, Tongji Chen, Sijia Li 0003, Yangdong Ye, Shizhe Hu |
AAAI | 6 |
| 2026 | Cross-modal information propagation for contrastive multi-modal clustering
Tongji Chen, Guoliang Zou, Shizhe Hu, Yangdong Ye |
Inf. Process. Manag. | 3 |
| 2026 | Reliable continual multi-modal clustering
Guoliang Zou, Shizhe Hu, Sijia Li 0003, Tongji Chen, Yangdong Ye |
Pattern Recognit. | 2 |
| 2026 | To the Best of Trust: Full-Stage Trusted Multi-Modal ClusteringabstractMulti-modal clustering aims to integrate complementary information from different modalities to uncover latent consistent structures and improve clustering performance. However, existing methods mainly rely on predictive (result) uncertainty to improve robustness, while often neglecting aleatoric (data) uncertainty introduced by sample noise and epistemic (model) uncertainty induced by model parameters and structural variations. To this end, we propose a novel Full-Stage Trusted Multi-modal Clustering (FSTMC) method. To the best of trust, we jointly utilize aleatoric, epistemic, and predictive uncertainties to optimize the model and learn more reliable feature representations and clustering results. In the representation learning phase, probabilistic modeling is used to capture stable latent representations and estimate aleatoric uncertainty, while structured random perturbations are present to estimate epistemic uncertainty. In the clustering stage, instead of conventional feature-level fusion, we design an evidence-based fusion strategy, where soft labels from each modality are first mapped into categorical evidence while cluster distributions are parameterized via a Dirichlet model, with finally dynamic multi-modal fusion achieved by Dempster-Shafer theory. To mitigate overconfidence and modal conflicts, prior constraints guided by aleatoric and epistemic uncertainty are imposed, resulting in calibrated predictive uncertainty. Finally, we exploit predictive uncertainty to selectively incorporate pseudo labels for optimization. Benchmark experiments on a number of multi-modal datasets demonstrate that our approach significantly improves accuracy compared to state-of-the-art methods. Shizhe Hu, Yucong Wu, Jinlan Wang, Xiaoheng Jiang, Pei Lv, Mingliang Xu 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Granular Information Bottleneck for Deep Multi-Modal ClusteringabstractDeep multi-modal clustering generally focuses on improving clustering accuracy by leveraging information from different modalities. However, existing methods are designed around the finest-grained points as input, neglecting the relationships and information integration across different granularity levels, which negatively affects the clustering results. To this end, we propose a novel granular information bottleneck (GIB) for deep multi-modal clustering, which embeds a dual-tiered information bottleneck constraint mechanism that operates synergistically at both granular and sample levels, thereby learning discriminative feature representations with enhanced inter-cluster separability. Specifically, GIB adaptively represents and covers the sample points through granular balls of different granularity levels, which effectively captures the feature distribution within each cluster. Simultaneously, information compression and preservation are used to exploit the independence and complementarity of modalities while optimizing cluster assignments alignment. Finally, the objectives of GIB are formulated as a target function based on mutual information, and we propose a variational optimization method to ensure its convergence. Extensive experimental results validate the effectiveness of the proposed GIB model in accuracy and reliability. Zhengzheng Lou, Yuhan Zhan, Yingxuan Li, Shizhe Hu |
IEEE Trans. Image Process. | 6 |
| 2026 | Structure-Enhanced Self-Supervised Weighted Information Bottleneck for Multiview ClusteringabstractMultiview clustering (MVC) is a popular research topic in the fields of data mining and pattern recognition, which focuses on fully exploring and employing the correlations between views to jointly discover a consistent cluster structure across data. Typically, weighted MVC is a common clustering method aimed at learning the importance or weights of each view and applying them to explore the complementary information between views. However, current weighted MVCs primarily focus on the quality of each view while overlooking the crucial role of pseudo-label-based self-supervision in weight learning. In addition, most weighted MVCs only use a weighting mechanism to utilize complementary features without sufficiently considering the consistency relationship between the clustering results of individual views and the final clustering result. Aiming to solve the above problems, this article proposes a structure-enhanced self-supervised weighted information bottleneck (S2WIB) method for MVC. Specifically, the S2WIB method establishes a view-weight learning mechanism that leverages both the view-contained information and the self-supervised information to learn view weights, and then integrates the weighted information from different views. Meanwhile, based on the information bottleneck (IB) theory, it explores view correlations from two perspectives, namely complementary information and consistent view cluster structure information, thereby fully exploiting the potential information contained in multiview data. Experiments on various multiview text datasets, multifeature image datasets, multiangle video datasets, multimodal text-image datasets, as well as large-scale datasets and biological multiomics datasets, demonstrate that the S2WIB method performs effectively and exhibits superiority across various fields. Zhengzheng Lou, Yucong Wu, Shizhe Hu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Multi-aspect Self-guided Deep Information Bottleneck for Multi-modal ClusteringabstractDeep multi-modal clustering can extract useful information among modals, thus benefiting the final clustering and many related fields. However, existing multi-modal clustering methods have two major limitations. First, they often ignore different levels of guiding information from both the feature representations and cluster assignments, which thus are difficult in learning discriminative representations. Second, most methods fail to effectively eliminate redundant information between multi-modal data, negatively affecting clustering results. In this paper, we propose a novel multi-aspect self-guided deep information bottleneck (MSDIB) method for multi-modal clustering, which can effectively employ different aspects of guiding information for learning cluster-friendly information among modals. MSDIB mainly contains two parts: information compression and information preservation. In information compression, we extract from the private information of each modality to obtain the compact representation and meanwhile conduct mutual compression between them. In information preservation, the aim is to preserve the shared information among modals and the self-supervised information from the clustering results in each iteration. In the above process, there are mainly three aspects of self-guiding information, the modality-private information, the modality-shared information and the self-supervised pseudo label information. By minimizing the mutual information based objective function with a variational optimization method, we can fully extract useful discriminative information while eliminating the irrelevant parts. Extensive experimental results demonstrate that our method outperforms state-of-the-art multi-modal clustering methods, showcasing its superior performance and broad application prospects. Shizhe Hu, Guoliang Zou, Yangdong Ye |
AAAI | 1 |
| 2025 | Self-supervised Trusted Contrastive Multi-view Clustering with Uncertainty RefinedabstractMulti-view clustering (MVC), especially contrastive MVC, has demonstrated promising potential in many fields and practical scenarios. However, existing contrastive MVC methods still ignore the reliability of clustering results and the impact of false negative pairs, which limits the application of methods in critical security areas. To solve the above challenges, we propose a Self-supervised Trusted Contrastive Multi-view Clustering with Uncertainty Refined (STCMC-UR) method, which integrates clustering results and uncertainty learning to guide the self-supervised contrastive learning (CL). First, the belief of a specific view is generated in the evidence generation module. Afterwards, the belief mass and uncertainty of each view are learned using the Dirichlet distribution and we fuse multiple views with the Dempster-Shafer theory to generate the final clustering result and the uncertainty of the view. Then, the view weight is further quantified to adjust the belief of each view. Different from existing methods, with the clustering result and uncertainty generated by the fusion, we design a feature-level uncertainty-refined self-supervised CL module, where the pseudo-label is selectively employed in each iteration to conduct more accurate CL. As a result, the modules are mutually beneficial, which is conducive to more effective feature learning and clustering structure discovery, and more accurate learning results are obtained. Extensive experiments on five datasets show that the proposed method has significant improvements in effectiveness compared with the latest methods. Shizhe Hu, Binyan Tian, Yangdong Ye |
AAAI | 1 |
| 2025 | A Peer-review Look on Multi-modal Clustering: An Information Bottleneck Realization MethodabstractDespite the superior capability in complementary information exploration and consistent clustering structure learning, most current weight-based multi-modal clustering methods still contain three limitations: 1) lack of trustworthiness in learned weights; 2) isolated view weight learning; 3) extra weight parameters. Motivated by the peer-review mechanism in the academia, we in this paper give a new peer-review look on the multi-modal clustering problem and propose to iteratively treat one modality as "author" and the remaining modalities as "reviewers" so as to reach a peer-review score for each modality. It essentially explores the underlying relationships among modalities. To improve the trustworthiness, we further design a new trustworthy score with a self-supervision working mechanism. Following that, we propose a novel Peer-review Trustworthy Information Bottleneck (PTIB) method for weighted multi-modal clustering, where both the above scores are simultaneously taken into account for accurate and parameter-free modality weight learning. Extensive experiments on eight multi-modal datasets suggest that PTIB can outperform the state-of-the-art multi-modal clustering methods. Zhengzheng Lou, Hang Xue, Shizhe Hu |
ICML | 4 |
| 2025 | Super Deep Contrastive Information Bottleneck for Multi-modal ClusteringabstractIn an era of increasingly diverse information sources, multi-modal clustering (MMC) has become a key technology for processing multi-modal data. It can apply and integrate the feature information and potential relationships of different modalities. Although there is a wealth of research on MMC, due to the complexity of datasets, a major challenge remains in how to deeply explore the complex latent information and interdependencies between modalities. To address this issue, this paper proposes a method called super deep contrastive information bottleneck (SDCIB) for MMC, which aims to explore and utilize all types of latent information to the fullest extent. Specifically, the proposed SDCIB explicitly introduces the rich information contained in the encoder’s hidden layers into the loss function for the first time, thoroughly mining both modal features and the hidden relationships between modalities. Moreover, the proposed SDCIB performs dual optimization by simultaneously considering consistency information from both the feature distribution and clustering assignment perspectives, the proposed SDCIB significantly improves clustering accuracy and robustness. We conducted experiments on 4 multi-modal datasets and the accuracy of the method on the ESP dataset improved by 9.3%. The results demonstrate the superiority and clever design of the proposed SDCIB. The source code is available on https://github.com/ShizheHu. Zhengzheng Lou, Yucong Wu, Shizhe Hu |
ICML | 4 |
| 2025 | Diversity-oriented Deep Multi-modal ClusteringabstractDeep multi-modal clustering (DMC) aims to explore the correlated information from different modalities to improve the clustering performance. Most existing DMCs attempt to investigate the consistency or/and complementarity information by fusing all modalities, but this will lead to the following challenges: 1) Information conflicts between modalities emerge. 2) Information-rich modalities may be weakened. To address the above challenges, we propose a diversity-oriented deep multi-modal clustering (DDMC) method, where the core is dominant modality enhancement instead of multi-modal fusion. Specifically, we select the modality with the highest average silhouette coefficient as the dominant modality, then learn the diversity information between the dominant madality and the remaining ones with diversity learning, and finally enhance the dominant modality for clustering. Extensive experiments show the superiority of the proposed method over several compared DMC methods. To our knowledge, this is the first work to perform multi-modal clustering by enhancing the dominant modality instead of fusion. Yanzheng Wang, Shizhe Hu |
NeurIPS | 4 |
| 2025 | Dual global information guidance for deep contrastive multi-modal clustering
Guoliang Zou, Shizhe Hu, Tongji Chen, Yunpeng Wu, Yangdong Ye |
Inf. Sci. | 2 |
| 2025 | Parameter-Free Deep Multi-Modal Clustering With Reliable Contrastive LearningabstractDeep multi-modal clustering (DMC) expects to improve clustering performance by exploiting abundant information available from multiple modalities. However, different modalities usually have heterogeneous distribution with uneven quality. This may lead to limited performance, especially for contrastive multi-modal clustering, which inevitably performs contrastive learning between high-quality and low-quality modalities. To tackle this challenge, we propose a novel framework named parameter-free deep multi-modal clustering with reliable contrastive learning (PDMC-RCL). Specifically, the reliable contrastive learning quantifies the relationship between contrastive modality pairs with weight values that will promote the discriminative features learning from useful modality pairs and slow down or even prevent the learning from unreliable modality pairs. Moreover, the reliable contrastive learning is imposed simultaneously at both the feature-level and cluster-level in this framework so that the feature representation learning can benefit from multi-level contrastive learning. It is worth noting that our PDMC-RCL method is parameter-free, which can achieve promising performance without additional hyperparameter tuning. Experimental results on various datasets show the effectiveness of our method over typical state-of-the-art compared DMCs. The source code is available on https://github.com/ShizheHu. Zhengzheng Lou, Hang Xue, Yanzheng Wang, Shizhe Hu |
IEEE Trans. Image Process. | 6 |
| 2025 | Deep Multiview Clustering by Pseudo-Label Guided Contrastive Learning and Dual Correlation LearningabstractDeep multiview clustering (MVC) is to learn and utilize the rich relations across different views to enhance the clustering performance under a human-designed deep network. However, most existing deep MVCs meet two challenges. First, most current deep contrastive MVCs usually select the same instance across views as positive pairs and the remaining instances as negative pairs, which always leads to inaccurate contrastive learning (CL). Second, most deep MVCs only consider learning feature or cluster correlations across views, failing to explore the dual correlations. To tackle the above challenges, in this article, we propose a novel deep MVC framework by pseudo-label guided CL and dual correlation learning. Specifically, a novel pseudo-label guided CL mechanism is designed by using the pseudo-labels in each iteration to help removing false negative sample pairs, so that the CL for the feature distribution alignment can be more accurate, thus benefiting the discriminative feature learning. Different from most deep MVCs learning only one kind of correlation, we investigate both the feature and cluster correlations among views to discover the rich and comprehensive relations. Experiments on various datasets demonstrate the superiority of our method over many state-of-the-art compared deep MVCs. The source implementation code will be provided at https://github.com/ShizheHu/Deep-MVC-PGCL-DCL. Shizhe Hu, Guoliang Zou, Zhengzheng Lou, Yangdong Ye |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | A High Performance Detailed Router Based on Integer Programming with Adaptive Route GuidesabstractDetailed routing is a crucial and time-consuming stage for ASIC design. As the number and complexity of design rules increase, it is challenging to achieve high solution quality and fast speed at the same time in detailed routing. In this work, a high performance detailed routing algorithm named IPAG with integer programming (IP) is proposed. The IP formulation uses the selection of candidate routes as decision variables. High quality candidate routes are generated by queue-based rip-up and reroute with adaptive global route guidance. A design rule checking engine which can simultaneously process nets with multiple routes is designed, to efficiently construct penalty parameters in the IP formulation. Experimental results on ISPD 2018 detailed routing benchmark show that IPAG achieves better solution quality in shorter or comparable runtime, as compared to the state-of-the-art academic detailed router. Zhongdong Qi, Shizhe Hu, Qi Peng 0003, Hailong You, Zhangming Zhu |
ASPDAC | 2 |
| 2024 | Self-supervised Weighted Information Bottleneck for Multi-view Clustering
Zhengzheng Lou, Hang Xue, Yangdong Ye, Qinglei Zhou, Shizhe Hu |
IJCAI | 6 |
| 2024 | Learning Dual Enhanced Representation for Contrastive Multi-view ClusteringabstractContrastive multi-view clustering is widely recognized for its effectiveness in mining feature representation across views via contrastive learning (CL), gaining significant attention in recent years. Most existing methods mainly focus on the feature-level or/and cluster-level CL, but there are still two shortcomings. Firstly, feature-level CL is limited by the influence of anomalies and large noise data, resulting in insufficient mining of discriminative feature representation. Secondly, cluster-level CL lacks the guidance of global information and is always restricted by the local diversity information. We in this paper Learn dUal enhanCed rEpresentation for Contrastive Multi-view Clustering (LUCE-CMC) to effectively addresses the above challenges, and it mainly contains two parts, i.e., enhanced feature-level CL (En-FeaCL) and enhanced cluster-level CL (En-CluCL). Specifically, we first adopt a shared encoder to learn shared feature representations between multiple views and then obtain cluster-relevant information that is beneficial to the clustering results. Moreover, we design a reconstitution approach to force the model to concentrate on learning features that are critical to reconstructing the input data, reducing the impact of noisy data and maximizing the sufficient discriminative information of different views in helping the En-FeaCL part. Finally, instead of contrasting the view-specific clustering result like most existing methods do, we in the En-CluCL part make the information at the cluster-level more richer by contrasting the cluster assignment from each view and the cluster assignment obtained from the shared fused features. The end-to-end training methods of the proposed model are mutually reinforcing and beneficial. Extensive experiments conducted on multi-view datasets show that the proposed LUCE-CMC outperforms established baselines to a considerable extent. The source code is released at https://github.com/ShizheHu. Guoliang Zou, Yangdong Ye, Tongji Chen, Shizhe Hu |
ACM Multimedia | 4 |
| 2024 | Clustering scRNA-seq data with the cross-view collaborative information fusion strategyabstractSingle-cell RNA sequencing (scRNA-seq) technology has revolutionized biological research by enabling high-throughput, cellular-resolution gene expression profiling. A critical step in scRNA-seq data analysis is cell clustering, which supports downstream analyses. However, the high-dimensional and sparse nature of scRNA-seq data poses significant challenges to existing clustering methods. Furthermore, integrating gene expression information with potential cell structure data remains largely unexplored. Here, we present scCFIB, a novel information bottleneck (IB)-based clustering algorithm that leverages the power of IB for efficient processing of high-dimensional sparse data and incorporates a cross-view fusion strategy to achieve robust cell clustering. scCFIB constructs a multi-feature space by establishing two distinct views from the original features. We then formulate the cell clustering problem as a target loss function within the IB framework, employing a collaborative information fusion strategy. To further optimize scCFIB's performance, we introduce a novel sequential optimization approach through an iterative process. Benchmarking against established methods on diverse scRNA-seq datasets demonstrates that scCFIB achieves superior performance in scRNA-seq data clustering tasks. Availability: the source code is publicly available on GitHub: https://github.com/weixiaojiao/scCFIB. Zhengzheng Lou, Xiaojiao Wei, Yuanhao Hu, Shizhe Hu, Yucong Wu, Zhen Tian 0004 |
Briefings Bioinform. | 4 |
| 2024 | Nice to meet images with Big Clusters and Features: A cluster-weighted multi-modal co-clustering method
Hang Xue, Xihui Wu, Zhengzheng Lou, Shouyi Yang, Qinglei Zhou, Shizhe Hu |
Inf. Process. Manag. | 8 |
| 2024 | A Survey on Information BottleneckabstractThis survey is for the remembrance of one of the creators of the information bottleneck theory, Prof. Naftali Tishby, passing away at the age of 68 on August, 2021. Information bottleneck (IB), a novel information theoretic approach for pattern analysis and representation learning, has gained widespread popularity since its birth in 1999. It provides an elegant balance between data compression and information preservation, and improves its prediction or representation ability accordingly. This survey summarizes both the theoretical progress and practical applications on IB over the past 20-plus years, where its basic theory, optimization, extensive models and task-oriented algorithms are systematically explored. Existing IB methods are roughly divided into two parts: traditional and deep IB, where the former contains the IBs optimized by traditional machine learning analysis techniques without involving any neural networks, and the latter includes the IBs involving the interpretation, optimization and improvement of deep neural works (DNNs). Specifically, based on the technique taxonomy, traditional IBs are further classified into three categories: Basic, Informative and Propagating IB; While the deep IBs, based on the taxonomy of problem settings, contain Debate: Understanding DNNs with IB, Optimizing DNNs Using IB, and DNN-based IB methods. Furthermore, some potential issues deserving future research are discussed. This survey attempts to draw a more complete picture of IB, from which the subsequent studies can benefit. Shizhe Hu, Zhengzheng Lou, Yangdong Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Contrastive cross-modal clustering with twin network
Yiqiao Mao, Shizhe Hu, Yangdong Ye |
Pattern Recognit. | 3 |
| 2024 | Multiview Clustering With Propagating Information BottleneckabstractIn many practical applications, massive data are observed from multiple sources, each of which contains multiple cohesive views, called hierarchical multiview (HMV) data, such as image-text objects with different types of visual and textual features. Naturally, the inclusion of source and view relationships offers a comprehensive view of the input HMV data and achieves an informative and correct clustering result. However, most existing multiview clustering (MVC) methods can only process single-source data with multiple views or multisource data with single type of feature, failing to consider all the views across multiple sources. Observing the rich closely related multivariate (i.e., source and view) information and the potential dynamic information flow interacting among them, in this article, a general hierarchical information propagation model is first built to address the above challenging problem. It describes the process from optimal feature subspace learning (OFSL) of each source to final clustering structure learning (CSL). Then, a novel self-guided method named propagating information bottleneck (PIB) is proposed to realize the model. It works in a circulating propagation fashion, so that the resulting clustering structure obtained from the last iteration can "self-guide" the OFSL of each source, and the learned subspaces are in turn used to conduct the subsequent CSL. We theoretically analyze the relationship between the cluster structures learned in the CSL phase and the preservation of relevant information propagated from the OFSL phase. Finally, a two-step alternating optimization method is carefully designed for optimization. Experimental results on various datasets show the superiority of the proposed PIB method over several state-of-the-art methods. Shizhe Hu, Zenglin Shi, Zhengzheng Lou, Yangdong Ye |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Self-supervised temporal autoencoder for egocentric action segmentation
Shizhe Hu, Zhongchuan Sun, Yangdong Ye |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Mutual Boost Network for attributed graph clustering
Xiangyu Yu, Shizhe Hu, Yangdong Ye |
Expert Syst. Appl. | 3 |
| 2023 | Joint contrastive triple-learning for deep multi-view clustering
Shizhe Hu, Guoliang Zou, Zhengzheng Lou, Ruilin Geng, Yangdong Ye |
Inf. Process. Manag. | 1 |
| 2023 | Deep purified feature mining model for joint named entity recognition and relation extraction
Zhongchuan Sun, Shizhe Hu, Yangdong Ye |
Inf. Process. Manag. | 5 |
| 2023 | A simple multiple-fold correlation-based multi-view multi-label learning
Changming Zhu, Shizhe Hu, Yilin Dong 0001, Lei Cao 0002, Yuhu Shi, Lai Wei 0001, Rigui Zhou |
Neural Comput. Appl. | 3 |
| 2023 | Multi-View Clustering via Triplex Information MaximizationabstractIn this paper, we address the problem of multi-view clustering (MVC), integrating the close relationships among views to learn a consistent clustering result, via triplex information maximization (TIM). TIM works by proposing three essential principles, each of which is realized by a formulation of maximization of mutual information. 1) Principle 1: Contained. The first and foremost thing for MVC is to fully employ the self-contained information in each view. 2) Principle 2: Complementary. The feature-level complementary information across pairwise views should be first quantified and then integrated for improving clustering. 3) Principle 3: Compatible. The rich cluster-level shared compatible information among individual clustering of each view is significant for ensuring a better final consistent result. Following these principles, TIM can enjoy the best of view-specific, cross-view feature-level, and cross-view cluster-level information within/among views. For principle 2, we design an automatic view correlation learning (AVCL) mechanism to quantify how much complementary information across views by learning the cross-view weights between pairwise views automatically, instead of view-specific weights as most existing MVCs do. Specifically, we propose two different strategies for AVCL, i.e., feature-based and cluster-based strategy, for effective cross-view weight learning, thus leading to two versions of our method, TIM-F and TIM-C, respectively. We further present a two-stage method for optimization of the proposed methods, followed by the theoretical convergence and complexity analysis. Extensive experimental results suggest the effectiveness and superiority of our methods over many state-of-the-art methods. Zhengzheng Lou, Qinglei Zhou, Shizhe Hu |
IEEE Trans. Image Process. | 4 |
| 2023 | Attentive Adversarial Collaborative FilteringabstractGenerative adversarial nets (GANs) have enjoyed considerable success in computer vision and attracted much attention from recommender systems. However, due to the discrete nature of items, it is infeasible to graft GANs directly onto recommendation models. Although several methods have taken steps forward, their training processes are slow-convergent, time-consuming, or even unstable. This article proposes a novel framework named attentive adversarial collaborative filtering (AACF) and an efficient training strategy to improve GANs in recommender systems. There are two distinct novelties over previous work. First, AACF is a differentiable generative adversarial framework that introduces an attention mechanism and “virtual items” to bridge the gap between the generator and the discriminator. Owing to the intrinsic differentiability, AACF can be stably optimized with gradient descent methods. Second, the efficient training strategy substantially reduces computational complexity. It is capable of efficiently training and scaling up the AACF model to large datasets. Extensive experiments on various datasets demonstrate the effectiveness, fast convergence, stability, and scalability of AACF. Since our ideas are general in nature, they will open a path to stably and efficiently train GANs in the research areas with discrete data. The implementation code is available athttps://github.com/zhongchuansun/AACF. Zhongchuan Sun, Bin Wu 0019, Shizhe Hu, Yangdong Ye |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | A Parameter-free Multi-view Information Bottleneck Clustering Method by Cross-view WeightingabstractWith the fast-growing multi-modal/media data in the Big Data era, multi-view clustering (MVC) has attracted lots of attentions lately. Most MVCs focus on integrating and utilizing the complementary information among views by linear sum of the learned view weights and have shown great success in some fields. However, they fail to quantify how complementary the information across views actually utilized for benefiting final clustering. Additionally, most of them contain at least one parameter for regularization without prior knowledge, which puts pressure on the parameter-tuning and thus makes them impractical. In this paper, we propose a novel parameter-free multi-view information bottleneck (PMIB) clustering method to automatically identify and exploit useful complementary information among views, thus reducing the negative impact from the harmful views. Specifically, we first discover the informative view by measuring the relevant information preserved by the original data and the compact clusters with mutual information. Then, a new cross-view weight learning scheme is designed to learn how complementary between the informative view and remaining views. Finally, the quantitative correlations among views are fully exploited to improve the clustering performance without needing any additional parameters or prior knowledge. Experimental results on different kinds of multi-view datasets show the effectiveness of the proposed method. Shizhe Hu, Ruilin Geng, Zhaoxu Cheng, Guoliang Zou, Zhengzheng Lou, Yangdong Ye |
ACM Multimedia | 1 |
| 2022 | DMIB: Dual-Correlated Multivariate Information Bottleneck for Multiview ClusteringabstractMultiview clustering (MVC) has recently been the focus of much attention due to its ability to partition data from multiple views via view correlations. However, most MVC methods only learn either interfeature correlations or intercluster correlations, which may lead to unsatisfactory clustering performance. To address this issue, we propose a novel dual-correlated multivariate information bottleneck (DMIB) method for MVC. DMIB is able to explore both interfeature correlations (the relationship among multiple distinct feature representations from different views) and intercluster correlations (the close agreement among clustering results obtained from individual views). For the former, we integrate both view-shared feature correlations discovered by learning a shared discriminative feature subspace and view-specific feature information to fully explore the interfeature correlation. This allows us to attain multiple reliable local clustering results of different views. Following this, we explore the intercluster correlations by learning the shared mutual information over different local clusterings for an improved global partition. By integrating both correlations, we formulate the problem as a unified information maximization function and further design a two-step method for optimization. Moreover, we theoretically prove the convergence of the proposed algorithm, and discuss the relationships between our method and several existing clustering paradigms. The experimental results on multiple datasets demonstrate the superiority of DMIB compared to several state-of-the-art clustering methods. Shizhe Hu, Zenglin Shi, Yangdong Ye |
IEEE Trans. Cybern. | 1 |
| 2022 | View-Wise Versus Cluster-Wise Weight: Which Is Better for Multi-View Clustering?abstractWeighted multi-view clustering (MVC) aims to combine the complementary information of multi-view data (such as image data with different types of features) in a weighted manner to obtain a consistent clustering result. However, when the cluster-wise weights across views are vastly different, most existing weighted MVC methods may fail to fully utilize the complementary information, because they are based on view-wise weight learning and can not learn the fine-grained cluster-wise weights. Additionally, extra parameters are needed for most of them to control the weight distribution sparsity or smoothness, which are hard to tune without prior knowledge. To address these issues, in this paper we propose a novel and effective Cluster-weighted mUlti-view infoRmation bottlEneck (CURE) clustering algorithm, which can automatically learn the cluster-wise weights to discover the discriminative clusters across multiple views and thus can enhance the clustering performance by properly exploiting the cluster-level complementary information. To learn the cluster-wise weights, we design a new weight learning scheme by exploring the relation between the mutual information of the joint distribution of a specific cluster (containing a group of data samples) and the weight of this cluster. Finally, a novel draw-and-merge method is presented to solve the optimization problem. Experimental results on various multi-view datasets show the superiority and effectiveness of our cluster-wise weighted CURE over several state-of-the-art methods. Shizhe Hu, Zhengzheng Lou, Yangdong Ye |
IEEE Trans. Image Process. | 1 |
| 2021 | Multi-view content-context information bottleneck for image clustering
Shizhe Hu, Zhengzheng Lou, Yangdong Ye |
Expert Syst. Appl. | 1 |
| 2021 | Deep multi-view learning methods: A review
Shizhe Hu, Yiqiao Mao, Yangdong Ye, Hui Yu 0001 |
Neurocomputing | 2 |
| 2021 | Learning a deep network with cross-hierarchy aggregation for crowd counting
Qiang Guo 0012, Shizhe Hu, Sonephet Phoummixay, Yangdong Ye |
Knowl. Based Syst. | 3 |
| 2021 | Multi-Task Image Clustering through Correlation PropagationabstractTraditional image clustering algorithms deal with single-task clustering (STC) problem on a single domain. However, with the increasing number of related images on the Web, it is challenging for STCs to perform related image clustering tasks independently without considering the between-task relationship, which mainly consists of similar visual features and image patterns among tasks. Therefore, it is intuitive to resort to multi-task clustering (MTC) algorithms. However, most existing MTCs learn a shared feature subspace, which may lead to negative transfer when facing the image clustering tasks that are not strongly related. In this paper, we propose a novel multi-task image clustering algorithm, which performs multiple image clustering tasks simultaneously and propagates the task correlation to improve clustering performance. Specifically, we first extend the information bottleneck method to cluster tasks independently. The related and unrelated images between the pairwise clusters of different tasks are then discovered. Meanwhile, two corresponding types of correlations are propagated among the tasks, where only the positive correlation benefits the clustering of each task. A sequential and collaborative method is further designed to ensure an optimal solution. Moreover, we perform a theoretical analysis of the properties on correlation propagation and the convergence of our algorithm. The experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art clustering methods. Shizhe Hu, Yangdong Ye |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Content Vs Context: How About "Walking Hand-In-Hand" For Image Clustering?abstractImage clustering has been one of the most important issues in the field of pattern recognition. However, most of existing methods only focus on utilizing either content or context information of images, failing to consider both of them. In fact, the powerful algorithms can be realized by a combination of the rich content and context information. This paper proposes a novel content-context information bottleneck (C2IB) algorithm, which simultaneously explores and exploits the content and context information for discovering image clusters. The "content" describes the intrinsic characteristics contained in each image such as the appearance feature, and the "context" depicts the close correlations between images such as inter-image distance or similarity. Then, we formulate the problem as an information loss function by maximally preserving the content and context information while compressing the images. Finally, we design a new sequential method for the optimization. Experimental results show the superiority of the proposed method. Shizhe Hu, Zhenquan Hou, Zhengzheng Lou, Yangdong Ye |
ICASSP | 1 |
| 2020 | Heterogeneous Dual-Task Clustering with Visual-Textual InformationabstractExisting visual-textual cross-modal clustering techniques focus on finding a clustering partition of different modalities by dealing with each modality dependently or integrating multiple modalities into a shared space, which may results in unsatisfactory performance due to the heterogeneous gap of different modalities. Aiming at this problem, we propose a novel heterogeneous dual-task clustering (HDC) method, which is capable of exploring high-level relatedness between visual and textual data to improve the performance of individual task. Our intuition is that although the visual and textual data are heterogenous to each other, they may share related high-level semantics and rich latent correlations, which can lead to improved performance if we treat the clustering of visual and textual data as different but related learning tasks. Specifically, the problem of heterogeneous dual-task clustering is formulated as an information-theoretic function, in which the low-level information in each modality and high-level relatedness between multiple modalities are maximally preserved. Then, a progressive optimization method is proposed to ensure a local optimal solution. Extensive experiments show noticeable performance of the HDC approach in comparison with several state-of-the-art baselines. Yiqiao Mao, Shizhe Hu, Yangdong Ye |
SDM | 3 |
| 2020 | DSPNet: Deep scale purifier network for dense crowd counting
Yunpeng Wu, Shizhe Hu, Ruobin Wang, Yangdong Ye |
Expert Syst. Appl. | 3 |
| 2020 | Joint specific and correlated information exploration for multi-view action clustering
Shizhe Hu, Yangdong Ye |
Inf. Sci. | 1 |
| 2020 | Dynamic auto-weighted multi-view co-clustering
Shizhe Hu, Yangdong Ye |
Pattern Recognit. | 1 |
| 2020 | Multi-task Information Bottleneck Co-clustering for Unsupervised Cross-view Human Action CategorizationabstractThe widespread adoption of low-cost cameras generates massive amounts of videos recorded from different viewpoints every day. To cope with this vast amount of unlabeled and heterogeneous data, a new multi-task information bottleneck co-clustering (MIBC) approach is proposed to automatically categorize human actions in collections of unlabeled cross-view videos. Our motivation is that, if a learning action category from each view is seen as a single task, it is reasonable to assume that the tasks of learning action patterns from the videos recorded by multiple cameras are dependent and inter-related, since the actions of the same subjects synchronously recorded from different camera viewpoints are complementary to each other. MIBC aims to transfer the shared view knowledge across multiple tasks (i.e., camera viewpoints) to boost the performance of each task. Specifically, MIBC involves the following two parts: (1) extracting action categories for each task by independently maintaining its own relevant information, and (2) allowing the feature representations of all tasks to be compressed into a common feature space, which is utilized to capture the relatedness of multiple tasks and transfer the shared knowledge across different camera viewpoints. These two parts of MIBC work simultaneously and can be solved in a novel co-clustering mechanism. Our experimental evaluation on several cross-view action collections shows that the MIBC algorithm outperforms the existing state-of-the-art baselines. Zhengzheng Lou, Shizhe Hu, Yangdong Ye |
ACM Trans. Knowl. Discov. Data | 3 |
| 2017 | Multi-task Clustering of Human Actions by Sharing Information
Shizhe Hu, Yangdong Ye |
CVPR | 2 |