VLDB 2026 Research / reviewers in the wild / expert
Man-Sheng Chen
dblp:238/2498
· DBLP profile ↗
32ranked-venue papers
16as first author
30since 2021 · last 2026
0000-0001-6578-0616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 16 since 2021Databases, data management, data science and information retrieval · 11 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | P3L: Patent Prediction With Prompt LearningabstractPatents are crucial for protecting technological innovations and fostering competitive advancements in industry. Patent prediction, a novel task in the field of patent mining, aims to forecast future technological trends, providing valuable insights for strategic planning and innovation in the industry. However, the complexity of patent data and the diversity of technological fields make effective patent prediction a significant challenge. Existing methods for predicting scientific research trends struggle to effectively model patent structures and capture dependencies between patents, resulting in suboptimal patent trend predictions. In this article, we propose a novel method, patent prediction with prompt learning (P3L), to achieve effective and accurate prediction of future patent developments based on a pretrained language model (PLM). P3L includes a patent similarity path extraction module to extract multiple patent development paths from extensive datasets. Following this, we design a patent prompt learning approach that integrates patent development paths, keywords, and patent similarities into the prompts. To mitigate potential noise introduced by this integration, we introduce an attention mask matrix for prompt denoising. Finally, we introduce three patent datasets with rich structures, and conduct extensive experiments on these datasets as well as a public dataset, demonstrating the superiority of the proposed method. The dataset and code have been made publicly available athttps://github.com/AllminerLab/P3L Yi-Hong Lu, Pei-Yuan Lai, Man-Sheng Chen, Huan-Tao Cai, Zeng-Hui Wang, Shuang-Yin Liu, Chang-Dong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2026 | Knowledge-Aware ClusteringabstractData clustering aims to partition the input data entities into several disjoint categories, where similar entities are grouped together while dissimilar ones are pulled apart. In general, the existing data clustering methods merely depend on the attribute information or constructed similarity graph information of the input data entities, and such one-sided cues may result in an incomplete understanding and potentially biased conclusions. How to well consider the inherent heterogeneous knowledge of data while maintaining the local semantics remains a challenging problem. To address this, we propose a novel knowledge-aware clustering (KAC) method, where the attribute information and inherent heterogeneous knowledge of data are jointly considered for better cluster structure recovery. Within this framework, there are four major modules. They are: 1) the target attribute autoencoder to capture a customizable target feature representation; 2) the meta-path aggregated knowledge graph encoder to learn an informative graph feature representation; 3) the dual information maximization to ensure the consistent semantics between two embedding representations; and 4) the self-training clustering to self-supervise the clustering learning. Experiments on several real-world datasets are conducted to validate the superiority of KAC, indicating the significance of explicitly considering the heterogeneous knowledge while maintaining local semantics. Man-Sheng Chen, Li-An Ren, Wudong Xi, Chang-Dong Wang 0001, Philip S. Yu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2025 | Motif-aware curriculum learning for node classification
Xiaosha Cai, Man-Sheng Chen, Chang-Dong Wang 0001, Haizhang Zhang |
Neural Networks | 2 |
| 2025 | Graph Prompt ClusteringabstractDue to the wide existence of unlabeled graph-structured data (e.g., molecular structures), the graph-level clustering has recently attracted increasing attention, whose goal is to divide the input graphs into several disjoint groups. However, the existing methods habitually focus on learning the graphs embeddings with different graph reguralizations, and seldom refer to the obvious differences in data distributions of distinct graph-level datasets. How to characteristically consider multiple graph-level datasets in a general well-designed model without prior knowledge is still challenging. In view of this, we propose a novel Graph Prompt Clustering (GPC) method. Within this model, there are two main modules, i.e., graph model pretraining as well as prompt and finetuning. In the graph model pretraining module, the graph model is pretrained by a selected source graph-level dataset with mutual information maximization and self-supervised clustering regularization. In the prompt and finetuning module, the network parameters of the pretrained graph model are frozen, and a groups of learnable prompt vectors assigned to each graph-level representation are trained for adapting different target graph-level datasets with various data distributions. Experimental results across six benchmark datasets demonstrate the impressive generalization capability and effectiveness of GPC compared with the state-of-the-art methods. Man-Sheng Chen, Pei-Yuan Lai, De-Zhang Liao, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Knowledge-Aware Synergistic Discovery of Drug Combinations: A Large Language Model PerspectiveabstractDrug combination therapy with significant advantages is a well-established concept in cancer treatment. Some related efforts have been made with multiple artful deep learning techniques. However, they are usually based on data for drug synergy prediction, ignoring the professional characteristics of data and the systematic knowledge accumulation. Meanwhile, integrating the dispersed professional knowledge and effectively utilizing it in data remains a crucial technical challenge. In this study, we propose KSDDC, a novel model for knowledge-aware synergistic discovery of drug combinations from a large language model (LLM) perspective (i.e., from the continuously learnable and refined large database). Within this framework, three main modules are well-designed, i.e., knowledge-aware drug feature auto-encoding, knowledge-aware cell line feature encoding and drug-drug synergy prediction. Informative embeddings of samples are discovered and combined to make accurate drug synergy prediction. Overall, KSDDC is superior compared with the other shallow machine learning based methods and deep learning based methods on several synergistic prediction benchmarks, where about 19% F1-score improvements over the second best method on DrugComb_1 can be observed. Starting with drug synergy prediction, our studies with knowledge-enabled data mining offer valuable insights and serve as a reference method for future research in this field. Pei-Yuan Lai, Man-Sheng Chen, De-Zhang Liao, Chang-Dong Wang 0001, Min Chen 0003 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Homophily Induced Contrastive Attributed Graph ClusteringabstractAttributed graph clustering, aiming to discover the underlying graph structure and partition the graph nodes into several disjoint categories, is a basic task in graph data analysis. Although recent efforts over graph contrastive clustering have achieved decent performances, most of them get accustomed to construct the positive neighbor set by the generated pseudo clustering information, directly ignoring the ready-made neighbor nodes and the underlying semantics of edges in a graph. How to well deal with the graph neighbor-specific information to facilitate the performance of graph contrastive clustering is still a challenging problem. Therefore, in this work, we propose a novel Homophily Induced Contrastive Attributed Graph Clustering (HomoCAGC) method, where the power of homophily is exploited in facilitating the performance of contrastive attributed graph clustering, while the pseudo homophily in a graph is also explored and distinguished. Especially, the node feature as reliable information guidance is used to compute the underlying feature-oriented pair-wise node similarity, based on which the positive node pairs in contrastive regularizer are adjusted for better node representation characterization. According to the refined node representations, a triplet self-supervised clustering objective is well-designed to ensure the output embedding is cluster-oriented, and suitable for the downstream clustering task. Extensive experiments on seven benchmarks are conducted to demonstrate the effectiveness of HomoCAGC. Man-Sheng Chen, Pei-Yuan Lai, De-Zhang Liao, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | "Pre-Train, Prompt" Framework to Boost Graph Neural Networks Performance in EEG AnalysisabstractElectroencephalography (EEG) is a vital non-invasive technique used in neuroscience research and clinical diagnosis. However, EEG data have a complex non-Euclidean structure and are often scarce, making training effective graph neural network (GNN) models difficult. We propose a "pre-train, prompt" framework in graph neural networks for EEG analysis, called GNN-based EEG Prompt Learning (GEPL). The framework first uses unsupervised contrastive learning to pre-train on a large-scale EEG dataset. It then transfers the generic EEG knowledge learned by the model to target EEG datasets through graph prompt learning, thereby enhancing the model's performance with a limited amount of EEG data from the target domain. We tested the framework on five EEG datasets, and the results showed that GEPL outperformed traditional fine-tuning methods in classification accuracy and area under the ROC curve (AUC). GEPL demonstrated improved generalization, robustness, and computational efficiency, thereby significantly reducing the overfitting risks associated with limited EEG data. Moreover, the model provided interpretable results, highlighting relevant brain regions during classification tasks. This research suggests that the "pre-train, prompt" paradigm is well-suited for EEG analysis and offers potential applications in other domains where data are limited. Canming Cui, Man-Sheng Chen, Zhaopeng Tong, Chanmei Fang, Chang-Dong Wang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Interpretable Staging Prediction of Liver Cancer Based on Joint-Knowledge NetworkabstractClinical staging is crucial for treatment strategies and improving 5-year survival rates in hepatocellular carcinoma (HCC) patients. However, existing methods struggle to distinguish stages with highly similar textual features. Additionally, their lack of interpretability hampers their practical application in medical scenarios. Here, we introduce KnowST, a joint-knowledge network designed to leverage task relevance to explore implicit knowledge for interpretable staging prediction of liver cancer. First, the relevance of auxiliary tasks and the main task is established from two perspectives to guide the model's focus on staging-related implicit knowledge in radiology reports. Stages-to-stages: KnowST learns the inter-stage distinctions between different stages and the similarities within the same stages, using these as important references for staging differentiation. Factors-to-stages: Clinically, staging is determined by multiple tumor factors. These factors can serve as effective clues to assist KnowST in predicting the correct stage, especially in the case of confusing stages. Second, domain-specific word embeddings are introduced to bridge the gap between pre-trained language models and Chinese radiology reports. Lastly, tumor factor prediction enhances the credibility of the deep model in staging prediction, and its visualized results effectively demonstrate the model's interpretability. Overall, KnowST leverages the joint-knowledge from these two perspectives, effectively utilizing implicit information in radiology reports to achieve interpretable clinical staging. Compared to the optimal baselines, KnowST improves AUC by 7.69% and achieves 90.52% accuracy on 573 real-world radiology reports, while also demonstrating superior stage identification and stable performance across various metrics. Xuecong Zheng, Ya Li 0008, Zhiqi Wu, Yiyang Tang, Pei-Yuan Lai, Man-Sheng Chen, Chang-Dong Wang 0001, Jiaping Li |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Knowledge Graph-Based Patent ClusteringabstractPatent data generally includes information from different perspectives or different types, and its heterogeneous attributes can be greatly beneficial to data clustering analysis. However, the existing patent analysis method always focus on the patent text cues, and such a strategy merely depends on the feature information to capture the data characteristics, failing to multi-type informative patent representation. Therefore, in this paper, to model the underlying structure/relationships of patent data, we employ the knowledge graph to depict the heterogeneous attributes of patent, and propose a novel Knowledge Graph-based Patent Clustering (KGPC) method, where the relationship reconstruction in knowledge graph as well as clustering-oriented representation refinement for patent clustering are jointly considered. With this model, there are three components, i.e., entity representation refinement, relationship reconstruction and self-supervised entity clustering. Given a patent knowledge graph as input, the entity representation refinement can be mutually boosted by the relationship reconstruction and self-supervised clustering objective, thereby leading to a balanced clustering-oriented output. Extensive experiments on several real-world patent knowledge graph datasets validate the effectiveness of KGPC while compared with the state-of-the-art. Pei-Yuan Lai, Man-Sheng Chen, Chang-Dong Wang 0001, Min Chen 0003, Mohsen Guizani |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Smoothness-Induced Efficient Incomplete Multi-View Clustering
Tianchuan Yang, Haiqiang Chen, Man-Sheng Chen, Xiangcheng Li 0001, Youming Sun, Chang-Dong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Contrastive Ensemble ClusteringabstractEnsemble clustering aims to combine different base clusterings into a better clustering than that of the individual one. In general, a co-association matrix depicting the pairwise affinity between different data samples is constructed by average fusion or weighted fusion of the connective matrices from multiple base clusterings. Despite the significant success, the existing works fail to capture the global structure information from multiple noisy connective matrices. Meanwhile, the locality property of the resulting representation matrix could not be explicitly preserved. In this article, we propose a novel contrastive ensemble clustering (CEC) method. Specifically, a consensus mapping model is designed for the discovery of the latent representation from the noisy observations with distinct confidences. Furthermore, a contrastive regularizer is dexterously formulated to refine the latent representation while preserving its locality property. Extensive experiments conducted on several benchmark datasets demonstrate the superiority of the proposed CEC method. To the best of our knowledge, it is the first time to explore the potential of latent representation learning and contrastive components for the ensemble clustering task. Man-Sheng Chen, Jia-Qi Lin 0001, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Contrastive Multiview Attribute Graph Clustering With Adaptive EncodersabstractMultiview attribute graph clustering aims to cluster nodes into disjoint categories by taking advantage of the multiview topological structures and the node attribute values. However, the existing works fail to explicitly discover the inherent relationships in multiview topological graph matrices while considering different properties between the graphs. Besides, they cannot well handle the sparse structure of some graphs in the learning procedure of graph embeddings. Therefore, in this article, we propose a novel contrastive multiview attribute graph clustering (CMAGC) with adaptive encoders method. Within this framework, the adaptive encoders concerning different properties of distinct topological graphs are chosen to integrate multiview attribute graph information by checking whether there exists high-order neighbor information or not. Meanwhile, the number of layers of the GCN encoders is selected according to the prior knowledge related to the characteristics of different topological graphs. In particular, the feature-level and cluster-level contrastive learning are conducted on the multiview soft assignment representations, where the union of the first-order neighbors from the corresponding graph pairs is regarded as the positive pairs for data augmentation and the sparse neighbor information problem in some graphs can be well dealt with. To the best of our knowledge, it is the first time to explicitly deal with the inherent relationships from the interview and intraview perspectives. Extensive experiments are conducted on several datasets to verify the superiority of the proposed CMAGC method compared with the state-of-the-art methods. Man-Sheng Chen, Xi-Ran Zhu, Jia-Qi Lin 0001, Chang-Dong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Dual Information Enhanced Multiview Attributed Graph ClusteringabstractMultiview attributed graph clustering is an important approach to partition multiview data based on the attribute characteristics and adjacent matrices from different views. Some attempts have been made in using graph neural network (GNN), which have achieved promising clustering performance. Despite this, few of them pay attention to the inherent specific information embedded in multiple views. Meanwhile, they are incapable of recovering the latent high-level representation from the low-level ones, greatly limiting the downstream clustering performance. To fill these gaps, a novel dual information enhanced multiview attributed graph clustering (DIAGC) method is proposed in this article. Specifically, the proposed method introduces the specific information reconstruction (SIR) module to disentangle the explorations of the consensus and specific information from multiple views, which enables graph convolutional network (GCN) to capture the more essential low-level representations. Besides, the contrastive learning (CL) module maximizes the agreement between the latent high-level representation and low-level ones and enables the high-level representation to satisfy the desired clustering structure with the help of the self-supervised clustering (SC) module. Extensive experiments on several real-world benchmarks demonstrate the effectiveness of the proposed DIAGC method compared with the state-of-the-art baselines. Jia-Qi Lin 0001, Man-Sheng Chen, Xi-Ran Zhu, Chang-Dong Wang 0001, Haizhang Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | HomoMGC: Homophily-Enhanced Adaptive Graph Refinement for Multi-View Graph ClusteringabstractDue to the emergency of multi-view graph data, considerable attention is focused on the multi-view graph clustering. Although great efforts have been made in developing the multi-view graph clustering methods, most of them implicitly follow the homophily assumption, where the connected nodes with edges tend to be in the same category. As a matter of fact, such an ideal assumption is hard to be satisfied in the real-world graph data, and there are some heterogeneous edges connecting dissimilar nodes in graph. How to well consider the homophily and refine the noisy/heterogeneous edges in multi-view graph clustering still remains an under-explored challenge. Therefore, in this paper, we propose a Homophily-enhanced Adaptive Graph Refinement for Multi-view Graph Clustering (HomoMGC) method, where an adaptive graph refinement strategy is seamlessly designed. Specifically, a feature-oriented graph is constructed based on the shared feature, and an integrated graph is computed by averagely fusing all the input adjacent graphs. Then, the feature-oriented graph and integrated graph are stacked into a graph tensor with a low-rank tensor constraint, where a refined affinity probability matrix can be adaptively recovered from the integrated graph by considering multiple graph information as well as the semantics features. Extensive experiments on several benchmark datasets demonstrate the superiority of HomoMGC compared with the state-of-the-art graph clustering methods. For the code reproducibility, the source code of HomoMGC is public available at https://github.com/ManshengChen/Code-for-HomoMGc-master. Man-Sheng Chen, Xiaosha Cai, Chang-Dong Wang 0001, Dong Huang 0001, Min Chen 0003, Mohsen Guizani |
ICDM | 1 |
| 2024 | Periodic Prompt on Dynamic Heterogeneous Graph for Next Basket RecommendationabstractIn next basket recommendation, baskets are usually formed through a large number of user interactions with items in the early stage. In general, the existing methods for next basket recommendation primarily focus on historical purchase behavior of users, assuming that user purchase interests are static, and overlook the dynamic and diverse changes in user purchase interests. In order to fully capture dynamic user interests and provide users with more diverse recommendations, we propose our method, Dynamic Heterogeneous Graph Prompt (DHGP), for next basket recommendation. By constructing a dynamic heterogeneous graph, we can adequately consider the influence of various interactive behaviors on the user's baskets at different times. Furthermore, we introduce a periodic dynamic heterogeneous prompt strategy to capture the interest directions between baskets from different users and provide users with more diverse interest directions. Extensive experimental validation on six real world datasets demonstrates that our method shows strong applicability across datasets under various conditions and outperforms several state-of-the-art recommendation methods. To the best of our knowledge, DHGP is the first next basket recommendation method that effectively combines dynamic and heterogeneous information. The implementation code is accessible at https://github.com/AllminerLab. Ru-Bin Li, Man-Sheng Chen, Xin-Yu Ding, Chang-Dong Wang 0001, Sihong Xie, Shuangyin Liu, Min Chen 0003, Mohsen Guizani |
ICDM | 2 |
| 2024 | MIGP: Metapath Integrated Graph Prompt Neural Network
Pei-Yuan Lai, Yi-Hong Lu, Zeng-Hui Wang, Man-Sheng Chen, Chang-Dong Wang 0001 |
Neural Networks | 5 |
| 2024 | A Tensor Approach for Uncoupled Multiview ClusteringabstractMultiview clustering plays an important part in unsupervised learning. Although the existing methods have shown promising clustering performances, most of them assume that the data is completely coupled between different views, which is unfortunately not always ensured in real-world applications. The clustering performance of these methods drops dramatically when handling the uncoupled data. The main reason is that: 1) cross-view correlation of uncoupled data is unclear, which limits the existing multiview clustering methods to explore the complementary information between views and 2) features from different views are uncoupled with each other, which may mislead the multiview clustering methods to partition data into wrong clusters. To address these limitations, we propose a tensor approach for uncoupled multiview clustering (T-UMC) in this article. Instead of pairwise correlation, T-UMC chooses a most reliable view by view-specific silhouette coefficient (VSSC) at first, and then couples the self-representation matrix of each view with it by pairwise cross-view coupling learning. After that, by integrating recoupled self-representation matrices into a third-order tensor, the high-order correlations of all views are explored with tensor singular value decomposition (t-SVD)-based tensor nuclear norm (TNN). And the view-specific local structures of each individual view are also preserved with the local structure learning scheme with manifold learning. Besides, the physical meaning of view-specific coupling matrix is also discussed in this article. Extensive experiments on six commonly used benchmark datasets have demonstrated the superiority of the proposed method compared with the state-of-the-art multiview clustering methods. Jia-Qi Lin 0001, Man-Sheng Chen, Chang-Dong Wang 0001, Haizhang Zhang |
IEEE Trans. Cybern. | 2 |
| 2024 | Concept Factorization Based Multiview Clustering for Large-Scale DataabstractMost existing large-scale multiview clustering algorithms attempt to capture data distribution in multiple views by selecting view-wise anchor representations beforehand with$k$-means, or by direct matrix factorization on the original observations. Despite impressive performance, few of them have paid attention to the semantic correlations between anchor bases and cluster centroids, or even the underlying relations between clusters and data samples. In view of this, we propose aConceptFactorization basedMultiviewClustering for Large-scale Data (CFMC) method with nearly linear complexity. The anchor bases learning, coefficient expression with clear semantic cues and partitioning are integrated together in this unified model. Meanwhile, explicit connections among multiview data, anchor bases and clusters are modeled via coefficient representations with semantic meanings. A four-step alternate minimizing algorithm is designed to handle the optimization problem, which is proved to have linear time complexityw.r.t.the sample size. Extensive experiments conducted on several challenging large-scale datasets confirm the superiority of the method compared with the state-of-the-art methods. Man-Sheng Chen, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Refining Graph Structure for Incomplete Multi-View ClusteringabstractAs a challenging problem, incomplete multi-view clustering (MVC) has drawn much attention in recent years. Most of the existing methods contain the feature recovering step inevitably to obtain the clustering result of incomplete multi-view datasets. The extra target of recovering the missing feature in the original data space or common subspace is difficult for unsupervised clustering tasks and could accumulate mistakes during the optimization. Moreover, the biased error is not taken into consideration in the previous graph-based methods. The biased error represents the unexpected change of incomplete graph structure, such as the increase in the intra-class relation density and the missing local graph structure of boundary instances. It would mislead those graph-based methods and degrade their final performance. In order to overcome these drawbacks, we propose a new graph-based method named Graph Structure Refining for Incomplete MVC (GSRIMC). GSRIMC avoids recovering feature steps and just fully explores the existing subgraphs of each view to produce superior clustering results. To handle the biased error, the biased error separation is the core step of GSRIMC. In detail, GSRIMC first extracts basic information from the precomputed subgraph of each view and then separates refined graph structure from biased error with the help of tensor nuclear norm. Besides, cross-view graph learning is proposed to capture the missing local graph structure and complete the refined graph structure based on the complementary principle. Extensive experiments show that our method achieves better performance than other state-of-the-art baselines. Xiang-Long Li, Man-Sheng Chen, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Incomplete Data Meets Uncoupled Case: A Challenging Task of Multiview ClusteringabstractIncomplete multiview clustering (IMC) methods have achieved remarkable progress by exploring the complementary information and consensus representation of incomplete multiview data. However, to our best knowledge, none of the existing methods attempts to handle the uncoupled and incomplete data simultaneously, which affects their generalization ability in real-world scenarios. For uncoupled incomplete data, the unclear and partial cross-view correlation introduces the difficulty to explore the complementary information between views, which results in the unpromising clustering performance for the existing multiview clustering methods. Besides, the presence of hyperparameters limits their applications. To fill these gaps, a novel uncoupled IMC (UIMC) method is proposed in this article. Specifically, UIMC develops a joint framework for feature inferring and recoupling. The high-order correlations of all views are explored by performing a tensor singular value decomposition (t-SVD)-based tensor nuclear norm (TNN) on recoupled and inferred self-representation matrices. Moreover, all hyperparameters of the UIMC method are updated in an exploratory manner. Extensive experiments on six widely used real-world datasets have confirmed the superiority of the proposed method in handling the uncoupled incomplete multiview data compared with the state-of-the-art methods. Jia-Qi Lin 0001, Xiang-Long Li, Man-Sheng Chen, Chang-Dong Wang 0001, Haizhang Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | On Regularizing Multiple Clusterings for Ensemble Clustering by Graph Tensor LearningabstractEnsemble clustering has shown its promising ability in fusing multiple base clusterings into a probably better and more robust clustering result. Typically, the co-association matrix based ensemble clustering methods attempt to integrate multiple connective matrices from base clusterings by weighted fusion to acquire a common graph representation. However, few of them are aware of the potential noise or corruption from the common representation by direct integration of different connective matrices with distinct cluster structures, and further consider the mutual information propagation between the input observations. In this paper, we propose a Graph Tensor Learning based Ensemble Clustering (GTLEC) method to refine multiple connective matrices by the substantial rank recovery and graph tensor learning. Within this framework, each input connective matrix is dexterously refined to approximate a graph structure by obeying the theoretical rank constraint with an adaptive weight coefficient. Further, we stack multiple refined connective matrices into a three-order tensor to extract their higher-order similarities via graph tensor learning, where the mutual information propagation across different graph matrices will also be promoted. Extensive experiments on several challenging datasets have confirmed the superiority of GTLEC compared with the state-of-the-art. Man-Sheng Chen, Jia-Qi Lin 0001, Chang-Dong Wang 0001, Wudong Xi, Dong Huang 0001 |
ACM Multimedia | 1 |
| 2023 | Signal Contrastive Enhanced Graph Collaborative Filtering for RecommendationabstractAbstract Graph collaborative filtering methods have shown great performance improvements compared with deep neural network-based models. However, these methods suffer from data sparsity and data noise problems. To address these issues, we propose a new contrastive learning-based graph collaborative filtering method to learn more robust representations. The proposed method is called signal contrastive enhanced graph collaborative filtering (SC-GCF), which conducts contrastive learning on graph signals. It has been proved that graph neural networks correspond to low-pass filters on the graph signals from the graph convolution perspective. Different from the previous contrastive learning-based methods, we first pay attention to the diversity of graph signals to directly optimize the informativeness of the graph signals. We introduce a hypergraph module to strengthen the representation learning ability of graph neural networks. The hypergraph learning module utilizes a learnable hypergraph structure to model the latent global dependency relations that graph neural networks cannot depict. Experiments are conducted on four public datasets, and the results show significant improvements compared with the state-of-the-art methods, which confirms the importance of considering signal-level contrastive learning and hypergraph learning. Man-Sheng Chen, Yuefang Gao, Chang-Dong Wang 0001 |
Data Sci. Eng. | 2 |
| 2023 | Low-Rank Tensor Based Proximity Learning for Multi-View ClusteringabstractGraph-oriented multi-view clustering methods have achieved impressive performances by employing relationships and complex structures hidden in multi-view data. However, most of them still suffer from the following two common problems. (1) They target at studying a common representation or pairwise correlations between views, neglecting the comprehensiveness and deeper higher-order correlations among multiple views. (2) The prior knowledge of view-specific representation can not be taken into account to obtain the consensus indicator graph in a unified graph construction and clustering framework. To deal with these problems, we propose a novel Low-rank Tensor Based Proximity Learning (LTBPL) approach for multi-view clustering, where multiple low-rank probability affinity matrices and consensus indicator graph reflecting the final performances are jointly studied in a unified framework. Specifically, multiple affinity representations are stacked in a low-rank constrained tensor to recover their comprehensiveness and higher-order correlations. Meanwhile, view-specific representation carrying different adaptive confidences is jointly linked with the consensus indicator graph. Extensive experiments on nine real-world datasets indicate the superiority of LTBPL compared with the state-of-the-art methods. Man-Sheng Chen, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Efficient Orthogonal Multi-view Subspace ClusteringabstractMulti-view subspace clustering targets at clustering data lying in a union of low-dimensional subspaces. Generally, an n X n affinity graph is constructed, on which spectral clustering is then performed to achieve the final clustering. Both graph construction and graph partitioning of spectral clustering suffer from quadratic or even cubic time and space complexity, leading to difficulty in clustering large-scale datasets. Some efforts have recently been made to capture data distribution in multiple views by selecting key anchor bases beforehand with k-means or uniform sampling strategy. Nevertheless, few of them pay attention to the algebraic property of the anchors. How to learn a set of high-quality orthogonal bases in a unified framework, while maintaining its scalability for very large datasets, remains a big challenge. In view of this, we propose an Efficient Orthogonal Multi-view Subspace Clustering (OMSC) model with almost linear complexity. Specifically, the anchor learning, graph construction and partition are jointly modeled in a unified framework. With the mutual enhancement of each other, a more discriminative and flexible anchor representation and cluster indicator can be jointly obtained. An alternate minimizing strategy is developed to deal with the optimization problem, which is proved to have linear time complexity w.r.t. the sample number. Extensive experiments have been conducted to confirm the superiority of the proposed OMSC method. The source codes and data are available at https://github.com/ManshengChen/Code-for-OMSC-master. Man-Sheng Chen, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai, Philip S. Yu |
KDD | 1 |
| 2022 | Adaptively-weighted Integral Space for Fast Multiview ClusteringabstractMultiview clustering has been extensively studied to take advantage of multi-source information to improve the clustering performance. In general, most of the existing works typically compute an $n\times n$ affinity graph by some similarity/distance metrics (e.g. the Euclidean distance) or learned representations, and explore the pairwise correlations across views. But unfortunately, a quadratic or even cubic complexity is often needed, bringing about difficulty in clustering large-scale datasets. Some efforts have been made recently to capture data distribution in multiple views by selecting view-wise anchor representations with k-means, or by direct matrix factorization on the original observations. Despite the significant success, few of them have considered the view-insufficiency issue, implicitly holding the assumption that each individual view is sufficient to recover the cluster structure. Moreover, the latent integral space as well as the shared cluster structure from multiple insufficient views is not able to be simultaneously discovered. In view of this, we propose an \underlineA daptively-weighted \underlineI ntegral Space for Fast \underlineM ultiview \underlineC lustering (AIMC) with nearly linear complexity. Specifically, view generation models are designed to reconstruct the view observations from the latent integral space with diverse adaptive contributions. Meanwhile, a centroid representation with orthogonality constraint and cluster partition are seamlessly constructed to approximate the latent integral space. An alternate minimizing algorithm is developed to solve the optimization problem, which is proved to have linear time complexityw.r.t. the sample size. Extensive experiments conducted on several real-world datasets confirm the superiority of the proposed AIMC method compared with the state-of-the-art methods. Man-Sheng Chen, Tuo Liu, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
ACM Multimedia | 1 |
| 2022 | Basket Booster for Prototype-based Contrastive Learning in Next Basket Recommendation
Ting-Ting Su, Zhenyu He 0009, Man-Sheng Chen, Chang-Dong Wang 0001 |
ECML/PKDD (1) | 3 |
| 2022 | Coupled Learning for Kernel Representation and Graph Tensor in Multi-view Subspace Clustering
Man-Sheng Chen, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
PRCV (2) | 1 |
| 2022 | Representation Learning in Multi-view Clustering: A Literature ReviewabstractAbstract Multi-view clustering (MVC) has attracted more and more attention in the recent few years by making full use of complementary and consensus information between multiple views to cluster objects into different partitions. Although there have been two existing works for MVC survey, neither of them jointly takes the recent popular deep learning-based methods into consideration. Therefore, in this paper, we conduct a comprehensive survey of MVC from the perspective of representation learning. It covers a quantity of multi-view clustering methods including the deep learning-based models, providing a novel taxonomy of the MVC algorithms. Furthermore, the representation learning-based MVC methods can be mainly divided into two categories, i.e., shallow representation learning-based MVC and deep representation learning-based MVC, where the deep learning-based models are capable of handling more complex data structure as well as showing better expression. In the shallow category, according to the means of representation learning, we further split it into two groups, i.e., multi-view graph clustering and multi-view subspace clustering. To be more comprehensive, basic research materials of MVC are provided for readers, containing introductions of the commonly used multi-view datasets with the download link and the open source code library. In the end, some open problems are pointed out for further investigation and development. Man-Sheng Chen, Jia-Qi Lin 0001, Xiang-Long Li, Bao-Yu Liu, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
Data Sci. Eng. | 1 |
| 2022 | Multiview Subspace Clustering With Grouping EffectabstractMultiview subspace clustering (MVSC) is a recently emerging technique that aims to discover the underlying subspace in multiview data and thereby cluster the data based on the learned subspace. Though quite a few MVSC methods have been proposed in recent years, most of them cannot explicitly preserve the locality in the learned subspaces and also neglect the subspacewise grouping effect, which restricts their ability of multiview subspace learning. To address this, in this article, we propose a novel MVSC with grouping effect (MvSCGE) approach. Particularly, our approach simultaneously learns the multiple subspace representations for multiple views with smooth regularization, and then exploits the subspacewise grouping effect in these learned subspaces by means of a unified optimization framework. Meanwhile, the proposed approach is able to ensure the cross-view consistency and learn a consistent cluster indicator matrix for the final clustering results. Extensive experiments on several benchmark datasets have been conducted to validate the superiority of the proposed approach. Man-Sheng Chen, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001, Philip S. Yu |
IEEE Trans. Cybern. | 1 |
| 2021 | Smoothness Regularized Multiview Subspace Clustering With Kernel LearningabstractMultiview subspace clustering has attracted an increasing amount of attention in recent years. However, most of the existing multiview subspace clustering methods assume linear relations between multiview data points when learning the affinity representation by means of the self-expression or fail to preserve the locality property of the original feature space in the learned affinity representation. To address the above issues, in this article, we propose a new multiview subspace clustering method termed smoothness regularized multiview subspace clustering with kernel learning (SMSCK). To capture the nonlinear relations between multiview data points, the proposed model maps the concatenated multiview observations into a high-dimensional kernel space, in which the linear relations reflect the nonlinear relations between multiview data points in the original space. In addition, to explicitly preserve the locality property of the original feature space in the learned affinity representation, the smoothness regularization is deployed in the subspace learning in the kernel space. Theoretical analysis has been provided to ensure that the optimal solution of the proposed model meets the grouping effect. The unique optimal solution of the proposed model can be obtained by an optimization strategy and the theoretical convergence analysis is also conducted. Extensive experiments are conducted on both image and document data sets, and the comparison results with state-of-the-art methods demonstrate the effectiveness of our method. Chang-Dong Wang 0001, Man-Sheng Chen, Ling Huang 0002, Jian-Huang Lai, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Multi-View Clustering in Latent Embedding SpaceabstractPrevious multi-view clustering algorithms mostly partition the multi-view data in their original feature space, the efficacy of which heavily and implicitly relies on the quality of the original feature presentation. In light of this, this paper proposes a novel approach termed Multi-view Clustering in Latent Embedding Space (MCLES), which is able to cluster the multi-view data in a learned latent embedding space while simultaneously learning the global structure and the cluster indicator matrix in a unified optimization framework. Specifically, in our framework, a latent embedding representation is firstly discovered which can effectively exploit the complementary information from different views. The global structure learning is then performed based on the learned latent embedding representation. Further, the cluster indicator matrix can be acquired directly with the learned global structure. An alternating optimization scheme is introduced to solve the optimization problem. Extensive experiments conducted on several real-world multi-view datasets have demonstrated the superiority of our approach. Man-Sheng Chen, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001 |
AAAI | 1 |
| 2019 | Multi-view Spectral Clustering via Multi-view Weighted Consensus and Matrix-Decomposition Based Discretization
Man-Sheng Chen, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001 |
DASFAA (1) | 1 |