EDBT 2026 Demo / reviewers in the wild / expert
Jinyu Cai
dblp:223/9427
· DBLP profile ↗
31ranked-venue papers
17as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 14 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unifying Multi-View Knowledge for Graph Learning via Model CollaborationabstractWith the increasing scale and complexity of graph data, node attributes are also becoming richer and more complex, particularly in the form of informative text. Classic GNNs equipped with shallow attribute encoders are no longer sufficient to handle such data independently, making model collaboration across heterogeneous architectures an inevitable trend. Recently, the integration of Large Language Models (LLMs) and GNNs has attracted significant attention, yet the inherent disparity between these models remains a key challenge. Promising solutions have considered fine-tuning Small Language Models (SLMs) to bridge the gap between GNNs and frozen LLMs. However, this introduces another problem: these heterogeneous models bring complementary knowledge, but how to effectively integrate them and allow mutual refinement becomes a significant research gap. To address these challenges, we introduce COLA, a collaborative large–small model framework that enables seamless cooperation among semantic LLMs, task-specific fine-tuned SLMs, and structure-aware GNNs. COLA features a unique Consensus–Complement Coordination Mechanism (C3M), wherein its Mixture-of-Coordinators (MoC) architecturally aligns the LLM and SLM. Built upon this, a flexible graph-knowledge infusion strategy encourages the joint alignment and graph knowledge learning of textual representations. Extensive evaluations across nine diverse datasets show that COLA consistently achieves state-of-the-art performance, validating the effectiveness and generality of our collaborative paradigm. Zhihao Wu 0003, Jielong Lu, Jinyu Cai, Guangyong Chen, Jiajun Bu, Haishuai Wang |
AAAI | 4 |
| 2026 | SEAR: LLM-Powered Sequential Recommendation via Fusion of Collaborative, Semantic, and Rating InformationabstractAs users' preferences evolve over time, personalized online services increasingly rely on sequential recommender systems to predict future interactions by modeling patterns in historical user behavior. However, existing methods for sequential recommendation (SR) face two key challenges: they struggle to simultaneously leverage collaborative, semantic, and rating information, and the use of hard labels during training provides limited supervision. In this paper, we introduce SEAR, an LLM-powered Sequential recommEndation framework via fusion of collAborative, semantic, and Rating information. The proposed deep model comprises an embedding layer and a sequence encoder. The embedding layer transforms user-item interactions into three types of embeddings: collaborative, semantic, and rating. The sequence encoder then integrates these embeddings and identifies sequential patterns to model user representations. To enhance the utilization of item semantics, we integrate a large language model (LLM) to extract LLM embeddings. These embeddings are then employed to initialize the semantic embedding layer, collaborative embedding layer, and item embeddings. To capture more nuanced user behavior patterns, we generate preference-weighted soft labels based on the next k interactions. Extensive experiments validate the effectiveness of SEAR, and ablation studies further highlight the distinct contributions of the collaborative, semantic, and rating information. Wei Guan 0006, Jian Cao 0001, Qiqi Cai, Jianqi Gao 0001, Jinyu Cai, See-Kiong Ng |
WWW | 5 |
| 2026 | SAGE: Semantic-aware gray-box game regression testing with large language models
Jinyu Cai, Jialong Li 0001, Nianyu Li, Zhenyu Mao, Mingyue Zhang 0002, Kenji Tei |
Autom. Softw. Eng. | 1 |
| 2026 | MoEGAD: A Mixture-of-Experts Framework With Pseudo-Anomaly Generation for Graph-Level Anomaly DetectionabstractGraph-level anomaly detection (GLAD) aims to identify graphs that significantly deviate from the norm. Despite remarkable advancements in recent years, existing GLAD approaches struggle with the scarcity of labeled anomalies. Although some semi-supervised approaches leverage a small fraction of anomalous graphs during training, the limited diversity of these anomalies poses challenges in learning robust decision boundaries. Additionally, the detection of multi-task graph anomalies, a prevalent challenge in real-world scenarios, remains largely unexplored. To bridge these gaps, we propose MoEGAD, a novel framework leveraging a mixture of experts (MoE) architecture for GLAD. MoEGAD introduces an iterative anomalous graph generation module to produce pseudo-anomalous graphs, which facilitates the subsequent decision boundary learning. An early stopping mechanism is incorporated to ensure that the generated anomalies preserve sufficient dissimilarity from normal graphs. More importantly, we also propose a latent MoE module comprising multiple expert networks alongside a specialized gating network, which promotes cross-task adaptability for diverse GLAD problems. To the best of our knowledge, this is the first work exploring the potential of MoE architecture in the context of GLAD. Extensive experiments across single-task, large-scale, and multi-task scenarios demonstrate that MoEGAD significantly outperforms state-of-the-art GLAD baselines. Jinyu Cai, Yunhe Zhang 0001, Pengyang Wang, See-Kiong Ng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | discDC: Unsupervised discriminative deep image clustering via confidence-driven self-labeling
Jinyu Cai, Wenzhong Guo, Yunhe Zhang 0001, Jicong Fan 0001 |
Pattern Recognit. | 1 |
| 2025 | Mixture of Experts as Representation Learner for Deep Multi-View ClusteringabstractMulti-view clustering (MVC) aims to integrate information from diverse data sources to facilitate the clustering process, which has achieved considerable success in various real-world applications. However, previous MVC methods typically employ one of two strategies: (1) designing separate feature extraction pipelines for each view, which restricts their ability to fully exploit collaborative potential; or (2) employing a single shared representation module, which hinders the capture of diverse, view-specific representations. To tackle these challenges, we introduce Deep Multi-View Clustering via Collaborative Experts (DMVC-CE), a novel MVC approach that employs the Mixture of Experts (MoE) framework. DMVC-CE incorporates a gating network that dynamically selects multiple experts for handling each data sample, capturing diverse and complementary information from different views. Additionally, to ensure balanced expert utilization and maintain their diversity, we introduce an equilibrium loss and a multi-expert distinctiveness enhancer. The equilibrium loss prevents excessive reliance on specific experts, while the distinctiveness enhancer encourages each expert to specialize in different aspects of the data, thereby promoting diversity in learned representations. Comprehensive experiments on various multi-view benchmark datasets demonstrate the superiority of DMVC-CE compared to state-of-the-art MVC baselines. Yunhe Zhang 0001, Jinyu Cai, Zhihao Wu 0002, Pengyang Wang, See-Kiong Ng |
AAAI | 2 |
| 2025 | Multi-to-Single: Reducing Multimodal Dependency in Emotion Recognition Through Contrastive LearningabstractMultimodal emotion recognition is a crucial research area in the field of affective brain-computer interfaces. However, in practical applications, it is often challenging to obtain all modalities simultaneously. To deal with this problem, researchers focus on using cross-modal methods to learn multimodal representations with fewer modalities. However, due to the significant differences in the distribution of different modalities, it is challenging to enable any modality to fully learn multimodal features. To address this limitation, we propose a Multi-to-Single (M2S) emotion recognition model, leveraging contrastive learning and incorporating two innovative modules: 1) a spatial and temporal-sparse (STS) attention mechanism that enhances the encoders' ability to extract features from data; 2) a novel Multi-to-Multi Contrastive Predictive Coding (M2M CPC) that learns and fuses features across different modalities. In the final testing, we only use a single modality for emotion recognition, reducing the dependence on multimodal data. Extensive experiments on five public multimodal emotion datasets demonstrate that our model achieves the state-of-the-art performance in the cross-modal tasks and maintains multimodal performance using only a single modality. Yan-Kai Liu, Jinyu Cai, Bao-Liang Lu, Wei-Long Zheng |
AAAI | 2 |
| 2025 | Self-Discriminative Modeling for Anomalous Graph DetectionabstractIdentifying anomalous graphs is essential in real-world scenarios such as molecular and social network analysis, yet anomalous samples are generally scarce and unavailable. This paper proposes a Self-Discriminative Modeling (SDM) framework that trains a deep neural network only on normal graphs to detect anomalous graphs. The neural network simultaneously learns to construct pseudo-anomalous graphs from normal graphs and learns an anomaly detector to recognize these pseudo-anomalous graphs. As a result, these pseudo-anomalous graphs interpolate between normal graphs and real anomalous graphs, which leads to a reliable decision boundary of anomaly detection. In this framework, we develop three algorithms with different computational efficiencies and stabilities for anomalous graph detection. Extensive experiments on 12 different graph benchmarks demonstrated that the three variants of SDM consistently outperform the state-of-the-art GLAD baselines. The success of our methods stems from the integration of the discriminative classifier and the well-posed pseudo-anomalous graphs, which provided new insights for graph-level anomaly detection. Jinyu Cai, Yunhe Zhang 0001, Jicong Fan 0001 |
ICML | 1 |
| 2025 | Leveraging Diffusion Model as Pseudo-Anomalous Graph Generator for Graph-Level Anomaly DetectionabstractA fundamental challenge in graph-level anomaly detection (GLAD) is the scarcity of anomalous graph data, as the training dataset typically contains only normal graphs or very few anomalies. This imbalance hinders the development of robust detection models. In this paper, we propose Anomalous Graph Diffusion (AGDiff), a framework that explores the potential of diffusion models in generating pseudo-anomalous graphs for GLAD. Unlike existing diffusion-based methods that focus on modeling data normality, AGDiff leverages the latent diffusion framework to incorporate subtle perturbations into graph representations, thereby generating pseudo-anomalous graphs that closely resemble normal ones. By jointly training a classifier to distinguish these generated graph anomalies from normal graphs, AGDiff learns more discriminative decision boundaries. The shift from solely modeling normality to explicitly generating and learning from pseudo graph anomalies enables AGDiff to effectively identify complex anomalous patterns that other approaches might overlook. Comprehensive experimental results demonstrate that the proposed AGDiff significantly outperforms several state-of-the-art GLAD baselines. Jinyu Cai, Yunhe Zhang 0001, Fusheng Liu, See-Kiong Ng |
ICML | 1 |
| 2025 | Self-Perturbed Anomaly-Aware Graph Dynamics for Multivariate Time-Series Anomaly DetectionabstractDetecting anomalies in multivariate time-series data is an essential task across various domains, yet there are unresolved challenges such as (1) severe class imbalance between normal and anomalous data due to rare anomaly availability in the real world; (2) limited adaptability of the static graph-based methods to dynamically changing inter-variable correlations; and (3) neglect of subtle anomalies due to overfitting to normal patterns in reconstruction-based methods. To tackle these issues, we propose Self-Perturbed Anomaly-Aware Graph Dynamics (SPAGD), a framework for time-series anomaly detection. SPAGD employs a self-perturbation module that generates self-perturbed time series from the reconstruction process of normal ones, which provide auxiliary signals to alleviate class imbalance during training. Concurrently, an anomaly-aware graph construction module is proposed to dynamically adjust the graph structure by leveraging the reconstruction residuals of self-perturbed time series, thereby emphasizing the inter-variable disruptions induced by anomalous candidates. A unified spatio-temporal anomaly detection module then integrates both spatial and temporal convolutions to train a classifier that distinguishes normal time series from the auxiliary self-perturbed samples. Extensive experiments across multiple benchmark datasets demonstrate the effectiveness of SPAGD compared to state-of-the-art baselines. Jinyu Cai, Glynnis Lim, Yifang Yin, Roger Zimmermann, See-Kiong Ng |
NeurIPS | 1 |
| 2025 | Where Graph Meets Heterogeneity: Multi-View Collaborative Graph ExpertsabstractThe convergence of graph learning and multi-view learning has propelled the emergence of multi-view graph neural networks (MGNNs), offering strong capabilities to address complex real-world data characterized by heterogeneous yet interconnected information.
While existing MGNNs exploit the potential of multi-view graphs, the inherent conflict persists between the two critical inductive biases of multi-view learning, consistency and complementarity. Consequently, the challenge of defining and resolving this tension in the new context of multi-view graphs remains largely underexplored. To bridge this gap, we propose Multi-view Collaborative Graph Experts (MvCGE), a novel framework grounded in the Mixture-of-Experts (MoE) paradigm. MvCGE establishes architectural consistency through shared parameters while preserving complementarity via layer-wise collaborative graph experts, which are dynamically activated by a graph-aware routing mechanism that adapts to the structural nuances of each view. This dual-level design is further reinforced by two novel components: a load equilibrium loss to prevent expert collapse and ensure balanced specialization, and a graph discrepancy loss based on distributional divergence to enhance inter-view complementarity. Extensive experiments on diverse datasets demonstrate MvCGE’s superiority. Zhihao Wu 0003, Jinyu Cai, Yunhe Zhang 0001, Jielong Lu, Zhaoliang Chen, Shuman Zhuang, Haishuai Wang |
NeurIPS | 2 |
| 2025 | Detection and pose measurement of underground drill pipes based on GA-PointNet++
Jiangnan Luo, Jinyu Cai, Deyi Zhang, Jiuhua Gao, Liu Lei, Mengda Hao |
Appl. Intell. | 2 |
| 2025 | MULTI: multimodal understanding leaderboard with text and images
Lu Chen 0002, Jingkai Yang, Yichuan Ma, Hailin Wen, Jinyu Cai, Yingzi Ma, Situo Zhang, Zihan Zhao 0001, Liangtai Sun, Kai Yu 0004 |
Sci. China Inf. Sci. | 9 |
| 2024 | Language Evolution for Evading Social Media Regulation via LLM-Based Multi-Agent SimulationabstractSocial media platforms such as Twitter, Reddit, and Sina Weibo playa crucial role in global communication but often encounter strict regulations in geopolitically sensitive regions. This situation has prompted users to ingeniously modify their way of communicating, frequently resorting to coded language in these regulated social media environments. This shift in communication is not merely a strategy to counteract regulation, but a vivid manifestation of language evolution, demonstrating how language naturally evolves under societal and technological pressures. Studying the evolution of language in regulated social media contexts is of significant importance for ensuring freedom of speech, optimizing content moderation, and advancing linguistic research. This paper proposes a multi-agent simulation frame-work using Large Language Models (LLMs) to explore the evolution of user language in regulated social media environments. The framework employs LLM-driven agents: supervisory agent who enforce dialogue supervision and participant agents who evolve their language strategies while engaging in conversation, simulating the evolution of communication styles under strict regulations aimed at evading social media regulation. The study evaluates the framework's effectiveness through a range of scenarios from abstract scenarios to real-world situations. Key findings indicate that LLMs are capable of simulating nuanced language dynamics and interactions in constrained settings, showing improvement in both evading supervision and information accuracy as evolution progresses. Furthermore, it was found that LLM agents adopt different strategies for different scenarios. The reproduction kit can be accessed at https://github.com/BlueLinkXlGA-MAS. Jinyu Cai, Jialong Li 0001, Mingyue Zhang 0002, Munan Li, Chen-Shu Wang, Kenji Tei |
CEC | 1 |
| 2024 | Deep Orthogonal Hypersphere Compression for Anomaly DetectionabstractMany well-known and effective anomaly detection methods assume that a reasonable decision boundary has a hypersphere shape, which however is difficult to obtain in practice and is not sufficiently compact, especially when the data are in high-dimensional spaces. In this paper, we first propose a novel deep anomaly detection model that improves the original hypersphere learning through an orthogonal projection layer, which ensures that the training data distribution is consistent with the hypersphere hypothesis, thereby increasing the true positive rate and decreasing the false negative rate. Moreover, we propose a bi-hypersphere compression method to obtain a hyperspherical shell that yields a more compact decision region than a hyperball, which is demonstrated theoretically and numerically. The proposed methods are not confined to common datasets such as image and tabular data, but are also extended to a more challenging but promising scenario, graph-level anomaly detection, which learns graph representation with maximum mutual information between the substructure and global structure features while exploring orthogonal single- or bi-hypersphere anomaly decision boundaries. The numerical and visualization results on benchmark datasets demonstrate the superiority of our methods in comparison to many baselines and state-of-the-art methods. Yunhe Zhang 0001, Jinyu Cai, Jicong Fan 0001 |
ICLR | 3 |
| 2024 | Dual Contrastive Graph-Level Clustering with Multiple Cluster Perspectives Alignment
Jinyu Cai, Yunhe Zhang 0001, Jicong Fan 0001, Yali Du 0001, Wenzhong Guo |
IJCAI | 1 |
| 2024 | LG-FGAD: An Effective Federated Graph Anomaly Detection Framework
Jinyu Cai, Yunhe Zhang 0001, Jicong Fan 0001, See-Kiong Ng |
IJCAI | 1 |
| 2024 | Kernel Readout for Graph Neural Networks
Jiajun Yu, Zhihao Wu 0002, Jinyu Cai, Adele Lu Jia, Jicong Fan 0001 |
IJCAI | 3 |
| 2024 | Towards Effective Federated Graph Anomaly Detection via Self-boosted Knowledge DistillationabstractGraph anomaly detection (GAD) aims to identify anomalous graphs that significantly deviate from other ones, which has raised growing attention due to the broad existence and complexity of graph-structured data in many real-world scenarios. However, existing GAD methods usually execute with centralized training, which may lead to privacy leakage risk in some sensitive cases, thereby impeding collaboration among organizations seeking to collectively develop robust GAD models. Although federated learning offers a promising solution, the prevalent non-IID problems and high communication costs present significant challenges, particularly pronounced in collaborations with graph data distributed among different participants. To tackle these challenges, we propose an effective federated graph anomaly detection framework (FGAD). We first introduce an anomaly generator to perturb the normal graphs to be anomalous and train a powerful anomaly detector by distinguishing generated anomalous graphs from normal ones. We subsequently leverage a student model to distill knowledge from the trained anomaly detector (teacher model), which aims to maintain the personality of local models and alleviate the adverse impact of non-IID problems. Additionally, we design an effective collaborative learning mechanism that facilitates the personalization preservation of local models and significantly reduces communication costs among clients. Empirical results of diverse GAD tasks demonstrate the superiority and efficiency of FGAD. Jinyu Cai, Yunhe Zhang 0001, Zhoumin Lu, Wenzhong Guo, See-Kiong Ng |
ACM Multimedia | 1 |
| 2024 | Deep graph-level clustering using pseudo-label-guided mutual information maximization network
Jinyu Cai, Wenzhong Guo, Jicong Fan 0001 |
Neural Comput. Appl. | 1 |
| 2024 | Deep Masked Graph Node ClusteringabstractIn recent years, reconstructing features and learning node representations by graph autoencoders (GAE) have attracted much attention in deep graph node clustering. However, existing works often overemphasize structural information and overlook the impact of real-world prevalent noise on feature learning and clustering with graph data, which may be detrimental to robust training. To address these issues, the utilization of a masking strategy that specifically focuses on feature reconstruction may mitigate these limitations. In this article, we propose a graph node clustering generative method named deep masked graph node clustering (DMGNC), which leverages a masked autoencoder to effectively reconstruct node features, enabling the discovery of latent information crucial for accurate node clustering. Additionally, a clustering self-optimization module is designed to guide the iterative update of our end-to-end clustering framework. Further, we extend the masked graph autoencoder (MGA) and develop a contrastive method called deep masked graph node contrastive clustering (DMGNCC), which applies the MGA to graph node contrastive learning at both the node level and the class level in a united model. Extensive experimental results on real-world graph benchmark datasets demonstrate the effectiveness and superiority of the proposed method. Jinbin Yang, Jinyu Cai, Luying Zhong, Yueyang Pi, Shiping Wang |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Wasserstein Embedding Learning for Deep Clustering: A Generative ApproachabstractDeep learning-based clustering methods, especially those incorporating deep generative models, have recently shown noticeable improvement on many multimedia benchmark datasets. However, existing generative models still suffer from unstable training, and the gradient vanishes, which results in the inability to learn desirable embedded features for clustering. In this paper, we aim to tackle this problem by exploring the capability of Wasserstein embedding in learning representative embedded features and introducing a new clustering module for jointly optimizing embedding learning and clustering. To this end, we propose Wasserstein embedding clustering (WEC), which integrates robust generative models with clustering. By directly minimizing the discrepancy between the prior and marginal distribution, we transform the optimization problem of Wasserstein distance from the original data space into embedding space, which differs from other generative approaches that optimize in the original data space. Consequently, it naturally allows us to construct a joint optimization framework with the designed clustering module in the embedding layer. Due to the substitutability of the penalty term in Wasserstein embedding, we further propose two types of deep clustering models by selecting different penalty terms. Comparative experiments conducted on nine publicly available multimedia datasets with several state-of-the-art methods demonstrate the effectiveness of our method. Jinyu Cai, Yunhe Zhang 0001, Shiping Wang, Jicong Fan 0001, Wenzhong Guo |
IEEE Trans. Multim. | 1 |
| 2023 | Label correction using contrastive prototypical classifier for noisy label learning
Chaoyang Xu, Renjie Lin, Jinyu Cai, Shiping Wang |
Inf. Sci. | 3 |
| 2022 | Efficient Deep Embedded Subspace ClusteringabstractRecently deep learning methods have shown significant progress in data clustering tasks. Deep clustering methods (including distance-based methods and subspace-based methods) integrate clustering and feature learning into a unified framework, where there is a mutual promotion between clustering and representation. However, deep subspace clustering methods are usually in the framework of self-expressive model and hence have quadratic time and space complexities, which prevents their applications in large-scale clustering and real-time clustering. In this paper, we propose a new mechanism for deep clustering. We aim to learn the subspace bases from deep representation in an iterative refining manner while the refined subspace bases help learning the representation of the deep neural networks in return. The proposed method is out of the self-expressive framework, scales to the sample size linearly, and is applicable to arbitrarily large datasets and online clustering scenarios. More importantly, the clustering accuracy of the proposed method is much higher than its competitors. Extensive comparison studies with state-of-the-art clustering approaches on benchmark datasets demonstrate the superiority of the proposed method. Jinyu Cai, Jicong Fan 0001, Wenzhong Guo, Shiping Wang, Yunhe Zhang 0001, Zhao Zhang 0001 |
CVPR | 1 |
| 2022 | Value Iteration Residual Network with Self-attention
Jinyu Cai, Jialong Li 0001, Zhenyu Mao, Kenji Tei |
ISDA (3) | 1 |
| 2022 | Perturbation Learning Based Anomaly DetectionabstractThis paper presents a simple yet effective method for anomaly detection. The main idea is to learn small perturbations to perturb normal data and learn a classifier to classify the normal data and the perturbed data into two different classes. The perturbator and classifier are jointly learned using deep neural networks. Importantly, the perturbations should be as small as possible but the classifier is still able to recognize the perturbed data from unperturbed data. Therefore, the perturbed data are regarded as abnormal data and the classifier provides a decision boundary between the normal data and abnormal data, although the training data do not include any abnormal data.Compared with the state-of-the-art of anomaly detection, our method does not require any assumption about the shape (e.g. hypersphere) of the decision boundary and has fewer hyper-parameters to determine. Empirical studies on benchmark datasets verify the effectiveness and superiority of our method. Jinyu Cai, Jicong Fan 0001 |
NeurIPS | 1 |
| 2022 | Deep image clustering by fusing contrastive learning and neighbor relation mining
Chaoyang Xu, Renjie Lin, Jinyu Cai, Shiping Wang |
Knowl. Based Syst. | 3 |
| 2022 | Unsupervised deep clustering via contractive feature representation and focal loss
Jinyu Cai, Shiping Wang, Chaoyang Xu, Wenzhong Guo |
Pattern Recognit. | 1 |
| 2021 | Unsupervised embedded feature learning for deep clustering with stacked sparse auto-encoder
Jinyu Cai, Shiping Wang, Wenzhong Guo |
Expert Syst. Appl. | 1 |
| 2020 | Unsupervised discriminative feature representation via adversarial auto-encoder
Wenzhong Guo, Jinyu Cai, Shiping Wang |
Appl. Intell. | 2 |
| 2019 | An Overview of Unsupervised Deep Feature Representation for Text CategorizationabstractHigh-dimensional features are extensively accessible in machine learning and computer vision areas. How to learn an efficient feature representation for specific learning tasks is invariably a crucial issue. Due to the absence of class label information, unsupervised feature representation is exceedingly challenging. In the last decade, deep learning has captured growing attention from researchers in a broad range of areas. Most of the deep learning methods are supervised, which is required to be fed with a large amount of accurately labeled data points. Nevertheless, acquiring sufficient accurately labeled data is unaffordable in numerous real-world applications, which is suggestive of the needs of unsupervised learning. Toward this end, quite a few unsupervised feature representation approaches based on deep learning have been proposed in recent years. In this paper, we attempt to provide a comprehensive overview of unsupervised deep learning methods and compare their performances in text categorization. Our survey starts with the autoencoder and its representative variants, including sparse autoencoder, stacked autoencoder, contractive autoencoder, denoising autoencoder, variational autoencoder, graph autoencoder, convolutional autoencoder, adversarial autoencoder, and residual autoencoder. Aside from autoencoders, deconvolutional networks, restricted Boltzmann machines, and deep belief nets are introduced. Then, the reviewed unsupervised feature representation methods are compared in terms of text clustering. Extensive experiments in eight publicly available data sets of text documents are conducted to provide a fair test bed for the compared methods. Shiping Wang, Jinyu Cai, Qihao Lin, Wenzhong Guo |
IEEE Trans. Comput. Soc. Syst. | 2 |