Tao Du 0002

dblp:51/3026-2 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-6346-5636ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Density Peak Clustering via Shared-Neighbor Markov Transition Matrix
Yaru Zhang, Rui Wang 0199, Jin Zhou 0003, Tao Du 0002, Dongmei Niu, Shi-Yuan Han, Yingxu Wang 0002
ICIC (13)5
2026 Collaborative multi-view fuzzy clustering based on Gaussian mixture model
Shi-Yuan Han, Jin Zhou 0003, C. L. Philip Chen, Tong Zhang 0015, Yuehui Chen, Lin Wang 0004, Tao Du 0002
Neurocomputing8
2026 Expanded Deep Embedding Clustering With Adversarial Learning and Adaptive Graph Constraint
abstract
The autoencoder (AE) is an efficient feature extraction tool that learns latent representations from raw data by minimizing the reconstruction loss. Building upon the AE architecture, deep clustering models are designed to jointly optimize the deep neural network and perform unsupervised clustering. However, existing methods directly impose the clustering objective on the latent features produced by the AE network, thereby neglecting the potential conflict between data clustering and data representation. Specifically, data clustering aims to enhance data aggregation, whereas data representation focuses on ensuring that latent features faithfully reflect the manifold structure of the raw data. To address this issue, this article proposes an innovative expanded deep embedding clustering (E-DEC) model, in which the AE network is employed to seek better latent representations, and a novel residual expansion module (REM) is integrated to construct an expanded feature space that better serves clustering tasks. Furthermore, adversarial learning between the soft cluster assignments and a prior one-hot distribution is adopted in lieu of the conventional Kullback–Leibler (KL) divergence, so as to enhance the discrimination of different clusters and avoid the degeneracy problem. Finally, an entropy regularization technique is incorporated to adaptively refine the affinity graph throughout the clustering process, thereby reducing the sensitivity of clustering performance to the initial affinity graph. Extensive experiments on real-world benchmark datasets demonstrate the superiority of the proposed model over state-of-the-art deep clustering methods.
Shi-Yuan Han, Jin Zhou 0003, C. L. Philip Chen, Yingxu Wang 0002, Yuehui Chen, Lin Wang 0004, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013
IEEE Trans. Syst. Man Cybern. Syst.8
2025 Incomplete Data Clustering Based on Multiple Imputation and Autoencoders
Jin Zhou 0003, Shi-Yuan Han, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013
ICIC (20)4
2025 Expanded Feature for Deep Embedding Clustering
Jin Zhou 0003, Shi-Yuan Han, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013
ICIC (20)4
2025 Cross-View Representation Learning-Based Deep Multiview Clustering With Adaptive Graph Constraint
abstract
Deep multiview clustering provides an efficient way to analyze the data consisting of multiple modalities and features. Recently, the autoencoder (AE)-based deep multiview clustering algorithms have attracted intensive attention by virtue of their rewarding capabilities of extracting inherent features. Nevertheless, most existing methods are still confronted by several problems. First, the multiview data usually contains abundant cross-view information, thus parallel performing an individual AE for each view and directly combining the extracted latent together can hardly construct an informative view-consensus feature space for clustering. Second, the intrinsic local structures of multiview data are complicated, hence simply embedding a preset graph constraint into multiview clustering models cannot guarantee expected performance. Third, current methods commonly utilize the Kullback-Leibler (KL) divergence as clustering loss and accordingly may yield appalling clusters that lack discriminate characters. To solve these issues, in this article we propose two new AE-based deep multiview clustering algorithms named AE-based deep multiview clustering model incorporating graph embedding (AG-DMC) and deep discriminative multiview clustering algorithm with adaptive graph constraint (ADG-DMC). In AG-DMC, a novel cross-view representation learning model is established delicately by performing decoding processes based on the cascaded view-specific latent to learn sound view-consensus features for inspiring clustering results. In addition, an entropy-regularized adaptive graph constraint is imposed on the obtained soft assignments of data to precisely preserve potential local structures. Furthermore, in the improved model ADG-DMC, the adversarial learning mechanism is adopted as clustering loss to strengthen the discrimination of different clusters for better performance. In the comprehensive experiments carried out on eight real-world datasets, the proposed algorithms have achieved superior performance in the comparison with other advanced multiview clustering algorithms.
Yingxu Wang 0002, Xuesong Wang 0001, C. L. Philip Chen, Long Chen 0001, Yuehui Chen, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013, Jin Zhou 0003
IEEE Trans. Neural Networks Learn. Syst.7
2024 Graph Embedding-Based Deep Multi-view Clustering
Jin Zhou 0003, Shi-Yuan Han, Yingxu Wang 0002, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013
ICIC (2)5
2023 BYOL Network Based Contrastive Clustering
Xuehao Chen, Jin Zhou 0003, Yingxu Wang 0002, Shi-Yuan Han, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013
ICIC (1)6
2023 Graph-Based Short Text Clustering via Contrastive Learning with Graph Embedding
Jin Zhou 0003, Yingxu Wang 0002, Shi-Yuan Han, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013
ICIC (1)6
2023 Deep Multi-view Clustering Based on Graph Embedding
Jin Zhou 0003, Yingxu Wang 0002, Shi-Yuan Han, Tao Du 0002, Cheng Yang 0011, Bowen Liu 0013
ICIC (1)6
2023 A domain density peak clustering algorithm based on natural neighbor
abstract
Density peaks clustering (DPC) is as an efficient algorithm due for the cluster centers can be found quickly. However, this approach has some disadvantages. Firstly, it is sensitive to the cutoff distance; secondly, the neighborhood information of the data is not considered when calculating the local density; thirdly, during allocation, one assignment error may cause more errors. Considering these problems, this study proposes a domain density peak clustering algorithm based on natural neighbor (NDDC). At first, natural neighbor is introduced innovatively to obtain the neighborhood of each point. Then, based on the natural neighbors, several new methods are proposed to calculate corresponding metrics of the points to identify the centers. At last, this study proposes a new two-step assignment strategy to reduce the probability of data misclassification. A series of experiments are conducted that the NDDC offers higher accuracy and robustness than other methods.
Tao Du 0002, Jin Zhou 0003, Tianyu Shen
Intell. Data Anal.2
2023 Transfer-Learning-Based Gaussian Mixture Model for Distributed Clustering
abstract
Distributed clustering based on the Gaussian mixture model (GMM) has exhibited excellent clustering capabilities in peer-to-peer (P2P) networks. However, more iterative numbers and communication overhead are required to achieve the consensus in existing distributed GMM clustering algorithms. In addition, the truth that it cannot find a closed form for the update of parameters in GMM causes the imprecise clustering accuracy. To solve these issues, by utilizing the transfer learning technique, a general transfer distributed GMM clustering framework is exploited to promote the clustering performance and accelerate the clustering convergence. In this work, each node is treated as both the source domain and the target domain, and these nodes can learn from each other to complete the clustering task in distributed P2P networks. Based on this framework, the transfer distributed expectation-maximization algorithm with the fixed learning rate is first presented for data clustering. Then, an improved version is designed to obtain the stable clustering accuracy, in which an adaptive transfer learning strategy is adopted to adjust the learning rate automatically instead of a fixed value. To demonstrate the extensibility of the proposed framework, a representative GMM clustering method, the entropy-type classification maximum-likelihood algorithm, is further extended to the transfer distributed counterpart. Experimental results verify the effectiveness of the presented algorithms in contrast with the existing GMM clustering approaches.
Shi-Yuan Han, Jin Zhou 0003, Yuehui Chen, Lin Wang 0004, Tao Du 0002, Ke Ji, Ya-ou Zhao, Kun Zhang 0013
IEEE Trans. Cybern.6
2023 Transfer Learning-Based Collaborative Multiview Clustering
abstract
Collaborative multiview clustering methods can efficiently realize the view fusion by exploring complementary and consistent information among multiple views. However, these studies ignore all the differences between multiple views in fusion. In fact, in the multiview clustering, the data are diverse from view to view. The larger the difference between any two views is, the more the fusion of these views is required. Moreover, a global tradeoff parameter is generally adopted to restrain the penalty related to the disagreement of all views, which is often defined empirically. Inspired by the idea of transfer learning, a series of novel collaborative multiview clustering algorithms are proposed to tackle these challenges. In the most basic one, each view performs clustering independently and learns from others to improve its own clustering performance, in which a global learning factor is defined to control the interaction between multiple views. The fuzzy memberships are regarded as the important knowledge to provide guidance between views, and the consensus constraint is defined to ensure the consistent partitions of all views. In addition, the local adaptive learning factors between any two views instead of a global fixed one are adopted in an improved version to emphasize the difference between views, and the adjustment strategy for the learning factor is further designed to guarantee the stability of multiview clustering without the influence of initial values. Finally, to identify the significance of different views to the clustering, the extended versions are excavated with the assignment of view weights and the maximum entropy regularization technique is employed to optimize the weights. Experiments on various real-world multiview datasets verify the superiority of the presented approaches.
Xiangdao Liu, Jin Zhou 0003, C. L. Philip Chen, Tong Zhang 0015, Yuehui Chen, Shi-Yuan Han, Tao Du 0002, Ke Ji, Kun Zhang 0013
IEEE Trans. Fuzzy Syst.8
2022 Kernel Fuzzy Clustering based on Quasi-Monte Carlo Feature Map with Neighbor Affinity Constraint
abstract
In recent years, kernel-based fuzzy clustering has attracted significant attention, primarily benefiting from the outstanding performance of capturing the potential non-linear structure in data clustering. However, many existing kernel clustering methods are not available for large datasets due to computational costs. To overcome this limitation, the low-rank random feature map is utilized to approximate the kernel space. Nevertheless, this kind of feature approximation method ignores the graph structure information hidden in the data and does not take the correlations between data samples in the clustering into account. Thus, we present a new kernel fuzzy clustering based on Quasi-Monte Carlo feature map with neighbor affinity constraint (Na_QMC_KFC). In this scheme, the Quasi-Monte Carlo method is adopted to approximate the Gaussian kernel function so as to reduce the computational costs. Meanwhile, the neighbor affinity constraint is designed to maintain the graph structure information of the data and further facilitate the consistency of the membership degrees and the raw data. What’s more, the Alternating Direction Method of Multipliers method is utilized to optimize the problem with respect to the neighbor affinity lasso. The experiments on several non-linear and real-world datasets exhibits the efficiency of the presented algorithm.
Wenpu Zhang, Jin Zhou 0003, Shi-Yuan Han, Lin Wang 0004, Tao Du 0002, Ke Ji
FUZZ-IEEE8
2022 DCE-IVI: Density-based clustering ensemble by selecting internal validity index
abstract
As each clustering algorithm cannot efficiently partition datasets with arbitrary shapes, the thought of clustering ensemble is proposed to consistently integrate clustering results to obtain better division. Most of ensemble research employs a single algorithm with different parameters to clustering. And this can be easily integrated, however it is hardly to divide complex datasets. Other available methods integrate different algorithms, it can divide datasets from different aspects, but fail to take outliers into account, which produces negative effects on the partition results. In order to solve these problems, we clustering datasets with three different density-based algorithms. The innovation of this paper is described as: (1) by setting dynamic thresholds, lower frequency evidence in the co-association matrix is gradually deleted to obtain multiple reconstructed matrices; (2) these reconstructed matrices are analyzed by hierarchical clustering to obtain basic clustering results; (3) an internal validity index is designed by the compactness within clusters and the correlation between clusters, which is used to select the final clustering result. By this innovation, the clustering effect is significantly improved. Finally, a series of experiments are designed, and the results verify the improvement and effectiveness of the proposed technique (DCE-IVI).
Qinlu Li, Tao Du 0002, Jin Zhou 0003, Shouning Qu
Intell. Data Anal.2
2021 Adaptive density-based clustering algorithm with shared KNN conflict game
Tao Du 0002, Shouning Qu
Inf. Sci.2
2019 Research on Parallel Incremental Association Rule Algorithm based on Data Stream
Tao Du 0002, Shouning Qu, Tao Mi
DATA2
2017 A Data Stream Clustering Algorithm Based on Density and Extended Grid
Tao Du 0002, Shouning Qu, Guodong Mou
ICIC (2)2