VLDB 2026 Research / reviewers in the wild / expert
Dong Huang 0001
dblp:94/3756-1
· DBLP profile ↗
82ranked-venue papers
17as first author
53since 2021 · last 2026
0000-0003-3923-8828ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 10 first-author · 33 since 2021Databases, data management, data science and information retrieval · 22 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision transformer for contrastive clustering
Hua-Bao Ling, Dong Huang 0001, Ding-Hua Chen, Chang-Dong Wang 0001, Jian-Huang Lai |
Knowl. Based Syst. | 3 |
| 2026 | One-step bipartite graph cut: A normalized formulation and its application to scalable subspace clustering
Si-Guo Fang, Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai |
Neural Networks | 2 |
| 2026 | Structure-preserving contrastive graph clustering with dual-channel label alignmentabstractThe past few years have witnessed the rapid development of contrastive graph clustering (CGC). Although a series of achievements have been made, there still remain two challenging problems in the literature. First, previous works typically generate different views via some pre-defined graph augmentation strategies, but inappropriate augmentations may alter the latent semantics of the original data. Second, they often overlook the discriminative unsupervised information when constructing positive and negative sample pairs, resulting in compromised clustering performance. Third, some of them are restricted to only static neighborhood connections for contrastive learning, which neglect the dynamical structural relationship via robust neighboring graph learning. To cope with these issues, this paper proposes a Structure-preserving Contrastive Graph Clustering approach with Dual-channel Label Alignment (SCGC-DLA). In terms of the high-and-low frequency issues, the low-pass and hybrid graph filters are designed for generating two views of reliable augmentations, which can supply rich and complementary information to each other. Further, we construct a structure-preserving matrix, which is derived from the edge betweenness centrality (EBC) perspective design and allows us to efficiently capture the topological relationships among different embedding representations. Under the guidance of the non-dominated sorting theory, the clustering distribution information of dual-channel is used to construct high-confidence pseudo labels. Especially, the generated high-confidence pseudo labels are aligned with latent semantic labels. Finally, the overall network is guided by a self-supervised learning scheme and therefore the final clustering could be obtained. Substantial results on five benchmarks prove the robustness and effectiveness of our approach compared to several state-of-the-arts. Yan-Di Huang, Dong Huang 0001, Chang-Dong Wang 0001, Yang Liu 0084, Enbo Huang |
Neural Networks | 3 |
| 2026 | REC-GCN: Robust ensemble clustering with graph convolutional networks
Dong Huang 0001, Qiang Lai, Yuankun Xu, Chang-Bin Guan, Chang-Dong Wang 0001 |
Pattern Recognit. | 1 |
| 2026 | Mutual Neighborhood-Enhanced Protypical Contrastive Learning for ClusteringabstractDeep clustering has achieved remarkable progress due to the strong representation learning capabilities of deep neural networks. However, many popular deep clustering algorithms based on contrastive or non-contrastive strategies focus solely on instance-level or cluster-level representation learning, lacking the ability to explore cross-level information necessary for effective clustering. Moreover, these algorithms often overlook the semantic neighborhood information around instances or clusters, which may lead to suboptimal performance. To address these issues, we propose a novel and efficient end-to-end deep clustering framework termed Mutual Neighborhood-Enhanced Representation Learning for Clustering (MNC), which integrates Neighborhood-Enhanced Prototype-level Contrastive Learning (NPCL) with an Aligning with Consistency Samples (ACS) strategy. In particular, the network architecture comprises four main modules, namely, the feature extraction module, the consistency samples search module, the class-level contrastive module, and the instance-level non-contrastive module. Notably, we extend the conventional non-contrastive instance-level representation learning by incorporating operations on deeper neighborhood structures, leveraging mutual neighborhood relationships from dual views. Furthermore, we perform contrastive learning at the cluster-level based on the compact neighborhood structure in a spherical feature space, which encourages intra-cluster compactness and inter-cluster separability. Extensive experiments on five challenging image benchmarks demonstrate the superiority of our approach over the state-of-the-art. In particular, our approach achieves an ACC of 0.342 (0.804) on the Tiny-ImageNet (STL-10) dataset, which surpasses the second-best method by 7.9% and 4.5%, respectively. The code is available at https://github.com/SandLYJ. Dong Huang 0001, Chang-Dong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | ShrimpFormer-X: A Transformer-Based Framework for Counting and Localization of Shrimp Larvae
Yuefang Gao, Dong Huang 0001 |
WISA | 3 |
| 2025 | Multi-scale Multi-order Attributed Graph Clustering
Dong Huang 0001, Chang-Dong Wang 0001 |
ICIC (9) | 3 |
| 2025 | Multi-View Clustering via Flexible Dual-Level FusionabstractThe anchor-based multi-view clustering is a prominent focus in massive data analysis. Although it has witnessed considerable progress, two critical limitations can still be observed in recent research. First, most existing clustering frameworks construct one anchor graph independently for each view, which may which may overlook the versatile diversity across multi-view latent distributions. In additional, they typically suffer from the two-stage issue, further restricting their flexibility in complex practical scenarios. By considering of this, we derive a new Multi-view Clustering approach via Flexible Dual-level Fusion (MC-FDF). Unlike the traditional one-anchor-graph-per-view practice, the proposed approach first produces a series of diversified anchor graphs from multiple views. Subsequently, these diversified anchor graphs are further fused from a scale-aware perspective, followed by the rotation of scale alignments to capture the consistent clustering structure across different views. With multiple views extended to the multi-scale multi-view paradigm, two levels of granularity (i.e. diversified anchor graph fusion and multi-scale late fusion) are seamlessly formulated into a mutually beneficial framework, in especially a facilitated optimization algorithm with linear time complexity is also provided. Comprehensive evaluations on heterogeneous benchmark datasets confirm that the proposed approach achieves superior performance and scalability over the state-of-the-art baselines. Jiaming Deng, Dong Huang 0001, Chang-Dong Wang 0001 |
ICPADS | 4 |
| 2025 | Dual-Level Facilitated Multi-View Contrastive Graph ClusteringabstractMulti-view attributed graph clustering (MAGC) has recently experienced impressive attention in the graph exploration literature. Although several excellent achievements have been made, previous MVAGC approaches merely consider the homogeneous information across different views, easily resulting in the compromised results when faced with the heterogeneous graph scenarios. Further, many of them rely on the static neighborhood connection from original attributed graphs, which ignores the dynamical structural relationship for enhancing cross-view contrastive learning. To deal with these drawbacks, this paper derives a Dual-level Facilitated Multi-view Contrastive Graph Clustering (DF-MCGC) approach. Specifically, we design a hybrid graph filter by considering the homogeneity hidden in individual view. Further, the view-consistent topology invariant matrix is derived to exploit the topological relationship among different embedded representations. This design progressively helps to construct the cluster-wise sample pairs for cross-view contrastive learning. Especially, the overall network is incorporated with a dual-level self-supervised paradigm, among which the self-supervised signals can be efficiently facilitated in a mutually enhanced manner. Experiments on heterogeneous benchmarks have confirmed the advantages of our DF-MCGC approach in comparison with the advanced competitors. Zi-Ying Li, Dong Huang 0001, Chang-Dong Wang 0001, Yang Liu 0084 |
ICPADS | 4 |
| 2025 | Towards multi-fusion graph neural network for single-cell RNA sequence clustering
Chen-Min Yang, Dong Huang 0001, Yuankun Xu, Xiuting He, Chang-Dong Wang 0001 |
Neurocomputing | 2 |
| 2025 | Scalable tri-factorization guided multi-view subspace clustering
Chang-Bin Guan, Dong Huang 0001, Chang-Dong Wang 0001 |
Knowl. Based Syst. | 3 |
| 2025 | Pyramid contrastive learning for clustering
Zi-Feng Zhou, Dong Huang 0001, Chang-Dong Wang 0001 |
Neural Networks | 2 |
| 2025 | Simple One-Step Multi-View Clustering With Fast Similarity and Cluster Structure LearningabstractMulti-view clustering (MVC) is essential for integrating heterogeneous data from multiple sources. However, many existing approaches are hindered by high computational complexity and the separate optimization of similarity and cluster structures. In light of these challenges, this paper presents a novel anchor-based MVC method termed simple one-step multi-view clustering with fast similarity and cluster structure learning (SONIC), which models adaptive anchor learning, multi-view similarity structure learning, and discrete cluster structure learning in a joint framework. In particular, we employ the anchor-based multi-view similarity learning to capture the consensus manifold structure latent in multiple views, thereby constructing a unified bipartite graph with adaptive anchor learning and view weighting. Then we impose a low-rank constraint on the bipartite graph structure to directly reveal the desired number of clusters without additional post-processing. An efficient alternating minimization algorithm is developed to optimize the model, resulting in a computational complexity that scales linearly with the number of samples. Extensive experiments on eight benchmark datasets demonstrate the superior performance of SONIC in both clustering quality and computational efficiency. Code available:https://github.com/huangdonghere/SONIC. Xianxian Xia, Dong Huang 0001, Chen-Min Yang, Chaobo He, Chang-Dong Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Large-Scale Tensorized Multi-View Kernel Subspace ClusteringabstractThe anchor-based multi-view subspace clustering (AMSC) has turned into a favorable tool for large-scale multi-view clustering. However, there still exist some limitations to the current AMSC approaches. First, they typically recover anchor graph structure in the original linear space, restricting their feasibility for nonlinear scenarios. Second, they usually overlook the potential benefits of jointly capturing the inter-view and intra-view information for enhancing the anchor representation learning. Third, these approaches mostly perform anchor-based subspace learning by a specific matrix norm, neglecting the latent high-order correlation across different views. To overcome these limitations, this article presents an efficient and effective approach termed Large-Scale Tensorized Multi-View Kernel Subspace Clustering (LTKMSC). Different from the existing AMSC approaches, our LTKMSC approach exploits both inter-view and intra-view awareness for anchor-based representation building. Concretely, the low-rank tensor learning is leveraged to capture the high-order correlation (i.e., the inter-view complementary information) among distinct views, upon which the \(l_{1,2}\) norm is imposed to explore the intra-view anchor graph structure in each view. Moreover, the kernel learning technique is leveraged to explore the nonlinear anchor–sample relationships embedded in multiple views. With the unified objective function formulated, an efficient optimization algorithm that enjoys low computational complexity is further designed. Extensive experiments on a variety of multi-view datasets have confirmed the efficiency and effectiveness of our approach when compared with the other competitive approaches. Dong Huang 0001, Chang-Dong Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | Contrastive Ensemble ClusteringabstractEnsemble clustering aims to combine different base clusterings into a better clustering than that of the individual one. In general, a co-association matrix depicting the pairwise affinity between different data samples is constructed by average fusion or weighted fusion of the connective matrices from multiple base clusterings. Despite the significant success, the existing works fail to capture the global structure information from multiple noisy connective matrices. Meanwhile, the locality property of the resulting representation matrix could not be explicitly preserved. In this article, we propose a novel contrastive ensemble clustering (CEC) method. Specifically, a consensus mapping model is designed for the discovery of the latent representation from the noisy observations with distinct confidences. Furthermore, a contrastive regularizer is dexterously formulated to refine the latent representation while preserving its locality property. Extensive experiments conducted on several benchmark datasets demonstrate the superiority of the proposed CEC method. To the best of our knowledge, it is the first time to explore the potential of latent representation learning and contrastive components for the ensemble clustering task. Man-Sheng Chen, Jia-Qi Lin 0001, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Knowledge-Reinforced Cross-Domain RecommendationabstractOver the past few years, cross-domain recommendation has gained great attention to resolve the cold-start issue. Many existing cross-domain recommendation methods model a preference bridge between the source and target domains to transfer preferences by the overlapping users. However, when there are insufficient cross-domain users available to bridge the two domains, it will negatively impact the recommender system's accuracy (ACC) and performance. Therefore, in this article, we propose to create a link between the source and the target domains by leveraging knowledge graph (KG) as the auxiliary information, and propose a novel knowledge-reinforced cross-domain recommendation (KR-CDR) method. First of all, we construct a new cross-domain KG (CDKG) by using the KGs that represent the source and target domains, respectively. Additionally, we employ reinforcement learning (RL) with meta learning on CDKG to discover meta-paths between the source and target domains. With these meta-paths, we obtain meta-path aggregated embedding vectors for cold-start users. Ultimately, the predicted rating can be acquired from the user meta-path aggregated embedding vector and item embedding vector. Experiments carried out on five real-world datasets show that the proposed method performs better than the state-of-the-art methods. Ling Huang 0002, Dong Huang 0001, Han Zou, Yuefang Gao, Chang-Dong Wang 0001, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | HomoMGC: Homophily-Enhanced Adaptive Graph Refinement for Multi-View Graph ClusteringabstractDue to the emergency of multi-view graph data, considerable attention is focused on the multi-view graph clustering. Although great efforts have been made in developing the multi-view graph clustering methods, most of them implicitly follow the homophily assumption, where the connected nodes with edges tend to be in the same category. As a matter of fact, such an ideal assumption is hard to be satisfied in the real-world graph data, and there are some heterogeneous edges connecting dissimilar nodes in graph. How to well consider the homophily and refine the noisy/heterogeneous edges in multi-view graph clustering still remains an under-explored challenge. Therefore, in this paper, we propose a Homophily-enhanced Adaptive Graph Refinement for Multi-view Graph Clustering (HomoMGC) method, where an adaptive graph refinement strategy is seamlessly designed. Specifically, a feature-oriented graph is constructed based on the shared feature, and an integrated graph is computed by averagely fusing all the input adjacent graphs. Then, the feature-oriented graph and integrated graph are stacked into a graph tensor with a low-rank tensor constraint, where a refined affinity probability matrix can be adaptively recovered from the integrated graph by considering multiple graph information as well as the semantics features. Extensive experiments on several benchmark datasets demonstrate the superiority of HomoMGC compared with the state-of-the-art graph clustering methods. For the code reproducibility, the source code of HomoMGC is public available at https://github.com/ManshengChen/Code-for-HomoMGc-master. Man-Sheng Chen, Xiaosha Cai, Chang-Dong Wang 0001, Dong Huang 0001, Min Chen 0003, Mohsen Guizani |
ICDM | 4 |
| 2024 | Confidence-oriented Contrastive Graph ClusteringabstractContrastive clustering has recently been an emerging topic in deep unsupervised learning. Nevertheless, the previous works mostly adopt the stochastic data augmentations, which easily leads to the semantic drift problem by limited transformations. Moreover, these approaches ignore the data distribution information when generating the positive and negative pair-wise samples. In light of this, this paper proposes a simple yet effective unsupervised clustering network termed Confidence-oriented Contrastive Graph Clustering (CoCGC). Particularly, we design an end-to-end network paradigm with un-shared weights, among which a hybrid graph filter is utilized to generate two views of reliable augmentations. Guided by the non-dominated sorting theory, we further construct a confidence-oriented sample set from the latent data distribution perspective. By considering the local density and cluster distribution of the embedding representations, the discriminative sample pairs can be derived from the confidence-oriented sets in a two-view contrastive manner. Finally, a cross-view neighbor contrastive loss is devised for better exploiting the self-supervised network signals. Extensive experimental results on five benchmark datasets demonstrate the effectiveness of our method against the existing state-of-the-art deep graph clustering methods. Yan-Di Huang, Dong Huang 0001, Chang-Dong Wang 0001, Yang Liu 0084, Enbo Huang |
IJCNN | 3 |
| 2024 | Learning clustering-friendly representations via partial information discrimination and cross-level interaction
Hai-Xin Zhang, Dong Huang 0001, Hua-Bao Ling, Weijun Sun |
Neural Networks | 2 |
| 2024 | Tensorized Incomplete Multi-view Kernel Subspace Clustering
Dong Huang 0001, Chang-Dong Wang 0001 |
Neural Networks | 2 |
| 2024 | Deep image clustering with contrastive learning and multi-scale graph convolutional networks
Yuankun Xu, Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai |
Pattern Recognit. | 2 |
| 2024 | Motorcyclist helmet detection in single images: a dual-detection framework with multi-head self-attention
Chun-Hong Li, Dong Huang 0001, Jinrong Cui |
Soft Comput. | 2 |
| 2024 | Towards Scalable Multi-View Clustering via Joint Learning of Many Bipartite GraphsabstractThis paper focuses on two limitations to previous multi-view clustering approaches. First, they frequently suffer from quadratic or cubic computational complexity, which restricts their feasibility for large-scale datasets. Second, they often rely on a single graph on each view, yet lack the ability to jointly explore many versatile graph structures for enhanced multi-view information exploration. In light of this, this paper presents a new Scalable Multi-view Clustering via Many Bipartite graphs (SMCMB) approach, which is capable of jointly learning and fusing many bipartite graphs from multiple views while maintaining high efficiency for very large-scale datasets. Different from the one-anchor-set-per-view paradigm, we first produce multiple diversified anchor sets on each view and thus obtain many anchor sets on multiple views, based on which the anchor-based subspace representation learning is enforced and many bipartite graphs are simultaneously learned. Then these bipartite graphs are efficiently partitioned to produce the base clusterings, which are further re-formulated into a unified bipartite graph for the final clustering. Note that SMCMB has almost linear time and space complexity. Extensive experiments on twenty general-scale and large-scale multi-view datasets confirm its superiority in scalability and robustness over the state-of-the-art. Jinghuan Lao, Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Big Data | 2 |
| 2024 | Deep Clustering With Hybrid-Grained Contrastive and Discriminative LearningabstractDeep contrastive clustering has recently gained significant attention due to its advantageous ability to leverage the contrastive learning paradigm for joint representation learning and clustering. However, previous deep contrastive clustering approaches mostly focus on instance discrimination or cluster discrimination, which often overlook the rich semantic information latent in the vastintermediatelevels of granularity between instances and clusters. Moreover, they are typically prone to utilizing relationships only within the same level of granularity, e.g., instance-instance relationships and cluster-cluster relationships, but frequently neglect the interactions between different granularity-levels that are ubiquitous in real-world scenarios. To tackle these issues, this paper presents a novel end-to-end deep contrastive clustering approach termed Deep Clustering with Hybrid-Grained Contrastive and Discriminative Learning (DCHL). Particularly, the instance-level contrastive learning and cluster-level contrastive learning are first formulated, where the cluster-level contrastive learning is further split into fine-grained and coarse-grained branches. To capture the global dependencies, the cluster-level contrastiveness is explored on the coarse-grained cluster branch. Meanwhile, to capture the hybrid-grained relationships, the dual-level instance-group discrimination learning is enforced between the instance branch and the fine-grained cluster branch, where the self instance-group discrimination and the cross instance-group discrimination are simultaneously optimized for enhancing the deep clustering performance. Experimental results on five challenging image datasets confirm the superiority of DCHL over state-of-the-art. Code available: https://github.com/dengxiaozhi/DCHL. Dong Huang 0001, Xiaozhi Deng, Ding-Hua Chen, Weijun Sun, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Learning Attention in the Frequency Domain for Flexible Real Photograph DenoisingabstractRecent advancements in deep learning techniques have pushed forward the frontiers of real photograph denoising. However, due to the inherent pooling operations in the spatial domain, current CNN-based denoisers are biased towards focusing on low-frequency representations, while discarding the high-frequency components. This will induce a problem for suboptimal visual quality as the image denoising tasks target completely eliminating the complex noises and recovering all fine-scale and salient information. In this work, we tackle this challenge from the frequency perspective and present a new solution pipeline, coined as frequency attention denoising network (FADNet). Our key idea is to build a learning-based frequency attention framework, where the feature correlations on a broader frequency spectrum can be fully characterized, thus enhancing the representational power of the network across multiple frequency channels. Based on this, we design a cascade of adaptive instance residual modules (AIRMs). In each AIRM, we first transform the spatial-domain features into the frequency space. Then, a learning-based frequency attention framework is devised to explore the feature inter-dependencies converted in the frequency domain. Besides this, we introduce an adaptive layer by leveraging the guidance of the estimated noise map and intermediate features to meet the challenges of model generalization in the noise discrepancy. The effectiveness of our method is demonstrated on several real camera benchmark datasets, with superior denoising performance, generalization capability, and efficiency versus the state-of-the-art. Ruijun Ma 0001, Yaoxuan Zhang, Bob Zhang 0001, Leyuan Fang, Dong Huang 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Concept Factorization Based Multiview Clustering for Large-Scale DataabstractMost existing large-scale multiview clustering algorithms attempt to capture data distribution in multiple views by selecting view-wise anchor representations beforehand with$k$-means, or by direct matrix factorization on the original observations. Despite impressive performance, few of them have paid attention to the semantic correlations between anchor bases and cluster centroids, or even the underlying relations between clusters and data samples. In view of this, we propose aConceptFactorization basedMultiviewClustering for Large-scale Data (CFMC) method with nearly linear complexity. The anchor bases learning, coefficient expression with clear semantic cues and partitioning are integrated together in this unified model. Meanwhile, explicit connections among multiview data, anchor bases and clusters are modeled via coefficient representations with semantic meanings. A four-step alternate minimizing algorithm is designed to handle the optimization problem, which is proved to have linear time complexityw.r.t.the sample size. Extensive experiments conducted on several challenging large-scale datasets confirm the superiority of the method compared with the state-of-the-art methods. Man-Sheng Chen, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Efficient Multi-View Clustering via Unified and Discrete Bipartite Graph LearningabstractAlthough previous graph-based multi-view clustering (MVC) algorithms have gained significant progress, most of them are still faced with three limitations. First, they often suffer from high computational complexity, which restricts their applications in large-scale scenarios. Second, they usually perform graph learning either at the single-view level or at the view-consensus level, but often neglect the possibility of the joint learning of single-view and consensus graphs. Third, many of them rely on the k -means for discretization of the spectral embeddings, which lack the ability to directly learn the graph with discrete cluster structure. In light of this, this article presents an efficient MVC approach via u nified and d iscrete b ipartite g raph l earning (UDBGL). Specifically, the anchor-based subspace learning is incorporated to learn the view-specific bipartite graphs from multiple views, upon which the bipartite graph fusion is leveraged to learn a view-consensus bipartite graph with adaptive weight learning. Furthermore, the Laplacian rank constraint is imposed to ensure that the fused bipartite graph has discrete cluster structures (with a specific number of connected components). By simultaneously formulating the view-specific bipartite graph learning, the view-consensus bipartite graph learning, and the discrete cluster structure learning into a unified objective function, an efficient minimization algorithm is then designed to tackle this optimization problem and directly achieve a discrete clustering solution without requiring additional partitioning, which notably has linear time complexity in data size. Experiments on a variety of multi-view datasets demonstrate the robustness and efficiency of our UDBGL approach. The code is available at https://github.com/huangdonghere/UDBGL. Si-Guo Fang, Dong Huang 0001, Xiaosha Cai, Chang-Dong Wang 0001, Chaobo He, Yong Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Multi-View Graph Learning by Joint Modeling of Consistency and InconsistencyabstractGraph learning has emerged as a promising technique for multi-view clustering due to its ability to learn a unified and robust graph from multiple views. However, existing graph learning methods mostly focus on the multi-view consistency issue, yet often neglect the inconsistency between views, which makes them vulnerable to possibly low-quality or noisy datasets. To overcome this limitation, we propose a new multi-view graph learning framework, which for the first time simultaneously and explicitly models multi-view consistency and inconsistency in a unified objective function, through which the consistent and inconsistent parts of each single-view graph as well as the unified graph that fuses the consistent parts can be iteratively learned. Though optimizing the objective function is NP-hard, we design a highly efficient optimization algorithm that can obtain an approximate solution with linear time complexity in the number of edges in the unified graph. Furthermore, our multi-view graph learning approach can be applied to both similarity graphs and dissimilarity graphs, which lead to two graph fusion-based variants in our framework. Experiments on 12 multi-view datasets have demonstrated the robustness and efficiency of the proposed approach. The code is available at https://github.com/youweiliang/Multi-view_Graph_Learning. Youwei Liang, Dong Huang 0001, Chang-Dong Wang 0001, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | On Regularizing Multiple Clusterings for Ensemble Clustering by Graph Tensor LearningabstractEnsemble clustering has shown its promising ability in fusing multiple base clusterings into a probably better and more robust clustering result. Typically, the co-association matrix based ensemble clustering methods attempt to integrate multiple connective matrices from base clusterings by weighted fusion to acquire a common graph representation. However, few of them are aware of the potential noise or corruption from the common representation by direct integration of different connective matrices with distinct cluster structures, and further consider the mutual information propagation between the input observations. In this paper, we propose a Graph Tensor Learning based Ensemble Clustering (GTLEC) method to refine multiple connective matrices by the substantial rank recovery and graph tensor learning. Within this framework, each input connective matrix is dexterously refined to approximate a graph structure by obeying the theoretical rank constraint with an adaptive weight coefficient. Further, we stack multiple refined connective matrices into a three-order tensor to extract their higher-order similarities via graph tensor learning, where the mutual information propagation across different graph matrices will also be promoted. Extensive experiments on several challenging datasets have confirmed the superiority of GTLEC compared with the state-of-the-art. Man-Sheng Chen, Jia-Qi Lin 0001, Chang-Dong Wang 0001, Wudong Xi, Dong Huang 0001 |
ACM Multimedia | 5 |
| 2023 | HOLT-Net: Detecting smokers via human-object interaction with lite transformer network
Hua-Bao Ling, Dong Huang 0001, Jinrong Cui, Chang-Dong Wang 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Facilitated low-rank multi-view subspace clustering
Dong Huang 0001, Chang-Dong Wang 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Incomplete multi-view clustering network via nonlinear manifold embedding and probability-induced loss
Jinrong Cui, Yulu Fu, Dong Huang 0001, Lusi Li |
Neural Networks | 4 |
| 2023 | Heterogeneous Tri-stream Clustering Network
Xiaozhi Deng, Dong Huang 0001, Chang-Dong Wang 0001 |
Neural Process. Lett. | 2 |
| 2023 | Deep Temporal Contrastive Clustering
Dong Huang 0001, Chang-Dong Wang 0001 |
Neural Process. Lett. | 2 |
| 2023 | Strongly augmented contrastive clustering
Xiaozhi Deng, Dong Huang 0001, Ding-Hua Chen, Chang-Dong Wang 0001, Jian-Huang Lai |
Pattern Recognit. | 2 |
| 2023 | Cross-Subject Tinnitus Diagnosis Based on Multi-Band EEG Contrastive Representation LearningabstractElectroencephalogram (EEG) is an important technology to explore the central nervous mechanism of tinnitus. However, it is hard to obtain consistent results in many previous studies for the high heterogeneity of tinnitus. In order to identify tinnitus and provide theoretical guidance for the diagnosis and treatment, we propose a robust, data-efficient multi-task learning framework called Multi-band EEG Contrastive Representation Learning (MECRL). In this study, we collect resting-state EEG data from 187 tinnitus patients and 80 healthy subjects to generate a high-quality large-scale EEG dataset on tinnitus diagnosis, and then apply the MECRL framework on the generated dataset to obtain a deep neural network model which can distinguish tinnitus patients from the healthy controls accurately. Subject-independent tinnitus diagnosis experiments are conducted and the result shows that the proposed MECRL method is significantly superior to other state-of-the-art baselines and can be well generalized to unseen topics. Meanwhile, visual experiments on key parameters of the model indicate that the high-classification weight electrodes of tinnitus' EEG signals are mainly distributed in the frontal, parietal and temporal regions. In conclusion, this study facilitates our understanding of the relationship between electrophysiology and pathophysiology changes of tinnitus and provides a new deep learning method (MECRL) to identify the neuronal biomarkers in tinnitus. Chang-Dong Wang 0001, Xi-Ran Zhu, Xueqing Zhou, Liping Lan, Dong Huang 0001, Yiqing Zheng |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Fast Multi-View Clustering Via Ensembles: Towards Scalability, Superiority, and SimplicityabstractDespite significant progress, there remain three limitations to the previous multi-view clustering algorithms. First, they often suffer from high computational complexity, restricting their feasibility for large-scale datasets. Second, they typically fuse multi-view information via one-stage fusion, neglecting the possibilities in multi-stage fusions. Third, dataset-specific hyperparameter-tuning is frequently required, further undermining their practicability. In light of this, we propose afastmulti-viewclustering viaensembles (FastMICE) approach. Particularly, the concept of random view groups is presented to capture the versatile view-wise relationships, through which the hybrid early-late fusion strategy is designed to enable efficient multi-stage fusions. Withmultipleviews extended tomanyview groups, three levels of diversity (w.r.t. features, anchors, and neighbors, respectively) are jointly leveraged for constructing the view-sharing bipartite graphs in the early-stage fusion. Then, a set of diversified base clusterings for different view groups are obtained via fast graph partitioning, which are further formulated into a unified bipartite graph for final clustering in the late-stage fusion. Notably, FastMICE has almost linear time and space complexity, and is free of dataset-specific tuning. Experiments on 22 multi-view datasets demonstrate its advantages in scalability (for extremely large datasets), superiority (in clustering performance), and simplicity (to be applied) over the state-of-the-art. Code available:https://github.com/huangdonghere/FastMICE. Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | D-TRACE: Deep Triply-Aligned Clustering
Ding-Hua Chen, Dong Huang 0001, Haiyan Cheng, Chang-Dong Wang 0001 |
ICANN (1) | 2 |
| 2022 | Efficient Orthogonal Multi-view Subspace ClusteringabstractMulti-view subspace clustering targets at clustering data lying in a union of low-dimensional subspaces. Generally, an n X n affinity graph is constructed, on which spectral clustering is then performed to achieve the final clustering. Both graph construction and graph partitioning of spectral clustering suffer from quadratic or even cubic time and space complexity, leading to difficulty in clustering large-scale datasets. Some efforts have recently been made to capture data distribution in multiple views by selecting key anchor bases beforehand with k-means or uniform sampling strategy. Nevertheless, few of them pay attention to the algebraic property of the anchors. How to learn a set of high-quality orthogonal bases in a unified framework, while maintaining its scalability for very large datasets, remains a big challenge. In view of this, we propose an Efficient Orthogonal Multi-view Subspace Clustering (OMSC) model with almost linear complexity. Specifically, the anchor learning, graph construction and partition are jointly modeled in a unified framework. With the mutual enhancement of each other, a more discriminative and flexible anchor representation and cluster indicator can be jointly obtained. An alternate minimizing strategy is developed to deal with the optimization problem, which is proved to have linear time complexity w.r.t. the sample number. Extensive experiments have been conducted to confirm the superiority of the proposed OMSC method. The source codes and data are available at https://github.com/ManshengChen/Code-for-OMSC-master. Man-Sheng Chen, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai, Philip S. Yu |
KDD | 3 |
| 2022 | Adaptively-weighted Integral Space for Fast Multiview ClusteringabstractMultiview clustering has been extensively studied to take advantage of multi-source information to improve the clustering performance. In general, most of the existing works typically compute an $n\times n$ affinity graph by some similarity/distance metrics (e.g. the Euclidean distance) or learned representations, and explore the pairwise correlations across views. But unfortunately, a quadratic or even cubic complexity is often needed, bringing about difficulty in clustering large-scale datasets. Some efforts have been made recently to capture data distribution in multiple views by selecting view-wise anchor representations with k-means, or by direct matrix factorization on the original observations. Despite the significant success, few of them have considered the view-insufficiency issue, implicitly holding the assumption that each individual view is sufficient to recover the cluster structure. Moreover, the latent integral space as well as the shared cluster structure from multiple insufficient views is not able to be simultaneously discovered. In view of this, we propose an \underlineA daptively-weighted \underlineI ntegral Space for Fast \underlineM ultiview \underlineC lustering (AIMC) with nearly linear complexity. Specifically, view generation models are designed to reconstruct the view observations from the latent integral space with diverse adaptive contributions. Meanwhile, a centroid representation with orthogonality constraint and cluster partition are seamlessly constructed to approximate the latent integral space. An alternate minimizing algorithm is developed to solve the optimization problem, which is proved to have linear time complexityw.r.t. the sample size. Extensive experiments conducted on several real-world datasets confirm the superiority of the proposed AIMC method compared with the state-of-the-art methods. Man-Sheng Chen, Tuo Liu, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
ACM Multimedia | 4 |
| 2022 | Coupled Learning for Kernel Representation and Graph Tensor in Multi-view Subspace Clustering
Man-Sheng Chen, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
PRCV (2) | 3 |
| 2022 | Kernelized multi-view subspace clustering via auto-weighted graph learning
Xiao-Wei Chen, Chang-Dong Wang 0001, Dong Huang 0001, Xiao-Yu He |
Appl. Intell. | 5 |
| 2022 | Representation Learning in Multi-view Clustering: A Literature ReviewabstractAbstract Multi-view clustering (MVC) has attracted more and more attention in the recent few years by making full use of complementary and consensus information between multiple views to cluster objects into different partitions. Although there have been two existing works for MVC survey, neither of them jointly takes the recent popular deep learning-based methods into consideration. Therefore, in this paper, we conduct a comprehensive survey of MVC from the perspective of representation learning. It covers a quantity of multi-view clustering methods including the deep learning-based models, providing a novel taxonomy of the MVC algorithms. Furthermore, the representation learning-based MVC methods can be mainly divided into two categories, i.e., shallow representation learning-based MVC and deep representation learning-based MVC, where the deep learning-based models are capable of handling more complex data structure as well as showing better expression. In the shallow category, according to the means of representation learning, we further split it into two groups, i.e., multi-view graph clustering and multi-view subspace clustering. To be more comprehensive, basic research materials of MVC are provided for readers, containing introductions of the commonly used multi-view datasets with the download link and the open source code library. In the end, some open problems are pointed out for further investigation and development. Man-Sheng Chen, Jia-Qi Lin 0001, Xiang-Long Li, Bao-Yu Liu, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai |
Data Sci. Eng. | 6 |
| 2022 | Multiview Subspace Clustering With Grouping EffectabstractMultiview subspace clustering (MVSC) is a recently emerging technique that aims to discover the underlying subspace in multiview data and thereby cluster the data based on the learned subspace. Though quite a few MVSC methods have been proposed in recent years, most of them cannot explicitly preserve the locality in the learned subspaces and also neglect the subspacewise grouping effect, which restricts their ability of multiview subspace learning. To address this, in this article, we propose a novel MVSC with grouping effect (MvSCGE) approach. Particularly, our approach simultaneously learns the multiple subspace representations for multiple views with smooth regularization, and then exploits the subspacewise grouping effect in these learned subspaces by means of a unified optimization framework. Meanwhile, the proposed approach is able to ensure the cross-view consistency and learn a consistent cluster indicator matrix for the final clustering results. Extensive experiments on several benchmark datasets have been conducted to validate the superiority of the proposed approach. Man-Sheng Chen, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001, Philip S. Yu |
IEEE Trans. Cybern. | 4 |
| 2022 | Toward Multidiversified Ensemble Clustering of High-Dimensional Data: From Subspaces to Metrics and BeyondabstractThe rapid emergence of high-dimensional data in various areas has brought new challenges to current ensemble clustering research. To deal with the curse of dimensionality, recently considerable efforts in ensemble clustering have been made by means of different subspace-based techniques. However, besides the emphasis on subspaces, rather limited attention has been paid to the potential diversity in similarity/dissimilarity metrics. It remains a surprisingly open problem in ensemble clustering how to create and aggregate a large population of diversified metrics, and furthermore, how to jointly investigate the multilevel diversity in the large populations of metrics, subspaces, and clusters in a unified framework. To tackle this problem, this article proposes a novel multidiversified ensemble clustering approach. In particular, we create a large number of diversified metrics by randomizing a scaled exponential similarity kernel, which are then coupled with random subspaces to form a large set of metric-subspace pairs. Based on the similarity matrices derived from these metric-subspace pairs, an ensemble of diversified base clusterings can be thereby constructed. Furthermore, an entropy-based criterion is utilized to explore the cluster wise diversity in ensembles, based on which three specific ensemble clustering algorithms are presented by incorporating three types of consensus functions. Extensive experiments are conducted on 30 high-dimensional datasets, including 18 cancer gene expression datasets and 12 image/speech datasets, which demonstrate the superiority of our algorithms over the state of the art. The source code is available at https://github.com/huangdonghere/MDEC. Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai, Chee Keong Kwoh 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Node Pair Information Preserving Network Embedding Based on Adversarial NetworksabstractNetwork embedding aims to learn the low-dimensional node representations for networks, which has attracted an increasing amount of attention in recent years. Most existing efforts in this field attempt to embed the network based on node similarity, which generally relies on edge existence statistics of the network. Instead of relying on the global edge existence statistics for every node pair, in this article, we utilize the information between a pair of nodes in a local way and propose a model, called node pair information preserving network embedding (NINE), based on adversarial networks. The main idea lies in preserving the node pair information (NI) by means of adversarial networks. The architecture of the proposed NINE model consists of three main components, namely: 1) NI embedder; 2) NI generator; and 3) NI discriminator. In the NI embedder, to avoid the complicated similarity calculation for a pair of nodes, the original NI vector calculated from the direct neighbor information of the two nodes is adopted as features, and the edge existence information is taken as labels to learn the embedded NI vector in a supervised learning manner. The second component is the NI generator, which takes the original node representation vectors of a node pair as input and outputs the generated NI vector. In order to make the generated NI vector follow the same distribution of the corresponding embedded NI vector, the generative adversarial network (GAN) is adopted, resulting in the third component, called the NI discriminator. Extensive experiments are conducted on seven real-world datasets in three downstream tasks, namely: 1) network reconstruction; 2) link prediction; and 3) node classification. Comparison results with seven state-of-the-art models demonstrate the effectiveness, efficiency, and rationality of our model. Chang-Dong Wang 0001, Ling Huang 0002, Kun-Yu Lin, Dong Huang 0001, Philip S. Yu |
IEEE Trans. Cybern. | 5 |
| 2021 | Consistency- and Inconsistency-Aware Multi-view Subspace Clustering
Xiao-Wei Chen, Chang-Dong Wang 0001, Dong Huang 0001 |
DASFAA (2) | 5 |
| 2021 | Link-Based Consensus Clustering with Random Walk Propagation
Xiaosha Cai, Dong Huang 0001 |
ICONIP (5) | 2 |
| 2021 | Detecting Helmets on Motorcyclists by Deep Neural Networks with a Dual-Detection Scheme
Chun-Hong Li, Dong Huang 0001 |
ICONIP (2) | 2 |
| 2021 | Single-Image Smoker Detection by Human-Object Interaction with Post-refinement
Hua-Bao Ling, Dong Huang 0001 |
ICONIP (2) | 2 |
| 2021 | Joint representation learning for multi-view subspace clustering
Chang-Dong Wang 0001, Dong Huang 0001, Xiaoyu He 0001 |
Expert Syst. Appl. | 4 |
| 2021 | Attributed Network Embedding with Micro-Meso StructureabstractRecently, network embedding has received a large amount of attention in network analysis. Although some network embedding methods have been developed from different perspectives, on one hand, most of the existing methods only focus on leveraging the plain network structure, ignoring the abundant attribute information of nodes. On the other hand, for some methods integrating the attribute information, only the lower-order proximities (e.g., microscopic proximity structure) are taken into account, which may suffer if there exists the sparsity issue and the attribute information is noisy. To overcome this problem, the attribute information and mesoscopic community structure are utilized. In this article, we propose a novel network embedding method termed Attributed Network Embedding with Micro-Meso structure, which is capable of preserving both the attribute information and the structural information including the microscopic proximity structure and mesoscopic community structure. In particular, both the microscopic proximity structure and node attributes are factorized by Nonnegative Matrix Factorization (NMF), from which the low-dimensional node representations can be obtained. For the mesoscopic community structure, a community membership strength matrix is inferred by a generative model (i.e., BigCLAM) or modularity from the linkage structure, which is then factorized by NMF to obtain the low-dimensional node representations. The three components are jointly correlated by the low-dimensional node representations, from which two objective functions (i.e., ANEM_B and ANEM_M) can be defined. Two efficient alternating optimization schemes are proposed to solve the optimization problems. Extensive experiments have been conducted to confirm the superior performance of the proposed models over the state-of-the-art network embedding methods. Juanhui Li, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001, Jian-Huang Lai, Pei Chen 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2021 | Enhanced Ensemble Clustering via Fast Propagation of Cluster-Wise SimilaritiesabstractEnsemble clustering has been a popular research topic in data mining and machine learning. Despite its significant progress in recent years, there are still two challenging issues in the current ensemble clustering research. First, most of the existing algorithms tend to investigate the ensemble information at the object-level, yet often lack the ability to explore the rich information at higher levels of granularity. Second, they mostly focus on the direct connections (e.g., direct intersection or pair-wise co-occurrence) in the multiple base clusterings, but generally neglect the multiscale indirect relationship hidden in them. To address these two issues, this paper presents a novel ensemble clustering approach based on fast propagation of cluster-wise similarities via random walks. We first construct a cluster similarity graph with the base clusters treated as graph nodes and the cluster-wise Jaccard coefficient exploited to compute the initial edge weights. Upon the constructed graph, a transition probability matrix is defined, based on which the random walk process is conducted to propagate the graph structural information. Specifically, by investigating the propagating trajectories starting from different nodes, a new cluster-wise similarity matrix can be derived by considering the trajectory relationship. Then, the newly obtained cluster-wise similarity matrix is mapped from the cluster-level to the object-level to achieve an enhanced co-association matrix, which is able to simultaneously capture the object-wise co-occurrence relationship as well as the multiscale cluster-wise relationship in ensembles. Finally, two novel consensus functions are proposed to obtain the consensus clustering result. Extensive experiments on a variety of real-world datasets have demonstrated the effectiveness and efficiency of our approach. Dong Huang 0001, Chang-Dong Wang 0001, Hongxing Peng, Jian-Huang Lai, Chee Keong Kwoh 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Multi-View Clustering in Latent Embedding SpaceabstractPrevious multi-view clustering algorithms mostly partition the multi-view data in their original feature space, the efficacy of which heavily and implicitly relies on the quality of the original feature presentation. In light of this, this paper proposes a novel approach termed Multi-view Clustering in Latent Embedding Space (MCLES), which is able to cluster the multi-view data in a learned latent embedding space while simultaneously learning the global structure and the cluster indicator matrix in a unified optimization framework. Specifically, in our framework, a latent embedding representation is firstly discovered which can effectively exploit the complementary information from different views. The global structure learning is then performed based on the learned latent embedding representation. Further, the cluster indicator matrix can be acquired directly with the learned global structure. An alternating optimization scheme is introduced to solve the optimization problem. Extensive experiments conducted on several real-world multi-view datasets have demonstrated the superiority of our approach. Man-Sheng Chen, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001 |
AAAI | 4 |
| 2020 | Subspace-Weighted Consensus Clustering for High-Dimensional Data
Xiaosha Cai, Dong Huang 0001 |
ADMA | 2 |
| 2020 | Spectral Clustering by Subspace Randomization and Graph Fusion for High-Dimensional Data
Xiaosha Cai, Dong Huang 0001, Chang-Dong Wang 0001, Chee Keong Kwoh 0001 |
PAKDD (1) | 2 |
| 2020 | One-step Kernel Multi-view Subspace Clustering
Xiaoyu He 0001, Chang-Dong Wang 0001, Dong Huang 0001 |
Knowl. Based Syst. | 5 |
| 2020 | Community Detection by Motif-Aware Label PropagationabstractCommunity detection (or graph clustering) is crucial for unraveling the structural properties of complex networks. As an important technique in community detection, label propagation has shown the advantage of finding a good community structure with nearly linear time complexity. However, despite the progress that has been made, there are still several important issues that have not been properly addressed. First, the label propagation typically proceeds over the lower order structure of the network and only the direct one-hop connections between nodes are taken into consideration. Unfortunately, the higher order structure that may encode design principle of the network and be crucial for community detection is neglected under this regime. Second, the stability of the identified community structure may also be seriously affected by the inherent randomness in the label propagation process. To tackle the above issues, this article proposes a Motif-Aware Weighted Label Propagation method for community detection. We focus on triangles within the network, but our technique extends to other kinds of motifs as well. Specifically, the motif-based higher order structure mining is conducted to capture structural characteristics of the network. First, the motif of interest (locally meaningful pattern) is identified, and then, the motif-based hypergraph can be constructed to encode the higher order connections. To further utilize the structural information of the network, a re-weighted network is designed, which unifies both the higher order structure and the original lower order structure. Accordingly, a novel voting strategy termed NaS (considering both Number and Strength of connections) is proposed to update node labels during the label propagation process. In this way, the random label selection can be effectively eliminated, yielding more stable community structures. Experimental results on multiple real-world datasets have shown the superiority of the proposed method. Pei-Zhen Li, Ling Huang 0002, Chang-Dong Wang 0001, Jian-Huang Lai, Dong Huang 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2020 | Ultra-Scalable Spectral Clustering and Ensemble ClusteringabstractThis paper focuses on scalability and robustness of spectral clustering for extremely large-scale datasets with limited resources. Two novel algorithms are proposed, namely, ultra-scalable spectral clustering (U-SPEC) and ultra-scalable ensemble clustering (U-SENC). In U-SPEC, a hybrid representative selection strategy and a fast approximation method for K-nearest representatives are proposed for the construction of a sparse affinity sub-matrix. By interpreting the sparse sub-matrix as a bipartite graph, the transfer cut is then utilized to efficiently partition the graph and obtain the clustering result. In U-SENC, multiple U-SPEC clusterers are further integrated into an ensemble clustering framework to enhance the robustness of U-SPEC while maintaining high efficiency. Based on the ensemble generation via multiple U-SEPC's, a new bipartite graph is constructed between objects and base clusters and then efficiently partitioned to achieve the consensus clustering result. It is noteworthy that both U-SPEC and U-SENC have nearly linear time and space complexity, and are capable of robustly and efficiently partitioning 10-million-level nonlinearly-separable datasets on a PC with 64 GB memory. Experiments on various large-scale datasets have demonstrated the scalability and robustness of our algorithms. The MATLAB code and experimental data are available at https://www.researchgate.net/publication/330760669. Dong Huang 0001, Chang-Dong Wang 0001, Jian-Sheng Wu, Jian-Huang Lai, Chee Keong Kwoh 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Multi-view Spectral Clustering via Multi-view Weighted Consensus and Matrix-Decomposition Based Discretization
Man-Sheng Chen, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001 |
DASFAA (1) | 4 |
| 2019 | Consistency Meets Inconsistency: A Unified Graph Learning Framework for Multi-view ClusteringabstractGraph Learning has emerged as a promising technique for multi-view clustering, and has recently attracted lots of attention due to its capability of adaptively learning a unified and probably better graph from multiple views. However, the existing multi-view graph learning methods mostly focus on the multi-view consistency, but neglect the potential multi-view inconsistency (which may be incurred by noise, corruptions, or view-specific characteristics). To address this, this paper presents a new graph learning-based multi-view clustering approach, which for the first time, to our knowledge, simultaneously and explicitly formulates the multi-view consistency and the multi-view inconsistency in a unified optimization model. To solve this model, a new alternating optimization scheme is designed, where the consistent and inconsistent parts of each single-view graph as well as the unified graph that fuses the consistent parts of all views can be iteratively learned. It is noteworthy that our multi-view graph learning model is applicable to both similarity graphs and dissimilarity graphs, leading to two graph fusion-based variants, namely, distance (dissimilarity) graph fusion and similarity graph fusion. Experiments on various multi-view datasets demonstrate the superiority of our approach. The MATLAB source code is available at https://github.com/youweiliang/ConsistentGraphLearning. Youwei Liang, Dong Huang 0001, Chang-Dong Wang 0001 |
ICDM | 2 |
| 2019 | LSCD: Low-rank and sparse cross-domain recommendation
Ling Huang 0002, Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Dong Huang 0001, Hongyang Chao |
Neurocomputing | 4 |
| 2019 | Unsupervised feature selection with multi-subspace randomization and collaboration
Dong Huang 0001, Xiaosha Cai, Chang-Dong Wang 0001 |
Knowl. Based Syst. | 1 |
| 2018 | Attributed Network Embedding with Micro-meso Structure
Juanhui Li, Chang-Dong Wang 0001, Ling Huang 0002, Dong Huang 0001, Jian-Huang Lai, Pei Chen 0001 |
DASFAA (1) | 4 |
| 2018 | Low-Rank and Sparse Cross-Domain Recommendation Algorithm
Zhi-Lin Zhao 0001, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001 |
DASFAA (1) | 4 |
| 2018 | A Generalized Predictive Framework for Data Driven Prognostics and Diagnostics using Machine LogsabstractMalfunctions in machines require equipment engineers to conduct fault diagnostic. The fault diagnostics is traditionally reliant on the skills and experiences of the equipment operators and maintenance engineers heavily, which creates an unnecessary technical barrier and results in extra cost in downtime cost and operation overhead. Meanwhile, there have been rich machine logs (sensory readings, performance logs, system logs, context data, process data) of the machine as well as the maintenance data and post-service reports. Such data provide an opportunity of leveraging intelligent data-driven technologies to reduce the maintenance cost through developing automated solutions on machine fault diagnostics and prognostic for critical component failures.In this paper, a data driven framework is proposed for machine diagnostics and prognostics to relieve the maintenance cost and increase the efficiency. It has been validated with real-world big data from complex vending machines. The proposed framework addresses the data size issue effectively by deriving applicable features and subsequently sustaining the top attributable ones only in the model. An accurate data labeling methodology is developed for supervised learning via comparing the serial number of target components in the adjacent dates. Two predictive models have been developed in this work whereby the first one is in the domain of binary classification for diagnostics, and the second one is a generalized two-stage prognostics model for multi-class classification for preventive maintenance. Cross-validated simulation results have shown that our developed diagnostics model can achieve above 80% accuracy in terms of precision, recall, and F-measure. It has also been shown that the proposed two-stage prognostics framework can outperform the conventional one-stage multiclass prediction models. Shili Xiang, Dong Huang 0001, Xiaoli Li 0001 |
TENCON | 2 |
| 2018 | TW-Co-k-means: Two-level weighted collaborative k-means for multi-view clustering
Chang-Dong Wang 0001, Dong Huang 0001, Wei-Shi Zheng 0001 |
Knowl. Based Syst. | 3 |
| 2018 | Locally Weighted Ensemble ClusteringabstractDue to its ability to combine multiple base clusterings into a probably better and more robust clustering, the ensemble clustering technique has been attracting increasing attention in recent years. Despite the significant success, one limitation to most of the existing ensemble clustering methods is that they generally treat all base clusterings equally regardless of their reliability, which makes them vulnerable to low-quality base clusterings. Although some efforts have been made to (globally) evaluate and weight the base clusterings, yet these methods tend to view each base clustering as an individual and neglect the local diversity of clusters inside the same base clustering. It remains an open problem how to evaluate the reliability of clusters and exploit the local diversity in the ensemble to enhance the consensus performance, especially, in the case when there is no access to data features or specific assumptions on data distribution. To address this, in this paper, we propose a novel ensemble clustering approach based on ensemble-driven cluster uncertainty estimation and local weighting strategy. In particular, the uncertainty of each cluster is estimated by considering the cluster labels in the entire ensemble via an entropic criterion. A novel ensemble-driven cluster validity measure is introduced, and a locally weighted co-association matrix is presented to serve as a summary for the ensemble of diverse clusters. With the local diversity in ensembles exploited, two novel consensus functions are further proposed. Extensive experiments on a variety of real-world datasets demonstrate the superiority of the proposed approach over the state-of-the-art. Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai |
IEEE Trans. Cybern. | 1 |
| 2017 | LWMC: A Locally Weighted Meta-Clustering Algorithm for Ensemble Clustering
Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai |
ICONIP (5) | 1 |
| 2017 | Community Detection in Graph Streams by Pruning Zombie Nodes
Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001 |
PAKDD (1) | 4 |
| 2017 | Multi-view collaborative locally adaptive clustering with Minkowski metric
Chang-Dong Wang 0001, Dong Huang 0001, Wei-Shi Zheng 0001 |
Expert Syst. Appl. | 3 |
| 2017 | CCMS: A nonlinear clustering method based on crowd movement and selection
Kun-Yu Lin, Chang-Dong Wang 0001, Jian-Bo Liu, Dong Huang 0001 |
Neurocomputing | 5 |
| 2016 | Ensemble-driven support vector clustering: From ensemble learning to automatic parameter estimationabstractSupport vector clustering (SVC) is a versatile clustering technique that is able to identify clusters of arbitrary shapes by exploiting the kernel trick. However, one hurdle that restricts the application of SVC lies in its sensitivity to the kernel parameter and the trade-off parameter. Although many extensions of SVC have been developed, to the best of our knowledge, there is still no algorithm that is able to effectively estimate the two crucial parameters in SVC without supervision. In this paper, we propose a novel support vector clustering approach termed ensemble-driven support vector clustering (EDSVC), which for the first time tackles the automatic parameter estimation problem for SVC based on ensemble learning, and is capable of producing robust clustering results in a purely unsupervised manner. Experimental results on multiple real-world datasets demonstrate the effectiveness of our approach. Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai, Yun Liang 0003, Shan Bian |
ICPR | 1 |
| 2016 | FTMF: Recommendation in social network with Feature Transfer and Probabilistic Matrix FactorizationabstractIt is well-known that recommendation system which is widely used in many e-commerce platforms to recommend items to the right users suffers from data sparsity, imbalanced rating and cold start problems. Matrix factorization is a good way to deal with the sparsity and imbalance problems, which is however unable to make prediction for new users due to the lack of auxiliary information. With the advent of online social networks, the trust relation in the social network can be utilized as auxiliary data to solve the aforementioned problems since consumers' buying behavior is usually affected by the people around them. This paper reports a study of exploiting the trust relationship in social network for personalized recommendation. Although previous studies have paid attention to this topic, we improve the quality of recommendation further and solve the cold start problem better. To this end, we propose a recommendation system in social network with Feature Transfer and Probabilistic Matrix Factorization (FTMF). The auxiliary data and matrix factorization technique are integrated to learn a social latent feature vector of users which represents the features transferred from trusted people. And an adaptive firm factor is introduced to balance the impact from user's own factors and trusted people on buying behavior for each user. The experimental results show that our model can effectively use the auxiliary data and outperforms the existing state-of-the-art social network based recommendation algorithms. Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Yuan-Yu Wan, Jian-Huang Lai, Dong Huang 0001 |
IJCNN | 5 |
| 2016 | Ensembling over-segmentations: From weak evidence to strong segmentation
Dong Huang 0001, Jian-Huang Lai, Chang-Dong Wang 0001, Pong C. Yuen |
Neurocomputing | 1 |
| 2016 | Ensemble clustering using factor graph
Dong Huang 0001, Jian-Huang Lai, Chang-Dong Wang 0001 |
Pattern Recognit. | 1 |
| 2016 | Robust Ensemble Clustering Using Probability TrajectoriesabstractAlthough many successful ensemble clustering approaches have been developed in recent years, there are still two limitations to most of the existing approaches. First, they mostly overlook the issue of uncertain links, which may mislead the overall consensus process. Second, they generally lack the ability to incorporate global information to refine the local links. To address these two limitations, in this paper, we propose a novel ensemble clustering approach based on sparse graph representation and probability trajectory analysis. In particular, we present the elite neighbor selection strategy to identify the uncertain links by locally adaptive thresholds and build a sparse graph with a small number of probably reliable links. We argue that a small number of probably reliable links can lead to significantly better consensus results than using all graph links regardless of their reliability. The random walk process driven by a new transition probability matrix is utilized to explore the global information in the graph. We derive a novel and dense similarity measure from the sparse graph by analyzing the probability trajectories of the random walkers, based on which two consensus functions are further proposed. Experimental results on multiple real-world datasets demonstrate the effectiveness and efficiency of our approach. Dong Huang 0001, Jian-Huang Lai, Chang-Dong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Achieving Accuracy Guarantee for Answering Batch Queries with Differential Privacy
Dong Huang 0001, Shuguo Han, Xiaoli Li 0001 |
PAKDD (2) | 1 |
| 2015 | Orthogonal mechanism for answering batch queries with differential privacyabstractDifferential privacy has recently become very promising in achieving data privacy guarantee. Typically, one can achieve ε-differential privacy by adding noise based on Laplace distribution to a query result. To reduce the noise magnitude for higher accuracy, various techniques have been proposed. They generally require high computational complexity, making them inapplicable to large-scale datasets. In this paper, we propose a novel orthogonal mechanism (OM) to represent a query set Q with a linear combination of a new query set Q, where Q consists of orthogonal query sets and is derived by exploiting the correlations between queries in Q. As a result of orthogonality of the derived queries, the proposed technique not only greatly reduces computational complexity, but also achieves better accuracy than the existing mechanisms. Extensive experimental results demonstrate the effectiveness and efficiency of the proposed technique. Dong Huang 0001, Shuguo Han, Xiaoli Li 0001, Philip S. Yu |
SSDBM | 1 |
| 2015 | Combining multiple clusterings via crowd agreement estimation and multi-granularity link analysis
Dong Huang 0001, Jian-Huang Lai, Chang-Dong Wang 0001 |
Neurocomputing | 1 |
| 2013 | SVStream: A Support Vector-Based Algorithm for Clustering Data StreamsabstractIn this paper, we propose a novel data stream clustering algorithm, termed SVStream, which is based on support vector domain description and support vector clustering. In the proposed algorithm, the data elements of a stream are mapped into a kernel space, and the support vectors are used as the summary information of the historical elements to construct cluster boundaries of arbitrary shape. To adapt to both dramatic and gradual changes, multiple spheres are dynamically maintained, each describing the corresponding data domain presented in the data stream. By allowing for bounded support vectors (BSVs), the proposed SVStream algorithm is capable of identifying overlapping clusters. A BSV decaying mechanism is designed to automatically detect and remove outliers (noise). We perform experiments over synthetic and real data streams, with the overlapping, evolving, and noise situations taken into consideration. Comparison results with state-of-the-art data stream clustering methods demonstrate the effectiveness and efficiency of the proposed method. Chang-Dong Wang 0001, Jian-Huang Lai, Dong Huang 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Incremental support vector clustering with outlier detection
Dong Huang 0001, Jian-Huang Lai, Chang-Dong Wang 0001 |
ICPR | 1 |