VLDB 2026 Research / reviewers in the wild / expert
Hongjun Wang 0002
dblp:65/3627-2
· DBLP profile ↗
53ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0001-7280-2852ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 19 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-view Semantic Contrastive Alignment for Multimodal RecommendationabstractMultimodal recommendation advocates integrating the multimodal features of items with historical user behaviors to enhance recommendation accuracy across various online media platforms. The majority of existing methods concentrate on leveraging cross-modal learning over multimodal features to augment node representations. However, these approaches are confronted with two key challenges: i) augmented representations offer limited information gain for interactive prediction in the collaborative view, and ii) semantic discrepancy between the collaborative view and modality-augmented features remains inadequately addressed. To overcome these obstacles, we present a new Multi-view Semantic Contrastive Alignment (MSCA) approach for multimodal recommendation, which models and aligns node representations from multiple views. Specifically, we introduce a multi-view semantic pattern encoder that learns basic embeddings from the collaborative view and independently captures augmented semantic patterns from the item-item structural view and intra-modal view. Furthermore, a semantic contrastive alignment task is designed to mitigate the semantic divergence between collaborative embeddings and augmented representations by maximizing the mutual consistency between them, thereby facilitating an effective integration of both. Comprehensive experiments on three benchmark datasets confirm that the proposed MSCA consistently excels over diverse state-of-the-art baselines. Jiuqiang Li 0002, Hongjun Wang 0002 |
WWW | 2 |
| 2026 | Manifold-aware dual hypergraph co-clustering
Hongjun Wang 0002, Luqing Wang, Tianrui Li 0001 |
Neurocomputing | 2 |
| 2026 | General Adaptive Hypergraph Model for CoclusteringabstractCoclustering, as a crucial technique in data analysis, has proven effective in exploring the inherent data structures and revealing the synergistic effects of samples and features. However, most existing coclustering methods are limited to graph structures that primarily capture pairwise relationships between data points. In contrast, hyperedges in hypergraphs can connect multiple nodes, capturing complex associations between sets of points. Compared with the unidirectional clustering of hyperedges and nodes, bidirectional clustering of samples and features on hypergraphs presents a greater challenge. Moreover, unidirectional clustering overlooks the mutual dependencies between rows and columns, making it difficult to optimize clustering performance. Therefore, this article proposes a novel general and adaptive hypergraph model for coclustering (GAHGC), aiming to cocluster samples and features. It not only captures the similarity between samples but also considers the correlation information between samples and features. First, a general hypergraph construction strategy is proposed to overcome the limitation of existing methods that are often constrained by specific hypergraph data types. Next, an adaptive hypergraph optimization mechanism is designed to enhance coclustering performance. Finally, extensive comparative experiments on benchmark datasets demonstrate the effectiveness and superiority of the proposed model, further revealing the potential of hypergraph coclustering in capturing complex multilevel interactions. Xueyao Wang 0002, Hongjun Wang 0002, Tianrui Li 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | AGCI2L: Adversarial Graph Contrastive Information Invariant Learning
Zhicheng Gao, Hongjun Wang 0002, Tianrui Li 0001 |
Knowl. Based Syst. | 3 |
| 2025 | Weakly-supervised locally linear embedding model for discriminant feature learning
Luqing Wang, Chengsu Wang, Hongjun Wang 0002, Chongshou Li, Jie Hu 0007, Tianrui Li 0001 |
Knowl. Based Syst. | 3 |
| 2025 | Multidimensional Scaling Orienting Discriminative Co-Representation LearningabstractCo-representation, which co-represents samples and features, has been widely used in various machine learning tasks, such as document clustering, gene expression analysis, and recommendation systems. It not only reveals the cluster structure of both samples and features, but also reveals the sample–feature correlation. Given a tabular data matrix, co-representation usually exhibits as the co-occurrence structures of rows and columns. However, identifying such structured patterns in complex real-world data can be very challenging. To address this problem, we propose an unsupervised discriminative co-representation learning model based on multidimensional scaling (DCLMDS). The main novelty is that DCLMDS introduces a co-representation learning term to ensure the discriminability between co-occurrence structures. As a result, the co-representation learned by DCLMDS contains richer information of the underlying correlation between samples and features within data. This could subsequently enhance the capacity of machines and systems for processing complex real-world information more proficiently. Furthermore, inspired by the fuzzy set theory, we integrate fuzzy membership degree that can accurately capture the uncertainty within data, thus enabling DCLMDS to learn a more effective co-representation in a soft manner. To evaluate the performance of DCLMDS, we conduct extensive experiments on 18 datasets, and the results demonstrate that DCLMDS can generate both accurate and discriminative co-representation, which well meets our desired outcomes. Zhang Qin, Yinghui Zhang 0005, Hongjun Wang 0002, Chongshou Li, Tianrui Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2024 | Graph Diffusive Self-Supervised Learning for Social RecommendationabstractSocial recommendation aims at augmenting user-item interaction relationships and boosting recommendation quality by leveraging social information. Recently, self-supervised learning (SSL) has gained widespread adoption for social recommender. However, most existing methods exhibit poor robustness when faced with sparse user behavior data and are susceptible to inevitable social noise. To overcome the aforementioned limitations, we introduce a new Graph Diffusive Self-Supervised Learning (GDSSL) paradigm for social recommendation. Our approach involves the introduction of a guided social graph diffusion model that can adaptively mitigate the impact of social relation noise commonly found in real-world scenarios. This model progressively introduces random noise to the initial social graph and then iteratively restores it to recover the original structure. Additionally, to enhance robustness against noise and sparsity, we propose graph diffusive self-supervised learning, which utilizes the denoised social relation graph generated by our diffusion model for contrastive learning. The extensive experimental outcomes consistently indicate that our proposed GDSSL outmatches existing advanced solutions in social recommendation. Jiuqiang Li 0002, Hongjun Wang 0002 |
SIGIR | 2 |
| 2024 | A nondominated sorting genetic model for co-clustering
Wuchun Yang, Hongjun Wang 0002, Yinghui Zhang 0005, Tianrui Li 0001 |
Inf. Sci. | 2 |
| 2024 | Information bottleneck fusion for deep multi-view clustering
Jie Hu 0007, Hongjun Wang 0002, Bo Peng 0006, Tianrui Li 0001 |
Knowl. Based Syst. | 4 |
| 2024 | T-Distributed Stochastic Neighbor Embedding for Co-Representation LearningabstractCo-clustering is the simultaneous clustering of the samples and attributes of a data matrix that provides deeper insight into data than traditional clustering. However, there is a lack of representation learning algorithms that serve this mechanism of co-clustering, and the current representation learning algorithms are limited to the sample perspective and lack the use of information in the attribute perspective. To solve this problem, in this article, ctSNE , a co-representation learning model based on t-distributed stochastic neighbor embedding, is proposed for unsupervised co-clustering, where ctSNE makes the dataset representation outputted more discriminative of row and column clusters (i.e. co-discrimination). On the basis of t-distributed stochastic neighbor embedding retaining the sample data distribution and local data structure, the philosophy of collaboration is introduced (i.e., row and column hidden relationship information) so that the ctSNE model is equipped with co-representation learning capability, which can effectively improve the performance of co-clustering. To prove the effectiveness of the ctSNE model, several classic co-clustering algorithms are used to check the co-representation performance of ctSNE, and a novel internal index based on an internal clustering index, known as total inertia, is proposed to demonstrate the effect of co-clustering. The numerous experimental results show that ctSNE has tremendous co-representation capability and can significantly improve the performance of co-clustering algorithms. Wei Chen 0141, Hongjun Wang 0002, Yinghui Zhang 0005, Ping Deng 0002, Tianrui Li 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2024 | A Survey of Co-ClusteringabstractCo-clustering is to cluster samples and features simultaneously, which can also reveal the relationship between row clusters and column clusters. Therefore, lots of scientists have drawn much attention to conduct extensive research on it, and co-clustering is widely used in recommendation systems, gene analysis, medical data analysis, natural language processing, image analysis, and social network analysis. In this article, we survey the entire research aspect of co-clustering, especially the latest advances in co-clustering, and discover the current research challenges and future directions. First, due to different views from researchers on the definition of co-clustering, this article summarizes the definition of co-clustering and its extended definitions, as well as related issues, based on the perspectives of various scientists. Second, existing co-clustering techniques are approximately categorized into four classes: information-theory-based, graph-theory-based, matrix-factorization-based, and other theories-based. Third, co-clustering is applied in various aspects such as recommendation systems, medical data analysis, natural language processing, image analysis, and social network analysis. Furthermore, 10 popular co-clustering algorithms are empirically studied on 10 benchmark datasets with 4 metrics—accuracy, purity, block discriminant index, and running time, and their results are objectively reported. Finally, future work is provided to get insights into the research challenges of co-clustering. Hongjun Wang 0002, Wei Chen 0141, Chongshou Li, Tianrui Li 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2024 | Hierarchical Active Learning With Label Proportions on Data RegionsabstractLearning classification models from real-world data often requires substantial human effort devoted to instance annotation. As the instance-based annotating process can be very time-consuming and costly, we propose a novel active learning framework that builds classification models from human-annotatedregions. A region is defined by a set of conjunctive patterns that are formed by value ranges over the input features. A region label is a human assessment of the classproportionin the data population covered by the region. By leveraginglearning from label proportionsalgorithms, regions and their class proportions can be used to train instance-based classification models. However, the key challenge is that in practice, very few regions are defined already. Therefore, to identify regions important for model learning, we design ahierarchical active learning(HAL) framework, which actively builds a hierarchy of regions. Similar to the decision-tree learning process, our approach progressively divides the input data space into smaller sub-regions, solicits labels for the new regions, and retrains the base classification model with all the leaf regions. And we further develop amulti-hierarchy(forest) solution, which builds multiple shallower hierarchies that have more informative, diverse, and simpler regions. We evaluate our HAL framework on numerous impactful classification datasets as well as on a real user study - on the survival analysis of colorectal cancer patients. The results demonstrate that region-based active learning methods can learn high-quality classifiers from very few labeled regions. Hence, our framework is shown very effective in reducing the human annotation effort needed for building classification models. Qiang Gao 0003, Yazhou He, Hongjun Wang 0002, Milos Hauskrecht, Tianrui Li 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Unsupervised Feature Learning Architecture with Multi-clustering Integration RBMabstractFeature learning is a crucial phase machine learning [1] – [3] . How to obtain appropriate features distribution without any background is still a hard problem in machine learning. In this paper, we present a novel unsupervised feature learning architecture (see Fig. 1 ), which consists of a multi-clustering integration module and a variant of RBM termed multi-clustering integration RBM (MIRBM). In the multi-clustering integration module, we choose three clusterers to obtain three different global clustering partitions (CPs). Then, an unanimous voting strategy is used to generate the local clustering partition (LCP) of visible layer data. Hence, the LCP only has partial visible layer data. The novel MIRBM model is a core feature encoding part of the proposed unsupervised feature learning architecture. The novelty of it is that the LCP as an unsupervised guidance is integrated into the CD 1 learning to guide the distribution of the hidden layer features. For the instance in the same LCP cluster, the hidden and reconstructed hidden layer features of the MIRBM model in the proposed architecture tend to constrict together in the training process. Meanwhile, each LCP center tends to disperse from each other as much as possible in the hidden and reconstructed hidden layer during training. This work has three main contributions: 1) A novel unsupervised feature learning architecture is proposed, which consists of a multi-clustering integration module and an MIRBM model. 2) In the multi-clustering integration module of the proposed architecture, three unsupervised algorithms are employed to obtain three different global CPs without any background knowledge or label. 3) The MIRBM model in the proposed architecture uses the LCP as an unsupervised guidance to guide the distribution of the hidden layer features by integrating the LCP into the CD 1 learning. Jielei Chu, Hongjun Wang 0002, Zhiguo Gong, Tianrui Li 0001 |
ICDE | 2 |
| 2023 | Multi-view clustering guided by unconstrained non-negative matrix factorization
Ping Deng 0002, Tianrui Li 0001, Dexian Wang 0001, Hongjun Wang 0002, Shi-Jinn Horng |
Knowl. Based Syst. | 4 |
| 2023 | Micro-Supervised Disturbance Learning: A Perspective of Representation Probability DistributionabstractThe instability is shown in the existing methods of representation learning based on Euclidean distance under a broad set of conditions. Furthermore, the scarcity and high cost of labels prompt us to explore more expressive representation learning methods which depends on as few labels as possible. To address above issues, the small-perturbation ideology is firstly introduced on the representation learning model based on the representation probability distribution. The positive small-perturbation information (SPI) which only depend on two labels of each cluster is used to stimulate the representation probability distribution and then two variant models are proposed to fine-tune the expected representation distribution of Restricted Boltzmann Machine (RBM), namely, Micro-supervised Disturbance Gaussian-binary RBM (Micro-DGRBM) and Micro-supervised Disturbance RBM (Micro-DRBM) models. The Kullback-Leibler (KL) divergence of SPI is minimized in the same cluster to promote the representation probability distributions to become more similar in Contrastive Divergence (CD) learning. In contrast, the KL divergence of SPI is maximized in the different clusters to enforce the representation probability distributions to become more dissimilar in CD learning. To explore the representation learning capability under the continuous stimulation of the SPI, we present a deep Micro-supervised Disturbance Learning (Micro-DL) framework based on the Micro-DGRBM and Micro-DRBM models and compare it with a similar deep structure which has no external stimulation. Experimental results demonstrate that the proposed deep Micro-DL architecture shows better performance in comparison to the baseline method, the most related shallow models and deep frameworks for clustering. Jielei Chu, Hongjun Wang 0002, Hua Meng 0001, Zhiguo Gong, Tianrui Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Auto-attention mechanism for multi-view deep embedding clustering
Bassoma Diallo, Jie Hu 0007, Tianrui Li 0001, Ghufran Ahmad Khan, Xinyan Liang, Hongjun Wang 0002 |
Pattern Recognit. | 6 |
| 2023 | Graph Regularized Sparse Non-Negative Matrix Factorization for ClusteringabstractThe graph regularized nonnegative matrix factorization (GNMF) algorithms have received a lot of attention in the field of machine learning and data mining, as well as the square loss method is commonly used to measure the quality of reconstructed data. However, noise is introduced when data reconstruction is performed; and the square loss method is sensitive to noise, which leads to degradation in the performance of data analysis tasks. To solve this problem, a novel graph regularized sparse NMF (GSNMF) is proposed in this article. To obtain a cleaner data matrix to approximate the high-dimensional matrix, the$l_{1}$-norm to the low-dimensional matrix is added to achieve the adjustment of data eigenvalues in the matrix and sparsity constraint. In addition, the corresponding inference and alternating iterative update algorithm to solve the optimization problem are given. Then, an extension of GSNMF, namely, graph regularized sparse nonnegative matrix trifactorization (GSNMTF), is proposed, and the detailed inference procedure is also shown. Finally, the experimental results on eight different datasets demonstrate that the proposed model has a good performance. Ping Deng 0002, Tianrui Li 0001, Hongjun Wang 0002, Dexian Wang 0001, Shi-Jinn Horng |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Hierarchical Active Learning With Qualitative Feedback on RegionsabstractLearning classification models in practice usually requires numerous labeled data for training. However, instance-based annotation can be inefficient for humans to perform. In this article, we propose and study a new type of human supervision that is fast to perform and useful for model learning. Instead of labeling individual instances, humans provide supervision to dataregions, which are subspaces of the input data space, representing subpopulations of data. Since labeling now is performed on a region level, 0/1 labeling becomes imprecise. Thus, we design the region label to be aqualitativeassessment of the class proportion, which coarsely preserves the labeling precision but is also easy for humans to do. To identify informative regions for labeling and learning, we further devise ahierarchical active learningprocess that recursively constructs a region hierarchy. This process is semisupervised in the sense that it is driven by both active learning strategies and human expertise, where humans can provide discriminative features. To evaluate our framework, we conducted extensive experiments on nine datasets as well as a real user study on a survival analysis of colorectal cancer patients. The results have clearly demonstrated the superiority of our region-based active learning framework against many instance-based active learning methods. Yazhou He, Yanbing Xue, Hongjun Wang 0002, Milos Hauskrecht, Tianrui Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2023 | Self-supervised Discriminative Representation Learning by Fuzzy AutoencoderabstractRepresentation learning based on autoencoders has received great concern for its potential ability to capture valuable latent information. Conventional autoencoders pursue minimal reconstruction error, but in most machine learning tasks such as classification and clustering, the discrimination of feature representation is also important. To address this limitation, an enhanced self-supervised discriminative fuzzy autoencoder (FAE) is innovatively proposed, which focuses on exploring information within data to guide the unsupervised training process and enhancing feature discrimination in a self-supervised manner. In FAE, fuzzy membership is applied to provide a means of self-supervised, which allows FAE can not only utilize AE’s outstanding representation learning capabilities but can also transform the original data into another space with improved discrimination. First, the objective function corresponding to FAE is proposed by reconstruction loss and clustering oriented loss simultaneously. Subsequently, Mini-Batch Gradient Descent is applied to infer the objective function and the detailed process is illustrated step by step. Finally, empirical studies on clustering tasks have demonstrated the superiority of FAE over the state of the art. Wenlu Yang, Hongjun Wang 0002, Yinghui Zhang 0005, Zehao Liu 0003, Tianrui Li 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2023 | Fast Flexible Bipartite Graph Model for Co-ClusteringabstractCo-clustering methods make use of the correlation between samples and attributes to explore the co-occurrence structure in data. These methods have played a significant role in gene expression analysis, image segmentation, and document clustering. In bipartite graph partition-based co-clustering methods, the relationship between samples and attributes is described by constructing a diagonal symmetric bipartite graph matrix, which is clustered by the philosophy of spectral clustering. However, this not only has high time complexity but also the same number of row and column clusters. In fact, the number of categories of rows and columns often changes in the real world. To address these problems, this paper proposes a novel fast flexible bipartite graph model for the co-clustering method (FBGPC) that directly uses the original matrix to construct the bipartite graph. Then, it uses the inflation operation to partition the bipartite graph in order to learn the co-occurrence structure of the original data matrix based on the inherent relationship between bipartite graph partitioning and co-clustering. Finally, hierarchical clustering is used to obtain the clustering results according to the set relationship of the co-occurrence structure. Extensive empirical results show the effectiveness of our proposed model and verify the faster performance, generality, and flexibility of our model. Wei Chen 0141, Hongjun Wang 0002, Zhiguo Long, Tianrui Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | The detection of distributional discrepancy for language GANsabstractA pre-trained neural language model (LM) is usually used to generate texts. Due to exposure bias, the generated text is not as good as real text. Many researchers claimed they employed the Generative Adversarial Nets (GAN) to alleviate this issue by feeding reward signals from a discriminator to update the LM (generator). However, some researchers argued that GAN did not work by evaluating the generated texts with a quality-diversity metric such as Bleu versus self-Bleu, and language model score versus reverse language model score. Unfortunately, these two-dimension metrics are not reliable. Furthermore, the existing methods only assessed the final generated texts, thus neglecting the dynamic evaluating the adversarial learning process. Different from the above-mentioned methods, we adopted the most recent metric functions, which measure the distributional discrepancy between real and generated text. Besides that, we design a comprehensive experiment to investigate the performance during the learning process. First, we evaluate a language model with two functions and identify a large discrepancy. Then, several methods with the detected discrepancy signal to improve the generator were tried. Experimenting with two language GANs on two benchmark datasets, we found that the distributional discrepancy increases with more adversarial learning rounds. Our research provides convicted evidence that the language GANs fail. Ping Cai, Hongjun Wang 0002, Xinyu Dai, Jiajun Chen 0001 |
Connect. Sci. | 4 |
| 2022 | Dual graph-regularized sparse concept factorization for clustering
Dexian Wang 0001, Tianrui Li 0001, Ping Deng 0002, Hongjun Wang 0002, Pengfei Zhang 0016 |
Inf. Sci. | 4 |
| 2022 | Multi-local Collaborative AutoEncoder
Jielei Chu, Hongjun Wang 0002, Zhiguo Gong, Tianrui Li 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Gaussian gravitation for cluster ensembles
Kai Cong, Hongjun Wang 0002 |
Knowl. Based Syst. | 3 |
| 2022 | Biased unconstrained non-negative matrix factorization for clustering
Ping Deng 0002, Fan Zhang 0108, Tianrui Li 0001, Hongjun Wang 0002, Shi-Jinn Horng |
Knowl. Based Syst. | 4 |
| 2022 | Bilateral discriminative autoencoder model orienting co-representation learning
Zehao Liu 0003, Hongjun Wang 0002, Wei Chen 0141, Luqing Wang, Tianrui Li 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Markov clustering ensemble
Luqing Wang, Hongjun Wang 0002, Tianrui Li 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Unsupervised Feature Learning Architecture With Multi-Clustering Integration RBMabstractIn this paper, we present a novel unsupervised feature learning architecture, which consists of a multi-clustering integration module and a variant of RBM termed multi-clustering integration RBM (MIRBM). In the multi-clustering integration module, we apply three clusterers (K-means, affinity propagation and spectral clustering algorithms) to obtain three different clustering partitions (CPs) without any background knowledge or label. Then, an unanimous voting strategy is used to generate a local clustering partition (LCP). The novel MIRBM model is a core feature encoding part of the proposed unsupervised feature learning architecture. The novelty of it is that the LCP as an unsupervised guidance is integrated into one step contrastive divergence (${\mathtt{{CD}}}_{1}$) learning to guide the distribution of the hidden layer features. For the instance in the same LCP cluster, the hidden and reconstructed hidden layer features of the MIRBM model in the proposed architecture tend to constrict together in the training process. Meanwhile, each LCP center tends to disperse from each other as much as possible in the hidden and reconstructed hidden layer during training. The experiments demonstrate that the proposed unsupervised feature learning architecture has more powerful feature representation and generalization capability than the state-of-the-art models for clustering tasks in the Microsoft Research Asia Multimedia (MSRA-MM)2.0 dataset. Jielei Chu, Hongjun Wang 0002, Zhiguo Gong, Tianrui Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Enhanced clustering embedded in curvilinear distance analysis guided by pairwise constraints
Yinghui Zhang 0005, Hongjun Wang 0002, Ping Deng 0002, Tianrui Li 0001 |
Inf. Sci. | 3 |
| 2021 | Distributional discrepancy: A metric for unconditional text generation
Ping Cai, Hongjun Wang 0002, Tianrui Li 0001 |
Knowl. Based Syst. | 4 |
| 2021 | Tri-regularized nonnegative matrix tri-factorization for co-clustering
Ping Deng 0002, Tianrui Li 0001, Hongjun Wang 0002, Shi-Jinn Horng, Zeng Yu 0001 |
Knowl. Based Syst. | 3 |
| 2021 | Hybrid genetic model for clustering ensemble
Wenlu Yang, Yinghui Zhang 0005, Hongjun Wang 0002, Ping Deng 0002, Tianrui Li 0001 |
Knowl. Based Syst. | 3 |
| 2019 | Linear discriminant analysis guided by unsupervised ensemble learning
Ping Deng 0002, Hongjun Wang 0002, Tianrui Li 0001, Shi-Jinn Horng, Xinwen Zhu |
Inf. Sci. | 2 |
| 2019 | A factor graph model for unsupervised feature selection
Hongjun Wang 0002, Yinghui Zhang 0005, Ji Zhang 0012, Tianrui Li 0001, Lingxi Peng |
Inf. Sci. | 1 |
| 2019 | Nonnegative matrix factorization for clustering ensemble based on dark knowledge
Wenting Ye, Hongjun Wang 0002, Shan Yan, Tianrui Li 0001, Yan Yang 0001 |
Knowl. Based Syst. | 2 |
| 2019 | Improved Gaussian-Bernoulli restricted Boltzmann machine for learning discriminative representations
Ji Zhang 0012, Hongjun Wang 0002, Jielei Chu, Shudong Huang, Tianrui Li 0001, Qigang Zhao |
Knowl. Based Syst. | 2 |
| 2019 | Particle Subswarms Collaborative ClusteringabstractCollaborative clustering aims to find a common data structure between several distributed data sets governed by different privacy constraints and technical limitations that prohibit a central collection of data for processing. Therefore, it is required to process the data sets separately using collaboration, which allows clustering algorithms to work locally on an individual data set while exchanging information about the finding with algorithms in other data locations. Thus, the different data locations share information to improve individual clustering result amidst technical and privacy limitations but without breaching privacy. In this article, we present a framework of collaborative clustering that does not require interaction coefficients to regulate the effect of collaboration. We further adapt the framework to cluster distributed data using crisp and fuzzy clustering algorithms. We use particle swarm optimization techniques to inference the framework and, therefore, call it particle subswarms. Moreover, the collaboration increases the number of particles in the swarm without increasing the number of clusters in the data set. This article, therefore, provides the theoretical foundations of particle subswarms and some experimental results on several data sets. Collins Census, Hongjun Wang 0002, Ji Zhang 0012, Ping Deng 0002, Tianrui Li 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2019 | Restricted Boltzmann Machines With Gaussian Visible Units Guided by Pairwise ConstraintsabstractRestricted Boltzmann machines (RBMs) and their variants are usually trained by contrastive divergence (CD) learning, but the training procedure is an unsupervised learning approach, without any guidances of the background knowledge. To enhance the expression ability of traditional RBMs, in this paper, we propose pairwise constraints (PCs) RBM with Gaussian visible units (pcGRBM) model, in which the learning procedure is guided by PCs and the process of encoding is conducted under these guidances. The PCs are encoded in hidden layer features of pcGRBM. Then, some pairwise hidden features of pcGRBM flock together and another part of them are separated by the guidances. In order to deal with real-valued data, the binary visible units are replaced by linear units with Gaussian noise in the pcGRBM model. In the learning process of pcGRBM, the PCs are iterated transitions between visible and hidden units during CD learning procedure. Then, the proposed model is inferred by approximative gradient descent method and the corresponding learning algorithm is designed. In order to compare the availability of pcGRBM and traditional RBMs with Gaussian visible units, the features of the pcGRBM and RBMs hidden layer are used as input "data" for K -means, spectral clustering (SP) and affinity propagation (AP) algorithms, respectively. We also use tenfold cross-validation strategy to train and test pcGRBM model to obtain more meaningful results with PCs which are derived from incremental sampling procedures. A thorough experimental evaluation is performed with 12 image datasets of Microsoft Research Asia Multimedia. The experimental results show that the clustering performance of K -means, SP, and AP algorithms based on pcGRBM model are significantly better than traditional RBMs. In addition, the pcGRBM model for clustering tasks shows better performance than some semi-supervised clustering algorithms. Jielei Chu, Hongjun Wang 0002, Hua Meng 0001, Tianrui Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2018 | Robust graph regularized nonnegative matrix factorization for clustering
Shudong Huang, Hongjun Wang 0002, Tao Li 0001, Tianrui Li 0001, Zenglin Xu |
Data Min. Knowl. Discov. | 2 |
| 2018 | Parallel Semi-Supervised Multi-Ant Colonies Clustering Ensemble Based on MapReduce MethodologyabstractSemi-supervised clustering ensemble has emerged as an important elaboration of classical clustering problem that improves quality and robustness in clustering by combining the results of different clustering components with user provided constraints. MapReduce is a parallel programming model for processing big data using large numbers of distributed computers (nodes). In this paper, we propose a novel semi-supervised multi-ant colonies consensus clustering algorithm and implement the parallelization of this algorithm using MapReduce on Hadoop platform. Our method incorporates pairwise constraints not only in each ant colony clustering process, but also in computing new similarity matrix during the process of the multi-ant colonies ensemble. In addition, it enhances the computational efficiency for big data by adopting a MapReduce Framework. Experimental results demonstrate the effectiveness of the proposed method. Yan Yang 0001, Fei Teng 0001, Tianrui Li 0001, Hao Wang 0068, Hongjun Wang 0002 |
IEEE Trans. Cloud Comput. | 5 |
| 2016 | Semi-supervised hierarchical clustering ensemble and its application
Wenchao Xiao, Yan Yang 0001, Hongjun Wang 0002, Tianrui Li 0001, Huanlai Xing |
Neurocomputing | 3 |
| 2016 | Hierarchical cluster ensemble model based on knowledge granulation
Jie Hu 0007, Tianrui Li 0001, Hongjun Wang 0002, Hamido Fujita |
Knowl. Based Syst. | 3 |
| 2016 | Constraint Co-Projections for Semi-Supervised Co-ClusteringabstractCo-clustering aims to simultaneously cluster the objects and features to explore intercorrelated patterns. However, it is usually difficult to obtain good co-clustering results by just analyzing the object-feature correlation data due to the sparsity of the data and the noise. Meanwhile, most co-clustering algorithms cannot take the prior information into consideration and may produce unmeaningful results. Semi-supervised co-clustering aims to incorporate the known prior knowledge into the co-clustering algorithm. In this paper, a new technique named constraint co-projections for semi-supervised co-clustering (CPSSCC) is presented. Constraint co-projections can not only make use of two popular techniques including pairwise constraints and constraint projections, but also simultaneously perform the object constraint projections and feature constraint projections. The two popular techniques are illustrated for semi-supervised co-clustering when some objects and features are believed to be in the same cluster a priori. Furthermore, we also prove that the co-clustering problem can be formulated as a typical eigen-problem and can be efficiently solved with the selected eigenvectors. To the best of our knowledge, constraint co-projections is first stated in this paper and this is the first work on using CPSSCC. Extensive experiments on benchmark data sets demonstrate the effectiveness of the proposed method. This paper also shows that CPSSCC has some favorable features compared with previous related co-clustering algorithms. Shudong Huang, Hongjun Wang 0002, Tao Li 0001, Yan Yang 0001, Tianrui Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Belief Revision over Infinite Propositional LanguageabstractThere are different models to characterize AGM belief revision framework. When the background language is finite propositional language, Katsuno and Mendelzon (KM) proposed in 1991 a representation model using total preorder on worlds. This ‘preorder’ model is very influential and has been extended to characterize epistemic state in iterated belief revision. KM showed an approach how to construct the preorder via a belief set and an AGM belif revision operator, however, this approach does not work well when the language is infinite. In this paper, we argue when the language is infinite propositional language, how to construct a preorder on world to model AGM belief revision framework, and then we generalize the representation theorem of KM over an infinite language. Hua Meng 0001, Yayan Yuan, Jielei Chu, Hongjun Wang 0002 |
KSEM | 4 |
| 2015 | Spectral co-clustering ensemble
Shudong Huang, Hongjun Wang 0002, Dingcheng Li, Yan Yang 0001, Tianrui Li 0001 |
Knowl. Based Syst. | 2 |
| 2015 | Semi-supervised evolutionary ensembles for Web video categorization
Amjad Mahmood, Tianrui Li 0001, Yan Yang 0001, Hongjun Wang 0002, Mehtab Afzal |
Knowl. Based Syst. | 4 |
| 2014 | Hyper-ellipsoidal clustering technique for evolving data stream
Muhammad Zia-ur Rehman 0003, Tianrui Li 0001, Yan Yang 0001, Hongjun Wang 0002 |
Knowl. Based Syst. | 4 |
| 2014 | Bayesian image segmentation fusion
Hongjun Wang 0002, Yinghui Zhang 0005, Ruihua Nie, Yan Yang 0001, Bo Peng 0006, Tianrui Li 0001 |
Knowl. Based Syst. | 1 |
| 2014 | Constraint Neighborhood Projections for Semi-Supervised ClusteringabstractSemi-supervised clustering aims to incorporate the known prior knowledge into the clustering algorithm. Pairwise constraints and constraint projections are two popular techniques in semi-supervised clustering. However, both of them only consider the given constraints and do not consider the neighbors around the data points constrained by the constraints. This paper presents a new technique by utilizing the constrained pairwise data points and their neighbors, denoted as constraint neighborhood projections that requires fewer labeled data points (constraints) and can naturally deal with constraint conflicts. It includes two steps: 1) the constraint neighbors are chosen according to the pairwise constraints and a given radius so that the pairwise constraint relationships can be extended to their neighbors, and 2) the original data points are projected into a new low-dimensional space learned from the pairwise constraints and their neighbors. A CNP-Kmeans algorithm is developed based on the constraint neighborhood projections. Extensive experiments on University of California Irvine (UCI) datasets demonstrate the effectiveness of the proposed method. Our study also shows that constraint neighborhood projections (CNP) has some favorable features compared with the previous techniques. Hongjun Wang 0002, Tao Li 0001, Tianrui Li 0001, Yan Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2013 | Semi-supervised Clustering Ensemble Evolved by Genetic Algorithm for Web Video Categorization
Amjad Mahmood, Tianrui Li 0001, Yan Yang 0001, Hongjun Wang 0002 |
ADMA (2) | 4 |
| 2013 | Efficient Complex Event Processing under Boolean Model
Shanglian Peng, Tianrui Li 0001, Hongjun Wang 0002, Jia He 0003 |
WAIM | 3 |
| 2012 | Exemplars-Constraints for Semi-supervised Clustering
Hongjun Wang 0002, Tao Li 0001, Tianrui Li 0001, Yan Yang 0001 |
ADMA | 1 |
| 2012 | Constraint projections for semi-supervised affinity propagation
Hongjun Wang 0002, Ruihua Nie, Xingnian Liu, Tianrui Li 0001 |
Knowl. Based Syst. | 1 |