Shudong Huang

dblp:48/2141 · DBLP profile ↗
← Back
67ranked-venue papers
26as first author
49since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 18 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 23 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing
abstract
Dynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynchronous models preserve the event stream's native format, they often neglect spatial information, compromising their adaptability and efficiency. To address these limitations, we propose a Spatiotemporally Separated Sparse Network (S3Net) for efficient event stream encoding and learning. Specifically, we employ a learnable sparse encoding scheme to construct a voxel-structured representation that effectively extracts spatiotemporal relationships among event data. After that, we propose a dual-branch architecture to capture localized spatial dependencies and dynamic temporal patterns of event data. By explicitly decoupling spatial and temporal modeling, S3Net enables end-to-end asynchronous processing of variable-length event sequences, achieving both strong representational capacity and high computational efficiency. Experimental results on six event-based datasets demonstrate that S3Net achieves state-of-the-art performance. Compared to frame-based methods, it significantly reduces computational overhead and model complexity, while also outperforming existing asynchronous approaches in inference speed without compromising accuracy. Extensive experiments across six event-based datasets show that S3Net establishes new state-of-the-art performance. Our method reduces computational costs by 35% and model parameters by 27% compared to frame-based approaches, while delivering 1.58× faster inference than existing point-based methods at comparable accuracy levels.
Rong Xiao 0001, Wanying Xu, Chenwei Tang, Shudong Huang, Huajin Tang
AAAI5
2026 PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence Learning
Zhuoyao Liu, Wentao Feng, Shudong Huang
AAAI4
2026 CasAD: Adaptive irregular events modeling and temporal dynamics disentangling for information popularity prediction
Haoyue Zheng, Lanlan Yu, Shudong Huang, Tao Zhou 0001, Jiancheng Lv 0001, Quanhui Liu
Knowl. Based Syst.4
2026 FDNet: High-frequency disentanglement network with information-theoretic guidance for multivariate time series forecasting
Ao Hu, Liangjian Wen, Jiang Duan, Yong Dai 0001, Dongkai Wang, Shudong Huang, Jun Wang 0089, Zenglin Xu
Pattern Recognit.6
2026 Optimal transport filtering for robust cross-modal retrieval with open-set noisy labels
Xinliu Liu, Ruitao Pu, Yuan Sun 0016, Yingke Chen, Shudong Huang, Dezhong Peng, Yongsheng Sang
Pattern Recognit.5
2025 Learning Dynamic Similarity by Bidirectional Hierarchical Sliding Semantic Probe for Efficient Text Video Retrieval
abstract
Text-video retrieval is a foundation task in multi-modal research which aims to align texts and videos in the embedding space. The key challenge is to learn the similarity between videos and texts. A conventional approach involves directly aligning video-text pairs using cosine similarity. However, due to the disparity in the information conveyed by videos and texts, i.e., a single video can be described from multiple perspectives, the retrieval accuracy is suboptimal. An alternative approach employs cross-modal interaction to enable videos to dynamically acquire distinct features from various texts, thus facilitating similarity calculations. Nevertheless, this solution incurs a computational complexity of O(n^2) during retrieval. To this end, this paper proposes a novel method called Bidirectional Hierarchical Sliding Semantic Probe (BiHSSP), which calculates dynamic similarity between videos and texts with O(n) complexity during retrieval. We introduce a hierarchical semantic probe module that learns semantic probes at different scales for both video and text features. Semantic probe involves a sliding calculation of the cross-correlation between semantic probes at different scales and embeddings from another modality, allowing for dynamic similarity computation between video and text descriptions from various perspectives. Specifically, for text descriptions from different angles, we calculate the similarity at different locations within the video features and vice versa. This approach preserves the complete information of the video while addressing the issue of unequal information between video and text without requiring cross-modal interaction. Additionally, our method can function as a plug-and-play module across various methods, thereby enhancing the corresponding performance. Experimental results demonstrate that our BiHSSP significantly outperforms the baseline.
Yang Liu 0264, Shudong Huang, Deng Xiong, Jiancheng Lv 0001
AAAI2
2025 Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
abstract
Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it difficult to compute the similarity between these two modalities accurately and efficiently. In this paper, we propose a novel framework called Asymmetric Visual Semantic Embedding (AVSE) to dynamically select features from various regions of images tailored to different textual inputs for similarity calculation. To capture information from different views in the image, we design a radial bias sampling module to sample image patches and obtain image features from various views, Furthermore, AVSE introduces a novel module for efficient computation of visual semantic similarity between asymmetric image and text embeddings. Central to this module is the presumption of foundational semantic units within the embeddings, denoted as ``meta-semantic embeddings." It segments all embeddings into meta-semantic embeddings with the same dimension and calculates visual semantic similarity by finding the optimal match of meta-semantic embeddings of two modalities. Our proposed AVSE model is extensively evaluated on the large-scale MS-COCO and Flickr30K datasets, demonstrating its superiority over recent state-of-the-art methods.
Yang Liu 0264, Mengyuan Liu 0001, Shudong Huang, Jiancheng Lv 0001
AAAI3
2025 Multi-view Granular-ball Contrastive Clustering
abstract
Previous multi-view contrastive learning methods typically operate at two scales: instance-level and cluster-level. The former generally constructs positive and negative pairs based on the correspondence between samples and view instances. These methods aim to bring positive pairs closer and push negative pairs further apart in the latent space. This kind of approaches has the drawback of inevitably introducing false negatives in an unsupervised setting, leading to reduced model discriminability. The latter usually involves calculating cluster assignments for samples under each view and maximizing view consensus by reducing distribution discrepancies through methods like optimizing the KL divergence between different view distributions or maximizing mutual information. However, clusters represent a macro structure that overlooks the local structure within the sample set, and the relationships between clusters across different views cannot be explicitly measured. To overcome the shortcomings of these two types of methods, we propose a method named Multi-view Granular-ball Contrastive Clustering (MGBCC). This method segments the sample set into coarse-grained granular balls, and establishes associations between intra-view and cross-view granular balls. These associations are reinforced in a shared latent space, thereby achieving multi-granularity contrastive learning. Granular balls lie between instances and clusters, naturally preserving the local topological structure of the sample set. We conduct extensive experiments to validate the effectiveness of the proposed method.
Shudong Huang, Weihong Ma, Deng Xiong, Jiancheng Lv 0001
AAAI2
2025 Crocodile: Cross Experts Covariance for Disentangled Learning in Multi-Domain Recommendation
abstract
Multi-domain learning (MDL) has become a prominent topic in enhancing the quality of personalized services. It's critical to learn commonalities between domains and preserve the distinct characteristics of each domain. However, this leads to a challenging dilemma in MDL. On the one hand, a model needs to leverage domain-aware modules such as experts or embeddings to preserve each domain's distinctiveness. On the other hand, real-world datasets often exhibit long-tailed distributions across domains, where some domains may lack sufficient samples to effectively train their specific modules. Unfortunately, nearly all existing work falls short of resolving this dilemma. To this end, we propose a novel Cross-experts Covariance Loss for Disentangled Learning model (Crocodile), which employs multiple embedding tables to make the model domain-aware at the embeddings which consist most parameters in the model, and a covariance loss upon these embeddings to disentangle them, enabling the model to capture diverse user interests among domains. Empirical analysis demonstrates that our method successfully addresses both challenges and outperforms all state-of-the-art methods on public datasets. During online A/B testing in Tencent's advertising platform, Crocodile achieves 0.72% CTR lift and 0.73% GMV lift on a primary advertising scenario. The code is openly accessible at: https://github.com/SkylerLinn/Crocodile.
Zhutian Lin, Junwei Pan, Xi Xiao 0001, Ximei Wang, Zhixiang Feng, Shifeng Wen, Shudong Huang, Lei Xiao 0001
CIKM8
2025 Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching
Yang Liu 0264, Wentao Feng, Zhuoyao Liu, Shudong Huang, Jiancheng Lv 0001
ICCV4
2025 DONIS: Importance Sampling for Training Physics-Informed DeepONet
abstract
Deep Operator Network (DeepONet) effectively learns complex operator mappings, especially for systems governed by differential equations. Physics-informed DeepONet (PI-DeepONet) extends these capabilities by integrating physical constraints, enabling robust performance with limited or no labeled data. However, combining operator learning with these constraints increases computational complexity, which makes training more difficult and convergence slower, particularly for nonlinear or high-dimensional problems. In this work, we present an enhanced PI-DeepONet framework, that applies importance sampling to both of DeepONet inputs (i.e., the functions and the collocation points) to alleviate these training challenges. By focusing on critical data regions in both input domains, our approach showcases accelerated convergence and improved accuracy across various complex applications.
Shudong Huang, Wentao Feng
IJCAI1
2025 Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
abstract
Recently, prompt learning has achieved remarkable success in adapting pre-trained Vision-Language Models (VLMs) to downstream tasks such as image classification. However, its application to the downstream Image-Text Retrieval (ITR) task is more challenging. We find that the challenge lies in discriminating both fine-grained attributes and similar subcategories of the downstream data. To address this challenge, we propose Dual prompt Learning with Joint Category-Attribute Reweighting (DCAR), a novel dual-prompt learning framework to achieve precise image-text matching. The framework dynamically adjusts prompt vectors from both semantic and visual dimensions to improve the performance of CLIP on the downstream ITR task. Based on the prompt paradigm, DCAR jointly optimizes attribute and category features to enhance fine-grained representation learning. Specifically, (1) at the attribute level, it dynamically updates the weights of attribute descriptions based on text-image mutual information correlation; and (2) at the category level, it introduces negative samples from multiple perspectives with category-matching weighting to learn subcategory distinctions. To validate our method, we construct the Fine-class Described Retrieval Dataset (FDRD), which serves as a challenging benchmark for ITR in downstream data domains. It covers over 1,500 downstream fine categories and 230,000 image-caption pairs with detailed attribute annotations. Extensive experiments on FDRD demonstrate that DCAR achieves state-of-the-art performance over existing baselines. The code and data are available at https://github.com/wyf202322/DCAR.
Tao Wang 0053, Chenwei Tang, Caiyang Yu, Zhengqing Zang, Mengmi Zhang, Shudong Huang, Jiancheng Lv 0001
ACM Multimedia7
2025 Towards Unifying Feature Interaction Models for Click-Through Rate Prediction
Junwei Pan, Jipeng Jin, Shudong Huang, Xiaofeng Gao 0001, Lei Xiao 0001
ECML/PKDD (5)4
2025 DSAIS-PINN: Dynamic seeds allocation importance sampling for physics-informed neural networks
Wentao Feng, Chenwei Tang, Shudong Huang, Jiancheng Lv 0001
Neurocomputing5
2025 Implicit Multi-Behavior Generative Recommendation With Mixture of Quantization
abstract
Generative recommendation systems have recently seen a surge in interest, largely due to the promising advancements in generative AI. As a competitive solution for multi-behavior sequence recommendations, much of the recent research has concentrated on predicting the next item a user will likely interact with using a generative approach. However, these methods often 1). assign multiple residual quantization layers to obtain item codes, which leads to extra storage costs of more codebooks. And 2). explicitly utilize behavior sequences leading to longer sequences, potentially increasing the training time as well as inference time compared with original sequences. In response to these challenges, we introduce theImplicitMulti-BehaviorGenerative recommendation with a mixture of quantization (IMBGen) approach in this paper. Specifically, we have devised aMixtureofQuantization (MoQ) that combines the merits of both residual and parallel quantization for a more effective tokenization process. Additionally, we propose an Implicit Behavior Modeling (IBM) framework, allowing for more efficient integration of users' behaviors into the interacted items. Finally, we conducted extensive experiments on two widely used benchmark datasets and further confirmed our findings with an online A/B test. The results consistently demonstrate the advantages of our approach over other baseline methods.
Yuze Tan, Yanjie Gou, Kouying Xue, Shudong Huang, Ivor W. Tsang, Jiancheng Lv 0001
IEEE Trans. Knowl. Data Eng.4
2025 Variational Graph Generator for Multiview Graph Clustering
abstract
Multiview graph clustering (MGC) methods are increasingly being studied due to the explosion of multiview data with graph structural information. The critical point of MGC is to better utilize view-specific and view-common information in features and graphs of multiple views. However, existing works have an inherent limitation that they are unable to concurrently utilize the consensus graph information across multiple graphs and the view-specific feature information. To address this issue, we propose a variational graph generator for MGC (VGMGC). Specifically, a novel variational graph generator is proposed to extract common information among multiple graphs. This generator infers a reliable variational consensus graph based on a priori assumption over multiple graphs. Then, a simple yet effective graph encoder in conjunction with the multiview clustering objective is presented to learn the desired graph embeddings for clustering, which embeds the inferred view-common graph and view-specific graphs together with features. Finally, theoretical results illustrate the rationality of the VGMGC by analyzing the uncertainty of the inferred consensus graph with the information bottleneck (IB) principle. Extensive experiments demonstrate the superior performance of our VGMGC over state-of-the-art methods (SOTAs). The source code is publicly available at: https://github.com/cjpcool/VGMGC.
Jianpeng Chen, Yawen Ling, Jie Xu 0044, Yazhou Ren 0001, Shudong Huang, Xiaorong Pu, Zhifeng Hao 0004, Philip S. Yu, Lifang He 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 STSF: Spiking Time Sparse Feedback Learning for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are biologically plausible models known for their computational efficiency. A significant advantage of SNNs lies in the binary information transmission through spike trains, eliminating the need for multiplication operations. However, due to the spatio-temporal nature of SNNs, direct application of traditional backpropagation (BP) training still results in significant computational costs. Meanwhile, learning methods based on unsupervised synaptic plasticity provide an alternative for training SNNs but often yield suboptimal results. Thus, efficiently training high-accuracy SNNs remains a challenge. In this article, we propose a highly efficient and biologically plausible spiking time sparse feedback (STSF) learning method. This algorithm modifies synaptic weights by incorporating a neuromodulator for global supervised learning using sparse direct feedback alignment (DFA) and local homeostasis learning with vanilla spike-timing-dependent plasticity (STDP). Such neuromorphic global-local learning focuses on instantaneous synaptic activity, enabling independent and simultaneous optimization of each network layer, thereby improving biological plausibility, enhancing parallelism, and reducing storage overhead. Incorporating sparse fixed random feedback connections for global error modulation, which uses selection operations instead of multiplication operations, further improves computational efficiency. Experimental results demonstrate that the proposed algorithm markedly reduces the computational cost with significantly higher accuracy comparable to current state-of-the-art algorithms across a wide range of classification tasks. Our implementation codes are available at https://github.com/hppeace/STSF.
Rong Xiao 0001, Chenwei Tang, Shudong Huang, Jiancheng Lv 0001, Huajin Tang
IEEE Trans. Neural Networks Learn. Syst.4
2025 Partial Differential Equations Meet Deep Neural Networks: A Survey
abstract
Many problems in science and engineering can be mathematically modeled using partial differential equations (PDEs), which are essential for fields like computational fluid dynamics (CFD), molecular dynamics, and dynamical systems. Although traditional numerical methods like the finite difference/element method are widely used, their computational inefficiency, due to the large number of iterations required, has long been a challenge. Recently, deep learning (DL) has emerged as a promising alternative for solving PDEs, offering new paradigms beyond conventional methods. Despite the growing interest in techniques like physics-informed neural networks (PINNs), a systematic review of the diverse neural network (NN) approaches for PDEs is still missing. This survey fills that gap by categorizing and reviewing the current progress of deep NNs (DNNs) for PDEs. Unlike previous reviews focused on specific methods like PINNs, we offer a broader taxonomy and analyze applications across scientific, engineering, and medical fields. We also provide a historical overview, key challenges, and future trends, aiming to serve both researchers and practitioners with insights into how DNNs can be effectively applied to solve PDEs.
Shudong Huang, Wentao Feng, Chenwei Tang, Zhenan He 0001, Caiyang Yu, Jiancheng Lv 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 An Effective Augmented Lagrangian Method for Fine-Grained Multi-View Optimization
abstract
The significance of multi-view learning in effectively mitigating the intricate intricacies entrenched within heterogeneous data has garnered substantial attention in recent years. Notwithstanding the favorable achievements showcased by recent strides in this area, a confluence of noteworthy challenges endures. To be specific, a majority of extant methodologies unceremoniously assign weights to data points view-wisely. This ineluctably disregards the intrinsic reality that disparate views confer diverse contributions to each individual sample, consequently neglecting the rich wellspring of sample-level structural insights harbored within the dataset. In this paper, we proposed an effective Augmented Lagrangian MethOd for fiNe-graineD (ALMOND) multi-view optimization. This innovative approach scrutinizes the interplay among multiple views at the granularity of individual samples, thereby fostering the enhanced preservation of local structural coherence. The Augmented Lagrangian Method (ALM) is elaborately incorporated into our framework, which enables us to achieve an optimal solution without involving an inexplicable intermediate variable as previous methods do. Empirical experiments on multi-view clustering tasks across heterogeneous datasets serve to incontrovertibly showcase the effectiveness of our proposed methodology, corroborating its preeminence over incumbent state-of-the-art alternatives.
Yuze Tan, Hecheng Cai, Shudong Huang, Shuping Wei, Jiancheng Lv 0001
AAAI3
2024 Multi-View Clustering by Inter-cluster Connectivity Guided Reward
abstract
Multi-view clustering has been widely explored for its effectiveness in harmonizing heterogeneity along with consistency in different views of data. Despite the significant progress made by recent works, the performance of most existing methods is heavily reliant on strong priori information regarding the true cluster number $\textit{K}$, which is rarely feasible in real-world scenarios. In this paper, we propose a novel graph-based multi-view clustering algorithm to infer unknown $\textit{K}$ through a graph consistency reward mechanism. To be specific, we evaluate the cluster indicator matrix during each iteration with respect to diverse $\textit{K}$. We formulate the inference process of unknown $\textit{K}$ as a parsimonious reinforcement learning paradigm, where the reward is measured by inter-cluster connectivity. As a result, our approach is capable of independently producing the final clustering result, free from the input of a predefined cluster number. Experimental results on multiple benchmark datasets demonstrate the effectiveness of our proposed approach in comparison to existing state-of-the-art methods.
Yang Liu 0264, Hecheng Cai, Shudong Huang, Jiancheng Lv 0001
ICML5
2024 With a Little Help from Language: Semantic Enhanced Visual Prototype Framework for Few-Shot Learning
Hecheng Cai, Yang Liu 0264, Shudong Huang, Jiancheng Lv 0001
IJCAI3
2024 Robust Contrastive Multi-view Kernel Clustering
Yixi Liu, Shujian Li, Shudong Huang, Jiancheng Lv 0001
IJCAI4
2024 Understanding the Ranking Loss for Recommendation with Sparse User Feedback
abstract
Click-through rate (CTR) prediction is a crucial area of research in online advertising. While binary cross entropy (BCE) has been widely used as the optimization objective for treating CTR prediction as a binary classification problem, recent advancements have shown that combining BCE loss with an auxiliary ranking loss can significantly improve performance. However, the full effectiveness of this combination loss is not yet fully understood. In this paper, we uncover a new challenge associated with the BCE loss in scenarios where positive feedback is sparse: the issue of gradient vanishing for negative samples. We introduce a novel perspective on the effectiveness of the auxiliary ranking loss in CTR prediction: it generates larger gradients on negative samples, thereby mitigating the optimization difficulties when using the BCE loss only and resulting in improved classification ability. To validate our perspective, we conduct theoretical analysis and extensive empirical evaluations on public datasets. Additionally, we successfully integrate the ranking loss into Tencent's online advertising system, achieving notable lifts of 0.70% and 1.26% in Gross Merchandise Value (GMV) for two main scenarios. The code is openly accessible at: https://github.com/SkylerLinn/Understanding-the-Ranking-Loss.
Zhutian Lin, Junwei Pan, Shangyu Zhang, Ximei Wang, Xi Xiao 0001, Shudong Huang, Lei Xiao 0001, Jie Jiang 0015
KDD6
2024 Adaptive Instance-wise Multi-view Clustering
abstract
Multi-view clustering has garnered attention for its effectiveness in addressing heterogeneous data by unsupervisedly revealing underlying correlations between different views. As a mainstream method, multi-view graph clustering has attracted increasing attention in recent years. Despite its success, it still has some limitations. Notably, many methods construct the similarity graph without considering the local geometric structure and exploit coarse-grained complementary and consensus information from different views at the view level. To solve the shortcomings, we focus on local structure consistency and fine-grained representations across multiple views. Specifically, each view's local consistency similarity graph is obtained through the adaptive neighbor. Subsequently, the multi-view similarity tensor is rotated and sliced into fine-grained instance-wise slices. Finally, these slices are fused into the final similarity matrix. Consequently, cross-view consistency can be captured by exploring the intersections of multiple views in an instance-wise manner. We design a collaborative framework with the augmented Lagrangian method to refine all subtasks towards optimal solutions iteratively. Extensive experiments on several multi-view datasets confirm the significant enhancement in clustering accuracy achieved by our method.
Shudong Huang, Hecheng Cai, Wentao Feng, Jiancheng Lv 0001
ACM Multimedia1
2024 Multi-Sequence Attentive User Representation Learning for Side-information Integrated Sequential Recommendation
abstract
Side-information integrated sequential recommendation incorporates supplementary information to alleviate the issue of data sparsity. The state-of-the-art works mainly leverage some side information to improve the attention calculation to learn user representation more accurately. However, there are still some limitations to be addressed in this topic. Most of them merely learn the user representation at the item level and overlook the association of the item sequence and the side-information sequences when calculating the attentions, which results in the incomprehensive learning of user representation. Some of them learn the user representations at both the item and side-information levels, but they still face the problem of insufficient optimization of multiple user representations. To address these limitations, we propose a novel model, i.e., Multi-Sequence Sequential Recommender (MSSR), which learns the user's multiple representations from diverse sequences. Specifically, we design a multi-sequence integrated attention layer to learn more attentive pairs than the existing works and adaptively fuse these pairs to learn user representation. Moreover, our user representation alignment module constructs the self-supervised signals to optimize the representations. Subsequently, they are further refined by our side information predictor during training. For item prediction, our MSSR extra considers the side information of the candidate item, enabling a comprehensive measurement of the user's preferences. Extensive experiments on four public datasets show that our MSSR outperforms eleven state-of-the-art baselines. Visualization and case study also demonstrate the rationality and interpretability of our MSSR.
Xiaolin Lin, Jinwei Luo, Junwei Pan, Weike Pan, Zhong Ming 0001, Shudong Huang, Jie Jiang 0015
WSDM7
2024 Euclidean Distance is Not Your Swiss Army Knife
abstract
Graph-based multi-view learning, which has hitherto been used to discover the intrinsic patterns of graph data giving the credit to its convenience of implementation and effectiveness. Note that even though these approaches have been increasingly adopted in multi-view clustering and have generated promising outcomes, they are still faced with the sub-optimal solution. For one thing, multi-view data can be corrupted in the raw feature space. For the other, most existing approaches normally utilize euclidean distance to obtain the similarity between two samples, which can not be the best option for all types of real-world data and leads to inferior results. Therefore, to overcome the aforementioned issues, we integrate multi-metric learning, graph filtering, and subspace learning into a collaborative learning framework for multi-view clustering. Particularly, we prefer to recover a smooth representation of data by graph filtering, which can reserve the geometric structure of the original multi-view data and discard the corruptions simultaneously. Furthermore, instead of using euclidean distance as a Swiss army knife, multiple metrics are utilized to fully exploit the correlation of data based on the smooth representation, hence finally facilitating the downstream clustering task. Extensive experiments on multi-view clustering tasks validate our theoretical findings of ours and prove the improvement of our method over the SOTA approaches.
Yuze Tan, Yixi Liu, Hongjie Wu, Shudong Huang, Zenglin Xu, Ivor W. Tsang, Jiancheng Lv 0001
IEEE Trans. Knowl. Data Eng.4
2024 Self-Weighted Contrastive Fusion for Deep Multi-View Clustering
abstract
Multi-view clustering can explore consensus information from multiple views and has attracted increasing attention in the past two decades. However, existing works face two major challenges: i) how to deal with the conflict between learning view-consensus information and reconstructing inconsistent viewprivate information, and ii) how to mitigate representation degeneration caused by implementing the consistency objective for multi-view data. To address these challenges, we propose a novel framework of self-weighted contrastive fusion for deep multi-view clustering (SCMVC). First, our method establishes a hierarchical feature fusion framework, effectively segregating the consistency objective from the reconstruction objective. Then, multi-view contrastive fusion is implemented via maximizing consistency expression between the view-consensus representation and global representation, fully exploring the view consistency and complementary. More importantly, we propose to measure the discrepancy between pairwise representations, and then introduce a self-weighting method, which adaptively strengthens useful views in feature fusion and weakens unreliable views, to mitigate representation degeneration. Extensive experiments on nine public datasets demonstrate that our proposed method achieves state-of-the-art clustering performance. The code is available athttps://github.com/SongwuJob/SCMVC.
Yazhou Ren 0001, Jing He 0004, Xiaorong Pu, Shudong Huang, Zhifeng Hao 0004, Lifang He 0001
IEEE Trans. Multim.6
2024 CGDD: Multiview Graph Clustering via Cross-Graph Diversity Detection
abstract
Multiview graph clustering has emerged as an important yet challenging technique due to the difficulty of exploiting the similarity relationships among multiple views. Typically, the similarity graph for each view learned by these methods is easily corrupted because of the unavoidable noise or diversity among views. To recover a clean graph, existing methods mainly focus on the diverse part within each graph yet overlook the diversity across multiple graphs. In this article, instead of merely considering the sparsity of diversity within a graph as previous methods do, we incline to a more suitable consideration that the diversity should be sparse across graphs. It is intuitive that the divergent parts are supposed to be inconsistent with each other, otherwise it would contradict the definition of diversity. By simultaneously and explicitly detecting the multiview consistency and cross-graph diversity, a pure graph for each view can be expected. The multiple pure graphs are further fused to the structured consensus graph with exactly r connected components where r is the number of clusters. Once the consensus graph is obtained, the cluster label to each instance can be directly allocated as each connected component precisely corresponds to an individual cluster. An alternating iterative algorithm is designed to optimize the subtasks of learning the similarity graphs adaptively, detecting the consistency as well as cross-graph diversity, fusing the multiple pure graphs, and assigning cluster label to each instance in a mutual reinforcement manner. Extensive experimental results on several benchmark multiview datasets demonstrate the effectiveness of our model, in comparison to several state-of-the-art algorithms.
Shudong Huang, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Self-Supervised Graph Attention Networks for Deep Weighted Multi-View Clustering
abstract
As one of the most important research topics in the unsupervised learning field, Multi-View Clustering (MVC) has been widely studied in the past decade and numerous MVC methods have been developed. Among these methods, the recently emerged Graph Neural Networks (GNN) shine a light on modeling both topological structure and node attributes in the form of graphs, to guide unified embedding learning and clustering. However, the effectiveness of existing GNN-based MVC methods is still limited due to the insufficient consideration in utilizing the self-supervised information and graph information, which can be reflected from the following two aspects: 1) most of these models merely use the self-supervised information to guide the feature learning and fail to realize that such information can be also applied in graph learning and sample weighting; 2) the usage of graph information is generally limited to the feature aggregation in these models, yet it also provides valuable evidence in detecting noisy samples. To this end, in this paper we propose Self-Supervised Graph Attention Networks for Deep Weighted Multi-View Clustering (SGDMC), which promotes the performance of GNN-based deep MVC models by making full use of the self-supervised information and graph information. Specifically, a novel attention-allocating approach that considers both the similarity of node attributes and the self-supervised information is developed to comprehensively evaluate the relevance among different nodes. Meanwhile, to alleviate the negative impact caused by noisy samples and the discrepancy of cluster structures, we further design a sample-weighting strategy based on the attention graph as well as the discrepancy between the global pseudo-labels and the local cluster assignment. Experimental results on multiple real-world datasets demonstrate the effectiveness of our method over existing approaches.
Zongmo Huang, Yazhou Ren 0001, Xiaorong Pu, Shudong Huang, Zenglin Xu, Lifang He 0001
AAAI4
2023 Metric Multi-View Graph Clustering
abstract
Graph-based methods have hitherto been used to pursue the coherent patterns of data due to its ease of implementation and efficiency. These methods have been increasingly applied in multi-view learning and achieved promising performance in various clustering tasks. However, despite their noticeable empirical success, existing graph-based multi-view clustering methods may still suffer the suboptimal solution considering that multi-view data can be very complicated in raw feature space. Moreover, existing methods usually adopt the similarity metric by an ad hoc approach, which largely simplifies the relationship among real-world data and results in an inaccurate output. To address these issues, we propose to seamlessly integrates metric learning and graph learning for multi-view clustering. Specifically, we employ a useful metric to depict the inherent structure with linearity-aware of affinity graph representation learned based on the self-expressiveness property. Furthermore, instead of directly utilizing the raw features, we prefer to recover a smooth representation such that the geometric structure of the original data can be retained. We model the above concerns into a unified learning framework, and hence complements each learning subtask in a mutual reinforcement manner. The empirical studies corroborate our theoretical findings, and demonstrate that the proposed method is able to boost the multi-view clustering performance.
Yuze Tan, Yixi Liu, Hongjie Wu, Jiancheng Lv 0001, Shudong Huang
AAAI5
2023 Sample-level Multi-view Graph Clustering
abstract
Multi-view clustering has hitherto been studied due to their effectiveness in dealing with heterogeneous data. Despite the empirical success made by recent works, there still exists several severe challenges. Particularly, previous multi-view clustering algorithms seldom consider the topological structure in data, which is essential for clustering data on manifold. Moreover, existing methods cannot fully explore the consistency of local structures between different views as they uncover the clustering structure in a intra-view way instead of a inter-view manner. In this paper, we propose to exploit the implied data manifold by learning the topological structure of data. Besides, considering that the consistency of multiple views is manifested in the generally similar local structure while the inconsistent structures are the minority, we further explore the intersections of multiple views in the sample level such that the cross-view consistency can be better maintained. We model the above concerns in a unified framework and design an efficient algorithm to solve the corresponding optimization problem. Experimental results on various multi-view datasets certificate the effectiveness of the proposed method and verify its superiority over other SOTA approaches.
Yuze Tan, Yixi Liu, Shudong Huang, Wentao Feng, Jiancheng Lv 0001
CVPR3
2023 Lifelong Multi-view Spectral Clustering
abstract
In recent years, spectral clustering has become a well-known and effective algorithm in machine learning. However, traditional spectral clustering algorithms are designed for single-view data and fixed task setting. This can become a limitation when dealing with new tasks in a sequence, as it requires accessing previously learned tasks. Hence it leads to high storage consumption, especially for multi-view datasets. In this paper, we address this limitation by introducing a lifelong multi-view clustering framework. Our approach uses view-specific knowledge libraries to capture intra-view knowledge across different tasks. Specifically, we propose two types of libraries: an orthogonal basis library that stores cluster centers in consecutive tasks, and a feature embedding library that embeds feature relations shared among correlated tasks. When a new clustering task is coming, the knowledge is iteratively transferred from libraries to encode the new task, and knowledge libraries are updated according to the online update formulation. Meanwhile, basis libraries of different views are further fused into a consensus library with adaptive weights. Experimental results show that our proposed method outperforms other competitive clustering methods on multi-view datasets by a large margin.
Hecheng Cai, Yuze Tan, Shudong Huang, Jiancheng Lv 0001
IJCAI3
2023 Preserving Local and Global Information: An Effective Metric-based Subspace Clustering
abstract
Subspace clustering, which recoveries the subspace representation in the form of an affinity graph, has drawn tons of attention due to its effectiveness in various clustering tasks. However, existing subspace clustering methods are usually fed with raw data, which may lead to a suboptimal result since it is difficult to directly and accurately depict the inherent relation between data points. In this paper, we propose a novel subspace clustering method by holistically utilizing the pairwise similarity and graph geometric structure. Our model first constructs an initial subspace representation by means of self-expression, which is able to depict the global structure of data. Then, we use an effective metric to recover an intrinsic matrix with pairwise similarity based on the obtained representation, which further preserves the local structure. Besides, we propose to facilitate the downstream subspace learning task by searching for a smooth representation of the original data, which is obtained by applying a low-pass filter to retain the graph geometric features. By leveraging the subtasks of learning the smooth representation, performing the subspace learning, and recovering the intrinsic similarity matrix in a unified learning framework, each subtask can be alternately boosted. Experiments on several benchmark data sets have been conducted to verify the proposed method.
Yixi Liu, Yuze Tan, Hongjie Wu, Shudong Huang, Yazhou Ren 0001, Jiancheng Lv 0001
ACM Multimedia4
2023 Multimodal Physiological Signals Fusion for Online Emotion Recognition
abstract
Multimodal physiological-based emotion recognition is one of the most available but challenging studies due to complexity of emotions and individual differences in physiological signals. However, existing studies mainly combine multimodal data to fuse multimodal information in offline scenarios, ignoring data/modalities correlation among multimodal data and individual differences of non-stationary physiological signals in online scenarios. In this paper, we propose a novel Online Multimodal HyperGraph Learning (OMHGL) method to fuse multimodal information for emotion recognition based on time-series physiological signals. Our method consists of multimodal hypergraph fusion and online hypergraph learning. Specifically, the multimodal hypergraph fusion can fuse multimodal physiological signals to effectively obtain emotionally dependent information via leveraging multimodal information and higher-order correlations among multimodal data/modalities. The online hypergraph learning is designed to learn new information from online data by updating hypergraph projection. As a result, the proposed online emotion recognition model can be more effective for emotion recognition of target subjects when target data arrive in an online manner. Experimental results have demonstrated that the proposed method significantly outperforms the baselines and compared state-of-the-art methods in online emotion recognition tasks.
Tongjie Pan, Yalan Ye, Hecheng Cai, Shudong Huang, Yang Yang 0002, Guoqing Wang 0001
ACM Multimedia4
2023 PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning
abstract
Classical federated learning (FL) enables training machine learning models without sharing data for privacy preservation, but heterogeneous data characteristic degrades the performance of the localized model. Personalized FL (PFL) addresses this by synthesizing personalized models from a global model via training on local data. Such a global model may overlook the specific information that the clients have been sampled. In this paper, we propose a novel scheme to inject personalized prior knowledge into the global model in each client, which attempts to mitigate the introduced incomplete information problem in PFL. At the heart of our proposed approach is a framework, the $\textit{PFL with Bregman Divergence}$ (pFedBreD), decoupling the personalized prior from the local objective function regularized by Bregman divergence for greater adaptability in personalized scenarios. We also relax the mirror descent (RMD) to extract the prior explicitly to provide optional strategies. Additionally, our pFedBreD is backed up by a convergence analysis. Sufficient experiments demonstrate that our method reaches the $\textit{state-of-the-art}$ performances on 5 datasets and outperforms other methods by up to 3.5% across 8 benchmarks. Extensive analyses verify the robustness and necessity of proposed designs. The code will be made public.
Mingjia Shi, Yuhao Zhou 0004, Kai Wang 0036, Huaizheng Zhang, Shudong Huang, Jiancheng Lv 0001
NeurIPS5
2023 Evaluating the impact of test-trace-isolate for COVID-19 management and alternative strategies
abstract
There are many contrasting results concerning the effectiveness of Test-Trace-Isolate (TTI) strategies in mitigating SARS-CoV-2 spread. To shed light on this debate, we developed a novel static-temporal multiplex network characterizing both the regular (static) and random (temporal) contact patterns of individuals and a SARS-CoV-2 transmission model calibrated with historical COVID-19 epidemiological data. We estimated that the TTI strategy alone could not control the disease spread: assuming R0 = 2.5, the infection attack rate would be reduced by 24.5%. Increased test capacity and improved contact trace efficiency only slightly improved the effectiveness of the TTI. We thus investigated the effectiveness of the TTI strategy when coupled with reactive social distancing policies. Limiting contacts on the temporal contact layer would be insufficient to control an epidemic and contacts on both layers would need to be limited simultaneously. For example, the infection attack rate would be reduced by 68.1% when the reactive distancing policy disconnects 30% and 50% of contacts on static and temporal layers, respectively. Our findings highlight that, to reduce the overall transmission, it is important to limit contacts regardless of their types in addition to identifying infected individuals through contact tracing, given the substantial proportion of asymptomatic and pre-symptomatic SARS-CoV-2 transmission.
Zhichu Xia, Shudong Huang, Gui-Quan Sun, Jiancheng Lv 0001, Marco Ajelli, Keisuke Ejima, Quanhui Liu
PLoS Comput. Biol.3
2023 Pure graph-guided multi-view subspace clustering
Hongjie Wu, Shudong Huang, Chenwei Tang, Yancheng Zhang, Jiancheng Lv 0001
Pattern Recognit.2
2023 Multi-View Subspace Clustering by Joint Measuring of Consistency and Diversity
abstract
In multi-view subspace clustering, it is significant to find a common latent space in which the multi-view datasets are located. A number of multi-view subspace clustering methods have been proposed to explore the common latent subspace and achieved promising performance. However, previous multi-view subspace clustering algorithms seldom consider the multi-view consistency and multi-view diversity, let alone take them into consideration simultaneously. In this paper, we propose a novel multi-view subspace clustering by joint measuring the consistency and diversity, which is able to exploit these two complementary criteria seamlessly into a holistic design of clustering algorithms. The proposed model first searches a pure graph for each view by detecting the intrinsic consistent and diverse parts. A consensus graph is then obtained by fusing the multiple pure graphs. Moreover, the consensus graph is structurized to contain exactly$c$connected components where$c$is the number of clusters. In this way, the final clustering result can be obtained directly since each connected component precisely corresponds to an individual cluster. Extensive experimental studies on various datasets manifest that our model achieves comparable performance than the other state-of-the-art methods.
Shudong Huang, Yixi Liu, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001
IEEE Trans. Knowl. Data Eng.1
2023 Latent Representation Guided Multi-View Clustering
abstract
Multi-view clustering aims to reveal the correlation between different input modalities in an unsupervised way. Similarity between data samples can be described by a similarity graph, which governs the quality of multi-view clustering. However, existing multi-view graph learning methods mainly construct similarity graph based on raw features, which are unreliable as real-world datasets usually contain noises, outliers, or even redundant information. In this paper, we formulate a novel model to simultaneously learn a robust structured similarity graph and perform multi-view clustering. The similarity graph is adaptively learned based on a latent representation that is invulnerable to noises and outliers. Furthermore, the similarity graph is enforced to contain a clear structure, i.e., the number of connected components of the target graph is exactly equal to the ground-truth class number. Consequently, the label to each data sample can be directly assigned without any postprocessing. As a result, our model aims at accomplishing three subtasks: latent representation extraction, similarity graph learning, and cluster label allocation, in a unified framework. These three subtasks are seamlessly integrated and can be mutually boosted by each other towards the overall optimal solution. An efficient alternation algorithm is proposed to solve the optimization problem. Experimental results on several benchmark datasets illustrate the effectiveness of the proposed model.
Shudong Huang, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001
IEEE Trans. Knowl. Data Eng.1
2022 Multi-View Clustering on Topological Manifold
abstract
Multi-view clustering has received a lot of attentions in data mining recently. Though plenty of works have been investigated on this topic, it is still a severe challenge due to the complex nature of the multiple heterogeneous features. Particularly, existing multi-view clustering algorithms fail to consider the topological structure in the data, which is essential for clustering data on manifold. In this paper, we propose to exploit the implied data manifold by learning the topological relationship between data points. Our method coalesces multiple view-wise graphs with the topological relevance considered, and learns the weights as well as the consensus graph interactively in a unified framework. Furthermore, we manipulate the consensus graph by a connectivity constraint such that the data points from the same cluster are precisely connected into the same component. Substantial experiments on both toy data and real datasets are conducted to validate the effectiveness of the proposed method, compared to the state-of-the-art algorithms over the clustering performance.
Shudong Huang, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001, Quanhui Liu
AAAI1
2022 Learning Smooth Representation for Multi-view Subspace Clustering
abstract
Multi-view subspace clustering aims to exploit data correlation consensus among multiple views, which essentially can be treated as graph-based approach. However, existing methods usually suffer from suboptimal solution as the raw data might not be separable into subspaces. In this paper, we propose to achieve a smooth representation for each view and thus facilitate the downstream clustering task. It is based on a assumption that a graph signal is smooth if nearby nodes on the graph have similar features representations. Specifically, our mode is able to retain the graph geometric features by applying a low-pass filter to extract the smooth representations of multiple views. Besides, our method achieves the smooth representation learning as well as multi-view clustering interactively in a unified framework, hence it is an end-to-end single-stage learning problem. Substantial experiments on benchmark multi-view datasets are performed to validate the effectiveness of the proposed method, compared to the state-of-the-arts over the clustering performance.
Shudong Huang, Yixi Liu, Yazhou Ren 0001, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001
ACM Multimedia1
2022 Self-Paced Label Distribution Learning for In-The-Wild Facial Expression Recognition
abstract
Label distribution learning (LDL) has achieved great progress in facial expression recognition (FER), where the generating label distribution is a key procedure for LDL-based FER. However, many existing researches have shown the common problem with noisy samples in FER, especially on in-the-wild datasets. This issue may lead to generating unreliable label distributions (which can be seen as label noise), and will further negatively affect the FER model. To this end, we propose a play-and-plug method of self-paced label distribution learning (SPLDL) for in-the-wild FER. Specifically, a simple yet efficient label distribution generator is adopted to generate label distributions to guide label distribution learning. We then introduce self-paced learning (SPL) paradigm and develop a novel self-paced label distribution learning strategy, which considers both classification losses and distribution losses. SPLDL first learns easy samples with reliable label distributions and gradually steps to complex ones, effectively suppressing the negative impact introduced by noisy samples and unreliable label distributions. Extensive experiments on in-the-wild FER datasets (\emphi.e., RAF-DB and AffectNet) based on three backbone networks demonstrate the effectiveness of the proposed method.
Jianjian Shao, Zhenqian Wu, Yuanyan Luo, Shudong Huang, Xiaorong Pu, Yazhou Ren 0001
ACM Multimedia4
2022 Multi-view Subspace Clustering on Topological Manifold
abstract
Multi-view subspace clustering aims to exploit a common affinity representation by means of self-expression. Plenty of works have been presented to boost the clustering performance, yet seldom considering the topological structure in data, which is crucial for clustering data on manifold. Orthogonal to existing works, in this paper, we argue that it is beneficial to explore the implied data manifold by learning the topological relationship between data points. Our model seamlessly integrates multiple affinity graphs into a consensus one with the topological relevance considered. Meanwhile, we manipulate the consensus graph by a connectivity constraint such that the connected components precisely indicate different clusters. Hence our model is able to directly obtain the final clustering result without reliance on any label discretization strategy as previous methods do. Experimental results on several benchmark datasets illustrate the effectiveness of the proposed model, compared to the state-of-the-art competitors over the clustering performance.
Shudong Huang, Hongjie Wu, Yazhou Ren 0001, Ivor W. Tsang, Zenglin Xu, Wentao Feng, Jiancheng Lv 0001
NeurIPS1
2022 Spaks: Self-paced multiple kernel subspace clustering with feature smoothing regularization
Zhao Kang 0001, Zenglin Xu, Shudong Huang, Hongguang Fu
Knowl. Based Syst.4
2022 Multiple partitions alignment via spectral rotation
Shudong Huang, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001
Mach. Learn.1
2022 Efficient federated multi-view learning
Shudong Huang, Zenglin Xu, Ivor W. Tsang, Jiancheng Lv 0001
Pattern Recognit.1
2022 Measuring Diversity in Graph Learning: A Unified Framework for Structured Multi-View Clustering
abstract
Graph learning has emerged as a promising technique for multi-view clustering due to its efficiency of learning a unified graph from multiple views. Previous multi-view graph learning methods mainly try to exploit the multi-view consistency to boost learning performance. However, these methods ignore the prevalent multi-view diversity which may be induced by noise, corruptions, or even view-specific attributes. In this paper, we propose to simultaneously and explicitly leverage the multi-view consistency and the multi-view diversity in a unified framework. The consistent parts are further fused to our target graph with a clear clustering structure, on which the cluster label to each instance can be directly allocated without any postprocessing such as$k$-means in classical spectral clustering. In addition, our model can automatically assign suitable weight for each view based on its clustering capacity. By leveraging the subtasks of measuring the diversity of graphs, integrating the consistent parts with automatically learned weights, and allocating cluster label to each instance in a joint framework, each subtask can be alternately boosted by utilizing the results of the others towards an overall optimal solution. Extensive experimental results on several benchmark multi-view datasets demonstrate the effectiveness of our model in comparison to several state-of-the-art algorithms.
Shudong Huang, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001
IEEE Trans. Knowl. Data Eng.1
2021 CDD: Multi-view Subspace Clustering via Cross-view Diversity Detection
abstract
The goal of multi-view subspace clustering is to explore a common latent space where the multi-view data points lying on. Myriads of subspace learning algorithms have been investigated to boost the performance of multi-view clustering, but seldom exploiting both the multi-view consistency and multi-view diversity, let alone taking them into consideration simultaneously. To do so, we lodge a novel multi-view subspace clustering via cross-view diversity detection (CDD). CDD is able to exploit these two complementary criteria seamlessly into a holistic design of clustering algorithms. With the consistent part and diverse part being detected, a pure graph can be derived for each view. The consistent pure parts of different views are further fused to a consensus structured graph with exactly k connected components where k is the number of clusters. Thus we can directly obtain the final clustering result without any postprocessing as each connected component precisely corresponds to an individual cluster. We model the above concerns into a unified optimization framework. Our empirical studies validate that the proposed model outperforms several other state-of-the-art methods.
Shudong Huang, Ivor W. Tsang, Zenglin Xu, Jiancheng Lv 0001, Quanhui Liu
ACM Multimedia1
2021 Robust deep k-means: An effective and simple method for data clustering
Shudong Huang, Zhao Kang 0001, Zenglin Xu, Quanhui Liu
Pattern Recognit.1
2020 Regularized nonnegative matrix factorization with adaptive local structure learning
Shudong Huang, Zenglin Xu, Zhao Kang 0001, Yazhou Ren 0001
Neurocomputing1
2020 Self-paced and auto-weighted multi-view clustering
Yazhou Ren 0001, Shudong Huang, Minghao Han, Zenglin Xu
Neurocomputing2
2020 Auto-weighted multi-view co-clustering with bipartite graphs
Shudong Huang, Zenglin Xu, Ivor W. Tsang, Zhao Kang 0001
Inf. Sci.1
2020 Multi-graph fusion for multi-view spectral clustering
Zhao Kang 0001, Guoxin Shi, Shudong Huang, Wenyu Chen 0001, Xiaorong Pu, Joey Tianyi Zhou, Zenglin Xu
Knowl. Based Syst.3
2020 Auto-weighted multi-view clustering via deep matrix decomposition
Shudong Huang, Zhao Kang 0001, Zenglin Xu
Pattern Recognit.1
2019 Multiple Partitions Aligned Clustering
abstract
Multi-view clustering is an important yet challenging task due to the difficulty of integrating the information from multiple representations. Most existing multi-view clustering methods explore the heterogeneous information in the space where the data points lie. Such common practice may cause significant information loss because of unavoidable noise or inconsistency among views. Since different views admit the same cluster structure, the natural space should be all partitions. Orthogonal to existing techniques, in this paper, we propose to leverage the multi-view information by fusing partitions. Specifically, we align each partition to form a consensus cluster indicator matrix through a distinct rotation matrix. Moreover, a weight is assigned for each view to account for the clustering capacity differences of views. Finally, the basic partitions, weights, and consensus clustering are jointly learned in a unified framework. We demonstrate the effectiveness of our approach on several real datasets, where significant improvement is found over other state-of-the-art multi-view clustering methods.
Zhao Kang 0001, Zipeng Guo, Shudong Huang, Siying Wang 0002, Wenyu Chen 0001, Yuanzhang Su, Zenglin Xu
IJCAI3
2019 Self-paced and soft-weighted nonnegative matrix factorization for data representation
Shudong Huang, Yazhou Ren 0001, Tianrui Li 0001, Zenglin Xu
Knowl. Based Syst.1
2019 Improved Gaussian-Bernoulli restricted Boltzmann machine for learning discriminative representations
Ji Zhang 0012, Hongjun Wang 0002, Jielei Chu, Shudong Huang, Tianrui Li 0001, Qigang Zhao
Knowl. Based Syst.4
2019 Auto-weighted multi-view clustering via kernelized graph learning
Shudong Huang, Zhao Kang 0001, Ivor W. Tsang, Zenglin Xu
Pattern Recognit.1
2018 Robust graph regularized nonnegative matrix factorization for clustering
Shudong Huang, Hongjun Wang 0002, Tao Li 0001, Tianrui Li 0001, Zenglin Xu
Data Min. Knowl. Discov.1
2018 Robust multi-view data clustering with multi-view capped-norm K-means
Shudong Huang, Yazhou Ren 0001, Zenglin Xu
Neurocomputing1
2018 Self-weighted multi-view clustering with soft capped norm
Shudong Huang, Zhao Kang 0001, Zenglin Xu
Knowl. Based Syst.1
2018 Adaptive local structure learning for document co-clustering
Shudong Huang, Zenglin Xu, Jiancheng Lv 0001
Knowl. Based Syst.1
2018 Overview of the NIST 2016 LoReHLT evaluation
Audrey Tong, Lukas L. Diduch, Jon Fiscus, Yasaman Haghpanah, Shudong Huang, David Joy, Kay Peterson, Ian Soboroff
Mach. Transl.5
2017 Nonnegative matrix factorization with adaptive neighbors
abstract
Nonnegative Matrix factorization (NMF) and its graph regularized extensions have been playing an outstanding role in machine learning and data mining. Recent studies of graph regularized NMF have focused on the application for clustering algorithms. The clustering results of these methods highly depend on the data similarity learning since the data group is utilized based on the input data similarity matrix. Previous graph regularized NMF usually construct the data graph by considering the K-Nearest Neighbors (KNN) which may mislead the factorization since the nearest neighbors may belong to different clusters. That is, it is not a good similarity measurement to construct the data graph by considering the KNN. In this paper, we present NMF with Adaptive Neighbors (NMFAN) for clustering. NMFAN learns the data similarity matrix by assigning the adaptive and optimal neighbors for each data point by exploring the local connectivity of data. It is based on the assumption that the data points with a smaller distance should have a larger probability to be neighbors. Furthermore, in order to achieve the ideal neighbors assignment, we constrain the data similarity matrix such that the neighbors assignment becomes an adaptive process, thus an ideal neighbors assignment can be expected. In order to solve the optimization problem of our method, an efficient iterative updating algorithm is proposed and its convergence is also guaranteed theoretically. Experiments on benchmark data sets demonstrate the effectiveness of the proposed method.
Shudong Huang, Zenglin Xu, Fei Wang 0001
IJCNN1
2016 Constraint Co-Projections for Semi-Supervised Co-Clustering
abstract
Co-clustering aims to simultaneously cluster the objects and features to explore intercorrelated patterns. However, it is usually difficult to obtain good co-clustering results by just analyzing the object-feature correlation data due to the sparsity of the data and the noise. Meanwhile, most co-clustering algorithms cannot take the prior information into consideration and may produce unmeaningful results. Semi-supervised co-clustering aims to incorporate the known prior knowledge into the co-clustering algorithm. In this paper, a new technique named constraint co-projections for semi-supervised co-clustering (CPSSCC) is presented. Constraint co-projections can not only make use of two popular techniques including pairwise constraints and constraint projections, but also simultaneously perform the object constraint projections and feature constraint projections. The two popular techniques are illustrated for semi-supervised co-clustering when some objects and features are believed to be in the same cluster a priori. Furthermore, we also prove that the co-clustering problem can be formulated as a typical eigen-problem and can be efficiently solved with the selected eigenvectors. To the best of our knowledge, constraint co-projections is first stated in this paper and this is the first work on using CPSSCC. Extensive experiments on benchmark data sets demonstrate the effectiveness of the proposed method. This paper also shows that CPSSCC has some favorable features compared with previous related co-clustering algorithms.
Shudong Huang, Hongjun Wang 0002, Tao Li 0001, Yan Yang 0001, Tianrui Li 0001
IEEE Trans. Cybern.1
2015 Spectral co-clustering ensemble
Shudong Huang, Hongjun Wang 0002, Dingcheng Li, Yan Yang 0001, Tianrui Li 0001
Knowl. Based Syst.1
2006 The Mixer and Transcript Reading Corpora: Resources for Multilingual, Crosschannel Speaker Recognition Research
Christopher Cieri, Walter D. Andrews, Joseph P. Campbell, George R. Doddington, John J. Godfrey, Shudong Huang, Mark Y. Liberman, Alvin F. Martin, Hirotaka Nakasone, Mark A. Przybocki, Kevin Walker
LREC6