Shubin Ma

dblp:252/9290 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0002-9794-2661ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Sample Weighted Incomplete Multimodal Clustering Based on Graph Coarsening Label Extraction
abstract
Multimodal data is typically collected through heterogeneous sensors and processing pipelines. However, due to variations in acquisition environments, device capabilities, and feature extraction methods, such data often suffers from incompleteness and inconsistent quality across modalities. To address these challenges, prior studies have explored modality selection and data completion strategies to improve information fusion. Nevertheless, these approaches face two main limitations: (1) they struggle to simultaneously ensure computational efficiency for large-scale graph data and maintain structural and semantic consistency across heterogeneous modality graphs; and (2) most of them operate at the modality level and fail to capture fine-grained, sample-specific quality variations. To overcome these issues, we propose a novel clustering framework, Sample Weighted Incomplete Multimodal Clustering Based on Graph Coarsening Label Extraction (IMC-GCSW). The proposed method introduces a graph coarsening-based label extraction strategy. It significantly reduces the computational cost of multimodal graph processing, while preserving key node information and local topological structures. Furthermore, a quality-aware sample weighting strategy is designed to enable fine-grained modeling of modality-specific data quality, allowing the model to dynamically suppress the influence of low-quality modalities on individual samples. Experiments on both general-purpose datasets and the Fructus Aurantii Disease and Pest Datasets demonstrate that the proposed method exhibits superior performance and strong adaptability in handling multimodal data with incompleteness and quality inconsistency.
Zhenjiao Liu, Jiao Xue, Shubin Ma, Liang Zhao 0005
AAAI5
2026 KNNDA: A New Perspective of Alignment Recovery for Partially View-Aligned Clustering
abstract
In multi-view clustering (MVC), complementary and consistent information from multiple views is integrated to improve clustering performance. However, inter-view sample correspondences may be partially missing in practice, making it difficult to learn cross-view consistency, which leads to the partially view-aligned problem (PVP). Most existing partially view-aligned clustering (PVC) methods first learn cross-view consistent representations based on known alignments, and then recover missing correspondences by measuring cross-view similarity between samples. However, such an indirect alignment recovery process depends on high-quality consistent representations and lacks effective utilization of known alignments, often resulting in sub-optimal outcomes. To address this, we propose a novel direct alignment recovery perspective, instantiated as K-Nearest Neighbors Direct Alignment (KNNDA). Specifically, we first construct an alignment domain by mapping the aligned neighbors of each unaligned sample into the aligned view. Then, we compute alignment confidence based on the similarity between known aligned pairs of neighbors. In particular, we use a dynamic threshold to filter out unreliable alignments. Finally, new alignments are generated within the high-confidence alignment domain. Contrastive loss is used to learn consistent representations for clustering. Comprehensive experiments on several real-world datasets show the effectiveness and superiority of our module in partially view-aligned clustering.
Liang Zhao 0005, Tianqi Yue, Shubin Ma, Bo Xu 0008
AAAI3
2026 Dual-Selection and Optimal Transport for Robust Incomplete and Unaligned Multi-view Clustering
Liang Zhao 0005, Shubin Ma, Chuanye He, Chenhui Yao
Pattern Recognit.3
2026 Multi-View Aligned Clustering via Sample-Bundled Optimization: Anchor Graph Enhancement and Contrastive Propagation
abstract
Multi-view representation is powerful to capture the complex characteristics of real-world data by integrating complementary information from various modalities. However, in many use cases, such as boiler combustion monitoring, factors including sensor sampling frequency, equipment malfunctions, and network delays can lead to temporal asynchrony in data collection. This asynchrony leads misaligned multi-modal data, furthering the difficulty of learning optimal fused representation. To address this misalignment in multi-view data, a body of methods based on autoencoders and non-negative matrix factorization (NMF) have been presented. However, those methods are incapable of jointly exploring the underlying structure inherent in each view as well as the semantic consistency and structural similarity across views. To these ends, we propose a novel sample-bundled optimization for multi-view aligned clustering, which is based on Enhanced Anchor Graph and Contrastive Propagation (termed asEAGCP). We state that there is semantic consistency among intra-class samples (with the same and cross views) and that the global structures across different views demonstrate similarity. By leveraging these associations, we introduce an enhanced anchor graph with learnable sample correlation and a contrastive graph with feature propagation. Specifically, each anchor graphs preserves the semantic relationship among same-view samples, while the contrastive graph propagates feature information across multi-view samples. Experimental results demonstrate the superiority of our method on benchmark datasets by validating its effectiveness in aligning and clustering multi-view data.
Shubin Ma, Zhikui Chen, Lin Wu 0001, Liang Zhao 0005
IEEE Trans. Multim.1
2025 Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search
abstract
Multi-modal representation is faithful and highly effective in describing real-world data samples' characteristics by describing their complementary information. However, the collected data often exhibits incomplete and misaligned characteristics due to factors such as inconsistent sensor frequencies and device malfunctions. Existing research has not effectively addressed the issue of filling missing data in scenarios where multiview data are both imbalanced and misaligned. Instead, it relies on class-level alignment of the available data. Thus, it results in some data samples not being well-matched, thereby affecting the quality of data fusion. In this paper, we propose the Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search(CAPIMAC) to tackle the problem of filling imbalanced and misaligned data in multi-modal datasets. Specifically, we propose a self-repellent greedy anchor search module(SRGASM), which employs a self-repellent random walk combined with a greedy algorithm to identify anchor points for re-representing incomplete and misaligned multi-modal data. Subsequently, based on noise-contrastive learning, we design a consistency-aware padding module (CAPM) to effectively interpolate and align imbalanced and misaligned data, thereby improving the quality of multi-modal data fusion. Experimental results demonstrate the superiority of our method over benchmark datasets. The code will be publicly released at https://github.com/bestow09090/-CAPIMAC.git.
Shubin Ma, Liang Zhao 0005, Mingdong Lu, Bo Xu 0008
IJCAI1
2025 Dual-Learning based Penalized Multi-Align Clustering for Multi-View Incomplete and Disorderly Data
abstract
Multimodal feature fusion, by integrating the complementary information from each modality, can effectively capture complex features in real-world data. However, in many use cases, such as boiler combustion monitoring, factors including equipment failure, inconsistent sensor sampling frequencies, and network delays often cause data collected from different modalities to suffer from missing modality and temporal asynchrony. This leads to the incompleteness and disorderliness of multimodal data. To address these issues, previous studies have proposed several data fusion methods that align the cluster centers before fusion. However, these approaches have two key limitations: 1) they do not guarantee a high alignment accuracy of data pairs at the sample level, and 2) they do not address the issue of significant discrepancies in data sizes across different classes, which impacts the subsequent data fusion performance.
Liang Zhao 0005, Shubin Ma, Bo Xu 0008, Qingchen Zhang 0001
ACM Multimedia2
2025 Pseudo-Label Guided Incomplete Partial View-Aligned Clustering
abstract
Addressing the challenges of incomplete and misaligned data in multi-view learning is critical, giventhe inherent uncertainties and complexities of real-world data collection. These challenges often result in significant discrepancies in the quantity, quality, and completeness of data across different views. However, previous research has predominantly focused on addressing either incompleteness or misalignment in isolation. To address both incompleteness and misalignment simultaneously, we propose a novel model, Pseudo-Label Guided Incomplete Partial View-aligned Clustering (PGIPVC). Specifically, A pseudo-label acquisition module based on Cauchy divergence is proposed to preliminarily train the clustering structure of data in a single view, thereby obtaining pseudo-labels for each sample in the view. Subsequently, an incomplete partial alignment clustering module is designed to obtain discriminative latent representations through contrastive learning with positive-negative pairs selected based on KNN and pseudo-labeled samples. Extensive experiments on benchmark datasets demonstrate the superiority of our method compared to other state-of-the-art approaches.
Shubin Ma, Liang Zhao 0005, Songtao Wu, Bo Xu 0008
IEEE Signal Process. Lett.1
2023 An End-to-End Framework for Partial View-Aligned Clustering with Graph Structure
abstract
Over the last decade, many multi-view clustering (MVC) methods have achieved promising results with intact and completely correct correspondence of multi-view data, which is hard to satisfy in practice leading to the problem of partially view-aligned clustering. In this paper, we propose a novel method to tackle it, termed An End-to-end Framework for Partial View-aligned Clustering with Graph structure(EGPVC). It employs Dykstra’s cyclic constraint projection algorithm to obtain the correspondence between two views. In particular, EGPVC develops an end-to-end framework for partially view-aligned clustering, in which representation learning and clustering process can benefit from each other through the deep embedded clustering layer. Moreover, a cross-view graph regularization term is designed to improve the quality of the learned common representation with graph structure information. Experimental results on several real-world datasets show our promising results comparing with the state-of-the-art methods in partially view-aligned clustering.
Liang Zhao 0005, Qiongjie Xie, Sontao Wu, Shubin Ma
ICASSP4