Zhenglai Li

dblp:251/8321 · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-9151-3631ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Balanced Multi-View Clustering
abstract
Multi-view clustering (MvC) aims to integrate information from different views to enhance the capability of the model in capturing the underlying data structures. The widely used joint training paradigm in MvC potentially does not fully leverage the multi-view information, due to the imbalanced and under-optimized view-specific features caused by the uniform learning objective for all views. For instance, particular views with more discriminative information could dominate the learning process in the joint training paradigm, leading to other views being under-optimized. To alleviate this issue, we first analyze the imbalanced phenomenon in the joint-training paradigm of multi-view clustering from the perspective of gradient descent for each view-specific feature extractor. Then, we propose a novel balanced multi-view clustering (BMvC) method, which introduces a view-specific contrastive regularization (VCR) to modulate the optimization of each view. Concretely, VCR preserves the sample similarities captured from the joint features and view-specific ones into the clustering distributions corresponding to view-specific features to enhance the learning process of view-specific feature extractors. Additionally, an analysis is provided to illustrate that VCR adaptively modulates the magnitudes of gradients for updating the parameters of view-specific feature extractors to achieve a balanced multi-view learning procedure. In such a manner, BMvC achieves a better trade-off between the exploitation of view-specific patterns and the exploration of view-invariance patterns to fully learn the multi-view information for the clustering task. Finally, a set of experiments are conducted to verify the superiority of the proposed method compared with state-of-the-art approaches both on eight benchmark MvC datasets and two spatially resolved transcriptomics datasets.
Zhenglai Li, Jun Wang 0118, Chang Tang, Xinzhong Zhu, Wei Zhang 0049, Xinwang Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Learning Disentangled Representations for Generalized Multi-View Clustering
abstract
Multi-View Clustering (MVC) has gained significant attention for its ability to leverage complementary information across diverse views. However, existing deep MVC methods often struggle with view-distribution entanglement during cross-view fusion, which hampers the quality of the shared latent space and leads to suboptimal clustering performance. To address this issue, we propose the Generalized Multi-view Auto-Encoder (GMAE), a framework designed to preserve cross-view complementarity through disentangled representation learning. Specifically, GMAE employs dual-path autoencoders to decouple source features into view-specific and view-common embeddings, facilitating the discovery of clearer clustering structures. We further construct cross-view adversarial discriminators to guide view-specific encoders in capturing more discriminative features. By strategically modulating mutual information, GMAE effectively aligns distributions and prevents representation collapse, ensuring the generation of robust, non-trivial embeddings. Comprehensive experiments on 13 benchmark datasets demonstrate that GMAE consistently outperforms state-of-the-art methods in both complete and incomplete MVC tasks.
Xin Zou 0001, Ruimeng Liu, Chang Tang, Zhenglai Li, Xinwang Liu 0002, Kunlun He, Wanqing Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Tensor Multi-Rank Constraint Guided Anchor-Wise Adaptive Alignment for Multi-View Clustering
abstract
Anchor graph learning has become a widely used technique for significantly reducing the computational complexity in existing multi-view clustering methods. However, most existing approaches select anchors independently for each view and then generate the consensus graph by directly fusing all anchor graphs. This process overlooks the correspondence between anchor sets across different views, i.e., the column order correspondence of the anchor graphs. To address this limitation, we propose a novel anchor-based tensor multi-rank constraint multi-view clustering method (TMC). Specifically, TMC captures the high-order structural information of the original data by constructing an anchor graph tensor and enforcing a multi-rank constraint to induce a block-diagonal structure. Additionally, to enhance anchor consistency across all view, we construct the anchor graph of each view into an anchor tensor and impose a low-rank constraint on it. In this way, the block-diagonal structure of each anchor graph maintains an approximate alignment between anchors. Furthermore, we provide theoretical proof that the generated anchor graphs inherently exhibit a block-diagonal structure. Extensive experimental results on six multi-view datasets demonstrate that TMC outperforms existing state-of-the-art methods, highlighting its effectiveness in multi-view clustering task.
Jun Wang 0118, Miaomiao Li 0001, Zhenglai Li, Hao Yu 0017, Suyuan Liu, Dayu Hu, Chang Tang, Xinwang Liu 0002
IEEE Trans. Knowl. Data Eng.3
2025 Scalable Cross-View Sample Alignment for Multi-View Clustering with View Structure Similarity
abstract
Most existing multi-view clustering methods aim to generate a consensus partition across all views, based on the assumption that all views share the same sample arrangement. However, in real-world scenarios, the collected data across different views is often unsynchronized, making it difficult to ensure consistent sample correspondence between views. To address this issue, we propose a scalable sample-alignment-based multi-view clustering method, referred to as SSA-MVC. Specifically, we first employ a cluster-label matching (CLM) algorithm to select the view whose clustering labels best match those of the others as the benchmark view. Then, for each of the remaining views, we construct representations of non-aligned samples by computing their similarities with aligned samples. Based on these representations, we build a similarity graph between the non-aligned samples of each view and those in the benchmark view, which serves as the alignment criterion. This alignment criterion is then integrated into a late-fusion framework to enable clustering without requiring aligned samples. Notably, the learned sample alignment matrix can be used to enhance existing multi-view clustering methods in scenarios where sample correspondence is unavailable. The effectiveness of the proposed SSA-MVC algorithm is validated through extensive experiments conducted on eight real-world multi-view datasets.
Jun Wang 0118, Zhenglai Li, Chang Tang, Suyuan Liu, Hao Yu 0017, Chuan Tang, Miaomiao Li 0001, Xinwang Liu 0002
NeurIPS2
2025 Mask-Informed Deep Contrastive Incomplete Multi-View Clustering
abstract
Multi-view clustering (MvC) utilizes information from multiple views to uncover the underlying structures of data. Despite significant advancements in MvC, mitigating the impact of missing samples in specific views on the integration of knowledge from different views remains a critical challenge. This paper proposes a novel Mask-informed Deep Contrastive Incomplete Multi-view Clustering (Mask-IMvC) method, which elegantly identifies a view-common representation for clustering. Specifically, we introduce a mask-informed fusion network that aggregates incomplete multi-view information while considering the observation status of samples across various views as a mask, thereby reducing the adverse effects of missing values. Additionally, we design a prior knowledge-assisted contrastive learning loss that boosts the representation capability of the aggregated view-common representation by injecting neighborhood information of samples from different views. Finally, extensive experiments are conducted to demonstrate the superiority of the proposed Mask-IMvC method over state-of-the-art approaches across multiple MvC datasets, both in complete and incomplete scenarios. The demo code for our work will be publicly available at https://github.com/guanyuezhen/Mask-IMvC.
Zhenglai Li, Yuqi Shi, Xiao He 0010, Chang Tang
IEEE Trans. Circuits Syst. Video Technol.1
2024 Heterogeneous Graph Guided Contrastive Learning for Spatially Resolved Transcriptomics Data
abstract
Spatial transcriptomics provides revolutionary insights into cellular interactions and disease development mechanisms by combining high-throughput gene sequencing and spatially resolved imaging technologies to analyze genes naturally associated with spatially variable tissue genes. However, existing methods typically map aggregated multi-view features into a unified representation, ignoring the heterogeneity and view independence of genes and spatial information. To this end, we construct a heterogeneous Graph guided Contrastive Learning (stGCL) for aggregating spatial transcriptomics data. The method is guided by the inherent heterogeneity of cellular molecules by dynamically coordinating triple-level node attributes through comparative learning loss distributed across view domains, thus maintaining view independence during the aggregation process. In addition, we introduce a cross-view hierarchical feature alignment module employing a parallel approach to decouple spatial and genetic views on molecular structures while aggregating multi-view features according to information theory, thereby enhancing the integrity of inter- and intra-views. Rigorous experiments demonstrate that stGCL outperforms existing methods in various tasks and related downstream applications.
Xiao He 0010, Chang Tang, Xinwang Liu 0002, Chuankun Li, Shan An, Zhenglai Li
ACM Multimedia6
2024 MS-Former: Memory-Supported Transformer for Weakly Supervised Change Detection With Patch-Level Annotations
abstract
Fully supervised change detection methods have achieved significant advancements in performance, yet they depend severely on acquiring costly pixel-level labels. Considering that the patch-level annotations also contain abundant information corresponding to both changed and unchanged objects in bi-temporal images, an intuitive solution is to segment the changes with patch-level annotations. How to capture the semantic variations associated with the changed and unchanged regions from the patch-level annotations to obtain promising change results is the critical challenge for the weakly supervised change detection task. In this paper, we propose a memory-supported transformer (MS-Former), a novel framework consisting of a bi-directional attention block (BAB) and a patch-level supervision scheme (PSS) tailored for weakly supervised change detection with patch-level annotations. More specifically, the BAB captures contexts associated with the changed and unchanged regions from the temporal difference features to construct informative prototypes stored in the memory bank. On the other hand, the BAB extracts useful information from the prototypes as supplementary contexts to enhance the temporal difference features, thereby better distinguishing changed and unchanged regions. After that, the PSS guides the network learning valuable knowledge from the patch-level annotations, thus further elevating the performance. Experimental results on three benchmark datasets demonstrate the effectiveness of our proposed method in the change detection task. The demo code for our work will be publicly available at https://github.com/guanyuezhen/MS-Former.
Zhenglai Li, Chang Tang, Xinwang Liu 0002, Changdong Li, Xianju Li, Wei Zhang 0049
IEEE Trans. Geosci. Remote. Sens.1
2024 Multiple Kernel Clustering With Adaptive Multi-Scale Partition Selection
abstract
Multiple kernel clustering (MKC) enhances clustering performance by deriving a consensus partition or graph from a predefined set of kernels. Despite many advanced MKC methods proposed in recent years, the prevalent approaches involve incorporating all kernels by default to capture diverse information within the data. However, learning from all kernels may not be better than one of a few kernels, particularly since some kernels exhibit a higher proportion of noise than semantic content. Additionally, existing MKC methods, whether based on early-fusion or late-fusion approaches, predominantly rely on pairwise relationships among samples or cluster structures, neglecting potential correlations between these two aspects. To this end, we propose a multiple kernel clustering with an adaptive multi-scale partition selection method (MPS), which exploits multiple-dimensional representations and the pairwise cluster structure for clustering. By the proposed kernel selection framework, potentially harmful kernels are dynamically excluded during the kernel fusion process, and then the multi-scale partitions and similarity graphs derived from the retained kernels are utilized to facilitate the improved consensus partition generation. Finally, extensive experiments are conducted to demonstrate the effectiveness of MPS on eight benchmark datasets.
Jun Wang 0118, Zhenglai Li, Chang Tang, Suyuan Liu, Xinhang Wan, Xinwang Liu 0002
IEEE Trans. Knowl. Data Eng.2
2023 DPNET: Dynamic Poly-attention Network for Trustworthy Multi-modal Classification
abstract
With advances in sensing technology, multi-modal data collected from different sources are increasingly available. Multi-modal classification aims to integrate complementary information from multi-modal data to improve model classification performance. However, existing multi-modal classification methods are basically weak in integrating global structural information and providing trustworthy multi-modal fusion, especially in safety-sensitive practical applications (e.g., medical diagnosis). In this paper, we propose a novel Dynamic Poly-attention Network (DPNET) for trustworthy multi-modal classification. Specifically, DPNET has four merits: (i) To capture the intrinsic modality-specific structural information, we design a structure-aware feature aggregation module to learn the corresponding structure-preserved global compact feature representation. (ii) A transparent fusion strategy based on the modality confidence estimation strategy is induced to track information variation within different modalities for dynamical fusion. (iii) To facilitate more effective and efficient multi-modal fusion, we introduce a cross-modal low-rank fusion module to reduce the complexity of tensor-based fusion and activate the implication of different rank-wise features via a rank attention mechanism. (iv) A label confidence estimation module is devised to drive the network to generate more credible confidence. An intra-class attention loss is introduced to supervise the network training. Extensive experiments on four real-world multi-modal biomedical datasets demonstrate that the proposed method achieves competitive performance compared to other state-of-the-art ones.
Xin Zou 0001, Chang Tang, Zhenglai Li, Xiao He 0010, Shan An, Xinwang Liu 0002
ACM Multimedia4
2023 Mutual structure learning for multiple kernel clustering
Zhenglai Li, Chang Tang, Zhiguo Wan, Kun Sun 0002, Wei Zhang 0049, Xinzhong Zhu
Inf. Sci.1
2023 Lightweight Remote Sensing Change Detection With Progressive Feature Aggregation and Supervised Attention
abstract
Remote sensing change detection (RSCD) aims to explore surface changes from co-registered pair of images. However, the high cost of memory and computation in previous convolutional neural network (CNN)-based methods prevent their successes from being applied to real-world applications. Therefore, we propose a novel lightweight network, which identifies changes based on the features extracted by mobile networks via progressive feature aggregation and supervised attention, termed as A2Net. Considering the less powerful representation capability of mobile networks, we design a neighbor aggregation module (NAM) to fuse features within nearby stages of the backbone to strengthen the representation capability of temporal features. Then, we propose a progressive change identifying module (PCIM) to extract temporal difference information from bitemporal features. Besides, we design a supervised attention module (SAM) to reweight features for effectively aggregating multilevel features from high levels to low levels. With NAM, PCIM, and SAM incorporated, A2Net can achieve favorable results compared with the state-of-the-art methods on three challenging RSCD datasets with fewer parameters (3.78 M) and lower computation costs (6.02 G). The demo code of this work is publicly available athttps://github.com/guanyuezhen/A2Net.
Zhenglai Li, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Jie Dou, Lizhe Wang 0001, Albert Y. Zomaya
IEEE Trans. Geosci. Remote. Sens.1
2023 Unified One-Step Multi-View Spectral Clustering
abstract
Multi-view spectral clustering, which exploits the complementary information among graphs of diverse views to obtain superior clustering results, has attracted intensive attention recently. However, most existing multi-view spectral clustering methods obtain the clustering partitions in a two-step scheme, i.e., spectral embedding and subsequent$k$-means. This two-step scheme inevitably seeks sub-optimal clustering results due to the information loss during the two-steps processes. Besides, existing multi-view spectral clustering methods do not jointly utilize the information of graphs and embedding matrices, which also degrades final clustering results. To solve these issues, we propose a unified one-step multi-view spectral clustering method, which integrates the spectral embedding and$k$-means into a unified framework to obtain discrete clustering labels with a one-step strategy. Under the observation that the inner product of the embedding matrix is a low-rank approximation of the graph, we combine graphs and embedding matrices of different views to obtain a unified graph. Then, we directly capture the discrete clustering indicator matrix from the unified graph. Furthermore, we design an effective optimization algorithm to solve the resultant problem. Finally, a set of experiments on various datasets are conducted to verify the effectiveness of the proposed method. The demo code of this work is publicly available atrgb]0,0,1https://github.com/guanyuezhen/UOMvSC.
Chang Tang, Zhenglai Li, Jun Wang 0118, Xinwang Liu 0002, Wei Zhang 0049, En Zhu
IEEE Trans. Knowl. Data Eng.2
2022 Efficient Multiple Kernel Clustering via Spectral Perturbation
abstract
Clustering is a fundamental task in the machine learning and data mining community. Among existing clustering methods, multiple kernel clustering (MKC) has been widely investigated due to its effectiveness to capture non-linear relationships among samples. However, most of the existing MKC methods bear intensive computational complexity in learning an optimal kernel and seeking the final clustering partition. In this paper, based on the spectral perturbation theory, we propose an efficient MKC method that reduces the computational complexity from O(n3) to O(nk2 + k3), with n and k denoting the number of data samples and the number of clusters, respectively. The proposed method recovers the optimal clustering partition from base partitions by maximizing the eigen gaps to approximate the perturbation errors. An equivalent optimization objective function is introduced to obtain base partitions. Furthermore, a kernel weighting scheme is embedded to capture the diversity among multiple kernels. Finally, the optimal partition, base partitions, and kernel weights are jointly learned in a unified framework. An efficient alternate iterative optimization algorithm is designed to solve the resultant optimization problem. Experimental results on various benchmark datasets demonstrate the superiority of the proposed method when compared to other state-of-the-art ones in terms of both clustering efficacy and efficiency.
Chang Tang, Zhenglai Li, Weiqing Yan, Guanghui Yue 0001, Wei Zhang 0049
ACM Multimedia2
2022 Unified K-means coupled self-representation and neighborhood kernel learning for clustering single-cell RNA-sequencing data
Chang Tang, Zhenglai Li, Wei Zhang 0049, Lijuan Cao
Neurocomputing4
2022 Remote Sensing Change Detection via Temporal Feature Interaction and Guided Refinement
abstract
Remote sensing change detection (RSCD), which identifies the changed and unchanged pixels from a registered pair of remote sensing images, has enjoyed remarkable success recently. However, locating changed objects with fine structural details is still a challenging problem in RSCD. In this paper, we propose a novel remote sensing change detection network via temporal feature interaction and guided refinement (TFI-GR) to solve this issue. Specifically, unlike previous methods, which just employ one single concatenation or subtraction operation for bi-temporal feature fusion, we design a temporal feature interaction module (TFIM) to enhance interaction between bi-temporal features and capture temporal difference information at diverse feature levels. Afterword, a guided refinement modules (GRM), which aggregates both low- and high-level temporal difference representations to polish the location information of high-level features and filter the background clutters of low-level features, is repeatedly performed. Finally, the multi-level temporal difference features are progressively fused to generate change maps for change detection. To demonstrate the effectiveness of the proposed TFI-GR, comprehensive experiments are performed on three high spatial resolution remote sensing change detection datasets. Experimental results indicate that the proposed method is superior to other state-of-the-art change detection methods. The demo code of this work is publicly available at https://github.com/guanyuezhen/TFI-GR.
Zhenglai Li, Chang Tang, Lizhe Wang 0001, Albert Y. Zomaya
IEEE Trans. Geosci. Remote. Sens.1
2022 High-Order Correlation Preserved Incomplete Multi-View Subspace Clustering
abstract
Incomplete multi-view clustering aims to exploit the information of multiple incomplete views to partition data into their clusters. Existing methods only utilize the pair-wise sample correlation and pair-wise view correlation to improve the clustering performance but neglect the high-order correlation of samples and that of views. To address this issue, we propose a high-order correlation preserved incomplete multi-view subspace clustering (HCP-IMSC) method which effectively recovers the missing views of samples and the subspace structure of incomplete multi-view data. Specifically, multiple affinity matrices constructed from the incomplete multi-view data are treated as a third-order low rank tensor with a tensor factorization regularization which preserves the high-order view correlation and sample correlation. Then, a unified affinity matrix can be obtained by fusing the view-specific affinity matrices in a self-weighted manner. A hypergraph is further constructed from the unified affinity matrix to preserve the high-order geometrical structure of the data with incomplete views. Then, the samples with missing views are restricted to be reconstructed by their neighbor samples under the hypergraph-induced hyper-Laplacian regularization. Furthermore, the learning of view-specific affinity matrices as well as the unified one, tensor factorization, and hyper-Laplacian regularization are integrated into a unified optimization framework. An iterative algorithm is designed to solve the resultant model. Experimental results on various benchmark datasets indicate the superiority of the proposed method. The code is implemented by using MATLAB R2018a and MindSpore library: https://github.com/ChangTang/HCP-IMSC.
Zhenglai Li, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, En Zhu
IEEE Trans. Image Process.1
2022 Consensus Graph Learning for Multi-View Clustering
abstract
Multi-view clustering, which exploits the multi-view information to partition data into their clusters, has attracted intense attention. However, most existing methods directly learn a similarity graph from original multi-view features, which inevitably contain noises and redundancy information. The learned similarity graph is inaccurate and is insufficient to depict the underlying cluster structure of multi-view data. To address this issue, we propose a novel multi-view clustering method that is able to construct an essential similarity graph in a spectral embedding space instead of the original feature space. Concretely, we first obtain multiple spectral embedding matrices from the view-specific similarity graphs, and reorganize the gram matrices constructed by the inner product of the normalized spectral embedding matrices into a tensor. Then, we impose a weighted tensor nuclear norm constraint on the tensor to capture high-order consistent information among multiple views. Furthermore, we unify the spectral embedding and low rank tensor learning into a unified optimization framework to determine the spectral embedding matrices and tensor representation jointly. Finally, we obtain the consensus similarity graph from the gram matrices via an adaptive neighbor manner. An efficient optimization algorithm is designed to solve the resultant optimization problem. Extensive experiments on six benchmark datasets are conducted to verify the efficacy of the proposed method. The code is implemented by using MATLAB R2018a and MindSpore library[1]:https://github.com/guanyuezhen/CGL.
Zhenglai Li, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, En Zhu
IEEE Trans. Multim.1
2021 Tensor-Based Multi-View Block-Diagonal Structure Diffusion for Clustering Incomplete Multi-View Data
abstract
In this paper, we propose a novel incomplete multi-view clustering method, in which a tensor nuclear norm regularizer elegantly diffuses the information of multi-view block-diagonal structure across different views. By exploring the membership between observed and missing samples and that between missing ones in each incomplete view with the guidance of the high-order view consistency, a global block-diagonal structure is well preserved in multiple spectral embedding matrices. Meanwhile, a consensus representation with strong separability is obtained for clustering. An iterative algorithm based on Augmented Lagrange Multiplier (ALM) is designed to solve the resultant model. Experimental results on six benchmark datasets indicate the superiority of the proposed method. http://github.com/ChangTang/TMBSD
Zhenglai Li, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, En Zhu
ICME1
2019 Diversity and consistency learning guided spectral embedding for multi-view clustering
Zhenglai Li, Chang Tang, Jiajia Chen 0010, Weiqing Yan, Xinwang Liu 0002
Neurocomputing1