Yan Chen 0036

dblp:88/2827-36 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-3898-2399ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 P-MOEBC: A Pairwise Evolutionary Framework for Balanced Clustering
abstract
Balanced clustering aims to partition data into cohesive groups while satisfying size constraints. It is important in applications such as load balancing, resource allocation, and capacity-aware data analysis. In practice, structural quality and size balance often compete, which makes balanced clustering difficult for methods that commit to a single operating point. To address this issue, we propose Pairwise Multi-Objective Evolutionary Balanced Clustering (P-MOEBC), a permutation-invariant evolutionary framework that searches for a diverse set of structure-balance trade-offs. We formulate the task with a pairwise graph-cut objective and a flexible balance-violation objective, and then design two label-free operators tailored to this formulation: (1) a Consensus Block Crossover that recombines reliable co-membership structures without label alignment, and (2) a Balance-Aware Mutation with exact local updates that evaluates candidate moves in O(knn) time. Experiments on synthetic data show that scalarized baselines can be sensitive to penalty selection, while real-world benchmarks show that P-MOEBC is competitive across strict and relaxed feasibility regimes and remains scalable. Code is available at https://github.com/zjh308/P-MOEBC.
Jianhao Zhu, Yunhui Liang, Yan Chen 0036, Peng Zhou 0006, Liang Du 0003
GECCO3
2026 FairFBC: Scalable Fair Fuzzy Clustering via Group-Balanced Anchor Graphs
abstract
Clustering is widely used to organize large-scale multimedia collections, but standard clustering methods can inherit and amplify demographic imbalance in the underlying data. Existing fair clustering methods still face two practical limitations: many scale poorly to high-dimensional visual datasets, and many rely on a fairness-weight coefficient that must be tuned for each dataset. We propose FairFBC, a scalable fair fuzzy clustering framework for large-scale data. FairFBC first constructs a group-balanced anchor graph to obtain a sparse and scalable representation. It then learns a shared fuzzy partition by maximizing the sum of group-specific trace-sqrt quality terms, which encourages balanced structure across sensitive groups without introducing an explicit fairness-accuracy trade-off coefficient. Finally, a confidence-aware quota-constrained assignment step converts fuzzy memberships into a discrete partition while preserving consistency with the learned structure. Experiments on ten datasets from vision, vision-language, and tabular domains, including FairFace, CelebA, and ChestX-ray, show that FairFBC achieves a strong accuracy-fairness trade-off while remaining stable on large datasets where several competitive baselines fail to run. The code for our method is publicly available at https://github.com/Whale-Waves/FairFBC.
Tongzheng Zhao, Yan Chen 0036, Peng Zhou 0006, Liang Du 0003
ICMR2
2026 Clustering Ensembles: A Data Perspective Survey
Wenjun He, Yan Chen 0036, Peng Zhou 0006, Liang Du 0003
PAKDD (4)2
2025 Sharper Error Bounds in Late Fusion Multi-view Clustering with Eigenvalue Proportion Optimization
abstract
Multi-view clustering (MVC) aims to integrate complementary information from multiple views to enhance clustering performance. Late Fusion Multi-View Clustering (LFMVC) has shown promise by synthesizing diverse clustering results into a unified consensus. However, current LFMVC methods struggle with noisy and redundant partitions and often fail to capture high-order correlations across views. To address these limitations, we present a novel theoretical framework for analyzing the generalization error bounds of multiple kernel k-means, leveraging local Rademacher complexity and principal eigenvalue proportions. Our analysis establishes a convergence rate of O(1/n), significantly improving upon the existing rate in the order of O(sqrt(k/n)). Building on this insight, we propose a low-pass graph filtering strategy within a multiple linear K-means framework to mitigate noise and redundancy, further refining the principal eigenvalue proportion and enhancing clustering accuracy. Experimental results on benchmark datasets confirm that our approach outperforms state-of-the-art methods in clustering performance and robustness.
Liang Du 0003, Henghui Jiang, Yiqing Guo, Yan Chen 0036, Feijiang Li, Peng Zhou 0006
AAAI5
2025 Consensus Graph Filter Learning for Multiple Graph Clustering
abstract
Multi-view Clustering (MVC) has gained significant attention for its ability to utilize consistent and complementary information from multiple views. Graph filter-based MVC methods have recently demonstrated promising performance, attracting growing interest. However, existing graph filter-based methods typically rely on a single filter for each view. These filters are usually derived from either a specific view or a consensus graph across all views, which limits their effectiveness in fully integrating multi-view information. To address this limitation, we propose a novel method, Consensus Graph Filter Learning for Multiple Graph Clustering (CGFMVC). Unlike existing methods that rely on a single filter per view, we construct multiple graph filters by integrating information from all views. For each view, we generate intra-view graph filters to intermediate high-order information and bridge local-global structures. This approach facilitates neighborhood smoothing and preserves local consistency within each view. By learning a consensus graph filter from these multiple filters, we effectively capture global complementary information across views while preserving local consistency. CGFMVC demonstrates excellent efficiency and effectiveness on various benchmark datasets, outperforming state-of-the-art methods. The code is publicly available at https://github.com/Sean-zjl/CGFMVC.
Jiale Zou, Yan Chen 0036, Peng Zhou 0006, Liang Du 0003
ICASSP2
2025 Scalable Multi-View Clustering via Bipartite Graph Consensus Filtering
abstract
As data sources and modalities become more diverse, existing multi-view graph clustering methods face high computational complexity, hindering scalability. Bipartite graph clustering overcomes this by using anchors to build association graphs, reducing complexity to linear time. However, the quality of bipartite graphs and the neglect of higher-order correlations remain significant limitations. To tackle these challenges, we propose MCBGF, a scalable bipartite graph-based multi-view clustering method enhanced by consensus graph filtering. By integrating the original bipartite structure to counteract degradation from over-smoothing in high-order filters, MCBGF achieves robust and consistent performance. With linear time complexity, MCBGF efficiently processes large-scale datasets and consistently outperforms state-of-the-art methods in experimental evaluations. The code has been released at https://github.com/sxuHui/MCBGF.
Henghui Jiang, Yiqing Guo, Yan Chen 0036, Liang Du 0003
ICIP3
2025 Scalable Multi-Kernel Clustering with Dynamic Procrustes
abstract
Multi-Kernel Clustering (MKC) uses several kernels or graphs to boost clustering performance. Traditional early fusion methods try to combine these into one consensus structure, but this often ends up with an averaging effect that weakens the final clustering result and depends on heavy matrix operations. In contrast, late fusion methods create separate clustering results for each kernel and then merge them, yet they suffer from high computational costs and do not make full use of the complementary information among kernels. Many researchers have assumed that generating separate partitions for every kernel at each iteration is too expensive because of the intensive matrix computations required. To overcome these issues, we introduce a new framework called Scalable MKC with Dynamic Procrustes (SMKCDP), which learns both the individual base partitions and the final consensus clustering at the same time. SMKCDP not only avoids the dilution seen in early fusion but also improves the quality of the base partitions in late fusion. An interesting finding in our approach is that the optimization only needs the more efficient Singular Value Decomposition (SVD) of rectangular matrices, which greatly reduces the overall computational burden. Experiments on benchmark datasets show that SMKCDP significantly improves both clustering accuracy and efficiency, opening up new directions for scalable MKC methods. The code has been released at https://github.com/WuLizhu-SXU/SMKCDP.
Lizhu Wu, Yan Chen 0036, Peng Zhou 0006, Liang Du 0003
ICME2
2025 Balanced Multiple Kernel Clustering with Discrete Partition Entropy Auto Regularization
abstract
Clustering, a fundamental task in machine learning and data mining, is essential for uncovering patterns by grouping data points with similar characteristics. Traditional methods struggle with nonlinear data structures, but kernel-based approaches alleviate this issue by mapping data to high dimensional spaces. Multiple Kernel Clustering (MKC) further improves clustering by automating kernel selection and integration. However, MKC faces challenges related to kernel graph quality, information loss during relax-and-discretization, neglect of balanced clustering constraints, and the trade-off between high clustering quality and balance. To address these challenges, we introduce Balanced Multiple Kernel Clustering (BMKC). BMKC utilizes local kernel reconstruction and advanced high-order diffusion techniques for comprehensive kernel graph learning. It directly learns a discrete partition matrix using a robust L1-induced local reconstruction criterion, eliminating the two step process. BMKC incorporates an automatic mechanism for trade-off control between clustering and balance, supported by a versatile optimization algorithm accommodating various balance regularization choices. Experimental validation demonstrates the superior performance of MKC on benchmarks data sets, showcasing its effectiveness. The code for our method is publicly available at https://github.com/ChenYan01TYUT/BMKC-ACM-MM-2025.
Yan Chen 0036, Bingbing Jiang 0001, Peng Zhou 0006, Lei Duan, Liang Du 0003
ACM Multimedia1
2025 Robust Tensor Learning with Graph Diffusion for Scalable Multi-view Graph Clustering
abstract
The rapid proliferation of multi-view data has necessitated robust and scalable clustering techniques capable of capturing complex, high-dimensional patterns. While Multi-view Bipartite Graph Clustering (MVBGC) has shown promising results, existing approaches often overlook that the generated bipartite graph is susceptible to disturbances from complex structures and noise. To address these challenges, we propose RTGD-MVC, a novel framework for Robust Tensor Learning with Graph Diffusion tailored for efficient and scalable multi-view graph clustering. RTGD-MVC integrates a graph diffusion mechanism to suppress noise propagation and employs cross-view diffusion to enhance global consistency while capturing complementary information across views. Additionally, a non-convex Tensor Exponential Norm (TEN) is introduced as a tighter surrogate for the tensor rank, enabling the learning of more discriminative and noise-robust representations. By embedding these components into a unified optimization model with linear computational complexity, RTGD-MVC achieves both theoretical efficiency and practical scalability. Extensive experiments on diverse benchmark datasets demonstrate that RTGD-MVC significantly outperforms state-of-the-art methods, highlighting its superior ability to capture intricate multi-view correlations and structural patterns.
Jiale Zou, Yan Chen 0036, Bingbing Jiang 0001, Peng Zhou 0006, Liang Du 0003, Lei Duan
ACM Multimedia2
2025 Late Fusion Multiple Kernel Clustering Refined via Optimal Linear Graph Filtering
Henghui Jiang, Yiqing Guo, Yan Chen 0036, Liang Du 0003
ECML/PKDD (1)3
2025 A scalable Consensus Fast Graph Filtering approach for late fusion multi-view clustering
Yiqing Guo, Henghui Jiang, Yan Chen 0036, Liang Du 0003
Signal Process.3
2024 Higher Order Multiple Graph Filtering for Structured Graph Learning
abstract
In the field of machine learning, multi-view clustering aims to reveal hidden clustering patterns across different data perspectives. However, traditional methods often struggle due to their reliance on low-order similarity data. To overcome this, we propose a new approach that integrates the learning of multiple graph filters, approximated through Chebyshev polynomials, with consensus structural graph learning into a unified framework. This method fully utilizes high-order statistical information from multiple data sources, thereby enhancing multi-view clustering. Comprehensive experiments conducted on multiple datasets consistently demonstrate significant performance improvements over traditional methods. Code is available at https://github.com/lxd1204/HMGC.
Liang Du 0003, Yan Chen 0036, Gui Yang, Mian Ilyas Ahmad, Peng Zhou 0006
ICASSP3
2024 Fast and Scalable Incomplete Multi-View Clustering with Duality Optimal Graph Filtering
abstract
Incomplete Multi-View Clustering (IMVC) is crucial for multi-media data analysis. While graph learning-based IMVC methods have shown promise, they still have limitations. The prevalent first-order affinity graph often misclassifies out-neighborhood intra-cluster and in-neighbor inter-cluster samples, worsened by data incompleteness. These inaccuracies, combined with high computational demands, restrict their suitability for large-scale IMVC tasks. To address these issues, we propose a novel Fast and Scalable IMVC with duality Optimal graph Filtering (FSIMVC-OF). Specifically, we refine the clustering-friendly structure of the bipartite graph by learning an optimal filter within a consensus clustering framework. Instead of learning a sample-side filter, we optimize an anchor-side graph filter and apply it to the anchor side, ensuring computational efficiency with linear complexity, supported by the provable equivalence between these two types of graph filters. We present an alternative optimization algorithm with linear complexity. Extensive experimental analysis demonstrates the superior performance of FSIMVC-OF over current IMVC methods. The codes of this article are released in https://github.com/sroytik/FSIMVC-OF.
Liang Du 0003, Yukai Shi, Yan Chen 0036, Peng Zhou 0006
ACM Multimedia3
2024 Heat Kernel Diffusion for Enhanced Late Fusion Multi-View Clustering
abstract
Recent advancements in Multi-view Clustering (MVC) have highlighted the benefits of late fusion techniques. However, existing late fusion-based MVC (LFMVC) approaches often struggle with intrinsic noise and redundancy within base clustering embeddings generated by traditional method, and fail to capture higher-order correlations among samples and across views. We propose a novel method that incorporates an optimal consensus heat kernel-induced graph filter to address these issues. Our approach leverages diffusion processes to construct filters that capture high-order information, achieving neighborhood smoothing while maintaining local consistency within each view. By integrating a consensus graph filter with downstream clustering tasks within a unified learning framework, our method enhances performance through mutual reinforcement. In particular, our method does not require additional hyperparameter tuning, simplifying the optimization process. Experimental results on various benchmark datasets demonstrate the superior effectiveness of our approach, outperforming state-of-the-art methods in terms of clustering performance.
Gui Yang, Jiale Zou, Yan Chen 0036, Liang Du 0003, Peng Zhou 0006
IEEE Signal Process. Lett.3
2022 A Trace Ratio Maximization Method for Parameter Free Multiple Kernel Clustering
Yan Chen 0036, Liang Du 0003, Lei Duan
DASFAA (2)1