Geping Yang

dblp:319/3574 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-1403-3324ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 EVA-MVC: Equitable View-weight Allocation for Generic Multi-View Clustering
abstract
Contemporary datasets sourced from the web often adopt a multiview format, collecting data from diverse sources, domains, or modules.Existing methodologies employed to analyze such datasets frequently overlook or inaccurately allocate the view-weights, pivotal metrics reflecting each view's significance.This work introduces EVA-MVC, a simple yet effective algorithm designed for Equitable View-weight Allocation (EVA) seamlessly integrated with arbitrary Multi-view Clustering (MVC) methods.Within the EVA module, we establish theoretical connections between view supplementarity and Multi-view Subspace Learning (MSL), leading to the partition of views into View Communities (VCs) based on these foundational principles.These VCs exhibit internal supplementarity similarities, facilitating Equitable View-weights Allocation through VCspecific MSL.The proposed EVA process precedes and operates independently of traditional or SOTA MVC approaches, requiring no additional processing or specialized design, making it an ideal preprocessing step for MVC applications.Through comprehensive evaluations across diverse multi-view datasets, our findings reveal that our EVA significantly enhances the effectiveness of mainstream MVC frameworks, resulting in a notable performance improvement.
Yuan Fang 0001, Xiaofeng Feng, Geping Yang, Ruichu Cai, Yiyang Yang, Zhiguo Gong, Zhifeng Hao 0004
WWW3
2025 MSC-DOLES: Multi-View Subspace Clustering in Diverse Orthogonal Latent Embedding Spaces
abstract
In the domain of Multi-view Subspace Clustering (MSC) in Latent Embedding Space (LES), existing methods aim to capture and leverage critical multi-view information by mapping it into a low-dimensional LES. However, several aspects can be further improved: (i) Fusion Strategy: Existing methods adopt either early fusion or late fusion to integrate multi-view information, limiting the effectiveness of the fusion. (ii) Diversity: Current methods often overlook the inherent diversity in the multi-view data by focusing on a single LES. (iii) Efficiency: LES-based methods exhibit high computational complexity, with cubic time and quadratic space requirements based on the number of samples. To address these issues, we propose a novel framework called MSC-DOLES (Multi-view Subspace Clustering in Diverse Orthogonal Latent Embedding Spaces), a novel framework designed to tackle these challenges. MSC-DOLES incorporates a two-stage fusion approach that generates and learns from multiple LES to maximize cross-view diversity. Orthogonality constraints on individual LES ensure view-internal diversity, resulting in a set of Diverse Orthogonal Latent Embedding Spaces (DOLES). The DOLES are then fused into a consensus anchor graph using learnable anchors. The final clustering is induced by partitioning the obtained graph without pre-processing. We develop an eight-step optimization algorithm for MSC-DOLES, which exhibits nearly linear time and space complexities relative to the number of samples. Extensive experiments demonstrate the superiority of MSC-DOLES over state-of-the-art methods.
Yuan Fang 0001, Geping Yang, Ruichu Cai, Yiyang Yang, Zhiguo Gong, Zhifeng Hao 0004
IEEE Trans. Knowl. Data Eng.2
2025 SPGMVC: Multiview Clustering via Partitioning the Signed Prototype Graph
abstract
Multiview clustering (MVC) has been widely studied in machine learning and data mining for its capability of improving clustering performance by fusing the information from multiview data. In the past decade, a large number of MVC methods have made impressive progress, but most of them suffer from computational burdens, especially in large-scale tasks. Binary MVC (BMVC) is proposed to address this issue by representing the large-scale high-dimensional dataset as a group of consensus and low-dimensional binary codes. However, current BMVC-based approaches generate the clustering by executing binary k-means on the obtained binary codes, which fail to capture the embedded geometric information, leading to poor clustering performance. In addition, parameter selection is another "mission impossible" in unsupervised learning tasks including MVC. To tackle these challenges, a framework of multiview clustering via partitioning the signed prototype graph (SPGMVC) is proposed in this work. The SPGMVC framework offers several contributions. First, SPGMVC is designed as a unified framework for MVC. It combines effective technologies, such as consensus binary coding, code compression (CC), signed prototype graph (SPG) partitioning, and prototype-based cluster assignment. Second, SPGMVC partitions the signed graph (SG) based on the relationships between positive and negative edges. By capturing the underlying structure of the data, this partitioning strategy improves clustering accuracy (ACC). CC techniques are applied to reduce the graph's scale, enabling further partitioning and enhancing computational efficiency. Third, SPGMVC employs an alternate minimizing strategy to efficiently handle the optimization problem. This strategy has nearly linear time and space complexity with respect to the data volume, making it suitable for large-scale tasks. Fourth, SPGMVC proposes an automatic parameter selection strategy, eliminating the need for extensive parameter exploration. Comprehensive experiments illustrate the superiority of our model. The implementation of SPGMVC is available at: https://github.com/gepingyang/PSGMVC.
Geping Yang, Shusen Yang, Yiyang Yang, Xiang Chen 0007, Zhiguo Gong, Zhifeng Hao 0004
IEEE Trans. Neural Networks Learn. Syst.1
2024 QFINCH: Quick hierarchical clustering using k-means and first neighbor relations
abstract
This paper introduces QFINCH, a fast hierarchical clustering framework that leverages k-means and first neighbor relations of samples. QFINCH achieves a computational complexity of $\mathcal{O}(N\log (N))$ without the use of any index technology. We efficiently utilize k-means to construct an initial coarsened partition and employ the centers of the partitions as input for the first nearest neighbor merging. Additionally, we assign the true labels to samples in the coarse partition. Through experimental validation, we demonstrate the effectiveness of our algorithm and show that it significantly reduces time consumption. Finally, we assess the algorithm’s performance in various real-world large-scale datasets.
Geping Yang, Yiyang Yang
CSCWD2
2024 Landmark-based k-factorization multi-view subspace clustering
Geping Yang, Zhiguo Gong, Yiyang Yang
Inf. Sci.2
2024 UP-DPC: Ultra-scalable parallel density peak clustering
Geping Yang, Yiyang Yang, Xiang Chen 0007, Zhiguo Gong, Zhifeng Hao 0004
Inf. Sci.2
2024 Module-based graph pooling for graph classification
Sucheng Deng, Geping Yang, Yiyang Yang, Zhiguo Gong, Xiang Chen 0007, Zhifeng Hao 0004
Pattern Recognit.2
2023 RESKM: A General Framework to Accelerate Large-Scale Spectral Clustering
Geping Yang, Sucheng Deng, Xiang Chen 0007, Yiyang Yang, Zhiguo Gong, Zhifeng Hao 0004
Pattern Recognit.1
2023 LiteWSEC: A Lightweight Framework for Web-Scale Spectral Ensemble Clustering
abstract
Spectral Clustering (SC) is an effective clustering method for its excellent performance in partitioning non-linearly distributed data. On the other hand, Ensemble Clustering (EC), a different clustering technology, can promote cluster quality by ensembling the results of base clusterings. In this work, we concentrate on an EC framework that utilizes SC as the base method. Nevertheless, SC suffers from scalability due to its high computational complexity in constructing the Laplacian graph and computing the corresponding eigendecomposition. In the past decades, many efforts have been made to it. However, SC suffers from the scalability issue in processing extensive data, especially in web-scale scenarios. Additionally, EC requires multiple clustering results as the ensemble bases, which further aggravates resource consumption. To address this issue, LiteWSEC, a simple yet efficient Lightweight Framework for Web-scale Spectral Ensemble Clustering, is proposed to cluster web-scale data with limited resource requirements. It adopts the Web-scale Spectral Clustering (WSC) as the base method, which has minimal space overhead without computing overall embedding explicitly. LiteWSEC is highly flexible in the memory requirement, which is adaptive to the available resource. It can partition web-scale data (e.g.,$n $= 8,000 k) in an resource-limited host (e.g., memory is restricted to 1 GB). Experiments on real-world, large-scale, and web-scale datasets demonstrate both the efficiency and effectiveness of LiteWSEC over state-of-the-art SC and EC methods.
Geping Yang, Sucheng Deng, Yiyang Yang, Zhiguo Gong, Xiang Chen 0007, Zhifeng Hao 0004
IEEE Trans. Knowl. Data Eng.1
2022 LiteWSC: A Lightweight Framework for Web-Scale Spectral Clustering
Geping Yang, Sucheng Deng, Yiyang Yang, Zhiguo Gong, Xiang Chen 0007, Zhifeng Hao 0004
DASFAA (2)1
2022 FastDEC: Clustering by Fast Dominance Estimation
Geping Yang, Hongzhang Lv, Yiyang Yang, Zhiguo Gong, Xiang Chen 0007, Zhifeng Hao 0004
ECML/PKDD (1)1