Jiazhen Hong

dblp:311/3583 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
0009-0003-3098-4012ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
1.012026
A Geometric Approach to $k$k-Means Clustering · IEEE Trans. Knowl. Data Eng. 2026
Data mining › clustering
k-means clustering
1.012026
A Geometric Approach to $k$k-Means Clustering · IEEE Trans. Knowl. Data Eng. 2026

Methods — techniques the papers use, named apart from their topics

non-local operations · 1.0geometric analysis · 1.0
YearPublicationVenuePosition
2026 A Geometric Approach to $k$k-Means Clustering
abstract
$k$-meansclustering is a fundamental problem in many scientific and engineering domains. The optimization problem associated with$k$-means clustering is nonconvex, for which standard algorithms are only guaranteed to find a local optimum. Leveraging the hidden structure of local solutions, we propose a general algorithmic framework for escaping undesirable local solutions and recovering the global solution or the ground truth clustering. This framework consists of iteratively alternating between two steps: (i) detect mis-specified clusters in a local solution, and (ii) improve the local solution by non-local operations. We discuss specific implementation of these steps, and elucidate how the proposed framework unifies many existing variants of$k$-means algorithms through a geometric perspective. We also present two natural variants of the proposed framework, where the initial number of clusters may be over- or under-specified. We provide theoretical justifications and extensive experiments to demonstrate the efficacy of the proposed approach.
Jiazhen Hong, Yudong Chen 0001
IEEE Trans. Knowl. Data Eng.1