VLDB 2026 Research / reviewers in the wild / expert
Georgios Vardakas
dblp:322/4052
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-1352-2062ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Clustering Using the Soft Silhouette Score: Towards Compact and Well-Separated ClustersabstractAbstract Unsupervised learning has gained prominence in the big data era, offering a means to extract valuable insights from unlabeled datasets. Deep clustering has emerged as an important unsupervised category, aiming to exploit the non-linear mapping capabilities of neural networks in order to enhance clustering performance. The majority of deep clustering literature focuses on minimizing the inner-cluster variability in some embedded space while keeping the learned representation consistent with the original high-dimensional dataset. In this work, we propose soft silhouette , a probabilistic formulation of the silhouette coefficient. Soft silhouette rewards compact and distinctly separated clustering solutions such as the conventional silhouette coefficient. When optimized within a deep clustering framework, soft silhouette guides the learned representations towards forming compact and well-separated clusters. In addition, we introduce an autoencoder-based deep learning architecture that is suitable for optimizing the soft silhouette objective function. The proposed deep clustering method has been tested and compared with several well-studied deep clustering methods on various benchmark datasets, yielding very satisfactory clustering results. Georgios Vardakas, Ioannis Papakostas, Aristidis Likas |
Mach. Learn. | 1 |
| 2026 | UniForCE: The Unimodality Forest method for Clustering and Estimation of the number of clustersabstractEstimating the number of clusters k while clustering the data is a challenging task. An incorrect cluster assumption indicates that the number of clusters k gets wrongly estimated. Consequently, the model fitting becomes less important. In this work, we focus on the concept of unimodality and propose a flexible cluster definition called locally unimodal cluster . A locally unimodal cluster extends for as long as unimodality is locally preserved across pairs of subclusters of the data. Then, we propose the UniForCE method for locally unimodal clustering. The method starts with an initial overclustering of the data and relies on the unimodality graph that connects subclusters forming unimodal pairs. Such pairs are identified using an appropriate statistical test. UniForCE identifies maximal locally unimodal clusters that are statistically significant by computing a spanning forest in the unimodality graph. Experimental results on both real and synthetic datasets illustrate that the proposed methodology is particularly flexible and robust in discovering regular and highly complex cluster shapes. Most importantly, it automatically provides an adequate estimation of the number of clusters. Georgios Vardakas, Argyris Kalogeratos, Aristidis Likas |
Pattern Recognit. | 1 |
| 2025 | Evaluating Clustering Quality in Centroid-Based Clustering Using Counterfactual Distances
Georgios Vardakas, Antonia Karra, Evaggelia Pitoura, Aristidis Likas |
DS | 1 |
| 2025 | Counterfactual Explanations for k-Means and Gaussian ClusteringabstractCounterfactuals have been recognized as an effective approach to explain classifier decisions. In this work, we focus on the use of counterfactuals to explain clustering solutions. First, we present a general definition for counterfactuals for model-based clustering that includes plausibility and feasibility constraints. Then we consider the counterfactual generation problem for$k$-means and Gaussian clustering assuming Euclidean distance. Our approach takes as input the factual, the target cluster, a binary mask indicating actionable or immutable features and a plausibility factor specifying how far from the cluster boundary the counterfactual should be placed. In the$k$-means clustering case, analytical mathematical formulas are presented for computing the optimal solution, while in the Gaussian clustering case (assuming full, diagonal, or spherical covariances) our method requires the numerical solution of a nonlinear equation with a single parameter only. We demonstrate the advantages of our approach through illustrative examples and quantitative experimental comparisons. Georgios Vardakas, Antonia Karra, Evaggelia Pitoura, Aristidis Likas |
ICTAI | 1 |
| 2025 | Efficient error minimization in kernel k-means clusteringabstractAbstract Kernel k -means extends the k -means algorithm to identify non-linearly separable clusters but is inherently sensitive to cluster initialization. To address this challenge, we first formulate the kernel k-means ++ method, which conveys the efficient center initialization strategy of k -means++ from Euclidean to kernel space. Building on this, we propose global kernel k-means ++ ( $$\text {GK}k\text {M}$$ ++), a novel clustering algorithm designed to balance clustering error minimization with reduced computational cost. $$\text {GK}k\text {M}$$ ++ extends the well-established global kernel k -means algorithm by incorporating the stochastic initialization strategy of kernel k -means++. This approach significantly reduces computational complexity while preserving superior clustering error minimization capabilities akin to traditional global kernel k -means. The experimental results on synthetic, real, and graph datasets indicate that $$\text {GK}k\text {M}$$ ++ consistently outperforms both kernel k -means with random initialization and kernel k -means++, while achieving solutions comparable to those provided by the exhaustive and computational intensive global kernel k -means method. Georgios Vardakas, Ioannis Papakostas, Aristidis Likas |
Pattern Anal. Appl. | 1 |
| 2024 | Revisiting Silhouette Aggregation
John Pavlopoulos, Georgios Vardakas, Aristidis Likas |
DS (1) | 2 |
| 2024 | Global k-means++: an effective relaxation of the global k-means clustering algorithm
Georgios Vardakas, Aristidis Likas |
Appl. Intell. | 1 |
| 2024 | Explainable dating of greek papyri imagesabstractAbstract Greek literary papyri, which are unique witnesses of antique literature, do not usually bear a date. They are thus currently dated based on palaeographical methods, with broad approximations which often span more than a century. We created a dataset of 242 images of papyri written in “bookhand” scripts whose date can be securely assigned, and we used it to train algorithms for the task of dating, showing its challenging nature. To address data scarcity, we extended our dataset by segmenting each image into its respective text lines. By using the line-based version of our dataset, we trained a Convolutional Neural Network, equipped with a fragmentation-based augmentation strategy, and we achieved a mean absolute error of 54 years. The results improve further when the task is cast as a multi-class classification problem, predicting the century. Using our network, we computed precise date estimations for papyri whose date is disputed or vaguely defined, employing explainability to understand dating-driving features. John Pavlopoulos, Maria Konstantinidou, Elpida Perdiki, Isabelle Marthot-Santaniello, Holger Essler, Georgios Vardakas, Aristidis Likas |
Mach. Learn. | 6 |
| 2023 | Explaining the Chronological Attribution of Greek Papyri Images
John Pavlopoulos, Maria Konstantinidou, Georgios Vardakas, Isabelle Marthot-Santaniello, Elpida Perdiki, Dimitris Koutsianos, Aristidis Likas, Holger Essler |
DS | 3 |
| 2023 | Neural clustering based on implicit maximum likelihoodabstractAbstract Clustering is one of the most fundamental unsupervised learning tasks with numerous applications in various fields. Clustering methods based on neural networks, called deep clustering methods, leverage the representational power of neural networks to enhance clustering performance. ClusterGan constitutes a generative deep clustering method that exploits generative adversarial networks (GANs) to perform clustering. However, it inherits some deficiencies of GANs, such as mode collapse, vanishing gradients and training instability. In order to tackle those deficiencies, the generative approach of implicit maximum likelihood estimation (IMLE) has been recently proposed. In this paper, we present a clustering method based on generative neural networks, called neural implicit maximum likelihood clustering, which adopts ideas from both ClusterGAN and IMLE. The proposed method has been compared with ClusterGAN and other neural clustering methods on both synthetic and real datasets, demonstrating promising results. Georgios Vardakas, Aristidis Likas |
Neural Comput. Appl. | 1 |