Andrii Shkabrii

dblp:360/7677 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 61% Representation and self-supervised learning · 30% Image recognition and object detection · 9%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
perturbation robustness
0.912025
H-SPLID: HSIC-based Saliency Preserving Latent Information Decomposition · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
H-SPLID: HSIC-based Saliency Preserving Latent Information Decomposition · NeurIPS 2025
Data mining
clustering
0.912025
Breaking the Reclustering Barrier in Centroid-based Deep Clustering · ICLR 2025
Data mining › clustering
deep clustering
0.912025
Breaking the Reclustering Barrier in Centroid-based Deep Clustering · ICLR 2025
Computer vision › Image recognition and object detection
image classification
0.312025
H-SPLID: HSIC-based Saliency Preserving Latent Information Decomposition · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

representation decomposition · 0.9hilbert-schmidt independence criterion · 0.9contrastive loss · 0.9
YearPublicationVenuePosition
2025 Breaking the Reclustering Barrier in Centroid-based Deep Clustering
abstract
This work investigates an important phenomenon in centroid-based deep clustering (DC) algorithms: Performance quickly saturates after a period of rapid early gains. Practitioners commonly address early saturation with periodic reclustering, which we demonstrate to be insufficient to address performance plateaus. We call this phenomenon the “reclustering barrier” and empirically show when the reclustering barrier occurs, what its underlying mechanisms are, and how it is possible to Break the Reclustering Barrier with our algorithm BRB. BRB avoids early over-commitment to initial clusterings and enables continuous adaptation to reinitialized clustering targets while remaining conceptually simple. Applying our algorithm to widely-used centroid-based DC algorithms, we show that (1) BRB consistently improves performance across a wide range of clustering benchmarks, (2) BRB enables training from scratch, and (3) BRB performs competitively against state-of-the-art DC algorithms when combined with a contrastive loss. We release our code and pre-trained models at https://github.com/Probabilistic-and-Interactive-ML/breaking-the-reclustering-barrier .
Lukas Miklautz, Timo Klein, Kevin Sidak, Collin Leiber, Thomas Lang, Andrii Shkabrii, Sebastian Tschiatschek, Claudia Plant
ICLR6
2025 H-SPLID: HSIC-based Saliency Preserving Latent Information Decomposition
abstract
We introduce H-SPLID, a novel algorithm for learning salient feature representations through the explicit decomposition of salient and non-salient features into separate spaces. We show that H-SPLID promotes learning low-dimensional, task-relevant features. We prove that the expected prediction deviation under input perturbations is upper-bounded by the dimension of the salient subspace and the Hilbert-Schmidt Independence Criterion (HSIC) between inputs and representations. This establishes a link between robustness and latent representation compression in terms of the dimensionality and information preserved. Empirical evaluations on image classification tasks show that models trained with H-SPLID primarily rely on salient input components, as indicated by reduced sensitivity to perturbations affecting non-salient features, such as image backgrounds.
Lukas Miklautz, Chengzhi Shi, Andrii Shkabrii, Theodoros-Thirimachos Davarakis, Prudence Lam, Claudia Plant, Jennifer G. Dy, Stratis Ioannidis
NeurIPS3
2023 Non-Redundant Image Clustering of Early Medieval Glass Beads
abstract
Glass beads were among the most common grave goods in the Early Middle Ages, with an estimated number in the millions. The color, size, shape and decoration of the beads are diverse leading to many different archaeological classification systems that depend on the subjective decisions of individual experts. The lack of an agreed upon expert categorization leads to a pressing problem in archaeology, as the categorization of archaeological artifacts, like glass beads, is important to learn about cultural trends, manufacturing processes or economic relationships (e.g., trade routes) of historical times. An automated, objective and reproducible classification system is therefore highly desirable. We present a high-quality data set of images of Early Medieval beads and propose a clustering pipeline to learn a classification system in a data-driven way. The pipeline consists of a novel extension of deep embedded non-redundant clustering to identify multiple, meaningful clusterings of glass bead images. During the cluster analysis we address several challenges associated with the data and as a result identify high-quality clusterings that overlap with archaeological domain expertise. To the best of our knowledge this is the first application of non-redundant image clustering for archaeological data.
Lukas Miklautz, Andrii Shkabrii, Collin Leiber, Bendeguz Tobias, Benedict Seidl, Elisabeth Weissensteiner, Andreas Rausch 0001, Christian Böhm 0001, Claudia Plant
DSAA2