Matej Zorek

dblp:281/6702 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
0.912025
Differentiable Distance Between Hierarchically-Structured Data · ICDM 2025
Data mining › predictive modeling
classification
0.912025
Differentiable Distance Between Hierarchically-Structured Data · ICDM 2025
Data mining › predictive modeling › classification › nearest neighbor classification
distance-based classification
0.912025
Differentiable Distance Between Hierarchically-Structured Data · ICDM 2025
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.912025
Differentiable Distance Between Hierarchically-Structured Data · ICDM 2025

Methods — techniques the papers use, named apart from their topics

differentiable distance functions · 0.9
YearPublicationVenuePosition
2026 Agility for Steganalysis: Dealing with Distribution Drift in JPEG File Structures
Matej Zorek, Tomás Pevný, Rainer Böhme
EuroS&P1
2025 Differentiable Distance Between Hierarchically-Structured Data
abstract
Many popular machine learning methods for classification, anomaly detection, clustering, and dimensionality reduction rely on distance functions. However, these methods' theoretical foundations and practical performance critically depend on well-defined and meaningful distance functions on the input space. While distances are well-defined in Euclidean spaces, extending them to popular structured data stored in formats such as JSON, XML, ProtoBuffer, or MessagePack remains challenging. To fill this gap, this work proposes the Hierarchically-Structured Tree Distance (HTD), a fully differentiable distance tailored to these heterogeneous data formats. HTD is flexible, parameterized by weights, and supports automatic recursive construction, enabling it to be used on various datasets and tasks. This effectiveness and generality is demonstrated on a wide range of tasks, such as classification, anomaly detection, fast retrieval (indexing), and clustering, and on diverse datasets. The experimental comparison shows that classical distance-based methods with the HTD often rival or outperform state-of-the-art neural network models with orders of magnitude more parameters. HTD thus opens the door to scalable, interpretable, and efficient modeling of hierarchically structured data.
Matej Zorek, Tomás Pevný, Václav Smídl
ICDM1
2024 Deep anomaly detection on set data: Survey and comparison
Michaela Masková, Matej Zorek, Tomás Pevný, Václav Smídl
Pattern Recognit.2
2022 Comparison of Anomaly Detectors: Context Matters
abstract
Deep generative models are challenging the classical methods in the field of anomaly detection nowadays. Every newly published method provides evidence of outperforming its predecessors, sometimes with contradictory results. The objective of this article is twofold: to compare anomaly detection methods of various paradigms with a focus on deep generative models and identification of sources of variability that can yield different results. The methods were compared on popular tabular and image datasets. We identified that the main sources of variability are the experimental conditions: 1) the type of dataset (tabular or image) and the nature of anomalies (statistical or semantic) and 2) strategy of selection of hyperparameters, especially the number of available anomalies in the validation set. Methods perform differently in different contexts, i.e., under a different combination of experimental conditions together with computational time. This explains the variability of the previous results and highlights the importance of careful specification of the context in the publication of a new method. All our code and results are available for download.
Vít Skvára, Jan Francu, Matej Zorek, Tomás Pevný, Václav Smídl
IEEE Trans. Neural Networks Learn. Syst.3