EDBT 2026 Demo / reviewers in the wild / expert
Tomás Pevný
dblp:20/1317
· DBLP profile ↗
53ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-5768-9713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 2 first-author · 13 since 2021Security and privacy · 23 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Computer networks · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distillation of a tractable model from the VQ-VAEabstractDeep generative models with a discrete latent space, such as the Vector-Quantized Variational Autoencoder (VQ-VAE), offer excellent data generation capabilities, but-due to the large size of their latent space-their probabilistic inference is deemed intractable.We demonstrate that the VQ-VAE can be distilled into a tractable model by selecting a subset of latent variables with high probability under the prior.We frame the distilled model as a probabilistic circuit, and show that it preserves the expressiveness of the VQ-VAE while providing tractable probabilistic inference.Experiments illustrate competitive performance in both density estimation and conditional generation tasks, challenging the view of the VQ-VAE as an inherently intractable model. Armin Hadzic, Milan Papez, Tomás Pevný |
ESANN | 3 |
| 2026 | Agility for Steganalysis: Dealing with Distribution Drift in JPEG File Structures
Matej Zorek, Tomás Pevný, Rainer Böhme |
EuroS&P | 2 |
| 2026 | Tackle CSM in JPEG Steganalysis with Data AdaptationabstractSteganalysis models excel on benchmark datasets but struggle in the wild when analyzed images are produced by a processing pipeline unseen during training. This problem known as Cover Source Mismatch (CSM) is particularly hard in realistic settings where practitioners (1) have access to only a small, unlabeled dataset, (2) are unsure of the processing techniques applied to these images, and (3) lack information on the proportion of covers and stegos in that set. To answer this challenge, we introduce TADA (Target Alignment through Data Adaptation), a framework learning to emulate the unknown processing pipeline from a small unlabeled target set. This architecture is trained with a loss combining residual covariance alignment, residual distribution matching, and a ℓ2 loss constraining the emulator to produce realistic images. Across toy and operational targets, TADA yields substantial gains in robustness to CSM and improves operational generalization compared to strong holistic and atomistic baselines. Additional resources are available at this link: https://github.com/RonyAbecidan/TADA. Rony Abecidan, Vincent Itier, Jérémie Boulanger, Patrick Bas, Tomás Pevný |
IH&MMSec | 5 |
| 2026 | Cover-Source Mismatch in Deepfake Detection: A Systematic StudyabstractAI-generated image detectors achieve near-perfect accuracy under controlled conditions but fail dramatically when deployed on images with different processing pipelines. This failure is analogous to cover-source mismatch (CSM), well-understood in steganalysis for decades but unexplored in deepfake detection. We investigate this bottleneck through controlled experiments where image processing pipelines and semantic content are precisely manipulated. Using contemporary detectors as case studies, we characterize robustness across three distribution shift dimensions: pipeline shifts, semantic shifts, and image-space shifts. We find that detectors rely predominantly on pipeline artifacts rather than semantic signals, a dependence that renders them brittle under real-world deployment. Notably, matching semantic content between generated and camera images paradoxically degrades robustness—evidence that detectors exploit fragile pixel-level artifacts. Our work formalizes the domain adaptation bottleneck in deepfake detection and suggests that understanding image processing pipelines is fundamental to building generalizable detectors. Jan Francu, Milan Papez, Tomás Pevný |
IH&MMSec | 3 |
| 2026 | Better Inversion of Diffusion Models for Generative SteganographyabstractTraditional inversion algorithms attempt to directly invert the diffusion sampling equation. In this work, built on Latent Diffusion Models (LDMs), we propose a family of algorithms with varying time complexities that perform the search of an antecedent within the latent space and/or the Variational Autoencoder (VAE) decoder. Aurélien Noirault, Tomás Pevný, Jan Butora, Vincent Itier, Patrick Bas |
IH&MMSec | 2 |
| 2025 | State Encodings for GNN-Based Lifted PlannersabstractThe application of graph neural networks (GNNs) to learn heuristic functions in classical planning is gaining traction. Despite the variety of methods proposed in the literature to encode classical planning tasks for GNNs, a comparative study evaluating their relative performances has been lacking. Moreover, some encodings have been assessed solely for their expressiveness rather than practical effectiveness in planning. This paper provides an extensive comparative analysis of existing encodings. Our results indicate that the smallest encoding based on Gaifman graphs, not yet applied in planning, outperforms the rest due to its fast evaluation times and the informativeness of the resulting heuristic. The overall coverage measured on the IPC almost reaches that of the state-of-the-art planner LAMA while exhibiting rather complementary strengths across different domains. Rostislav Horcík, Gustav Sír, Vítezslav Simek, Tomás Pevný |
AAAI | 4 |
| 2025 | Differentiable Distance Between Hierarchically-Structured DataabstractMany popular machine learning methods for classification, anomaly detection, clustering, and dimensionality reduction rely on distance functions. However, these methods' theoretical foundations and practical performance critically depend on well-defined and meaningful distance functions on the input space. While distances are well-defined in Euclidean spaces, extending them to popular structured data stored in formats such as JSON, XML, ProtoBuffer, or MessagePack remains challenging. To fill this gap, this work proposes the Hierarchically-Structured Tree Distance (HTD), a fully differentiable distance tailored to these heterogeneous data formats. HTD is flexible, parameterized by weights, and supports automatic recursive construction, enabling it to be used on various datasets and tasks. This effectiveness and generality is demonstrated on a wide range of tasks, such as classification, anomaly detection, fast retrieval (indexing), and clustering, and on diverse datasets. The experimental comparison shows that classical distance-based methods with the HTD often rival or outperform state-of-the-art neural network models with orders of magnitude more parameters. HTD thus opens the door to scalable, interpretable, and efficient modeling of hierarchically structured data. Matej Zorek, Tomás Pevný, Václav Smídl |
ICDM | 2 |
| 2025 | Generating Likely Counterfactuals Using Sum-Product NetworksabstractThe need to explain decisions made by AI systems is driven by both recent regulation and user demand. The decisions are often explainable only post hoc. In counterfactual explanations, one may ask what constitutes the best counterfactual explanation. Clearly, multiple criteria must be taken into account, although "distance from the sample" is a key criterion. Recent methods that consider the plausibility of a counterfactual seem to sacrifice this original objective. Here, we present a system that provides high-likelihood explanations that are, at the same time, close and sparse. We show that the search for the most likely explanations satisfying many common desiderata for counterfactual explanations can be modeled using Mixed-Integer Optimization (MIO). We use a Sum-Product Network (SPN) to estimate the likelihood of a counterfactual. To achieve that, we propose an MIO formulation of an SPN, which can be of independent interest. The source code with examples is available at https://github.com/Epanemu/LiCE. Jiri Nemecek 0002, Tomás Pevný, Jakub Marecek |
ICLR | 2 |
| 2025 | Bias Detection via Maximum Subgroup DiscrepancyabstractBias evaluation is fundamental to trustworthy AI, both in terms of checking data quality and in terms of checking the outputs of AI systems.In testing data quality, for example, one may study the distance of a given dataset, viewed as a distribution, to a given ground-truth reference dataset.However, classical metrics, such as the Total Variation and the Wasserstein distances, are known to have high sample complexities and, therefore, may fail to provide a meaningful distinction in many practical scenarios.In this paper, we propose a new notion of distance, the Maximum Subgroup Discrepancy (MSD).In this metric, two distributions are close if, roughly, discrepancies are low for all feature subgroups.While the number of subgroups may be exponential, we show that the sample complexity is linear in the number of features, thus making it feasible for practical applications.Moreover, we provide a practical algorithm for evaluating the distance based on Mixedinteger optimization (MIO).We also note that the proposed distance is easily interpretable, thus providing clearer paths to fixing the biases once they have been identified.Finally, we describe a natural general bias detection framework, termed MSDD distances, and show that MSD aligns well with this framework.We empirically evaluate MSD by comparing it with other metrics and by demonstrating the above properties of MSD on real-world datasets. Jiri Nemecek 0002, Mark Kozdoba, Illia Kryvoviaz, Tomás Pevný, Jakub Marecek |
KDD (2) | 4 |
| 2025 | Probabilistic Graph Circuits: Deep Generative Models for Tractable Probabilistic Inference over GraphsabstractDeep generative models (DGMs) have recently demonstrated remarkable success in capturing complex probability distributions over graphs. Although their excellent performance is attributed to powerful and scalable deep neural networks, it is, at the same time, exactly the presence of these highly non-linear transformations that makes DGMs intractable. Indeed, despite representing probability distributions, intractable DGMs deny probabilistic foundations by their inability to answer even the most basic inference queries without approximations or design choices specific to a very narrow range of queries. To address this limitation, we propose probabilistic graph circuits (PGCs), a framework of tractable DGMs that provide exact and efficient probabilistic inference over (arbitrary parts of) graphs. Nonetheless, achieving both exactness and efficiency is challenging in the permutation-invariant setting of graphs. We design PGCs that are inherently invariant and satisfy these two requirements, yet at the cost of low expressive power. Therefore, we investigate two alternative strategies to achieve the invariance: the first sacrifices the efficiency, and the second sacrifices the exactness. We demonstrate that ignoring the permutation invariance can have severe consequences in anomaly detection, and that the latter approach is competitive with, and sometimes better than, existing intractable DGMs in the context of molecular graph generation. Milan Papez, Martin Rektoris, Václav Smídl, Tomás Pevný |
UAI | 4 |
| 2024 | Sum-Product-Set Networks: Deep Tractable Models for Tree-Structured GraphsabstractDaily internet communication relies heavily on tree-structured graphs, embodied by popular data formats such as XML and JSON. However, many recent generative (probabilistic) models utilize neural networks to learn a probability distribution over undirected cyclic graphs. This assumption of a generic graph structure brings various computational challenges, and, more importantly, the presence of non-linearities in neural networks does not permit tractable probabilistic inference. We address these problems by proposing sum-product-set networks, an extension of probabilistic circuits from unstructured tensor data to tree-structured graph data. To this end, we use random finite sets to reflect a variable number of nodes and edges in the graph and to allow for exact and efficient inference. We demonstrate that our tractable model performs comparably to various intractable models based on neural networks. Milan Papez, Martin Rektoris, Václav Smídl, Tomás Pevný |
ICLR | 4 |
| 2024 | Classification with costly features in hierarchical deep setsabstractAbstract Classification with costly features (CwCF) is a classification problem that includes the cost of features in the optimization criteria. Individually for each sample, its features are sequentially acquired to maximize accuracy while minimizing the acquired features’ cost. However, existing approaches can only process data that can be expressed as vectors of fixed length. In real life, the data often possesses rich and complex structure, which can be more precisely described with formats such as XML or JSON. The data is hierarchical and often contains nested lists of objects. In this work, we extend an existing deep reinforcement learning-based algorithm with hierarchical deep sets and hierarchical softmax, so that it can directly process this data. The extended method has greater control over which features it can acquire and, in experiments with seven datasets, we show that this leads to superior performance. To showcase the real usage of the new method, we apply it to a real-life problem of classifying malicious web domains, using an online service. Jaromír Janisch, Tomás Pevný, Viliam Lisý |
Mach. Learn. | 2 |
| 2024 | Anomaly detection in multifactor data
Vít Skvára, Václav Smídl, Tomás Pevný |
Neural Comput. Appl. | 3 |
| 2024 | Deep anomaly detection on set data: Survey and comparison
Michaela Masková, Matej Zorek, Tomás Pevný, Václav Smídl |
Pattern Recognit. | 3 |
| 2024 | Malicious Internet Entity Detection Using Local Graph InferenceabstractDetection of malicious behavior in a large network is a challenging problem for machine learning in computer security, since it requires a model with high expressive power and scalable inference. Existing solutions struggle to achieve this feat—current cybersec-tailored approaches are still limited in expressivity, and methods successful in other domains do not scale well for large volumes of data, rendering frequent retraining impossible. This work proposes a new perspective for learning from graph data that is modeling network entity interactions as a large heterogeneous graph. High expressivity of the method is achieved with neural network architecture HMILnet that naturally models this type of data and provides theoretical guarantees. The scalability is achieved by pursuing local graph inference, i.e., classifying individual vertices and their neighborhood as independent samples. Our experiments exhibit improvement over the state-of-the-art Probabilistic Threat Propagation (PTP) algorithm, show a further threefold accuracy improvement when additional data is used, which is not possible with the PTP algorithm, and demonstrate the generalization capabilities of the method to new, previously unseen entities. Simon Mandlík, Tomás Pevný, Václav Smídl, Lukás Bajer |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | On the Economics of Adversarial Machine LearningabstractGiven the widespread deployment of machine learning algorithms, the security of these algorithms and thus, the field of adversarial machine learning gained popularity in the research community. In this article, we loosen several unrealistic restrictions found in prior art and bring economical-inspired adversarial machine learning one step closer to being applicable in the real world. First, we extend our own game-theoretical framework such that it allows any arbitrary number of actions for both actors, and analytically determine equilibrium strategies and conditions where mixed strategies are expected for the specific case in which both actors choose from any two arbitrary actions. Then, we pay special attention to an adversary’s knowledge about the attacked system by modeling them as a white-, gray-, or black-box adversary. We conduct extensive experiments for three architectures, two training procedures, and four adversarial attacks in different variations as direct and transfer attacks, resulting in 300 data points consisting of the respective accuracy and robustness values and the computational costs for both actors. We then instantiate our model with this data and explore the structure of the game for a wide range of each game parameter, overcoming the complexity by applying algorithmic game theory. We discover surprising properties in the actors’ strategies, such as the feasibility of cheap attacks that have been dismissed as practically irrelevant so far - examples include universal adversarial perturbations or (transfer) attacks utilizing only few optimization steps. For the defender, we find that given recent attacks and countermeasures, a rational defender would try to hide as much as possible from their infrastructure. Florian Merkle, Maximilian Samsinger, Pascal Schöttle, Tomás Pevný |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Optimize Planning Heuristics to Rank, not to Estimate Cost-to-GoalabstractIn imitation learning for planning, parameters of heuristic functions are optimized against a set of solved problem instances. This work revisits the necessary and sufficient conditions of strictly optimally efficient heuristics for forward search algorithms, mainly A* and greedy best-first search, which expand only states on the returned optimal path. It then proposes a family of loss functions based on ranking tailored for a given variant of the forward search algorithm. Furthermore, from a learning theory point of view, it discusses why optimizing cost-to-goal h* is unnecessarily difficult. The experimental comparison on a diverse set of problems unequivocally supports the derived theory. Leah Chrestien, Stefan Edelkamp, Antonín Komenda, Tomás Pevný |
NeurIPS | 4 |
| 2023 | The Non-Zero-Sum Game of Steganography in Heterogeneous EnvironmentsabstractThe highly heterogeneous nature of images found in real-world environments, such as online sharing platforms, has been one of the long-standing obstacles to the transition of steganalysis techniques outside the laboratory. Recent advances in identifying the properties of images relevant to steganalysis as well as the effectiveness of deep neural networks on highly heterogeneous datasets have laid some groundwork for resolving this problem. Despite this progress, we argue that the way the game played between the steganographer and the steganalyst is currently modeled lacks some important features expected in a real-world environment: 1) the steganographer can adapt her cover source choice to the environment and/or to the steganalyst’s classifier, 2) the distribution of cover sources in the environment impacts the optimal threshold for a given classifier, and 3) the steganalyst and steganographer have different goals, hence different utilities. We propose to take these facts into account using a two-player non-zero-sum game constrained by an environment composed of multiple cover sources. We then show how to convert this non-zero-sum game into an equivalent zero-sum game, allowing us to propose two methods to find Nash equilibria for this game: a standard method using the double oracle algorithm and a minimum regret method based on approximating a set of atomistic classifiers. Applying these methods to contemporary steganography and steganalysis in a realistic environment, we show that classifiers which do not adapt to the environment severely underperform when the steganographer is allowed to select into which cover source to embed. Eva Giboulot, Tomás Pevný, Andrew D. Ker |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | JsonGrinder.jl: automated differentiable neural architecture for embedding arbitrary JSON dataabstractStandard machine learning (ML) problems are formulated on data converted into a suitable tensor representation. However, there are data sources, for example in cybersecurity, that are naturally represented in a unifying hierarchical structure, such as XML, JSON, and Protocol Buffers. Converting this data to a tensor representation is usually done by manual feature engineering, which is laborious, lossy, and prone to bias originating from the human inability to correctly judge the importance of particular features. JsonGrinder.jl is a library automating various ML tasks on these difficult sources. Starting with an arbitrary set of JSON samples, it automatically creates a differentiable ML model (called hmilnet), which embeds raw JSON samples into a fixed-size tensor representation. This embedding network can be naturally extended by an arbitrary ML model expecting tensor inputs in order to perform classification, regression, or clustering. Simon Mandlík, Matej Racinsky, Viliam Lisý, Tomás Pevný |
J. Mach. Learn. Res. | 4 |
| 2022 | Backpack: A Backpropagable Adversarial Embedding SchemeabstractA min max protocol offers a general method to automatically optimize steganographic algorithm against a wide class of steganalytic detectors. The quality of the resulting steganograhic algorithm depends on the ability to find an “adversarial” stego image undetectable by a set of detectors while communicating a given message. Despite min max protocol instantiated with ADV-EMB scheme leading to unexpectedly good results, we show it suffers a significant flaw and we present a theoretically sound solution called Backpack. Extensive experimental verification of min max protocol with Backpack shows superior performance to ADV-EMB, the generality of the tool by targeting a new JPEG QF100 compatibility attack and further improves the security of steganographic algorithms. Solène Bernard, Patrick Bas, John Klein, Tomás Pevný |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Comparison of Anomaly Detectors: Context MattersabstractDeep generative models are challenging the classical methods in the field of anomaly detection nowadays. Every newly published method provides evidence of outperforming its predecessors, sometimes with contradictory results. The objective of this article is twofold: to compare anomaly detection methods of various paradigms with a focus on deep generative models and identification of sources of variability that can yield different results. The methods were compared on popular tabular and image datasets. We identified that the main sources of variability are the experimental conditions: 1) the type of dataset (tabular or image) and the nature of anomalies (statistical or semantic) and 2) strategy of selection of hyperparameters, especially the number of available anomalies in the validation set. Methods perform differently in different contexts, i.e., under a different combination of experimental conditions together with computational time. This explains the variability of the previous results and highlights the importance of careful specification of the context in the publication of a new method. All our code and results are available for download. Vít Skvára, Jan Francu, Matej Zorek, Tomás Pevný, Václav Smídl |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Optimizing Additive Approximations of Non-additive Distortion FunctionsabstractThe progress in steganography is hampered by a gap between non-additive distortion functions, which capture well complex dependencies in natural images, and their additive counterparts, which are efficient for data embedding. This paper proposes a theoretically justified method to approximate the former by the latter. The proposed method, called Backpack (for BACKPropagable AttaCK), combines new results in the approximation of gradients of discrete distributions with a gradient of implicit functions in order to derive a gradient w.r.t. the distortion of each JPEG coefficient. Backpack combined with the min max iterative protocol leads to a very secure steganographic algorithm. For example, the error rate of XuNet on 512 X 512 JPEG images, compressed with quality factor 100 and a payload of 0.4 bits per non-zero AC coefficient is 37.3% with Backpack, compared to a 26.5% error rate using ADV-EMB with minmax (considered state of the art in this work) and a 16.9% error rate with J-UNIWARD. Solène Bernard, Patrick Bas, Tomás Pevný, John Klein |
IH&MMSec | 3 |
| 2021 | Explicit Optimization of min max Steganographic GameabstractThis article proposes an algorithm which allows Alice to simulate the game played between her and Eve. Under the condition that the set of detectors that Alice assumes Eve to have is sufficiently rich (e.g. CNNs), and that she has an algorithm enabling to avoid detection by a single classifier (e.g adversarial embedding, gibbs sampler, dynamic STCs), the proposed algorithm converges to an efficient steganographic algorithm. This is possible by using a min max strategy which consists at each iteration in selecting the least detectable stego image for the best classifier among the set of Eve's learned classifiers. The algorithm is extensively evaluated and compared to prior arts and results show the potential to increase the practical security of classical steganographic methods. For example the error probability Perr of XU-Net on detecting stego images with payload of 0.4 bpnzAC embedded by J-Uniward and QF 75 starts at 7.1% and is increased by +13.6% to reach 20.7% after eight iterations. For the same embedding rate and for QF 95, undetectability by XU-Net with J-Uniward embedding is 23.4%, and it jumps by +25.8% to reach 49.2% at iteration 3. Solène Bernard, Patrick Bas, John Klein, Tomás Pevný |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Neural Power UnitsabstractConventional Neural Networks can approximate simple arithmetic operations, but fail to generalize beyond the range of numbers that were seen during training. Neural Arithmetic Units aim to overcome this difficulty, but current arithmetic units are either limited to operate on positive numbers or can only represent a subset of arithmetic operations. We introduce the Neural Power Unit (NPU) that operates on the full domain of real numbers and is capable of learning arbitrary power functions in a single layer. The NPU thus fixes the shortcomings of existing arithmetic units and extends their expressivity. We achieve this by using complex arithmetic without requiring a conversion of the network to complex numbers. A simplification of the unit to the RealNPU yields a highly transparent model. We show that the NPUs outperform their competitors in terms of accuracy and sparsity on artificial arithmetic datasets, and that the RealNPU can discover the governing equations of a dynamical system only from data. Niklas Heim, Tomás Pevný, Václav Smídl |
NeurIPS | 2 |
| 2020 | Anomaly explanation with random forests
Martin Kopp, Tomás Pevný, Martin Holena |
Expert Syst. Appl. | 2 |
| 2020 | Classification with costly features as a sequential decision-making problem
Jaromír Janisch, Tomás Pevný, Viliam Lisý |
Mach. Learn. | 2 |
| 2019 | Classification with Costly Features Using Deep Reinforcement LearningabstractWe study a classification problem where each feature can be acquired for a cost and the goal is to optimize a trade-off between the expected classification error and the feature cost. We revisit a former approach that has framed the problem as a sequential decision-making problem and solved it by Q-learning with a linear approximation, where individual actions are either requests for feature values or terminate the episode by providing a classification decision. On a set of eight problems, we demonstrate that by replacing the linear approximation with neural networks the approach becomes comparable to the state-of-the-art algorithms developed specifically for this problem. The approach is flexible, as it can be improved with any new reinforcement learning enhancement, it allows inclusion of pre-trained high-performance classifier, and unlike prior art, its performance is robust across all evaluated datasets. Jaromír Janisch, Tomás Pevný, Viliam Lisý |
AAAI | 2 |
| 2019 | Exploiting Adversarial Embeddings for Better SteganographyabstractThis work proposes a protocol to iteratively build a distortion function for adaptive steganography while increasing its practical security after each iteration. It relies on prior art on targeted attacks and iterative design of steganalysis schemes. It combines targeted attacks on a given detector with a \min\max strategy, which dynamically selects the most difficult stego content associated with the best classifier at each iteration. We theoretically prove the convergence, which is confirmed by the practical results. Applied on J-Uniward this new protocol increases \perr from 7% to 20% estimated by Xu-Net, and from 10% to 23% for a non-targeted steganalysis by a linear classifier with GFR features. Solène Bernard, Tomás Pevný, Patrick Bas, John Klein |
IH&MMSec | 2 |
| 2019 | Joint detection of malicious domains and infected clients
Paul Prasse, René Knaebel, Lukás Machlica, Tomás Pevný, Tobias Scheffer |
Mach. Learn. | 4 |
| 2018 | Exploring Non-Additive Distortion in SteganographyabstractLeading steganography systems make use of the Syndrome-Trellis Code (STC) algorithm to minimize a distortion function while encoding the desired payload, but this constrains the distortion function to be additive. The Gibbs Embedding algorithm works for a certain class of non-additive distortion functions, but has its own limitations and is highly complex. Tomás Pevný, Andrew D. Ker |
IH&MMSec | 1 |
| 2018 | Probabilistic analysis of dynamic malware traces
Jan Stiborek, Tomás Pevný, Martin Rehák |
Comput. Secur. | 2 |
| 2018 | Multiple instance learning for malware classification
Jan Stiborek, Tomás Pevný, Martin Rehák |
Expert Syst. Appl. | 2 |
| 2018 | Network Traffic Fingerprinting Based on Approximated Kernel Two-Sample TestabstractMany applications and communication protocols exhibit unique communication patterns that can be exploited to identify them in network traffic. This paper proposes a method to represent these patterns compactly, such that they can be used in different analytical tasks. The method treats each communication as a set of observations of a random variable with unknown probability distribution. This view allows us to derive the representation from a distance between two probability distributions used in maximum mean discrepancy—a non-parametric kernel test. The representation (and distance) can be then easily used in various algorithms for identification of communicating application and data analysis, independently of the specific type of input data. Jan Kohout, Tomás Pevný |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Using Neural Network Formalism to Solve Multiple-Instance Problems
Tomás Pevný, Petr Somol |
ISNN (1) | 1 |
| 2017 | Malware Detection by Analysing Encrypted Network Traffic with Neural Networks
Paul Prasse, Lukás Machlica, Tomás Pevný, Jirí Havelka, Tobias Scheffer |
ECML/PKDD (2) | 3 |
| 2017 | Malware Discovery Using Behaviour-Based Exploration of Network Traffic
Jakub Lokoc, Tomás Grosup, Premysl Cech, Tomás Pevný, Tomás Skopal |
SISAP | 4 |
| 2017 | Reducing false positives of network anomaly detection by local adaptive multivariate smoothing
Martin Grill, Tomás Pevný, Martin Rehák |
J. Comput. Syst. Sci. | 2 |
| 2016 | Rethinking Optimal EmbeddingabstractAt present, almost all leading steganographic techniques for still images use a distortion minimization paradigm, where each potential change is assigned a cost ci and the change probabilities πi chosen to minimize the average total cost ∑iπici. However, some detectors have exploited knowledge of this adaptivity and the embedding cannot be considered optimal. In this work we prove a theoretical result suggesting that, against a knowing attacker, the embedder should simply minimize ∑iπ2ici instead, for the same costs ci, which is the minimax and equilibrium strategy. This aligns with some special case results that have appeared in recent literature. We then test some simple steganographic methods in theoretical and real settings, showing that naive (average cost) adaptivity is exploitable, but the equilibrium probabilities cannot be exploited. However, it is essential to determine statistically well-founded costs ci. Andrew D. Ker, Tomás Pevný, Patrick Bas |
IH&MMSec | 2 |
| 2016 | Feature Extraction and Malware Detection on Large HTTPS Data Using MapReduce
Premysl Cech, Jan Kohout, Jakub Lokoc, Tomás Komárek, Jakub Marousek, Tomás Pevný |
SISAP | 6 |
| 2016 | Learning combination of anomaly detectors for security domain
Martin Grill, Tomás Pevný |
Comput. Networks | 2 |
| 2016 | Loda: Lightweight on-line detector of anomalies
Tomás Pevný |
Mach. Learn. | 1 |
| 2015 | Unsupervised detection of malware in persistent web trafficabstractPersistent network communication can be found in many instances of malware. In this paper, we analyse the possibility of leveraging low variability of persistent malware communication for its detection. We propose a new method for capturing statistical fingerprints of connections and employ outlier detection to identify the malicious ones. Emphasis is put on using minimal information possible to make our method very lightweight and easy to deploy. Anomaly detection is commonly used in network security, yet to our best knowledge, there are not many works focusing on the persistent communication itself, without making further assumptions about its purpose. Jan Kohout, Tomás Pevný |
ICASSP | 2 |
| 2015 | Automatic discovery of web servers hosting similar applicationsabstractIncreasingly more popular cloud services have frequently many functional parts, which makes their structure rather complex yet its understanding improves network monitoring for security purposes, traffic routing, etc. Since the structure of third-party services is typically unknown, automated tools for its discovery are of great need. In this work, we propose such tool relying only on high-level statistics of servers' usage, such as volumes and times of interactions with the servers. Without looking into the communication contents, the method works for encrypted channels as well, which is experimentally demonstrated on Dropbox service and Windows Live platform. Jan Kohout, Tomás Pevný |
IM | 2 |
| 2014 | Steganographic key leakage through payload metadataabstractThe only steganalysis attack which can provide absolute certainty about the presence of payload is one which finds the embedding key. In this paper we consider refined versions of the key exhaustion attack exploiting metadata such as message length or decoding matrix size, which must be stored along with the payload. We show simple errors of implementation lead to leakage of key information and powerful inference attacks; furthermore, complete absence of information leakage seems difficult to avoid. This topic has been somewhat neglected in the literature for the last ten years, but must be considered in real-world implementations. Tomás Pevný, Andrew D. Ker |
IH&MMSec | 1 |
| 2014 | Randomized Operating Point Selection in Adversarial Classification
Viliam Lisý, Robert Kessl, Tomás Pevný |
ECML/PKDD (2) | 3 |
| 2014 | The Steganographer is the Outlier: Realistic Large-Scale SteganalysisabstractWe present a method for a completely new kind of steganalysis to determine who, out of a large number of actors each transmitting a large number of objects, is hiding payload inside some of them. It has significant challenges, including unknown embedding parameters and natural deviation between innocent cover sources, which are usually avoided in steganalysis tested under laboratory conditions. Our method uses standard steganalysis features, the maximum mean discrepancy measure of distance, and ranks the actors by their degree of deviation from the rest: we show that it works reliably, completely unsupervised, when tested against some of the standard steganography methods available to nonexperts. We also determine good parameters for the detector and show that it creates a two-player game between the guilty actor and the steganalyst. Andrew D. Ker, Tomás Pevný |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Attacking the IDS learning processesabstractWe study the problem of directed attacks on the learning process of an anomaly-based Intrusion Detection System (IDS). We assume that the attack is performed by a knowledgeable attacker with an access to system's inputs, outputs, and all internal states. The attacker uses his knowledge of the IDS (implemented as an ensemble of anomaly detection algorithms) and its internal states to design the strongest undetectable attack of a particular type. We have experimented with different attacks against several anomaly detection algorithms individually, and against their combination. We show that while the individual anomaly detection algorithms can be easily avoided by the worst-case attacker that we assume, it is nearly impossible to avoid them simultaneously. These results were achieved during the experiments performed on university network traffic and are consistent with theoretical hypothesis grounded in steganalysis and watermarking. Tomás Pevný, Martin Komon, Martin Rehák |
ICASSP | 1 |
| 2013 | Moving steganography and steganalysis from the laboratory into the real worldabstractThere has been an explosion of academic literature on steganography and steganalysis in the past two decades. With a few exceptions, such papers address abstractions of the hiding and detection problems, which arguably have become disconnected from the real world. Most published results, including by the authors of this paper, apply "in laboratory conditions" and some are heavily hedged by assumptions and caveats; significant challenges remain unsolved in order to implement good steganography and steganalysis in practice. This position paper sets out some of the important questions which have been left unanswered, as well as highlighting some that have already been addressed successfully, for steganography and steganalysis to be used in the real world. Andrew D. Ker, Patrick Bas, Rainer Böhme, Rémi Cogranne, Scott Craver, Tomás Filler, Jessica J. Fridrich, Tomás Pevný |
IH&MMSec | 8 |
| 2012 | From Blind to Quantitative SteganalysisabstractA quantitative steganalyzer is an estimator of the number of embedding changes introduced by a specific embedding operation. Since for most algorithms the number of embedding changes correlates with the message length, quantitative steganalyzers are important forensic tools. In this paper, a general method for constructing quantitative steganalyzers from features used in blind detectors is proposed. The core of the method is a support vector regression, which is used to learn the mapping between a feature vector extracted from the investigated object and the embedding change rate. To demonstrate the generality of the proposed approach, quantitative steganalyzers are constructed for a variety of steganographic algorithms in both JPEG transform and spatial domains. The estimation accuracy is investigated in detail and compares favorably with state-of-the-art quantitative steganalyzers. Tomás Pevný, Jessica J. Fridrich, Andrew D. Ker |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2010 | Steganalysis by subtractive pixel adjacency matrixabstractThis paper presents a method for detection of steganographic methods that embed in the spatial domain by adding a low-amplitude independent stego signal, an example of which is least significant bit (LSB) matching. First, arguments are provided for modeling the differences between adjacent pixels using first-order and second-order Markov chains. Subsets of sample transition probability matrices are then used as features for a steganalyzer implemented by support vector machines. The major part of experiments, performed on four diverse image databases, focuses on evaluation of detection of LSB matching. The comparison to prior art reveals that the presented feature set offers superior accuracy in detecting LSB matching. Even though the feature set was developed specifically for spatial domain steganalysis, by constructing steganalyzers for ten algorithms for JPEG images, it is demonstrated that the features detect steganography in the transform domain as well. Tomás Pevný, Patrick Bas, Jessica J. Fridrich |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2008 | Detection of Double-Compression in JPEG Images for Applications in SteganographyabstractThis paper presents a method for the detection of double JPEG compression and a maximum-likelihood estimator of the primary quality factor. These methods are essential for construction of accurate targeted and blind steganalysis methods for JPEG images. The proposed methods use support vector machine classifiers with feature vectors formed by histograms of low-frequency discrete cosine transformation coefficients. The performance of the algorithms is compared to selected prior art. Tomás Pevný, Jessica J. Fridrich |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2008 | Multiclass Detector of Current Steganographic Methods for JPEG FormatabstractThe aim of this paper is to construct a practical forensic steganalysis tool for JPEG images that can properly analyze both single- and double-compressed stego images and classify them to selected current steganographic methods. Although some of the individual modules of the steganalyzer were previously published by the authors, they were never tested as a complete system. The fusion of the modules brings its own challenges and problems whose analysis and solution is one of the goals of this paper. By determining the stego algorithm, this tool provides the first step needed for extracting the secret message. Given a JPEG image, the detector assigns it to 6 popular steganographic algorithms. The detection is based on feature extraction and supervised training of two banks of multi-classifiers realized using support vector machines. For accurate classification of single-compressed images, a separate multi-classifier is trained for each JPEG quality factor from a certain range. Another bank of multiclassifiers is trained for double-compressed images for the same range of primary quality factors. The image under investigation is first analyzed using a pre-classifier that detects selected cases of double-compression and estimates the primary quantization table. It then sends the image to the appropriate single- or double-compression multiclassifier. The error is estimated from more than 2.6 million images. The steganalyzer is also tested on two previously unseen methods to examine its ability to generalize. Tomás Pevný, Jessica J. Fridrich |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2005 | Towards Multi-class Blind Steganalyzer for JPEG Images
Tomás Pevný, Jessica J. Fridrich |
IWDW | 1 |