EDBT 2026 Demo / reviewers in the wild / expert
Pavel Laskov
dblp:63/4310
· DBLP profile ↗
31ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0002-3212-7167ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-authorSecurity and privacy · 11 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Ephemeral Threat: Assessing the Security of Algorithmic Trading Systems powered by Deep LearningabstractWe study the security of stock price forecasting using Deep Learning (DL) in computational finance.Despite abundant prior research on vulnerability of DL to adversarial perturbations, such work has hitherto hardly addressed practical adversarial threat models in the context of DL-powered algorithmic trading systems (ATS).Specifically, we investigate the vulnerability of ATS to adversarial perturbations launched by a realistically constrained attacker.We first show that existing literature has paid limited attention to DL security in the financial domain-which is naturally attractive for adversaries.Then, we formalize the concept of ephemeral perturbations (EP), which can be used to stage a novel type of attack tailored for DL-based ATS.Finally, we carry out an end-to-end evaluation of our EP against a profitable ATS.Our results reveal that the introduction of small changes to the input stock-prices not only (i) induces the DL model to behave incorrectly but also (ii) leads to the whole ATS to make suboptimal buy/sell decisions, resulting in a worse financial performance of the targeted ATS. Advije Rizvani, Giovanni Apruzzese, Pavel Laskov |
CODASPY | 3 |
| 2023 | SoK: Pragmatic Assessment of Machine Learning for Network Intrusion DetectionabstractMachine Learning (ML) has become a valuable asset to solve many real-world tasks. For Network Intrusion Detection (NID), however, scientific advances in ML are still seen with skepticism by practitioners. This disconnection is due to the intrinsically limited scope of research papers, many of which primarily aim to demonstrate new methods "outperforming" prior work—oftentimes overlooking the practical implications for deploying the proposed solutions in real systems. Unfortunately, the value of ML for NID depends on a plethora of factors, such as hardware, that are often neglected in scientific literature.This paper aims to reduce the practitioners’ skepticism towards ML for NID by changing the evaluation methodology adopted in research. After elucidating which factors influence the operational deployment of ML in NID, we propose the notion of pragmatic assessment, which enable practitioners to gauge the real value of ML methods for NID. Then, we show that the state-of-research hardly allows one to estimate the value of ML for NID. As a constructive step forward, we carry out a pragmatic assessment. We re-assess existing ML methods for NID, focusing on the classification of malicious network traffic, and consider: hundreds of configuration settings; diverse adversarial scenarios; and four hardware platforms. Our large and reproducible evaluations enable estimating the quality of ML for NID. We also validate our claims through a user-study with security practitioners. Giovanni Apruzzese, Pavel Laskov, Johannes Schneider 0002 |
EuroS&P | 2 |
| 2022 | SoK: The Impact of Unlabelled Data in Cyberthreat DetectionabstractMachine learning (ML) has become an important paradigm for cyberthreat detection (CTD) in the recent years. A substantial research effort has been invested in the development of specialized algorithms for CTD tasks. From the operational perspective, however, the progress of ML-based CTD is hindered by the difficulty in obtaining the large sets of labelled data to train ML detectors. A potential solution to this problem are semisupervised learning (SsL) methods, which combine small labelled datasets with large amounts of unlabelled data. This paper is aimed at systematization of existing work on SsL for CTD and, in particular, on understanding the utility of unlabelled data in such systems. To this end, we analyze the cost of labelling in various CTD tasks and develop a formal cost model for SsL in this context. Building on this foundation, we formalize a set of requirements for evaluation of SsL methods, which elucidates the contribution of unlabelled data. We review the state-of-the-art and observe that no previous work meets such requirements. To address this problem, we propose a framework for assessing the benefits of unlabelled data in SsL. We showcase an application of this framework by performing the first benchmark evaluation that highlights the tradeoffs of 9 existing SsL methods on 9 public datasets. Our findings verify that, in some cases, unlabelled data provides a small, but statistically significant, performance gain. This paper highlights that SsL in CTD has a lot of room for improvement, which should stimulate future research in this field. Giovanni Apruzzese, Pavel Laskov, Aliya Tastemirova |
EuroS&P | 2 |
| 2022 | Towards Understanding the Skill Gap in CybersecurityabstractGiven the ongoing "arms race" in cybersecurity, the shortage of skilled professionals in this field is one of the strongest in computer science. The currently unmet staffing demand in cybersecurity is estimated at over 3 million jobs worldwide. Furthermore, the qualifications of the existing workforce are largely believed to be insufficient. We attempt to gain deeper insights into the nature of the current skill gap in cybersecurity. To this end, we correlate data from job ads and academic curricula using two kinds of skill characterizations: manual definitions from established skill frameworks as well as "skill topics" automatically derived by text mining tools. Our analysis shows a strong agreement between these two analysis techniques and reveals a substantial undersupply in several crucial skill categories, e.g., software and application security, security management, requirements engineering, compliance and certification. Based on the results of our analysis, we provide recommendations for future curricula development in cybersecurity so as to decrease the identified skill gaps. François Goupil, Pavel Laskov, Irdin Pekaric, Michael Felderer, Alexander Dürr, Frédéric Thiesse |
ITiCSE (1) | 2 |
| 2022 | Wild Networks: Exposure of 5G Network Infrastructures to Adversarial ExamplesabstractFifth Generation (5G) networks must support billions of heterogeneous devices while guaranteeing optimal Quality of Service (QoS). Such requirements are impossible to meet with human effort alone, and Machine Learning (ML) represents a core asset in 5G. ML, however, is known to be vulnerable to adversarial examples; moreover, as our paper will show, the 5G context is exposed to a yet another type of adversarial ML attacks that cannot be formalized with existing threat models. Proactive assessment of such risks is also challenging due to the lack of ML-powered 5G equipment available for adversarial ML research. To tackle these problems, we propose a novel adversarial ML threat model that is particularly suited to 5G scenarios, and is agnostic to the precise function solved by ML. In contrast to existing ML threat models, our attacks do not require any compromise of the target 5G system while still being viable due to the QoS guarantees and the open nature of 5G networks. Furthermore, we propose an original framework for realistic ML security assessments based on public data. We proactively evaluate our threat model on 6 applications of ML envisioned in 5G. Our attacks affect both the training and the inference stages, can degrade the performance of state-of-the-art ML systems, and have a lower entry barrier than previous attacks. Giovanni Apruzzese, Rodion Vladimirov, Aliya Tastemirova, Pavel Laskov |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2021 | Detection of illicit cryptomining using network metadataabstractAbstract Illicit cryptocurrency mining has become one of the prevalent methods for monetization of computer security incidents. In this attack, victims’ computing resources are abused to mine cryptocurrency for the benefit of attackers. The most popular illicitly mined digital coin is Monero as it provides strong anonymity and is efficiently mined on CPUs.Illicit mining crucially relies on communication between compromised systems and remote mining pools using the de facto standard protocol Stratum. While prior research primarily focused on endpoint-based detection of in-browser mining, in this paper, we address network-based detection of cryptomining malware in general. We propose XMR-Ray, a machine learning detector using novel features based on reconstructing the Stratum protocol from raw NetFlow records. Our detector is trained offline using only mining traffic and does not require privacy-sensitive normal network traffic, which facilitates its adoption and integration.In our experiments, XMR-Ray attained 98.94% detection rate at 0.05% false alarm rate, outperforming the closest competitor. Our evaluation furthermore demonstrates that it reliably detects previously unseen mining pools, is robust against common obfuscation techniques such as encryption and proxies, and is applicable to mining in the browser or by compiled binaries. Finally, by deploying our detector in a large university network, we show its effectiveness in protecting real-world systems. Michele Russo, Nedim Srndic, Pavel Laskov |
EURASIP J. Inf. Secur. | 3 |
| 2016 | Hidost: a static machine-learning-based detector of malicious filesabstractMalicious software, i.e., malware, has been a persistent threat in the information security landscape since the early days of personal computing. The recent targeted attacks extensively use non-executable malware as a stealthy attack vector. There exists a substantial body of previous work on the detection of non-executable malware, including static, dynamic, and combined methods. While static methods perform orders of magnitude faster, their applicability has been hitherto limited to specific file formats. This paper introduces Hidost, the first static machine-learning-based malware detection system designed to operate on multiple file formats . Extending a previously published, highly effective method, it combines the logical structure of files with their content for even better detection accuracy. Our system has been implemented and evaluated on two formats, PDF and SWF (Flash). Thanks to its modular design and general feature set, it is extensible to other formats whose logical structure is organized as a hierarchy. Evaluated in realistic experiments on timestamped datasets comprising 440,000 PDF and 40,000 SWF files collected during several months, Hidost outperformed all antivirus engines deployed by the website VirusTotal to detect the highest number of malicious PDF files and ranked among the best on SWF malware. Nedim Srndic, Pavel Laskov |
EURASIP J. Inf. Secur. | 2 |
| 2014 | Practical Evasion of a Learning-Based Classifier: A Case StudyabstractLearning-based classifiers are increasingly used for detection of various forms of malicious data. However, if they are deployed online, an attacker may attempt to evade them by manipulating the data. Examples of such attacks have been previously studied under the assumption that an attacker has full knowledge about the deployed classifier. In practice, such assumptions rarely hold, especially for systems deployed online. A significant amount of information about a deployed classifier system can be obtained from various sources. In this paper, we experimentally investigate the effectiveness of classifier evasion using a real, deployed system, PDFrate, as a test case. We develop a taxonomy for practical evasion strategies and adapt known evasion algorithms to implement specific scenarios in our taxonomy. Our experimental results reveal a substantial drop of PDFrate's classification scores and detection accuracy after it is exposed even to simple attacks. We further study potential defense mechanisms against classifier evasion. Our experiments reveal that the original technique proposed for PDFrate is only effective if the executed attack exactly matches the anticipated one. In the discussion of the findings of our study, we analyze some potential techniques for increasing robustness of learning-based systems against adversarial manipulation of data. Nedim Srndic, Pavel Laskov |
IEEE Symposium on Security and Privacy | 2 |
| 2013 | Detection of Malicious PDF Files Based on Hierarchical Document Structure
Nedim Srndic, Pavel Laskov |
NDSS | 2 |
| 2013 | Evasion Attacks against Machine Learning at Test Time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, Fabio Roli |
ECML/PKDD (3) | 6 |
| 2012 | Poisoning Attacks against Support Vector Machines
Battista Biggio, Blaine Nelson, Pavel Laskov |
ICML | 3 |
| 2012 | Security analysis of online centroid anomaly detection
Marius Kloft, Pavel Laskov |
J. Mach. Learn. Res. | 2 |
| 2011 | Static detection of malicious JavaScript-bearing PDF documentsabstractDespite the recent security improvements in Adobe's PDF viewer, its underlying code base remains vulnerable to novel exploits. A steady flow of rapidly evolving PDF malware observed in the wild substantiates the need for novel protection instruments beyond the classical signature-based scanners. In this contribution we present a technique for detection of JavaScript-bearing malicious PDF documents based on static analysis of extracted JavaScript code. Compared to previous work, mostly based on dynamic analysis, our method incurs an order of magnitude lower run-time overhead and does not require special instrumentation. Due to its efficiency we were able to evaluate it on an extremely large real-life dataset obtained from the VirusTotal malware upload portal. Our method has proved to be effective against both known and unknown malware and suitable for large-scale batch processing. Pavel Laskov, Nedim Srndic |
ACSAC | 1 |
| 2010 | A factorization method for the classification of infrared spectraabstractBACKGROUND: Bioinformatics data analysis often deals with additive mixtures of signals for which only class labels are known. Then, the overall goal is to estimate class related signals for data mining purposes. A convenient application is metabolic monitoring of patients using infrared spectroscopy. Within an infrared spectrum each single compound contributes quantitatively to the measurement. RESULTS: In this work, we propose a novel factorization technique for additive signal factorization that allows learning from classified samples. We define a composed loss function for this task and analytically derive a closed form equation such that training a model reduces to searching for an optimal threshold vector. Our experiments, carried out on synthetic and clinical data, show a sensitivity of up to 0.958 and specificity of up to 0.841 for a 15-class problem of disease classification. Using class and regression information in parallel, our algorithm outperforms linear SVM for training cases having many classes and few data. CONCLUSIONS: The presented factorization method provides a simple and generative model and, therefore, represents a first step towards predictive factorization methods. Carsten Henneges, Pavel Laskov, Endang Darmawan, Juergen Backhaus, Bernd Kammerer, Andreas Zell |
BMC Bioinform. | 2 |
| 2010 | Machine learning in adversarial environments
Pavel Laskov, Richard Lippmann |
Mach. Learn. | 1 |
| 2009 | Cyber-Critical Infrastructure Protection Using Real-Time Payload-Based Anomaly Detection
Patrick Düssel, Christian Gehl, Pavel Laskov, Jens-Uwe Bußer, Christof Störmann, Jan Kästner |
CRITIS | 3 |
| 2009 | Efficient and Accurate Lp-Norm Multiple Kernel LearningabstractLearning linear combinations of multiple kernels is an appealing strategy when the right choice of features is unknown. Previous approaches to multiple kernel learning (MKL) promote sparse kernel combinations and hence support interpretability. Unfortunately, L1-norm MKL is hardly observed to outperform trivial baselines in practical applications. To allow for robust kernel mixtures, we generalize MKL to arbitrary Lp-norms. We devise new insights on the connection between several existing MKL formulations and develop two efficient interleaved optimization strategies for arbitrary p>1. Empirically, we demonstrate that the interleaved optimization strategies are much faster compared to the traditionally used wrapper approaches. Finally, we apply Lp-norm MKL to real-world problems from computational biology, showing that non-sparse MKL achieves accuracies that go beyond the state-of-the-art. Marius Kloft, Ulf Brefeld, Sören Sonnenburg, Pavel Laskov, Klaus-Robert Müller, Alexander Zien |
NIPS | 4 |
| 2008 | Learning and Classification of Malware Behavior
Konrad Rieck, Thorsten Holz, Carsten Willems, Patrick Düssel, Pavel Laskov |
DIMVA | 5 |
| 2008 | Stopping conditions for exact computation of leave-one-out error in support vector machinesabstractWe propose a new stopping condition for a Support Vector Machine (SVM) solver which precisely reflects the objective of the Leave-One-Out error computation. The stopping condition guarantees that the output on an intermediate SVM solution is identical to the output of the optimal SVM solution with one data point excluded from the training set. A simple augmentation of a general SVM training algorithm allows one to use a stopping criterion equivalent to the proposed sufficient condition. A comprehensive experimental evaluation of our method shows consistent speedup of the exact LOO computation by our method, up to the factor of 13 for the linear kernel. The new algorithm can be seen as an example of constructive guidance of an optimization algorithm towards achieving the best attainable expected risk at optimal computational cost. Vojtech Franc, Pavel Laskov, Klaus-Robert Müller |
ICML | 2 |
| 2008 | Linear-Time Computation of Similarity Measures for Sequential Data
Konrad Rieck, Pavel Laskov |
J. Mach. Learn. Res. | 2 |
| 2006 | Detecting Unknown Network Attacks Using Language Models
Konrad Rieck, Pavel Laskov |
DIMVA | 2 |
| 2006 | Computation of Similarity Measures for Sequential Data using Generalized Suffix TreesabstractWe propose a generic algorithm for computation of similarity measures for se- quential data. The algorithm uses generalized suffix trees for efficient calculation of various kernel, distance and non-metric similarity functions. Its worst-case run-time is linear in the length of sequences and independent of the underlying embedding language, which can cover words, k-grams or all contained subse- quences. Experiments with network intrusion detection, DNA analysis and text processing applications demonstrate the utility of distances and similarity coeffi- cients for sequences as alternatives to classical kernel functions. Konrad Rieck, Pavel Laskov, Sören Sonnenburg |
NIPS | 2 |
| 2006 | Incremental Support Vector Learning: Analysis, Implementation and ApplicationsabstractIncremental Support Vector Machines (SVM) are instrumental in practical applications of online learning. This work focuses on the design and analysis of efficient incremental SVM learning, with the aim of providing a fast, numerically stable and robust implementation. A detailed analysis of convergence and of algorithmic complexity of incremental SVM learning is carried out. Based on this analysis, a new design of storage and numerical operations is proposed, which speeds up the training of an incremental SVM by a factor of 5 to 20. The performance of the new algorithm is demonstrated in two scenarios: learning with limited resources and active learning. Various applications of the algorithm, such as in drug discovery, online monitoring of industrial devices and and surveillance of network traffic, can be foreseen. Pavel Laskov, Christian Gehl, Stefan Krüger, Klaus-Robert Müller |
J. Mach. Learn. Res. | 1 |
| 2004 | Using classification to determine the number of finger strokes on a multi-touch tactile device
Caspar von Wrede, Pavel Laskov |
ESANN | 2 |
| 2004 | A Fast Algorithm for Joint Diagonalization with Non-orthogonal Transformations and its Application to Blind Source Separation
Andreas Ziehe, Pavel Laskov, Guido Nolte, Klaus-Robert Müller |
J. Mach. Learn. Res. | 2 |
| 2003 | 3D nonrigid motion analysis under small deformations
Chandra Kambhamettu, Dmitry B. Goldgof, Matthew He, Pavel Laskov |
Image Vis. Comput. | 4 |
| 2003 | Curvature-Based Algorithms for Nonrigid Motion and Correspondence EstimationabstractWe present a novel technique for utilizing the Gaussian curvature information in 3D nonrigid motion estimation in the absence of known correspondence. Differential-geometric constraints derived in the paper allow one to estimate parameters of the local affine motion model given the values of Gaussian curvature before and after motion. These constraints can be further combined with the previously known constraints based on the unit normals before and after motion. Our experiments demonstrate that the resulting hybrid algorithm is more accurate than each of its constituents and more accurate than the classical ICP algorithm. We also present a technique for curvilinear orthogonalization of quadratic Monge patches that is essential in our derivation and useful in other applications. Pavel Laskov, Chandra Kambhamettu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Feasible Direction Decomposition Algorithms for Training Support Vector Machines
Pavel Laskov |
Mach. Learn. | 1 |
| 2001 | Comparison of 3D Algorithms for Non-rigid Motion and Correspondence EstimationabstractWe address the problem of non-rigid motion and correspondence estimation in 3D images in the absense of prior domain information. A generic framework is utilized in which a solution is approached by hypothesizing correspondence and evaluting the motion models constructed under each hypothesis. We present and evaluate experimentally ve algorithms that can be used in this approach. Our experiments were carried out on synthetic and real data with ground truth correspondence information. 1 Pavel Laskov, Chandra Kambhamettu |
BMVC | 1 |
| 1999 | An Improved Decomposition Algorithm for Regression Support Vector Machines
Pavel Laskov |
NIPS | 1 |
| 1996 | Recognition approach to gesture language understandingabstractWe explore recognition implications of understanding gesture communication, having chosen American sign language as an example of a gesture language. An instrumented glove and specially developed software have been used for data collection and labeling. We address the problem of recognizing dynamic signing, i.e. signing performed at natural speed. Two neural network architectures have been used for recognition of different types of finger-spelled sentences. Experimental results are presented suggesting that two features of signing affect recognition accuracy: signing frequency which to a large extent can be accounted for by training a network on the samples of the respective frequency; and coarticulation effect which a network fails to identify. As a possible solution to coarticulation problem two post-processing algorithms for temporal segmentation are proposed and experimentally evaluated. Roman Erenshteyn, Pavel Laskov, Richard A. Foulds, Lynn Messing, Garland Stern |
ICPR | 2 |