EDBT 2026 Demo / reviewers in the wild / expert
Quan Le
dblp:09/439
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-6513-8340ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Certified Adversarial Robustness of End-to-End Malware Detectors via (De)Randomized SmoothingabstractEnd-to-end machine learning malware detectors are vulnerable to adversarial EXEmples, carefully-crafted malicious programs that evade detection through minimal perturbations. Such attacks typically operate by either replacing unused content (patch attacks), or injecting new patterns (content-injection attacks). To counter these attacks, recent research has focused on certification methods for end-to-end models that aim to prove robustness guarantees within bounded perturbation sizes. However, existing approaches (i) are not robust against content-injection manipulations and (ii) provide only probabilistic guarantees for perturbations that are negligible relative to the overall program size. Hence, in this paper we address these limitations through a novel deterministic certification schema based on (de)randomized smoothing. Our defense splits each executable into non-overlapping chunks and classifies them independently. The final decision is obtained via majority voting across all chunks, ensuring that localized modifications, such as injected or patched code, influence only a limited subset of chunks and have minimal impact in the overall classification. This design guarantees that each chunk either contains or does not contain an adversarial perturbation, enabling us to (i) handle manipulations occurring at arbitrary locations within the program, and (ii) compute deterministic estimates of the perturbation magnitude required to evade detection. We demonstrate the effectiveness of our certification schema through extensive experimental analysis, comparing our defense against a range of state-of-the-art attacks and defenses. The results show that our approach achieves unmatched robustness across all tested attack scenarios, substantially outperforming competing defenses. Daniel Gibert, Luca Demetrio, Giulio Zizzo, Quan Le, Jordi Planes, Battista Biggio |
ACM Trans. Priv. Secur. | 4 |
| 2025 | Assessing the impact of packing on static machine learning-based malware detection and classification systemsabstractThe proliferation of malware, particularly through the use of packing, presents a significant challenge to static analysis and signature-based malware detection techniques. Applying packing to the original executable code renders extracting meaningful features and signatures challenging. To deal with the increasing amount of malware in the wild, researchers and anti-malware companies started harnessing machine learning capabilities with very promising results. However, little is known about the effects of packing on static machine learning-based malware detection and classification systems. This work addresses this gap by investigating the impact of packing on the performance of static machine learning-based models used for malware detection and classification, with a particular focus on those using visualization techniques. To this end, we present a comprehensive analysis of various packing techniques and their effects on the performance of machine learning-based detectors and classifiers. Our findings highlight the limitations of current static detection and classification systems and underscore the need to be proactive to effectively counteract the evolving tactics of malware authors. Daniel Gibert, Nikolaos Totosis, Constantinos Patsakis, Quan Le, Giulio Zizzo |
Comput. Secur. | 4 |
| 2024 | Using GitHub Analytics to Assess the Quality of Collaboration in Software Engineering TeamsabstractThis research-to-practice full paper investigates using team process analytics from GitHub to support team management. Effective teamwork is essential in higher education learning and workplace success. The role of educators in supporting proper team functioning includes helping students learn how to participate actively and communicate effectively in meetings, delegate work fairly, manage high-quality work throughput, and resolve conflicts if problems arise. Problems often emerge when team members have differing visions or individuals do not contribute equally to the work output. These problems are exacerbated in large classes involving many teams. Detecting these potential issues in teams and helping students work through them is important for team success. In software engineering projects, monitoring individual contributions can begin with mining activities on programming platforms such as GitHub, which makes much of the individual contributions more visible and quantifiable. In this work, we propose a framework for fairly assessing teamwork and present the development of team analytics to assist educators in detecting potential issues in team collaboration. We describe a pilot study involving this tool in the context of a software engineering capstone course with 104 students split into 22 teams managed by 4 teaching assistants. Our results show that the reports offer value in guiding the evaluation process and identifying where problems may be, but do not replace the actual repository analysis where needed. We discuss the potential value of using this tool to improve collaboration. Quan Le, Kiet Phan, Bowen Hui, Adara Putri |
FIE | 1 |
| 2023 | American Stories: A Large-Scale Structured Text Dataset of Historical U.S. NewspapersabstractExisting full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other layout regions. OCR quality can also be low. This study develops a novel, deep learning pipeline for extracting full article texts from newspaper images and applies it to the nearly 20 million scans in Library of Congress's public domain Chronicling America collection. The pipeline includes layout detection, legibility classification, custom OCR, and association of article texts spanning multiple bounding boxes. To achieve high scalability, it is built with efficient architectures designed for mobile phones. The resulting American Stories dataset provides high quality data that could be used for pre-training a large language model to achieve better understanding of historical English and historical world knowledge. The dataset could also be added to the external database of a retrieval-augmented language model to make historical information - ranging from interpretations of political events to minutiae about the lives of people's ancestors - more widely accessible. Furthermore, structured article texts facilitate using transformer-based methods for popular social science applications like topic classification, detection of reproduced content, and news story clustering. Finally, American Stories provides a massive silver quality dataset for innovating multimodal layout analysis models and other multimodal applications. Melissa Dell, Jacob Carlson, Tom Bryan, Emily Silcock, Abhishek Arora 0003, Shannon Shen 0001, Luca D'Amico-Wong, Quan Le, Pablo Querubin, Leander Heldring |
NeurIPS | 8 |
| 2022 | Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
Daniel Gibert, Matt Fredrikson, Carles Mateu, Jordi Planes, Quan Le |
Comput. Secur. | 5 |
| 2022 | Fusing feature engineering and deep learning: A case study for malware classificationabstractMachine learning has become an appealing signature-less approach to detect and classify malware because of its ability to generalize to never-before-seen samples and to handle large volumes of data. While traditional feature-based approaches rely on the manual design of hand-crafted features based on experts’ knowledge of the domain, deep learning approaches replace the manual feature engineering process by an underlying system, typically consisting of a neural network with multiple layers, that perform both feature learning and classification altogether. However, the combination of both approaches could substantially enhance detection systems. In this paper we present an hybrid approach to address the task of malware classification by fusing multiple types of features defined by experts and features learned through deep learning from raw data. In particular, our approach relies on deep learning to extract N-gram like features from the assembly language instructions and the bytes of malware, and texture patterns and shapelet-based features from malware’s grayscale image representation and structural entropy, respectively. These deep features are later passed as input to a gradient boosting model that combines the deep features and the hand-crafted features using an early-fusion mechanism. The suitability of our approach has been evaluated on the Microsoft Malware Classification Challenge benchmark and results show that the proposed solution achieves state-of-the-art performance and outperforms gradient boosting and deep learning methods in the literature. Daniel Gibert, Jordi Planes, Carles Mateu, Quan Le |
Expert Syst. Appl. | 4 |
| 2017 | Protein multiple sequence alignment benchmarking through secondary structure predictionabstractMotivation: Multiple sequence alignment (MSA) is commonly used to analyze sets of homologous protein or DNA sequences. This has lead to the development of many methods and packages for MSA over the past 30 years. Being able to compare different methods has been problematic and has relied on gold standard benchmark datasets of 'true' alignments or on MSA simulations. A number of protein benchmark datasets have been produced which rely on a combination of manual alignment and/or automated superposition of protein structures. These are either restricted to very small MSAs with few sequences or require manual alignment which can be subjective. In both cases, it remains very difficult to properly test MSAs of more than a few dozen sequences. PREFAB and HomFam both rely on using a small subset of sequences of known structure and do not fairly test the quality of a full MSA. Results: In this paper we describe QuanTest, a fully automated and highly scalable test system for protein MSAs which is based on using secondary structure prediction accuracy (SSPA) to measure alignment quality. This is based on the assumption that better MSAs will give more accurate secondary structure predictions when we include sequences of known structure. SSPA measures the quality of an entire alignment however, not just the accuracy on a handful of selected sequences. It can be scaled to alignments of any size but here we demonstrate its use on alignments of either 200 or 1000 sequences. This allows the testing of slow accurate programs as well as faster, less accurate ones. We show that the scores from QuanTest are highly correlated with existing benchmark scores. We also validate the method by comparing a wide range of MSA alignment options and by including different levels of mis-alignment into MSA, and examining the effects on the scores. Availability and Implementation: QuanTest is available from http://www.bioinf.ucd.ie/download/QuanTest.tgz. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Quan Le, Fabian Sievers, Desmond G. Higgins |
Bioinform. | 1 |
| 2003 | Client Dependent GMM-SVM Models for Speaker Verification
Quan Le, Samy Bengio |
ICANN | 1 |