Raphaël Khoury

dblp:98/7410 · DBLP profile ↗
← Back
29ranked-venue papers
10as first author
14since 2021 · last 2025
0000-0002-7625-3384ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 5 first-author · 5 since 2021Security and privacy · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Mining Five Years of Actively Exploited Vulnerabilities
abstract
Companies and researchers rely heavily on the National Vulnerability Database (NVD) in order to be cognizant of the threat environment they face. However, research shows that only about 5% of reported vulnerabilities are eventually exploited. Our study compares exploited and non-exploited vulnerabilities to provide valuable insights for advancing research on effective vulnerability prioritization. To achieve this, we compile a database of over 4,000 exploited vulnerabilities, spanning 5 years, from 2019 to 2023. We further mine this dataset to uncover trends and patterns that relate to how exploited vulnerabilities differ from non-exploited ones and how exploited vulnerabilities evolved over the 5-year spans of our study. We found that exploited vulnerabilities differ from non-exploited vulnerabilities with respect to combinations of CVSS score attributes, but not with respect to the attributed CVSS score when considered in isolation. We further found that the CVSS scores of exploited vulnerabilities were largely stable over time for the duration of the study, and that widely used security resources are not concordant with the data we observed.
Raphaël Khoury, Kobra Khanmohammadi, Justin Vallé, Abdelwahab Hamou-Lhadj
COMPSAC1
2025 Evaluating the Effectiveness of ChatGPT for Analyzing Real-World Multi-File Cryptographic Vulnerabilities
Rezika Bouzid, Raphaël Khoury
CRiSIS2
2025 Beyond Detection: Evaluating LLMs' Semantic Understanding of Code Vulnerabilities
Charmant Nicolas Shangwe, Alan Davoust, Raphaël Khoury
CRiSIS3
2025 Enhancing Network Intrusion Detection Systems: A Multi-Layer Ensemble Approach to Mitigate Adversarial Attacks
abstract
Adversarial examples can represent a serious threat to machine learning (ML) algorithms. If used to manipulate the behaviour of ML-based Network Intrusion Detection Systems (NIDS), they can jeopardize network security. In this work, we aim to mitigate such risks by increasing the robustness of NIDS towards adversarial attacks. To that end, we explore two adversarial methods for generating malicious network traffic. The first method is based on Generative Adversarial Networks (GAN) and the second one is the Fast Gradient Sign Method (FGSM). The adversarial examples generated by these methods are then used to evaluate a novel multilayer defense mechanism, specifically designed to mitigate the vulnerability of ML-based NIDS. Our solution consists of one layer of stacking classifiers and a second layer based on an autoencoder. If the incoming network data are classified as benign by the first layer, the second layer is activated to ensure that the decision made by the stacking classifier is correct. We also incorporated adversarial training to further improve the robustness of our solution. Experiments on two datasets, namely UNSW-NB15 and NSL-KDD, demonstrate that the proposed approach increases resilience to adversarial attacks.
Nasim Soltani, Shayan Nejadshamsi, Zakaria Abou El Houda, Raphaël Khoury, Kelton A. P. Costa, Tiago H. Falk, Anderson R. Avila
SMC4
2024 Distributed and Verifiable Digital Badges
Raphaël Khoury, Sylvain Hallé, Muni Venkateswarlu K.
CRiSIS1
2023 How Secure is Code Generated by ChatGPT?
abstract
In recent years, large language models have been responsible for great advances in the field of artificial intelligence (AI). ChatGPT in particular, an AI chatbot developed and recently released by OpenAI, has taken the field to the next level. The conversational model is able not only to process human-like text, but also to translate natural language into code. However, the safety of programs generated by ChatGPT should not be overlooked. In this paper, we perform an experiment to address this issue. Specifically, we ask ChatGPT to generate a number of computer programs in order to evaluate the security of the resulting source code. We further investigate whether ChatGPT can be prodded to improve code security by appropriate prompts, and discuss the ethical aspects of using AI to generate code. Results suggest that ChatGPT is aware of potential vulnerabilities, but nonetheless often generates source code that are not robust to certain attacks.
Raphaël Khoury, Anderson R. Avila, Jacob Brunelle, Baba Mamadou Camara
SMC1
2023 Comparing the Effectiveness of Static, Dynamic and Hybrid Malware Detection on a Common Dataset
abstract
The detection of malicious Android applications is a major security challenge. A number of machine learning-based techniques have been put forth, and some of them have attained great accuracy. However, the diversity of apps and frequency at which new malware families are found means that the issue remains unresolved. In this paper, we use both static, dynamic and hybrid analysis to automatically classify Android apps as benign or infected. We compare all three approaches on a common dataset — the TwinDroid dataset which contains over 15,000 system call traces from over 9,000 benign and infected app. This method allows comparison on equal footing. We make further contributions on the topic of feature selection and trace abstraction.
Asma Razgallah, Raphaël Khoury, Kobra Khanmohammadi, Christophe Pere
SMC2
2022 TwinDroid: A Dataset of Android app System call traces and Trace Generation Pipeline
abstract
System call traces are an invaluable source of information about a program's runtime behavior and be particularly useful for malware detection in Android apps. However, the paucity of publicly available high-quality datasets hinders the development of the field. In this paper, we introduce TwinDroid, a dataset of over 1000 system calls traces, from both benign and infected Android apps. A large part of the apps used to create the dataset is from benign-malicious app pairs, identical apart from the inclusion of malware in the latter. This makes TwinDroid an ideal basis for security research, and an earlier version of TwinDroid has already been used for this purpose. In addition to a dataset of traces, TwinDroid includes a fully automated traces generation pipeline, which allows users to generate new traces in a standardized manner seamlessly. This pipeline will enable the dataset to remain up-to-date and relevant despite the rapid pace of change that characterizes Android security.
Asma Razgallah, Raphaël Khoury, Jean-Baptiste Poulet
MSR2
2022 A Taxonomy of Software Flaws Leading to Buffer Overflows
abstract
The buffer overflow attack has been dubbed ‘the vulnerability of the century’, because of the frequency and impact of this class of vulnerability. The wide variety of situations where this vulnerability can arise makes it particularly difficult to assess their occurrence or prevent them. In this paper, we present a novel taxonomy of programming errors which can lead to buffer overflows. This taxonomy easily translates into preconditions that ensure the code’s safe execution. We also illustrate each taxonomic class with a real-life example. Finally, from these examples, we draw a series of principles that developers can immediately incorporate in their programming habits in order to improve the security of their code.
Raphaël Khoury
QRS1
2021 Behavioral classification of Android applications using system calls
abstract
The exponential growth in the number of Android applications on the market has been matching with a corresponding growth in malicious application. Of particular concern is the risk of application repackaging, a process by which cy-bercriminals downloads, modifies and republishes an application that already exists on the store with the addition of malicious code. Dynamic detection in system call traces, based on machine learning models has emerged as a promising solution. In this paper, we introduce a novel abstraction process, and demonstrate that it improves the classification process by replicating multiples malware detection techniques from the literature. We further propose a novel classification method, based on our observation that malware triggers specific system calls at different points than benign programs. We further make our dataset available for future researchers.
Asma Razgallah, Raphaël Khoury
APSEC2
2021 Detecting trend deviations with generic stream processing patterns
Massiva Roudjane, Djamal Rebaïne, Raphaël Khoury, Sylvain Hallé
Inf. Syst.3
2021 The evolution of IoT Malwares, from 2008 to 2019: Survey, taxonomy, process simulator and perspectives
Benjamin Vignau, Raphaël Khoury, Sylvain Hallé, Abdelwahab Hamou-Lhadj
J. Syst. Archit.2
2021 Use of Security Logs for Data Leak Detection: A Systematic Literature Review
abstract
Security logs are widely used to monitor data, networks, and computer activities. By analyzing them, security experts can pick out anomalies that reveal the presence of cyber attacks or information leaks and stop them quickly before serious damage occurs. This paper presents a systematic literature review on the use of security logs for data leak detection. Our findings are fourfold: (i) we propose a new classification of information leaks, which uses the GDPR principles; (ii) we identify the twenty most widely used publicly available datasets in threat detection; (iii) we describe twenty types of attacks present in public datasets; and (iv) we describe thirty algorithms used for data leak detection. The selected papers point to many opportunities that can be investigated by researchers interested in contributing to this area of research.
Ricardo Ávila, Raphaël Khoury, Richard Khoury, Fábio Petrillo
Secur. Commun. Networks2
2021 Automata-based monitoring for LTL-FO+
Raphaël Khoury, Sylvain Hallé, Yannick Lebrun
Int. J. Softw. Tools Technol. Transf.1
2019 Predictive Analytics for Event Stream Processing
abstract
Historical data contained in event logs can reveal important insights about the execution of a business process. In particular, the trends computed by processing and analyzing the sequence of events generated by multiple instances of the same process serve as the basis to produce forecasts about current executions of the process. In this paper, we join the concepts of event stream processing and machine learning to create a framework that allows the computation of various kinds of predictions on event logs. The proposed framework is generic: by providing different definitions to a handful of event functions, multiple different types of predictions can be computed using the same basic workflow. The approach has been implemented and experimentally evaluated by extending an existing event stream processing engine.
Massiva Roudjane, Djamal Rebaïne, Raphaël Khoury, Sylvain Hallé
EDOC3
2019 Empirical study of android repackaged applications
Kobra Khanmohammadi, Neda Ebrahimi Koopaei, Abdelwahab Hamou-Lhadj, Raphaël Khoury
Empir. Softw. Eng.4
2018 Real-Time Data Mining for Event Streams
abstract
Information systems produce different types of event logs; in many situations, it may be desirable to look for trends inside these logs. We show how trends of various kinds can be computed over such logs in real time, using a generic framework called the trend distance workflow. Many common computations on event streams turn out to be special cases of this workflow, depending on how a handful of workflow parameters are defined. This process has been implemented and tested in a real-world event stream processing tool, called BeepBeep. Experimental results show that deviations from a reference trend can be detected in realtime for streams producing up to thousands of events per second.
Massiva Roudjane, Djamal Rebaïne, Raphaël Khoury, Sylvain Hallé
EDOC3
2018 Writing Domain-Specific Languages for BeepBeep
Sylvain Hallé, Raphaël Khoury
RV2
2018 Decentralized enforcement of document lifecycle constraints
Sylvain Hallé, Raphaël Khoury, Quentin Betti, Antoine El-Hokayem, Yliès Falcone
Inf. Syst.2
2017 SealTest: a simple library for test sequence generation
abstract
SealTest is a Java library for generating test sequences based on a formal specification. It allows a user to easily define a wide range of coverage metrics using multiple specification languages. Its simple and generic architecture makes it a useful testing tool for dynamic software systems, as well as an appropriate research testbed for implementing and experimentally comparing test sequence generation algorithms.
Sylvain Hallé, Raphaël Khoury
ISSTA2
2017 Event Stream Processing with Multiple Threads
Sylvain Hallé, Raphaël Khoury, Sébastien Gaboury
RV2
2016 A glue language for event stream processing
abstract
This paper describes the design and implementation of an SQL-like language for performing complex queries on event streams. The Event Stream Query Language (eSQL) aims at providing a simple, intuitive and fully non-procedural syntax, while still preserving backwards compatibility with traditional SQL. More importantly, eSQL's core syntax is designed to be extended by user-defined grammatical constrcts. These new constructs can form domain-specific sub-languages, with eSQL being used as the “glue” to form very expressive queries. These concepts have been implemented in BeepBeep 3, an open source event stream query engine.
Sylvain Hallé, Sébastien Gaboury, Raphaël Khoury
IEEE BigData3
2016 Decentralized Enforcement of Artifact Lifecycles
abstract
Artifact-centric workflows describe possible executions of a business process through constraints expressed from the point of view of the documents exchanged between principals. A sequence of manipulations is deemed valid as long as every document in the workflow follows its prescribed lifecycle at all steps of the process. So far, establishing that a given workflow complies with artifact lifecycles has mostly been done through static verification, or by assuming a centralized access to all artifacts where these constraints can be monitored and enforced. We present in this paper an alternate method of enforcing document lifecycles that requires neither static verification nor single-point access. Rather, the document itself is designed to carry fragments of its history, protected from tampering using hashing and public-key encryption. Any principal involved in the process can verify at any time that a document's history complies with a given lifecycle. Moreover, the proposed system also enforces access permissions: not all actions are visible to all principals, and one can only modify and verify what one is allowed to observe.
Sylvain Hallé, Raphaël Khoury, Antoine El-Hokayem, Yliès Falcone
EDOC2
2016 Key Elements Extraction and Traces Comprehension Using Gestalt Theory and the Helmholtz Principle
abstract
Trace analysis techniques are used by software engineers to understand the behaviour of large systems. This understanding can facilitate various software maintenance activities including debugging and feature enhancement. However, traces usually tend to be very large, which makes it difficult for software engineers to unveil the key logic and functionalities embedded in a program's execution. Hence, it is necessary to develop methods and tools that can efficiently identify the important information contained in a large trace. In this paper, we propose an approach that builds on the concept of trace segmentation to extract the major components of a traced scenario. Our approach is based on Gestalt theory and the Helmholtz principle. We show the effectiveness of our approach by applying it to a dataset of large traces.
Raphaël Khoury, Abdelwahab Hamou-Lhadj
ICSME1
2016 Execution Trace Analysis Using LTL-FO ^+
Raphaël Khoury, Sylvain Hallé, Omar Waldmann
ISoLA (2)1
2015 Equivalence-preserving corrective enforcement of security properties
abstract
Runtime monitoring is a widely used approach for the enforcement of security policies. It allows the safe execution of untrusted code by observing the execution and reacting if needed to prevent a violation of a user-defined security policy. Previous studies have determined that the set of security properties enforceable by monitors is greatly extended by giving the monitor some licence to transform its target execution. In this study, we present a new framework to model and study the behaviour of such monitors. In order to assure that the enforcement is meaningful, we bound the monitor's ability to transform the target execution by a restriction stating that any transformation must preserve equivalence between the monitor's input and output. We proceed by giving examples of meaningful equivalence relations and identify the security policies that are enforceable with their use. We also relate our work to previous work in this field. Finally, we investigate how an a priori knowledge of the target program's behaviour would increase the monitor's enforcement power.
Raphaël Khoury, Nadia Tawbi
Int. J. Inf. Comput. Secur.1
2012 Towards a formal framework for evaluating the effectiveness of system diversity when applied to security
abstract
N-version programming has been shown to be an effective way to increase the reliability of systems. In this study, we examine the possibility of extending this approach to address security, rather than reliability concerns. We focus specifically on how to evaluate the efficiency of the use of diversity for security. We show that while several key elements must be taken into account when N-version programming is used for security rather than reliability, it is nonetheless possible to devise a reasoning framework to evaluate the efficiency of this development paradigm in a security context. This framework allows us to reason about the most effective way to use diversity for security.
Raphaël Khoury, Abdelwahab Hamou-Lhadj, Mario Couture
CISDA1
2012 Corrective Enforcement: A New Paradigm of Security Policy Enforcement by Monitors
abstract
Runtime monitoring is an increasingly popular method to ensure the safe execution of untrusted codes. Monitors observe and transform the execution of these codes, responding when needed to correct or prevent a violation of a user-defined security policy. Prior research has shown that the set of properties monitors can enforce correlates with the latitude they are given to transform and alter the target execution. But for enforcement to be meaningful this capacity must be constrained, otherwise the monitor can enforce any property, but not necessarily in a manner that is useful or desirable. However, such constraints have not been significantly addressed in prior work. In this article, we develop a new paradigm of security policy enforcement in which the behavior of the enforcement mechanism is restricted to ensure that valid aspects present in the execution are preserved notwithstanding any transformation it may perform. These restrictions capture the desired behavior of valid executions of the program, and are stated by way of a preorder over sequences. The resulting model is closer than previous ones to what would be expected of a real-life monitor, from which we demand a minimal footprint on both valid and invalid executions. We illustrate this framework with examples of real-life security properties. Since several different enforcement alternatives of the same property are made possible by the flexibility of this type of enforcement, our study also provides metrics that allow the user to compare monitors objectively and choose the best enforcement paradigm for a given situation.
Raphaël Khoury, Nadia Tawbi
ACM Trans. Inf. Syst. Secur.1
2011 Extending the enforcement power of truncation monitors using static analysis
Hugues Chabot, Raphaël Khoury, Nadia Tawbi
Comput. Secur.2