VLDB 2026 Research / reviewers in the wild / expert
Jérémie Guiochet
dblp:00/905
· DBLP profile ↗
27ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-1285-8974ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 9 · 6 since 2021Security and privacy · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unifying Runtime Monitoring Approaches for Safety-Critical Machine Learning: Application to Vision-Based Landing
Mathieu Dario, Florent Chenevier, Kevin Delmas, Joris Guérin, Jérémie Guiochet |
ICPR (5) | 5 |
| 2025 | Safety Monitoring of Machine Learning Perception Functions: A SurveyabstractABSTRACT Machine Learning (ML) models, such as deep neural networks, are widely applied in autonomous systems to perform complex perception tasks. New dependability challenges arise when ML predictions are used in safety‐critical applications, like autonomous cars and surgical robots. Thus, the use of fault tolerance mechanisms, such as safety monitors, is essential to ensure the safe behavior of the system despite the occurrence of faults. This paper presents an extensive literature review on safety monitoring of perception functions using ML in a safety‐critical context. In this review, we structure the existing literature to highlight key factors to consider when designing such monitors: threat identification, requirements elicitation, detection of failure, reaction, and evaluation. We also highlight the ongoing challenges associated with safety monitoring and suggest directions for future research. Raul Sena Ferreira, Joris Guérin, Kevin Delmas, Jérémie Guiochet, Hélène Waeselynck |
Comput. Intell. | 4 |
| 2024 | Can we Defend Against the Unknown? An Empirical Study About Threshold Selection for Neural Network MonitoringabstractWith the increasing use of neural networks in critical systems, runtime monitoring becomes essential to reject unsafe predictions during inference. Various techniques have emerged to establish rejection scores that maximize the separability between the distributions of safe and unsafe predictions. The efficacy of these approaches is mostly evaluated using threshold-agnostic metrics, such as the area under the receiver operating characteristic curve. However, in real-world applications, an effective monitor also requires identifying a good threshold to transform these scores into meaningful binary decisions. Despite the pivotal importance of threshold optimization, this problem has received little attention. A few studies touch upon this question, but they typically assume that the runtime data distribution mirrors the training distribution, which is a strong assumption as monitors are supposed to safeguard a system against potentially unforeseen threats. In this work, we present rigorous experiments on various image datasets to investigate: 1. The effectiveness of monitors in handling unforeseen threats, which are not available during threshold adjustments. 2. Whether integrating generic threats into the threshold optimization scheme can enhance the robustness of monitors. Khoi Tran Dang, Kevin Delmas, Jérémie Guiochet, Joris Guérin |
UAI | 3 |
| 2024 | Confidence assessment in safety argument structure - Quantitative vs. qualitative approaches
Yassir Idmessaoud, Didier Dubois, Jérémie Guiochet |
Int. J. Approx. Reason. | 3 |
| 2023 | Out-of-Distribution Detection Is Not All You NeedabstractThe usage of deep neural networks in safety-critical systems is limited by our ability to guarantee their correct behavior. Runtime monitors are components aiming to identify unsafe predictions and discard them before they can lead to catastrophic consequences. Several recent works on runtime monitoring have focused on out-of-distribution (OOD) detection, i.e., identifying inputs that are different from the training data. In this work, we argue that OOD detection is not a well-suited framework to design efficient runtime monitors and that it is more relevant to evaluate monitors based on their ability to discard incorrect predictions. We call this setting out-of-model-scope detection and discuss the conceptual differences with OOD. We also conduct extensive experiments on popular datasets from the literature to show that studying monitors in the OOD setting can be misleading: 1. very good OOD results can give a false impression of safety, 2. comparison under the OOD setting does not allow identifying the best monitor to detect errors. Finally, we also show that removing erroneous training data samples helps to train better monitors. Joris Guérin, Kevin Delmas, Raul Sena Ferreira, Jérémie Guiochet |
AAAI | 4 |
| 2023 | SENA: Similarity-Based Error-Checking of Neural ActivationsabstractIn this work, we propose SENA, a run-time monitor focused on detecting unreliable predictions from machine learning (ML) classifiers. The main idea is that instead of trying to detect when an image is out-of-distribution (OOD), which will not always result in a wrong output, we focus on detecting if the prediction from the ML model is not reliable, which will most of the time result in a wrong output, independently of whether it is in-distribution (ID) or OOD. The verification is done by checking the similarity between the neural activations of an incoming input and a set of representative neural activations recorded during training. SENA uses information from true-positive and false-negative examples collected during training to verify if a prediction is reliable or not. Our approach achieves results comparable to state-of-the-art solutions without requiring any prior OOD information and without hyperparameter tuning. Besides, the code is publicly available for easy reproducibility at https://github.com/raulsenaferreira/SENA. Raul Sena Ferreira, Joris Guérin, Jérémie Guiochet, Hélène Waeselynck |
ECAI | 3 |
| 2023 | Pairwise Testing Revisited for Structured Data With ConstraintsabstractPairwise testing (PT) exercises the interactions of pairs of input parameters. The approach is classically defined for a flat set of parameters, the number of which is fixed. Such a definition does not fit well with applications that process structured data like XML and JSON documents. This paper revisits the PT concepts to accommodate hierarchical data structures. The choices and pairs are created by considering the multiplicity of data instances, their access paths and common ancestors. The revised PT approach is implemented on top of on a recent data generation tool, TAF. TAF mixes random sampling and constraint solving to produce diverse data from XML-based models. Our PT implementation interacts with TAF by inserting pair coverage constraints into the models. It monitors overall coverage progress by XPath queries on the data returned by TAF. The approach is demonstrated for two data models: a 3D scene for an agricultural robot, and a population of taxpayers for a tax management system. Luca Vittorio Sartori, Hélène Waeselynck, Jérémie Guiochet |
ICST | 3 |
| 2022 | Integration of Test Generation Into Simulation-Based Platforms: An Experience ReportabstractField-testing is costly and time-consuming, hence, simulation-based testing is becoming more and more important to validate autonomous systems. Since autonomous systems can be deployed in diverse environments, a significant amount of diversified test cases has to be created. TAF (Testing Automation Framework) is a test generation tool we developed to serve this purpose. It produces the test cases from a data model that specifies the virtual environments of interest. This paper presents a practitioner's view of the integration of TAF into simulation-based test platforms, through two industrial case studies. The first one is for testing an agricultural robot developed by Naio Technologies, and the second one for a static perception system by SICK AG that surveils a road crossing to support connected vehicles with tracking data in complex urban scenarios. We report on our experience in the design of the data models, as well as in the automation of the execution, logging, and analysis of the generated tests. We conclude with lessons learned. Luca Vittorio Sartori, Jérémie Guiochet, Hélène Waeselynck, Aizar Antonio Berlanga Galvan, Simon Hébert-Vernhes, Magnus Albert |
AST | 2 |
| 2022 | Evaluation of Runtime Monitoring for UAV Emergency LandingabstractTo certify UAV operations in populated areas, risk mitigation strategies - such as Emergency Landing (EL) - must be in place to account for potential failures. EL aims at reducing ground risk by finding safe landing areas using on-board sensors. The first contribution of this paper is to present a new EL approach, in line with safety requirements introduced in recent research. In particular, the proposed EL pipeline includes mechanisms to monitor learning based components during execution. This way, another contribution is to study the behavior of Machine Learning Runtime Monitoring (MLRM) approaches within the context of a real-world critical system. A new evaluation methodology is introduced, and applied to assess the practical safety benefits of three MLRM mechanisms. The proposed approach is compared to a default mitigation strategy (open a parachute when a failure is detected), and appears to be much safer. Joris Guérin, Kevin Delmas, Jérémie Guiochet |
ICRA | 3 |
| 2022 | Unifying Evaluation of Machine Learning Safety MonitorsabstractWith the increasing use of Machine Learning (ML) in critical autonomous systems, runtime monitors have been developed to detect prediction errors and keep the system in a safe state during operations. Monitors have been proposed for different applications involving diverse perception tasks and ML models, and specific evaluation procedures and metrics are used for different contexts. This paper introduces three unified safety-oriented metrics, representing the safety benefits of the monitor (Safety Gain), the remaining safety gaps after using it (Residual Hazard), and its negative impact on the system's performance (Availability Cost). To compute these metrics, one requires to define two return functions, representing how a given ML prediction will impact expected future rewards and hazards. Three use-cases (classification, drone landing, and autonomous driving) are used to demonstrate how metrics from the literature can be expressed in terms of the proposed metrics. Experimental results on these examples show how different evaluation choices impact the perceived performance of a monitor. As our formalism requires us to formulate explicit safety assumptions, it allows us to ensure that the evaluation conducted matches the high-level system requirements. Joris Guérin, Raul Sena Ferreira, Kevin Delmas, Jérémie Guiochet |
ISSRE | 4 |
| 2022 | SiMOOD: Evolutionary Testing Simulation With Out-Of-Distribution ImagesabstractTesting perception functions for safety-critical autonomous systems is a crucial task. The reason is that accurate machine learning (ML) models applied in computer vision tasks still fail in scenarios where humans perform well. Out-of-distribution (OOD) images are usually a source of such failures. For this reason, literature usually applies data augmentation techniques or runtime monitors such as OOD detectors to increase robustness. Evaluating such solutions is usually performed by analyzing metrics based on positive and negative rates over a dataset containing several perturbations. However, using such metrics on such datasets can be misleading since not all OOD data lead to failures in the perception system. Hence, testing a perception system cannot be reduced to measuring ML performances on a dataset but rely on the images captured by the system at runtime. However, the amount of time spent to generate diverse test cases during a simulation of perception components can grow quickly since it is a combinatorial optimization problem. Aiming to provide a solution for this challenging task, we present SiMOOD, an evolutionary simulation testing of safety-critical perception systems, which comes integrated into the CARLA simulator. Unlike related works that simulate scenarios that raise failures for control or specific perception problems such as adversarial and novelty, we provide an approach that finds the most relevant OOD perturbations that can lead to hazards in safety-critical perception systems. Moreover, our approach can decrease, at least ten times, the amount of time to find a set of hazards in safety-critical scenarios such as autonomous emergency braking system simulation. Besides, code is publicly available for use. Raul Sena Ferreira, Joris Guérin, Jérémie Guiochet, Hélène Waeselynck |
PRDC | 3 |
| 2022 | Uncertainty Elicitation and Propagation in GSN Models of Assurance Cases
Yassir Idmessaoud, Didier Dubois, Jérémie Guiochet |
SAFECOMP | 3 |
| 2021 | A Fault Tolerant Control Architecture Based on Fault Trees for an Underwater Robot Executing Transect MissionsabstractRobotic systems evolving in hazardous and harsh environment are prone to mission failure or system loss in presence of faults. This paper presents a fault tolerant methodology, implemented into a control architecture of an underwater robot that executes biological monitoring missions. High level constraint violations (mission, safety, energy, time and localization) and low level faults (software and hardware faults) are considered using a method based on fault trees. These undesirable events are detected and treated by a fault tolerant module that decides to recover at low level or to give a feedback to the mission manager which selects the high level reaction. This fault tolerant architecture has been tested on real field conditions, and we illustrate our methodology on a set of selected events. We conclude about reliability improvement of low cost underwater robots for complex and long missions. Adrien Hereau, Karen Godary-Dejean, Jérémie Guiochet, Didier Crestani |
ICRA | 3 |
| 2021 | Benchmarking Safety Monitors for Image Classifiers with Machine LearningabstractHigh-accurate machine learning (ML) image classifiers cannot guarantee that they will not fail at operation. Thus, their deployment in safety-critical applications such as autonomous vehicles is still an open issue. The use of fault tolerance mechanisms such as safety monitors is a promising direction to keep the system in a safe state despite errors of the ML classifier. As the prediction from the ML is the core information directly impacting safety, many works are focusing on monitoring the ML model itself. Checking the efficiency of such monitors in the context of safety-critical applications is thus a significant challenge. Therefore, this paper aims at establishing a baseline framework for benchmarking monitors for ML image classifiers. Furthermore, we propose a framework covering the entire pipeline, from data generation to evaluation. Our approach measures monitor performance with a broader set of metrics than usually proposed in the literature. Moreover, we benchmark three different monitor approaches in 79 benchmark datasets containing five categories of out-of-distribution data for image classifiers: class novelty, noise, anomalies, distributional shifts, and adversarial attacks. Our results indicate that these monitors are no more accurate than a random monitor. We also release the code of all experiments for reproducibility. Raul Sena Ferreira, Jean Arlat, Jérémie Guiochet, Hélène Waeselynck |
PRDC | 3 |
| 2021 | TAF: a Tool for Diverse and Constrained Test Case GenerationabstractThe generation of test cases may have to accommodate size-varying data structures and semantic constraints between the data elements. This often requires the development of custom generators. In this paper, we introduce a novel generic tool to generate constrained and diverse test cases from a data model. First, the user defines the model using an XML-based domain-specific language. Then TAF generates diverse test cases by combining random sampling with the use of an SMT solver. The capabilities of the tool are demonstrated by four examples of models coming from various application domains: virtual crop fields for testing an agriculture robot, bitmap images with a graduated background, a population of taxpayers in a tax management system, and tree structures of diverse sizes and heights. We show how TAF performs in terms of data diversity and execution time. We also provide some comparison results with an UML-based tool using SMT solving. Clément Robert, Jérémie Guiochet, Hélène Waeselynck, Luca Vittorio Sartori |
QRS | 2 |
| 2020 | The virtual lands of Oz: testing an agribot in simulation
Clément Robert, Thierry Sotiropoulos, Hélène Waeselynck, Jérémie Guiochet, Simon Vernhes |
Empir. Softw. Eng. | 4 |
| 2019 | Safety case confidence propagation based on Dempster-Shafer theory
Rui Wang 0042, Jérémie Guiochet, Gilles Motet, Walter Schön |
Int. J. Approx. Reason. | 2 |
| 2018 | SMOF: A Safety Monitoring Framework for Autonomous SystemsabstractSafety-critical systems with decisional abilities, such as autonomous robots, are about to enter our everyday life. Nevertheless, confidence in their behavior is still limited, particularly regarding safety. Considering the variety of hazards that can affect these systems, many techniques might be used to increase their safety. Among them, active safety monitors are a means to maintain the system safety in spite of faults or adverse situations. The specification of the safety rules implemented in such devices is of crucial importance, but has been hardly explored so far. In this paper, we propose a complete framework for the generation of these safety rules based on the concept of safety margin. The approach starts from a hazard analysis, and uses formal verification techniques to automatically synthesize the safety rules. It has been successfully applied to an industrial use case, a mobile manipulator robot for co-working. Mathilde Machin, Jérémie Guiochet, Hélène Waeselynck, Jean-Paul Blanquart, Matthieu Roy, Lola Masson |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2017 | Can Robot Navigation Bugs Be Found in Simulation? An Exploratory StudyabstractThe ability to navigate in diverse and previously unknown environments is a critical service of autonomous robots. The validation of the navigation software typically involves test campaigns in the field, which are costly and potentially risky for the robot itself or its environment. An alternative approach is to perform simulation-based testing, by immersing the software in virtual worlds. A question is then whether the bugs revealed in real worlds can also be found in simulation. The paper reports on an exploratory study of bugs in an academic software for outdoor robots navigation. The detailed analysis of the triggers and effects of these bugs shows that most of them can be revealed in low-fidelity simulation. It also provides insights into interesting navigation scenarios to test as well as into how to address the test oracle problem. Thierry Sotiropoulos, Hélène Waeselynck, Jérémie Guiochet, Félix Ingrand |
QRS | 3 |
| 2017 | Confidence Assessment Framework for Safety Arguments
Rui Wang 0042, Jérémie Guiochet, Gilles Motet |
SAFECOMP | 2 |
| 2015 | A Model for Safety Case Confidence Assessment
Jérémie Guiochet, Quynh Anh Do Hoang, Mohamed Kaâniche |
SAFECOMP | 1 |
| 2014 | Specifying Safety Monitors for Autonomous Systems Using Model-Checking
Mathilde Machin, Fanny Dufossé, Jean-Paul Blanquart, Jérémie Guiochet, David Powell, Hélène Waeselynck |
SAFECOMP | 4 |
| 2014 | Towards privacy-driven design of a dynamic carpooling system
Jesus Friginal, Sébastien Gambs, Jérémie Guiochet, Marc-Olivier Killijian |
Pervasive Mob. Comput. | 3 |
| 2013 | Towards a Privacy Risk Assessment Methodology for Location-Based Systems
Jesus Friginal, Jérémie Guiochet, Marc-Olivier Killijian |
MobiQuitous | 2 |
| 2012 | Safety Trigger Conditions for Critical Autonomous SystemsabstractA systematic process for eliciting safety trigger conditions is presented. Starting from a risk analysis of the monitored system, critical transitions to catastrophic system states are identified and handled in order to specify safety margins on them. The conditions for existence of such safety margins are given and an alternative solution is proposed if no safety margin can be defined. The proposed process is illustrated on a robotic rollator. Amina Mekki-Mokhtar, Jean-Paul Blanquart, Jérémie Guiochet, David Powell, Matthieu Roy |
PRDC | 3 |
| 2007 | Fault Tolerant Planning for Critical RobotsabstractAutonomous robots offer alluring perspectives in numerous application domains: space rovers, satellites, medical assistants, tour guides, etc. However, a severe lack of trust in their dependability greatly reduces their possible usage. In particular, autonomous systems make extensive use of decisional mechanisms that are able to take complex and adaptative decisions, but are very hard to validate. This paper proposes a fault tolerance approach for decisional planning components, which are almost mandatory in complex autonomous systems. The proposed mechanisms focus on development faults in planning models and heuristics, through the use of diversification. The paper presents an implementation of these mechanisms on an existing autonomous robot architecture, and evaluates their impact on performance and reliability through the use of fault injection. Benjamin Lussier, Matthieu Gallien, Jérémie Guiochet, Félix Ingrand, Marc-Olivier Killijian, David Powell |
DSN | 3 |
| 2003 | Integration of UML in human factors analysis for safety of a medical robot for tele-echographyabstractFor new robot applications, as medical robots, safety has became a major concern. The human sharing the working area with the robot led to integrate the field of human factors in the development. Hence, the human component has to be integrated in the early steps of the development process. Regards to the complexity of today's robotic application, and to the requirements of a teamwork, we choose UML as the language. This paper focuses on the UML modeling contribution to the human factors analysis of a medical robot. A first section presents the function allocation and task analysis step, and a second section deals with human error. Each section is illustrated by a case study of a system for robotic tele-echography (ultrasound scan examination). Jérémie Guiochet, Bertrand Tondu, Claude Baron |
IROS | 1 |