EDBT 2026 Demo / reviewers in the wild / expert
Hadrien Pouget
dblp:274/0753
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2021
0000-0002-0678-1616ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 50% Reinforcement learning · 25% Robot navigation and mapping · 25% | |
| Software engineering, system software, and programming languages
2 papers |
Software testing · 87% Debugging and program repair · 13% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot navigation and mapping
fault detection |
0.5 | 1 | 2021 | Exposing previously undetectable faults in deep neural networks · ISSTA 2021 |
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
policy explanation |
0.5 | 1 | 2021 | Ranking Policy Decisions · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Exposing previously undetectable faults in deep neural networks · ISSTA 2021 |
Software testing
deep learning testing |
0.5 | 1 | 2021 | Exposing previously undetectable faults in deep neural networks · ISSTA 2021 |
Software testing
test input generation |
0.5 | 1 | 2021 | Exposing previously undetectable faults in deep neural networks · ISSTA 2021 |
Debugging and program repair › fault localization
statistical debugging |
0.1 | 1 | 2021 | Ranking Policy Decisions · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
statistical fault localization · 1.0generative machine learning · 1.0black-box ranking · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Exposing previously undetectable faults in deep neural networksabstractExisting methods for testing DNNs solve the oracle problem by constraining the raw features (e.g. image pixel values) to be within a small distance of a dataset example for which the desired DNN output is known. But this limits the kinds of faults these approaches are able to detect. In this paper, we introduce a novel DNN testing method that is able to find faults in DNNs that other methods cannot. The crux is that, by leveraging generative machine learning, we can generate fresh test inputs that vary in their high-level features (for images, these include object shape, location, texture, and colour). We demonstrate that our approach is capable of detecting deliberately injected faults as well as new faults in state-of-the-art DNNs, and that in both cases, existing methods are unable to find these faults. Isaac Dunn, Hadrien Pouget, Daniel Kroening, Tom Melham |
ISSTA | 2 |
| 2021 | Ranking Policy DecisionsabstractPolicies trained via Reinforcement Learning (RL) without human intervention are often needlessly complex, making them difficult to analyse and interpret. In a run with $n$ time steps, a policy will make $n$ decisions on actions to take; we conjecture that only a small subset of these decisions delivers value over selecting a simple default action. Given a trained policy, we propose a novel black-box method based on statistical fault localisation that ranks the states of the environment according to the importance of decisions made in those states. We argue that among other things, the ranked list of states can help explain and understand the policy. As the ranking method is statistical, a direct evaluation of its quality is hard. As a proxy for quality, we use the ranking to create new, simpler policies from the original ones by pruning decisions identified as unimportant (that is, replacing them by default actions) and measuring the impact on performance. Our experimental results on a diverse set of standard benchmarks demonstrate that pruned policies can perform on a level comparable to the original policies. We show that naive approaches for ranking policies, e.g. ranking based on the frequency of visiting a state, do not result in high-performing pruned policies. To the best of our knowledge, there are no similar techniques for ranking RL policies' decisions. Hadrien Pouget, Hana Chockler, Youcheng Sun, Daniel Kroening |
NeurIPS | 1 |