EDBT 2026 Demo / reviewers in the wild / expert
Helge Spieker
dblp:169/5121
· DBLP profile ↗
37ranked-venue papers
12as first author
28since 2021 · last 2026
0000-0003-2494-4279ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 15 since 2021Software engineering, systems software and programming languages · 15 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Metamorphic Relations Families as Types: Towards a New Foundation of Metamorphic Testing
Arnaud Gotlieb, Mathieu Le Louedec, Helge Spieker |
COMPSAC | 3 |
| 2026 | Metamorphic Testing with the Rashomon Set: Explanation Faithfulness in Machine LearningabstractMultiple machine learning models can achieve near-equivalent predictive performance on the same task, yet provide divergent feature-based explanations. This is called the Rashomon effect of (explainable) machine learning, and it raises the question of which explanations, if any, are trustworthy. We propose a framework based on metamorphic testing that assesses explanation faithfulness without requiring ground-truth labels by exploring attributed feature importance from post-hoc explanation methods. Five metamorphic relations formalize expected consistency properties between model behavior and feature attributions. We apply this general framework to two tabular regression datasets and two post-hoc explainers (SHAP and LIME) to demonstrate the approach. The framework offers a practical, model-agnostic tool for selecting accurate models with reliable and trustworthy explanations. Helge Spieker, Jørn Eirik Betten, Arnaud Gotlieb |
COMPSAC | 1 |
| 2026 | ScenaGen: A CP Model for Grounding Qualitative Driving ScenariosabstractValidating Automated Driving Systems (ADS) requires generating various kinematically executable traffic scenarios. The grounding of qualitative descriptions into concrete trajectories is a combinatorial task poorly addressed by learning-based methods. We propose ScenaGen, a CP model operating on qualitative explainable graphs (QXGs) to encode spatio-temporal relations between traffic entities. Formulated over integer position variables, ScenaGen enforces qualitative spatial constraints, distance thresholds, and inter-frame kinematic consistency. A single QXG acts as a formal template for systematically enumerating distinct, quantitatively varied concrete scenarios. Evaluation of synthetic and real-world benchmarks demonstrates that ScenaGen provides a robust and efficient alternative for scenario instantiation, outperforming standard search baselines in both scalability and solution diversity. Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker |
CP | 4 |
| 2026 | Context-Aware Autoencoders for Anomaly Detection in Maritime Surveillance
Divya Acharya, Pierre Bernabé, Antoine Chevrot, Helge Spieker, Arnaud Gotlieb, Bruno Legeard |
ICAART (3) | 4 |
| 2025 | Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies
Helge Spieker |
ICAART (2) | 2 |
| 2025 | Bounded PCTL Model Checking of Large Language Model OutputsabstractIn this paper, we introduce LLMchecker, a model-checking-based verification method to verify the probabilistic computation tree logic (PCTL) properties of an LLM text generation process. We empirically show that only a limited number of tokens are typically chosen during text generation, which are not always the same. This insight drives the creation of$\alpha-k$-bounded text generation, narrowing the focus to the$\alpha$maximal cumulative probability on the top-$k$tokens at every step of the text generation process. Our verification method considers an initial string and the subsequent top-$k$tokens while accommodating diverse text quantification methods, such as evaluating text quality and biases. The threshold$\alpha$further reduces the selected tokens, only choosing those that exceed or meet it in cumulative probability. LLMCHECKER then allows us to formally verify the PCTL properties of$\alpha-k$-bounded LLMs. We demonstrate the applicability of our method in several LLMs, including Llama, Gemma, Mistral, Genstruct, and BERT. To our knowledge, this is the first time PCTL-based model checking has been used to check the consistency of the LLM text generation process. Helge Spieker, Arnaud Gotlieb |
ICTAI | 2 |
| 2025 | Prompting for Performance: Exploring LLMs for Configuring SoftwareabstractSoftware systems usually provide numerous configuration options that can affect performance metrics such as execution time, memory usage, binary size, or bitrate. On the one hand, making informed decisions is challenging and requires domain expertise in options and their combinations. On the other hand, machine learning techniques can search vast configuration spaces, but with a high computational cost, since concrete executions of numerous configurations are required. In this exploratory study, we investigate whether large language models (LLMs) can assist in performance-oriented software configuration through prompts. We evaluate several LLMs on tasks including identifying relevant options, ranking configurations, and recommending performant configurations across various configurable systems, such as compilers, video encoders, and SAT solvers. Our preliminary results reveal both positive abilities and notable limitations: depending on the task and systems, LLMs can well align with expert knowledge, whereas hallucinations or superficial reasoning can emerge in other cases. These findings represent a first step toward systematic evaluations and the design of LLM-based solutions to assist with software configuration. Helge Spieker, Théo Matricon, Nassim Belmecheri, Jørn Eirik Betten, Gauthier Le Bartz Lyan, Heraldo Borges, Quentin Mazouni, Arnaud Gotlieb, Mathieu Acher |
ICTAI | 1 |
| 2025 | Automatic Cause Determination in Road Scene Understanding Using Qualitative Reasoning and Four-Valued LogicabstractRoad scene understanding in automated driving (AD) aims to build a comprehensive analysis of video sequences taken on the road by embedded or fixed cameras (e.g., mounted on vertical road signals). One goal is to identify the relevant actors in the scene and another goal is to determine the causes that have triggered a specific action of the ego car (i.e., stop, slow down, turn left, etc.). In a complex urban environment, these causes can be multiple, confusing, possibly contradictory to other causes and not easily expressible using simplistic reasoning. Still, providing accurate automatic cause determination supports a) user acceptance by providing appropriate explanations to the car passengers and road users; b) increased road safety by providing detailed road scene understanding to traffic. In this paper, we propose using spatiotemporal reasoning and Belnap's four-valued logic to formulate complex causes of AD action in a road scene. We compute these causes by analysing a Qualitative eXplainable Graph (QXG), which is an abstract representation of the road scene capturing spatiotemporal relations between road entities. Starting from a QXG, our approach called CAIDLOGIC, is targeted to determine complex causes of a selected AD action occurring in a specific frame of a road scene. The usefulness of CAIDLOGIC is demonstrated on several scenes extracted from the well-known NuScene dataset. Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker |
IV | 4 |
| 2025 | Explainable Scene Understanding with Qualitative Representations and Graph Neural NetworksabstractThis paper investigates the integration of graph neural networks (GNNs) with Qualitative Explainable Graphs (QXGs) for scene understanding in automated driving. Scene understanding is the basis for any further reactive or proactive decision-making. Scene understanding and related reasoning is inherently an explanation task: why is another traffic participant doing something, what or who caused their actions? While previous work demonstrated QXGs' effectiveness using shallow machine learning models, these approaches were limited to analysing single relation chains between object pairs, disregarding the broader scene context. We propose a novel GNN architecture that processes entire graph structures to identify relevant objects in traffic scenes. We evaluate our method on the nuScenes dataset enriched with DriveLM's human-annotated relevance labels. Experimental results show that our GNN-based approach achieves superior performance compared to baseline methods. The model effectively handles the inherent class imbalance in relevant object identification tasks while considering the complete spatial-temporal relationships between all objects in the scene. Our work demonstrates the potential of combining qualitative representations with deep learning approaches for explainable scene understanding in autonomous driving systems. Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker |
IV | 4 |
| 2025 | Reusable Test Suites for Reinforcement Learning
Jørn Eirik Betten, Quentin Mazouni, Pedro G. Lind, Helge Spieker |
ICTSS | 5 |
| 2025 | Metamorphic Testing of Multimodal Human Trajectory Prediction
Helge Spieker, Nadjib Lazaar, Arnaud Gotlieb, Nassim Belmecheri |
Inf. Softw. Technol. | 1 |
| 2025 | Mutation-Guided Metamorphic Testing of Optimality in AI PlanningabstractABSTRACT Autonomous systems such as space‐ or underwater‐exploration robots or elderly people assistance robots often include an artificial intelligence (AI) planner as a component. Starting from the initial state of a system, an AI planner automatically generates sequential plans to reach final states that satisfy user‐specified goals. Generating plans having a minimum number of intermediate steps or taking the least time to execute is usually strongly desired, as these plans exhibit minimal costs. Unfortunately, testing if an AI planner generates optimal plans is almost impossible because the expected cost of these plans is usually unknown. Based on mutation adequacy test suite selection, this article proposes a novel metamorphic testing framework for detecting the lack of optimality in AI planners. The general idea is to perform a systematic but non‐exhaustive state space exploration from the initial state and to select mutant‐adequate states to instantiate new planning tasks as follow‐up test cases. We then check a metamorphic relation between the automatically generated solutions of the AI planner for these new test cases and the cost of the initial plan. We implemented this metamorphic testing framework in a tool called MorphinPlan. Our experimental evaluation shows that MorphinPlan can detect non‐optimal behaviour in both mutated AI planners and off‐the‐shelf, configurable planners. It also shows that our proposed mutation adequacy test selection strategy outperforms three alternative test generation and selection strategies, including both random state selection and random walks through the state space in terms of mutation scores. Quentin Mazouni, Arnaud Gotlieb, Helge Spieker, Mathieu Acher, Benoît Combemale |
Softw. Test. Verification Reliab. | 3 |
| 2024 | Testing for Fault Diversity in Reinforcement LearningabstractReinforcement Learning is the premier technique to approach sequential decision problems, including complex tasks such as driving cars and landing spacecraft. Among the software validation and verification practices, testing for functional fault detection is a convenient way to build trustworthiness in the learned decision model. While recent works seek to maximise the number of detected faults, none consider fault characterisation during the search for more diversity. We argue that policy testing should not find as many failures as possible (e.g., inputs that trigger similar car crashes) but rather aim at revealing as informative and diverse faults as possible in the model. In this paper, we explore the use of quality diversity optimisation to solve the problem of fault diversity in policy testing. Quality diversity (QD) optimisation is a type of evolutionary algorithm to solve hard combinatorial optimisation problems where high-quality diverse solutions are sought. We define and address the underlying challenges of adapting QD optimisation to the test of action policies. Furthermore, we compare classical QD optimisers to state-of-the-art frameworks dedicated to policy testing, both in terms of search efficiency and fault diversity. We show that QD optimisation, while being conceptually simple and generally applicable, finds effectively more diverse faults in the decision model, and conclude that QD-based policy testing is a promising approach. Quentin Mazouni, Helge Spieker, Arnaud Gotlieb, Mathieu Acher |
AST | 2 |
| 2024 | Safety-Oriented Pruning and Interpretation of Reinforcement Learning PoliciesabstractPruning neural networks (NNs) can streamline them but risks removing vital parameters from safe reinforcement learning (RL) policies.We introduce an interpretable RL method called VERINTER, which combines NN pruning with model checking to ensure interpretable RL safety.VERINTER exactly quantifies the effects of pruning and the impact of neural connections on complex safety properties by analyzing changes in safety measurements.This method maintains safety in pruned RL policies and enhances understanding of their safety dynamics, which has proven effective in multiple RL settings. Helge Spieker |
ESANN | 2 |
| 2024 | Probabilistic Model Checking of Stochastic Reinforcement Learning Policies
Helge Spieker |
ICAART (3) | 2 |
| 2024 | Enhancing Manufacturing Quality Prediction Models Through the Integration of Explainability Methods
Helge Spieker, Arnaud Gotlieb, Ricardo Knoblauch |
ICAART (3) | 2 |
| 2024 | Policy Testing with MDPFuzz (Replicability Study)abstractIn recent years, following tremendous achievements in Reinforcement Learning, a great deal of interest has been devoted to ML models for sequential decision-making. Together with these scientific breakthroughs/advances, research has been conducted to develop automated functional testing methods for finding faults in black-box Markov decision processes. Pang et al. (ISSTA 2022) presented a black-box fuzz testing framework called MDPFuzz. The method consists of a fuzzer whose main feature is to use Gaussian Mixture Models (GMMs) to compute coverage of the test inputs as the likelihood to have already observed their results. This guidance through coverage evaluation aims at favoring novelty during testing and fault discovery in the decision model. Pang et al. evaluated their work with four use cases, by comparing the number of failures found after twelve-hour testing campaigns with or without the guidance of the GMMs (ablation study). In this paper, we verify some of the key findings of the original paper and explore the limits of MDPFuzz through reproduction and replication. We re-implemented the proposed methodology and evaluated our replication in a large-scale study that extends the original four use cases with three new ones. Furthermore, we compare MDPFuzz and its ablated counterpart with a random testing baseline. We also assess the effectiveness of coverage guidance for different parameters, something that has not been done in the original evaluation. Despite this parameter analysis and unlike Pang et al.’s original conclusions, we find that in most cases, the aforementioned ablated Fuzzer outperforms MDPFuzz, and conclude that the coverage model proposed does not lead to finding more faults. Quentin Mazouni, Helge Spieker, Arnaud Gotlieb, Mathieu Acher |
ISSTA | 2 |
| 2024 | Enhancing RL Safety with Counterfactual LLM Reasoning
Helge Spieker |
ICTSS | 2 |
| 2024 | Query-driven Qualitative Constraint AcquisitionabstractMany planning, scheduling or multi-dimensional packing problems involve the design of subtle logical combinations of temporal or spatial constraints. Recently, we introduced GEQCA-I, which stands for Generic Qualitative Constraint Acquisition, as a new active constraint acquisition method for learning qualitative constraints using qualitative queries. In this paper, we revise and extend GEQCA-I to GEQCA-II with a new type of query, universal query, for qualitative constraint acquisition, with a deeper query-driven acquisition algorithm. Our extended experimental evaluation shows the efficiency and usefulness of the concept of universal query in learning randomly-generated qualitative networks, including both temporal networks based on Allen’s algebra and spatial networks based on region connection calculus. We also show the effectiveness of GEQCA-II in learning the qualitative part of real scheduling problems. Mohamed-Bachir Belaid, Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker |
J. Artif. Intell. Res. | 5 |
| 2024 | Learning input-aware performance models of configurable systems: An empirical evaluation
Luc Lesoil, Helge Spieker, Arnaud Gotlieb, Mathieu Acher, Paul Temple, Arnaud Blouin, Jean-Marc Jézéquel |
J. Syst. Softw. | 2 |
| 2024 | Detecting Intentional AIS Shutdown in Open Sea Maritime Surveillance Using Self-Supervised Deep LearningabstractIn maritime traffic surveillance, detecting illegal activities, such as illegal fishing or transshipment of illicit products is a crucial task of the coastal administration. In the open sea, one has to rely on Automatic Identification System (AIS) message transmitted by on-board transponders, which are captured by surveillance satellites. However, insincere vessels often intentionally shut down their AIS transponders to hide illegal activities. In the open sea, it is very challenging to differentiate intentional AIS shutdowns from missing reception due to protocol limitations, bad weather conditions or restricting satellite positions. This paper presents a novel approach for the detection of abnormal AIS missing reception based on self-supervised deep learning techniques and transformer models. Using historical data, the trained model predicts if a message should be received in the upcoming minute or not. Afterwards, the model reports on detected anomalies by comparing the prediction with what actually happens. Our method can process AIS messages in real-time, in particular, more than 500 Millions AIS messages per month, corresponding to the trajectories of more than 60 000 ships. The method is evaluated on 1-year of real-world data coming from four Norwegian surveillance satellites. Using related research results, we validated our method by rediscovering already detected intentional AIS shutdowns. Pierre Bernabé, Arnaud Gotlieb, Bruno Legeard, Dusica Marijan, Frank Olaf Sem-Jacobsen, Helge Spieker |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Constraint-Guided Test Execution Scheduling: An Experience Report at ABB Robotics
Arnaud Gotlieb, Morten Mossige, Helge Spieker |
SAFECOMP | 3 |
| 2022 | GEQCA: Generic Qualitative Constraint AcquisitionabstractMany planning, scheduling or multi-dimensional packing problems involve the design of subtle logical combinations of temporal or spatial constraints. On the one hand, the precise modelling of these constraints, which are formulated in various relation algebras, entails a number of possible logical combinations and requires expertise in constraint-based modelling. On the other hand, active constraint acquisition (CA) has been used successfully to support non-experienced users in learning conjunctive constraint networks through the generation of a sequence of queries. In this paper, we propose GEACQ, which stands for Generic Qualitative Constraint Acquisition, an active CA method that learns qualitative constraints via the concept of qualitative queries. GEACQ combines qualitative queries with time-bounded path consistency (PC) and background knowledge propagation to acquire the qualitative constraints of any scheduling or packing problem. We prove soundness, completeness and termination of GEACQ by exploiting the jointly exhaustive and pairwise disjoint property of qualitative calculus and we give an experimental evaluation that shows (i) the efficiency of our approach in learning temporal constraints and, (ii) the use of GEACQ on real scheduling instances. Mohamed-Bachir Belaid, Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker |
AAAI | 5 |
| 2022 | FoCA: Failure-oriented Class Augmentation for Robust Image ClassificationabstractImage classification with classes of varying difficulty can cause performance disparity in deep learning models and reduce the overall performance and reliability of the predictions. In this paper, we introduce a failure-oriented class augmentation (FoCA) technique to address the problem of imbalanced performance in image classification, where the trained model has performance deficits in some of the dataset's classes. By employing Generative Adversarial Networks (GANs) to augment these deficit classes, we finetune the model towards a balanced performance among the different classes and an overall better performance on the whole dataset. Unlike earlier works, during training, our method focuses on those classes with the lowest accuracy after the initial training phase. Only these classes are augmented to boost the accuracy, which leads to better performance. FoCA is designed to be used with a light-weight GAN method to make the GAN-based augmentation viable and effective, even for datasets with only few images per class, while simultaneously requiring less computation than other, more complex GAN methods. Our implementation of FoCA combines this light-weight GAN method for class-wise data augmentation with state-of-the-art deep neural network techniques for training. Experiments show an overall improvement from FoCA with competitive or better accuracy than the previous state-of-the-art on five datasets with different sizes and image resolutions. Mohit Kumar Ahuja, Sahil Sahil, Helge Spieker |
ICTAI | 3 |
| 2022 | A fine-grained data set and analysis of tangling in bug fixing commitsabstractAbstract Context Tangled commits are changes to software that address multiple concerns at once. For researchers interested in bugs, tangled commits mean that they actually study not only bugs, but also other concerns irrelevant for the study of bugs. Objective We want to improve our understanding of the prevalence of tangling and the types of changes that are tangled within bug fixing commits. Methods We use a crowd sourcing approach for manual labeling to validate which changes contribute to bug fixes for each line in bug fixing commits. Each line is labeled by four participants. If at least three participants agree on the same label, we have consensus. Results We estimate that between 17% and 32% of all changes in bug fixing commits modify the source code to fix the underlying problem. However, when we only consider changes to the production code files this ratio increases to 66% to 87%. We find that about 11% of lines are hard to label leading to active disagreements between participants. Due to confirmed tangling and the uncertainty in our data, we estimate that 3% to 47% of data is noisy without manual untangling, depending on the use case. Conclusion Tangled commits have a high prevalence in bug fixes and can lead to a large amount of noise in the data. Prior research indicates that this noise may alter results. As researchers, we should be skeptics and assume that unvalidated data is likely very noisy, until proven otherwise. Steffen Herbold, Alexander Trautsch, Benjamin Ledel, Alireza Aghamohammadi, Taher Ahmed Ghaleb, Kuljit Kaur Chahal, Tim Bossenmaier, Bhaveet Nagaria, Philip Makedonski, Matin Nili Ahmadabadi, Kristóf Szabados, Helge Spieker, Matej Madeja, Nathaniel Hoy, Valentina Lenarduzzi, Shangwen Wang, Gema Rodríguez-Pérez, Ricardo Colomo-Palacios, Roberto Verdecchia, Paramvir Singh, Yihao Qin, Debasish Chakroborti, Willard Davis, Vijay Walunj, Diego Marcilio, Omar Alam, Abdullah Aldaeej, Idan Amit, Burak Turhan, Simon Eismann, Anna-Katharina Wickert, Ivano Malavolta, Matús Sulír, Fatemeh Hendijani Fard, Austin Z. Henley, Stratos Kourtzanidis, Eray Tüzün, Christoph Treude, Simin Maleki Shamasbi, Ivan Pashchenko, Marvin Wyrich, James C. Davis 0001, Alexander Serebrenik, Ella Albrecht, Ethem Utku Aktas, Daniel Strüber 0001, Johannes Erbel |
Empir. Softw. Eng. | 12 |
| 2021 | Encoding Temporal and Spatial Vessel Context using Self-Supervised Learning Model (Student Abstract)abstractMaritime surveillance is essential to avoid illegal activities and for environmental protection. However, the unlabeled, noisy, irregular time-series data and the large area to be covered make it challenging to detect illegal activities. Existing solutions focus only on trajectory reconstruction and probabilistic models that do ignore the context, such as the neighboring vessels. We propose a novel representation learning method that considers both temporal and spatial contexts learned in a self-supervised manner, using a selection of pretext tasks that do not require to be labeled manually. The underlying model encodes the representation of maritime vessel data compactly and effectively. This generic encoder can then be used as input for more complex tasks lacking labeled data. Pierre Bernabé, Helge Spieker, Bruno Legeard, Arnaud Gotlieb |
AAAI | 2 |
| 2021 | Summary of: Adaptive Metamorphic Testing with Contextual BanditsabstractMetamorphic Testing (MT) is a software testing paradigm that aims at using user-specified properties of a program under test to either check its expected outputs or to generate new test cases [1] , [2] . More precisely, MT tackles the so-called oracle problem which occurs whenever predicting the expected outputs of a system is just too difficult or even impossible. A typical example where MT has been successfully deployed is for testing machine learning models. For instance, in supervised machine learning, we train models for classification problems, but testing these models is hard as only stochastic behaviors of these models can be specified [3] . Indeed, we initially train these models with existing labelled datasets and then we exploit them to classify new data samples. Testing these models means only to reserve some portion of the labelled datasets to control that the correct classification is given for these reserved datasets. However, nothing is really available to test these models on unlabelled data samples. Helge Spieker, Arnaud Gotlieb |
ICST | 1 |
| 2021 | Constraint-Guided Reinforcement Learning: Augmenting the Agent-Environment-InteractionabstractReinforcement Learning (RL) agents have great successes in solving tasks with large observation and action spaces from limited feedback. Still, training the agents is data-intensive and there are no guarantees that the learned behavior is safe and does not violate rules of the environment, which has limitations for the practical deployment in real-world scenarios. This paper discusses the engineering of reliable agents via the integration of deep RL with constraint-based augmentation models to guide the RL agent towards safe behavior. Within the constraints set, the RL agent is free to adapt and explore, such that its effectiveness to solve the given problem is not hindered. However, once the RL agent leaves the space defined by the constraints, the outside models can provide guidance to still work reliably. We discuss integration points for constraint guidance within the RL process and perform experiments on two case studies: a strictly constrained card game and a grid world environment with additional combinatorial subgoals. Our results show that constraint-guidance does both provide reliability improvements and safer behavior, as well as accelerated training. Helge Spieker |
IJCNN | 1 |
| 2020 | Adaptive metamorphic testing with contextual bandits
Helge Spieker, Arnaud Gotlieb |
J. Syst. Softw. | 1 |
| 2019 | Towards Sequence-to-Sequence Reinforcement Learning for Constraint Solving with Constraint-Based Local Search
Helge Spieker |
AAAI | 1 |
| 2019 | Rotational Diversity in Multi-Cycle Assignment ProblemsabstractIn multi-cycle assignment problems with rotational diversity, a set of tasks has to be repeatedly assigned to a set of agents. Over multiple cycles, the goal is to achieve a high diversity of assignments from tasks to agents. At the same time, the assignments’ profit has to be maximized in each cycle. Due to changing availability of tasks and agents, planning ahead is infeasible and each cycle is an independent assignment problem but influenced by previous choices. We approach the multi-cycle assignment problem as a two-part problem: Profit maximization and rotation are combined into one objective value, and then solved as a General Assignment Problem. Rotational diversity is maintained with a single execution of the costly assignment model. Our simple, yet effective method is applicable to different domains and applications. Experiments show the applicability on a multi-cycle variant of the multiple knapsack problem and a real-world case study on the test case selection and assignment problem, an example from the software engineering domain, where test cases have to be distributed over compatible test machines. Helge Spieker, Arnaud Gotlieb, Morten Mossige |
AAAI | 1 |
| 2018 | Different Cycle, Different Assignment: Diversity in Assignment Problems With Multiple CyclesabstractWe present approaches to handle diverse assignments in multi-cycle assignment problems. The goal is to assign a task to different agents in each cycle, such that all possible combinations are made over time. Our method combines the original profit value, that is to be optimized by the assignment problem with an additional assignment preference. By merging both, we steer the optimization towards diverse assignments without large trade-offs in the original profits. Helge Spieker, Arnaud Gotlieb, Morten Mossige |
AAAI | 1 |
| 2018 | Stratified Constructive Disjunction and Negation in Constraint ProgrammingabstractConstraint Programming (CP) is a powerful declarative programming paradigm combining inference and search in order to find solutions to various type of constraint systems. Dealing with highly disjunctive constraint systems is notoriously difficult in CP. Apart from trying to solve each disjunct independently from each other, there is little hope and effort to succeed in constructing intermediate results combining the knowledge originating from several disjuncts. In this paper, we propose If-Then-Else (ITE), a lightweight approach for implementing stratified constructive disjunction and negation on top of an existing CP solver, namely SICStus Prolog clpfd. Although constructive disjunction is known for more than three decades, it does not have straightforward implementations in most CP solvers. ITE is a freely available library proposing stratified and constructive reasoning for various operators, including disjunction and negation, implication and conditional. Our preliminary experimental results show that ITE is competitive with existing approaches that handle disjunctive constraint systems. Arnaud Gotlieb, Dusica Marijan, Helge Spieker |
ICTAI | 3 |
| 2017 | Time-Aware Test Case Execution Scheduling for Cyber-Physical Systems
Morten Mossige, Arnaud Gotlieb, Helge Spieker, Hein Meling, Mats Carlsson |
CP | 3 |
| 2017 | Reinforcement learning for automatic test case prioritization and selection in continuous integrationabstractTesting in Continuous Integration (CI) involves test case prioritization, selection, and execution at each cycle. Selecting the most promising test cases to detect bugs is hard if there are uncertainties on the impact of committed code changes or, if traceability links between code and tests are not available. This paper introduces Retecs, a new method for automatically learning test case selection and prioritization in CI with the goal to minimize the round-trip time between code commits and developer feedback on failed test cases. The Retecs method uses reinforcement learning to select and prioritize test cases according to their duration, previous last execution and failure history. In a constantly changing environment, where new test cases are created and obsolete test cases are deleted, the Retecs method learns to prioritize error-prone test cases higher under guidance of a reward function and by observing previous CI cycles. By applying Retecs on data extracted from three industrial case studies, we show for the first time that reinforcement learning enables fruitful automatic adaptive test case selection and prioritization in CI and regression testing. Helge Spieker, Arnaud Gotlieb, Dusica Marijan, Morten Mossige |
ISSTA | 1 |
| 2017 | Multi-stage evolution of single- and multi-objective MCLP - Successive placement of charging stations
Helge Spieker, Alexander Hagg, Adam Gaier, Stefanie Meilinger, Alexander Asteroth |
Soft Comput. | 1 |
| 2015 | Successive evolution of charging station placementabstractAn evolving strategy for a multi-stage placement of charging stations for electrical cars is developed. Both an incremental as well as a decremental placement decomposition are evaluated on this Maximum Covering Location Problem. We show that an incremental Genetic Algorithm benefits from problem decomposition effects of having multiple stages and shows greedy behaviour. Helge Spieker, Alexander Hagg, Alexander Asteroth, Stefanie Meilinger, Volker Jacobs, Alexander Oslislo |
INISTA | 1 |