VLDB 2026 Research / reviewers in the wild / expert
Markus Bohlin
dblp:87/6323
· DBLP profile ↗
22ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-1597-6738ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-authorArtificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Theory of computation · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Requirements Ambiguity Detection and Explanation with LLMS: An Industrial StudyabstractDeveloping large-scale industrial systems requires high-quality requirements to avoid costly rework and project delays. However, linguistic ambiguities in natural language (NL) requirements have been a long-standing challenge, often introducing misinterpretations and inconsistencies that propagate throughout the development lifecycle. Such ambiguous NL requirements necessitate early detection and well-reasoned explanations to clarify and prevent further misunderstandings among stakeholders. While solutions have been developed to detect ambiguities in NL requirements, the advent of generative large language models (LLMs) offers new avenues for explanation-augmented requirements ambiguity detection. This paper empirically investigates LLMs for ambiguity detection and explanation in real-world industrial requirements by adopting an in-context learning paradigm. Our results from three industrial datasets show that LLMs achieve a 20.2% average performance increase in classifying ambiguous requirements when prompted with ten relevant in-context demonstrations (10 -shot), compared to no demonstrations (0 -shot). Additionally, we conducted human evaluations of the LLM-generated outputs with eight industry experts along four dimensions-naturalness, adequacy, usefulness and relevance-to gain practical insights. The results show an average rating of 3.84 out of 5 across evaluation criteria, indicating that the approach is effective in providing supporting explanations for requirement ambiguities. Sarmad Bashir, Alessio Ferrari 0001, Per Erik Strandberg, Zulqarnain Haider, Mehrdad Saadatmand, Markus Bohlin |
ICSME | 7 |
| 2025 | Efficient Torque Prediction for Digital Twins in Quarry Operations: A Data-Driven and Expert-Guided ApproachabstractQuarry sites present unique operational challenges where the performance of heavy machinery is critical for maintaining efficiency and safety. In such environments, accurate torque prediction is essential for effective engine management and optimal task execution. This work addresses the torque prediction challenge for a wheel loader operating in quarry conditions by proposing a structured three-phase approach to feature selection that reduces model complexity while preserving predictive accuracy. In the first phase, features are selected based on domain expertise to capture the physical and operational realities of quarry machinery. A comprehensive set of features is then employed to establish a robust performance baseline. In the final phase, a data-driven analysis using SHapley Additive Explanations (SHAP) identifies the top five features that most significantly impact torque prediction. Model efficacy was validated via cross-validation, with R-squared and mean-squared error serving as the key performance indicators. Comparative analysis reveals that while SHAP-ranked features yield statistically optimal results, the expert-selected features are more aligned with the practical requirements of quarry operations. These findings support the design of efficient, interpretable digital twins for real-time decisions in challenging environments. Abdulkarim Habbab, Anas Fattouh, Mohammad Loni, Koteshwar Chirumalla, Bobbie Frank, Markus Bohlin |
INDIN | 6 |
| 2024 | Machine learning testing in an ADAS case study using simulation-integrated bio-inspired search-based testingabstractSummary This paper presents an extended version of Deeper, a search‐based simulation‐integrated test solution that generates failure‐revealing test scenarios for testing a deep neural network‐based lane‐keeping system. In the newly proposed version, we utilize a new set of bio‐inspired search algorithms, genetic algorithm (GA), and evolution strategies (ES), and particle swarm optimization (PSO), that leverage a quality population seed and domain‐specific crossover and mutation operations tailored for the presentation model used for modeling the test scenarios. In order to demonstrate the capabilities of the new test generators within Deeper, we carry out an empirical evaluation and comparison with regard to the results of five participating tools in the cyber‐physical systems testing competition at SBST 2021. Our evaluation shows the newly proposed test generators in Deeper not only represent a considerable improvement on the previous version but also prove to be effective and efficient in provoking a considerable number of diverse failure‐revealing test scenarios for testing an ML‐driven lane‐keeping system. They can trigger several failures while promoting test scenario diversity, under a limited test time budget, high target failure severity, and strict speed limit constraints. Mahshid Helali Moghadam, Markus Borg, Mehrdad Saadatmand, Seyed Jalaleddin Mousavirad, Markus Bohlin, Björn Lisper |
J. Softw. Evol. Process. | 5 |
| 2023 | Requirement or Not, That is the Question: A Case from the Railway Industry
Sarmad Bashir, Muhammad Abbas 0002, Mehrdad Saadatmand, Eduard Paul Enoiu, Markus Bohlin, Pernilla Lindberg |
REFSQ | 5 |
| 2022 | An autonomous performance testing framework using self-adaptive fuzzy reinforcement learningabstractAbstract Test automation brings the potential to reduce costs and human effort, but several aspects of software testing remain challenging to automate. One such example is automated performance testing to find performance breaking points. Current approaches to tackle automated generation of performance test cases mainly involve using source code or system model analysis or use-case-based techniques. However, source code and system models might not always be available at testing time. On the other hand, if the optimal performance testing policy for the intended objective in a testing process instead could be learned by the testing system, then test automation without advanced performance models could be possible. Furthermore, the learned policy could later be reused for similar software systems under test, thus leading to higher test efficiency. We propose SaFReL, a self-adaptive fuzzy reinforcement learning-based performance testing framework. SaFReL learns the optimal policy to generate performance test cases through an initial learning phase, then reuses it during a transfer learning phase, while keeping the learning running and updating the policy in the long term. Through multiple experiments in a simulated performance testing setup, we demonstrate that our approach generates the target performance test cases for different programs more efficiently than a typical testing process and performs adaptively without access to source code and performance models. Mahshid Helali Moghadam, Mehrdad Saadatmand, Markus Borg, Markus Bohlin, Björn Lisper |
Softw. Qual. J. | 4 |
| 2021 | Performance Testing Using a Smart Reinforcement Learning-Driven Test AgentabstractPerformance testing with the aim of generating an efficient and effective workload to identify performance issues is challenging. Many of the automated approaches mainly rely on analyzing system models, source code, or extracting the usage pattern of the system during the execution. However, such information and artifacts are not always available. Moreover, all the transactions within a generated workload do not impact the performance of the system the same way, a finely tuned workload could accomplish the test objective in an efficient way. Model-free reinforcement learning is widely used for finding the optimal behavior to accomplish an objective in many decision-making problems without relying on a model of the system. This paper proposes that if the optimal policy (way) for generating test workload to meet a test objective can be learned by a test agent, then efficient test automation would be possible without relying on system models or source code. We present a self-adaptive reinforcement learning-driven load testing agent, RELOAD, that learns the optimal policy for test workload generation and generates an effective workload efficiently to meet the test objective. Once the agent learns the optimal policy, it can reuse the learned policy in subsequent testing activities. Our experiments show that the proposed intelligent load test agent can accomplish the test objective with lower test cost compared to common load testing procedures, and results in higher test efficiency. Mahshid Helali Moghadam, Golrokh Hamidi, Markus Borg, Mehrdad Saadatmand, Markus Bohlin, Björn Lisper, Pasqualina Potena |
CEC | 5 |
| 2020 | Poster: Performance Testing Driven by Reinforcement LearningabstractPerformance testing remains a challenge, particularly for complex systems. Different application-, platform- and workload-based factors can influence the performance of software under test. Common approaches for generating platform- and workload-based test conditions are often based on system model or source code analysis, real usage modeling and use-case based design techniques. Nonetheless, creating a detailed performance model is often difficult, and also those artifacts might not be always available during the testing. On the other hand, test automation solutions such as automated test case generation can enable effort and cost reduction with the potential to improve the intended test criteria coverage. Furthermore, if the optimal way (policy) to generate test cases can be learnt by testing system, then the learnt policy can be reused in further testing situations such as testing variants, evolved versions of software, and different testing scenarios. This capability can lead to additional cost and computation time saving in the testing process. In this research, we present an autonomous performance testing framework which uses a model-free reinforcement learning augmented by fuzzy logic and self-adaptive strategies. It is able to learn the optimal policy to generate platform- and workload-based test conditions which result in meeting the intended testing objective without access to system model and source code. The use of fuzzy logic and self-adaptive strategy helps to tackle the issue of uncertainty and improve the accuracy and adaptivity of the proposed learning. Our evaluation experiments show that the proposed autonomous performance testing framework is able to generate the test conditions efficiently and in a way adaptive to varying testing situations. Mahshid Helali Moghadam, Mehrdad Saadatmand, Markus Borg, Markus Bohlin, Björn Lisper |
ICST | 4 |
| 2018 | Using Mutant Stubbornness to Create Minimal and Prioritized Test SetsabstractIn testing, engineers want to run the most useful tests early (prioritization). When tests are run hundreds or thousands of times, minimizing a test set can result in significant savings (minimization). This paper proposes a new analysis technique to address both the minimal test set and the test case prioritization problems. This paper precisely defines the concept of mutant stubbornness, which is the basis for our analysis technique. We empirically compare our technique with other test case minimization and prioritization techniques in terms of the size of the minimized test sets and how quickly mutants are killed. We used seven C language subjects from the Siemens Repository, specifically the test sets and the killing matrices from a previous study. We used 30 different orders for each set and ran every technique 100 times over each set. Results show that our analysis technique performed significantly better than prior techniques for creating minimal test sets and was able to establish new bounds for all cases. Also, our analysis technique killed mutants as fast or faster than prior techniques. These results indicate that our mutant stubbornness technique constructs test sets that are both minimal in size, and prioritized effectively, as well or better than other techniques. Loreto Gonzalez-Hernandez, Birgitta Lindström, A. Jefferson Offutt, Sten F. Andler, Pasqualina Potena, Markus Bohlin |
QRS | 6 |
| 2018 | ESPRET: A tool for execution time estimation of manual test cases
Sahar Tahvili, Wasif Afzal, Mehrdad Saadatmand, Markus Bohlin, Sharvathul Hasan Ameerjan |
J. Syst. Softw. | 4 |
| 2018 | Similarity-based prioritization of test case automationabstractThe importance of efficient software testing procedures is driven by an ever increasing system complexity as well as global competition. In the particular case of manual test cases at the system integration level, where thousands of test cases may be executed before release, time must be well spent in order to test the system as completely and as efficiently as possible. Automating a subset of the manual test cases, i.e, translating the manual instructions to automatically executable code, is one way of decreasing the test effort. It is further common that test cases exhibit similarities, which can be exploited through reuse when automating a test suite. In this paper, we investigate the potential for reducing test effort by ordering the test cases before such automation, given that we can reuse already automated parts of test cases. In our analysis, we investigate several approaches for prioritization in a case study at a large Swedish vehicular manufacturer. The study analyzes the effects with respect to test effort, on four projects with a total of 3919 integration test cases constituting 35,180 test steps, written in natural language. The results show that for the four projects considered, the difference in expected manual effort between the best and the worst order found is on average 12 percentage points. The results also show that our proposed prioritization method is nearly as good as more resource demanding meta-heuristic approaches at a fraction of the computational time. Based on our results, we conclude that the order of automation is important when the set of test cases contain similar steps (instructions) that cannot be removed, but are possible to reuse. More precisely, the order is important with respect to how quickly the manual test execution effort decreases for a set of test cases that are being automated. Daniel Flemström, Pasqualina Potena, Daniel Sundmark, Wasif Afzal, Markus Bohlin |
Softw. Qual. J. | 5 |
| 2017 | Towards Execution Time Prediction for Manual Test Cases from Test SpecificationabstractKnowing the execution time of test cases is important to perform test scheduling, prioritization and progress monitoring. This work in progress paper presents a novel approach for predicting the execution time of test cases based on test specifications and available historical data on previously executed test cases. Our approach works by extracting timing information (measured and maximum execution time)for various steps in manual test cases. This information is then used to estimate the maximum time for test steps that have not previously been executed, but for which textual specifications exist. As part of our approach, natural language parsing of the specifications is performed to identify word combinations to check whether existing timing information on various test activities is already available or not. Finally, linear regression is used to predict the actual execution time for test cases. A proof-of-concept use case at Bombardier Transportation serves to evaluate the proposed approach. Sahar Tahvili, Mehrdad Saadatmand, Markus Bohlin, Wasif Afzal, Sharvathul Hasan Ameerjan |
SEAA | 3 |
| 2016 | Cost-Benefit Analysis of Using Dependency Knowledge at Integration Testing
Sahar Tahvili, Markus Bohlin, Mehrdad Saadatmand, Stig Larsson 0002, Wasif Afzal, Daniel Sundmark |
PROFES | 2 |
| 2014 | Search-Based Testing for Embedded Telecom Software with Complex Input Structures
Kivanc Doganay, Sigrid Eldh, Wasif Afzal, Markus Bohlin |
ICTSS | 4 |
| 2012 | Optimal Freight Train Classification using Column GenerationabstractWe consider planning of freight train classification at hump yards using integer programming. The problem involves the formation of departing freight trains from arriving trains subject to scheduling and capacity constraints. To increase yard capacity, we allow the temporary storage of early freight cars on specific mixed-usage tracks. The problem has previously been modeled using a direct integer programming model, but this approach did not yield lower bounds of sufficient quality to prove optimality. In this paper, we formulate a new extended integer programming model and design a column generation approach based on branch-and-price to solve problem instances of industrial size. We evaluate the method on historical data from the Hallsberg hump yard in Sweden, and compare the results with previous approaches. The new method managed to find optimal solutions in all of the 192 problem instances tried. Furthermore, no instance took more than 13 minutes to solve to optimality using fairly standard computer hardware. Markus Bohlin, Florian Dahms, Holger Flier, Sara Gestrelius |
ATMOS | 1 |
| 2012 | Statistical Anomaly Detection for Train FleetsabstractWe have developed a method for statistical anomaly detection which has been deployed in a tool for condition monitoring of train fleets. The tool is currently used by several railway operators over the world to inspect and visualize the occurrence of “event messages” generated on the trains. The anomaly detection component helps the operators to quickly find significant deviations from normal behavior and to detect early indications for possible problems. The savings in maintenance costs comes mainly from avoiding costly breakdowns, and have been estimated to several million Euros per year for the tool. In the long run, it is expected that maintenance costs can be reduced with between 5 and 10% by using the tool. Anders Holst, Markus Bohlin, Jan Ekman, Ola Sellin, Björn Lindström, Stefan Larsen |
IAAI | 2 |
| 2011 | Track Allocation in Freight-Train Classification with Mixed TracksabstractWe consider the process of forming outbound trains from cars of inbound trains at rail-freight hump yards. Given the arrival and departure times as well as the composition of the trains, we study the problem of allocating classification tracks to outbound trains such that every outbound train can be built on a separate classification track. We observe that the core problem can be formulated as a special list coloring problem in interval graphs, which is known to be NP-complete. We focus on an extension where individual cars of different trains can temporarily be stored on a special subset of the tracks. This problem induces several new variants of the list-coloring problem, in which the given intervals can be shortened by cutting off a prefix of the interval. We show that in case of uniform and sufficient track lengths, the corresponding coloring problem can be solved in polynomial time, if the goal is to minimize the total cost associated with cutting off prefixes of the intervals. Based on these results, we devise two heuristics as well as an integer program to tackle the problem. As a case study, we consider a real-world problem instance from the Hallsberg Rangerbangard hump yard in Sweden. Planning over horizons of seven days, we obtain feasible solutions from the integer program in all scenarios, and from the heuristics in most scenarios. Markus Bohlin, Holger Flier, Jens Maue, Matús Mihalák |
ATMOS | 1 |
| 2009 | MILP formulations of cumulative constraints for railway scheduling - A comparative study
Martin Aronsson, Markus Bohlin, Per Kreuger |
ATMOS | 2 |
| 2009 | A Tool for Gas Turbine Maintenance Scheduling
Markus Bohlin, Kivanc Doganay, Per Kreuger, Rebecca Steinert, Mathias Wärja |
IAAI | 1 |
| 2009 | Simulation-Based Timing Analysis of Complex Real-Time SystemsabstractThis paper presents an efficient best-effort approach for simulation-based timing analysis of complex real-time systems. The method can handle in principle any software design that can be simulated, and is based on controlling simulation input using a simple yet novel hill-climbing algorithm. Unlike previous approaches, the new algorithm directly manipulates simulation parameters such as execution times, arrival jitter and input. An evaluation is presented using six different simulation models, and two other simulation methods as reference: Monte Carlo simulation and MABERA. The new method proposed in this paper was 4-11% more accurate while at the same time 42 times faster, on average, than the reference methods. Markus Bohlin, Yue Lu 0005, Johan Kraft, Per Kreuger, Thomas Nolte |
RTCSA | 1 |
| 2008 | Bounding Shared-Stack Usage in Systems with Offsets and PrecedencesabstractThe paper presents two novel methods to bound the stack memory used in preemptive, shared stack, real-time systems. The first method is based on branch-and-bound search for possible preemption patterns, and the second one approximates the first in polynomial time. The work extends previous methods by considering a more general task-model, in which all tasks can share the same stack. In addition, the new methods account for precedence and offset relations. Thus, the methods give tight bounds for a large set of realistic systems. The methods have been implemented and a comprehensive evaluation, comparing our new methods against each other and against existing methods, is presented. The evaluation shows that our exact method can significantly reduce the amount of stack memory needed. Markus Bohlin, Kaj Hänninen, Jukka Mäki-Turja, Jan Carlson, Mikael Nolin |
ECRTS | 1 |
| 2006 | Determining Maximum Stack Usage in Preemptive Shared Stack SystemsabstractThis paper presents a novel method to determine the maximum stack memory used in preemptive, shared stack, real-time systems. We provide a general and exact problem formulation applicable for any preemptive system model based on dynamic (run-time) properties. We also show how to safely approximate the exact stack usage by using static (compile time) information about the system model and the underlying run-time system on a relevant and commercially available system model: a hybrid, statically and dynamically, scheduled system. Comprehensive evaluations show that our technique significantly reduces the amount of stack memory needed compared to existing analysis techniques. For typical task sets a decrease in the order of 70% is typical Kaj Hänninen, Jukka Mäki-Turja, Markus Bohlin, Jan Carlson, Mikael Nolin |
RTSS | 3 |
| 2002 | Improving Cost Calculations for Global Constraints in Local Search
Markus Bohlin |
CP | 1 |