VLDB 2026 Research / reviewers in the wild / expert
Felix Dobslaw
dblp:91/8284
· DBLP profile ↗
16ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0001-9372-3416ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding on the Edge: LLM-generated Boundary Test ExplanationsabstractBoundary value analysis and testing (BVT) is fundamental in software quality assurance because faults tend to cluster at input extremes, yet testers often struggle to understand and justify why certain input-output pairs represent meaningful behavioral boundaries. Large Language Models (LLMs) could help by producing natural-language rationales, but their value for BVT has not been empirically assessed. We therefore conducted an exploratory study on LLM-generated boundary explanations: in a survey, twenty-seven software professionals rated GPT-4.1 explanations for twenty boundary pairs on clarity, correctness, completeness and perceived usefulness, and six of them elaborated in follow-up interviews. Overall, 63.5% of all ratings were positive (4–5 on a five-point Likert scale) compared to 17% negative (1–2), indicating general agreement but also variability in perceptions. Participants favored explanations that followed a clear structure, cited authoritative sources, and adapted their depth to the reader’s expertise; they also stressed the need for actionable examples to support debugging and documentation. From these insights, we distilled a seven-item requirement checklist that defines concrete design criteria for future LLM-based boundary explanation tools. The results suggest that, with further refinement, LLM-based tools can support testing workflows by making boundary explanations more actionable and trustworthy. Sabinakhon Akbarova, Felix Dobslaw, Robert Feldt |
AST | 2 |
| 2026 | Automatic techniques for issue report classification: A systematic mapping studyabstractSeveral studies have evaluated automatic techniques for classifying software issue reports into bugs and non-bugs to assist practitioners in effectively assigning relevant resources based on the type of issue. Currently, no comprehensive overview of this area has been published. A comprehensive overview will help identify future research directions and provide an extensive collection of potentially relevant existing solutions. This study aims to provide a comprehensive overview of the use of automatic techniques to classify issue reports. We conducted a systematic mapping study and identified 46 studies on the topic. The study results indicate that the existing literature applies various techniques for classifying issue reports, including traditional machine learning and deep learning-based techniques, and more advanced large language models. Furthermore, we observe that these studies (a) lack the involvement of practitioners, (b) do not consider other potentially relevant adoption factors beyond prediction accuracy, such as the explainability, scalability, and generalizability of the techniques, and (c) mainly rely on archival data from open-source repositories only. Therefore, future research should focus on real industrial evaluations, consider other potentially relevant adoption factors, and actively involve practitioners. Muhammad Laiq, Felix Dobslaw |
Autom. Softw. Eng. | 2 |
| 2026 | Addressing trust requirements in the design of an open-source multi-agent LLM-based domain-specific chatbotabstractAbstract Large Language Models (LLMs) have the potential to automate knowledge-intensive interactions in enterprise systems, yet their adoption is often limited. One reason is a lack of user trust. This study examines how trust can be systematically engineered into an LLM-driven, multi-agent chatbot that handles routine human-resources (HR) queries. We follow a two-cycle Design Science Research methodology. Cycle 1 triangulated a systematic literature review with a thematic analysis over semi-structured interviews of six employees at a global firm and a confirmatory workshop with five AI experts to elicit and validate trust requirements . Cycle II instantiated these requirements in a multi-agent LLM chatbot prototype artifact and evaluated whether the artifact satisfies them through controlled user sessions and expert walkthroughs, emphasizing perceived usefulness and trust captured in post-task interviews ( $$n = 11$$ ) and operationalizing trust via alignment-oriented measures (faithfulness, answer relevancy, and adversarial robustness). The study yields a refined taxonomy of external (transparency, organizational safeguards, third-party security) and internal (model provenance, bias risk, reliability) trust factors, identifying reliability as the primary determinant of adoption. The implemented design achieved $$\ge 0.86$$ on trust-aligned metrics and was endorsed by 9/11 participants as ready for field deployment. These findings demonstrate that trust can be proactively addressed through design and offer prescriptive guidelines for software engineers seeking to embed LLMs safely and responsibly in socio-technical contexts. Jonatan Axetorn, Felix Edholm, Felix Dobslaw, Lucas Gren |
Requir. Eng. | 3 |
| 2025 | An empirical investigation of the impact of architectural smells on software maintainabilityabstractIn recent years, interest in the potential influence of architectural smells on software maintainability has grown. Yet, empirical evidence directly linking maintainability quality attributes — such as modularity and testability — with architectural smells is scarce. This study analyzes seven architectural smells across 378 versions of eight open-source projects. We developed a tool to gather data on architectural smells, associated projects, and package-level quality attributes. The empirical findings reinforce that certain architectural smells indeed correlate with specific quality attributes - a notion previously backed merely by argument. Most architectural smells negatively correlated with project-level testability but not project-level modularity. While there is not a consistent negative correlation with testability at the package level, many smells display a pronounced negative association with the maintainability quality attributes. • The size of projects has a significant impact on both modularity and testability. • Architectural smells show a negative relationship with testability. • Addressing Dense Structure smell can improve modularity and testability. • Addressing God Component smell can improve modularity and testability. • Addressing Scattered Functionality smell can improve modularity and testability. Rodi Jolak, Simon Karlsson, Felix Dobslaw |
J. Syst. Softw. | 3 |
| 2024 | On current limitations of online eye-tracking to study the visual processing of source codeabstractEye-tracking is an increasingly popular instrument to study how programmers process and comprehend source code. While most studies are conducted in controlled environments with lab-grade hardware, it would be desirable to simplify and scale participation in experiments for users sitting remotely, leveraging home equipment. This study investigates the possibility of performing eye-tracking studies remotely using open-source algorithms and consumer-grade webcams. It establishes the technology’s current limitations and evaluates the quality of the data collected by it. We conclude by recommending ways forward to address the shortcomings and make remote code-reading studies in support of eye-tracking feasible in the future. We gathered eye-gaze data remotely from 40 participants performing a code reading experiment on a purpose-built web application. The utilized eye-tracker worked client-side and used ridge regression to generate x- and y-coordinates in real-time predicting the participants’ on-screen gaze points without the need to collect and save video footage. We processed and analysed the collected data according to common practices for isolating eye-movement events and deriving metrics used in software engineering eye-tracking studies. In response to the lack of an algorithm explicitly developed for detecting oculomotor fixation events in low-frequency webcam data, we also introduced a dispersion threshold algorithm for that purpose. The quality of the collected data was subsequently assessed to determine the adequacy and validity of the methodology for eye-tracking. The collected data was found to be of varying quality despite extensive calibration and graphical user guidance. We present our results highlighting both the negative and positive observations from which the community hopefully can learn. Both accuracy and precision were low and ultimately deemed insufficient for drawing valid conclusions in a high-precision empirical study. We nonetheless contribute to identifying critical limitations to be addressed in future research. Apart from the overall challenge of vastly diverse equipment, setup, and configuration, we found two main problems with the current webcam eye-tracking technology. The first was the absence of a validated algorithm to isolate fixations in low-frequency data, compromising the assurance of the accuracy of the data derived from it. The second problem was the lack of algorithmic support for head movements when predicting gaze location. Unsupervised participants do not always keep their heads still, even if instructed to do so. Consequently, we frequently observed spatial shifts that corrupted many collected datasets. Three encouraging observations resulted from the study. Even when shifted, gaze points were consistently dispersed in patterns resembling both the shape and size of the stimuli without extreme deviations. We could also distinguish recognizable reading patterns. Linearity was significantly different when participants were reading source code compared to natural text, and we could detect the expected left-to-right and top-to-bottom reading directions for participants reading natural text snippets. The accuracy and precision levels were not sufficient for a word-by-word analysis of code reading but could be adequate for a broader, coarse-grained precision study. Additionally we identified two main issues compromising the collected data validity and contributed a fixation detection algorithm to approach one of these issues. With suitable solutions to the identified issues, remote eye-tracking studies with webcams on code reading could eventually be feasible. Eva Thilderkvist, Felix Dobslaw |
Inf. Softw. Technol. | 2 |
| 2023 | Investigating Software Testing and Maintenance of Open-Source Distributed LedgerabstractA distributed ledger is the backbone of all blockchain solutions. It provides a shared database spreading across a network of nodes. The number of DL solutions and their implementations has grown in recent years. Besides the architectural and performance promises of thesesolutions, organizations seekingto implement DL also need to consider the overall quality of the software available and its ecosystem. Particularly, previous research has identified the need to better understand the testing and maintenance practices behind these types of technologies. This paper investigates the testing and maintenance of 18 different open-source projects that implement distributed ledgers. We perform a manual inspection of test artefacts and mine the history of commits, issues and contributors of the chosen projects to understand the landscape of testing and maintenance in these projects. Our findings suggest that unit and integration tests are present in most projects, they do not follow a holistic system testing approach. Moreover, projects rely on a small team of core contributors (5 on average). While the projects are continuously maintained, larger changes are uncommon. Our results can be used for benchmarking and pinpointing areas of improvement for the development of distributed ledgers. Petya Hristova Cvitic, Felix Dobslaw, Francisco Gomes de Oliveira Neto |
SANER | 2 |
| 2023 | Generic and industrial scale many-criteria regression test selectionabstractWhile several test case selection algorithms (heuristic and optimal) and formulations (linear and non-linear) have been proposed, no multi-criteria framework enables Pareto search — the state-of-the-art approach of doing multi-criteria optimization. Therefore, we introduce the highly parallelizable, openly available Many-Criteria Test-Optimization Algorithm (MC-TOA) framework that combines heuristic Pareto search and optimality gap knowledge per criterion. MC-TOA is largely agnostic to the criteria formulations and can incorporate many criteria where existing approaches offer limited scope (single or few objectives/constraints), lack flexibility in the expression and assurance of constraints, or run into problem complexity issues. For two large-scale systems with up to seven criteria and thousands of system test cases, MC-TOA not only produces, over the board, superior Pareto fronts in terms of HVI score compared to the state-of-the-art many-objective heuristic baseline, it also does that within minutes of runtime for worst-case executions, i.e., assuming that a regression affects the entire test-suite. MC-TOA depends on convex solvers. We find that the evaluated open-source solvers are slower but suffice for smaller systems, while being less robust for larger systems. Linear formulations execute faster and obtain near-optimal results, which led to faster and better overall convergence of MC-TOA compared to integer formulations. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Felix Dobslaw, Ruiyuan Wan, Yuechan Hao |
J. Syst. Softw. | 1 |
| 2021 | Linking Developer Experience to Coding Style in Open-Source RepositoriesabstractWe, humans, gain experience by doing something for a prolonged period. In this paper, we address whether we can link the use of advanced language features in software to a greater developer experience, indicating higher software quality. The coding style we chose to measure is the usage of lambdas and whether it correlates with programming experience since previous research has shown that less experienced C++ developers face difficulties with lambdas despite them being a language feature for almost ten years. If we established that lambda use and developer experience correlate positively, we could further investigate whether a good understanding of lambdas has a lasting contribution to software quality. Further, we could use it as an indicator within code quality metrics and promote the teaching of functional programming and lambda use in software engineering curricula. To measure experience, we introduce the Mean Repository Experience (MRE) metric, a novel but straightforward way of measuring what we here call repository experience, i.e., the combined assessed experience of contributors in a repository. We use this metric to analyze 500 C++ repositories. The proposed MRE metric shows potential as a proxy for software quality and could further extend to advanced language features other than lambdas. Our results suggest that the developer experience positively correlates with lambda usage. Future research includes understanding how well MRE reflects actual developer experience and further implications. Heidi Hokka, Felix Dobslaw, Jonathan Bengtsson |
SANER | 2 |
| 2020 | Free the Bugs: Disclosing Blocking Violations in Reactive ProgrammingabstractIn programming, concurrency allows threads to share processing units interleaving and seemingly simultaneous to improve resource utilization and performance. Previous research has found that concurrency faults are hard to avoid, hard to find, often leading to undesired and unpredictable behavior. Further, with the growing availability of multi-core devices and adaptation of concurrency features in high-level languages, concurrency faults occur reportedly often, which is why countermeasures must be investigated to limit harm. Reactive programming provides an abstraction to simplify complex concurrent and asynchronous tasks through reactive language extensions such as the RxJava and Project Reactor libraries for Java. Still, blocking violations are possibly resulting in concurrency faults with no Java compiler warnings. BlockHound is a tool that detects incorrect blocking by wrapping the original code and intercepting blocking calls to provide appropriate runtime errors. In this study, we seek an understanding of how common blocking violations are and whether a tool such as BlockHound can give us insight into the root-causes to highlight them as pitfalls to developers. The investigated Softwares are Java-based open-source projects using reactive frameworks selected based on high star ratings and large fork quantities that indicate high adoption. We activated BlockHound in the project's test-suites and analyzed log files for common patterns to reveal blocking violations in 7/29 investigated open-source projects with 5024 stars and 1437 forks. A small number of system calls could be identified as root-causes. We here present countermeasures that successfully removed the uncertainty of blocking violations. The code's intentional logic was retained in all validated projects through passing unit-tests. Felix Dobslaw, Morgan Vallin, Robin Sundström |
SCAM | 1 |
| 2019 | Estimating Return on Investment for GUI Test Automation FrameworksabstractAutomated graphical user interface (GUI) tests can reduce manual testing activities and increase test frequency. This motivates the conversion of manual test cases into automated GUI tests. However, it is not clear whether such automation is cost-effective given that GUI automation scripts add to the code base and demand maintenance as a system evolves. In this paper, we introduce a method for estimating maintenance cost and Return on Investment (ROI) for Automated GUI Testing (AGT). The method utilizes the existing source code change history and has the potential to be used for the evaluation of other testing or quality assurance automation technologies. We evaluate the method for a real-world, industrial software system and compare two fundamentally different AGT frameworks, namely Selenium and EyeAutomate, to estimate and compare their ROI. We also report on their defect-finding capabilities and usability. The quantitative data is complemented by interviews with employees at the company the study has been conducted at. The method was successfully applied, and estimated maintenance cost and ROI for both frameworks are reported. Overall, the study supports earlier results showing that implementation time is the leading cost for introducing AGT. The findings further suggest that, while EyeAutomate tests are significantly faster to implement, Selenium tests require more of a programming background but less maintenance. Felix Dobslaw, Robert Feldt, David Michaelsson, Patrick Haar, Francisco Gomes de Oliveira Neto, Richard Torkar |
ISSRE | 1 |
| 2019 | Towards Automated Boundary Value Testing with Program Derivatives and Search
Robert Feldt, Felix Dobslaw |
SSBSE | 2 |
| 2016 | End-to-End Reliability-Aware Scheduling for Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are gaining popularity as a flexible and economical alternative to field-bus installations for monitoring and control applications. For mission-critical applications, communication networks must provide end-to-end reliability guarantees, posing substantial challenges for WSN. Reliability can be improved by redundancy, and is often addressed on the MAC layer by resubmission of lost packets, usually applying slotted scheduling. Recently, researchers have proposed a strategy to optimally improve the reliability of a given schedule by repeating the most rewarding slots in a schedule incrementally until a deadline. This Incrementer can be used with most scheduling algorithms but has scalability issues which narrows its usability to offline calculations of schedules, for networks that are rather static. In this paper, we introduce SchedEx, a generic heuristic scheduling algorithm extension, which guarantees a user-defined end-to-end reliability. SchedEx produces competitive schedules to the existing approach, and it does that consistently more than an order of magnitude faster. The harsher the end-to-end reliability demand of the network, the better the SchedEx performs compared to the Incrementer. We further show that SchedEx has a more evenly distributed improvement impact on the scheduling algorithms, whereas the Incrementer favors schedules created by certain scheduling algorithms. Felix Dobslaw, Mikael Gidlund |
IEEE Trans. Ind. Informatics | 1 |
| 2016 | QoS-Aware Cross-Layer Configuration for Industrial Wireless Sensor NetworksabstractIn many applications of industrial sensor networks, stringent reliability and maximum delay constraints paired with priority demands on a sensor-basis are present. These quality of service (QoS) requirements pose tough challenges for industrial wireless sensor networks that are deployed to an ever larger extent due to their flexibility and extendibility. In this paper, we introduce an integrated cross-layer framework, SchedEx-GA, spanning medium access control (MAC) layer and network layer. SchedEx-GA attempts to identify a network configuration that fulfills all application-specific process requirements over a topology including the sensor publish rates, maximum acceptable delay, service differentiation, and event transport reliabilities. The network configuration comprises the decision for routing, as well as scheduling. For many of the evaluated topologies it is not possible to find a valid configuration due to the physical conditions of the environment. We therefore introduce a converging algorithm on top of the framework which configures a given topology by additional sink positioning in order to build a backbone with the gateway that guarantees the application specific constraints. The results show that, in order to guarantee a high end-to-end reliability of 99.999% for all flows in a network containing emergency, control loop, and monitoring traffic, a backbone with multiple sinks is often required for the tested topologies. Additional features, such as multichannel utilization and aggregation, though, can substantially reduce the demand for required sinks. In its present version, the framework is used for centralized control, but with the potential to be extended for decentralized control in future work. Felix Dobslaw, Mikael Gidlund |
IEEE Trans. Ind. Informatics | 1 |
| 2013 | QoS assessment for mission-critical Wireless Sensor Network applicationsabstractWireless sensor networks (WSN) must ensure worst-case end-to-end delay and reliability guarantees for mission-critical applications. TDMA-based scheduling offers delay guarantees, thus it is used in industrial monitoring and automation. We propose to evolve pairs of TDMA schedule and routing-tree in a cross-layer in order to fulfill multiple conflicting QoS requirements, exemplified by latency and reliability. The genetic algorithm we utilize can be used as an analytical tool for both the feasibility and expected QoS in production. Near-optimal cross-layer solutions are found within seconds and can be directly enforced into the network. Felix Dobslaw, Mikael Gidlund |
LCN | 1 |
| 2013 | SAS-TDMA: a source aware scheduling algorithm for real-time communication in industrial wireless sensor networks
Mikael Gidlund, Felix Dobslaw |
Wirel. Networks | 4 |
| 2011 | Iteration-wise parameter learningabstractAdjusting the control parameters of population-based algorithms is a means for improving the quality of these algorithms' result when solving optimization problems. The difficulty lies in determining when to assign individual values to specific parameters during the run. This paper investigates the possible implications of a generic and computationally cheap approach towards parameter analysis for population-based algorithms. The effect of parameter settings was analyzed in the application of a genetic algorithm to a set of traveling salesman problem instances. The findings suggest that statistics about local changes of a search from iteration i to iteration i + 1 can provide valuable insight into the sensitivity of the algorithm to parameter values. A simple method for choosing static parameter settings has been shown to recommend settings competitive to those extracted from a state-of-the-art parameter tuner, paramlLS, with major time and setup advantages. Felix Dobslaw |
IEEE Congress on Evolutionary Computation | 1 |