Leon Moonen

dblp:m/LeonMoonen · DBLP profile ↗
← Back
65ranked-venue papers
10as first author
17since 2021 · last 2025
0000-0002-1761-6771ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 62 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models
abstract
Automatic program repair (APR) aims at reducing the manual efforts required to identify and fix errors in source code. Before the rise of Large Language Model (LLM)-based agents, a common strategy was simply to increase the number of generated patches, sometimes to the thousands, which usually yielded better repair results on benchmarks. More recently, self-iterative capabilities enabled LLMs to refine patches over multiple rounds guided by feedback. However, literature often focuses on many iterations and disregards different numbers of outputs.
Fernando Vallecillos Ruiz, Max Hort, Leon Moonen
EASE3
2025 The Impact of Fine-Tuning Large Language Models on Automated Program Repair
abstract
Automated Program Repair (APR) uses various tools and techniques to help developers achieve functional and errorfree code faster. In recent years, Large Language Models (LLMs) have gained popularity as components in APR tool chains because of their performance and flexibility. However, training such models requires a significant amount of resources. Fine-tuning techniques have been developed to adapt pre-trained LLMs to specific tasks, such as APR, and enhance their performance at far lower computational costs than training from scratch. In this study, we empirically investigate the impact of various fine-tuning techniques on the performance of llms used for APR. Our experiments provide insights into the performance of a selection of state-of-the-art LLMs pre-trained on code. The evaluation is done on three popular APR benchmarks (i.e., QuixBugs, Defects4J and HumanEval-Java) and considers six different LLMs with varying parameter sizes (resp. CodeGen, CodeT5, StarCoder, DeepSeekCoder, Bloom, and CodeLlama2). We consider three training regimens: no fine-tuning, full fine-tuning, and parameter-efficient fine-tuning (PEFT) using LoRA and IA3. We observe that full fine-tuning techniques decrease the benchmarking performance of various models due to different data distributions and overfitting. By using parameterefficient fine-tuning methods, we restrict models in the amount of trainable parameters and achieve better results.
Roman Machacek, Anastasiia Grishina, Max Hort, Leon Moonen
ICSME4
2025 Codehacks: A Dataset of Adversarial Tests for Competitive Programming Problems Obtained from Codeforces
abstract
Software is used in critical applications in our day-to-day life and it is important to ensure its correctness. One popular approach to assess correctness is to evaluate software on tests. If a test fails, it indicates a fault in the software under test; if all tests pass correctly, one may assume that the software is correct. However, the reliability of these results depends on the test suite considered, and there is a risk of false negatives (i.e. software that passes all available tests but contains bugs because some cases are not tested). Therefore, it is important to consider error-inducing test cases when evaluating software. To support data-driven creation of such a test-suite, which is especially of interest for testing software synthesized from large language models, we curate a dataset (Codehacks) of programming problems together with corresponding error-inducing test cases (i.e., “hacks”). This dataset is collected from the wild, in particular, from the Codeforces online judge platform. The dataset comprises 288,617 hacks for 5,578 programming problems, each with a natural language description, as well as the source code for 2,196 submitted solutions to these problems that can be broken with their corresponding hacks.
Max Hort, Leon Moonen
ICST2
2025 A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMs
abstract
Vulnerability databases, such as the National Vulnerability Database (NVD), offer detailed descriptions of Common Vulnerabilities and Exposures (CVEs), but often lack information on their real-world impact, such as the tactics, techniques, and procedures (TTPs) that adversaries may use to exploit the vulnerability. However, manually linking CVEs to their corresponding TTPs is a challenging and time-consuming task, and the high volume of new vulnerabilities published annually makes automated support desirable.This paper introduces Triage, a two-pronged automated approach that uses Large Language Models (LLMs) to map CVEs to relevant techniques from the Att&ck knowledge base. We first prompt an LLM with instructions based on MITRE’s CVE Mapping Methodology to predict an initial list of techniques. This list is then combined with the results from a second LLM-based module that uses in-context learning to map a CVE to relevant techniques. This hybrid approach strategically combines rule-based reasoning with data-driven inference. Our evaluation reveals that in-context learning outperforms the individual mapping methods, and the hybrid approach improves recall of exploitation techniques. We also find that GPT-4o-mini performs better than Llama3.3-70B on this task. Overall, our results show that LLMs can be used to automatically predict the impact of cybersecurity vulnerabilities and Triage makes the process of mapping CVEs to Att&ck more efficient.
Anders Mølmen Høst, Pierre Lison, Leon Moonen
TrustCom3
2025 Fully Autonomous Programming Using Iterative Multi-Agent Debugging with Large Language Models
abstract
Program synthesis with Large Language Models (LLMs) suffers from a “near-miss syndrome”: The generated code closely resembles a correct solution but fails unit tests due to minor errors. We address this with a multi-agent framework called Synthesize, Execute, Instruct, Debug, and Repair (SEIDR). Effectively applying SEIDR to instruction-tuned LLMs requires determining (a) optimal prompts for LLMs, (b) what ranking algorithm selects the best programs in debugging rounds, and (c) balancing the repair of unsuccessful programs with the generation of new ones. We empirically explore these tradeoffs by comparing replace-focused, repair-focused, and hybrid debug strategies. We also evaluate lexicase and tournament selection to rank candidates in each generation. On Program Synthesis Benchmark 2 (PSB2), our framework outperforms both conventional use of OpenAI Codex without a repair phase and traditional genetic programming approaches. SEIDR outperforms the use of an LLM alone, solving 18 problems in C++ and 20 in Python on PSB2 at least once across experiments. To assess generalizability, we employ GPT-3.5 and Llama 3 on the PSB2 and HumanEval-X benchmarks. Although SEIDR with these models does not surpass current state-of-the-art methods on the Python benchmarks, the results on HumanEval-C++ are promising. SEIDR with Llama 3-8B achieves an average pass@100 of 84.2%. Across all SEIDR runs, 163 of 164 problems are solved at least once with GPT-3.5 in HumanEval-C++, and 162 of 164 with the smaller Llama 3-8B. We conclude that SEIDR effectively overcomes the near-miss syndrome in program synthesis with LLMs.
Anastasiia Grishina, Vadim Liventsev, Aki Härmä, Leon Moonen
ACM Trans. Evol. Learn. Optim.4
2024 A Comparative Study on Large Language Models for Log Parsing
abstract
Background: Log messages provide valuable information about the status of software systems. This information is provided in an unstructured fashion and automated approaches are applied to extract relevant parameters. To ease this process, log parsing can be applied, which transforms log messages into structured log templates. Recent advances in language models have led to several studies that apply ChatGPT to the task of log parsing with promising results. However, the performance of other state-of-the-art large language models (LLMs) on the log parsing task remains unclear.
Merve Astekin, Max Hort, Leon Moonen
ESEM3
2024 The Impact of Program Reduction on Automated Program Repair
abstract
Correcting bugs using modern Automated Program Repair (APR) can be both time-consuming and resource-expensive. We describe a program repair approach that aims to improve the scalability of modern APR tools. The approach leverages program reduction in the form of program slicing to eliminate code irrelevant to fixing the bug, which improves the APR tool's overall performance. We investigate slicing's impact on all three phases of the repair process: fault localization, patch generation, and patch validation. Our empirical exploration finds that the proposed approach on average enhances the repair ability of the TBar APR tool, but we also discovered a few cases where it was less successful. Specifically, on examples from the widely used Defects4J dataset, we obtain a substantial reduction in median repair time, which falls from 80 minutes to just under 18 minutes. We conclude that program reduction can improve the performance of APR without degrading repair quality, but this improvement is not universal.
Linas Vidziunas, Dave W. Binkley, Leon Moonen
ICSME3
2024 Extending the range of bugs that automated program repair can handle
abstract
Modern automated program repair (APR) is well-tuned to finding and repairing bugs that introduce observable erroneous behavior to a program. However, a significant class of bugs does not lead to observable behavior (e.g., termination bugs and non-functional bugs). Such bugs can generally not be handled with current APR approaches, so complementary techniques are needed. To stimulate the systematic study of alternative approaches and hybrid combinations, we devise a novel bug classification system that enables methodical analysis of their bug detection power and bug repair capabilities. To demonstrate the benefits, we study the repair of termination bugs in sequential and concurrent programs. Our analysis shows that integrating dynamic APR with formal analysis techniques, such as termination provers and software model checkers, reduces complexity and improves the overall reliability of these repairs. We empirically investigate how well the hybrid approach can repair termination and performance bugs by experimenting with hybrids that integrate different APR approaches with termination provers and execution time monitors. Our findings indicate that hybrid repair holds promise for handling termination and performance bugs. However, the capability of the chosen tools and the completeness of the available correctness specification affects the quality of the patches that can be produced.
Omar I. Al-Bataineh, Leon Moonen, Linas Vidziunas
J. Syst. Softw.2
2023 An Exploratory Literature Study on Sharing and Energy Use of Language Models for Source Code
abstract
Context: Large language models trained on source code can support a variety of software development tasks, such as code recommendation and program repair. Large amounts of data for training such models benefit the models' performance. However, the size of the data and models results in long training times and high energy consumption. While publishing source code allows for replicability, users need to repeat the expensive training process if models are not shared. Goals: The main goal of the study is to investigate if publications that trained language models for software engineering (SE) tasks share source code and trained artifacts. The second goal is to analyze the transparency on training energy usage. Methods: We perform a snowballing-based literature search to find publications on language models for source code, and analyze their reusability from a sustainability standpoint. Results: From a total of 494 unique publications, we identified 293 relevant publications that use language models to address code-related tasks. Among them, 27% (79 out of 293) make artifacts available for reuse. This can be in the form of tools or IDE plugins designed for specific tasks or task-agnostic models that can be fine-tuned for a variety of downstream tasks. Moreover, we collect insights on the hardware used for model training, as well as training time, which together determine the energy consumption of the development process. Conclusion: We find that there are deficiencies in the sharing of information and artifacts for current studies on source code models for software engineering tasks, with 40% of the surveyed papers not sharing source code or trained artifacts. We recommend the sharing of source code as well as trained artifacts, to enable sustainable reproducibility. Moreover, comprehensive information on training times and hardware configurations should be shared for transparency on a model's carbon footprint.
Max Hort, Anastasiia Grishina, Leon Moonen
ESEM3
2023 Fully Autonomous Programming with Large Language Models
abstract
Current approaches to program synthesis with Large Language Models (LLMs) exhibit a "near miss syndrome": they tend to generate programs that semantically resemble the correct answer (as measured by text similarity metrics or human evaluation), but achieve a low or even zero accuracy as measured by unit tests due to small imperfections, such as the wrong input or output format. This calls for an approach known as Synthesize, Execute, Debug (SED), whereby a draft of the solution is generated first, followed by a program repair phase addressing the failed tests. To effectively apply this approach to instruction-driven LLMs, one needs to determine which prompts perform best as instructions for LLMs, as well as strike a balance between repairing unsuccessful programs and replacing them with newly generated ones. We explore these trade-offs empirically, comparing replace-focused, repair-focused, and hybrid debug strategies, as well as different template-based and model-based prompt-generation techniques. We use OpenAI Codex as the LLM and Program Synthesis Benchmark 2 as a database of problem descriptions and tests for evaluation. The resulting framework outperforms both conventional usage of Codex without the repair phase and traditional genetic programming approaches.
Vadim Liventsev, Anastasiia Grishina, Aki Härmä, Leon Moonen
GECCO4
2023 CHESS: A Framework for Evaluation of Self-Adaptive Systems Based on Chaos Engineering
abstract
There is an increasing need to assess the correct behavior of self-adaptive and self-healing systems due to their adoption in critical and highly dynamic environments. However, there is a lack of systematic evaluation methods for self-adaptive and self-healing systems. We proposed CHESS, a novel approach to address this gap by evaluating self-adaptive and self-healing systems through fault injection based on chaos engineering (CE).The artifact presented in this paper provides an extensive overview of the use of CHESS through two microservice-based case studies: a smart office case study and an existing demo application called Yelb. It comes with a managing system service, a self-monitoring service, as well as five fault injection scenarios covering infrastructure faults and functional faults. Each of these components can be easily extended or replaced to adopt the CHESS approach to a new case study, help explore its promises and limitations, and identify directions for future research.
Sehrish Malik, Moeen Ali Naqvi, Leon Moonen
SEAMS3
2023 The EarlyBIRD Catches the Bug: On Exploiting Early Layers of Encoder Models for More Efficient Code Classification
abstract
The use of modern Natural Language Processing (NLP) techniques has shown to be beneficial for software engineering tasks, such as vulnerability detection and type inference. However, training deep NLP models requires significant computational resources. This paper explores techniques that aim at achieving the best usage of resources and available information in these models.
Anastasiia Grishina, Max Hort, Leon Moonen
ESEC/SIGSOFT FSE3
2022 Towards Extending the Range of Bugs That Automated Program Repair Can Handle
abstract
Modern automated program repair (APR) is well-tuned to finding and repairing bugs that introduce observable erroneous behavior to a program. However, a significant class of bugs does not lead to such observable behavior (e.g., liveness/termination bugs, non-functional bugs, and information flow bugs). Such bugs can generally not be handled with current APR approaches, so, as a community, we need to develop complementary techniques.To stimulate the systematic study of alternative APR approaches and hybrid APR combinations, we devise a novel bug classification system that enables methodical analysis of their bug detection power and bug repair capabilities. To demonstrate the benefits, we analyze the repair of termination bugs in sequential and concurrent programs. The study shows that integrating dynamic APR with formal analysis techniques, such as termination provers and software model checkers, reduces complexity and improves the overall reliability of these repairs.
Omar I. Al-Bataineh, Leon Moonen
QRS2
2022 Assessing the Impact of Execution Environment on Observation-Based Slicing
abstract
Program slicing reduces a program to a smaller version that retains a chosen computation, referred to as a slicing criterion. One recent multi-lingual slicing approach, observation-based slicing (ORBS), speculatively deletes parts of the program and then executes the code. If the behavior of the slicing criteria is unchanged, the speculative deletion is made permanent. While this makes ORBS language agnostic, it can lead to the production of some non-intuitive slices. One particular challenge is when the execution environment plays a role. For example, ORBS will delete the line “$\mathrm{a}=0$” if the memory location assigned to a contains zero before executing this statement, because the deletion does not affect the value of a and thus the slicing criterion. Consequently, slices can differ between execution environments due to factors such as initialization and call stack reuse. The technique considered,$\boldsymbol{n}$VORBS, attempts to ameliorate this problem by validating a candidate slice in$\boldsymbol{n}$different execution environments. We conduct an empirical study to collect initial insights into how often the execution environment leads to slice differences. Specifically, we compare and contrast the slices produced by seven different instantiations of$\boldsymbol{n}$VORBS. Looking forward, the technique can be seen as a variation on metamorphic testing, and thus suggests how ideas from metamorphic testing might be used to improve dynamic program analysis.
Dave W. Binkley, Leon Moonen
SCAM2
2022 Featherweight assisted vulnerability discovery
abstract
Predicting vulnerable source code helps to focus the attention of a developer, or a program analysis technique, on those parts of the code that need to be examined with more scrutiny. Recent work proposed the use of function names as semantic cues that can be learned by a deep neural network (DNN) to aid in the hunt for vulnerability of functions. Combining identifier splitting, which we use to split each function name into its constituent words, with a novel frequency-based algorithm, we explore the extent to which the words that make up a function’s name can be used to predict potentially vulnerable functions. In contrast to the lightweight prediction provided by a DNN considering only function names, avoiding the need for a DNN provides featherweight prediction. The underlying idea is that function names that contain certain “dangerous” words are more likely to accompany vulnerable functions. Of course, this assumes that the frequency-based algorithm can be properly tuned to focus on truly dangerous words. Because it is more transparent than a DNN, which behaves as a “black box” and thus provides no insight into the rationalization underlying its decisions, the frequency-based algorithm enables us to investigate the inner workings of the DNN. If successful, this investigation into what the DNN does and does not learn will help us train more effective future models. We empirically evaluate our approach on a heterogeneous dataset containing over 73 000 functions labeled vulnerable, and over 950 000 functions labeled benign. Our analysis shows that words alone account for a significant portion of the DNN’s classification ability. We also find that words are of greatest value in the datasets with a more homogeneous vocabulary. Thus, when working within the scope of a given project, where the vocabulary is unavoidably homogeneous, our approach provides a cheaper, potentially complementary, technique to aid in the hunt for source-code vulnerabilities. Finally, this approach has the advantage that it is viable with orders of magnitude less training data.
Dave W. Binkley, Leon Moonen, Sibren Isaacman
Inf. Softw. Technol.2
2021 Towards More Reliable Automated Program Repair by Integrating Static Analysis Techniques
abstract
A long-standing open challenge for automated program repair is the overfitting problem, which is caused by having insufficient or incomplete specifications to validate whether a generated patch is correct or not. Most available repair systems rely on weak specifications (i.e., specifications that are synthesized from test cases) which limits the quality of generated repairs. To strengthen specifications and improve the quality of repairs, we propose to closer integrate static bug detection techniques with automated program repair. The integration combines automated program repair with static analysis techniques in such a way that bug detection patterns can be synthesized into specifications that the repair system can use. We explore the feasibility of such integration using two types of bugs: arithmetic bugs, such as integer overflow, and logical bugs, such as termination bugs. As part of our analysis, we make several observations that help to improve patch generation for these classes of bugs. Moreover, these observations assist with narrowing down the candidate patch search space, and inferring an effective search order.
Omar I. Al-Bataineh, Anastasiia Grishina, Leon Moonen
QRS3
2021 Adaptive Immunity for Software: Towards Autonomous Self-healing Systems
abstract
Testing and code reviews are known techniques to improve the quality and robustness of software. Unfortunately, the complexity of modern software systems makes it impossible to anticipate all possible problems that can occur at runtime, which limits what issues can be found using testing and reviews. Thus, it is of interest to consider autonomous self-healing software systems, which can automatically detect, diagnose, and contain unanticipated problems at runtime. Most research in this area has adopted a model-driven approach, where actual behavior is checked against a model specifying the intended behavior, and a controller takes action when the system behaves outside of the specification. However, it is not easy to develop these specifications, nor to keep them up-to-date as the system evolves. We pose that, with the recent advances in machine learning, such models may be learned by observing the system. Moreover, we argue that artificial immune systems (AISs) are particularly well-suited for building self-healing systems, because of their anomaly detection and diagnosis capabilities. We present the state-of-the-art in self-healing systems and in AISs, surveying some of the research directions that have been considered up to now. To help advance the state-of-the-art, we develop a research agenda for building self-healing software systems using AISs, identifying required foundations, and promising research directions.
Moeen Ali Naqvi, Merve Astekin, Sehrish Malik, Leon Moonen
SANER4
2020 Spectrum-Based Log Diagnosis
abstract
Background: Continuous Engineering practices are increasingly adopted in modern software development. However, a frequently reported need is for more effective methods to analyze the massive amounts of data resulting from the numerous build and test runs. Aims: We present and evaluate Spectrum-Based Log Diagnosis (SBLD), a method to help developers quickly diagnose problems found in complex integration and deployment runs. Inspired by Spectrum-Based Fault Localization, SBLD leverages the differences in event occurrences between logs for failing and passing runs, to highlight events that are stronger associated with failing runs.
Carl Martin Rosenberg, Leon Moonen
ESEM2
2020 On Adaptive Change Recommendation
abstract
As the complexity of a software system grows, it becomes harder for developers to be aware of all the dependencies between its artifacts (e.g., files or methods). Change impact analysis helps to overcome this challenge, by recommending relevant source-code artifacts related to a developer’s current changes. Association rule mining has shown promise in determining change impact by uncovering relevant patterns in the system’s change history. State-of-the-art change impact mining typically uses a change history of tens of thousands of transactions. For efficiency, targeted association rule mining constrains the transactions used to those potentially relevant to answering a particular query. However, it still considers all the relevant transactions in the history. This paper presents Atari, a new adaptive approach that further constrains targeted association rule mining by considering a dynamic selection of the relevant transactions. Our investigation of adaptive change impact mining empirically studies fourteen algorithm variants. We show that adaptive algorithms are viable, can be just as applicable as the start-of-the-art complete-history algorithms, and even outperform them for certain queries. However, more important than this direct comparison, our investigation motivates and lays the groundwork for the future study of adaptive techniques, and their application to challenges such as on-the-fly impact analysis at GitHub-scale.
Leon Moonen, Dave W. Binkley, Sydney Pugh
J. Syst. Softw.1
2018 On the Use of Automated Log Clustering to Support Effort Reduction in Continuous Engineering
abstract
Continuous engineering (CE) practices, such as continuous integration and continuous deployment, have become key to modern software development. They are characterized by short automated build and test cycles that give developers early feedback on potential issues. CE practices help to release software more frequently, and reduces risk by increasing incrementality. However, effective use of CE practices in industrial projects requires making sense of the vast amounts of data that results from the repeated build and test cycles. The goal of this paper is to investigate to what extent these data can be treated more effectively by automatically grouping logs of runs that failed for the same underlying reasons, and what effort reduction can be achieved. To this end, we replicate and extend earlier work on system log clustering to evaluate its efficacy in the CE context, and to investigate the impact of five alternative log vectorization techniques. We built a prototype tool that is used to conduct an empirical case study on continuous deployment logs provided by our industrial collaborator. Questions to be answered include: (1) Can we reduce the effort needed to discover all latent issues in a set of failing runs? (2) How to best leverage the contrast between passing and failing runs to increase accuracy? (3) What trade-offs are there between effort reduction and accuracy? We present a quantitative and qualitative analysis of the results of our study. We conclude by evaluating the trade-offs, and give recommendations for applying this approach in practice.
Carl Martin Rosenberg, Leon Moonen
APSEC2
2018 Improving problem identification via automated log clustering using dimensionality reduction
abstract
Background: Continuous engineering practices, such as continuous integration and continuous deployment, see increased adoption in modern software development. A frequently reported challenge for adopting these practices is the need to make sense of the large amounts of data that they generate.
Carl Martin Rosenberg, Leon Moonen
ESEM2
2018 [Research Paper] The Case for Adaptive Change Recommendation
abstract
As the complexity of a software system grows, it becomes increasingly difficult for developers to be aware of all the dependencies that exist between artifacts (e.g., files or methods) of the system. Change impact analysis helps to overcome this problem, as it recommends to a developer relevant source-code artifacts related to her current changes. Association rule mining has shown promise in determining change impact by uncovering relevant patterns in the system's change history. State-of-the-art change impact mining algorithms typically make use of a change history of tens of thousands of transactions. For efficiency, targeted association rule mining focuses on only those transactions potentially relevant to answering a particular query. However, even targeted algorithms must consider the complete set of relevant transactions in the history. This paper presents ATARI, a new adaptive approach to association rule mining that considers a dynamic selection of the relevant transactions. It can be viewed as a further constrained version of targeted association rule mining, in which as few as a single transaction might be considered when determining change impact. Our investigation of adaptive change impact mining empirically studies seven algorithm variants. We show that adaptive algorithms are viable, can be just as applicable as the start-of-the-art complete-history algorithms, and even outperform them for certain queries. However, more important than the direct comparison, our investigation lays necessary groundwork for the future study of adaptive techniques and their application to challenges such as the on-the-fly style of impact analysis that is needed at the GitHub-scale.
Sydney Pugh, Dave W. Binkley, Leon Moonen
SCAM3
2018 What are the effects of history length and age on mining software change impact?
Leon Moonen, Thomas Rolfsnes, Dave W. Binkley, Stefano Di Alesio
Empir. Softw. Eng.1
2018 Aggregating Association Rules to Improve Change Recommendation
Thomas Rolfsnes, Leon Moonen, Stefano Di Alesio, Razieh Behjati, Dave W. Binkley
Empir. Softw. Eng.2
2017 Predicting relevance of change recommendations
abstract
Software change recommendation seeks to suggest artifacts (e.g., files or methods) that are related to changes made by a developer, and thus identifies possible omissions or next steps. While one obvious challenge for recommender systems is to produce accurate recommendations, a complimentary challenge is to rank recommendations based on their relevance. In this paper, we address this challenge for recommendation systems that are based on evolutionary coupling. Such systems use targeted association-rule mining to identify relevant patterns in a software system's change history. Traditionally, this process involves ranking artifacts using interestingness measures such as confidence and support. However, these measures often fall short when used to assess recommendation relevance. We propose the use of random forest classification models to assess recommendation relevance. This approach improves on past use of various interestingness measures by learning from previous change recommendations. We empirically evaluate our approach on fourteen open source systems and two systems from our industry partners. Furthermore, we consider complimenting two mining algorithms: Co-Change and Tarmaq. The results find that random forest classification significantly outperforms previous approaches, receives lower Brier scores, and has superior trade-off between precision and recall. The results are consistent across software system and mining algorithm.
Thomas Rolfsnes, Leon Moonen, Dave W. Binkley
ASE2
2016 Practical guidelines for change recommendation using association rule mining
abstract
Association rule mining is an unsupervised learning technique that infers relationships among items in a data set. This technique has been successfully used to analyze a system's change history and uncover evolutionary coupling between system artifacts. Evolutionary coupling can, in turn, be used to recommend artifacts that are potentially affected by a given set of changes to the system. In general, the quality of such recommendations is affected by (1) the values selected for various parameters of the mining algorithm, (2) characteristics of the set of changes used to derive a recommendation, and (3) characteristics of the system's change history for which recommendations are generated.
Leon Moonen, Stefano Di Alesio, Dave W. Binkley, Thomas Rolfsnes
ASE1
2016 Improving change recommendation using aggregated association rules
abstract
Past research has proposed association rule mining as a means to uncover the evolutionary coupling from a system's change history. These couplings have various applications, such as improving system decomposition and recommending related changes during development. The strength of the coupling can be characterized using a variety of interestingness measures. Existing recommendation engines typically use only the rule with the highest interestingness value in situations where more than one rule applies. In contrast, we argue that multiple applicable rules indicate increased evidence, and hypothesize that the aggregation of such rules can be exploited to provide more accurate recommendations.
Thomas Rolfsnes, Leon Moonen, Stefano Di Alesio, Razieh Behjati, Dave W. Binkley
MSR2
2016 Exploring the Effects of History Length and Age on Mining Software Change Impact
abstract
The goal of Software Change Impact Analysis is to identify artifacts (typically source-code files) potentially affected by a change. Recently, there is an increased interest in mining software change impact based on evolutionary coupling. A particularly promising approach uses association rule mining to uncover potentially affected artifacts from patterns in the system's change history. Two main considerations when using this approach are the history length, the number of transactions from the change history used to identify the impact of a change, and history age, the number of transactions that have occurred since patterns were last mined from the history. Although history length and age can significantly affect the quality of mining results, few guidelines exist on how to best select appropriate values for these two parameters. In this paper, we empirically investigate the effects of history length and age on the quality of change impact analysis using mined evolutionary couplings. Specifically, we report on a series of systematic experiments involving the change histories of two large industrial systems and 17 large open source systems. In these experiments, we vary the length and age of the history used to mine software change impact, and assess how this affects precision and applicability. Results from the study are used to derive practical guidelines for choosing history length and age when applying association rule mining to conduct software change impact analysis.
Leon Moonen, Stefano Di Alesio, Thomas Rolfsnes, Dave W. Binkley
SCAM1
2016 Generalizing the Analysis of Evolutionary Coupling for Software Change Impact Analysis
abstract
Software change impact analysis aims to find artifacts potentially affected by a change. Typical approaches apply language-specific static or dynamic dependence analysis, and are thus restricted to homogeneous systems. This restriction is a major drawback given today's increasingly heterogeneous software. Evolutionary coupling has been proposed as a language-agnostic alternative that mines relations between source-code entities from the system's change history. Unfortunately, existing evolutionary coupling based techniques fall short. For example, using Singular Value Decomposition (SVD) quickly becomes computationally expensive. An efficient alternative applies targeted association rule mining, but the most widely known approach (ROSE) has restricted applicability: experiments on two large industrial systems, and four large open source systems, show that ROSE can only identify dependencies about 25% of the time. To overcome this limitation, we introduce TARMAQ, a new algorithm for mining evolutionary coupling. Empirically evaluated on the same six systems, TARMAQ performs consistently better than ROSE and SVD, is applicable 100% of the time, and runs orders of magnitude faster than SVD. We conclude that the proposed algorithm is a significant step forward towards achieving robust change impact analysis for heterogeneous systems.
Thomas Rolfsnes, Stefano Di Alesio, Razieh Behjati, Leon Moonen, Dave W. Binkley
SANER4
2016 Analyzing and visualizing information flow in heterogeneous component-based software systems
Leon Moonen, Amir Reza Yazdanshenas
Inf. Softw. Technol.1
2016 Introduction to the special issue on software maintenance and evolution
abstract
It is our pleasure to introduce you to the papers in this Special Issue based on the 30th International Conference on Software Maintenance and Evolution (ICSME 2014). ICSME is the premier international venue in software maintenance and evolution, where participants from academia, government, and industry gather to share and discuss their ideas on and experiences with solving critical software maintenance problems. In response to the call for research papers, we received 267 abstracts and 210 full paper submissions. Each submitted paper was reviewed by at least three members of the program committee (PC), who were selected through bidding to create a good match between paper topic and PC member expertise; each PC member reviewed 9–10 papers over several weeks, with collectively 632 reviews submitted. During the week-long discussion period, reviewers submitted over 1000 comments. In the end, 40 high-quality papers were accepted for publication in the conference proceedings, yielding an acceptance rate of 19%. The accepted papers covered a broad range of topics in software maintenance and evolution, including developer knowledge, evolving systems, developer support, technical debt, managing change, empirical studies, fault localization, software quality, patches, recommender systems, and software clones. Guided by the reviews and discussions, ICSME 2014 Program Co-Chairs Leon Moonen and Lori Pollock carefully selected nine outstanding papers of the 40 accepted for the conference and invited their authors to submit a significantly extended version of their conference paper to this special issue in the Journal of Software: Evolution and Process (JSEP). Six of the nine invited papers were extended by their authors and subjected to the rigorous JSEP reviewing process, thus undergoing additional rounds of reviews and revisions. Eventually, the following four papers successfully completed the review process and are contained in this special issue. The paper ‘A Simple, Efficient, Context Sensitive Approach for Code Completion’ by Muhammad Asaduzzaman, Chanchal K. Roy, Kevin A. Schneider, and Daqing Hou describes a technique for method call completion that uses the type name and context to search for method calls whose contexts match with that of the receiver object. A database of context–method pairs is created by collecting code examples from repositories. The proposed approach was shown to either outperform or perform as well as state-of-the-art techniques for code completion based on statistical language models. The paper ‘An Empirical Study on How Expert Knowledge Affects Bug Reports’ by Da Huo, Tao Ding, Collin McMillan, and Malcom Gethers describes an empirical study of the textual difference between bug reports written by experts and non-experts. The study showed that experts and non-experts wrote bug reports differently. The findings support the hypothesis that expert knowledge affects the way in which people write bug reports. The paper ‘How Does Code Obfuscation Impact Energy Usage?’ by Cagri Sahin, Philip Tornquist, Ryan Mckenna, Zachary Pearson, and James Clause describes an empirical study into the energy impacts of applying different code obfuscations. In addition to investigating how different obfuscations in four obfuscation tools alter the overall energy usage of an application, the paper also studies whether the impacts of obfuscations are likely to be meaningful for mobile application users. The results support the notion that developers can protect their applications without impacting the battery life of the devices where their applications execute. The paper ‘Empirical Analysis of the Relationship between CC and SLOC in a Large Corpus of Java Methods and C Functions’ by Davy Landman, Alexander Serebrenik, and Jurgen Vinju describes an extensive literature study of the Cyclomatic Complexity (CC) and Source Lines of Code (SLOC) correlation results, followed by a correlation study of CC/SLOC on large Java and C corpora. In contrast to the majority of the previous studies, this study did not observe a strong linear correlation between CC and SLOC of Java methods and C functions. We hope that readers will enjoy this special issue and gain useful insights from the four papers presented. We would like to thank all the authors who submitted papers to the conference and to this special issue. In addition, we would like to thank the members of the ICSME 2014 program committee and the external reviewers for their time, careful reviews, and active discussions of the submitted papers, which helped make this special issue special. This kind of service is important to the health of the community and the quality of its publications. Finally, we would like to thank the editorial board of the Journal of Software: Evolution and Process and the publisher Wiley for providing us with the opportunity to devote this issue to the best of ICSME 2014. We also thank JSEP Editor Gerardo Canfora for providing expert guidance and important advice throughout the process. Enjoy!
Leon Moonen, Lori L. Pollock
J. Softw. Evol. Process.1
2016 Guest editor's introduction to the Special Issue on Program Comprehension (ICPC 2014)
abstract
It is our pleasure to introduce you to the papers in this Special Issue based on the 22th International Conference on Program Comprehension (ICPC 2014).Program comprehension plays a central role in most of the phases of the software development life cycle, where it helps facilitate reuse, inspection, maintenance, reverse engineering, reengineering, migration, and extension of existing software systems.The International Conference on Program Comprehension (ICPC) is the primary venue for work in the area of program comprehension.It is also one of the leading venues for work in the areas of software analysis, reverse engineering, software evolution, and software visualization.ICPC provides an opportunity for researchers and industry practitioners to present and discuss the state-of-the-art and the state-of-the-practice in program comprehension and related areas.ICPC 2014 took place during June 2-3, 2014, in Hyderabad, India, and was co-located with the International Conference on Software Engineering (ICSE 2014).ICPC 2014 received a record number of submissions ( 76) from 19 different countries, which allowed us to assemble an excellent program that continues ICPC's tradition of providing a highquality venue for sharing the latest advances in program comprehension.The program included 20 full research papers, 11 short papers and 5 tool demonstration papers.Of these 20 full research papers, five were invited to submit an extended version to the Journal of Software: Evolution and Process.After a rigorous reviewing process with at least three reviewers per paper, four papers were accepted for publication in this special section.The paper entitled 'Framing Program Comprehension as Fault Localization' by Alexandre Perez and Rui Abreu proposes an approach, coined Spectrum-based Feature Comprehension (SFC), that borrows techniques from software-fault localization that were proven to be effective even when debugging large applications.SFC analyses the program by exploiting run-time information from test case executions to identify the components that are important for a given feature, helping software engineers to understand how a program is structured and each of the functionality's dependencies are.They present a toolset, coined PANGOLIN, that implements SFC and displays its report to the user using an intuitive visualization.The paper entitled 'Searching Crowd Knowledge to Recommend Solutions for API Usage Tasks' by Eduardo C. Campos, Lucas B. L. de Souza, and Marcelo de A. Maia presents an approach that makes use of 'crowd knowledge' in Stack Overflow to recommend information that can assist developer activities.This strategy recommends a ranked list of question-answer pairs from Stack Overflow based on a query.The criteria for ranking are based on three main aspects: the textual similarity of the pairs with respect to the query related to the developer's problem, the quality of the pairs, and a filtering mechanism that considers only 'how-to' posts.The paper entitled, 'AmaLgam+: Composing Rich Information Sources for Accurate Bug Localization' by Shaowei Wang and David Lo proposes AmaLgam+, which is a method for locating relevant buggy files that combines five sources of information, namely, version history, similar reports, structure, stack traces, and reporter information.They perform a large-scale experiment on four open source projects, namely, AspectJ, Eclipse, SWT, and ZXing to localize more than 3000 bugs and compare AmaLgam+ results with those of six state-of-the-art bug localization approaches.The study showed that the newly proposed method outperforms these existing approaches in terms of mean average precision.
Chanchal Kumar Roy, Andrew Begel, Leon Moonen
J. Softw. Evol. Process.3
2016 An Industrial Survey of Safety Evidence Change Impact Analysis Practice
abstract
Context. In many application domains, critical systems must comply with safety standards. This involves gathering safety evidence in the form of artefacts such as safety analyses, system specifications, and testing results. These artefacts can evolve during a system's lifecycle, creating a need for change impact analysis to guarantee that system safety and compliance are not jeopardised. Objective. We aim to provide new insights into how safety evidence change impact analysis is addressed in practice. The knowledge about this activity is limited despite the extensive research that has been conducted on change impact analysis and on safety evidence management. Method. We conducted an industrial survey on the circumstances under which safety evidence change impact analysis is addressed, the tool support used, and the challenges faced. Results. We obtained 97 valid responses representing 16 application domains, 28 countries, and 47 safety standards. The respondents had most often performed safety evidence change impact analysis during system development, from system specifications, and fully manually. No commercial change impact analysis tool was reported as used for all artefact types and insufficient tool support was the most frequent challenge. Conclusion. The results suggest that the different artefact types used as safety evidence co-evolve. In addition, the evolution of safety cases should probably be better managed, the level of automation in safety evidence change impact analysis is low, and the state of the practice can benefit from over 20 improvement areas.
Jose Luis de la Vara, Markus Borg, Krzysztof Wnuk, Leon Moonen
IEEE Trans. Software Eng.4
2015 Towards evidence-based recommendations to guide the evolution of component-based product families
Leon Moonen
Sci. Comput. Program.1
2014 Assembling multiple-case studies: potential, principles and practical considerations
abstract
Case studies are a research method aimed at holistically analyzing a phenomenon in its context. Despite the fact that they cannot be used to answer the same precise research questions as, e.g., can be addressed by controlled experiments, case studies can cope much better with situations having several variables of interest, multiple sources of evidence, or rich contexts that cannot be controlled or isolated. As such, case studies are a promising instrument to study the complex phenomena at play in Software Engineering.
Aiko Fallas Yamashita, Leon Moonen
EASE2
2013 Exploring the impact of inter-smell relations on software maintainability: an empirical study
abstract
Code smells are indicators of issues with source code quality that may hinder evolution. While previous studies mainly focused on the effects of individual code smells on maintainability, we conjecture that not only the individual code smells but also the interactions between code smells affect maintenance. We empirically investigate the interactions amongst 12 code smells and analyze how those interactions relate to maintenance problems. Professional developers were hired for a period of four weeks to implement change requests on four medium-sized Java systems with known smells. On a daily basis, we recorded what specific problems they faced and which artifacts were associated with them. Code smells were automatically detected in the pre-maintenance versions of the systems and analyzed using Principal Component Analysis (PCA) to identify patterns of co-located code smells. Analysis of these factors with the observed maintenance problems revealed how smells that were co-located in the same artifact interacted with each other, and affected maintainability. Moreover, we found that code smell interactions occurred across coupled artifacts, with comparable negative effects as same-artifact co-location. We argue that future studies into the effects of code smells on maintainability should integrate dependency analysis in their process so that they can obtain a more complete understanding by including such coupled interactions.
Aiko Fallas Yamashita, Leon Moonen
ICSE2
2013 Towards a Taxonomy of Programming-Related Difficulties during Maintenance
abstract
Empirical studies that investigate the relationship between source code characteristics and maintenance outcomes rarely use causal models to explain the relations between the code characteristics and the outcomes. We conjecture that the lack of a comprehensive catalogue of programming-related difficulties and their effects on different maintenance outcomes is one of the reasons behind this. This paper takes the first step in addressing this situation based on empirical evidence collected in a longitudinal maintenance study on four systems. Professional developers were hired to implement a number of changes in each of the systems. These activities were observed in detail over a period of 7 weeks, during which we recorded on a daily basis what specific problems they faced. The collected data was transcribed and analyzed using open and axial coding. Based on an analysis of these results, we propose a preliminary taxonomy to describe the programming-related difficulties that developers face during maintenance. Our intention is not to replace the existing categorizations/taxonomies, but to take the first steps towards an integrated, comprehensive catalogue by aligning our empirical observations and the earlier literature.
Aiko Fallas Yamashita, Leon Moonen
ICSM2
2013 First International Workshop on Multi Product Line Engineering (MultiPLE 2013)
abstract
In an industrial context, software systems are rarely developed by a single organization. For software product lines, this means that various organizations collaborate to provide and integrate the assets used in a product line. It is not uncommon that these assets themselves are built as product lines, a practice which is referred to as multi product lines. This cross-organizational distribution of reusable assets leads to numerous challenges, such as inconsistent configuration, costly and time-consuming integration, diverging evolution speed and direction, and inadequate testing.
Leon Moonen, Mithun Acharya, Razieh Behjati, Bedir Tekinerdogan, Rick Rabiser, Kyo Chul Kang
SPLC1
2013 To what extent can maintenance problems be predicted by code smell detection? - An empirical study
Aiko Fallas Yamashita, Leon Moonen
Inf. Softw. Technol.2
2012 Do code smells reflect important maintainability aspects?
abstract
Code smells are manifestations of design flaws that can degrade code maintainability. As such, the existence of code smells seems an ideal indicator for maintainability assessments. However, to achieve comprehensive and accurate evaluations based on code smells, we need to know how well they reflect factors affecting maintainability. After identifying which maintainability factors are reflected by code smells and which not, we can use complementary means to assess the factors that are not addressed by smells. This paper reports on an empirical study that investigates the extent to which code smells reflect factors affecting maintainability that have been identified as important by programmers. We consider two sources for our analysis: (1) expert-based maintainability assessments of four Java systems before they entered a maintenance project, and (2) observations and interviews with professional developers who maintained these systems during 14 working days and implemented a number of change requests.
Aiko Fallas Yamashita, Leon Moonen
ICSM2
2012 Fine-grained change impact analysis for component-based product families
abstract
Developing software product-lines based on a set of shared components is a proven tactic to enhance reuse, quality, and time to market in producing a portfolio of products. Large-scale product families face rapidly increasing maintenance challenges as their evolution can happen both as a result of collective domain engineering activities, and as a result of product-specific developments. To make informed decisions about prospective modifications, developers need to estimate what other sections of the system will be affected and need attention, which is known as change impact analysis. This paper contributes a method to carry out change impact analysis in a component-based product family, based on system-wide information flow analysis. We use static program slicing as the underlying analysis technique, and use model-driven engineering (MDE) techniques to propagate the ripple effects from a source code modification into all members of the product family. In addition, our approach ranks results based on an approximation of the scale of their impact. We have implemented our approach in a prototype tool, called Richter, which was evaluated on a real-world product family.
Amir Reza Yazdanshenas, Leon Moonen
ICSM2
2012 Tracking and visualizing information flow in component-based systems
abstract
Component-based software engineering is aimed at managing the complexity of large-scale software development by composing systems from reusable parts. In order to understand or validate the behavior of a given system, one needs to acquire understanding of the components involved in combination with understanding how these components are instantiated, initialized and interconnected in the particular system. In practice, this task is often hindered by the heterogeneous nature of source and configuration artifacts and there is little to no tool support to help software engineers with such a system-wide analysis. This paper contributes a method to track and visualize information flow in a component-based system at various levels of abstraction. We propose a hierarchy of 5 interconnected views to support the comprehension needs of both safety domain experts and developers from our industrial partner. We discuss the implementation of our approach in a prototype tool, and present an initial qualitative evaluation of the effectiveness and usability of the proposed views for software development and software certification. The prototype was already found to be very useful and a number of directions for further improvement were suggested. We conclude by discussing these improvements and lessons learned.
Amir Reza Yazdanshenas, Leon Moonen
ICPC2
2011 Crossing the boundaries while analyzing heterogeneous component-based software systems
abstract
One way to manage the complexity of software systems is to compose them from reusable components, instead of starting from scratch. Components may be implemented in different programming languages and are tied together using configuration files, or glue code, defining instantiation, initialization and interconnections. Although correctly engineering the composition and configuration of components is crucial for the overall behavior, there is surprisingly little support for incorporating this information in the static verification and validation of these systems. Analyzing the properties of programs within closed code boundaries has been studied for some decades and is well-established. This paper contributes a method to support analysis across the components of a component-based system. We build upon the Knowledge Discovery Metamodel to reverse engineer homogeneous models for systems composed of heterogeneous artifacts. Our method is implemented in a prototype tool that has been successfully used to track information flow across the components of a component-based system using program slicing.
Amir Reza Yazdanshenas, Leon Moonen
ICSM2
2011 Keynotes
abstract
Summary form only given. Program understanding is one of the core activities in software engineering, and one of the main challenges in getting a grip on large industrial systems is finding appropriate representations that support the comprehension process. In this talk, we will investigate the benefits and challenges of using a map metaphor to help software engineers explore and understand software systems. We will analyze what factors influence the legibitility of a software map, i.e. what makes the information contained in a map easy to understand, interpret and remember. In addition, we will look at what has been done in city planning and architecture to make it easier for people find their way in unknown terrain, and reflect on opportunities for using these results in program comprehension research. Leon Moonen is a research scientist at Simula Research Laboratory in Norway. His research is aimed at developing better techniques and tools for the exploration, assessment and evolution of large industrial software systems. His research interests include program comprehension, reverse engineering, program analysis, software visualisation and empirical software engineering. Current topics include the reconstruction and visualization of higher level abstractions (models) from the development artifacts of existing software systems, and the use of these models in software inspection, verification and validation. He is co-founder of the Software Improvement Group, a company that specializes in the use of source code analysis to help organizations get control over their software systems.
Leon Moonen
ICPC1
2009 Using concept mapping for maintainability assessments
abstract
Many important phenomena within software engineering are difficult to define and measure. One example is software maintainability, which has been the subject of considerable research and is believed to be a critical determinant of total software costs. We propose using concept mapping, a well-grounded method used in social research, to operationalize the concept of software maintainability according to a given goal and perspective in a concrete setting. We apply this method to describe four systems that were developed as part of an industrial multiple-case study. The outcome is a conceptual map that displays an arrangement of maintainability constructs, their interrelations, and corresponding measures. Our experience is that concept mapping (1) provides a structured way of combining static code analysis and expert judgment; (2) helps in the tailoring of the choice of measures to a particular system context; and (3) supports the mapping between software measures and aspects of software maintainability. As such, it constitutes a useful addition to existing frameworks for evaluating quality, such as ISO/IEC 9126 and GQM, and tools for static measurement of software code. Overall, concept mapping provides a systematic, structured, and repeatable method for developing constructs and measures, not only of maintainability, but also of software engineering phenomena in general.
Aiko Fallas Yamashita, Hans Christian Benestad, Bente Anda, Per Einar Arnstad, Dag I. K. Sjøberg, Leon Moonen
ESEM6
2009 Maintenance and agile development: Challenges, opportunities and future directions
abstract
Software entropy is a phenomenon where repeated changes gradually degrade the structure of the system, making it hard to understand and maintain. This phenomenon imposes challenges for organizations that have moved to agile methods from other processes, despite agile's focus on adaptability and responsiveness to change. We have investigated this issue through an industrial case study, and reviewed the literature on addressing software entropy, focussing on the detection of ldquocode smellsrdquo and their treatment by refactoring. We found that in order to remain agile despite of software entropy, developers need better support for understanding, planning and testing the impact of changes. However, it is exactly work on refactoring decision support and task complexity analysis that is lacking in literature. Based on our findings, we discuss strategies for dealing with entropy in this context and present avenues for future research.
Geir Kjetil Hanssen, Aiko Fallas Yamashita, Reidar Conradi, Leon Moonen
ICSM4
2009 Evaluating the relation between coding standard violations and faultswithin and across software versions
abstract
In spite of the widespread use of coding standards and tools enforcing their rules, there is little empirical evidence supporting the intuition that they prevent the introduction of faults in software. In previous work, we performed a pilot study to assess the relation between rule violations and actual faults, using the MISRA C 2004 standard on an industrial case. In this paper, we investigate three different aspects of the relation between violations and faults on a larger case study, and compare the results across the two projects. We find that 10 rules in the standard are significant predictors of fault location.
Cathal Boogerd, Leon Moonen
MSR2
2009 An integrated crosscutting concern migration strategy and its semi-automated application to JHotDraw
abstract
In this paper we propose a systematic strategy for migrating crosscutting concerns in existing object-oriented systems to aspect-oriented programming solutions. The proposed strategy consists of four steps: mining, exploration, documentation and refactoring of crosscutting concerns. We discuss in detail a new approach to refactoring to aspect-oriented programming that is fully integrated with our strategy, and apply the whole strategy to an object-oriented system, namely the JHotDraw framework. Moreover, we present a method to semi-automatically perform the aspect-introducing refactorings based on identified crosscutting concern sorts which is supported by a prototype tool called sair . We perform an exploratory case study in which we apply this tool on the same object-oriented system and compare its results with the results of manual migration in order to assess the feasibility of automated aspect refactoring. Both the refactoring tool sair and the results of the manual migration are made available as open-source, the latter providing the largest aspect-introducing refactoring available to date. We report on our experiences with conducting both case studies and reflect on the success and challenges of the migration process.
Marius Marin, Arie van Deursen, Leon Moonen, Robin van der Rijst
Autom. Softw. Eng.3
2009 A Systematic Survey of Program Comprehension through Dynamic Analysis
abstract
Program comprehension is an important activity in software maintenance, as software must be sufficiently understood before it can be properly modified. The study of a program's execution, known as dynamic analysis, has become a common technique in this respect and has received substantial attention from the research community, particularly over the last decade. These efforts have resulted in a large research body of which currently there exists no comprehensive overview. This paper reports on a systematic literature survey aimed at the identification and structuring of research on program comprehension through dynamic analysis. From a research body consisting of 4,795 articles published in 14 relevant venues between July 1999 and June 2008 and the references therein, we have systematically selected 176 articles and characterized them in terms of four main facets: activity, target, method, and evaluation. The resulting overview offers insight in what constitutes the main contributions of the field, supports the task of identifying gaps and opportunities, and has motivated our discussion of several important research directions that merit additional consideration in the near future.
Bas Cornelissen, Andy Zaidman, Arie van Deursen, Leon Moonen, Rainer Koschke
IEEE Trans. Software Eng.4
2008 Assessing the value of coding standards: An empirical study
abstract
In spite of the widespread use of coding standards and tools enforcing their rules, there is little empirical evidence supporting the intuition that they prevent the introduction of faults in software. Not only can compliance with a set of rules having little impact on the number of faults be considered wasted effort, but it can actually result in an increase in faults, as any modification has a non-zero probability of introducing a fault or triggering a previously concealed one. Therefore, it is important to build a body of empirical knowledge, helping us understand which rules are worthwhile enforcing, and which ones should be ignored in the context of fault reduction. In this paper, we describe two approaches to quantify the relation between rule violations and actual faults, and present empirical data on this relation for the MISRA C 2004 standard on an industrial case study.
Cathal Boogerd, Leon Moonen
ICSM2
2008 An assessmentmethodology for trace reduction techniques
abstract
Program comprehension is an important concern in software maintenance because these tasks generally require a degree of knowledge of the system at hand. While the use of dynamic analysis in this process has become increasingly popular, the literature indicates that dealing with the huge amounts of dynamic information remains a formidable challenge.
Bas Cornelissen, Leon Moonen, Andy Zaidman
ICSM2
2008 2nd International Workshop on Advanced Software Development Tools and Techniques (WASDeTT): Tools for software maintenance, visualization, and reverse engineering
abstract
The objective of the 2nd international workshop on advanced software development tools and techniques (WASDeTT) is to provide interested researchers with a forum to share their tool building experiences and to explore how tools can be built more effectively and efficiently. This workshop specifically focuses on tools for software maintenance and comprehension and addresses issues such as tool-building in an industrial context, component-based tool building, and tool building in teams.
Holger M. Kienle, Leon Moonen, Michael W. Godfrey, Hausi A. Müller
ICSM2
2008 On the Use of Data Flow Analysis in Static Profiling
abstract
Static profiling is a technique that produces estimates of execution likelihoods or frequencies based on source code analysis only. It is frequently used in determining cost/benefit ratios for certain compiler optimizations. In previous work,we introduced a simple algorithm to compute execution likelihoods,based on a control flow graph and heuristic branch prediction. In this paper we examine the benefits of using more involved analysis techniques in such a static profiler. In particular, we explore the use of value range propagation to improve the accuracy of the estimates, and we investigate the differences in estimating execution likelihoods and frequencies.
Cathal Boogerd, Leon Moonen
SCAM2
2008 Execution trace analysis through massive sequence and circular bundle views
Bas Cornelissen, Andy Zaidman, Danny Holten, Leon Moonen, Arie van Deursen, Jarke J. van Wijk
J. Syst. Softw.4
2007 SoQueT: Query-Based Documentation of Crosscutting Concerns
abstract
Understanding crosscutting concerns is difficult because their underlying relations remain hidden in a class-based decomposition of a system. Based on an extensive investigation of crosscutting concerns in existing systems and literature, we identified a number of typical implementation idioms and relations that allow us to group such concerns around so- called "sorts". In this paper, we present SoQueT, a tool that uses sorts to support the consistent description and documentation of crosscutting relations using pre-defined, sort- specific query templates.
Marius Marin, Leon Moonen, Arie van Deursen
ICSE2
2007 Understanding Execution Traces Using Massive Sequence and Circular Bundle Views
abstract
The use of dynamic information to aid in software understanding is a common practice nowadays. One of the many approaches concerns the comprehension of execution traces. A major issue in this context is scalability: due to the vast amounts of information, it is a very difficult task to successfully find your way through such traces without getting lost. In this paper, we propose the use of a novel trace visualization method based on a massive sequence and circular bundle view, constructed with scalability in mind. By means of three usage scenarios that were conducted on three different software systems, we show how our approach, implemented in a tool called EXTRAVIS, is applicable to the areas of trace exploration, feature location, and feature comprehension.
Bas Cornelissen, Danny Holten, Andy Zaidman, Leon Moonen, Jarke J. van Wijk, Arie van Deursen
ICPC4
2007 Special issue on source code analysis and manipulation (SCAM 2006)
abstract
This special issue features extended versions of selected papers from the 6th International Workshop on Source Code Analysis and Manipulation (SCAM 2006) that took place in Philadelphia, PA, USA on 27-29 September 2006, co-located with the 22nd IEEE International Conference on Software Maintenance (ICSM 2006). The aim of SCAM to bring together researchers and practitioners working on theory, techniques and applications, concerning analysis and/or manipulation of the source code of computer systems. While much attention in the wider software engineering community is properly directed towards other aspects of systems' development and evolution, such as specification, design and requirements engineering, it is the source code that contains the only precise description of the behavior of the system. The analysis and manipulation of source code thus remain pressing concern. Whereas several conferences and workshops address the applications of source code analysis and manipulation, the aim of SCAM is to focus on the algorithms and tools themselves—what they can achieve; and how they can be improved, refined, and combined. Held for the first time with ICSM 2001 in Florence, SCAM has been a very successful event during the last six years, with a continuously increasing attendance and number of submission; such that it has been transformed, starting from 2007, in a working conference. Over the years, SCAM has grown it's unique format, consisting of a two-day event filled with technical sessions that have plenty of time allocated to discussion. Each technical session is structured around three short presentations (15 min) followed by 45 min open discussion that is initiated by controversial questions and issues raised by the authors of the papers presented. All attendees are encouraged to write their ideas and comments on transparencies, rather than merely contributing verbally. These transparencies are then collected and scanned for publication on the SCAM Web site*. In 2006, out of 48 submissions, 20 excellent papers were selected to be presented at the workshop, together with an inspiring keynote on “Slicing concurrent Java programs with Indus” by John Hatcliff from Kansas State University. The accepted papers covered a broad range of topics in source code analysis and manipulation, such as slicing, refactoring,transformations, abstract interpretation, static analysis and verification. After the workshop, four outstanding papers were selected by the program chairs were invited for publication in this Journal of Software Maintenance and Evolution: Research and Practice special issue. The authors of the selected papers were asked to prepare a significant extension on the workshop version of their papers. Each paper underwent a rigorous reviewing process, involving three independent reviewers and two rounds of revisions. Eventually, three papers were accepted for inclusion in this special issue, covering three different directions of source code analysis and manipulation: the study of identifier well formedness, the construction of accurate call graphs and support for automatically validating annotations in Java. The use of well-formed variable names is crucial for code quality, particularly for supporting code comprehension processes. In particular, identifiers should be coincise and consistent, in that it is possible to build a mapping from the domain of identifiers to the domain of concepts. In the paper ‘An Empirical Study of Rules for Well-Formed Identifiers’, Lawrie, Feild and Binkley proposed an approach for verifying the well formedness of identifiers. While mappings between identifiers and concepts require domain experts, this work empirically explored whether syntactic violations are indicative of concept mapping violations. The authors report an empirical study featuring 48 million lines of code, showing that syntactic violations occur in practice, that these violations largely match with concept-based violations, and that open source systems tend to exhibit a higher percentage of violations than proprietary systems, due to the continuously increasing number of contributors for the former and to the greater engineering discipline for the latter. Call graphs are used for a wide number of software engineering tasks, such as program understanding and testing, as well as in the optimization phase of compilers. In particular, application call graphs represent calling relationships among methods, which can happen through libraries. Such call graphs are particularly useful since they provide a higher level of abstraction than whole call graphs. Construction of call graphs is relatively straightforward for procedural programming languages, while dynamic dispatch makes such a task more difficult for object-oriented languages. In the paper ‘Automatic Construction of Accurate Application Call Graph with Library Call Abstraction for Java’, Zhang and Ryder proposed an approach for the automatic construction of accurate application call graphs for Java. They limit the presence of spurious calls by introducing a data reachability algorithm and outline the applicability of the approach to program slicing and dataflow testing. In the paper ‘AVal: an Extensible Attribute-Oriented Programming Validator for Java’, Noguera and Pawlak proposed a framework for defining and applying constraints on (collections of) annotations in a Java program. In particular, the authors presented a generic method for specifying annotations and rules over them using a consistent higher-level notation. This allows the user to express certain domain-specific aspects of the program at a higher level of abstraction which can automatically be enforced. In addition, by hiding the implementation details of these annotations and their constraints from the program code, the program's complexity is reduced, resulting in simpler, more readable programs that are easier to evolve. We hope that readers will enjoy this special issue and gain useful insights from the three papers presented. We would like to thank all the authors who submitted papers to the workshop and to this special issue, as well as the SCAM 2006 program committee and the external reviewers who helped making this special issue possible. Finally, we would like to thank the editorial board of the Journal of Software Maintenance and Evolution: Research and Practice and the publisher Wiley for providing us with the opportunity to devote an issue of this journal to SCAM 2006.
Massimiliano Di Penta, Leon Moonen
J. Softw. Maintenance Res. Pract.2
2007 Identifying Crosscutting Concerns Using Fan-In Analysis
abstract
Aspect mining is a reverse engineering process that aims at finding crosscutting concerns in existing systems. This article proposes an aspect mining approach based on determining methods that are called from many different places, and hence have a high fan-in , which can be seen as a symptom of crosscutting functionality. The approach is semiautomatic, and consists of three steps: metric calculation, method filtering, and call site analysis. Carrying out these steps is an interactive process supported by an Eclipse plug-in called FINT. Fan-in analysis has been applied to three open source Java systems, totaling around 200,000 lines of code. The most interesting concerns identified are discussed in detail, which includes several concerns not previously discussed in the aspect-oriented literature. The results show that a significant number of crosscutting concerns can be recognized using fan-in analysis, and each of the three steps can be supported by tools.
Marius Marin, Arie van Deursen, Leon Moonen
ACM Trans. Softw. Eng. Methodol.3
2006 Documenting software systems using types
Arie van Deursen, Leon Moonen
Sci. Comput. Program.2
2006 Applying and combining three different aspect Mining Techniques
Mariano Ceccato, Marius Marin, Kim Mens, Leon Moonen, Paolo Tonella, Tom Tourwé
Softw. Qual. J.4
2005 A Classification of Crosscutting Concerns
abstract
Refactoring software to apply aspect oriented solutions requires a clear understanding of what are the potential crosscutting concerns and which aspect solutions to replace them with. This process can benefit from the recognition of recurring generic concerns and their reusable aspect solutions. In this paper, we propose a classification of crosscutting concerns in sorts based on the analysis of various refactoring efforts. We discuss how sorts help concern understanding and refactoring, how they support the identification of crosscutting concerns, and how they can contribute to the evolution of aspect languages.
Marius Marin, Leon Moonen, Arie van Deursen
ICSM2
2004 Symphony: View-Driven Software Architecture Reconstruction
abstract
Authentic descriptions of a software architecture are required as a reliable foundation for any but trivial changes to a system. Far too often, architecture descriptions of existing systems are out of sync with the implementation. If they are, they must be reconstructed. There are many existing techniques for reconstructing individual architecture views, but no information about how to select views for reconstruction, or about process aspects of architecture reconstruction in general. In this paper we describe view-driven process for reconstructing software architecture that fills this gap. To describe Symphony, we present and compare different case studies, thus serving a secondary goal of sharing real-life reconstruction experience. The Symphony process incorporates the state of the practice, where reconstruction is problem-driven and uses a rich set of architecture views. Symphony provides a common framework for reporting reconstruction experiences and for comparing reconstruction approaches. Finally, it is a vehicle for exposing and demarcating research problems in software architecture reconstruction.
Arie van Deursen, Christine Hofmeister, Rainer Koschke, Leon Moonen, Claudio Riva
WICSA4
2003 Exploring Software Systems
abstract
Software evolution is required to keep a software system in sync with the ever-changing needs of the system's users and environment. An unfortunate side-effect of evolution is that it often causes the knowledge about a system to degrade, which in turn impedes further evolution. In the dissertation, we investigate techniques and tools that help remedy this situation by supporting the exploration of a software system and improving its legibility (Moonen, 2002). We examine the analogy with urban exploration and present innovative techniques for the extraction, abstraction, and presentation of information needed for understanding software.
Leon Moonen
ICSM1
2001 The ASF+SDF Meta-environment: A Component-Based Language Development Environment
Mark van den Brand, Arie van Deursen, Jan Heering, Hayco de Jong, Merijn de Jonge, Tobias Kuipers, Paul Klint, Leon Moonen, Pieter A. Olivier, Jeroen Scheerder, Jurgen J. Vinju, Eelco Visser, Joost Visser 0001
CC8
2001 An empirical study into COBOL type inferencing
Arie van Deursen, Leon Moonen
Sci. Comput. Program.2