Erik Arisholm

dblp:27/5178 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 26 · 9 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Application of deep learning models to generate rich, dynamic and production-like test data
abstract
Abstract Traditionally, software development teams in many industries have used copies of production databases or their masked, anonymized, or obfuscated versions for testing. However, privacy protection regulations, for example, the General Data Protection Regulation (GDPR), prohibit such practices. In such a situation, there is often a need to generate production-like test data, i.e., test data that is statistically representative of the production data and conforms to the domain’s constraints. In this paper, we address this need by presenting a novel approach for generating production-like test data using deep learning techniques and studying the practical effectiveness of our proposed approach in industrial settings. We frame the problem of generating production-like test data as a Language Modeling problem. We then propose a general solution for test data generation and a framework for evaluating and comparing language models based on training effectiveness and the representativeness and validity of the generated data. To evaluate the practical effectiveness of our solution, we apply it to a case study: the Norwegian National Population Registry (NPR). Within the context of NPR, we experiment with three of the most successful Deep Learning algorithms for Language Modeling, namely Recurrent Neural Networks (RNN), Variational Autoencoders, and Generative Adversarial Networks (GANs). We furthermore evaluate and compare the effectiveness of these algorithms quantitatively, using the proposed evaluation framework. The results from our case study show that our approach can generate highly complex data that is statistically representative of the production data and conforms to the business rules of the domain. Moreover, test data generated with RNN, which outperforms the other two algorithms, are syntactically and semantically valid in more than 97% of the cases and are highly representative of the real NPR data. The practical applicability of our approach is evident from the fact that our approach was fully deployed in a test environment at the NPR, generating on-the-fly and scalable amounts of production-like test data that are used in the integration testing between NPR and its data consumers.
Razieh Behjati, Erik Arisholm
Empir. Softw. Eng.3
2023 Enhancing Synthetic Test Data Generation with Language Models Using a More Expressive Domain-Specific Language
Razieh Behjati, Erik Arisholm
ICTSS3
2011 Industrial experiences with automated regression testing of a legacy database application
abstract
This paper presents a practical approach and tool (DART) for functional black-box regression testing of complex legacy database applications. Such applications are important to many organizations, but are often difficult to change and consequently prone to regression faults during maintenance. They also tend to be built without particular considerations for testability and can be hard to control and observe. We have therefore devised a practical solution for functional regression testing that captures the changes in database state (due to data manipulations) during the execution of a system under test. The differences in changed database states between consecutive executions of the system under test, on different system versions, can help identify potential regression faults. In order to make the regression test approach scalable for large, complex database applications, classification tree models are used to prioritize test cases. The test case prioritization can be applied to reduce test execution costs and analysis effort. We report on how DART was applied and evaluated on business critical batch jobs in a legacy database application in an industrial setting, namely the Norwegian Tax Accounting System (SOFIE) at the Norwegian Tax Department (NTD). DART has shown promising fault detection capabilities and cost-effectiveness and has contributed to identify many critical regression faults for the past eight releases of SOFIE.
Erik Rogstad, Lionel C. Briand, Ronny Dalberg, Marianne Rynning, Erik Arisholm
ICSM5
2010 Assessing UML design metrics for predicting fault-prone classes in a Java system
abstract
Identifying and fixing software problems before implementation are believed to be much cheaper than after implementation. Hence, it follows that predicting fault-proneness of software modules based on early software artifacts like software design is beneficial as it allows software engineers to perform early predictions to anticipate and avoid faults early enough. Taking this motivation into consideration, in this paper we evaluate the usefulness of UML design metrics to predict fault-proneness of Java classes. We use historical data of a significant industrial Java system to build and validate a UML-based prediction model. Based on the case study we have found that level of detail of messages and import coupling-both measured from sequence diagrams, are significant predictors of class fault-proneness. We also learn that the prediction model built exclusively using the UML design metrics demonstrates a better accuracy than the one built exclusively using code metrics.
Ariadi Nugroho, Michel R. V. Chaudron, Erik Arisholm
MSR3
2010 Understanding cost drivers of software evolution: a quantitative and qualitative investigation of change effort in two evolving software systems
Hans Christian Benestad, Bente Anda, Erik Arisholm
Empir. Softw. Eng.3
2010 A systematic and comprehensive investigation of methods to build and evaluate fault prediction models
Erik Arisholm, Lionel C. Briand, Eivind B. Johannessen
J. Syst. Softw.1
2010 Effects of Personality on Pair Programming
abstract
Personality tests in various guises are commonly used in recruitment and career counseling industries. Such tests have also been considered as instruments for predicting the job performance of software professionals both individually and in teams. However, research suggests that other human-related factors such as motivation, general mental ability, expertise, and task complexity also affect the performance in general. This paper reports on a study of the impact of the Big Five personality traits on the performance of pair programmers together with the impact of expertise and task complexity. The study involved 196 software professionals in three countries forming 98 pairs. The analysis consisted of a confirmatory part and an exploratory part. The results show that: (1) Our data do not confirm a meta-analysis-based model of the impact of certain personality traits on performance and (2) personality traits, in general, have modest predictive value on pair programming performance compared with expertise, task complexity, and country. We conclude that more effort should be spent on investigating other performance-related predictors such as expertise, and task complexity, as well as other promising predictors, such as programming skill and learning. We also conclude that effort should be spent on elaborating on the effects of personality on various measures of collaboration, which, in turn, may be used to predict and influence performance. Insights into such malleable, rather than static, factors may then be used to improve pair programming performance.
Jo Erskine Hannay, Erik Arisholm, Harald Engvik, Dag I. K. Sjøberg
IEEE Trans. Software Eng.2
2009 Are We More Productive Now? Analyzing Change Tasks to Assess Productivity Trends during Software Evolution
Hans Christian Benestad, Bente Anda, Erik Arisholm
ENASE3
2009 Improving the effectiveness of root cause analysis in post mortem analysis: A controlled experiment
Finn Olav Bjørnson, Alf Inge Wang, Erik Arisholm
Inf. Softw. Technol.3
2009 The effectiveness of pair programming: A meta-analysis
Jo Erskine Hannay, Tore Dybå, Erik Arisholm, Dag I. K. Sjøberg
Inf. Softw. Technol.3
2009 The effect of task order on the maintainability of object-oriented software
Alf Inge Wang, Erik Arisholm
Inf. Softw. Technol.2
2009 Understanding software maintenance and evolution by analyzing individual changes: a literature review
abstract
Abstract Understanding, managing and reducing costs and risks inherent in change are key challenges of software maintenance and evolution, addressed in empirical studies with many different research approaches. Change‐based studies analyze data that describes the individual changes made to software systems. This approach can be effective in order to discover cost and risk factors that are hidden at more aggregated levels. However, it is not trivial to derive appropriate measures of individual changes for specific measurement goals. The purpose of this review is to improve change‐based studies by (1) summarizing how attributes of changes have been measured to reach specific study goals and (2) describing current achievements and challenges, leading to a guide for future change‐based studies. Thirty‐four papers conformed to the inclusion criteria. Forty‐three attributes of changes were identified, and classified according to a conceptual model developed for the purpose of this classification. The goal of each study was to either characterize the evolution process, to assess causal factors of cost and risk, or to predict costs and risks. Effective accumulation of knowledge across change‐based studies requires precise definitions of attributes and measures of change. We recommend that new change‐based studies base such definitions on the proposed conceptual model. Copyright © 2009 John Wiley & Sons, Ltd.
Hans Christian Benestad, Bente Anda, Erik Arisholm
J. Softw. Maintenance Res. Pract.3
2008 A Realistic Empirical Evaluation of the Costs and Benefits of UML in Software Maintenance
abstract
The Unified Modeling Language (UML) is the de facto standard for object-oriented software analysis and design modeling. However, few empirical studies exist that investigate the costs and evaluate the benefits of using UML in realistic contexts. Such studies are needed so that the software industry can make informed decisions regarding the extent to which they should adopt UML in their development practices. This is the first controlled experiment that investigates the costs of maintaining and the benefits of using UML documentation during the maintenance and evolution of a real, non-trivial system, using professional developers as subjects, working with a state-of-the-art UML tool during an extended period of time. The subjects in the control group had no UML documentation. In this experiment, the subjects in the UML group had on average a practically and statistically significant 54% increase in the functional correctness of changes (p=0.03), and an insignificant 7% overall improvement in design quality (p=0.22) - though a much larger improvement was observed on the first change task (56%) - at the expense of an insignificant 14% increase in development time caused by the overhead of updating the UML documentation (p=0.35).
Wojciech J. Dzidek, Erik Arisholm, Lionel C. Briand
IEEE Trans. Software Eng.2
2007 Educational Approach to an Experiment in a Software Architecture Course
abstract
This paper reports experiences from an experiment in a software architecture course where the focus was both on giving students valuable education as well as getting important empirical results. The paper describes how the experiment was integrated in the course, and presents an evaluation of the experiment from an educational point of view. Further, the paper reflects on the costs and the benefits of carrying out an experiment in the context of a software architecture course for the involved stakeholders namely the researchers, the students, and the instructors. We also describe some guidelines for planning and executing experiments as a part of a software engineering course.
Alf Inge Wang, Erik Arisholm, Letizia Jaccheri
CSEE&T2
2007 Data Mining Techniques for Building Fault-proneness Models in Telecom Java Software
abstract
This paper describes a study performed in an industrial setting that attempts to build predictive models to identify parts of a Java system with a high fault probability. The system under consideration is constantly evolving as several releases a year are shipped to customers. Developers usually have limited resources for their testing and inspections and would like to be able to devote extra resources to faulty system parts. The main research focus of this paper is two-fold: (1) use and compare many data mining and machine learning techniques to build fault-proneness models based mostly on source code measures and change/fault history data, and (2) demonstrate that the usual classification evaluation criteria based on confusion matrices may not be fully appropriate to compare and evaluate models.
Erik Arisholm, Lionel C. Briand, Magnus Fuglerud
ISSRE1
2007 Evaluating Pair Programming with Respect to System Complexity and Programmer Expertise
abstract
A total of 295 junior, intermediate, and senior professional Java consultants (99 individuals and 98 pairs) from 29 international consultancy companies in Norway, Sweden, and the UK were hired for one day to participate in a controlled experiment on pair programming. The subjects used professional Java tools to perform several change tasks on two alternative Java systems with different degrees of complexity. The results of this experiment do not support the hypotheses that pair programming in general reduces the time required to solve the tasks correctly or increases the proportion of correct solutions. On the other hand, there is a significant 84 percent increase in effort to perform the tasks correctly. However, on the more complex system, the pair programmers had a 48 percent increase in the proportion of correct solutions but no significant differences in the time taken to solve the tasks correctly. For the simpler system, there was a 20 percent decrease in time taken but no significant differences in correctness. However, the moderating effect of system complexity depends on the programmer expertise of the subjects. The observed benefits of pair programming in terms of correctness on the complex system apply mainly to juniors, whereas the reductions in duration to perform the tasks correctly on the simple system apply mainly to intermediates and seniors. It is possible that the benefits of pair programming will exceed the results obtained in this experiment for larger, more complex tasks and if the pair programmers have a chance to work together over a longer period of time
Erik Arisholm, Hans Gallis, Tore Dybå, Dag I. K. Sjøberg
IEEE Trans. Software Eng.1
2006 Assessing Software Product Maintainability Based on Class-Level Structural Measures
Hans Christian Benestad, Bente Anda, Erik Arisholm
PROFES3
2006 Empirical assessment of the impact of structural properties on the changeability of object-oriented software
Erik Arisholm
Inf. Softw. Technol.1
2006 The Impact of UML Documentation on Software Maintenance: An Experimental Evaluation
abstract
The Unified Modeling Language (UML) is becoming the de facto standard for software analysis and design modeling. However, there is still significant resistance to model-driven development in many software organizations because it is perceived to be expensive and not necessarily cost-effective. Hence, it is important to investigate the benefits obtained from modeling. As a first step in this direction, this paper reports on controlled experiments, spanning two locations, that investigate the impact of UML documentation on software maintenance. Results show that, for complex tasks and past a certain learning curve, the availability of UML documentation may result in significant improvements in the functional correctness of changes as well as the quality of their design. However, there does not seem to be any saving of time. For simpler tasks, the time needed to update the UML documentation may be substantial compared with the potential benefits, thus motivating the need for UML tools with better support for software maintenance
Erik Arisholm, Lionel C. Briand, Siw Elisabeth Hove, Yvan Labiche
IEEE Trans. Software Eng.1
2005 Collecting Feedback during Software Engineering Experiments
Amela Karahasanovic, Bente Anda, Erik Arisholm, Siw Elisabeth Hove, Magne Jørgensen, Dag I. K. Sjøberg, Ray Welland
Empir. Softw. Eng.3
2004 A Controlled Experiment Comparing the Maintainability of Programs Designed with and without Design Patterns-A Replication in a Real Programming Environment
Marek Vokác, Walter F. Tichy, Dag I. K. Sjøberg, Erik Arisholm, Magne Aldrin
Empir. Softw. Eng.4
2004 Dynamic Coupling Measurement for Object-Oriented Software
abstract
The relationships between coupling and external quality factors of object-oriented software have been studied extensively for the past few years. For example, several studies have identified clear empirical relationships between class-level coupling and class fault-proneness. A common way to define and measure coupling is through structural properties and static code analysis. However, because of polymorphism, dynamic binding, and the common presence of unused ("dead") code in commercial software, the resulting coupling measures are imprecise as they do not perfectly reflect the actual coupling taking place among classes at runtime. For example, when using static analysis to measure coupling, it is difficult and sometimes impossible to determine what actual methods can be invoked from a client class if those methods are overridden in the subclasses of the server classes. Coupling measurement has traditionally been performed using static code analysis, because most of the existing work was done on nonobject oriented code and because dynamic code analysis is more expensive and complex to perform. For modern software systems, however, this focus on static analysis can be problematic because although dynamic binding existed before the advent of object-orientation, its usage has increased significantly in the last decade. We describe how coupling can be defined and precisely measured based on dynamic analysis of systems. We refer to this type of coupling as dynamic coupling. An empirical evaluation of the proposed dynamic coupling measures is reported in which we study the relationship of these measures with the change proneness of classes. Data from maintenance releases of a large Java system are used for this purpose. Preliminary results suggest that some dynamic coupling measures are significant indicators of change proneness and that they complement existing coupling measures based on static analysis.
Erik Arisholm, Lionel C. Briand, Audun Føyen
IEEE Trans. Software Eng.1
2004 Evaluating the Effect of a Delegated versus Centralized Control Style on the Maintainability of Object-Oriented Software
abstract
A fundamental question in object-oriented design is how to design maintainable software. According to expert opinion, a delegated control style, typically a result of responsibility-driven design, represents object-oriented design at its best, whereas a centralized control style is reminiscent of a procedural solution, or a "bad" object-oriented design. We present a controlled experiment that investigates these claims empirically. A total of 99 junior, intermediate, and senior professional consultants from several international consultancy companies were hired for one day to participate in the experiment. To compare differences between (categories of) professionals and students, 59 students also participated. The subjects used professional Java tools to perform several change tasks on two alternative Java designs that had a centralized and delegated control style, respectively. The results show that the most skilled developers, in particular, the senior consultants, require less time to maintain software with a delegated control style than with a centralized control style. However, more novice developers, in particular, the undergraduate students and junior consultants, have serious problems understanding a delegated control style, and perform far better with a centralized control style. Thus, the maintainability of object-oriented software depends, to a large extent, on the skill of the developers who are going to maintain it. These results may have serious implications for object-oriented development in an industrial context: having senior consultants design object-oriented systems may eventually pose difficulties unless they make an effort to keep the designs simple, as the cognitive complexity of "expert" designs might be unmanageable for less skilled maintainers.
Erik Arisholm, Dag I. K. Sjøberg
IEEE Trans. Software Eng.1
2001 Program Understanding Behavior during Estimation of Enhancement Effort on Small Java Programs
Lars Bratthall, Erik Arisholm, Magne Jørgensen
PROFES2
2001 Assessing the Changeability of two Object-Oriented Design Alternatives - a Controlled Experiment
Erik Arisholm, Dag I. K. Sjøberg, Magne Jørgensen
Empir. Softw. Eng.1
2000 Towards a framework for empirical assessment of changeability decay
Erik Arisholm, Dag I. K. Sjøberg
J. Syst. Softw.1
1999 Empirical Studies of Object-Oriented Artifacts, Methods, and Processes: State of the Art and Future Directions
Lionel C. Briand, Erik Arisholm, Steve Counsell, Frank Houdek, Pascale Thévenod-Fosse
Empir. Softw. Eng.2