Dag I. K. Sjøberg

dblp:32/889 · DBLP profile ↗
← Back
61ranked-venue papers
11as first author
8since 2021 · last 2025
0000-0002-4941-7240ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 57 · 10 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Exploring the Interplay Between Learning Programming and Probability Simultaneously: A Case Study of a 9th-Grade Student
Sindre Mathias Strømnes Nordvoll, Ragnhild Kobro Runde, Quintin I. Cutts, Dag I. K. Sjøberg
ICER (1)4
2023 Improving the Reporting of Threats to Construct Validity
abstract
Background: Construct validity concerns the use of indicators to measure a concept that is not directly measurable. Aim: This study intends to identify, categorize, assess and quantify discussions of threats to construct validity in empirical software engineering literature and use the findings to suggest ways to improve the reporting of construct validity issues. Method: We analyzed 83 articles that report human-centric experiments published in five top-tier software engineering journals from 2015 to 2019. The articles’ text concerning threats to construct validity was divided into segments (the unit of analysis) based on predefined categories. The segments were then evaluated regarding whether they clearly discussed a threat and a construct. Results: Three-fifths of the segments were associated with topics not related to construct validity. Two-thirds of the articles discussed construct validity without using the definition of construct validity given in the article. The threats were clearly described in more than four-fifths of the segments, but the construct in question was clearly described in only two-thirds of the segments. The construct was unclear when the discussion was not related to construct validity but to other types of validity. Conclusions: The results show potential for improving the understanding of construct validity in software engineering. Recommendations addressing the identified weaknesses are given to improve the awareness and reporting of CV.
Dag I. K. Sjøberg, Gunnar R. Bergersen
EASE1
2023 On the roles of software testers: An exploratory study
abstract
Software development organizations need testers with high skill levels in a broad range of technical areas and application domains. Accordingly, we need a better understanding of how testers meet such skill demands in the practice of their role. This work aims to deepen the understanding of the typical tester role. We performed a thematic analysis of 19 in-depth, semi-structured interviews with software testers working in various industries. To investigate employers’ views on such roles, we conducted a thematic analysis of 400 job ads. From the interviews, we identified five subroles of software testers: domain-specific tester, test automation specialist, test infrastructure specialist, user experience tester, and test manager. Most of the practitioners preferred to develop skills and act in one subrole. In contrast, most of the job ads requested that testers act in multiple subroles. Our findings provide a deeper understanding of the tester role, which may guide testers in their acquisition of skills and employers in the recruiting of testers.
Raluca Florea, Viktoria Stray, Dag I. K. Sjøberg
J. Syst. Softw.3
2023 Construct Validity in Software Engineering
abstract
Empirical research aims to establish generalizable claims from data. Such claims may involve concepts that must be measured indirectly by using indicators. Construct validity is concerned with whether one can justifiably make claims at the conceptual level that are supported by results at the operational level. We report a quantitative analysis of the awareness of construct validity in the software engineering literature between 2000 and 2019 and a qualitative review of 83 articles about human-centric experiments published in five high-quality journals between 2015 and 2019. Over the two decades, the appearance in the literature of the term construct validity increased sevenfold. Some of the reviewed articles we reviewed employed various ways to ensure that the indicators span the concept in an unbiased manner. We also found articles that reuse formerly validated constructs. However, the articles disagree about how to define construct validity. Several interpret construct validity excessively by including threats to internal, external, or statistical conclusion validity. A few articles also include fundamental challenges of a study, such as cheating and misunderstanding of experiment material. The diversity of topics included as threats to construct validity calls for a more minimalist approach. Based on the review, we propose seven guidelines to improve how construct validity is handled and reported in software engineering.
Dag I. K. Sjøberg, Gunnar R. Bergersen
IEEE Trans. Software Eng.1
2022 Benefits management and Information Technology work distribution
abstract
Abstract Organisations spend much money on Information Technology (IT) development and maintenance activities with the intention that these activities will create results that enable benefits for the organisations. This paper seeks to understand potential associations between IT development and maintenance activities and the adoption of benefits management practices to realise value for the organization. The aim is also to uncover potential differences between public and private organisations. We surveyed 86 Norwegian public and private organisations, including data collected in similar surveys every five years since 1993. For the period between 1998 and 2018, we observe a stable pattern of IT work distribution. We found that organisations that managed benefits put more effort into advancing functionality for the end‐users than other organisations, and they realised more benefits. This advantage was particularly true for organisations that managed benefits beyond the early stages of the development lifecycle. Private organisations both managed and realised benefits to a larger extent than public organisations. Our findings can enable organisations to be evidence‐based when choosing management practices to achieve a higher return on investments in IT development and maintenance activities.
Knut Kjetil Holgeid, John Krogstie, Patrick Mikalef, Eirik E. Saur, Dag I. K. Sjøberg
IET Softw.5
2021 Reducing Incidents in Microservices by Repaying Architectural Technical Debt
abstract
Architectural technical debt (ATD) may create a substantial extra effort in software development, which is called interest. There is little evidence about whether repaying ATD in microservices reduces such interest. Objectives: We wanted to conduct a first study on investigating the effect of removing ATD on the occurrence of incidents in a microservices architecture. Method: We conducted a quantitative and qualitative case study of a project with approximately 1000 microservices in a large, international financing services company. We measured and compared the number of software incidents of different categories before and after repaying ATD. Results: The total number of incidents was reduced by 84%, and the numbers of critical- and high-priority incidents were both reduced by approximately 90% after the architectural refactoring. The number of incidents in the architecture with the ATD was mainly constant over time, but we observed a slight increase of low priority incidents related to inaccessibility and the environment in the architecture without the ATD. Conclusion: This study shows evidence that refactoring ATDs, such as lack of communication standards, poor management of dead-letter queues, and the use of inadequate technologies in microservices, reduces the number of critical- and high-priority incidents and, thus, part of its interest, although some low priority incidents may increase.
Saulo S. de Toledo, Antonio Martini 0001, Dag I. K. Sjøberg, Agata Przybyszewska, Johannes Skov Frandsen
SEAA3
2021 Benefits management in software development: A systematic review of empirical studies
abstract
Abstract Considerable resources are wasted on software projects delivering less than the planned benefits. Herein, the objective is to synthesize empirical evidence of the adoption and impact of benefits management (BM) in software development, and to suggest directions for future research. A systematic review of the literature is performed and identified 4836 scientific papers of which the authors found 47 to include relevant research. While most organizations identify and structure benefits at the outset of a project, fewer organizations report implementing BM as a continuous process throughout the project lifecycle. Empirical evidence gives support for positive impact on project outcome from the following BM practices: identifying and structuring benefits, planning benefits realization, BM during project execution, benefits evaluation and the practice of having people responsible for benefits realization. The authors suggest four research directions to understand (1) why BM practices sometimes not are adopted, (2) BM in relation to other management practices, (3) BM in agile software development and (4) BM in the context of organizations' value creation logics.
Knut Kjetil Holgeid, Magne Jørgensen, Dag I. K. Sjøberg, John Krogstie
IET Softw.3
2021 Identifying architectural technical debt, principal, and interest in microservices: A multiple-case study
abstract
Using a microservices architecture is a popular strategy for software organizations to deliver value to their customers fast and continuously. However, scientific knowledge on how to manage architectural debt in microservices is scarce. In the context of microservices applications, this paper aims to identify architectural technical debts (ATDs), their costs, and their most common solutions. We conducted an exploratory multiple case study by conducting 25 interviews with practitioners working with microservices in seven large companies. We found 16 ATD issues, their negative impact (interest), and common solutions to repay each debt together with the related costs (principal). Two examples of critical ATD issues found were the use of shared databases that, if not properly planned, leads to potential breaks on services every time the database schema changes and bad API designs, which leads to coupling among teams. We identified ATDs occurring in different domains and stages of development and created a map of the relationships among those debts. The findings may guide organizations in developing microservices systems that better manage and avoid architectural debts.
Saulo S. de Toledo, Antonio Martini 0001, Dag I. K. Sjøberg
J. Syst. Softw.3
2019 Architectural technical debt in microservices: a case study in a large company
abstract
Introduction: Software companies aim to achieve continuous delivery to constantly provide value to their customers. A popular strategy is to use microservices architecture. However, such an architecture is also subject to debt, which hinders the continuous delivery process and thus negatively affects the software released to the customers. Objectives: The aim of this study is to identify issues, solutions and risks related to Architecture Technical Debt in microservices. Method: We conducted an exploratory case study of a real life project with about 1000 services in a large, international company. Through qualitative analysis of documents and interviews, we investigated Architecture Technical Debt in the communication layer of a system with microservices architecture. Results: Our main contributions are a list of Architecture Technical Debt issues specific for the communication layer in a system with microservices architecture, as well as their associated negative impact (interest), a solution to repay the debt and the its cost (principal). Among the found Architecture Technical Debt issues were the existence of business logic in the communication layer and a high amount of point-to-point connections between services. The studied solution consists of the implementation of different canonical models specific to different domains, the removal of business logic from the communication layer, and migration from services to use the communication layer correctly. We also contributed with a list of possible risks that can affect the payment of the debt, as lack of funding and inadequate prioritization. Conclusion: We found issues, solutions and possible risks that are specific for microservices architectures not yet encountered in the current literature. Our results may be useful for practitioners that want to avoid or repay Technical Debt in their microservices architecture.
Saulo S. de Toledo, Antonio Martini 0001, Agata Przybyszewska, Dag I. K. Sjøberg
TechDebt@ICSE4
2018 An empirical study of WIP in kanban teams
abstract
Background: Limiting the amount of Work-In-Progress (WIP) is considered a fundamental principle in Kanban software development. However, no published studies from real cases exist that indicate what an optimal WIP limit should be. Aims: The primary aim is to study the effect of WIP on the performance of a Kanban team. The secondary aim is to illustrate methodological challenges when attempting to identify an optimal or appropriate WIP limit. Method: A quantitative case study was conducted in a software company that provided information about more than 8,000 work items developed over four years by five teams. Relationships between WIP, lead time and productivity were analyzed. Results: WIP correlates with lead time; that is, lower WIP indicates shorter lead times, which is consistent with claims in the literature. However, WIP also correlates with productivity, which is inconsistent with the claim in the literature that a low WIP (still above a certain threshold) will improve productivity. The collected data set did not include sufficient information to measure aspects of quality. There are several threats to the way productivity was measured. Conclusions: Indicating an optimal WIP limit is difficult in the studied company because a changing WIP gives contrasting results on different team performance variables. Because the effect of WIP has not been quantitatively examined before, this study clearly needs to be replicated in other contexts. In addition, studies that include other team performance variables, such as various aspects of quality, are requested. The methodological challenges illustrated in this paper need to be addressed.
Dag I. K. Sjøberg
ESEM1
2018 Teamwork Quality and Team Performance: Exploring Differences Between Small and Large Agile Projects
Yngve Lindsjørn, Gunnar R. Bergersen, Torgeir Dingsøyr, Dag I. K. Sjøberg
XP4
2018 Four commentaries on the use of students and professionals in empirical software engineering experiments
Robert Feldt, Thomas Zimmermann 0001, Gunnar R. Bergersen, Davide Falessi, Andreas Jedlitschka, Natalia Juristo Juzgado, Jürgen Münch, Markku Oivo, Per Runeson, Martin J. Shepperd, Dag I. K. Sjøberg, Burak Turhan
Empir. Softw. Eng.11
2016 The Relationship Between Software Process, Context and Outcome
Dag I. K. Sjøberg
PROFES1
2016 Incorrect results in software engineering experiments: How to improve research practices
Magne Jørgensen, Tore Dybå, Knut Liestøl, Dag I. K. Sjøberg
J. Syst. Softw.4
2016 Teamwork quality and project success in software development: A survey of agile development teams
abstract
Small, self-directed teams are central in agile development. This article investigates the effect of teamwork quality on team performance, learning and work satisfaction in agile software teams, and whether this effect differs from that of traditional software teams. A survey was administered to 477 respondents from 71 agile software teams in 26 companies and analyzed using structural equation modeling. A positive effect of teamwork quality on team performance was found when team members and team leaders rated team performance. In contrast, a negligible effect was found when product owners rated team performance. The effect of teamwork quality on team members´ learning and work satisfaction was strongly positive, but was only rated by the team members. Despite claims of the importance of teamwork in agile teams, this study did not find teamwork quality to be higher than in a similar survey on traditional teams. The effect of teamwork quality on team performance was only marginally greater for the agile teams than for the traditional teams.
Yngve Lindsjørn, Dag I. K. Sjøberg, Torgeir Dingsøyr, Gunnar R. Bergersen, Tore Dybå
J. Syst. Softw.2
2016 The daily stand-up meeting: A grounded theory study
Viktoria Stray, Dag I. K. Sjøberg, Tore Dybå
J. Syst. Softw.2
2014 Construction and Validation of an Instrument for Measuring Programming Skill
abstract
Skilled workers are crucial to the success of software development. The current practice in research and industry for assessing programming skills is mostly to use proxy variables of skill, such as education, experience, and multiple-choice knowledge tests. There is as yet no valid and efficient way to measure programming skill. The aim of this research is to develop a valid instrument that measures programming skill by inferring skill directly from the performance on programming tasks. Over two days, 65 professional developers from eight countries solved 19 Java programming tasks. Based on the developers' performance, the Rasch measurement model was used to construct the instrument. The instrument was found to have satisfactory (internal) psychometric properties and correlated with external variables in compliance with theoretical expectations. Such an instrument has many implications for practice, for example, in job recruitment and project allocation.
Gunnar R. Bergersen, Dag I. K. Sjøberg, Tore Dybå
IEEE Trans. Software Eng.2
2013 Obstacles to Efficient Daily Meetings in Agile Development Projects: A Case Study
abstract
Context: Most of the software organizations that use agile methods organize daily team meetings. Aim: Our aim was to understand how daily meetings are conducted and identify obstacles that reduce their efficiency. Method: We observed 56 daily meetings and conducted 21 interviews in three different teams in two countries. We used the repertory grid technique in the interviews and to analyze the results. Results: We identified thirteen obstacles. The four most prominent ones were: (1) The daily meetings lasted too long (on average, 22 minutes instead of the scheduled 15 minutes). (2) In the meetings that were not self-organized, team members reported to the Scrum Master instead of sharing information among all team members. (3) The interruption caused by daily meetings required substantially more time than the actual meeting time due to overhead before and after the meetings. (4) Several team members had negative attitudes towards the daily meetings. Conclusion: Organizers of daily meetings should evaluate whether the obstacles we have identified are present in their organization and consider our suggestions to remove or reduce these obstacles.
Viktoria Stray, Yngve Lindsjørn, Dag I. K. Sjøberg
ESEM3
2013 Trends in the Quality of Human-Centric Software Engineering Experiments-A Quasi-Experiment
abstract
Context: Several text books and papers published between 2000 and 2002 have attempted to introduce experimental design and statistical methods to software engineers undertaking empirical studies. Objective: This paper investigates whether there has been an increase in the quality of human-centric experimental and quasi-experimental journal papers over the time period 1993 to 2010. Method: Seventy experimental and quasi-experimental papers published in four general software engineering journals in the years 1992-2002 and 2006-2010 were each assessed for quality by three empirical software engineering researchers using two quality assessment methods (a questionnaire-based method and a subjective overall assessment). Regression analysis was used to assess the relationship between paper quality and the year of publication, publication date group (before 2003 and after 2005), source journal, average coauthor experience, citation of statistical text books and papers, and paper length. The results were validated both by removing papers for which the quality score appeared unreliable and using an alternative quality measure. Results: Paper quality was significantly associated with year, citing general statistical texts, and paper length (p <; 0.05). Paper length did not reach significance when quality was measured using an overall subjective assessment. Conclusions: The quality of experimental and quasi-experimental software engineering papers appears to have improved gradually since 1993.
Barbara A. Kitchenham, Dag I. K. Sjøberg, Tore Dybå, Pearl Brereton, David Budgen, Martin Höst, Per Runeson
IEEE Trans. Software Eng.2
2013 Quantifying the Effect of Code Smells on Maintenance Effort
abstract
Context: Code smells are assumed to indicate bad design that leads to less maintainable code. However, this assumption has not been investigated in controlled studies with professional software developers. Aim: This paper investigates the relationship between code smells and maintenance effort. Method: Six developers were hired to perform three maintenance tasks each on four functionally equivalent Java systems originally implemented by different companies. Each developer spent three to four weeks. In total, they modified 298 Java files in the four systems. An Eclipse IDE plug-in measured the exact amount of time a developer spent maintaining each file. Regression analysis was used to explain the effort using file properties, including the number of smells. Result: None of the 12 investigated smells was significantly associated with increased effort after we adjusted for file size and the number of changes; Refused Bequest was significantly associated with decreased effort. File size and the number of changes explained almost all of the modeled variation in effort. Conclusion: The effects of the 12 smells on maintenance effort were limited. To reduce maintenance effort, a focus on reducing code size and the work practices that limit the number of changes may be more beneficial than refactoring code smells.
Dag I. K. Sjøberg, Aiko Fallas Yamashita, Bente Anda, Audris Mockus, Tore Dybå
IEEE Trans. Software Eng.1
2012 Evaluating methods and technologies in software engineering with Respect to Developers' skill level
abstract
Background: It is trivial that the usefulness of a technology depends on the skill of the user. Several studies have reported an interaction between skill levels and different technologies, but the effect of skill is, for the most part, ignored in empirical, human-centric studies in software engineering. Aim: This paper investigates the usefulness of a technology as a function of skill. Method: An experiment that used students as subjects found recursive implementations to be easier to debug correctly than iterative implementations. We replicated the experiment by hiring 65 professional developers from nine companies in eight countries. In addition to the debugging tasks, performance on 17 other programming tasks was collected and analyzed using a measurement model that expressed the effect of treatment as a function of skill. Results: The hypotheses of the original study were confirmed only for the low-skilled subjects in our replication. Conversely, the high-skilled subjects correctly debugged the iterative implementations faster than the recursive ones, while the difference between correct and incorrect solutions for both treatments was negligible. We also found that the effect of skill (odds ratio = 9.4) was much larger than the effect of the treatment (odds ratio = 1.5). Conclusions: Claiming that a technology is better than another is problematic without taking skill levels into account. Better ways to assess skills as an integral part of technology evaluation are required.
Gunnar R. Bergersen, Dag I. K. Sjøberg
EASE2
2012 What works for whom, where, when, and why?: on the role of context in empirical software engineering
abstract
Context is a central concept in empirical software engineering. It is one of the distinctive features of the discipline and it is an in-dispensable part of software practice. It is likely responsible for one of the most challenging methodological and theoretical problems: study-to-study variation in research findings. Still, empirical software engineering research is mostly concerned with attempts to identify universal relationships that are independent of how work settings and other contexts interact with the processes important to software practice. The aim of this paper is to provide an overview of how context affects empirical research and how empirical software engineering research can be better 'contextualized' in order to provide a better understanding of what works for whom, where, when, and why. We exemplify the importance of context with examples from recent systematic reviews and offer recommendations on the way forward.
Tore Dybå, Dag I. K. Sjøberg, Daniela S. Cruzes
ESEM2
2012 Questioning software maintenance metrics: a comparative case study
abstract
Context: Many metrics are used in software engineering research as surrogates for maintainability of software systems. Aim: Our aim was to investigate whether such metrics are consistent among themselves and the extent to which they predict maintenance effort at the entire system level. Method: The Maintainability Index, a set of structural measures, two code smells (Feature Envy and God Class) and size were applied to a set of four functionally equivalent systems. The metrics were compared with each other and with the outcome of a study in which six developers were hired to perform three maintenance tasks on the same systems. Results: The metrics were not mutually consistent. Only system size and low cohesion were strongly associated with increased maintenance effort. Conclusion: Apart from size, surrogate maintainability measures may not reflect future maintenance effort. Surrogates need to be evaluated in the contexts for which they will be used. While traditional metrics are used to identify problematic areas in the code, the improvements of the worst areas may, inadvertently, lead to more problems for the entire system. Our results suggest that local improvements should be accompanied by an evaluation at the system level.
Dag I. K. Sjøberg, Bente Anda, Audris Mockus
ESEM1
2012 Three empirical studies on the agreement of reviewers about the quality of software engineering experiments
abstract
During systematic literature reviews it is necessary to assess the quality of empirical papers. Current guidelines suggest that two researchers should independently apply a quality checklist and any disagreements must be resolved. However, there is little empirical evidence concerning the effectiveness of these guidelines. This paper investigates the three techniques that can be used to improve the reliability (i.e. the consensus among reviewers) of quality assessments, specifically, the number of reviewers, the use of a set of evaluation criteria and consultation among reviewers. We undertook a series of studies to investigate these factors. Two studies involved four research papers and eight reviewers using a quality checklist with nine questions. The first study was based on individual assessments, the second study on joint assessments with a period of inter-rater discussion. A third more formal randomised block experiment involved 48 reviewers assessing two of the papers used previously in teams of one, two and three persons to assess the impact of discussion among teams of different size using the evaluations of the “teams” of one person as a control. For the first two studies, the inter-rater reliability was poor for individual assessments, but better for joint evaluations. However, the results of the third study contradicted the results of Study 2. Inter-rater reliability was poor for all groups but worse for teams of two or three than for individuals. When performing quality assessments for systematic literature reviews, we recommend using three independent reviewers and adopting the median assessment. A quality checklist seems useful but it is difficult to ensure that the checklist is both appropriate and understood by reviewers. Furthermore, future experiments should ensure participants are given more time to understand the quality checklist and to evaluate the research papers.
Barbara A. Kitchenham, Dag I. K. Sjøberg, Tore Dybå, Dietmar Pfahl, Pearl Brereton, David Budgen, Martin Höst, Per Runeson
Inf. Softw. Technol.2
2011 Inferring Skill from Tests of Programming Performance: Combining Time and Quality
abstract
The skills of software developers are important to the success of software projects. Also, when studying the general effect of a tool or method, it is important to control for individual differences in skill. However, the way skill is assessed is often ad hoc, or based on unvalidated methods. According to established test theory, validated tests of skill should infer skill levels from well-defined performance measures on multiple, small, representative tasks. In this respect, we show how time and quality, which are often analyzed separately, can be combined as task performance and subsequently be aggregated as an approximation of skill. Our results show significant positive correlations between our proposed measures of skill and other variables, such as seniority, lines of code written, and self-evaluated expertise. The method for combining time and quality is a promising first step to measuring programming skill in both industry and research settings.
Gunnar R. Bergersen, Jo Erskine Hannay, Dag I. K. Sjøberg, Tore Dybå, Amela Karahasanovic
ESEM3
2010 Can we evaluate the quality of software engineering experiments?
abstract
Context: The authors wanted to assess whether the quality of published human-centric software engineering experiments was improving. This required a reliable means of assessing the quality of such experiments. Aims: The aims of the study were to confirm the usability of a quality evaluation checklist, determine how many reviewers were needed per paper that reports an experiment, and specify an appropriate process for evaluating quality. Method: With eight reviewers and four papers describing human-centric software engineering experiments, we used a quality checklist with nine questions. We conducted the study in two parts: the first was based on individual assessments and the second on collaborative evaluations. Results: The inter-rater reliability was poor for individual assessments but much better for joint evaluations. Four reviewers working in two pairs with discussion were more reliable than eight reviewers with no discussion. The sum of the nine criteria was more reliable than individual questions or a simple overall assessment. Conclusions: If quality evaluation is critical, more than two reviewers are required and a round of discussion is necessary. We advise using quality criteria and basing the final assessment on the sum of the aggregated criteria. The restricted number of papers used and the relatively extensive expertise of the reviewers limit our results. In addition, the results of the second part of the study could have been affected by removing a time restriction on the review as well as the consultation process.
Barbara A. Kitchenham, Dag I. K. Sjøberg, Pearl Brereton, David Budgen, Tore Dybå, Martin Höst, Dietmar Pfahl, Per Runeson
ESEM2
2010 Are all code smells harmful? A study of God Classes and Brain Classes in the evolution of three open source systems
abstract
Code smells are particular patterns in object-oriented systems that are perceived to lead to difficulties in the maintenance of such systems. It is held that to improve maintainability, code smells should be eliminated by refactoring. It is claimed that classes that are involved in certain code smells are liable to be changed more frequently and have more defects than other classes in the code. We investigated the extent to which this claim is true for God Classes and Brain Classes, with and without normalizing the effects with respect to the class size. We analyzed historical data from 7 to 10 years of the development of three open-source software systems. The results show that God and Brain Classes were changed more frequently and contained more defects than other kinds of class. However, when we normalized the measured effects with respect to size, then God and Brain Classes were less subject to change and had fewer defects than other classes. Hence, under the assumption that God and Brain Classes contain on average as much functionality per line of code as other classes, the presence of God and Brain Classes is not necessarily harmful; in fact, such classes may be an efficient way of organizing code.
Steffen M. Olbrich, Daniela S. Cruzes, Dag I. K. Sjøberg
ICSM3
2010 The usability inspection performance of work-domain experts: An empirical study
abstract
Journal Article The usability inspection performance of work-domain experts: An empirical study Get access Asbjørn Følstad, Asbjørn Følstad * a Department of Cooperative and Trusted Systems, SINTEF, Oslo, Norway * Corresponding author. Address: SINTEF ICT, P.O. Box 124, Blindern, 0314 Oslo, Norway. Tel.: +47 22067515; fax: +47 22067350. E-mail addresses:[email protected] (A. Følstad), [email protected] (B.C.D. Anda), [email protected] (D.I.K. Sjøberg). Search for other works by this author on: Oxford Academic Google Scholar Bente C.D. Anda, Bente C.D. Anda b Department of Informatics, University of Oslo, Norway Search for other works by this author on: Oxford Academic Google Scholar Dag I.K. Sjøberg Dag I.K. Sjøberg b Department of Informatics, University of Oslo, Norway Search for other works by this author on: Oxford Academic Google Scholar Interacting with Computers, Volume 22, Issue 2, March 2010, Pages 75–87, https://doi.org/10.1016/j.intcom.2009.09.001 Published: 06 September 2009 Article history Received: 31 January 2008 Revision received: 20 August 2009 Accepted: 01 September 2009 Published: 06 September 2009
Asbjørn Følstad, Bente Anda, Dag I. K. Sjøberg
Interact. Comput.3
2010 Effects of Personality on Pair Programming
abstract
Personality tests in various guises are commonly used in recruitment and career counseling industries. Such tests have also been considered as instruments for predicting the job performance of software professionals both individually and in teams. However, research suggests that other human-related factors such as motivation, general mental ability, expertise, and task complexity also affect the performance in general. This paper reports on a study of the impact of the Big Five personality traits on the performance of pair programmers together with the impact of expertise and task complexity. The study involved 196 software professionals in three countries forming 98 pairs. The analysis consisted of a confirmatory part and an exploratory part. The results show that: (1) Our data do not confirm a meta-analysis-based model of the impact of certain personality traits on performance and (2) personality traits, in general, have modest predictive value on pair programming performance compared with expertise, task complexity, and country. We conclude that more effort should be spent on investigating other performance-related predictors such as expertise, and task complexity, as well as other promising predictors, such as programming skill and learning. We also conclude that effort should be spent on elaborating on the effects of personality on various measures of collaboration, which, in turn, may be used to predict and influence performance. Insights into such malleable, rather than static, factors may then be used to improve pair programming performance.
Jo Erskine Hannay, Erik Arisholm, Harald Engvik, Dag I. K. Sjøberg
IEEE Trans. Software Eng.4
2009 Using concept mapping for maintainability assessments
abstract
Many important phenomena within software engineering are difficult to define and measure. One example is software maintainability, which has been the subject of considerable research and is believed to be a critical determinant of total software costs. We propose using concept mapping, a well-grounded method used in social research, to operationalize the concept of software maintainability according to a given goal and perspective in a concrete setting. We apply this method to describe four systems that were developed as part of an industrial multiple-case study. The outcome is a conceptual map that displays an arrangement of maintainability constructs, their interrelations, and corresponding measures. Our experience is that concept mapping (1) provides a structured way of combining static code analysis and expert judgment; (2) helps in the tailoring of the choice of measures to a particular system context; and (3) supports the mapping between software measures and aspects of software maintainability. As such, it constitutes a useful addition to existing frameworks for evaluating quality, such as ISO/IEC 9126 and GQM, and tools for static measurement of software code. Overall, concept mapping provides a systematic, structured, and repeatable method for developing constructs and measures, not only of maintainability, but also of software engineering phenomena in general.
Aiko Fallas Yamashita, Hans Christian Benestad, Bente Anda, Per Einar Arnstad, Dag I. K. Sjøberg, Leon Moonen
ESEM5
2009 Comparing of feedback-collection and think-aloud methods in program comprehension studies
abstract
This paper reports an explorative experimental comparison of (i) an experience-sampling method called feedback collection and (ii) the think-aloud methods with respect to their usefulness in studies on program comprehension. Think-aloud methods are widely used in studies of cognitive processes, including program comprehension. Alternatively, as in the feedback-collection method (FCM), cognitive processes can be traced by collecting written feedback from the subjects at regular intervals. We compare FCM with concurrent think-aloud (CTA) and retrospective think-aloud (RTA) regarding type and usefulness of the collected information, costs related to analysis of the collected information and effects of the data collection methods on the subjects' performance. FCM allowed us to identify a greater number of comprehension problems that prevented progress or caused significant delay (FCM: 30 problems; CTA: 5; RTA: 15). It was less precise in identifying strategies for comprehension than CTA (92% correctness for FCM; 100% for CTA). FCM was less expensive in analysis (transcription and coding) than the other two methods (FCM: 0.7 h of analysis per protocol; CTA: 31 h; RTA: 7.9 h). The results indicate that all three methods of data collection were intrusive and affected the performance of the subjects with respect to time and correctness (small to medium effect size). This research confirms that FCM can be used beneficially in studies that trace the cognitive processes involved in, and identify problems related to, the comprehension of software applications. On the basis of our experience, we recommend that FCM be used in studies that have a large number of subjects and as a complement to other methods for tracing cognitive processes, such as user log files. We recommend a design with two groups (verbalisation and silent control) and a pretest task to be used in studies with FCM or CTA that focus on performances.
Amela Karahasanovic, Unni Nyhamar Hinkel, Dag I. K. Sjøberg, Richard C. Thomas
Behav. Inf. Technol.3
2009 The effectiveness of pair programming: A meta-analysis
Jo Erskine Hannay, Tore Dybå, Erik Arisholm, Dag I. K. Sjøberg
Inf. Softw. Technol.4
2009 A systematic review of quasi-experiments in software engineering
Vigdis By Kampenes, Tore Dybå, Jo Erskine Hannay, Dag I. K. Sjøberg
Inf. Softw. Technol.4
2009 Variability and Reproducibility in Software Engineering: A Study of Four Companies that Developed the Same System
abstract
The scientific study of a phenomenon requires it to be reproducible. Mature engineering industries are recognized by projects and products that are, to some extent, reproducible. Yet, reproducibility in software engineering (SE) has not been investigated thoroughly, despite the fact that lack of reproducibility has both practical and scientific consequences. We report a longitudinal multiple-case study of variations and reproducibility in software development, from bidding to deployment, on the basis of the same requirement specification. In a call for tender to 81 companies, 35 responded. Four of them developed the system independently. The firm price, planned schedule, and planned development process, had, respectively, “low,” “low,” and “medium” reproducibilities. The contractor's costs, actual lead time, and schedule overrun of the projects had, respectively, “medium,” “high,” and “low” reproducibilities. The quality dimensions of the delivered products, reliability, usability, and maintainability had, respectively, “low,” "high,” and “low” reproducibilities. Moreover, variability for predictable reasons is also included in the notion of reproducibility. We found that the observed outcome of the four development projects matched our expectations, which were formulated partially on the basis of SE folklore. Nevertheless, achieving more reproducibility in SE remains a great challenge for SE research, education, and industry.
Bente Anda, Dag I. K. Sjøberg, Audris Mockus
IEEE Trans. Software Eng.2
2007 Protocols in the use of empirical software engineering artifacts
Victor R. Basili, Marvin V. Zelkowitz, Dag I. K. Sjøberg, Anthony J. Cowling
Empir. Softw. Eng.3
2007 A systematic review of effect size in software engineering experiments
Vigdis By Kampenes, Tore Dybå, Jo Erskine Hannay, Dag I. K. Sjøberg
Inf. Softw. Technol.4
2007 Evaluating Pair Programming with Respect to System Complexity and Programmer Expertise
abstract
A total of 295 junior, intermediate, and senior professional Java consultants (99 individuals and 98 pairs) from 29 international consultancy companies in Norway, Sweden, and the UK were hired for one day to participate in a controlled experiment on pair programming. The subjects used professional Java tools to perform several change tasks on two alternative Java systems with different degrees of complexity. The results of this experiment do not support the hypotheses that pair programming in general reduces the time required to solve the tasks correctly or increases the proportion of correct solutions. On the other hand, there is a significant 84 percent increase in effort to perform the tasks correctly. However, on the more complex system, the pair programmers had a 48 percent increase in the proportion of correct solutions but no significant differences in the time taken to solve the tasks correctly. For the simpler system, there was a 20 percent decrease in time taken but no significant differences in correctness. However, the moderating effect of system complexity depends on the programmer expertise of the subjects. The observed benefits of pair programming in terms of correctness on the complex system apply mainly to juniors, whereas the reductions in duration to perform the tasks correctly on the simple system apply mainly to intermediates and seniors. It is possible that the benefits of pair programming will exceed the results obtained in this experiment for larger, more complex tasks and if the pair programmers have a chance to work together over a longer period of time
Erik Arisholm, Hans Gallis, Tore Dybå, Dag I. K. Sjøberg
IEEE Trans. Software Eng.4
2007 A Systematic Review of Theory Use in Software Engineering Experiments
abstract
Empirically based theories are generally perceived as foundational to science. However, in many disciplines, the nature, role and even the necessity of theories remain matters for debate, particularly in young or practical disciplines such as software engineering. This article reports a systematic review of the explicit use of theory in a comprehensive set of 103 articles reporting experiments, from of a total of 5,453 articles published in major software engineering journals and conferences in the decade 1993-2002. Of the 103 articles, 24 use a total of 40 theories in various ways to explain the cause-effect relationship(s) under investigation. The majority of these use theory in the experimental design to justify research questions and hypotheses, some use theory to provide post hoc explanations of their results, and a few test or modify theory. A third of the theories are proposed by authors of the reviewed articles. The interdisciplinary nature of the theories used is greater than that of research in software engineering in general. We found that theory use and awareness of theoretical issues are present, but that theory-driven research is, as yet, not a major issue in empirical software engineering. Several articles comment explicitly on the lack of relevant theory. We call for an increased awareness of the potential benefits of involving theory, when feasible. To support software engineering researchers who wish to use theory, we show which of the reviewed articles on which topics use which theories for what purposes, as well as details of the theories' characteristics
Jo Erskine Hannay, Dag I. K. Sjøberg, Tore Dybå
IEEE Trans. Software Eng.2
2006 Keynote Address: The Simula Approach to Experimentation in Software Engineering
abstract
The ultimate goal of software engineering research is to support the private and public software industry in developing higher quality systems with improved timeliness in a more cost-effective and predictable way. One contribution of the empirical software engineering community to this overall goal is the conducting of experiments to evaluate and compare technologies (processes, methods, techniques, languages and tools) for planning, building and maintaining software. However, the applicability of the experimental results to industrial practice is, in most cases, hampered by the experiments’ lack of realism and scale regarding subjects, tasks, systems and environments. In this talk, I will discuss Simula Research Laboratory’s strategy for addressing this challenge: (1) About 25% of our budget is used for hiring software consultants as experimental subjects, mainly at the expense of employing a larger number of researchers. In the last five years, about 800 professionals from 60 companies in several countries have participated in 25 experiments (some of them very large, in order to identify the variances between sub-populations) in which the professionals worked under various controlled circumstances, such as the complexity of tasks and systems, the tools used, whether they worked in pairs, and so on. (2) A large investment in infrastructures and apparatus has been made to support the logistics of running large experiments and surveys, and to collect and organise data with minimal overhead. (3) A senior project manager has been employed to organise the experiments and the resulting data. To increase flexibility and save administrative overhead, Simula hires people on a short-term basis for assistance with, for example, particularly large or complex experiments. They could be students for clerical work or consultants who are particularly qualified for certain tasks, for example, a statistician. (4) Active collaboration with industry (in addition to hiring consultants), such as taking part in industry-managed research projects on software process improvement, and giving seminars and courses, has been considered important. The focus on publicising our research in the media and disseminating it through teaching has also resulted in Simula becoming well known in the Norwegian software industry. (5) Software engineering is typically performed by humans in organisations. Hence, we have established research collaborations with other disciplines, such as psychology, sociology and management.
Dag I. K. Sjøberg
EASE1
2006 A systematic review of statistical power in software engineering experiments
Tore Dybå, Vigdis By Kampenes, Dag I. K. Sjøberg
Inf. Softw. Technol.3
2006 A longitudinal study of development and maintenance in Norway: Report from the 2003 investigation
John Krogstie, Arthur Jahr, Dag I. K. Sjøberg
Inf. Softw. Technol.3
2005 Investigating the Role of Use Cases in the Construction of Class Diagrams
Bente Anda, Dag I. K. Sjøberg
Empir. Softw. Eng.2
2005 Collecting Feedback during Software Engineering Experiments
Amela Karahasanovic, Bente Anda, Erik Arisholm, Siw Elisabeth Hove, Magne Jørgensen, Dag I. K. Sjøberg, Ray Welland
Empir. Softw. Eng.6
2005 A Survey of Controlled Experiments in Software Engineering
abstract
The classical method for identifying cause-effect relationships is to conduct controlled experiments. This paper reports upon the present state of how controlled experiments in software engineering are conducted and the extent to which relevant information is reported. Among the 5,453 scientific articles published in 12 leading software engineering journals and conferences in the decade from 1993 to 2002, 103 articles (1.9 percent) reported controlled experiments in which individuals or teams performed one or more software engineering tasks. This survey quantitatively characterizes the topics of the experiments and their subjects (number of subjects, students versus professionals, recruitment, and rewards for participation), tasks (type of task, duration, and type and size of application) and environments (location, development tools). Furthermore, the survey reports on how internal and external validity is addressed and the extent to which experiments are replicated. The gathered data reflects the relevance of software engineering experiments to industrial practice and the scientific maturity of software engineering research.
Dag I. K. Sjøberg, Jo Erskine Hannay, Ove Hansen, Vigdis By Kampenes, Amela Karahasanovic, Nils-Kristian Liborg, Anette C. Rekdal
IEEE Trans. Software Eng.1
2004 A Controlled Experiment Comparing the Maintainability of Programs Designed with and without Design Patterns-A Replication in a Real Programming Environment
Marek Vokác, Walter F. Tichy, Dag I. K. Sjøberg, Erik Arisholm, Magne Aldrin
Empir. Softw. Eng.3
2004 Evaluating the Effect of a Delegated versus Centralized Control Style on the Maintainability of Object-Oriented Software
abstract
A fundamental question in object-oriented design is how to design maintainable software. According to expert opinion, a delegated control style, typically a result of responsibility-driven design, represents object-oriented design at its best, whereas a centralized control style is reminiscent of a procedural solution, or a "bad" object-oriented design. We present a controlled experiment that investigates these claims empirically. A total of 99 junior, intermediate, and senior professional consultants from several international consultancy companies were hired for one day to participate in the experiment. To compare differences between (categories of) professionals and students, 59 students also participated. The subjects used professional Java tools to perform several change tasks on two alternative Java designs that had a centralized and delegated control style, respectively. The results show that the most skilled developers, in particular, the senior consultants, require less time to maintain software with a delegated control style than with a centralized control style. However, more novice developers, in particular, the undergraduate students and junior consultants, have serious problems understanding a delegated control style, and perform far better with a centralized control style. Thus, the maintainability of object-oriented software depends, to a large extent, on the skill of the developers who are going to maintain it. These results may have serious implications for object-oriented development in an industrial context: having senior consultants design object-oriented systems may eventually pose difficulties unless they make an effort to keep the designs simple, as the cognitive complexity of "expert" designs might be unmanageable for less skilled maintainers.
Erik Arisholm, Dag I. K. Sjøberg
IEEE Trans. Software Eng.2
2003 An effort prediction interval approach based on the empirical distribution of previous estimation accuracy
Magne Jørgensen, Dag I. K. Sjøberg
Inf. Softw. Technol.2
2003 Software effort estimation by analogy and "regression toward the mean"
Magne Jørgensen, Ulf Geir Indahl, Dag I. K. Sjøberg
J. Syst. Softw.3
2002 Towards an inspection technique for use case models
abstract
A use case model describes the functional requirements of a software system and is used as input to several activities in a software development project. The quality of the use case model therefore has an important impact on the quality of the resulting software product. Software inspection is regarded as one of the most efficient methods for verifying software documents. There are inspection techniques for most documents produced in a software development project, but no comprehensive inspection technique exists for use case models. This paper presents a taxonomy of typical defects in use case models and proposes a checklist-based inspection technique for detecting such defects. This inspection technique was evaluated in two studies with undergraduate students as subjects. The results from the evaluations indicate that inspections are useful for detecting defects in use case models and motivate further studies to improve the proposed inspection technique.
Bente Anda, Dag I. K. Sjøberg
SEKE2
2002 Impact of experience on maintenance skills
abstract
Abstract This study reports results from an empirical study of 54 software maintainers in the software maintenance department of a Norwegian company. The study addresses the relationship between amount of experience and maintenance skills. The findings were, amongst others, as follows. (1) While there may have been a reduction in the frequency of major unexpected problems from tasks solved by very inexperienced to medium experienced maintainers, additional years of general software maintenance experience did not lead to further reduction. More application specific experience, however, further reduced the frequency of major unexpected problems. (2) The most experienced maintainers did not predict maintenance problems better than maintainers with little or medium experience. (3) A simple one‐variable model outperformed the maintainers' predictions of maintenance problems, i.e. the average prediction performance of the maintainers seems poor. An important reason for the weak correlation between length of experience and ability to predict maintenance problems may be the lack of meaningful feedback on the predictions. Copyright © 2002 John Wiley & Sons, Ltd.
Magne Jørgensen, Dag I. K. Sjøberg
J. Softw. Maintenance Res. Pract.2
2001 Quality and Understandability of Use Case Models
Bente Anda, Dag I. K. Sjøberg, Magne Jørgensen
ECOOP2
2001 Software effort estimation by analogy and regression toward the mean
Magne Jørgensen, Ulf Geir Indahl, Dag I. K. Sjøberg
SEKE3
2001 Assessing the Changeability of two Object-Oriented Design Alternatives - a Controlled Experiment
Erik Arisholm, Dag I. K. Sjøberg, Magne Jørgensen
Empir. Softw. Eng.2
2001 Impact of effort estimates on software project work
Magne Jørgensen, Dag I. K. Sjøberg
Inf. Softw. Technol.2
2000 A study of development and maintenance in Norway: assessing the efficiency of information systems support using functional maintenance
Knut Kjetil Holgeid, John Krogstie, Dag I. K. Sjøberg
Inf. Softw. Technol.3
2000 Towards a framework for empirical assessment of changeability decay
Erik Arisholm, Dag I. K. Sjøberg
J. Syst. Softw.2
1999 Empirical Studies of Evolving Systems
Keith H. Bennett, Elizabeth Burd, Chris F. Kemerer, Meir M. Lehman, Raymond J. Madachy, C. Mair, Dag I. K. Sjøberg, Sandra Slaughter
Empir. Softw. Eng.8
1997 Software Constraints for Large Application Systems
abstract
As application systems live longer and grow in size and complexity, there is an ever increasing need for methods and tools that can support software builders in constructing maintainable, well-structured and consistent systems. This paper describes the notion of software constraints as an aid to developing such systems. Software constraints make rules and conventions commonly agreed to in a given programming environment explicit and automatically checkable. The potential usefulness of software constraints was investigated in both industrial and research environments. A framework for categorization of such constraints is defined. Constraints are proposed that are generally applicable and others that are tightly connected to and support a certain programming method. Tools for automatic checking are crucial if software constraints are to be used. An architecture for such tools and two realizations are described.
Dag I. K. Sjøberg, Ray Welland, Malcolm P. Atkinson 0001
Comput. J.1
1997 Exploiting Persistence in Build Management
abstract
A challenging issue in the construction and maintenance of large application systems is how to determine which components need to be rebuilt after change, when and in which order. Rebuilding is typically recompilation and linking, but may also include update of derivable components such as cross-reference databases and re-creation of library indexes. Type definitions or schema, and data values in a file store, database or persistent store may also need to be rebuilt. The main purpose of this paper is to describe how persistent language technology can be exploited to enhance build management. In particular, the paper describes a method for transactional, incremental linking and the implementation of its support. To help implement this method, and to make it safer and more efficient to carry out rebuild activities in general, we have defined a set of automatically checkable constraints on the software. The build management tool we have implemented, the Builder, derives rebuild dependencies automatically and offers partitioning of dependency graphs—a means to defer or avoid unnecessary rebuilding. The Builder is implemented in a persistent programming language and provides build management for applications written in the language. It exploits features such as strong typing, runtime linguistic reflection, and referential integrity provided by the language processing technology. The Builder operates over both programs and (complex) data, which is in contrast to conventional language-centred tools. © 1997 by John Wiley & Sons, Ltd.
Dag I. K. Sjøberg, Ray Welland, Malcolm P. Atkinson 0001, Paul Philbrow, Cathy Waite
Softw. Pract. Exp.1
1993 Quantifying schema evolution
Dag I. K. Sjøberg
Inf. Softw. Technol.1
1988 Database Concepts Discussed in Object-Oriented Perspective
Yngve Lindsjørn, Dag I. K. Sjøberg
ECOOP2