VLDB 2026 Research / reviewers in the wild / expert
Jeffrey C. Carver
dblp:c/JeffreyCCarver · also Jeff Carver 0001
· DBLP profile ↗
101ranked-venue papers
16as first author
19since 2021 · last 2026
0000-0002-7824-9151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 79 · 11 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 15 · 5 first-author · 4 since 2021Systems, architecture and hardware · 5 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Barriers to Use: Perspectives on Environmental Research Software
Yvette D. Hastings, Jeffrey C. Carver, Clemente Izurieta, Ann Marie Reinhold |
ICSOFT | 2 |
| 2026 | Peer code review in research software development: The research software engineer perspective
Md. Ariful Islam Malik, Jeffrey C. Carver, Nasir U. Eisty |
Empir. Softw. Eng. | 2 |
| 2026 | Preparing research software engineers to become security champions: Development and evaluation of a security awareness workshop
Matthew Armstrong, Jeffrey C. Carver, Reed Milewicz |
Future Gener. Comput. Syst. | 2 |
| 2026 | Characterizing the security culture of the research software engineering community: An empirical study
Matthew Armstrong, Jeffrey C. Carver, Reed Milewicz, Michael Meinel, Michael Felderer |
Future Gener. Comput. Syst. | 2 |
| 2025 | Sixth Annual Workshop on A/B Testing and Platform-Enabled Learning Engineering (PELE)abstractLearning engineering applies data and learning science principles to better understand outcomes and support improvement research. One important approach is A/B testing-common in large software companies and also represented academically at conferences like the Annual Conference on Digital Experimentation (CODE), and the International Consortium for Innovation and Collaboration in Learning Engineering (IEEE ICICLE). Several systems supporting A/B testing in educational applications have arisen recently, including UpGrade, E-TRIALS, and Terracotta. A/B testing can help improve educational platforms, yet there are challenging issues unique to conducting such work in these contexts. In response, a number of digital learning platforms have opened their systems to learning-improvement research by instructors and/or third-party researchers, with specific supports necessary for education-specific research designs. This workshop will explore how A/B testing is conducted in educational contexts, how digital learning platforms are accelerating education research, and how empirical approaches can be used to drive powerful gains in student learning. It will also discuss opportunities for funding to conduct platform-enabled learning engineering. April Murphy, Stephen Fancsali, Steven Ritter 0001, Neil T. Heffernan, Debshila Basu Mallick, Jeremy Roschelle, Danielle S. McNamara, Joseph Jay Williams, John C. Stamper, Norman L. Bier, Jeffrey C. Carver |
L@S | 11 |
| 2025 | Testing research software: an in-depth survey of practices, methods, and tools
Nasir U. Eisty, Upulee Kanewala, Jeffrey C. Carver |
Empir. Softw. Eng. | 3 |
| 2023 | Community smells - The sources of social debt: A systematic literature review
Eduardo Caballero-Espinosa, Jeffrey C. Carver, Kimberly Stowers |
Inf. Softw. Technol. | 2 |
| 2023 | Human error management in requirements engineering: Should we fix the people, the processes, or the environment?
Sweta Mahaju, Jeffrey C. Carver, Gary L. Bradshaw |
Inf. Softw. Technol. | 2 |
| 2022 | Training Computing Educators to Become Computing Education ResearchersabstractThe computing education community endeavors to consistently move forward, improving the educational experience of our students. As new innovations in computing education practice are learned and shared, however, these papers may not exhibit the desired qualities that move simple experience reports to true Scholarship of Teaching and Learning (SoTL). We report on our six years of experience in running professional development for computing educators in empirical research methods for social and behavioral studies in the classroom. Our goal is to have a direct impact on instructors who are in the beginning stages of transitioning their educational innovations from anecdotal to empirical results that can be replicated by instructors at other institutions. To achieve this, we created a year-long mentoring experience, beginning with a multi-day workshop on empirical research methods during the summer, followed by regular mentoring sessions with participants, and culminating in a follow-up session at the following year's SIGCSE Technical Symposium. From survey results and as evidenced by eventual research results and publications from participants, we believe that our method of structuring empirical research professional development was successful and could be a model for similar programs in other areas. Jeffrey C. Carver, Sarah Smith Heckman, Mark Sherriff |
SIGCSE (1) | 1 |
| 2022 | Testing Tutor: A Testing Pedagogical Active Learning PlatformabstractTesting Tutor is a web-based platform that helps instructors support software testing pedagogy by automatically diagnosing the fundamental testing concepts (e.g., boundary value) not covered in students' test suite and subsequently helping students initiate their own learning process about those concepts and systematically improve their test suites. The platform's differentiating features include 1) customizable feedback engine which allows instructors to scaffold the level of feedback (varying from conceptual to detailed), 2) a built-in repository of problems instructors can use, 3) access to digital learning content and 4) modes (learning and development) so instructors can scaffold the level of problems. This demo provides a brief overview of the platform and a walkthrough of two example use cases that illustrate the power of Testing Tutor from the perspectives of an instructor and a student. The two use cases will (1) demonstrate using Testing Tutor at different course levels by walking through the steps for an instructor to configure an assignment and the feedback engine and (2) demonstrate the student's experience submitting an assignment and receiving feedback. Information about Testing Tutor can be found at https://testingtutor.org. This work is supported by NSF IUSE grants 2013296 and 2013342. Lucas P. Cordova, Jeffrey C. Carver, Gursimran Singh Walia |
SIGCSE (2) | 2 |
| 2022 | Educating Students to be Better Citizens of Tech CommunitiesabstractIt well known that women are underrepresented in technology, holding less than 25% of IT positions. Reports of women and minorities being harassed, discriminated against, and abused in technology communities routinely appear in scholarly publications and popular media. These types of negative interactions add to the diversity problem by discouraging women and minorities from even considering participation in technology communities. To help address this diversity problem and encourage better citizenship in technology communities, we are focused on improving the soft skills of students. Research has identified the need for engineers with better soft skills because a large part of their job is non-technical and involves teamwork, effective communication, and conflict resolution. Engineers need the skill to be employable but also to be responsible, ethical, and aware of the societal impacts of engineering. We believe we can have a positive impact in teaching soft skills to engineers if students engage in a reflective, experiential, and human-centered experience. With this objective in mind, we are developing an undergraduate course that combines these elements by embedding students in OSS communities as contributors and providing them with a safe space for reflection and human-centered learning. Students can contribute by developing code, creating documentation, fixing bugs, improving usability, or performing testing. We plan to develop a positive classroom atmosphere that encourages sharing, reflection, and mutual support for the students while they are contributing to an OSS project. Vandana Singh, Jeffrey C. Carver |
SIGCSE (2) | 2 |
| 2022 | Developers perception of peer code review in research software development
Nasir U. Eisty, Jeffrey C. Carver |
Empir. Softw. Eng. | 2 |
| 2022 | Testing research software: a survey
Nasir U. Eisty, Jeffrey C. Carver |
Empir. Softw. Eng. | 2 |
| 2022 | Assessing expert system-assisted literature reviews with a case study
Zhe Yu 0002, Jeffrey C. Carver, Gregg Rothermel, Tim Menzies |
Expert Syst. Appl. | 2 |
| 2022 | A Systematic Literature Review of Empiricism and Norms of Reporting in Computing Education Research LiteratureabstractContext. Computing Education Research (CER) is critical to help the computing education community and policy makers support the increasing population of students who need to learn computing skills for future careers. For a community to systematically advance knowledge about a topic, the members must be able to understand published work thoroughly enough to perform replications, conduct meta-analyses, and build theories. There is a need to understand whether published research allows the CER community to systematically advance knowledge and build theories. Objectives. The goal of this study is to characterize the reporting of empiricism in Computing Education Research literature by identifying whether publications include content necessary for researchers to perform replications, meta-analyses, and theory building. We answer three research questions related to this goal: (RQ1) What percentage of papers in CER venues have some form of empirical evaluation? (RQ2) Of the papers that have empirical evaluation, what are the characteristics of the empirical evaluation? (RQ3) Of the papers that have empirical evaluation, do they follow norms (both for inclusion and for labeling of information needed for replication, meta-analysis, and, eventually, theory-building) for reporting empirical work? Methods. We conducted a systematic literature review of the 2014 and 2015 proceedings or issues of five CER venues: Technical Symposium on Computer Science Education (SIGCSE TS), International Symposium on Computing Education Research (ICER), Conference on Innovation and Technology in Computer Science Education (ITiCSE), ACM Transactions on Computing Education (TOCE), and Computer Science Education (CSE). We developed and applied the CER Empiricism Assessment Rubric to the 427 papers accepted and published at these venues over 2014 and 2015. Two people evaluated each paper using the Base Rubric for characterizing the paper. An individual person applied the other rubrics to characterize the norms of reporting, as appropriate for the paper type. Any discrepancies or questions were discussed between multiple reviewers to resolve. Results. We found that over 80% of papers accepted across all five venues had some form of empirical evaluation. Quantitative evaluation methods were the most frequently reported. Papers most frequently reported results on interventions around pedagogical techniques, curriculum, community, or tools. There was a split in papers that had some type of comparison between an intervention and some other dataset or baseline. Most papers reported related work, following the expectations for doing so in the SIGCSE and CER community. However, many papers were lacking properly reported research objectives, goals, research questions, or hypotheses; description of participants; study design; data collection; and threats to validity. These results align with prior surveys of the CER literature. Conclusions. CER authors are contributing empirical results to the literature; however, not all norms for reporting are met. We encourage authors to provide clear, labeled details about their work so readers can use the study methodologies and results for replications and meta-analyses. As our community grows, our reporting of CER should mature to help establish computing education theory to support the next generation of computing learners. Sarah Smith Heckman, Jeffrey C. Carver, Mark Sherriff, Ahmed Al-Zubidy |
ACM Trans. Comput. Educ. | 2 |
| 2022 | How do Practitioners Perceive the Relevance of Requirements Engineering Research?abstractContext: The relevance of Requirements Engineering (RE) research to practitioners is vital for a long-term dissemination of research results to everyday practice. Some authors have speculated about a mismatch between research and practice in the RE discipline. However, there is not much evidence to support or refute this perception.Objective: This article presents the results of a study aimed at gathering evidence from practitioners about their perception of the relevance of RE research and at understanding the factors that influence that perception.Method: We conducted a questionnaire-based survey of industry practitioners with expertise in RE. The participants rated the perceived relevance of 435 scientific papers presented at five top RE-related conferences.Results: The 153 participants provided a total of 2,164 ratings. The practitioners rated RE research as essential or worthwhile in a majority of cases. However, the percentage of non-positive ratings is still higher than we would like. Among the factors that affect the perception of relevance are the research's links to industry, the research method used, and respondents’ roles. The reasons for positive perceptions were primarily related to the relevance of the problem and the soundness of the solution, while the causes for negative perceptions were more varied. The respondents also provided suggestions for future research, including topics researchers have studied for decades, like elicitation or requirement quality criteria.Conclusions: The study is valuable for both researchers and practitioners. Researchers can use the reasons respondents gave for positive and negative perceptions and the suggested research topics to help make their research more appealing to practitioners and thus more prone to industry adoption. Practitioners can benefit from the overall view of contemporary RE research by learning about research topics that they may not be familiar with, and compare their perception with those of their colleagues to self-assess their positioning towards more academic research. Xavier Franch, Daniel Méndez 0001, Andreas Vogelsang, Rogardt Heldal, Eric Knauss, Marc Oriol, Guilherme Horta Travassos, Jeffrey C. Carver, Thomas Zimmermann 0001 |
IEEE Trans. Software Eng. | 8 |
| 2021 | A Comparison of Inquiry-Based Conceptual Feedback vs. Traditional Detailed Feedback Mechanisms in Software Testing Education: An Empirical InvestigationabstractThe feedback provided by current testing education tools about the deficiencies in a student's test suite either mimics industry code coverage tools or lists specific instructor test cases that are missing from the student's test suite. While useful in some sense, these types of feedback are akin to revealing the solution to the problem, which can inadvertently encourage students to pursue a trial-and-error approach to testing, rather than using a more systematic approach that encourages learning. In addition to not teaching students why their test suite is inadequate, this type of feedback may motivate students to become dependent on the feedback rather than thinking for themselves. To address this deficiency, there is an opportunity to investigate alternative feedback mechanisms that include a positive reinforcement of testing concepts. We argue that using an inquiry-based learning approach is better than simply providing the answers. To facilitate this type of learning, we present Testing Tutor, a web-based assignment submission platform that supports different levels of testing pedagogy via a customizable feedback engine. We evaluated the impact of the different types of feedback through an empirical study in two sophomore-level courses. We use Testing Tutor to provide students with different types of feedback, either traditional detailed code coverage feedback or inquiry-based learning conceptual feedback, and compare the effects. The results show that students that receive conceptual feedback had higher code coverage (by different measures), fewer redundant test cases, and higher programming grades than the students who receive traditional code coverage feedback. Lucas P. Cordova, Jeffrey C. Carver, Noah Gershmel, Gursimran Singh Walia |
SIGCSE | 2 |
| 2021 | Understanding peer review of software engineering papers
Neil A. Ernst, Jeffrey C. Carver, Daniel Méndez 0001, Marco Torchiano |
Empir. Softw. Eng. | 2 |
| 2021 | Software engineering practices for scientific software development: A systematic mapping studyabstractBackground: The development of scientific software applications is far from trivial, due to the constant increase in the necessary complexity of these applications, their increasing size, and their need for intensive maintenance and reuse. Aim: To this end, developers of scientific software (who usually lack a formal computer science background) need to use appropriate software engineering (SE) practices. This paper describes the results of a systematic mapping study on the use of SE for scientific application development and their impact on software quality. Method: To achieve this goal we have performed a systematic mapping study on 359 papers. We first describe a catalogue of SE practices used in scientific software development. Then, we discuss the quality attributes of interest that drive the application of these practices, as well as tentative side-effects of applying the practices on qualities. Results: The main findings indicate that scientific software developers are focusing on practices that improve implementation productivity, such as code reuse, use of third-party libraries, and the application of "good" programming techniques. In addition, apart from the finding that performance is a key-driver for many of these applications, scientific software developers also find maintainability and productivity to be important. Conclusions: The results of the study are compared to existing literature, are interpreted under a software engineering prism, and various implications for researchers and practitioners are provided. One of the key findings of the study, which is considered as important for driving future research endeavors is the lack of evidence on the trade-offs that need to be made when applying a software practice, i.e., negative (indirect) effects on other quality attributes. Elvira-Maria Arvanitou, Apostolos Ampatzoglou, Alexander Chatzigeorgiou, Jeffrey C. Carver |
J. Syst. Softw. | 4 |
| 2019 | Use of Software Process in Research Software Development: A SurveyabstractBackground: Developers face challenges in building high-quality research software due to its inherent complexity. These challenges can reduce the confidence users have in the quality of the result produced by the software. Use of a defined software development process, which divides the development into distinct phases, results in improved design, more trustworthy results, and better project management. Aims: This paper focuses on gaining a better understanding of the use of software development process for research software. Method: We surveyed research software developers to collect information about their use of software development processes. We analyze whether and demographic factors influence the respondents' use of and perceived value in defined process. Results: Based on 98 responses, research software developers appear to follow a defined software development process at least some of the time. The respondents also have a strong positive perception about the value of following processes. Conclusions: To produce high-quality and reliable research software, which is critical for many research domains, research software developers must follow a proper software development process. The results indicate a positive perception of value about using defined development processes that should lead to both short-term benefits through improved results and long-term benefits through more maintainable software. Nasir U. Eisty, George K. Thiruvathukal, Jeffrey C. Carver |
EASE | 3 |
| 2019 | FLOSS participants' perceptions about gender and inclusiveness: a surveyabstractBackground: While FLOSS projects espouse openness and acceptance for all, in practice, female contributors often face discriminatory barriers to contribution. Aims: In this paper, we examine the extent to which these problems still exist. We also study male and female contributors' perceptions of other contributors. Method: We surveyed participants from 15 FLOSS projects, asking a series of open-ended, closed-ended, and behavioral scale questions to gather information about the issue of gender in FLOSS projects. Results: Though many of those we surveyed expressed a positive sentiment towards females who participate in FLOSS projects, some were still strongly against their inclusion. Often, the respondents who were against inclusiveness also believed their own sentiments were the prevailing belief in the community, contrary to our findings. Others did not see the purpose of attempting to be inclusive, expressing the sentiment that a discussion of gender has no place in FLOSS. Conclusions: FLOSS projects have started to move forwards in terms of gender acceptance. However, there is still a need for more progress in the inclusion of gender-diverse contributors. Amanda Lee, Jeffrey C. Carver |
ICSE | 2 |
| 2019 | Software Testing in Introductory Programming Courses: A Systematic Mapping StudyabstractTraditionally, students learn about software testing during intermediate or advanced computing courses. However, it is widely advocated that testing should be addressed beginning in introductory programming courses. In this context, testing practices can help students think more critically while working on programming assignments. At the same time, students can develop testing skills throughout the computing curriculum. Considering this scenario, we conducted a systematic mapping of the literature about software testing in introductory programming courses, resulting in 293 selected papers. We mapped the papers to categories with respect to their investigated topic (curriculum, teaching methods, programming assignments, programming process, tools, program/test quality, concept understanding, and students' perceptions and behaviors) and evaluation method (literature review, exploratory study, descriptive/persuasive study, survey, qualitative study, experimental and experience report). We also identified the benefits and drawbacks of this teaching approach, as pointed out in the selected papers. The goal is to provide an overview of research performed in the area, highlighting gaps that should be further investigated. Lilian P. Scatalon, Jeffrey C. Carver, Rogério Eduardo Garcia, Ellen Francine Barbosa |
SIGCSE | 2 |
| 2019 | Identification and prioritization of SLR search tool requirements: an SLR and a survey
Ahmed Al-Zubidy, Jeffrey C. Carver |
Empir. Softw. Eng. | 2 |
| 2019 | Empirical research on concurrent software testing: A systematic mapping study
Silvana M. Melo, Jeffrey C. Carver, Paulo Sergio Lopes de Souza, Simone do Rócio Senger de Souza |
Inf. Softw. Technol. | 2 |
| 2018 | A Survey of Software Metric Use in Research Software DevelopmentabstractBackground: Breakthroughs in research increasingly depend on complex software libraries, tools, and applications aimed at supporting specific science, engineering, business, or humanities disciplines. The complexity and criticality of this software motivate the need for ensuring quality and reliability. Software metrics are a key tool for assessing, measuring, and understanding software quality and reliability. Aims: The goal of this work is to better understand how research software developers use traditional software engineering concepts, like metrics, to support and evaluate both the software and the software development process. One key aspect of this goal is to identify how the set of metrics relevant to research software corresponds to the metrics commonly used in traditional software engineering. Method: We surveyed research software developers to gather information about their knowledge and use of code metrics and software process metrics. We also analyzed the influence of demographics (project size, development role, and development stage) on these metrics. Results: The survey results, from 129 respondents, indicate that respondents have a general knowledge of metrics. However, their knowledge of specific SE metrics is lacking, their use even more limited. The most used metrics relate to performance and testing. Even though code complexity often poses a significant challenge to research software development, respondents did not indicate much use of code metrics. Conclusions: Research software developers appear to be interested and see some value in software metrics but may be encountering roadblocks when trying to use them. Further study is needed to determine the extent to which these metrics could provide value in continuous process improvement. Nasir U. Eisty, George K. Thiruvathukal, Jeffrey C. Carver |
eScience | 3 |
| 2018 | Designing Empirical Education Research Studies (DEERS): Creating an Answerable Research Question (Abstract Only)abstractOne of the most important, and difficult, aspects of starting an education research project is identifying an interesting, answerable, repeatable, measurable, and appropriately scoped research question. The lack of a valid research question reduces the potential impact of the work and could result in wasted effort. The goal of this workshop is to help educational researchers get off on the right foot by defining such a research question. This workshop is part of the larger Designing Empirical Education Research Studies (DEERS) project, which consists of an ongoing series of workshops in which researcher cohorts work with experienced empirical researchers to design, implement, evaluate, and publish empirical work in computer science education. In addition to instruction on the various aspects of good research questions, DEERS alumni will join us to mentor attendees in development of their own research questions in small group breakout sessions. At the end of the workshop, attendees will leave with a valid research question that can then be the start for designing a research study. Attendees will also receive information on how to apply to attend the full summer workshop, where they can fully flesh out the empirical study design, and join a DEERS research cohort. More information about DEERS can be found at http://empiricalcsed.org. Jeffrey C. Carver, Sarah Smith Heckman, Mark Sherriff |
SIGCSE | 1 |
| 2018 | Using human error information for error prevention
Jeffrey C. Carver, Vaibhav K. Anu, Gursimran Singh Walia, Gary L. Bradshaw |
Empir. Softw. Eng. | 2 |
| 2018 | Program comprehension of domain-specific and general-purpose languages: replication of a family of experiments using integrated development environments
Tomaz Kosar, Saso Gaberc, Jeffrey C. Carver, Marjan Mernik |
Empir. Softw. Eng. | 3 |
| 2018 | Development of a human error taxonomy for software requirements: A systematic literature review
Vaibhav K. Anu, Jeffrey C. Carver, Gursimran Singh Walia, Gary L. Bradshaw |
Inf. Softw. Technol. | 3 |
| 2018 | Attack surface definitions: A systematic literature review
Christopher Theisen, Nuthan Munaiah, Mahran Al-Zyoud, Jeffrey C. Carver, Andrew Meneely, Laurie A. Williams |
Inf. Softw. Technol. | 4 |
| 2017 | Issues and Opportunities for Human Error-Based Requirements Inspections: An Exploratory Studyabstract[Background] Software inspections are extensively used for requirements verification. Our research uses the perspective of human cognitive failures (i.e., human errors) to improve the fault detection effectiveness of traditional fault-checklist based inspections. Our previous evaluations of a formal human error based inspection technique called Error Abstraction and Inspection (EAI) have shown encouraging results, but have also highlighted a real need for improvement. [Aims and Method] The goal of conducting the controlled study presented in this paper was to identify the specific tasks of EAI that inspectors find most difficult to perform and the strategies that successful inspectors use when performing the tasks. [Results] The results highlighted specific pain points of EAI that can be addressed by improving the training and instrumentation. Vaibhav K. Anu, Gursimran Singh Walia, Jeffrey C. Carver, Gary L. Bradshaw |
ESEM | 4 |
| 2017 | Are One-Time Contributors Different? A Comparison to Core and Periphery Developers in FLOSS RepositoriesabstractContext: Free/Libre Open Source Software (FLOSS) communities consist of different types of contributors. Core contributors and peripheral contributors work together to create a successful project, each playing a different role. One-Time Contributors (OTCs), who are on the very fringe of the peripheral developers, are largely unstudied despite offering unique insights into the development process. In a prior survey, we identified OTCs and discovered their motivations and barriers. Aims: The objective of this study is to corroborate the survey results and provide a better understand of OTCs. We compare OTCs to other peripheral and core contributors to determine whether they are distinct. Method: We mined data from the same code-review repository used to identify survey respondents in our previous study. After identifying each contributor as core, periphery, or OTC, we compared them in terms of patch size, time interval from submission to decision, the nature of their conversations, and patch acceptance rates. Results: We identified a continuum between core developers and OTCs. OTCs create smaller patches, face longer time intervals between patch submission and rejection, have longer review conversations, and face lower patch acceptance rates. Conversely, core contributors create larger patches, face shorter time intervals for feedback, have shorter review conversations, and have patches accepted at the highest rate. The peripheral developers fall in between the OTCs and the core contributors. Conclusion: OTCs do, in fact, face the barriers identified in our prior survey. They represent a distinct group of contributors compared to core and peripheral developers. Amanda Lee, Jeffrey C. Carver |
ESEM | 2 |
| 2017 | Understanding the impressions, motivations, and barriers of one time code contributors to FLOSS projects: a surveyabstractSuccessful Free/Libre Open Source Software (FLOSS) projects must attract and retain high-quality talent. Researchers have invested considerable effort in the study of core and peripheral FLOSS developers. To this point, one critical subset of developers that have not been studied are One-Time code Contributors (OTC) - those that have had exactly one patch accepted. To understand why OTCs have not contributed another patch and provide guidance to FLOSS projects on retaining OTCs, this study seeks to understand the impressions, motivations, and barriers experienced by OTCs. We conducted an online survey of OTCs from 23 popular FLOSS projects. Based on the 184 responses received, we observed that OTCs generally have positive impressions of their FLOSS project and are driven by a variety of motivations. Most OTCs primarily made contributions to fix bugs that impeded their work and did not plan on becoming long term contributors. Furthermore, OTCs encounter a number of barriers that prevent them from continuing to contribute to the project. Based on our findings, there are some concrete actions FLOSS projects can take to increase the chances of converting OTCs into long-term contributors. Amanda Lee, Jeffrey C. Carver, Amiangshu Bosu |
ICSE | 2 |
| 2017 | How do Practitioners Perceive the Relevance of Requirements Engineering Research? An Ongoing StudyabstractThe relevance of Requirements Engineering (RE) research to practitioners is a prerequisite for problem-driven research in the area and key for a long-term dissemination of research results to everyday practice. To understand better how industry practitioners perceive the practical relevance of RE research, we have initiated the RE-Pract project, an international collaboration conducting an empirical study. This project opts for a replication of previous work done in two different domains and relies on survey research. To this end, we have designed a survey to be sent to several hundred industry practitioners at various companies around the world and ask them to rate their perceived practical relevance of the research described in a sample of 418 RE papers published between 2010 and 2015 at the RE, ICSE, FSE, ESEC/FSE, ESEM and REFSQ conferences. In this paper, we summarize our research protocol and present the current status of our study and the planned future steps. Xavier Franch, Daniel Méndez 0001, Marc Oriol, Andreas Vogelsang, Rogardt Heldal, Eric Knauss, Guilherme Horta Travassos, Jeffrey C. Carver, Óscar Dieste Tubío, Thomas Zimmermann 0001 |
RE | 8 |
| 2017 | Usefulness of a Human Error Identification Tool for Requirements Inspection: An Experience Report
Vaibhav K. Anu, Gursimran Singh Walia, Gary L. Bradshaw, Jeffrey C. Carver |
REFSQ | 5 |
| 2017 | Defect Prevention in Requirements Using Human Error Information: An Empirical Study
Jeffrey C. Carver, Vaibhav K. Anu, Gursimran Singh Walia, Gary L. Bradshaw |
REFSQ | 2 |
| 2017 | Designing Empirical Education Research Studies (DEERS): Creating an Answerable Research Question (Abstract Only)abstractOne of the most important, and difficult, aspects of starting an education research project is identifying an interesting, answerable, repeatable, measurable, and appropriately scoped research question. The lack of a valid research question reduces the potential impact of the work and could result in wasted effort. The goal of this workshop is to help educational researchers get off on the right foot by defining such a research question. This workshop is part of the larger Designing Empirical Education Research Studies (DEERS) project, which consists of an ongoing series of workshops in which researcher cohorts work with experienced empirical researchers to design, implement, evaluate, and publish empirical work in computer science education. In addition to instruction on the various aspects of good research questions, DEERS alumni will join us to mentor attendees in development of their own research questions in small group breakout sessions. At the end of the workshop, attendees will leave with a valid research question that can then be the start for designing a research study. Attendees will also receive information on how to apply to attend the full summer workshop, where they can fully flesh out the empirical study design, and join a DEERS research cohort. More information about DEERS can be found at http://empiricalcsed.org. Sarah Smith Heckman, Jeffrey C. Carver, Mark Sherriff |
SIGCSE | 2 |
| 2017 | Vision for SLR tooling infrastructure: Prioritizing value-added requirements
Ahmed Al-Zubidy, Jeffrey C. Carver, David P. Hale, Edgar E. Hassler |
Inf. Softw. Technol. | 2 |
| 2017 | Test-Driven Development in scientific software: a survey
Aziz Nanthaamornphong, Jeffrey C. Carver |
Softw. Qual. J. | 2 |
| 2017 | Process Aspects and Social Dynamics of Contemporary Code Review: Insights from Open Source Development and Industrial Practice at MicrosoftabstractMany open source and commercial developers practice contemporary code review, a lightweight, informal, tool-based code review process. To better understand this process and its benefits, we gathered information about code review practices via surveys of open source software developers and developers from Microsoft. The results of our analysis suggest that developers spend approximately 10-15 percent of their time in code reviews, with the amount of effort increasing with experience. Developers consider code review important, stating that in addition to finding defects, code reviews offer other benefits, including knowledge sharing, community building, and maintaining code quality. The quality of the code submitted for review helps reviewers form impressions about their teammates, which can influence future collaborations. We found a large amount of similarity between the Microsoft and OSS respondents. One interesting difference is that while OSS respondents view code review as an important method of impression formation, Microsoft respondents found knowledge dissemination to be more important. Finally, we found little difference between distributed and co-located Microsoft teams. Our findings identify the following key areas that warrant focused research: 1) exploring the non-technical benefits of code reviews, 2) helping developers in articulating review comments, and 3) assisting reviewers' program comprehension during code reviews. Amiangshu Bosu, Jeffrey C. Carver, Christian Bird, Jonathan D. Orbeck, Christopher Chockley |
IEEE Trans. Software Eng. | 2 |
| 2016 | How Practitioners Perceive the Relevance of ESEM ResearchabstractBackground: The relevance of ESEM research to industry practitioners is key to the long-term health of the conference. Aims: The goal of this work is to understand how ESEM research is perceived within the practitioner community and provide feedback to the ESEM community ensure our research remains relevant. Method: To understand how practitioners perceive ESEM research, we replicated previous work by sending a survey to several hundred industry practitioners at a number of companies around the world. We asked the survey participants to rate the relevance of the research described in 156 ESEM papers published between 2011 and 2015. Results: We received 9,941 ratings by 437 practitioners who labeled ideas as Essential, Worth-while, Unimportant, or Unwise. The results showed that overall, industrial practitioners find the work published in ESEM to be valuable: 67% of all ratings were essential or worthwhile. We found no correlation between citation count and perceived relevance of the papers. Through a qualitative analysis, we also identified a number of research themes on which practitioners would like to see an increased research focus. Conclusions: The work published in ESEM is generally relevant to industrial practitioners. There are a number of topics for which those practitioners would like to see additional research undertaken. Jeffrey C. Carver, Óscar Dieste Tubío, Nicholas A. Kraft, David Lo 0001, Thomas Zimmermann 0001 |
ESEM | 1 |
| 2016 | Detection of Requirement Errors and Faults via a Human Error Taxonomy: A Feasibility StudyabstractBackground: Developing correct software requirements is important for overall software quality. Most existing quality improvement approaches focus on detection and removal of faults (i.e. problems recorded in a document) as opposed identifying the underlying errors that produced those faults. Accordingly, developers are likely to make the same errors in the future and fail to recognize other existing faults with the same origins. Therefore, we have created a Human Error Taxonomy (HET) to help software engineers improve their software requirement specification (SRS) documents. Aims: The goal of this paper is to analyze whether the HET is useful for classifying errors and for guiding developers to find additional faults. Methods: We conducted a empirical study in a classroom setting to evaluate the usefulness and feasibility of the HET. Results: First, software developers were able to employ error categories in the HET to identify and classify the underlying sources of faults identified during the inspection of SRS documents. Second, developers were able to use that information to detect additional faults that had gone unnoticed during the initial inspection. Finally, the participants had a positive impression about the usefulness of the HET. Conclusions: The HET is effective for identifying and classifying requirements errors and faults, thereby helping to improve the overall quality of the SRS and the software. Jeffrey C. Carver, Vaibhav K. Anu, Gursimran Singh Walia, Gary L. Bradshaw |
ESEM | 2 |
| 2016 | Using a Cognitive Psychology Perspective on Errors to Improve Requirements Quality: An Empirical InvestigationabstractSoftware inspections are an effective method for early detection of faults present in software development artifacts (e.g., requirements and design documents). However, many faults are left undetected due to the lack of focus on the underlying sources of faults (i.e., what caused the injection of the fault?). To address this problem, research work done by Psychologists on analyzing the failures of human cognition (i.e., human errors) is being used in this research to help inspectors detect errors and corresponding faults (manifestations of errors) in requirements documents. We hypothesize that the fault detection performance will demonstrate significant gains when using a formal taxonomy of human errors (the underlying source of faults). This paper describes a newly developed Human Error Taxonomy (HET) and a formal Error-Abstraction and Inspection (EAI) process to improve fault detection performance of inspectors during the requirements inspection. A controlled empirical study evaluated the usefulness of HET and EAI compared to fault based inspection. The results verify our hypothesis and provide useful insights into commonly occurring human errors that contributed to requirement faults along with areas to further refine both the HET and the EAI process. Vaibhav K. Anu, Gursimran Singh Walia, Jeffrey C. Carver, Gary L. Bradshaw |
ISSRE | 4 |
| 2016 | Effectiveness of Human Error Taxonomy during Requirements Inspection: An Empirical InvestigationabstractSoftware inspections are an effective method for achieving high quality software.We hypothesize that inspections focused on identifying errors (i.e., root cause of faults) are better at finding requirements faults when compared to inspection methods that rely on checklists created using lessons-learned from historical fault-data.Our previous work verified that, error based inspections guided by an initial requirements errors taxonomy (RET) performed significantly better than standard fault-based inspections.However, RET lacked an underlying human information processing model grounded in Cognitive Psychology research.The current research reports results from a systematic literature review (SLR) of Software Engineering and Cognitive Science literature -Human Error Taxonomy (HET) that contains requirements phase human errors.The major contribution of this paper is a report of control group study that compared the fault detection effectiveness and usefulness of HET with the previously validated RET.Results of this study show that subjects using HET were not only more effective at detecting faults, but they found faults faster.Post-hoc analysis of HET also revealed meaningful insights into the most commonly occurring human errors at different points during requirements development.The results provide motivation and feedback for further refining HET and creating formal inspection tools based on HET. Vaibhav K. Anu, Gursimran Singh Walia, Jeffrey C. Carver, Gary L. Bradshaw |
SEKE | 4 |
| 2016 | A (Updated) Review of Empiricism at the SIGCSE Technical SymposiumabstractThe computer science education (CSEd) research community consists of a large group of passionate CS educators who often contribute to other disciplines of CS research. There has been a trend in other disciplines toward more rigorous and empirical evaluation of various hypotheses. Prior investigations of the then-current state of CSEd research showed a distinct lack of rigor in the top research publication venues, with most papers falling in the general category of experience reports. In this paper, we present our examination of the two most recent proceedings of the SIGCSE Technical Symposium, providing a snapshot of the current state of empiricism at the largest CSEd venue. Our goal to categorize the current state of empiricism in the SIGCSE Technical Symposium and identify where the community might benefit from increased empiricism when conducting CSEd research. We found an increase in empirical validation of CSEd research to over 70%; however, our findings suggest that current CSEd research minimizes replication precluding meta-analysis and theory building. Ahmed Al-Zubidy, Jeffrey C. Carver, Sarah Smith Heckman, Mark Sherriff |
SIGCSE | 2 |
| 2016 | Code clones and developer behavior: results of two surveys of the clone research community
Debarshi Chatterji, Jeffrey C. Carver, Nicholas A. Kraft |
Empir. Softw. Eng. | 2 |
| 2016 | Identification of SLR tool needs - results of a community workshop
Edgar E. Hassler, Jeffrey C. Carver, David P. Hale, Ahmed Al-Zubidy |
Inf. Softw. Technol. | 2 |
| 2016 | A Multi-Site Joint Replication of a Design Patterns Experiment Using Moderator Variables to Generalize across ContextsabstractContext.Several empirical studies have explored the benefits of software design patterns, but their collective results are highly inconsistent. Resolving the inconsistencies requires investigating moderators—i.e., variables that cause an effect to differ across contexts.Objectives.Replicate a design patterns experiment at multiple sites and identify sufficient moderators to generalize the results across prior studies.Methods.We perform a close replication of an experiment investigating the impact (in terms of time and quality) of design patterns (Decorator and Abstract Factory) on software maintenance. The experiment was replicated once previously, with divergent results. We execute our replication at four universities—spanning two continents and three countries—using a new method for performing distributed replications based on closely coordinated, small-scale instances (“joint replication”). We perform two analyses: 1) apost-hocanalysis of moderators, based on frequentist and Bayesian statistics; 2) ana priorianalysis of the original hypotheses, based on frequentist statistics.Results.The main effect differs across the previous instances of the experiment and across the sites in our distributed replication. Our analysis of moderators (including developer experience and pattern knowledge) resolves the differences sufficiently to allow for cross-context (and cross-study) conclusions. The final conclusions represent 126 participants from five universities and 12 software companies, spanning two continents and at least four countries.Conclusions.The Decorator pattern is found to be preferable to a simpler solution during maintenance, as long as the developer has at least some prior knowledge of the pattern. For Abstract Factory, the simpler solution is found to be mostly equivalent to the pattern solution. Abstract Factory is shown to require a higher level of knowledge and/or experience than Decorator for the pattern to be beneficial. Jonathan L. Krein, Lutz Prechelt, Natalia Juristo Juzgado, Aziz Nanthaamornphong, Jeffrey C. Carver, Sira Vegas, Charles D. Knutson, Kevin D. Seppi, Dennis Eggett |
IEEE Trans. Software Eng. | 5 |
| 2015 | SE4HPCS'15: The 2015 International Workshop on Software Engineering for High Performance Computing in ScienceabstractHPC software is developed and used in a wide variety of scientific domains including nuclear physics, computational chemistry, crash simulation, satellite data processing, fluid dynamics, climate modeling, bioinformatics, and vehicle development. The increase in the importance of this software motivates the need to identify and understand appropriate software engineering (SE) practices for HPC architectures. Because of the variety of the scientific domains addressed using HPC, existing SE tools and techniques developed for the business/IT community are often not efficient or effective. Appropriate SE solutions must account for the salient characteristics of the HPC, research oriented development environment. This situation creates a need for members of the SE community to interact with members of the scientific and HPC communities to address this need. This workshop facilitates that collaboration by bringing together members of the SE, the scientific, and the HPC communities to share perspectives and present findings relevant to research, practice, and education. A significant portion of the workshop is devoted to focused interaction among the participants with the goal of generating a research agenda to improve tools, techniques, and experimental methods regarding SE for HPC science. Jeffrey C. Carver, Neil P. Chue Hong, Paolo Ciancarini |
ICSE (2) | 1 |
| 2015 | Workshop on Applications of Human Error Research to Improve Software Engineering (WAHESE 2015)abstractAdvances in the psychological understanding of the origins and manifestations of human error have led to tremendous reductions in errors in fields such as medicine, aviation, and nuclear power plants. This workshop is intended to foster a better understanding of software engineering errors and how a psychological perspective can reduce them, improving software quality and reducing maintenance costs. The workshop goal is to develop a body of knowledge that can advance our understanding of the psychological processes (of human reasoning, planning, and problem solving) and how they fail during the software development. Applying human error research to software quality improvement will provide insights to the cognitive aspects of software development. The workshop will include interactive session to discuss common themes of errors in different fields, and structure software error information to detect and prevent software errors during the development. Gursimran Singh Walia, Jeffrey C. Carver, Gary L. Bradshaw |
ICSE (2) | 2 |
| 2015 | Claims about the use of software engineering practices in science: A systematic literature review
Dustin Heaton, Jeffrey C. Carver |
Inf. Softw. Technol. | 2 |
| 2014 | Replication types: towards a shared taxonomyabstractContext: The software engineering community is becoming more aware of the need for experimental replications. In spite of the importance of this topic, there is still much inconsistency in the terminology used to describe replications. Maria Teresa Baldassarre, Jeffrey C. Carver, Óscar Dieste Tubío, Natalia Juristo Juzgado |
EASE | 2 |
| 2014 | Outcomes of a community workshop to identify and rank barriers to the systematic literature review processabstractSystematic Literature Reviews (SLRs) are an important tool used by software engineering researchers to summarize the state of knowledge about a particular topic. Currently, SLR authors must perform the difficult, time-consuming task in largely manual fashion. To identify barriers faced by SLR authors, we conducted an interactive community workshop prior to ESEM'13. Workshop participants generated a total of 100 ideas that, through group discussions, formed 37 composite barriers to the SLR process. Further analysis reveals the barriers relate to latent themes regarding the SLR process, primary studies, the practitioner community, and tooling. This paper describes the barriers identified during the workshop along with a ranking of those barriers that is based on votes by workshop attendees. The paper concludes by describing the impact of these barriers on three important constituencies: SLR Methodology Researchers, SLR Authors and SLR consumers. Edgar E. Hassler, Jeffrey C. Carver, Nicholas A. Kraft, David P. Hale |
EASE | 2 |
| 2014 | Impact of developer reputation on code review outcomes in OSS projects: an empirical investigationabstractContext: Gaining an identity and building a good reputation are important motivations for Open Source Software (OSS) developers. It is unclear whether these motivations have any actual impact on OSS project success. Goal: To identify how an OSS developer's reputation affects the outcome of his/her code review requests. Method: We conducted a social network analysis (SNA) of the code review data from eight popular OSS projects. Working on the assumption that core developers have better reputation than peripheral developers, we developed an approach, Core Identification using K-means (CIK) to divide the OSS developers into core and periphery groups based on six SNA centrality measures. We then compared the outcome of the code review process for members of the two groups. Results: The results suggest that the core developers receive quicker first feedback on their review request, complete the review process in shorter time, and are more likely to have their code changes accepted into the project codebase. Peripheral developers may have to wait 2 - 19 times (or 12 - 96 hours) longer than core developers for the review process of their code to complete. Conclusion: We recommend that projects allocate resources or create tool support to triage the code review requests to motivate prospective developers through quick feedback. Amiangshu Bosu, Jeffrey C. Carver |
ESEM | 2 |
| 2014 | Identifying the characteristics of vulnerable code changes: an empirical studyabstractTo focus the efforts of security experts, the goals of this empirical study are to analyze which security vulnerabilities can be discovered by code review, identify characteristics of vulnerable code changes, and identify characteristics of developers likely to introduce vulnerabilities. Using a three-stage manual and automated process, we analyzed 267,046 code review requests from 10 open source projects and identified 413 Vulnerable Code Changes (VCC). Some key results include: (1) code review can identify common types of vulnerabilities; (2) while more experienced contributors authored the majority of the VCCs, the less experienced contributors' changes were 1.8 to 24 times more likely to be vulnerable; (3) the likelihood of a vulnerability increases with the number of lines changed, and (4) modified files are more likely to contain vulnerabilities than new files. Knowing which code changes are more prone to contain vulnerabilities may allow a security expert to concentrate on a smaller subset of submitted code changes. Moreover, we recommend that projects should: (a) create or adapt secure coding guidelines, (b) create a dedicated security review team, (c) ensure detailed comments during review to help knowledge dissemination, and (d) encourage developers to make small, incremental changes rather than large changes. Amiangshu Bosu, Jeffrey C. Carver, Munawar Hafiz, Patrick Hilley, Derek Janni |
SIGSOFT FSE | 2 |
| 2014 | Investigation of individual factors impacting the effectiveness of requirements inspections: a replicated experiment
Özlem Albayrak, Jeffrey C. Carver |
Empir. Softw. Eng. | 2 |
| 2014 | Replications of software engineering experiments
Jeffrey C. Carver, Natalia Juristo Juzgado, Maria Teresa Baldassarre, Sira Vegas |
Empir. Softw. Eng. | 1 |
| 2014 | Examination of the software architecture change characterization scheme using three empirical studies
Byron J. Williams, Jeffrey C. Carver |
Empir. Softw. Eng. | 2 |
| 2014 | Peer impressions in open source organizations: A survey
Amiangshu Bosu, Jeffrey C. Carver, Rosanna E. Guadagno, Blake Bassett, Debra McCallum, Lorin Hochstein |
J. Syst. Softw. | 2 |
| 2013 | Specification and reasoning in SE projects using a Web IDEabstractA key goal of our research is to introduce an approach that involves at the outset using analytical reasoning as a method for developing high quality software. This paper summarizes our experiences in introducing mathematical reasoning and formal specification-based development using a web-integrated environment in an undergraduate software engineering course at two institutions at different levels, with the goal that they will serve as models for other educators. At Alabama, the reasoning topics are introduced over a two-week period and are followed by a project. At Clemson, the topics are covered in more depth over a five-week period and are followed by specification-based software development and reasoning assignments. The courses and project assignments have been offered for multiple semesters. Evaluation of student performance indicates that the learning goals were met. Charles T. Cook, Svetlana V. Drachova, Yu-Shan Sun, Murali Sitaraman, Jeffrey C. Carver, Joseph E. Hollingsworth |
CSEE&T | 5 |
| 2013 | Impact of Peer Code Review on Peer Impression Formation: A SurveyabstractPeer code review has been adopted as an effective quality improvement practice by many Open Source Software (OSS) communities. In addition to increasing software quality, there is anecdotal evidence that peer code review has other benefits, including: sharing knowledge, sharing expertise, sharing development techniques, and most importantly building accurate peer impressions between the code review participants. To further investigate the presence of these benefits, we surveyed members of popular OSS communities who were involved with peer code review. We used established scales from Psychology, Information science, and Organizational Behavior to create survey questions. We also enforced multiple reliability and validity measures to ensure higher confidence in the survey results. In this paper, we present a subset of the surveys results focused on better understanding four aspects of peer impression formation: trust, reliability, perception of expertise, and friendship. The results indicate that there is indeed a high level of trust, reliability, perception of expertise, and friendship between OSS peers who have participated in code review for a period of time. Because code review involves examining someone else's code, unsurprisingly, peer code review helped most in building a perception of expertise between code review partners. Amiangshu Bosu, Jeffrey C. Carver |
ESEM | 2 |
| 2013 | Identifying Barriers to the Systematic Literature Review ProcessabstractConducting a systematic literature review (SLR) is difficult and time-consuming for an experienced researcher, and even more so for a novice graduate student. With a better understanding of the most common difficulties in the SLR process, mentors will be better prepared to guide novices through the process. This understanding will help researchers have more realistic expectations of the SLR process and will help mentors guide novices through its planning, execution, and documentation phases. Consequently, the objectives of this work are to identify the most difficult and time-consuming phases of the SLR process. Using data from two sources - 52 responses to an online survey sent to all authors of SLRs published in software engineering venues and qualitative experience reports from 8 PhD students who conducted SLRs as part of a course - we identified specific difficulties related to each phase of the SLR process. Our findings highlight the importance of planning, teamwork, and mentoring by an experienced researcher throughout the process. The paper also identifies implications for the teaching of the SLR process. Jeffrey C. Carver, Edgar E. Hassler, Elis Hernandes, Nicholas A. Kraft |
ESEM | 1 |
| 2013 | 5th international workshop on software engineering for computational science and engineering (SE-CSE 2013)
Jeffrey C. Carver, Tom Epperly, Lorin Hochstein, Valerie Maxville, Dietmar Pfahl, Jonathan Sillito |
ICSE | 1 |
| 2013 | Evaluating source code summarization techniques: Replication and expansionabstractDuring software evolution a developer must investigate source code to locate then understand the entities that must be modified to complete a change task. To help developers in this task, Haiduc et al. proposed text summarization based approaches to the automatic generation of class and method summaries, and via a study of four developers, they evaluated source code summaries generated using their techniques. In this paper we propose a new topic modeling based approach to source code summarization, and via a study of 14 developers, we evaluate source code summaries generated using the proposed technique. Our study partially replicates the original study by Haiduc et al. in that it uses the objects, the instruments, and a subset of the summaries from the original study, but it also expands the original study in that it includes more subjects and new summaries. The results of our study both support the findings of the original and provide new insights into the processes and criteria that developers use to evaluate source code summaries. Based on our results, we suggest future directions for research on source code summarization. Brian P. Eddy, Jeffrey A. Robinson, Nicholas A. Kraft, Jeffrey C. Carver |
ICPC | 4 |
| 2013 | Building reputation in StackOverflow: an empirical investigationabstractStackOverflow (SO) contributors are recognized by reputation scores. Earning a high reputation score requires technical expertise and sustained effort. We analyzed the SO data from four perspectives to understand the dynamics of reputation building on SO. The results of our analysis provide guidance to new SO contributors who want to earn high reputation scores quickly. In particular, the results indicate that the following activities can help to build reputation quickly: answering questions related to tags with lower expertise density, answering questions promptly, being the first one to answer a question, being active during off peak hours, and contributing to diverse areas. Amiangshu Bosu, Christopher S. Corley, Dustin Heaton, Debarshi Chatterji, Jeffrey C. Carver, Nicholas A. Kraft |
MSR | 5 |
| 2013 | Using error abstraction and classification to improve requirement quality: conclusions from a family of four empirical studies
Gursimran Singh Walia, Jeffrey C. Carver |
Empir. Softw. Eng. | 2 |
| 2012 | Application of kusumoto cost-metric to evaluate the cost effectiveness of software inspectionsabstractInspections and testing are two widely recommended techniques for improving software quality. While testing cannot be conducted until software is implemented, inspections can help find and fix the faults right after their injection in the requirements and design documents. It is estimated that majority of testing cost is spent on fault rework and can be saved by inspections of early software products. However there is a lack of evidence regarding the testing costs saved by performing inspections. This research analyzes the costs and benefits of inspections and testing to decide on whether to schedule an inspection. We also analyzed the effect of the team size on the decision of how to organize the inspections. Another aspect of our research evaluates the use of Capture Recapture (CR) estimation method when the actual fault count of software product is unknown. Using data from 73 inspectors, we applied the Kusumoto metric to evaluate the cost-effectiveness of the inspections with varying team size. Our results provide a detailed analysis of the number of inspectors required for varying levels of cost-effectiveness during inspections; and the number of inspectors required by the CR estimators to provide estimates within 5% to 20% of the actual. Narendar Mandala, Gursimran Singh Walia, Jeffrey C. Carver, Nachiappan Nagappan |
ESEM | 3 |
| 2012 | Teaching mathematical reasoning across the curriculumabstractNo abstract available. Joan Krone, Douglas Baldwin, Jeffrey C. Carver, Joseph E. Hollingsworth, Amruth N. Kumar, Murali Sitaraman |
SIGCSE | 3 |
| 2012 | Program comprehension of domain-specific and general-purpose languages: comparison using a family of experiments
Tomaz Kosar, Marjan Mernik, Jeffrey C. Carver |
Empir. Softw. Eng. | 3 |
| 2012 | PPModel: a modeling tool for source code maintenance and optimization of parallel programs
Ferosh Jacob, Jeffrey G. Gray, Jeffrey C. Carver, Marjan Mernik, Purushotham V. Bangalore |
J. Supercomput. | 3 |
| 2011 | Evaluating the testing ability of senior-level computer science studentsabstractTesting is a key skill for computer science students to acquire during their studies. To determine how well students are learning this skill, we conducted an empirical study in two offerings of a senior-level computer science course. The goal of the study was to determine whether students would be able to create a small, complete test suite for a simple program. The students created a test suite first without the aid of a coverage tool and then with the aid of a coverage tool. The results indicate that without a coverage tool, students achieved significantly less than 100% statement, branch or condition coverage. When provided with a code coverage tool, students increased coverage levels. Still, examination of the test suites indicated that they were significantly larger than the minimum required. These results indicate that students cannot conduct adequate testing of even a small program. To provide context for our results, we provide a literature survey summarizing various techniques proposed for teaching testing in the computer science curriculum. We discuss each technique, its strengths, and its weaknesses. Jeffrey C. Carver, Nicholas A. Kraft |
CSEE&T | 1 |
| 2011 | Measuring the Efficacy of Code Clone Information in a Bug Localization Task: An Empirical StudyabstractMuch recent research effort has been devoted to designing efficient code clone detection techniques and tools. However, there has been little human-based empirical study of developers as they use the outputs of those tools while performing maintenance tasks. This paper describes a study that investigates the usefulness of code clone information for performing a bug localization task. In this study 43 graduate students were observed while identifying defects in both cloned and non-cloned portions of code. The goal of the study was to understand how those developers used clone information to perform this task. The results of this study showed that participants who first identified a defect then used it to look for clones of the defect were more effective than participants who used the clone information before finding any defects. The results also show a relationship between the perceived efficacy of the clone information and effectiveness in finding defects. Finally, the results show that participants who had industrial experience were more effective in identifying defects than those without industrial experience. Debarshi Chatterji, Jeffrey C. Carver, Beverly Massengill, Jason Oslin, Nicholas A. Kraft |
ESEM | 2 |
| 2011 | Fourth international workshop on software engineering for computational science and engineering: (SE-CSE2011)abstractComputational Science and Engineering (CSE) software supports a wide variety of domains including nuclear physics, crash simulation, satellite data processing, fluid dynamics, climate modeling, bioinformatics, and vehicle development. The increase in the importance of CSE software motivates the need to identify and understand appropriate software engineering (SE) practices for CSE. Because of the uniqueness of CSE software development, existing SE tools and techniques developed for the business/IT community are often not efficient or effective. Appropriate SE solutions must account for the salient characteristics of the CSE development environment. This situation creates an opportunity for members of the SE community to interact with members of the CSE community to address this need. This workshop facilitates that collaboration by bringing together members of the SE community and the CSE community to share perspectives and present findings from research and practice relevant to CSE software. A significant portion of the workshop is devoted to focused interaction among the participants with the goal of generating a research agenda to improve tools, techniques, and experimental methods for studying CSE software engineering. Jeffrey C. Carver, Roscoe A. Bartlett, Ian Gorton, Lorin Hochstein, Diane Kelly 0002, Judith Segal |
ICSE | 1 |
| 2010 | Evaluating the Use of Requirement Error Abstraction and Classification Method for Preventing Errors during Artifact Creation: A Feasibility StudyabstractDefect prevention techniques can be used during the creation of software artifacts to help developers create high-quality artifacts. These artifacts should have fewer faults that must be removed during inspection and testing. The Requirement Error Taxonomy that we have developed helps focus developers' attention on common errors that can occur during requirements engineering. Our claim is that, by focusing on those errors, the developers will be less likely to commit them. This paper investigates the usefulness of the Requirement Error Taxonomy as a defect prevention technique. The goal was to determine if making requirements engineers' familiar with the Requirement Error Taxonomy would reduce the likelihood that they commit errors while developing a requirements document. We conducted an empirical study in which the participants were given the opportunity to learn how to use the Requirement Error Taxonomy by employing it during the inspection of a requirements document. Then, in teams of four, they developed their own requirements document. This requirements document was then evaluated by other students to identify any errors made. The hypothesis was that participants who find more errors during the inspection of a requirements document would make fewer errors when creating their own requirements document. The overall result supports this hypothesis. Gursimran Singh Walia, Jeffrey C. Carver |
ISSRE | 2 |
| 2010 | A checklist for integrating student empirical studies with research and teaching goals
Jeffrey C. Carver, Letizia Jaccheri, Sandro Morasca, Forrest Shull |
Empir. Softw. Eng. | 1 |
| 2010 | Characterizing software architecture changes: A systematic review
Byron J. Williams, Jeffrey C. Carver |
Inf. Softw. Technol. | 2 |
| 2009 | Modifiability measurement from a task complexity perspective: A feasibility studyabstractDespite the critical role of software modifiability, it has no universally accepted measurement model. Measuring modifiability in terms of maintenance effort is problematic because it confounds modifiability with the ability of individual maintainers. In this paper, we apply Wood's task complexity model to propose a general analytical model that describes the characteristics of maintenance tasks and the analytical dimensions of modifiability independent of the individual maintainers. The results of a case study demonstrate the construct validity of the model. Lulu He, Jeffrey C. Carver |
ESEM | 2 |
| 2009 | Cognitive factors in perspective-based reading (PBR): A protocol analysis studyabstractThe following study investigated cognitive factors involved in applying the Perspective-Based Reading (PBR) technique for defect detection in software inspections. Using the protocol analysis technique from cognitive science, the authors coded concurrent verbal reports from novice reviewers and used frequency-based analysis to consider existing research on cognition in software inspections from within a cognitive framework. The current coding scheme was able to describe over 98% of the cognitive activities reported during inspection at a level of detail capable of validating multiple hypotheses from literature. A number of threats to validity are identified for the protocol analysis method and the parameters of the current experiment. The authors conclude that protocol analysis is a useful tool for analyzing cognitively intense software engineering tasks such as software inspections. Bryan Robbins, Jeffrey C. Carver |
ESEM | 2 |
| 2009 | Evaluating the Effect of the Number of Naturally Occurring Faults on the Estimates Produced by Capture-Recapture ModelsabstractProject managers can use the capture-recapture models to estimate the number of faults in a software artifact. The capture-recapture estimates are calculated using the number of unique faults and the number of times each fault is found. The accuracy of the estimates is affected by the number of inspectors and the number of faults. Our earlier research investigated the effect that the number of inspectors had on the accuracy of the estimates. In this paper, we investigate the effect of the number of faults on the performance of the estimates using real requirement artifacts. These artifacts have an unknown amount of naturally occurring faults. The results show that while the estimators generally underestimate, they improve as the number of faults increases. The results also show that the capture-recapture estimators can be used to make correct re-inspection decisions. Gursimran Singh Walia, Jeffrey C. Carver |
ICST | 2 |
| 2009 | A visual analytic framework for exploring relationships in textual contents of digital forensics evidenceabstractWe describe the development of a set of tools for analyzing the textual contents of digital forensic evidence for the purpose of enhancing an investigator's ability to discover information quickly and efficiently. By examining the textual contents of files and unallocated space, relationships between sets of files and clusters can be formed based on the information that they contain. Using the information gathered from the evidence through the analysis tool, the visualization tool can be used to search through the evidence in an organized and efficient manner. The visualization depicts both the frequency of relevant terms and their location on disk. We also discuss a task analysis with forensics officers to motivate the design. T. J. Jankun-Kelly, David Wilson 0003, Andrew S. Stamps, Josh Franck, Jeffrey C. Carver, J. Edward Swan II |
VizSEC | 5 |
| 2009 | A systematic literature review to identify and classify software requirement errors
Gursimran Singh Walia, Jeffrey C. Carver |
Inf. Softw. Technol. | 2 |
| 2008 | Evaluation of capture-recapture models for estimating the abundance of naturally-occurring defectsabstractProject managers can use capture-recapture models to manage the inspection process by estimating the number of defects present in an artifact and determining whether a reinspection is necessary. Researchers have previously evaluated capture-recapture models on artifacts with a known number of defects. Before applying capture-recapture models in real development, an evaluation of those models on naturally-occurring defects is imperative. The data in this study is drawn from two inspections of real requirements documents (that later guided implementation) created as part of a capstone course (i.e. with naturally occurring defects). The major results show that: a) estimators improve from being negatively biased after one inspection to being positively biased after two inspections, b) the results contradict the earlier result that a model that includes two sources of variation is a significant improvement over models with one source of variation, and c) estimates are useful in determining the need for artifact reinspection. Gursimran Singh Walia, Jeffrey C. Carver |
ESEM | 2 |
| 2008 | A Framework for Software Engineering Experimental ReplicationsabstractExperimental replications are very important to the advancement of empirical software engineering. Replications are one of the key mechanisms to confirm previous experimental findings. They are also used to transfer experimental knowledge, to train people, and to expand a base of experimental evidence. Unfortunately, experimental replications are difficult endeavors. It is not easy to transfer experimental know-how and experimental findings. Based on our experience, this paper discusses this problem and proposes a Framework for Improving the Replication of Experiments (FIRE). The FIRE addresses knowledge sharing issues both at the intra-group (internal replications) and inter-group (external replications) levels. It encourages coordination of replications in order to facilitate knowledge transfer for lower cost, higher quality replications and more generalizable results. Manoel G. Mendonça, José Carlos Maldonado, Maria Cristina Ferreira de Oliveira, Jeffrey C. Carver, Sandra C. P. F. Fabbri, Forrest Shull, Guilherme Horta Travassos, Erika Nina Höhn, Victor R. Basili |
ICECCS | 4 |
| 2008 | The effect of the number of inspectors on the defect estimates produced by capture-recapture modelsabstractInspections can be made more cost-effective by using capture-recapture methods to estimate post-inspection defects. Previous capture-recapture studies of inspections used relatively small data sets compared with those used in biology and wildlife research (the origin of the models). A common belief is that capture-recapture models underestimate the number of defects but their performance can be improved with data from more inspectors. This increase has not been evaluated in detail. This paper evaluates new estimators from biology not been previously applied to inspections. Using a data from seventy-three inspectors, we analyze the effect of the number of inspectors on the quality of estimates. Contrary to previous findings indicating that Jackknife is the best estimator, our results show that the SC estimators are better suited to software inspections. Our results also provide a detailed analysis of the number of inspectors necessary to obtain estimates within 5% to 20% of the actual. Gursimran Singh Walia, Jeffrey C. Carver, Nachiappan Nagappan |
ICSE | 2 |
| 2008 | The Effect of the Number of Defects on Estimates Produced by Capture-Recapture ModelsabstractProject managers use inspection data as input to capture-recapture (CR) models to estimate the total number of faults present in a software artifact. The CR models use the number of faults found during an inspection and the overlap of faults among inspectors to calculate the estimate. A common belief is that CR models underestimate the number of faults but their performance can be improved with more input data. This paper investigates the minimum number of faults that has to be present in an artifact before the CR method can be used. The result shows that the minimum number of faults varies from ten faults to twenty-three faults for different CR estimators. Gursimran Singh Walia, Jeffrey C. Carver |
ISSRE | 2 |
| 2008 | Show Me How You See: Lessons from Studying Computer Forensics Experts for Visualization
T. J. Jankun-Kelly, Josh Franck, David Wilson 0003, Jeffrey C. Carver, David A. Dampier, J. Edward Swan II |
VizSEC | 4 |
| 2008 | The role of replications in Empirical Software Engineering
Forrest Shull, Jeffrey C. Carver, Sira Vegas, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 2 |
| 2008 | The Impact of Educational Background on the Effectiveness of Requirements Inspections: An Empirical StudyabstractWhile the inspection of various software artifacts increases the quality of the end product, the effectiveness of an inspection depends largely on the individual inspectors involved. To address that issue, a large-scale controlled inspection experiment with over 70 professionals was conducted at Microsoft Corporation that focused on the relationship between an inspector's background and their effectiveness during a requirements inspection. The results of the study showed that inspectors with university degrees in majors not related to computer science found significantly more defects than those with degrees in computer science majors. We also observed that level of education (Masters, PhD), prior industrial experience or other job related experiences did not significantly impact the effectiveness of an inspector. The only other type of experience that had a significant impact on effectiveness was experience in writing requirements, i.e. professionals with prior experience writing requirements found statistically significant more defects than their counterparts. Jeffrey C. Carver, Nachiappan Nagappan, Alan Page |
IEEE Trans. Software Eng. | 1 |
| 2007 | Increased Retention of Early Computer Science and Software Engineering Students Using Pair ProgrammingabstractAn important problem faced by many Computer Science and Software Engineering programs is declining enrollment. In an effort to reverse that trend at Mississippi State University, we have instituted pair programming for the laboratory exercises in the introductory programming course. This paper describes a study performed to analyze whether using pair programming would increase retention. An important goal of this study was not only to measure increased retention, but to provide insight into why retention increased or decreased. The results of the study showed that retention significantly increased for those students already majoring in Computer Science, Software Engineering, or Computer Engineering. In addition, survey results indicated that the students viewed many aspects of pair programming to be very beneficial to their learning experience. Jeffrey C. Carver, Lisa Henderson, Lulu He, Julia E. Hodges, Donna S. Reese |
CSEE&T | 1 |
| 2007 | An Empirical Study of the Effects of Gestalt Principles on Diagram UnderstandabilityabstractComprehension errors in software design must be detected at their origin to avoid propagation into later portions of the software lifecycle and also the final system. This research synthesizes software engineering and Gestalt principles of similarity, proximity, continuity for the purpose of discovering whether certain visual attributes of diagrams can affect the accuracy and efficiency of understanding the diagram. The experiment tested whether two dependent variables, accuracy and response time, were significantly affected by independent variables, diagram type (simple 1, simple2, complex), Gestalt principles (good vs. bad), and question order (forward/backward). The results of this study indicated that the Gestalt principles did affect the comprehension in the complex diagrams. Post-hoc analysis results indicated that number of bends per line, length of line in inches, number of lines crossing, boxes per diagram, and number of lines per diagram contributed to the ability of the subjects to comprehend the diagrams. Krystle Lemon, Edward B. Allen, Jeffrey C. Carver, Gary L. Bradshaw |
ESEM | 3 |
| 2007 | Characterizing Software Architecture Changes: An Initial StudyabstractWith today's ever increasing demands on software, developers must produce software that can be changed without the risk of degrading the software architecture. Degraded software architecture is problematic because it makes the system more prone to defects and increases the cost of making future changes. The effects of making changes to software can be difficult to measure. One way to address software changes is to characterize their causes and effects. This paper introduces an initial architecture change characterization scheme created to assist developers in measuring the impact of a change on the architecture of the system. It also presents an initial study conducted to gain insight into the validity of the scheme. The results of this study indicated a favorable view of the viability of the scheme by the subjects, and the scheme increased the ability of novice developers to assess and adequately estimate change effort. Byron J. Williams, Jeffrey C. Carver |
ESEM | 2 |
| 2007 | Software Development Environments for Scientific and Engineering Software: A Series of Case StudiesabstractThe need for high performance computing applications for computational science and engineering projects is growing rapidly, yet there have been few detailed studies of the software engineering process used for these applications. The DARPA High Productivity Computing Systems Program has sponsored a series of case studies of representative computational science and engineering projects to identify the steps involved in developing such applications (i.e. the life cycle, the workflows, technical challenges, and organizational challenges). Secondary goals were to characterize tool usage and identify enhancements that would increase the programmers' productivity. Finally, these studies were designed to develop a set of lessons learned that can be transferred to the general computational science and engineering community to improve the software engineering process used for their applications. Nine lessons learned from five representative projects are presented, along with their software engineering implications, to provide insight into the software development environments in this domain. Jeffrey C. Carver, Richard P. Kendall, Susan E. Squires, Douglass E. Post |
ICSE | 1 |
| 2007 | Requirement Error Abstraction and Classification: A Control Group Replicated StudyabstractThis paper is the second in a series of empirical studies about requirement error abstraction and classification as a quality improvement approach. The Requirement error abstraction and classification method supports the developers' effort in efficiently identifying the root cause of requirements faults. By uncovering the source of faults, the developers can locate and remove additional related faults that may have been overlooked, thereby improving the quality and reliability of the resulting system. This study is a replication of an earlier study that adds a control group to address a major validity threat. The approach studied includes a process for abstracting errors from faults and provides a requirement error taxonomy for organizing those errors. A unique aspect of this work is the use of research from human cognition to improve the process. The results of the replication are presented and compared with the results from the original study. Overall, the results from this study indicate that the error abstraction and classification approach improves the effectiveness and efficiency of inspectors. The requirement error taxonomy is viewed favorably and provides useful insights into the source of faults. In addition, human cognition research is shown to be an important factor that affects the performance of the inspectors. This study also provides additional evidence to motivate further research. Gursimran Singh Walia, Jeffrey C. Carver, Thomas Philip |
ISSRE | 2 |
| 2006 | Viope as a Tool for Teaching Introductory Programming: An Empirical InvestigationabstractIn this paper we describe the use of a tool from Viope for teaching introductory programming. We have noticed in our previous courses that the students often have trouble connecting the small classroom exercises with the larger laboratory projects. This tool allows the students to get extra practice with those concepts to help ensure they are understood. In this study data was collected using a survey. We surveyed students at the end of a semester in which the tool was not used to gather information about where the tool might be useful. Then we had students in a second semester use the tool and complete a similar survey. The results of our study showed that while there were some consistent complaints about the tool, overall the students found it useful enough to indicate they would like to use something similar in later semesters. 1. Jeffrey C. Carver, Lisa Henderson |
CSEE&T | 1 |
| 2006 | Can observational techniques help novices overcome the software inspection learning curve? An empirical investigation
Jeffrey C. Carver, Forrest Shull, Victor R. Basili |
Empir. Softw. Eng. | 1 |
| 2006 | Perspective-Based Reading: A Replicated Experiment Focused on Individual Reviewer Effectiveness
José Carlos Maldonado, Jeffrey C. Carver, Forrest Shull, Sandra C. P. F. Fabbri, Emerson Dória, Luciana Andréia Fondazzi Martimiano, Manoel G. Mendonça, Victor R. Basili |
Empir. Softw. Eng. | 2 |
| 2005 | Parallel Programmer Productivity: A Case Study of Novice Parallel ProgrammersabstractIn developing High-Performance Computing (HPC) software, time to solution is an important metric. This metric is comprised of two main components: the human effort required developing the software, plus the amount of machine time required to execute it. To date, little empirical work has been done to study the first component: the human effort required and the effects of approaches and practices that may be used to reduce it. In this paper, we describe a series of studies that address this problem. We instrumented the development process used in multiple HPC classroom environments. We analyzed data within and across such studies, varying factors such as the parallel programming model used and the application being developed, to understand their impact on the development process. Lorin Hochstein, Jeffrey C. Carver, Forrest Shull, Sima Asgari, Victor R. Basili |
SC | 2 |
| 2005 | Combining self-reported and automatic data to improve programming effort measurementabstractMeasuring effort accurately and consistently across subjects in a programming experiment can be a surprisingly difficult task. In particular, measures based on self-reported data may differ significantly from measures based on data which is recorded automatically from a subject's computing environment. Since self-reports can be unreliable, and not all activities can be captured automatically, a complete measure of programming effort should incorporate both classes of data. In this paper, we show how self-reported and automatic effort can be combined to perform validation and to measure total programming effort. Lorin Hochstein, Victor R. Basili, Marvin V. Zelkowitz, Jeffrey K. Hollingsworth, Jeffrey C. Carver |
ESEC/SIGSOFT FSE | 5 |
| 2004 | The Impact of Background and Experience on Software Inspections
Jeffrey C. Carver |
Empir. Softw. Eng. | 1 |
| 2004 | Knowledge-Sharing Issues in Experimental Software Engineering
Forrest Shull, Manoel G. Mendonça, Victor R. Basili, Jeffrey C. Carver, José Carlos Maldonado, Sandra C. P. F. Fabbri, Guilherme Horta Travassos, Maria Cristina Ferreira de Oliveira |
Empir. Softw. Eng. | 4 |
| 2001 | An empirical methodology for introducing software processesabstractThere is a growing interest in empirical study in software engineering, both for validating mature technologies and for guiding improvements of less-mature technologies. This paper introduces an empirical methodology, based on experiences garnered over more than two decades of work by the Empirical Software Engineering Group at the University of Maryland and related organizations, for taking a newly proposed improvement to development processes from the conceptual phase through transfer to industry. The methodology presents a series of questions that should be addressed, as well as the types of studies that best address those questions. The methodology is illustrated by a specific research program on inspection processes for Object-Oriented designs. Specific examples of the studies that were performed and how the methodology impacted the development of the inspection process are also described. Forrest Shull, Jeffrey C. Carver, Guilherme Horta Travassos |
ESEC / SIGSOFT FSE | 2 |