David C. Shepherd

dblp:153/5343 · DBLP profile ↗
← Back
55ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0003-2017-7842ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 36 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 17 · 13 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Towards ecological validity when assessing ADHD symptoms: Patterns in automatically collected, real-world PC activity data
abstract
Neuropsychological tests assessing attention and executive function (EF) in individuals with ADHD demonstrate little to no association with real-world ratings of ADHD behaviors. To address this critical gap, this study developed metrics that can analyze automatically collected data to measure levels of attention, motivation, and effort in emerging adults with ADHD. Specifically, we used virtual reality to simulate a study space and collect in-the-moment computer data activity while university students with ADHD (N = 21; 38% female) engaged in 12 sessions (total 180 hours) of real-world tasks. To identify common sequences we performed a qualitative analysis of this work session data (i.e., descriptive window titles, input levels, and window switches), resulting in four themes representing positive and negative work activity patterns. From these themes we derived four metrics, and a quantitative analyses showed that two predicted behavioral indices of attention, effort, and motivation with effects in the moderate range. To our knowledge, we are the first group to design and test such an approach, as well as validate identified computer metrics to behavioral indices of attention and EF. Given the automated nature of computer data collection and analysis, this approach represents a scalable, novel method for ADHD assessment and treatment.
Matheus B. da Costa, Elizabeth Chan, Joshua M. Langberg, Isabelle Cuber, Fatemeh Jamalinabijan, Aleksander Kurgan, Thomas Fritz 0001, David C. Shepherd
Int. J. Hum. Comput. Stud.8
2025 Block-based or graph-based? Why not both? Designing a hybrid programming environment for end-users
abstract
Abstract End-user programmers need programming tools that are easy to learn and use. Development environments for end-users often support one of two visual modalities: block-based programming or data-flow programming. In this work, we discuss differences in how these modalities represent programs, and why existing block-based programming tools are better suited for imperative tasks while data-flow programming better supports nested expressions. We focus on robot programming as an end-user scenario that requires both imperative and expressions-based code in the same program. To study how end-user tools can better support this scenario, we propose two programming system designs: one that changes how blocks represent nested expressions, and one that combines block-based and data-flow programming in the same hybrid environment. We compared these designs in a controlled experiment with 113 end-user participants who solved programming and program comprehension tasks using one of the two environments. Both groups indicated a small preference for the hybrid system in direct comparison, but participants who used blocks to solve tasks performed better on average than hybrid system users and gave higher usability ratings. These findings suggest that despite the appeal of data-flow programming, a well-adapted block-based programming interface can lead end-users to more programming success.
Nico Ritschel, Reid Holmes, Felipe Fronchetti, Ronald Garcia, David C. Shepherd
Interact. Comput.5
2024 Examining the Use of VR as a Study Aid for University Students with ADHD
abstract
Attention-deficit/hyperactivity disorder (ADHD) is a neurodevelopmental condition characterized by patterns of inattention and impulsivity, which lead to difficulties maintaining concentration and motivation while completing academic tasks. University settings, characterized by a high student-to-staff ratio, make treatments relying on human monitoring challenging. One potential replacement is Virtual Reality (VR) technology, which has shown potential to enhance learning outcomes and promote flow experience. In this study, we investigate the usage of VR with 27 university students with ADHD in an effort to improve their performance in ctableompleting homework, including an exploration of automated feedback via a technology probe. Quantitative results show significant increases in concentration, motivation, and effort levels during these VR sessions and qualitative data offers insight into considerations like comfort and deployment. Together, the results suggest that VR can be a valuable tool in leveling the playing field for university students with ADHD.
Isabelle Cuber, Juliana Gonçalves de Souza, Irene Jacobs, Caroline Lowman, David C. Shepherd, Thomas Fritz 0001, Joshua M. Langberg
CHI5
2024 An Exploratory Comparative Study on the Impacts of Technical Support on Student Successes in Computing Project-Based Learning
abstract
In this research paper, we conducted a comparative study to measure the effectiveness of the provided technical support in computing project-based learning (PjBL) courses. Students learn much better by solving authentic real-world problems through PjBL. PjBL in computing education has proven to boost student motivation and engagement while enhancing academic performance. Crucial to PjBL in computing is the technical support that the instructors can provide to students, which is required for sustained, successful learning during project tasks. Without adequate support, PjBL will fall short of accomplishing its goals, leading to a rise in student frustration, a loss of motivation and engagement, and compromised learning outcomes. Measuring the impacts and effectiveness of the provided support is imperative for fostering continuous improvement, informed decision-making, and student success. It enables instructors to assess the impacts of their strategies, improve their approaches, and utilize their resources more effectively. To measure the impacts of technical support on students during PjBL, we performed a comparative study on two undergraduate computing courses in Spring 2024, Fundamentals of Software Engineering and Database Systems. In both courses, students work on two assigned projects, one with little and inadequate support and the other with adequate support. We administered a post-survey after each project was completed. We analyzed students' selfreflection responses across four sub-scales, support satisfaction, motivation, self-efficacy, and project satisfaction. The results show a statistically significant increase in the supported project in the Fundamentals of Software Engineering course and no difference in the Database Systems course. This finding is likely due to other differences between the two projects for the Database Systems course beyond support, such as project scale. Qualitative analysis of students' responses also indicates the need for support by students in the less supported projects. Based on our experience, we reflect on the question of what would constitute a good design for studies that seek to compare two different student learning experiences.
Ahmad Daudu Suleiman, Jan E. DeWaters, Daqing Hou, Yu Liu 0037, David C. Shepherd
FIE5
2024 Providing Technical Support to Sustain Student Motivation and Engagement in Software Engineering Project-Based Learning
abstract
In this research paper, drawing from our own and other computing instructors' experiences, we highlight common technical challenges faced by students in software engineering project-based learning (PjBL) and discuss ways in which instructors can support students in overcoming them so that motivation is summoned and sustained. Through the use of practical hands-on experiences, PjBL has been shown to be an effective educational approach. However, unless projects are intentionally designed and supported in a way that summons and sustains student motivation, PjBL is likely to fail to accomplish its goals. Several factors influence student motivation, including their perception of the project's value and how confident they are in their ability to complete it. In particular, challenges that students perceive as insurmountable during the project can significantly weaken their motivation. On the other hand, supporting students to overcome such hurdles can be troublesome, especially in large classes as well as classes with diversity in student backgrounds. To generalize from our own experience, we designed a questionnaire targeted at PjBL computing instructors that contained closed and open questions on technical challenges faced by students, support instructors provided to overcome such challenges, and lessons learned by instructors on the effectiveness of their support. A total of 47 responses were collected from instructors with diverse backgrounds in terms of courses taught, students' years, and class sizes. We categorized the technical challenges into three main categories, namely (a) challenges in installing and configuring software packaged tools, (b) lack of prerequisite knowledge, and (c) challenges while completing project tasks. In this paper, we present the survey results from the three categories of technical challenges, their frequencies, importance, and effective support strategies instructors use to alleviate them.
Ahmad Daudu Suleiman, David C. Shepherd, Jan E. DeWaters, Yu Liu 0037, Daqing Hou
FIE2
2024 Block-based Programming for Two-Armed Robots: A Comparative Study
abstract
Programming industrial robots is difficult and expensive. Although recent work has made substantial progress in making it accessible to a wider range of users, it is often limited to simple programs and its usability remains untested in practice. In this article, we introduce Duplo, a block-based programming environment that allows end-users to program two-armed robots and solve tasks that require coordination. Duplo positions the program for each arm side-by-side, using the spatial relationship between blocks from each program to represent parallelism in a way that end-users can easily understand. This design was proposed by previous work, but not implemented or evaluated in a realistic programming setting. We performed a randomized experiment with 52 participants that evaluated Duplo on a complex programming task that contained several sub-tasks. We compared Duplo with RobotStudio Online YuMi, a commercial solution, and found that Duplo allowed participants to solve the same task faster and with greater success. By analyzing the information collected during our user study, we further identified factors that explain this performance difference, as well as remaining barriers, such as debugging issues and difficulties in interacting with the robot. This work represents another step towards allowing a wider audience of non-professionals to program, which might enable the broader deployment of robotics.
Felipe Fronchetti, Nico Ritschel, Logan Schorr, Chandler Barfield, Gabriella Chang, Rodrigo O. Spínola, Reid Holmes, David C. Shepherd
ICSE8
2024 Dear researchers - a new column sharing the perspective of software practitioners
Paris Avgeriou, David C. Shepherd
J. Syst. Softw.2
2023 The 2023 DREE Workshop on Designing and Running Project-Based Courses in Software Engineering Education
abstract
In this workshop, we introduce participants to the accomplishments and lessons learned from our ongoing NSF IUSE education research project, which is focused on supporting undergraduate project-based learning in computing education by developing and piloting a set of scaffolded course projects. The workshop has two main goals. One is to facilitate exchange of experiences on project-based learning among workshop participants. The other is to encourage adoption of the developed course projects by the broader computing education community.
Daqing Hou, Jan E. DeWaters, Mary Margaret Small, Yu Liu 0037, David C. Shepherd
FIE5
2023 The Importance of Project-Scale Scaffolding for Retention and Experience in Computing Courses
abstract
Teaching students complex problem-solving skills using large-scale, real-world problems is challenging for both students and teachers alike. As a result, most courses use small, well-specified, toy-like problems, which are not representative of what students will encounter in the workforce. One approach that allows teachers to use large-scale problems in class is by introducing scaffolding. Scaffolding breaks a larger problem into smaller steps, which students can solve independently, while deemphasizing tangential concepts such as the complex configuration files needed to compile open-source software systems. Strong scaffolding supports student learning, preventing them from getting bogged down with unnecessary tasks or overwhelmed by complexity. This work investigates a scaffolded problem-based-learning module for computing courses, using a realistically-sized project with characteristics representative of the industry. The project was implemented in a computer science course with roughly 100 students, and the results speak to the importance of scaffolding for student success. In fact, there were two student assignments that lacked sufficient scaffolding, compared with other tasks, and the reduction in student scoring and persistence shows that project scaffolding is necessary when implementing these types of assignments. Most students felt the project helped prepare them for a job in their chosen field.
Juliana Gonçalves de Souza, Mikaila Flavell, Ahmad Daudu Suleiman, David C. Shepherd, Jan E. DeWaters, Mary Margaret Small, Yu Liu 0037, Daqing Hou
FIE4
2023 Mapping Learning Objectives of Project-Based Undergraduate Software Engineering Courses to CC2020 Competency Model
abstract
This qualitative research performs a thematic analysis of the learning objectives in existing project-based undergraduate software engineering courses to align them with the competency model defined in the Computing Curricula 2020 reports (CC2020). This study identifies the trends, strengths, and gaps in how the reviewed course learning objectives cover the knowledge, skill, and disposition components of the CC2020 competency model. The learning objectives were categorized according to knowledge elements, skills, and dispositions as defined in the CC2020 competency model. Our analysis shows that 54% of knowledge elements from the reviewed learning objectives do not have any skill level specified and overall, only two out of the eleven dispositions in CC2020 are specified (“Collaboration” and “Professional”). We also find that technical knowledge elements from the software development category (e.g., software process, software design, and software quality, verification & validation) and systems modeling category (e.g., systems analysis & design, and requirements analysis and specification), probably unsurprisingly, are covered the most often. Similarly, collaboration & teamwork, and oral & written communication are unsurprisingly the most common professional & foundational knowledge elements in the reviewed course's learning objectives as they are essential to project-based learning. Although they are essential for the completion of a successful software project, knowledge elements such as time management, security technology & implementation, and user experience design are rarely mentioned. We discuss the implications of our findings on course design.
Ahmad Daudu Suleiman, Daqing Hou, Yu Liu 0037, Jan E. DeWaters, Mary Margaret Small, Juliana Gonçalves de Souza, David C. Shepherd
FIE7
2023 Using Domain-Specific, Immediate Feedback to Support Students Learning Computer Programming to Make Music
abstract
Broadening participation in computer science has been widely studied, creating many different techniques to attract, motivate, and engage students. A common meta-strategy is to use an outside domain as a hook, using the concepts in that domain to teach computer science. These domains are selected to interest the student, but students often lack a strong background in these domains. Therefore, a strategy designed to increase students' interest, motivation, and engagement could actually create more barriers for students, who now are faced with learning two new topics. To reduce this potential barrier in the domain of music, this paper presents the use of automated, immediate feedback during programming activities at a summer camp that uses music to teach foundational programming concepts. The feedback guides students musically, correcting notes that are out-of-key or rhythmic phrases that are too long or short, allowing students to focus their learning on the computer science concepts. This paper compares the correctness of students that received automated feedback with students that did not, which shows the effectiveness of the feedback. Follow up focus groups with students confirmed this quantitative data, with students claiming that the feedback was not only useful but that the activities would be much more challenging without the feedback.
Douglas Lusa Krug, Chrystalla Mouza, Taylor Barnett, Lori L. Pollock, David C. Shepherd
ITiCSE (1)6
2023 Attracting Adults to Computer Programming via Hip Hop
abstract
The demand for qualified computing professionals is high, with thousands of positions remaining unfilled each year. To create more qualified professionals, initiatives to attract and engage students in computer science have been proposed, but they tend to concentrate on primary, secondary (K-12), and post-secondary (college) level. With many adults looking for better career opportunities, it is surprising that few computer science initiatives focus on attracting adult learners to the field. This paper presents the results of an informal computer programming course that teaches the foundational concepts of computer programming to adults as they program hip-hop beats. This course is designed to attract adult learners that otherwise might have never considered computer programming, building their confidence and skills. We conducted this course online, two nights a week, for five weeks, for about 40 participants. Afterward, we conducted a qualitative analysis of written survey data. We found that the adult learners' perception of computer programming changed during the course, with many participants planning their next step in computing education.
Douglas Lusa Krug, Chrystalla Mouza, W. Monty Jones, Taylor Barnett, David C. Shepherd
SIGCSE (1)5
2023 Do CONTRIBUTING Files Provide Information about OSS Newcomers' Onboarding Barriers?
abstract
Effectively onboarding newcomers is essential for the success of open source projects. These projects often provide onboarding guidelines in their ’CONTRIBUTING’ files (e.g., CONTRIBUTING.md on GitHub). These files explain, for example, how to find open tasks, implement solutions, and submit code for review. However, these files often do not follow a standard structure, can be too large, and miss barriers commonly found by newcomers. In this paper, we propose an automated approach to parse these CONTRIBUTING files and assess how they address onboarding barriers. We manually classified a sample of files according to a model of onboarding barriers from the literature, trained a machine learning classifier that automatically predicts the categories of each paragraph (precision: 0.655, recall: 0.662), and surveyed developers to investigate their perspective of the predictions’ adequacy (75% of the predictions were considered adequate). We found that CONTRIBUTING files typically do not cover the barriers newcomers face (52% of the analyzed projects missed at least 3 out of the 6 barriers faced by newcomers; 84% missed at least 2). Our analysis also revealed that information about choosing a task and talking with the community, two of the most recurrent barriers newcomers face, are neglected in more than 75% of the projects. We made available our classifier as an online service that analyzes the content of a given CONTRIBUTING file. Our approach may help community builders identify missing information in the project ecosystem they maintain and newcomers can understand what to expect in CONTRIBUTING files.
Felipe Fronchetti, David C. Shepherd, Igor Scaliante Wiese, Christoph Treude, Marco Aurélio Gerosa, Igor Steinmacher
ESEC/SIGSOFT FSE2
2023 Ready Worker One? High-Res VR for the Home Office
abstract
Many employees prefer to work from home, yet struggle to squeeze their office into an already fully-utilized space. Virtual Reality (VR) seemingly offered a solution with its ability to transform even modest physical spaces into spacious, productive virtual offices, but hardware challenges—such as low resolution—have prevented this from becoming a reality. Now that hardware issues are being overcome, we are able to investigate the suitability of VR for daily work. To do so, we (1) studied the physical space that users typically dedicate to home offices and (2) conducted an exploratory study of users working in VR for one week. For (1) we used digital ethnography to study 430 self-published images of software developer workstations in the home, confirming that developers faced myriad space challenges. We used speculative design to re-envision these as VR workstations, eliminating many challenges. For (2) we asked 10 developers to work in their own home using VR for about two hours each day for four workdays, and then interviewed them. We found that working in VR improved focus and made mundane tasks more enjoyable. While some subjects reported issues—annoyances with the fit, weight, and umbilical cord of the headset—the vast majority of these issues seem to be addressable. Together, these studies show VR technology has the potential to address many key problems with home workstations, and, with continued improvements, may become an integral part of creating an effective workstation in the home.
Anastasia Ruvimova, Felipe Fronchetti, Boden A Kahn, Luiz Henrique Susin, Zekeya Hurley, Thomas Fritz 0001, Mark S. Hancock, David C. Shepherd
VRST8
2023 Introduction to the Special Issue on Software-Intensive Autonomous Systems: Methods and applications
abstract
Identifying self-admitted technical debt (SATD) plays an important role in maintaining software stability and improving software quality. Although existing methods can detect SATD and researchers have identified design debt and requirement debt, an approach to realize multiple classification of SATD, including defect, test, and documentation, is still lacking. In this paper, we combine text generation oversampling and the Convolutional Neural Networks-Gated Recurrent Unit (CNNGRU) model, and propose an approach called SCGRU to classify multiple debt, including defect, test, documentation, design, and requirement. First, SeqGAN-based text generation is employed to generate new samples by learning the original SATD data, thereby increasing the number of SATD samples such as defect debt and reducing data imbalance. Then, we apply the CNNGRU model to refine SATD into multiple classes. An experiment with cross-project identification of 10 projects shows that our approach is more effective than existing methods such as CNN and text mining. The proposed SCGRU approach has strong advantages especially in cases of flawed debt with very unbalanced data such as test debt and documention debt.
Nesrine Khabou, Ismael Bouassida Rodriguez, Khalil Drira, Paris Avgeriou, David C. Shepherd, Wing Kwong Chan, Raffaela Mirandola
J. Syst. Softw.5
2023 Training industrial end-user programmers with interactive tutorials
abstract
Abstract Newly released robot programming tools have made it feasible for end‐users to program industrial robots by combining block‐based languages and lead‐through programming. To use these systems effectively, end‐users, who usually have limited or no programming experience, require training. To train users, tutoring systems are often used for block‐based programming—some even for lead‐through programming—but no tutorial system combines these two types of programming. We present CoBlox Interactive Tutorials (CITs), a novel tutoring approach that teaches how to use both the hardware and software components that comprise a typical end‐user robot programming environment. As users switch between the two programming styles, CITs provide them with extensive scaffolding, give users immediate feedback on missteps, and provide guidance on next steps. To evaluate CITs, we conducted a study with 79 industrial end‐users using a programming environment released by ABB Robotics that compares our approach to training with training videos, the most commonly used training in industry. This study, one of the largest to date on training professional end‐users, found that CIT‐trained users authored more correct programs in less time than video‐trained users. This shows that a tight integration of hardware and software concepts is crucial to training end‐users to program industrial robots.
Nico Ritschel, Anand Ashok Sawant, David Weintrop, Reid Holmes, Alberto Bacchelli, Ronald Garcia, Chandrika K. R., Avijit Mandal, Patrick Francis, David C. Shepherd
Softw. Pract. Exp.10
2022 A Case Study of Middle Schoolers' Use of Computational Thinking Concepts and Practices during Coded Music Composition
abstract
Researchers and practitioners have demonstrated various benefits of introducing computational thinking (CT) through music composition coding. While researchers have studied the impacts on participant attitudes towards CT and their learning of CT concepts, more case studies are needed on both learning CT concepts as well as CT practices, i.e., the processes of constructing music coding projects. This paper presents a case study of middle schoolers in an informal learning environment focused on integrating music composition with coding in TunePad. Specifically, we collected and analyzed logs of coding events, final code products, and surveys to explore both CT concept use and CT practices exhibited by the participants as they completed open-ended music coding activities to create their own melodies with specific music and CT requirements and recommendations.
Douglas Lusa Krug, Chrystalla Mouza, David C. Shepherd, Lori L. Pollock
ITiCSE (1)4
2022 Can guided decomposition help end-users write larger block-based programs? a mobile robot experiment
abstract
Block-based programming environments, already popular in computer science education, have been successfully used to make programming accessible to end-users in domains like robotics, mobile apps, and even DevOps. Most studies of these applications have examined small programs that fit within a single screen, yet real-world programs often grow large, and editing these large block-based programs quickly becomes unwieldy. Traditional programming language features, like functions, allow programmers to decompose their programs. Unfortunately, both previous work, and our own findings, suggest that end-users rarely use these features, resulting in large monolithic code blocks that are hard to understand. In this work, we introduce a block-based system that provides users with a hierarchical, domain-specific program structure and requires them to decompose their programs accordingly. Through a user study with 92 users, we compared this approach, which we call guided program decomposition, to a traditional system that supports functions, but does not require decomposition. We found that while almost all users could successfully complete smaller tasks, those who decomposed their programs were significantly more successful as the tasks grew larger. As expected, most users without guided decomposition did not decompose their programs, resulting in poor performance on larger problems. In comparison, users of guided decomposition performed significantly better on the same tasks. Though this study investigated only a limited selection of tasks in one specific domain, it suggests that guided decomposition can benefit end-user programmers. While no single decomposition strategy fits all domains, we believe that similar domain-specific sub-hierarchies could be found for other application areas, increasing the scale of code end-users can create and understand.
Nico Ritschel, Felipe Fronchetti, Reid Holmes, Ronald Garcia, David C. Shepherd
Proc. ACM Program. Lang.5
2022 Comparing Block-Based Programming Models for Two-Armed Robots
abstract
Modern industrial robots can work alongside human workers and coordinate with other robots. This means they can perform complex tasks, but doing so requires complex programming. Therefore, robots are typically programmed by experts, but there are not enough to meet the growing demand for robots. To reduce the need for experts, researchers have tried to make robot programming accessible to factory workers without programming experience. However, none of that previous work supports coordinating multiple robot arms that work on the same task. In this paper we present four block-based programming language designs that enable end-users to program two-armed robots. We analyze the benefits and trade-offs of each design on expressiveness and user cognition, and evaluate the designs based on a survey of 273 professional participants of whom 110 had no previous programming experience. We further present an interactive experiment based on a prototype implementation of the design we deem best. This experiment confirmed that novices can successfully use our prototype to complete realistic robotics tasks. This work contributes to making coordinated programming of robots accessible to end-users. It further explores how visual programming elements can make traditionally challenging programming tasks more beginner-friendly.
Nico Ritschel, Vladimir Kovalenko, Reid Holmes, Ronald Garcia, David C. Shepherd
IEEE Trans. Software Eng.5
2021 Code Beats: A Virtual Camp for Middle Schoolers Coding Hip Hop
abstract
In spite of the efforts to provide computer science education for all, the percentage of Black and Latino Americans entering the computer science (CS) field has been stagnant for years. In an effort to attract and engage students many summer camps and after-school clubs use robotics, video-games, and even IoT devices, but these approaches seem to only attract those already considering STEM careers, a population low in Black and Latino students. To attract Black and Latino students to computer science a promising approach is to engage with their culture, making CS relevant to them personally. To this end, we present an approach that teaches middle school students to program using hip hop beats, intentionally leveraging a genre of music that appeals to a wide array of urban youth of color. This approach, called Code Beats, uses extensive scaffolding to support beginning students, authentic-sounding beats to engage students, and a expressive programming environment to support creative freedom. We present the results of our pilot camp, where students clearly showed an increase in computing enjoyment, confidence, belonging, and persistence. By the end of this course, all students were able to create their own, original beat from scratch, suggesting their progression to the Create phase of the Use-Modify-Create framework.
Douglas Lusa Krug, Edtwuan Bowman, Taylor Barnett, Lori L. Pollock, David C. Shepherd
SIGCSE5
2021 Observing and predicting knowledge worker stress, focus and awakeness in the wild
Mauricio Soto, Chris Satterfield, Thomas Fritz 0001, Gail C. Murphy, David C. Shepherd, Nicholas A. Kraft
Int. J. Hum. Comput. Stud.5
2021 What Predicts Software Developers' Productivity?
abstract
Organizations have a variety of options to help their software developers become their most productive selves, from modifying office layouts, to investing in better tools, to cleaning up the source code. But which options will have the biggest impact? Drawing from the literature in software engineering and industrial/organizational psychology to identify factors that correlate with productivity, we designed a survey that asked 622 developers across 3 companies about these productivity factors and about self-rated productivity. Our results suggest that the factors that most strongly correlate with self-rated productivity were non-technical factors, such as job enthusiasm, peer support for new ideas, and receiving useful feedback about job performance. Compared to other knowledge workers, our results also suggest that software developers' self-rated productivity is more strongly related to task variety and ability to work remotely.
Emerson R. Murphy-Hill, Ciera Jaspan, Caitlin Sadowski, David C. Shepherd, Michael Phillips, Collin Winter, Andrea Knight, Edward K. Smith, Matthew Jorde
IEEE Trans. Software Eng.4
2020 "Transport Me Away": Fostering Flow in Open Offices through Virtual Reality
abstract
Open offices are cost-effective and continue to be popular. However, research shows that these environments, brimming with distractions and sensory overload, frequently hamper productivity. Our research investigates the use of virtual reality (VR) to mitigate distractions in an open office setting and improve one's ability to be in flow. In a lab study, 35 participants performed visual programming tasks in four combinations of physical (open or closed office) and virtual environments (beach or virtual office). While participants both preferred and were in flow more in a closed office without VR, in an open office, the VR environments outperformed the no VR condition in all measures of flow, performance, and preference. Especially considering the recent rapid advancements in VR, our findings illustrate the potential VR has to improve flow and satisfaction in open offices.
Anastasia Ruvimova, Junhyeok Kim 0001, Thomas Fritz 0001, Mark S. Hancock, David C. Shepherd
CHI5
2020 Software documentation: the practitioners' perspective
abstract
In theory, (good) documentation is an invaluable asset to any software project, as it helps stakeholders to use, understand, maintain, and evolve a system. In practice, however, documentation is generally affected by numerous shortcomings and issues, such as insufficient and inadequate content and obsolete, ambiguous information. To counter this, researchers are investigating the development of advanced recommender systems that automatically suggest high-quality documentation, useful for a given task. A crucial first step is to understand what quality means for practitioners and what information is actually needed for specific tasks.
Emad Aghajani, Csaba Nagy 0001, Mario Linares-Vásquez, Laura Moreno, Gabriele Bavota, Michele Lanza 0001, David C. Shepherd
ICSE7
2020 Guest editorial: special section on software analysis, evolution, and reengineering
Massimiliano Di Penta, David C. Shepherd
Empir. Softw. Eng.2
2019 Modeling hierarchical usage context for software exceptions based on interaction data
Hui Chen 0001, Kostadin Damevski, David C. Shepherd, Nicholas A. Kraft
Autom. Softw. Eng.3
2019 Characterizing industry-academia collaborations in software engineering: evidence from 101 projects
abstract
Research collaboration between industry and academia supports improvement and innovation in industry and helps ensure the industrial relevance of academic research. However, many researchers and practitioners in the community believe that the level of joint industry-academia collaboration (IAC) projects in Software Engineering (SE) research is relatively low, creating a barrier between research and practice. The goal of the empirical study reported in this paper is to explore and characterize the state of IAC with respect to industrial needs, developed solutions, impacts of the projects and also a set of challenges, patterns and anti-patterns identified by a recent Systematic Literature Review (SLR) study. To address the above goal, we conducted an opinion survey among researchers and practitioners with respect to their experience in IAC. Our dataset includes 101 data points from IAC projects conducted in 21 different countries. Our findings include: (1) the most popular topics of the IAC projects, in the dataset, are: software testing, quality, process, and project managements; (2) over 90% of IAC projects result in at least one publication; (3) almost 50% of IACs are initiated by industry, busting the myth that industry tends to avoid IACs; and (4) 61% of the IAC projects report having a positive impact on their industrial context, while 31% report no noticeable impacts or were “not sure”. To improve this situation, we present evidence-based recommendations to increase the success of IAC projects, such as the importance of testing pilot solutions before using them in industry. This study aims to contribute to the body of evidence in the area of IAC, and benefit researchers and practitioners. Using the data and evidence presented in this paper, they can conduct more successful IAC projects in SE by being aware of the challenges and how to overcome them, by applying best practices (patterns), and by preventing anti-patterns.
Vahid Garousi, Dietmar Pfahl, João M. Fernandes 0001, Michael Felderer, Mika Mäntylä, David C. Shepherd, Andrea Arcuri, Ahmet Coskunçay, Bedir Tekinerdogan
Empir. Softw. Eng.6
2018 Evaluating CoBlox: A Comparative Study of Robotics Programming Environments for Adult Novices
abstract
A new wave of collaborative robots designed to work alongside humans is bringing the automation historically seen in large-scale industrial settings to new, diverse contexts. However, the ability to program these machines often requires years of training, making them inaccessible or impractical for many. This paper rethinks what robot programming interfaces could be in order to make them accessible and intuitive for adult novice programmers. We created a block-based interface for programming a one-armed industrial robot and conducted a study with 67 adult novices comparing it to two programming approaches in widespread use in industry. The results show participants using the block-based interface successfully implemented robot programs faster with no loss in accuracy while reporting higher scores for usability, learnability, and overall satisfaction. The contribution of this work is showing the potential for using block-based programming to make powerful technologies accessible to a wider audience.
David Weintrop, Afsoon Afzal, Jean Salac, Patrick Francis, Boyang Li 0002, David C. Shepherd, Diana Franklin
CHI6
2018 Predicting future developer behavior in the IDE using topic models
abstract
Interaction data, gathered from developers' daily clicks and key presses in the IDE, has found use in both empirical studies and in recommendation systems for software engineering. We observe that this data has several characteristics, common across IDEs:
Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Nicholas A. Kraft, Lori L. Pollock
ICSE3
2018 [Engineering Paper] An IDE for Easy Programming of Simple Robotics Tasks
abstract
Many robotic tasks in small manufacturing sites are quite simple. For example, a pick and place task requires only a few common commands. Unfortunately, the standard languages and programming environments for industrial robots are complex, making even these simple tasks nearly impossible for novices. To enable novices to program simple tasks we created a block-based programming language and environment focused on usability, learnability, and understandability and embedded its programming environment in a state-of-the-art robot simulator. By using this high-fidelity prototype over the course of a year in a case study, a user study, and for countless demonstrations we have gained many concrete insights. In this paper we discuss the details of the language, the design of its programming environment, and concrete insights gained via longitudinal usage.
David C. Shepherd, Patrick Francis, David Weintrop, Diana Franklin, Boyang Li 0002, Afsoon Afzal
SCAM1
2018 Predicting Future Developer Behavior in the IDE Using Topic Models
abstract
While early software command recommender systems drew negative user reaction, recent studies show that users of unusually complex applications will accept and utilize command recommendations. Given this new interest, more than a decade after first attempts, both the recommendation generation (backend) and the user experience (frontend) should be revisited. In this work, we focus on recommendation generation. One shortcoming of existing command recommenders is that algorithms focus primarily on mirroring the short-term past,-i.e., assuming that a developer who is currently debugging will continue to debug endlessly. We propose an approach to improve on the state of the art by modeling future task context to make better recommendations to developers. That is, the approach can predict that a developer who is currently debugging may continue to debug OR may edit their program. To predict future development commands, we applied Temporal Latent Dirichlet Allocation, a topic model used primarily for natural language, to software development interaction data (i.e., command streams). We evaluated this approach on two large interaction datasets for two different IDEs, Microsoft Visual Studio and ABB Robot Studio. Our evaluation shows that this is a promising approach for both predicting future IDE commands and producing empirically-interpretable observations.
Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Nicholas A. Kraft, Lori L. Pollock
IEEE Trans. Software Eng.3
2017 Reducing Interruptions at Work: A Large-Scale Field Study of FlowLight
abstract
Due to the high number and cost of interruptions at work, several approaches have been suggested to reduce this cost for knowledge workers. These approaches predominantly focus either on a manual and physical indicator, such as headphones or a closed office door, or on the automatic measure of a worker's interruptibilty in combination with a computer-based indicator. Little is known about the combination of a physical indicator with an automatic interruptibility measure and its long-term impact in the workplace. In our research, we developed the FlowLight, that combines a physical traffic-light like LED with an automatic interruptibility measure based on computer interaction data. In a large-scale and long-term field study with 449 participants from 12 countries, we found, amongst other results, that the FlowLight reduced the interruptions of participants by 46%, increased their awareness on the potential disruptiveness of interruptions and most participants never stopped using it.
Manuela Züger, Christopher S. Corley, André N. Meyer, Boyang Li 0002, Thomas Fritz 0001, David C. Shepherd, Vinay Augustine, Patrick Francis, Nicholas A. Kraft, Will Snipes
CHI6
2017 Behavior Metrics for Prioritizing Investigations of Exceptions
abstract
Many software development teams collect product defect reports, which can either be manually submitted or automatically created from product logs. Periodically, the teams use the collected defect reports to prioritize which defect to address next. We present a set of behavior-based metrics that can be used in this process. These metrics are based on the insight that development teams can estimate user inconvenience from user and application behavior in interaction logs. To estimate user inconvenience, the behavior metrics capture important user and application behavior after exceptions (the defects of interest in our case). We validated these metrics through a survey of how developers would incorporate the behavior metrics into their prioritization decisions. We found that developers change their priority of investigating an exception about 31% of the time after including the behavior metrics in the priority decision. These findings provide evidence that behavior metrics provide a promising advance towards prioritizing application exceptions.
Zack Coker, Kostadin Damevski, Claire Le Goues, Nicholas A. Kraft, David C. Shepherd, Lori L. Pollock
ICSME5
2017 On-demand Developer Documentation
abstract
We advocate for a paradigm shift in supporting the information needs of developers, centered around the concept of automated on-demand developer documentation. Currently, developer information needs are fulfilled by asking experts or consulting documentation. Unfortunately, traditional documentation practices are inefficient because of, among others, the manual nature of its creation and the gap between the creators and consumers. We discuss the major challenges we face in realizing such a paradigm shift, highlight existing research that can be leveraged to this end, and promote opportunities for increased convergence in research on software documentation.
Martin P. Robillard, Andrian Marcus, Christoph Treude, Gabriele Bavota, Oscar Chaparro, Neil A. Ernst, Marco Aurélio Gerosa, Michael W. Godfrey, Michele Lanza 0001, Mario Linares-Vásquez, Gail C. Murphy, Laura Moreno, David C. Shepherd, Edmund Wong
ICSME13
2017 Eye gaze and interaction contexts for change tasks - Observations and potential
Katja Kevic, Braden Walters, Timothy Shaffer, Bonita Sharif, David C. Shepherd, Thomas Fritz 0001
J. Syst. Softw.5
2017 Mining Sequences of Developer Interactions in Visual Studio for Usage Smells
abstract
In this paper, we present a semi-automatic approach for mining a large-scale dataset of IDE interactions to extract usage smells, i.e., inefficient IDE usage patterns exhibited by developers in the field. The approach outlined in this paper first mines frequent IDE usage patterns, filtered via a set of thresholds and by the authors, that are subsequently supported (or disputed) using a developer survey, in order to form usage smells. In contrast with conventional mining of IDE usage data, our approach identifies time-ordered sequences of developer actions that are exhibited by many developers in the field. This pattern mining workflow is resilient to the ample noise present in IDE datasets due to the mix of actions and events that these datasets typically contain. We identify usage patterns and smells that contribute to the understanding of the usability of Visual Studio for debugging, code search, and active file navigation, and, more broadly, to the understanding of developer behavior during these software development activities. Among our findings is the discovery that developers are reluctant to use conditional breakpoints when debugging, due to perceived IDE performance problems as well as due to the lack of error checking in specifying the conditional.
Kostadin Damevski, David C. Shepherd, Johannes Schneider 0002, Lori L. Pollock
IEEE Trans. Software Eng.2
2016 An empirical study of practitioners' perspectives on green software engineering
abstract
The energy consumption of software is an increasing concern as the use of mobile applications, embedded systems, and data center-based services expands. While research in green software engineering is correspondingly increasing, little is known about the current practices and perspectives of software engineers in the field. This paper describes the first empirical study of how practitioners think about energy when they write requirements, design, construct, test, and maintain their software. We report findings from a quantitative, targeted survey of 464 practitioners from ABB, Google, IBM, and Microsoft, which was motivated by and supported with qualitative data from 18 in-depth interviews with Microsoft employees. The major findings and implications from the collected data contextualize existing green software engineering research and suggest directions for researchers aiming to develop strategies and tools to help practitioners improve the energy usage of their applications.
Irene Manotas, Christian Bird, David C. Shepherd, Ciera Jaspan, Caitlin Sadowski, Lori L. Pollock, James Clause
ICSE4
2016 Interactive exploration of developer interaction traces using a hidden markov model
abstract
Using IDE usage data to analyze the behavior of software developers in the field, during the course of their daily work, can lend support to (or dispute) laboratory studies of developers. This paper describes a technique that leverages Hidden Markov Models (HMMs) as a means of mining high-level developer behavior from low-level IDE interaction traces of many developers in the field. HMMs use dual stochastic processes to model higher-level hidden behavior using observable input sequences of events. We propose an interactive approach of mining interpretable HMMs, based on guiding a human expert in building a high quality HMM in an iterative, one state at a time, manner. The final result is a model that is both representative of the field data and captures the field phenomena of interest. We apply our HMM construction approach to study debugging behavior, using a large IDE interaction dataset collected from nearly 200 developers at ABB, Inc. Our results highlight the different modes and constituent actions in debugging, exhibited by the developers in our dataset.
Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Lori L. Pollock
MSR3
2016 A field study of how developers locate features in source code
Kostadin Damevski, David C. Shepherd, Lori L. Pollock
Empir. Softw. Eng.2
2016 Roundtable: Research Opportunities and Challenges for Large-Scale Software Systems
Xusheng Xiao, Jian-Guang Lou, Shan Lu 0001, David C. Shepherd, Xin Peng 0001, Qianxiang Wang
J. Comput. Sci. Technol.4
2015 A Field Study on Fostering Structural Navigation with Prodet
abstract
Past studies show that developers who navigate code in a structural manner complete tasks faster and more correctly than those whose behavior is more opportunistic. The goal of this work is to move professional developers towards more effective program comprehension and maintenance habits by providing an approach that fosters structural code navigation. To this end, we created a Visual Studio plugin called Prodet that integrates an always-on navigable visualization of the most contextually relevant portions of the call graph. We evaluated the effectiveness of our approach by deploying it in a six week field study with professional software developers. The study results show a statistically significant increase in developers' use of structural navigation after installing Prodet. The results also show that developers continuously used the filtered and navigable call graph over the three week period in which it was deployed in production. These results indicate the maturity and value of our approach to increase developers' effectiveness in a practical and professional environment.
Vinay Augustine, Patrick Francis, Xiao Qu, David C. Shepherd, Will Snipes, Christoph Bräunlich, Thomas Fritz 0001
ICSE (2)4
2015 How and When to Transfer Software Engineering Research via Extensions
abstract
It is often reported that there is a large gap between software engineering research and practice, with little transfer from research to practice. While this is true in general, one transfer technique is increasingly breaking down this barrier: extensions to integrated development environments (IDEs). With the proliferation of app stores for IDEs and increasing transfer effort from researchers several research-based extensions have seen significant adoption. In this talk we'll discuss our experience transferring code search research, which currently is in the top 5% of Visual Studio extensions with over 13,000 downloads, as well as other research techniques transferred via extensions such as NCrunch, FindBugs, Code Recommenders, Mylyn, and Instasearch. We'll use the lessons learned from our transfer experience to provide case study evidence as to best practices for successful transfer, supplementing it with the quantitative evidence offered by app store and usage data across the broader set of extensions. The goal of this 30 minute talk is to provide researchers with a realistic view on which research techniques can be transferred to practice as well as concrete steps to execute such a transfer.
David C. Shepherd, Kostadin Damevski, Lori L. Pollock
ICSE (2)1
2015 Exploring the use of concern element role information in feature location evaluation
abstract
Before making changes, programmers need to locate and understand source code that corresponds to specific functionality, i.e., Perform concern or feature location. Numerous concern and feature location techniques have been proposed, but to the best of our knowledge, no existing techniques or evaluations report information on what role a code element plays in the larger concern. In this paper, we report on two case studies that investigate two hypotheses on how evaluation studies of concern location techniques can be strengthened by utilizing concern role information: (1) by increasing agreement among human annotators for gold set establishment and (2) by providing richer information about the elements ranked as relevant by concern location techniques, which could help further improve the tools. We conducted a case study of 6 Java developers annotating 3 concerns with role information. When the developers understood the task description, pair wise agreement increased by 20%, 25%, and 135% for the 3 concerns over a prior concern location study without role information. Our findings also suggest that there may be core element roles that need to be annotated by humans, but that the remaining roles may be automatically derived, which could facilitate more reliable concern location benchmarks in the future. We also conducted an exploratory study of the element roles represented in results returned by a state of the art feature location tool. The results of these two studies suggest that integrating concern element role information into evaluations can help to strengthen both the gold set establishment and the analysis of results returned by various tools.
Emily Hill 0001, David C. Shepherd, Lori L. Pollock
ICPC2
2015 Tracing software developers' eyes and interactions for change tasks
abstract
What are software developers doing during a change task? While an answer to this question opens countless opportunities to support developers in their work, only little is known about developers' detailed navigation behavior for realistic change tasks. Most empirical studies on developers performing change tasks are limited to very small code snippets or are limited by the granularity or the detail of the data collected for the study. In our research, we try to overcome these limitations by combining user interaction monitoring with very fine granular eye-tracking data that is automatically linked to the underlying source code entities in the IDE. In a study with 12 professional and 10 student developers working on three change tasks from an open source system, we used our approach to investigate the detailed navigation of developers for realistic change tasks. The results of our study show, amongst others, that the eye tracking data does indeed capture different aspects than user interaction data and that developers focus on only small parts of methods that are often related by data flow. We discuss our findings and their implications for better developer tool support.
Katja Kevic, Braden Walters, Timothy Shaffer, Bonita Sharif, David C. Shepherd, Thomas Fritz 0001
ESEC/SIGSOFT FSE5
2015 Scaling up evaluation of code search tools through developer usage metrics
abstract
Code search is a fundamental part of program understanding and software maintenance and thus researchers have developed many techniques to improve its performance, such as corpora preprocessing and query reformulation. Unfortunately, to date, evaluations of code search techniques have largely been in lab settings, while scaling and transitioning to effective practical use demands more empirical feedback from the field. This paper addresses that need by studying metrics based on automatically-gathered anonymous field data from code searches to infer user satisfaction. We describe techniques for addressing important concerns, such as how privacy is retained and how the overhead on the interactive system is minimized. We perform controlled user and field studies which identify metrics that correlate with user satisfaction, enabling the future evaluation of search tools through anonymous usage data. In comparing our metrics to similar metrics used in Internet search we observe differences in the relationship of some of the metrics to user satisfaction. As we further explore the data, we also present a predictive multi-metric model that achieves accuracy of over 70% in determining query satisfaction.
Kostadin Damevski, David C. Shepherd, Lori L. Pollock
SANER2
2014 CoMoGen: An Approach to Locate Relevant Task Context by Combining Search and Navigation
abstract
Developers spend a substantial amount of time searching and navigating source code to locate the relevant places for performing a change task. While the searching and navigating are highly intertwined and related, most current approaches focus either on search or on navigation support for developers, keeping the two distinct. In this paper, we present an approach called CoMoGen that combines search and navigation by expanding, ranking and visualizing search results with navigation context. In an experimental analysis we found that our approach is able to generate small task-relevant context models that locates more relevant search results than state-of-the-art and state-of-the-practice search approaches. A small, preliminary user study with ten participants further yields promising preliminary findings that CoMoGen supports developers in better understanding and assessing the relevance of search results and in reducing navigation steps.
Katja Kevic, Thomas Fritz 0001, David C. Shepherd
ICSME3
2014 Developers' code context models for change tasks
abstract
To complete a change task, software developers spend a substantial amount of time navigating code to understand the relevant parts. During this investigation phase, they implicitly build context models of the elements and relations that are relevant to the task. Through an exploratory study with twelve developers completing change tasks in three open source systems, we identified important characteristics of these context models and how they are created. In a second empirical analysis, we further examined our findings on data collected from eighty developers working on a variety of change tasks on open and closed source projects. Our studies uncovered, amongst other results, that code context models are highly connected, structurally and lexically, that developers start tasks using a combination of search and navigation and that code navigation varies substantially across developers. Based on these findings we identify and discuss design requirements to better support developers in the initial creation of code context models. We believe this work represents a substantial step in better understanding developers' code navigation and providing better tool support that will reduce time and effort needed for change tasks.
Thomas Fritz 0001, David C. Shepherd, Katja Kevic, Will Snipes, Christoph Bräunlich
SIGSOFT FSE2
2014 How developers use multi-recommendation system in local code search
abstract
Developers often start programming tasks by searching for relevant code in their local codebase. Previous research suggests that 88% of manually-composed queries retrieve no relevant results. Many searches fail because existing search tools depend solely on string matching with a manually-composed query, which cannot find semantically-related code. To solve this problem, researchers proposed query recommendation techniques to help developers compose queries without the extensive knowledge of the codebase under search. However, few of these techniques are empirically evaluated by the usage data from real-world developers. To fill this gap, we studied several query recommendation techniques by extending Sando and conducting a longitudinal field study. Our study shows that over 30% of all queries were adopted from recommendation; and recommended queries retrieved results 7% more often than manual queries.
Xi Ge, David C. Shepherd, Kostadin Damevski, Emerson R. Murphy-Hill
VL/HCC2
2013 Differentiating Roles of Program Elements in Action-Oriented Concerns
abstract
Many techniques have been developed to help programmers locate source code that corresponds to specific functionality, i.e., concern or feature location, as it is a frequent software maintenance activity. This paper proposes operational definitions for differentiating the roles that each program element of a concern plays with respect to the concern's implementation. By identifying the respective roles, we enable evaluations that provide more insight into comparative performance of concern location techniques. To provide definitions that are specific enough to be useful in practice, we focus on the subset of concerns that are action-oriented. We also conducted a case study that compares concern mappings derived from our role definitions with three developers' mappings across three concerns. The results suggest that our definitions capture the majority of developer-identified elements and that control-flow islands (i.e., groups of elements with little to no control flow connections) can cause developers to omit relevant elements.
Emily Hill 0001, David C. Shepherd, Lori L. Pollock, K. Vijay-Shanker
ICSM2
2012 Sando: an extensible local code search framework
abstract
Developers heavily rely on Local Code Search (LCS)---the execution of a text-based search on a single code base---to find starting points in software maintenance tasks. While LCS approaches commonly used by developers are based on lexical matching and often result in failed searches or irrelevant results, developers have not yet migrated to the various research approaches that have made significant advancements in LCS. We hypothesize that two of the major reasons for this lack of migration are as follows. First, developers do not know which approach is the best, due to a lack of comparative field studies and the discrepancies in the underlying LCS process that these research approaches address. Second, developers lack access to a stable implementation of most of the research approaches. To address these issues, we studied a number of LCS approaches, distilled the general component structure underlying these approaches and, based on this structure, developed a LCS tool and framework, called Sando. Currently used by developers at ABB, Inc. and elsewhere, Sando also supports the flexible extension of its components to rapidly disseminate research advancements, and allows for user-based evaluation of competing approaches.
David C. Shepherd, Kostadin Damevski, Bartosz Ropski, Thomas Fritz 0001
SIGSOFT FSE1
2009 Using activity traces to characterize programming behaviour beyond the lab
abstract
Systematically improving the efficiency of programmers requires understanding what activities occur during programming, which activities are inefficient and then assessing languages, tools and processes proposed to improve the situation. Conducting the experiments required to support a systematic approach is difficult for many reasons, including the lack of availability of experienced programmers and the common belief that individual programmer effectiveness varies greatly. In this paper, we investigate whether generic activity traces of how a programmer interacts with a development environment can help bridge between results gathered in the lab with how programming occurs in the field. We describe the kinds of information that can be gleaned from activity traces, consider whether positive indication of a behaviour seen in the lab translates to data collected from the field, and discuss challenges with gathering appropriate data and with using gathered data appropriately.
Gail C. Murphy, Petcharat Viriyakattiyaporn, David C. Shepherd
ICPC3
2007 Introducing natural language program analysis
abstract
This research group presentation focuses on our work in extracting and utilizing natural language clues from source code to improve software maintenance tools. We demonstrate the valuable information that can be gained from a software system's identifiers, literals, and comments. We then present an overview of our extraction process, program representation, and a set of tools we have developedusing this natural language program analysis.
Lori L. Pollock, K. Vijay-Shanker, David C. Shepherd, Emily Hill 0001, Zachary P. Fry, Kishen Maloor
PASTE3
2007 Case study: supplementing program analysis with natural language analysis to improve a reverse engineering task
abstract
Software maintainers often use reverse engineering tools to aid in the extremely difficult task of understanding unfamiliar code, especially within large, complex software systems. While traditional program analysis can provide detailed information for reverse engineering, often this information is not sufficient to assist the user with high-level program understanding tasks. To bridge the gap between current reverse engineering tools and the high-level questions that software maintainers want answered, we propose supplementing traditional program analysis with natural language analysis of program source code. This paper presents a case study where we have augmented an existing reverse engineering tool, an aspect miner, to complement the existing traditional program analysis-based miner with natural language analysis of method names, class names, and comments. Our quantitative and qualitative results strongly suggest that supplementing traditional program analysis with natural language analysis is a promising approach to raising the level of effectiveness of reverse engineering tools.
David C. Shepherd, Lori L. Pollock, K. Vijay-Shanker
PASTE1
2005 Timna: a framework for automatically combining aspect mining analyses
abstract
To realize the benefits of Aspect Oriented Programming (AOP), developers must refactor active and legacy code bases into an AOP language. When refactoring, developers first need to identify refactoring candidates, a process called aspect mining. Humans perform mining by using a variety of clues to determine which code to refactor. However, existing approaches to automating the aspect mining process focus on developing analyses of a single program characteristic. Each analysis often finds only a subset of possible refactoring candidates and is unlikely to find candidates which humans find by combining analyses. In this paper, we present Timna, a framework for enabling the automatic combination of aspect mining analyses. The key insight is the use of machine learning to learn when to refactor, from vetted examples. Experimental evaluation of the cost-effectiveness of Timna in comparison to Fan-in, a leading aspect mining analysis, indicates that such a framework for automatically combining analyses is very promising.
David C. Shepherd, Jeffrey Palm, Lori L. Pollock, Mark Chu-Carroll
ASE1
2003 Testing with Respect to Concerns
abstract
Often the code regions that are assigned for a maintenance task do not follow the modularization of the original application program, but instead include parts of code from many different units scattered throughout the application. In this paper, we investigate an approach to testing which we call concern-based testing, which leverages existing tools to help software maintainers identify the relevant code for their assigned task, their concern. The main contribution is a demonstration of the possible savings in test suite execution overhead and the increased precision in coverage information that can be obtained for a software maintainer if testing tasks are performed with respect to concerns. Based on a concern graph representation of the concern, a framework for guiding selective instrumentation for scalable coverage analysis is also presented.
Amie L. Souter, David C. Shepherd, Lori L. Pollock
ICSM2