Justin Smith 0001

dblp:75/2200 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
4since 2021 · last 2024
0000-0001-6987-5196ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 10 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 Exploring Experiences with Automated Program Repair in Practice
abstract
Automated program repair, also known as APR, is an approach for automatically repairing software faults. There is a large amount of research on automated program repair, but very little offers in-depth insights into how practitioners think about and employ APR in practice. To learn more about practitioners' perspectives and experiences with current APR tools and techniques, we administered a survey, which received valid responses from 331 software practitioners. We analyzed survey responses to gain insights regarding factors that correlate with APR awareness, experience, and use. We established a strong correlation between APR awareness and tool use and attributes including job position, company size, total coding experience, and preferred language of software practitioners. We also found that practitioners are using other forms of support, such as co-workers and ChatGPT, more frequently than APR tools when fixing software defects. We learned about the drawbacks that practitioners encounter while utilizing existing APR tools and the impact that each drawback has on their practice. Our findings provide implications for research and practice centered on development, adoption, and use of APR.
Fairuz Nawer Meem, Justin Smith 0001, Brittany Johnson
ICSE2
2024 Challenges and Opportunities for Survey Research in the Age of Generative AI: An Experience Report
abstract
Survey research, while common, has known challenges and limitations. In this paper, we discuss the challenges and concerns of conducting research using online survey in the age of widespread development and use of generative AI technologies. We base our discussion on our recent experiences conducting survey research to better understand software practitioners’ knowledge of and experience with automated program repair (APR). In our efforts, we encountered both anticipated and unexpected challenges, much of which stemmed from advancements in AI-assisted automation (e.g., generative AI). For example, we found that many of the open ended responses provided were likely generated by AI. Based on our experiences, we outline how we mitigated or handled these challenges and discuss opportunities for improving survey design and administration to ensure high quality research and outcomes in the age of advanced and widely available AI technologies.
Fairuz Nawer Meem, Justin Smith 0001, Brittany Johnson
VL/HCC2
2023 A Taxonomy of Machine Learning Fairness Tool Specifications, Features and Workflows
abstract
Biased machine learning (ML) models in real-world applications, such as healthcare and criminal justice, have led to significant societal harm. Several ML fairness tools promise to help create less biased ML models. However, little effort has been made to help practitioners or researchers reason about the growing space of available tools. In this work, we evaluated and categorized 14 existing fairness tools based on their features and usage workflows to develop a practical taxonomy of machine learning fairness tools. Our resulting taxonomy of fairness tools suggests the availability of an array of fairness tools, including tools that require little coding or allow for customization or extension. By structuring and organizing the landscape of fairness tools, we can identify gaps and explore options for supporting researchers and practitioners, such as automated tools for finding and selecting fairness tools.
Sadia Afrin Mim, Justin Smith 0001, Brittany Johnson
VL/HCC2
2021 How Students Unit Test: Perceptions, Practices, and Pitfalls
abstract
Unit testing is reported as one of the skills that graduating students lack, yet it is an essential skill for professional software developers. Understanding the challenges students face during testing can help inform practices for software testing education. To that end, we conduct an exploratory study to reveal students' perceptions of unit testing and challenges students encounter when practicing unit testing. We surveyed 54 students from two universities and gave them two testing tasks, one involving black-box test design and one involving white-box test implementation. For the tasks, we used two software projects from prior work in studying test-first development among software developers. We quantitatively analyzed the survey responses and test code properties, and qualitatively identified the mistakes and smells in the test code. We further report on our experience running this study with students.
Gina R. Bai, Justin Smith 0001, Kathryn T. Stolee
ITiCSE (1)2
2020 MatchingRef: Matching Variable Names in a Reference Page to Help Introductory CS Students Fix Compiler Errors
abstract
Debugging compiler errors is essential to programming and can be challenging for novice programmers. In introductory computer science courses, challenging errors can discourage students. One reason these errors are difficult to resolve is that most online help systems do not match a student's code. For example, online reference pages use different variable names, identifiers, and method names compared with a student's particular code. To utilize existing resources, students must wade through other people's code (which is often too advanced for novices to comprehend). This is time-consuming and does not provide novices with solutions.
Thuc Nhi Le, Shokhzodbek Saidov, Justin Smith 0001
ICER3
2020 A Case Study of Software Security Red Teams at Microsoft
abstract
The modern software security adversary employs persistent and evasive attack techniques, for example-using zero-day exploits that have not been disclosed publicly-to target high-profile companies for political and economic espionage or to exfiltrate sensitive data or intellectual property. To combat these threats, large organizations are adopting an emerging practice of staffing full-time offensive security teams, or red teams. To understand the workflows, culture, and day-to-day practices of software security engineers in red teams, we conducted 17 interviews with informants across five red teams within Microsoft. We found that software security engineers have substantial impact in the organization as they harden security practices, drawing from their diverse backgrounds. Software security engineers are both agile yet specialized in their activities, and closely emulate malicious adversaries-subject to some reasonable constraints. Although software security engineers are in some respects software engineers, they also have several consequential differences in how they write, maintain, and distribute software. The results of this work are applicable to practitioners, researchers, and toolsmiths who wish to understand how offensive security teams operate, situate, and collaborate with partner teams in their organization.
Justin Smith 0001, Christopher Theisen, Titus Barik
VL/HCC1
2019 How Developers Diagnose Potential Security Vulnerabilities with a Static Analysis Tool
abstract
While using security tools to resolve security defects, software developers must apply considerable effort. Success depends on a developer's ability to interact with tools, ask the right questions, and make strategic decisions. To build better security tools and subsequently help developers resolve defects more accurately and efficiently, we studied the defect resolution process-from the questions developers ask to their strategies for answering them. In this paper, we report on an exploratory study with novice and experienced software developers. We equipped them with Find Security Bugs, a security-oriented static analysis tool, and observed their interactions with security vulnerabilities in an open-source system that they had previously contributed to. We found that they asked questions not only about security vulnerabilities, associated attacks, and fixes, but also questions about the software itself, the social ecosystem that built the software, and related resources and tools. We describe the strategic successes and failures we observed and how future tools can leverage our findings to encourage better strategies.
Justin Smith 0001, Brittany Johnson, Emerson R. Murphy-Hill, Bill Chu, Heather Lipford
IEEE Trans. Software Eng.1
2018 Does ACM's code of ethics change ethical decision making in software development?
abstract
Ethical decisions in software development can substantially impact end-users, organizations, and our environment, as is evidenced by recent ethics scandals in the news. Organizations, like the ACM, publish codes of ethics to guide software-related ethical decisions. In fact, the ACM has recently demonstrated renewed interest in its code of ethics and made updates for the first time since 1992. To better understand how the ACM code of ethics changes software-related decisions, we replicated a prior behavioral ethics study with 63 software engineering students and 105 professional software developers, measuring their responses to 11 ethical vignettes. We found that explicitly instructing participants to consider the ACM code of ethics in their decision making had no observed effect when compared with a control group. Our findings suggest a challenge to the research community: if not a code of ethics, what techniques can improve ethical decision making in software engineering?
Andrew McNamara, Justin Smith 0001, Emerson R. Murphy-Hill
ESEC/SIGSOFT FSE2
2018 Supporting Effective Strategies for Resolving Vulnerabilities Reported by Static Analysis Tools
abstract
Static analysis tools detect potentially costly security defects early in the software development process. However, these defects can be difficult for developers to accurately and efficiently resolve. The goal of this work is to understand the vulnerability resolution process so that we can build tools that support more effective strategies for resolving vulnerabilities. In this work, I study developers as they resolve security vulnerabilities to identify their information needs and current strategies. Next, I study existing tools to understand how they support developers' strategies. Finally, I plan to demonstrate how strategy-aware tools can help developers resolve security vulnerabilities more accurately and efficiently.
Justin Smith 0001
VL/HCC1
2017 Do developers read compiler error messages?
abstract
In integrated development environments, developers receive compiler error messages through a variety of textual and visual mechanisms, such as popups and wavy red underlines. Although error messages are the primary means of communicating defects to developers, researchers have a limited understanding on how developers actually use these messages to resolve defects. To understand how developers use error messages, we conducted an eye tracking study with 56 participants from undergraduate and graduate software engineering courses at our university. The participants attempted to resolve common, yet problematic defects in a Java code base within the Eclipse development environment. We found that: 1) participants read error messages and the difficulty of reading these messages is comparable to the difficulty of reading source code, 2) difficulty reading error messages significantly predicts participants' task performance, and 3) participants allocate a substantial portion of their total task to reading error messages (13%-25%). The results of our study offer empirical justification for the need to improve compiler error messages for developers.
Titus Barik, Justin Smith 0001, Kevin Lubick, Elisabeth Holmes, Jing Feng 0004, Emerson R. Murphy-Hill, Chris Parnin
ICSE2
2017 Just-in-time static analysis
abstract
We present the concept of Just-In-Time (JIT) static analysis that interleaves code development and bug fixing in an integrated development environment. Unlike traditional batch-style analysis tools, a JIT analysis tool presents warnings to code developers over time, providing the most relevant results quickly, and computing less relevant results incrementally later. In this paper, we describe general guidelines for designing JIT analyses. We also present a general recipe for transforming static data-flow analyses to JIT analyses through a concept of layered analysis execution. We illustrate this transformation through CHEETAH, a JIT taint analysis for Android applications. Our empirical evaluation of CHEETAH on real-world applications shows that our approach returns warnings quickly enough to avoid disrupting the normal workflow of developers. This result is confirmed by our user study, in which developers fixed data leaks twice as fast when using CHEETAH compared to an equivalent batch-style analysis.
Lisa Nguyen Quang Do, Karim Ali 0001, Benjamin Livshits, Eric Bodden, Justin Smith 0001, Emerson R. Murphy-Hill
ISSTA5
2017 Flower: Navigating program flow in the IDE
abstract
Program navigation is a critical task for software developers. State-of-the-art tools have been shown to support effective program navigation strategies, and do so by adding widgets, secondary views, and visualizations to the screen. In this work, we build on prior work by exploring what types of navigation can be supported with relatively few interface elements. To that end, we designed and implemented a prototype tool, named Flower, that supports structural program navigation while maintaining a minimalistic interface. Flower enables developers to simultaneously navigate control flow and data flow within the Eclipse Integrated Development Environment. Based on a preliminary evaluation with eight programmers, Flower succeeds when call graphs contained relatively few branches, but was strained by complex program structures.
Justin Smith 0001, Chris Brown 0001, Emerson R. Murphy-Hill
VL/HCC1
2017 Spreadsheet practices and challenges in a large multinational conglomerate
abstract
Spreadsheets are ubiquitous. Thus, it is important to understand the challenges faced by spreadsheet users in practice. To better understand these challenges, we surveyed ABB employees and then interviewed a cross-section of survey respondents. We used a two-phase coding process to classify the challenges they described. Our survey findings demonstrate that practices in our single-company setting are consistent with practices in broader settings. Our interviews revealed both individual and organizational challenges. For instance, individual participants described data pipeline challenges related to importing data from external sources or storing and archiving spreadsheet data. Further, participants' collective responses revealed challenges pertaining to knowledge distribution within the organization. We outline possible interventions to address these challenges. Our results will help guide researchers and tool designers in addressing the practical challenges facing spreadsheet users.
Justin Smith 0001, Justin A. Middleton, Nicholas A. Kraft
VL/HCC1
2016 Resolving Input Validation Vulnerabilities by Retracing Taint Flow Through Source Code
abstract
The document was not made available for publication as part of the conference proceedings.
Justin Smith 0001
ICSME1
2016 Paradise unplugged: identifying barriers for female participation on stack overflow
abstract
It is no secret that females engage less in programming fields than males. However, in online communities, such as Stack Overflow, this gender gap is even more extreme: only 5.8% of contributors are female. In this paper, we use a mixed-methods approach to identify contribution barriers females face in online communities. Through 22 semi-structured interviews with a spectrum of female users ranging from non-contributors to a top 100 ranked user of all time, we identified 14 barriers preventing them from contributing to Stack Overflow. We then conducted a survey with 1470 female and male developers to confirm which barriers are gender related or general problems for everyone. Females ranked five barriers significantly higher than males. A few of these include doubts in the level of expertise needed to contribute, feeling overwhelmed when competing with a large number of users, and limited awareness of site features. Still, there were other barriers that equally impacted all Stack Overflow users or affected particular groups, such as industry programmers. Finally, we describe several implications that may encourage increased participation in the Stack Overflow community across genders and other demographics.
Denae Ford, Justin Smith 0001, Philip J. Guo, Chris Parnin
SIGSOFT FSE2
2016 A cross-tool communication study on program analysis tool notifications
abstract
Program analysis tools use notifications to communicate with developers, but previous research suggests that developers encounter challenges that impede this communication. This paper describes a qualitative study that identifies 10 kinds of challenges that cause notifications to miscommunicate with developers. Our resulting notification communication theory reveals that many challenges span multiple tools and multiple levels of developer experience. Our results suggest that, for example, future tools that model developer experience could improve communication and help developers build more accurate mental models.
Brittany Johnson, Rahul Pandita, Justin Smith 0001, Denae Ford, Sarah Elder, Emerson R. Murphy-Hill, Sarah Smith Heckman, Caitlin Sadowski
SIGSOFT FSE3
2016 Resolving input validation vulnerabilities by retracing taint flow through source code
abstract
Various security-oriented static analysis tools are designed to detect potential input validation vulnerabilities early in the development process. To verify and resolve these vulnerabilities, developers must retrace problematic data flows through the source code. My thesis proposes that existing tools do not adequately support the navigation of these traces. In this work I will explore the strategies developers use to navigate tainted data flow in source code and work toward solutions that support successful strategies.
Justin Smith 0001
VL/HCC1
2015 Fuse: A Reproducible, Extendable, Internet-Scale Corpus of Spreadsheets
abstract
Spreadsheets are perhaps the most ubiquitous form of end-user programming software. This paper describes a corpus, called Fuse, containing 2,127,284 URLs that return spreadsheets (and their HTTP server responses), and 249,376 unique spreadsheets, contained within a public web archive of over 26.83 billion pages. Obtained using nearly 60,000 hours of computation, the resulting corpus exhibits several useful properties over prior spreadsheet corpora, including reproducibility and extendability. Our corpus is unencumbered by any license agreements, available to all, and intended for wide usage by end-user software engineering researchers. In this paper, we detail the data and the spreadsheet extraction process, describe the data schema, and discuss the trade-offs of Fuse with other corpora.
Titus Barik, Kevin Lubick, Justin Smith 0001, John Slankas, Emerson R. Murphy-Hill
MSR3
2015 Questions developers ask while diagnosing potential security vulnerabilities with static analysis
abstract
Security tools can help developers answer questions about potential vulnerabilities in their code. A better understanding of the types of questions asked by developers may help toolsmiths design more effective tools. In this paper, we describe how we collected and categorized these questions by conducting an exploratory study with novice and experienced software developers. We equipped them with Find Security Bugs, a security-oriented static analysis tool, and observed their interactions with security vulnerabilities in an open-source system that they had previously contributed to. We found that they asked questions not only about security vulnerabilities, associated attacks, and fixes, but also questions about the software itself, the social ecosystem that built the software, and related resources and tools. For example, when participants asked questions about the source of tainted data, their tools forced them to make imperfect tradeoffs between systematic and ad hoc program navigation strategies.
Justin Smith 0001, Brittany Johnson, Emerson R. Murphy-Hill, Bill Chu, Heather Lipford
ESEC/SIGSOFT FSE1
2015 A study of interactive code annotation for access control vulnerabilities
abstract
While there are a variety of existing tools to help detect security vulnerabilities in code, they are seldom used by developers due to the time or security expertise required. We are investigating techniques integrated within the IDE to help developers detect and mitigate security vulnerabilities. In this paper, we examine using interactive annotation for access control vulnerabilities. We evaluated whether developers could indicate access control logic using interactive annotation and understand the vulnerabilities reported as a result. Our study indicates that developers can easily find and annotate access control logic but can struggle to use our tool to trace the cause of the vulnerability. Our results provide design guidance for improving the interaction and communication of such security tools with developers.
Tyler Thomas, Bill Chu, Heather Lipford, Justin Smith 0001, Emerson R. Murphy-Hill
VL/HCC4