EDBT 2026 Demo / reviewers in the wild / expert
Iftekhar Ahmed 0001
dblp:57/5553
· DBLP profile ↗
52ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0001-8221-5352ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 44 · 5 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Repair of Alloy Specifications in the Era of Large Language Models
Md Rashedul Hasan, Jiawei Li 0013, Iftekhar Ahmed 0001, Hamid Bagheri |
IEEE Trans. Software Eng. | 3 |
| 2025 | 'It's a spectrum': Exploring Autonomy, Competence, and Relatedness in Software Development Processes and Tools
Novia Wong, Nai-Yu Cheng, Bruna Oewel, Katherine E. Genuario, SarahElizabeth Stoeckl, Stephen M. Schueller, Iftekhar Ahmed 0001, André van der Hoek, Madhu C. Reddy |
CHI | 7 |
| 2025 | Context Conquers Parameters: Outperforming Proprietary Llm in Commit Message GenerationabstractCommit messages provide descriptions of the modifications made in a commit using natural language, making them crucial for software maintenance and evolution. Recent developments in Large Language Models (LLMs) have led to their use in generating high-quality commit messages, such as the Omniscient Message Generator (OMG). This method employs GPT-4 to produce state-of-the-art commit messages. However, the use of proprietary LLMs like GPT-4 in coding tasks raises privacy and sustainability concerns, which may hinder their industrial adoption. Considering that open-source LLMs have achieved competitive performance in developer tasks such as compiler validation, this study investigates whether they can be used to generate commit messages that are comparable with OMG. Our experiments show that an open-source LLM can generate commit messages comparable to those produced by OMG. In addition, through a series of contextual refinements, we propose OMEGA, a commit message generation approach that uses a 4-bit quantized 8B open-source LLM. OMEGA produces state-of-the-art commit messages, surpassing the performance of GPT-4 in practitioners' preference. Aaron Imani, Iftekhar Ahmed 0001, Mohammad Moshirpour |
ICSE | 2 |
| 2025 | An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far are We?abstractArtificial Intelligence (AI) techniques, especially Large Language Models (LLMs), have started gaining popularity among researchers and software developers for generating source code. However, LLMs have been shown to generate code with quality issues and also incurred copyright/licensing infringements. Therefore, detecting whether a piece of source code is written by humans or AI has become necessary. This study first presents an empirical analysis to investigate the effectiveness of the existing AI detection tools in detecting AI-generated code. The results show that they all perform poorly and lack sufficient generalizability to be practically deployed. Then, to improve the performance of AI-generated code detection, we propose a range of approaches, including fine-tuning the LLMs and machine learning-based classification with static code metrics or code embedding generated from Abstract Syntax Tree (AST). Our best model outperforms state-of-the-art AI-generated code detector (GPTSniffer) and achieves an F1 score of 82.55. We also conduct an ablation study on our best-performing model to investigate the impact of different source code features on its performance. Hyunjae Suh, Mahan Tafreshipour, Jiawei Li 0013, Adithya Bhattiprolu, Iftekhar Ahmed 0001 |
ICSE | 5 |
| 2025 | Prompting in the Wild: An Empirical Study of Prompt Evolution in Software RepositoriesabstractThe adoption of Large Language Models (LLMs) is reshaping software development as developers integrate these LLMs into their applications. In such applications, prompts serve as the primary means of interacting with LLMs. Despite the widespread use of LLM-integrated applications, there is limited understanding of how developers manage and evolve prompts. This study presents the first empirical analysis of prompt evolution in LLM-integrated software development. We analyzed 1,262 prompt changes across 243 GitHub repositories to investigate the patterns and frequencies of prompt changes, their relationship with code changes, documentation practices, and their impact on system behavior. Our findings show that developers primarily evolve prompts through additions and modifications, with most changes occurring during feature development. We identified key challenges in prompt engineering: only $21.9 \%$ of prompt changes are documented in commit messages, changes can introduce logical inconsistencies, and misalignment often occurs between prompt changes and LLM responses. These insights emphasize the need for specialized testing frameworks, automated validation tools, and improved documentation practices to enhance the reliability of LLM-integrated applications. Mahan Tafreshipour, Aaron Imani, Eduardo Santana de Almeida, Thomas Zimmermann 0001, Iftekhar Ahmed 0001 |
MSR | 6 |
| 2025 | Test smell: A parasitic energy consumer in software testingabstractTraditionally, energy efficiency research has focused on reducing energy consumption at the hardware level and, more recently, in the design and coding phases of the software development life cycle. However, software testing’s impact on energy consumption did not receive attention from the research community. Specifically, how test code design quality and test smell (e.g., sub-optimal design and bad practices in test code) impact energy consumption has not been investigated yet. This study aims to examine open-source software projects to analyze the association between test smell and its effects on energy consumption in software testing. We conducted a mixed-method empirical analysis from two perspectives; software (data mining in 12 Apache projects) and developers’ views (a survey of 62 software practitioners). Our findings show that: (1) test smell is associated with energy consumption in software testing. Specifically, the smelly part of a test case consumes more energy compared to the non-smelly part. (2) certain test smells are more energy-hungry than others, (3) refactored test cases tend to consume less energy than their smelly counterparts, and (4) most developers (45 % of the survey respondents) lack knowledge about test smells’ impact on energy consumption. Based on the results, we emphasize raising developers awareness regarding the impact of test smells on energy consumption. Additionally we present several observations that can direct future research and developments. Md Rakib Hossain Misu, Jiawei Li 0013, Adithya Bhattiprolu, Eduardo Santana de Almeida, Iftekhar Ahmed 0001 |
Inf. Softw. Technol. | 6 |
| 2025 | Using AI-based coding assistants in practice: State of affairs, perceptions, and ways forward
Agnia Sergeyuk, Yaroslav Golubev, Timofey Bryksin, Iftekhar Ahmed 0001 |
Inf. Softw. Technol. | 4 |
| 2025 | Artificial Intelligence for Software Engineering: The Journey So Far and the Road AheadabstractArtificial intelligence and recent advances in deep learning architectures, including transformer networks and large language models, change the way people think and act to solve problems. Software engineering, as an increasingly complex process to design, develop, test, deploy, and maintain large-scale software systems for solving real-world challenges, is profoundly affected by many revolutionary artificial intelligence tools in general and machine learning in particular. In this roadmap for artificial intelligence in software engineering, we highlight the recent deep impact of artificial intelligence on software engineering by discussing successful stories of applications of artificial intelligence to classic and new software development challenges. We identify the new challenges that the software engineering community has to address in the coming years to successfully apply artificial intelligence in software engineering, and we share our research roadmap toward the effective use of artificial intelligence in the software engineering profession, while still protecting fundamental human values. We spotlight three main areas that challenge the research in software engineering: the use of generative artificial intelligence and large language models for engineering large software systems, the need of large and unbiased datasets and benchmarks for training and evaluating deep learning and large language models for software engineering, and the need of a new code of digital ethics to apply artificial intelligence in software engineering. Iftekhar Ahmed 0001, Aldeida Aleti, Haipeng Cai, Alexander Chatzigeorgiou, Pinjia He, Xing Hu 0008, Mauro Pezzè, Denys Poshyvanyk, Xin Xia 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | What Makes a Great Software Quality Assurance Engineer?abstractSoftware Quality Assurance (SQA) Engineers play a critical role in evaluating products throughout the software development lifecycle to ensure that the outcomes of each phase and the final product possess the desired quality standards. In general, a great SQA engineer requires a different set of abilities from development engineers to effectively oversee the entire product development process. While recent empirical studies have explored the attributes of software engineers and managers, the quality assurance role is overlooked. As software quality gains increasing priority in the development cycles, both employers seeking skilled professionals and new graduates aspiring to excel in Software Quality Assurance (SQA) roles face a critical question: What makes a great SQA Engineer? To address this gap, we conducted 25 semi-structured interviews and surveyed 363 SQA engineers from diverse companies worldwide. We use the data collected from these activities to derive a comprehensive set of attributes for great SQA Engineers, categorized into five key areas: personal, social, technical, management, and decision-making attributes. Among these, curiosity, effective communication, and critical thinking emerged as defining characteristics of great SQA engineers. These findings offer valuable insights for future research with SQA practitioners, contextual considerations, and practical implications for research and practice. Roselane Silva, Iftekhar Ahmed 0001, Eduardo Santana de Almeida |
IEEE Trans. Software Eng. | 2 |
| 2024 | Ma11y: A Mutation Framework for Web Accessibility TestingabstractDespite the availability of numerous automatic accessibility testing solutions, web accessibility issues persist on many websites. Moreover, there is a lack of systematic evaluations of the efficacy of current accessibility testing tools. To address this gap, we present the first mutation analysis framework, called Ma11y, designed to assess web accessibility testing tools. Ma11y includes 25 mutation operators that intentionally violate various accessibility principles and an automated oracle to determine whether a mutant is detected by a testing tool. Evaluation on real-world websites demonstrates the practical applicability of the mutation operators and the framework’s capacity to assess tool performance. Our results demonstrate that the current tools cannot identify nearly 50% of the accessibility bugs injected by our framework, thus underscoring the need for the development of more effective accessibility testing tools. Finally, the framework’s accuracy and performance attest to its potential for seamless and automated application in practical settings. Mahan Tafreshipour, Anmol Vilas Deshpande, Forough Mehralian, Iftekhar Ahmed 0001, Sam Malek |
ISSTA | 4 |
| 2024 | Carving Out Control Code: Automated Identification of Control Software in Autopilot SystemsabstractCyber-physical systems interact with the world through software controlling physical effectors. Carefully designed controllers, implemented as safety-critical control software, also interact with other parts of the software suite, and may be difficult to separate, verify, or maintain. Moreover, some software changes, not intended to impact control system performance, do change controller response through a variety of means including interaction with external libraries or unmodeled changes only existing in the cyber system (e.g., exception handling). As a result, identifying safety-critical control software, its boundaries with other embedded software in the system, and the way in which control software evolves could help developers isolate, test, and verify control implementation, and improve control software development. In this work we present an automated technique, based on a novel application of machine learning, to detect commits related to control software, its changes, and how the control software evolves. We leverage messages from developers (e.g., commit comments), and code changes themselves to understand how control software is refined, extended, and adapted over time. We examine three distinct, popular, real-world, safety-critical autopilots—ArduPilot, Paparazzi UAV, and LibrePilot to test our method demonstrating an effective detection rate of 0.95 for control-related code changes. Balaji Balasubramaniam, Iftekhar Ahmed 0001, Hamid Bagheri, Justin M. Bradley |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2024 | Bug Analysis in Jupyter Notebook Projects: An Empirical StudyabstractComputational notebooks, such as Jupyter, have been widely adopted by data scientists to write code for analyzing and visualizing data. Despite their growing adoption and popularity, few studies have been found to understand Jupyter development challenges from the practitioners’ point of view. This article presents a systematic study of bugs and challenges that Jupyter practitioners face through a large-scale empirical investigation. We mined 14,740 commits from 105 GitHub open source projects with Jupyter Notebook code. Next, we analyzed 30,416 StackOverflow posts, which gave us insights into bugs that practitioners face when developing Jupyter Notebook projects. Next, we conducted 19 interviews with data scientists to uncover more details about Jupyter bugs and to gain insight into Jupyter developers’ challenges. Finally, to validate the study results and proposed taxonomy, we conducted a survey with 91 data scientists. We highlight bug categories, their root causes, and the challenges that Jupyter practitioners face. Taijara Loiola de Santana, Paulo Anselmo da Mota Silveira Neto, Eduardo Santana de Almeida, Iftekhar Ahmed 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Mental Wellbeing at Work: Perspectives of Software EngineersabstractSoftware engineers exhibit higher burnout and suicide rates compared to many other information workers. Consequently, mental wellbeing is a growing concern to technology organizations. To better understand the challenges of supporting mental wellbeing in the context of the work of software engineering, we conducted 14 interviews with software engineers. We examine the different aspects of their lived experiences with mental wellbeing at work, their strategies for managing mental wellbeing, the challenges they face in using these strategies, and recommendations they have for mental wellbeing technologies. We contribute to the HCI literature by discussing how mental wellbeing should be considered within the context of work across individual, team, and organization levels, and highlight the need for integrating mental wellbeing into the technologies employees use at work. Novia Wong, Victoria Jackson, André van der Hoek, Iftekhar Ahmed 0001, Stephen M. Schueller, Madhu C. Reddy |
CHI | 4 |
| 2023 | The Smelly Eight: An Empirical Study on the Prevalence of Code Smells in Quantum ComputingabstractQuantum Computing (QC) is a fast-growing field that has enhanced the emergence of new programming languages and frameworks. Furthermore, the increased availability of computational resources has also contributed to an influx in the development of quantum programs. Given that classical and QC are significantly different due to the intrinsic nature of quantum programs, several aspects of QC (e.g., performance, bugs) have been investigated, and novel approaches have been proposed. However, from a purely quantum perspective, maintenance, one of the major steps in a software development life-cycle, has not been considered by researchers yet. In this paper, we fill this gap and investigate the prevalence of code smells in quantum programs as an indicator of maintenance issues. We defined eight quantum-specific smells and validated them through a survey with 35 quantum developers. Since no tool specifically aims to detect quantum smells, we developed a tool called QSmell that supports the proposed quantum-specific smells. Finally, we conducted an empirical investigation to analyze the prevalence of quantum-specific smells in 15 open-source quantum programs. Our results showed that 11 programs (73.33%) contain at least one smell and, on average, a program has three smells. Furthermore, the long circuit is the most prevalent smell present in 53.33% of the programs. Qihong Chen, Rúben Câmara, José Campos 0001, André Souto, Iftekhar Ahmed 0001 |
ICSE | 5 |
| 2023 | Leveraging Feature Bias for Scalable Misprediction Explanation of Machine Learning ModelsabstractInterpreting and debugging machine learning models is necessary to ensure the robustness of the machine learning models. Explaining mispredictions can help significantly in doing so. While recent works on misprediction explanation have proven promising in generating interpretable explanations for mispredictions, the state-of-the-art techniques “blindly” deduce misprediction explanation rules from all data features, which may not be scalable depending on the number of features. To alleviate this problem, we propose an efficient misprediction explanation technique named Bias Guided Misprediction Diagnoser (BGMD), which leverages two prior knowledge about data: a) data often exhibit highly-skewed feature distributions and b) trained models in many cases perform poorly on subdataset with under-represented features. Next, we propose a technique named MAPS (Mispredicted Area UPweight Sampling). MAPS increases the weights of subdataset during model retraining that belong to the group that is prone to be mispredicted because of containing under-represented features. Thus, MAPS make retrained model pay more attention to the under-represented features. Our empirical study shows that our proposed BGMD outperformed the state-of-the-art misprediction diagnoser and reduces diagnosis time by 92%. Furthermore, MAPS outperformed two state-of-the-art techniques on fixing the machine learning model's performance on mispredicted data without compromising performance on all data. All the research artifacts (i.e., tools, scripts, and data) of this study are available in the accompanying website [1]. Jiri Gesi, Xinyun Shen, Yunfan Geng, Qihong Chen, Iftekhar Ahmed 0001 |
ICSE | 5 |
| 2023 | Commit Message Matters: Investigating Impact and Evolution of Commit Message QualityabstractCommit messages play an important role in communication among developers. To measure the quality of commit messages, researchers have defined what semantically constitutes a Good commit message: it should have both the summary of the code change (What) and the motivation/reason behind it (Why). The presence of the issue report/pull request links referenced in a commit message has been treated as a way of providing Why information. In this study, we found several quality issues that could hamper the links' ability to provide Why information. Based on this observation, we developed a machine learning classifier for automatically identifying whether a commit message has What and Why information by considering both the commit messages and the link contents. This classifier outperforms state-of-the-art machine learning classifiers by 12 percentage points improvement in the F1 score. With the improved classifier, we conducted a mixed method empirical analysis and found that: (1) Commit message quality has an impact on software defect proneness, and (2) the overall quality of the commit messages decreases over time, while developers believe they are writing better commit messages. All the research artifacts (i.e., tools, scripts, and data) of this study are available on the accompanying website [2]. Jiawei Li 0013, Iftekhar Ahmed 0001 |
ICSE | 2 |
| 2023 | Aligning Documentation and Q&A Forum through Constrained Decoding with Weak SupervisionabstractStack Overflow (SO) is a widely used question-and-answer (Q&A) forum dedicated to software development. It plays a supplementary role to official documentation (DOC for short) by offering practical examples and resolving uncertainties. However, the process of simultaneously consulting both the documentation and SO posts can be challenging and time-consuming due to their disconnected nature. In this study, we propose DOSA, a novel approach to automatically align SO and DOC, which inject domain-specific knowledge about the DOC structure into large language models (LLMs) through weak supervision and constrained decoding, thereby enhancing knowledge retrieval and streamlining task completion during the software development procedure. Our preliminary experiments find that DOSA outperforms various widely-used baselines, showing the promise of using generative retrieval models to perform low-resource software engineering tasks. Rohith Pudari, Shiyuan Zhou, Iftekhar Ahmed 0001, Zhuyun Dai, Shurui Zhou |
ICSME | 3 |
| 2023 | Let's Go to the Whiteboard (Again): Perceptions From Software Architects on Whiteboard Architecture MeetingsabstractThe whiteboard plays a crucial role in the day-to-day lives of software architects, as they frequently will organize meetings at the whiteboard to discuss a new architecture, some proposed changes to an existing architecture, a mismatch between a prescribed architecture and its code, and more. While much has been studied about software architects, the architectures they produce, and how they produce them, a detailed understanding of these whiteboards meetings is still lacking. In this paper, we contribute a mixed-methods study involving semi-structured interviews and a subsequent survey to understand the perceptions of software architects on whiteboard architecture meetings. We focus on four aspects: (1) why do they hold these meetings, (2) what is the impact of the experience levels of the participants in these meetings, (3) how do the architects document the meetings, and (4) what kinds of changes are made in downstream activities to the work produced after the meetings have concluded? In studying these aspects, we identify eleven observations related to both technical aspects and social aspects of the meetings. These insights have implications for further research, offer concrete advice to practitioners, and suggest ways of educating future software architects. Eduardo Santana de Almeida, Iftekhar Ahmed 0001, André van der Hoek |
IEEE Trans. Software Eng. | 2 |
| 2022 | A case study of implicit mentoring, its prevalence, and impact in ApacheabstractMentoring is traditionally viewed as a dyadic, top-down apprenticeship. This perspective, however, overlooks other forms of informal mentoring taking place in everyday activities in which developers invest time and effort. Here, we investigate informal mentoring taking place in Open Source Software (OSS). We define a specific type of informal mentoring—implicit mentoring—situations where contributors guide others through instructions and suggestions embedded in everyday (OSS) activities. We defined implicit mentoring by first performing a review of related work on mentoring, and then through formative interviews with OSS contributors and member-checking. Next, through an empirical investigation of Pull Requests (PRs) in 37 Apache Projects, we built a classifier to extract implicit mentoring. Our analysis of 107,895 PRs shows that implicit mentoring does occur through code reviews (27.41% of all PRs included implicit mentoring) and is beneficial for both mentors and mentees. We analyzed the impact of implicit mentoring on OSS contributors by investigating their contributions and learning trajectories in their projects. Through an online survey (N=231), we then triangulated these results and identified the potential benefits of implicit mentoring from OSS contributors’ perspectives. Amreeta Chatterjee, Anita Sarma, Iftekhar Ahmed 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2022 | Different, Really! A comparison of Highly-Configurable Systems and Single Systems
Raphael Pereira de Oliveira, Paulo Anselmo da Mota Silveira Neto, Qihong Chen, Eduardo Santana de Almeida, Iftekhar Ahmed 0001 |
Inf. Softw. Technol. | 5 |
| 2022 | An empirical study of emoji use in software development communication
Shiyue Rong, Weisheng Wang, Umme Ayda Mannan, Eduardo Santana de Almeida, Shurui Zhou, Iftekhar Ahmed 0001 |
Inf. Softw. Technol. | 6 |
| 2022 | Investigating replication challenges through multiple replications of an experiment
Daniel Amador dos Santos, Eduardo Santana de Almeida, Iftekhar Ahmed 0001 |
Inf. Softw. Technol. | 3 |
| 2022 | A Deep Dive into the Impact of COVID-19 on Software DevelopmentabstractThe COVID-19 pandemic is considered as the most crucial global health calamity of the century. It has impacted different business sectors around the world and software development is not an exception. This study investigates the impact of COVID-19 on software projects and software development professionals. We conducted a mining software repository study based on 100 GitHub projects developed in Java using ten different metrics. Next, we surveyed 279 software development professionals for better understanding the impact of COVID-19 on daily activities and wellbeing. We identified 12 observations related to productivity, code quality, and wellbeing. Our findings highlight that the impact of COVID-19 is not binary (reduce productivity versus increase productivity) but rather a spectrum. For many of our observations, substantial proportions of respondents have differing opinions from each other. We believe that more research is needed to uncover specific conditions that cause certain outcomes to be more prevalent. Paulo Anselmo da Mota Silveira Neto, Umme Ayda Mannan, Eduardo Santana de Almeida, Nachiappan Nagappan, David Lo 0001, Pavneet Singh Kochhar, Cuiyun Gao 0001, Iftekhar Ahmed 0001 |
IEEE Trans. Software Eng. | 8 |
| 2021 | Latte: Use-Case and Assistive-Service Driven Automated Accessibility Testing Framework for AndroidabstractFor 15% of the world population with disabilities, accessibility is arguably the most critical software quality attribute. The ever-growing reliance of users with disability on mobile apps further underscores the need for accessible software in this domain. Existing automated accessibility assessment techniques primarily aim to detect violations of predefined guidelines, thereby produce a massive amount of accessibility warnings that often overlook the way software is actually used by users with disability. This paper presents a novel, high-fidelity form of accessibility testing for Android apps, called Latte, that automatically reuses tests written to evaluate an app’s functional correctness to assess its accessibility as well. Latte first extracts the use case corresponding to each test, and then executes each use case in the way disabled users would, i.e., using assistive services. Our empirical evaluation on real-world Android apps demonstrates Latte’s effectiveness in detecting substantially more useful defects than prior techniques. Navid Salehnamadi, Abdulaziz Alshayban, Jun-Wei Lin, Iftekhar Ahmed 0001, Stacy M. Branham, Sam Malek |
CHI | 4 |
| 2021 | An Empirical Examination of the Impact of Bias on Just-in-time Defect PredictionabstractBackground: Just-In-Time (JIT) defect prediction models predict if a commit will introduce defects in the future. DeepJIT and CC2Vec are two state-of-the-art JIT Deep Learning (DL) techniques. Usually, defect prediction techniques are evaluated, treating all training data equally. However, data is usually imbalanced not only in terms of the overall class label (e.g., defect and non-defect) but also in terms of characteristics such as File Count, Edit Count, Multiline Comments, Inward Dependency Sum etc. Prior research has investigated the impact of class imbalance on prediction technique's performance but not the impact of imbalance of other characteristics. Aims: We aim to explore the impact of different commit related characteristic's imbalance on DL defect prediction. Method: We investigated different characteristic's impact on the overall performance of DeepJIT and CC2Vec. We also propose a Siamese network based few-shot learning framework for JIT defect prediction (SifterJIT) combining Siamese network and DeepJIT. Results: Our results show that DeepJIT and CC2Vec lose out on the performance by around 20% when trained and tested on imbalanced data. However, SifterJIT can outperform state-of-the-art DL techniques with an average of 8.65% AUC score, 11% precision, and 6% F1-score improvement. Conclusions: Our results highlight that dataset imbalanced in terms of commit characteristics can significantly impact prediction performance, and few-shot learning based techniques can help alleviate the situation. Jiri Gesi, Jiawei Li 0013, Iftekhar Ahmed 0001 |
ESEM | 3 |
| 2021 | AID: An automated detector for gender-inclusivity bugs in OSS project pagesabstractThe tools and infrastructure used in tech, including Open Source Software (OSS), can embed "inclusivity bugs"- features that disproportionately disadvantage particular groups of contributors. To see whether OSS developers have existing practices to ward off such bugs, we surveyed 266 OSS developers. Our results show that a majority (77%) of developers do not use any inclusivity practices, and 92% of respondents cited a lack of concrete resources to enable them to do so. To help fill this gap, this paper introduces AID, a tool that automates the GenderMag method to systematically find gender-inclusivity bugs in software. We then present the results of the tool's evaluation on 20 GitHub projects. The tool achieved precision of 0.69, recall of 0.92, an F-measure of 0.79 and even captured some inclusivity bugs that human GenderMag teams missed. Amreeta Chatterjee, Mariam Guizani, Catherine Stevens, Jillian Emard, Mary Evelyn May, Margaret M. Burnett, Iftekhar Ahmed 0001, Anita Sarma |
ICSE | 7 |
| 2021 | We'll Fix It in Post: What Do Bug Fixes in Video Game Update Notes Tell Us?abstractBugs that persist into releases of video games can have negative impacts on both developers and users, but particular aspects of testing in game development can lead to difficulties in effectively catching these missed bugs. It has become common practice for developers to apply updates to games in order to fix missed bugs. These updates are often accompanied by notes that describe the changes to the game included in the update. However, some bugs reappear even after an update attempts to fix them. In this paper, we develop a taxonomy for bug types in games that is based on prior work. We examine 12,122 bug fixes from 723 updates for 30 popular games on the Steam platform. We label the bug fixes included in these updates to identify the frequency of these different bug types, the rate at which bug types recur over multiple updates, and which bug types are treated as more severe. Additionally, we survey game developers regarding their experience with different bug types and what aspects of game development they most strongly associate with bug appearance. We find that Information bugs appear the most frequently in updates, while Crash bugs recur the most frequently and are often treated as more severe than other bug types. Finally, we find that challenges in testing, code quality, and bug reproduction have a close association with bug persistence. These findings should help developers identify which aspects of game development could benefit from greater attention in order to prevent bugs. Researchers can use our results in devising tools and methods to better identify and address certain bug types. Andrew Truelove, Eduardo Santana de Almeida, Iftekhar Ahmed 0001 |
ICSE | 3 |
| 2021 | PyNose: A Test Smell Detector For PythonabstractSimilarly to production code, code smells also occur in test code, where they are called test smells. Test smells have a detrimental effect not only on test code but also on the production code that is being tested. To date, the majority of the research on test smells has been focusing on programming languages such as Java and Scala. However, there are no available automated tools to support the identification of test smells for Python, despite its rapid growth in popularity in recent years. In this paper, we strive to extend the research to Python, build a tool for detecting test smells in this language, and conduct an empirical analysis of test smells in Python projects.We started by gathering a list of test smells from existing research and selecting test smells that can be considered language-agnostic or have similar functionality in Python’s standard Unittest framework. In total, we identified 17 diverse test smells. Additionally, we searched for Python-specific test smells by mining frequent code change patterns that can be considered as either fixing or introducing test smells. Based on these changes, we proposed our own test smell called Suboptimal Assert. To detect all these test smells, we developed a tool called PYNOSE in the form of a plugin to PyCharm, a popular Python IDE. Finally, we conducted a large-scale empirical investigation aimed at analyzing the prevalence of test smells in Python code. Our results show that 98% of the projects and 84% of the test suites in the studied dataset contain at least one test smell. Our proposed Suboptimal Assert smell was detected in as much as 70.6% of the projects, making it a valuable addition to the list. Tongjie Wang, Yaroslav Golubev, Oleg Smirnov, Jiawei Li 0013, Timofey Bryksin, Iftekhar Ahmed 0001 |
ASE | 6 |
| 2021 | Evaluating and Improving Static Analysis Tools Via Differential Mutation AnalysisabstractStatic analysis tools attempt to detect faults in code without executing it. Understanding the strengths and weaknesses of such tools, and performing direct comparisons of their ef-fectiveness, is difficult, involving either manual examination of differing warnings on real code, or the bias-prone construction of artificial test cases. This paper proposes a novel automated approach to comparing static analysis tools, based on producing mutants of real code, and comparing detection rates over these mutants. In addition to making tool differences quantitatively observable without extensive manual effort, this approach offers a new way to detect and fix omissions in a static analysis tool's set of detectors. We present an extensive comparison of three smart contract static analysis tools, and show how our approach allowed us to add three effective new detectors to the best of these. We also evaluate popular Java and Python static analysis tools and discuss their strengths and weaknesses. Alex Groce, Iftekhar Ahmed 0001, Josselin Feist, Gustavo Grieco, Jiri Gesi, Mehran Meidani, Qihong Chen |
QRS | 2 |
| 2020 | A Multiple Case Study of Artificial Intelligent System Development in IndustryabstractThere is a rapidly increasing amount of Artificial Intelligence (AI) systems developed in recent years, with much expectation on its capacity of innovation and business value generation. However, the promised value of AI systems in specific business contexts might not be understood, and further integrated into the development processes. We wanted to understand how software engineering processes and practices can be applied to develop AI systems in a fast-faced, business-driven manner. As the first step, we explored contextual factors of AI development and the connections between AI developments to business opportunities. We conducted 12 semi-structured interviews in seven companies in Brazil, Norway and Southeast Asia. Our investigation revealed different types of AI systems and different AI development approaches. However, it is common that business opportunities involving with AI systems are not validated and there is lack of business-driven metrics that guide the development of AI systems. The findings have implications for future research on business-driven AI development and supporting tools and practices. Anh Nguyen-Duc 0001, Ingrid Sundbø, Elizamary Nascimento, Tayana Conte, Iftekhar Ahmed 0001, Pekka Abrahamsson |
EASE | 5 |
| 2020 | Accessibility issues in Android apps: state of affairs, sentiments, and ways forwardabstractMobile apps are an integral component of our daily life. Ability to use mobile apps is important for everyone, but arguably even more so for approximately 15% of the world population with disabilities. This paper presents the results of a large-scale empirical study aimed at understanding accessibility of Android apps from three complementary perspectives. First, we analyze the prevalence of accessibility issues in over 1, 000 Android apps. We find that almost all apps are riddled with accessibility issues, hindering their use by disabled people. We then investigate the developer sentiments through a survey aimed at understanding the root causes of so many accessibility issues. We find that in large part developers are unaware of accessibility design principles and analysis tools, and the organizations in which they are employed do not place a premium on accessibility. We finally investigate user ratings and comments on app stores. We find that due to the disproportionately small number of users with disabilities, user ratings and app popularity are not indicative of the extent of accessibility issues in apps. We conclude the paper with several observations that form the foundation for future research and development. Abdulaziz Alshayban, Iftekhar Ahmed 0001, Sam Malek |
ICSE | 2 |
| 2020 | Planning for untangling: predicting the difficulty of merge conflictsabstractMerge conflicts are inevitable in collaborative software development and are disruptive. When they occur, developers have to stop their current work, understand the conflict and the surrounding code, and plan an appropriate resolution. However, not all conflicts are equally problematic---some can be easily fixed, while others might be complicated enough to need multiple people. Currently, there is not much support to help developers plan their conflict resolution. In this work, we aim to predict the difficulty of a merge conflict so as to help developers plan their conflict resolution. The ability to predict the difficulty of a merge conflict and to identify the underlying factors for its difficulty can help tool builders improve their conflict detection tools to prioritize and warn developers of difficult conflicts. In this work, we investigate the characteristics of difficult merge conflicts, and automatically classify them. We analyzed 6,380 conflicts across 128 java projects and found that merge conflict difficulty can be accurately predicted (AUC of 0.76) through machine learning algorithms, such as bagging. Caius Brindescu, Iftekhar Ahmed 0001, Rafael Leano, Anita Sarma |
ICSE | 2 |
| 2020 | ER Catcher: A Static Analysis Framework for Accurate and Scalable Event-Race Detection in AndroidabstractAndroid platform provisions a number of sophisticated concurrency mechanisms for the development of apps. The concurrency mechanisms, while powerful, are quite difficult to properly master by mobile developers. In fact, prior studies have shown concurrency issues, such as event-race defects, to be prevalent among real-world Android apps. In this paper, we propose a flow-, context-, and thread-sensitive static analysis framework, called ER Catcher, for detection of event-race defects in Android apps. ER Catcher introduces a new type of summary function aimed at modeling the concurrent behavior of methods in both Android apps and libraries. In addition, it leverages a novel, statically constructed Vector Clock for rapid analysis of happens-before relations. Altogether, these design choices enable ER Catcher to not only detect event-race defects with a substantially higher degree of accuracy, but also in a fraction of time compared to the existing state-of-the-art technique. Navid Salehnamadi, Abdulaziz Alshayban, Iftekhar Ahmed 0001, Sam Malek |
ASE | 3 |
| 2020 | A benchmark for event-race analysis in android appsabstractOver the past few years, researchers have proposed various program analysis tools for automated detection of event-race conditions in Android. However, to this date, it is not clear how these tools compare to one another, as they have been evaluated on arbitrary, disjointed set of Android apps, for which there is no ground truth, i.e., verified set of event races. To fill this gap and support future research in this area, we introduce BenchERoid, a set of 34 Android apps with injected event-race bugs. The current version of benchmark contains 36 types of event-race bugs that were identified by analyzing Android concurrency literature and publicly available issue repositories. We believe that our framework is a valuable resource for both developers and researchers interested in concurrency bug analysis in Android. BenchERoid is publicly available at: https://github.com/seal-hub/bencheroid. Navid Salehnamadi, Abdulaziz Alshayban, Iftekhar Ahmed 0001, Sam Malek |
MobiSys | 3 |
| 2020 | On the relationship between design discussions and design quality: a case study of Apache projectsabstractOpen design discussion is a primary mechanism through which open source projects debate, make and document design decisions. However, there are open questions regarding how design discussions are conducted and what effect they have on the design quality of projects. Recent work has begun to investigate design discussions, but has thus far focused on a single communication channel, whereas many projects use multiple channels. In this study, we examine 37 Apache projects and their design discussions, the project’s design quality evolution, and the relationship between design discussion and design quality. A mixed method empirical analysis (data mining and a survey of 130 developers) shows that: I) 89.51% of all design discussions occur in project mailing list, II) both core and non-core developers participate in design discussions, but core developers implement more design related changes (67.06%), and III) the correlation between design discussions and design quality is small. We conclude the paper with several observations that form the foundation for future research and development. Umme Ayda Mannan, Iftekhar Ahmed 0001, Carlos Jensen, Anita Sarma |
ESEC/SIGSOFT FSE | 2 |
| 2020 | An empirical investigation into merge conflicts and their effect on software quality
Caius Brindescu, Iftekhar Ahmed 0001, Carlos Jensen, Anita Sarma |
Empir. Softw. Eng. | 2 |
| 2020 | Using Relative Lines of Code to Guide Automated Test Generation for PythonabstractRaw lines of code (LOC) is a metric that does not, at first glance, seem extremely useful for automated test generation. It is both highly language-dependent and not extremely meaningful, semantically, within a language: one coder can produce the same effect with many fewer lines than another. However, relative LOC , between components of the same project, turns out to be a highly useful metric for automated testing. In this article, we make use of a heuristic based on LOC counts for tested functions to dramatically improve the effectiveness of automated test generation. This approach is particularly valuable in languages where collecting code coverage data to guide testing has a very high overhead. We apply the heuristic to property-based Python testing using the TSTL (Template Scripting Testing Language) tool. In our experiments, the simple LOC heuristic can improve branch and statement coverage by large margins (often more than 20%, up to 40% or more) and improve fault detection by an even larger margin (usually more than 75% and up to 400% or more). The LOC heuristic is also easy to combine with other approaches and is comparable to, and possibly more effective than, two well-established approaches for guiding random testing. Josie Holmes, Iftekhar Ahmed 0001, Caius Brindescu, Rahul Gopinath, He Zhang 0025, Alex Groce |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2019 | Land of Lost Knowledge: An Initial Investigation into Projects Lost KnowledgeabstractBackground: Software development teams adopt various communication tools to support coordination and team interaction during the software development process. Among many other communication channels, developers' use instant messaging to discuss ideas, decisions and other project related issues with team members. Due to the informal nature of instant messaging, many of these discussions and decisions are lost. This situation could be even more critical in startups and other software companies that rely more heavily on instant message tools or other informal communication channels.Aims: This work investigates the effectiveness of using a semiautomatic approach for identifying, extracting, and determining a project's lost knowledge that was discussed using unstructured communication tools such as instant message.Methodology: We employed data-mining techniques to automatically retrieve discussions from instant message logs and showed them to the project managers to identify lost knowledge from two startup companies.Results: Our results demonstrate that the data-mining technique was capable of retrieving sentences with relevant issues discussion; reaching a precision of 75% at the first 10 relevant sentences evaluated. Moreover, the qualitative analysis conducted involving project managers shows an association of retrieved sentences with the project's lost knowledge.Conclusion: Our findings indicate that automated approaches can be used to identify such lost knowledge in software development projects. Follow-up interviews revealed the interest of PMs in adopting such automated tools in other projects. Márcia Lima, Iftekhar Ahmed 0001, Tayana Conte, Elizamary Nascimento, Edson Oliveira 0001, Bruno Gadelha |
ESEM | 2 |
| 2019 | Understanding Development Process of Machine Learning Systems: Challenges and SolutionsabstractBackground: The number of Machine Learning (ML) systems developed in the industry is increasing rapidly. Since ML systems are different from traditional systems, these differences are clearly visible in different activities pertaining to ML systems software development process. These differences make the Software Engineering (SE) activities more challenging for ML systems because not only the behavior of the system is data dependent, but also the requirements are data dependent. In such scenario, how can Software Engineering better support the development of ML systems? Aim: Our objective is twofold. First, better understand the process that developers use to build ML systems. Second, identify the main challenges that developers face, proposing ways to overcome these challenges. Method: We conducted interviews with seven developers from three software small companies that develop ML systems. Based on the challenges uncovered, we proposed a set of checklists to support the developers. We assessed the checklists by using a focus group. Results: We found that the ML systems development follow a 4-stage process in these companies. These stages are: understanding the problem, data handling, model building, and model monitoring. The main challenges faced by the developers are: identifying the clients' business metrics, lack of a defined development process, and designing the database structure. We have identified in the focus group that our proposed checklists provided support during identification of the client's business metrics and in increasing visibility of the progress of the project tasks. Conclusions: Our research is an initial step towards supporting the development of ML systems, suggesting checklists that support developers in essential development tasks, and also serve as a basis for future research in the area. Elizamary Nascimento, Iftekhar Ahmed 0001, Edson Oliveira 0001, Márcio Piedade Palheta, Igor Steinmacher, Tayana Conte |
ESEM | 2 |
| 2019 | Collaboration in global software development: an investigation on research trends and evolutionabstractGlobal software development (GSD) done by geographically distributed teams of developers is one of the most common ways of developing software nowadays. Though GSD has various benefits, it also introduces challenges that have led to a plethora of research. This paper analyzes research papers published in top software engineering venues in recent years (2009-2018) focusing on team collaboration in order to understand the trend in GSD research. Out of 4,292 papers published in these venues, we found 33 papers that focused on team collaboration in the context of GSD. We study the kinds of data used in these papers and classify them into primary data (i.e., interview and observation data) and secondary data (i.e., repository and communication data) and found that interview data is the dominant type of data in these papers. We also found that the strength of evidence presented in most papers tends to be moderate. Yang Yue 0003, Iftekhar Ahmed 0001, Yi Wang 0013, David F. Redmiles |
ICGSE | 2 |
| 2018 | What Makes a Good Developer? An Empirical Study of Developers' Technical and Social CompetenciesabstractTechnical and social competencies are highly desirable for a protean developer. Managers make hiring decisions based on developer's contributions to online peer production sites like GitHub and Stack Overflow. These sites provide ample history regarding developers' technical and social skills. Although these histories are utilized by hiring tools to help managers make their hiring decisions, little is known empirically how developers' social skills affect their technical skills and vice versa. Without such knowledge, tools, research, and training might be flawed. We present an in-depth empirical study investigating the correlation between the technical and social skills of developers. Our quantitative analysis of factors influencing the social skills of developers compared with factors affecting their technical skills indicates that better collaboration competency skills are associated with enhanced coding abilities as well as the quality of code. Sandeep Kaur Kuttal, Iftekhar Ahmed 0001 |
VL/HCC | 3 |
| 2018 | How verified (or tested) is my code? Falsification-driven verification and testing
Alex Groce, Iftekhar Ahmed 0001, Carlos Jensen, Paul E. McKenney, Josie Holmes |
Autom. Softw. Eng. | 2 |
| 2017 | An Empirical Examination of the Relationship between Code Smells and Merge ConflictsabstractBackground: Merge conflicts are a common occurrence in software development. Researchers have shown the negative impact of conflicts on the resulting code quality and the development workflow. Thus far, no one has investigated the effect of bad design (code smells) on merge conflicts. Aims: We posit that entities that exhibit certain types of code smells are more likely to be involved in a merge conflict. We also postulate that code elements that are both "smelly" and involved in a merge conflict are associated with other undesirable effects (more likely to be buggy). Method: We mined 143 repositories from GitHub and recreated 6,979 merge conflicts to obtain metrics about code changes and conflicts. We categorized conflicts into semantic or non-semantic, based on whether changes affected the Abstract Syntax Tree. For each conflicting change, we calculate the number of code smells and the number of future bug-fixes associated with the affected lines of code. Results: We found that entities that are smelly are three times more likely to be involved in merge conflicts. Method-level code smells (Blob Operation and Internal Duplication) are highly correlated with semantic conflicts. We also found that code that is smelly and experiences merge conflicts is more likely to be buggy. Conclusion: Bad code design not only impacts maintainability, it also impacts the day to day operations of a project, such as merging contributions, and negatively impacts the quality of the resulting code. Our findings indicate that research is needed to identify better ways to support merge conflict resolution to minimize its effect on code quality. Iftekhar Ahmed 0001, Caius Brindescu, Umme Ayda Mannan, Carlos Jensen, Anita Sarma |
ESEM | 1 |
| 2017 | A case study of motivations for corporate contribution to FOSSabstractFree/Open Source Software developers come from a myriad of different backgrounds, and are driven to contribute to projects for a variety of different reasons, including compensation from corporations or foundations. Motivation can have a dramatic impact on how and what contribution an individual makes, as well as how tenacious they are. These contributions may align with the needs of the developer, the community, the organization funding the developer, or all of the above. Understanding how corporate sponsorship affects the social dynamics and evolution of Free/Open Source code and community is critical to fostering healthy communities. We present a case study of corporations contributing to the Linux Kernel. We find that corporate contributors contribute more code, but are less likely to participate in non-coding activities. This knowledge will help project leaders to better understand the dynamics of sponsorship, and help to steer resources. Iftekhar Ahmed 0001, Darren Forrest, Carlos Jensen |
VL/HCC | 1 |
| 2017 | Does choice of mutation tool matter?
Rahul Gopinath, Iftekhar Ahmed 0001, Mohammad Amin Alipour, Carlos Jensen, Alex Groce |
Softw. Qual. J. | 2 |
| 2017 | Mutation Reduction Strategies Considered HarmfulabstractMutation analysis is a well known yet unfortunately costly method for measuring test suite quality. Researchers have proposed numerous mutation reduction strategies in order to reduce the high cost of mutation analysis, while preserving the representativeness of the original set of mutants. As mutation reduction is an area of active research, it is important to understand the limits of possible improvements. We theoretically and empirically investigate the limits of improvement in effectiveness from using mutation reduction strategies compared to random sampling. Using real-world open source programs as subjects, we find an absolute limit in improvement of effectiveness over random sampling- 13.078%. Given our findings with respect to absolute limits, one may ask: How effective are the extant mutation reduction strategies? We evaluate the effectiveness of multiple mutation reduction strategies in comparison to random sampling. We find that none of the mutation reduction strategies evaluated-many forms of operator selection, and stratified sampling (on operators or program elements)-produced an effectiveness advantage larger than 5% in comparison with random sampling. Given the poor performance of mutation selection strategies-they may have a negligible advantage at best, and often perform worse than random sampling- we caution practicing testers against applying mutation reduction strategies without adequate justification. Rahul Gopinath, Iftekhar Ahmed 0001, Mohammad Amin Alipour, Carlos Jensen, Alex Groce |
IEEE Trans. Reliab. | 2 |
| 2016 | On the limits of mutation reduction strategiesabstractAlthough mutation analysis is considered the best way to evaluate the effectiveness of a test suite, hefty computational cost often limits its use. To address this problem, various mutation reduction strategies have been proposed, all seeking to reduce the number of mutants while maintaining the representativeness of an exhaustive mutation analysis. While research has focused on the reduction achieved, the effectiveness of these strategies in selecting representative mutants, and the limits in doing so have not been investigated, either theoretically or empirically. Rahul Gopinath, Mohammad Amin Alipour, Iftekhar Ahmed 0001, Carlos Jensen, Alex Groce |
ICSE | 3 |
| 2016 | Can testedness be effectively measured?abstractAmong the major questions that a practicing tester faces are deciding where to focus additional testing effort, and deciding when to stop testing. Test the least-tested code, and stop when all code is well-tested, is a reasonable answer. Many measures of "testedness" have been proposed; unfortunately, we do not know whether these are truly effective. In this paper we propose a novel evaluation of two of the most important and widely-used measures of test suite quality. The first measure is statement coverage, the simplest and best-known code coverage measure. The second measure is mutation score, a supposedly more powerful, though expensive, measure. Iftekhar Ahmed 0001, Rahul Gopinath, Caius Brindescu, Alex Groce, Carlos Jensen |
SIGSOFT FSE | 1 |
| 2015 | An Empirical Study of Design Degradation: How Software Projects Get Worse over TimeabstractContext: Software decay is a key concern for large, long-lived software projects. Systems degrade over time as design and implementation compromises and exceptions pile up. Goal: Quantify design decay and understand how software projects deal with this issue. Method: We conducted an empirical study on the presence and evolution of code smells, used as an indicator of design degradation in 220 open source projects. Results: The best approach to maintain the quality of a project is to spend time reducing both software defects (bugs) and design issues (refactoring). We found that design issues are frequently ignored in favor of fixing defects. We also found that design issues have a higher chance of being fixed in the early stages of a project, and that efforts to correct these stall as projects mature and the code base grows, leading to a build-up of problems. Conclusions: From studying a large set of open source projects, our research suggests that while core contributors tend to fix design issues more often than non-core contributors, there is no difference once the relative quantity of commits is accounted for. We also show that design issues tend to build up over time. Iftekhar Ahmed 0001, Umme Ayda Mannan, Rahul Gopinath, Carlos Jensen |
ESEM | 1 |
| 2015 | How hard does mutation analysis have to be, anyway?abstractMutation analysis is considered the best method for measuring the adequacy of test suites. However, the number of test runs required for a full mutation analysis grows faster than project size, which is not feasible for real-world software projects, which often have more than a million lines of code. It is for projects of this size, however, that developers most need a method for evaluating the efficacy of a test suite. Various strategies have been proposed to deal with the explosion of mutants. However, these strategies at best reduce the number of mutants required to a fraction of overall mutants, which still grows with program size. Running, e.g., 5% of all mutants of a 2MLOC program usually requires analyzing over 100,000 mutants. Similarly, while various approaches have been proposed to tackle equivalent mutants, none completely eliminate the problem, and the fraction of equivalent mutants remaining is hard to estimate, often requiring manual analysis of equivalence. In this paper, we provide both theoretical analysis and empirical evidence that a small constant sample of mutants yields statistically similar results to running a full mutation analysis, regardless of the size of the program or similarity between mutants. We show that a similar approach, using a constant sample of inputs can estimate the degree of stubbornness in mutants remaining to a high degree of statistical confidence, and provide a mutation analysis framework for Python that incorporates the analysis of stubbornness of mutants. Rahul Gopinath, Mohammad Amin Alipour, Iftekhar Ahmed 0001, Carlos Jensen, Alex Groce |
ISSRE | 3 |
| 2015 | How Verified is My Code? Falsification-Driven Verification (T)abstractFormal verification has advanced to the point that developers can verify the correctness of small, critical modules. Unfortunately, despite considerable efforts, determining if a "verification" verifies what the author intends is still difficult. Previous approaches are difficult to understand and often limited in applicability. Developers need verification coverage in terms of the software they are verifying, not model checking diagnostics. We propose a methodology to allow developers to determine (and correct) what it is that they have verified, and tools to support that methodology. Our basic approach is based on a novel variation of mutation analysis and the idea of verification driven by falsification. We use the CBMC model checker to show that this approach is applicable not only to simple data structures and sorting routines, and verification of a routine in Mozilla's JavaScript engine, but to understanding an ongoing effort to verify the Linux kernel Read-Copy-Update (RCU) mechanism. Alex Groce, Iftekhar Ahmed 0001, Carlos Jensen, Paul E. McKenney |
ASE | 2 |
| 2014 | The Impact of Automatic Crash Reports on Bug Triaging and Development in MozillaabstractFree/Open Source Software projects often rely on users submitting bug reports. However, reports submitted by novice users may lack information critical to developers, and the process may be intimidating and difficult. To gather more and better data, projects deploy automatic crash reporting tools, which capture stack traces and memory dumps when a crash occurs. These systems potentially generate large volumes of data, which may overwhelm developers, and their presence may discourage users from submitting traditional bug reports. In this paper, we examine Mozilla's automatic crash reporting system and how it affects their bug triaging process. We find that fewer than 0.00009% of crash reports end up in a bug report, but as many as 2.33% of bug reports have data from crash reports added. Feedback from developers shows that despite some problems, these systems are valuable. We conclude with a discussion of the pros and cons of automatic crash reporting systems. Iftekhar Ahmed 0001, Nitin Mohan, Carlos Jensen |
OpenSym | 1 |