VLDB 2026 Research / reviewers in the wild / expert
Esteban Parra
dblp:137/0576
· DBLP profile ↗
22ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0001-9813-9518ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 20 · 5 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What characteristics make ChatGPT effective for software issue resolution? An empirical study of task, project, and conversational signals in GitHub issuesabstractAbstract Conversational large-language models (LLMs), such as ChatGPT, are extensively used for issue resolution tasks, particularly for generating ideas to implement new features or resolve bugs. However, not all developer-LLM conversations are useful for effective issue resolution and it is still unknown what makes some of these conversations not helpful. In this paper, we analyze 686 developer-ChatGPT conversations shared within GitHub issue threads to identify characteristics that make these conversations effective for issue resolution. First, we empirically analyze the conversations and their corresponding issue threads to distinguish helpful from unhelpful conversations. We begin by categorizing the types of tasks developers seek help with (e.g., code generation , bug identification and fixing , test generation ), to better understand the scenarios in which ChatGPT is most effective. Next, we examine a wide range of conversational, project, and issue-related metrics to uncover statistically significant factors associated with helpful conversations. Finally, we identify common deficiencies in unhelpful ChatGPT responses to highlight areas that could inform the design of more effective developer-facing tools. We found that only 62% of the ChatGPT conversations were helpful for successful issue resolution. Among different tasks related to issue resolution, ChatGPT was most helpful in assisting with code generation, and tool/library/API recommendations, but struggled with generating code explanations. Our conversational metrics reveal that helpful conversations are shorter, more readable, and exhibit higher semantic and linguistic alignment. Our project metrics reveal that larger, more popular projects and experienced developers benefit more from ChatGPT’s assistance. Our issue metrics indicate that ChatGPT is more effective on simpler issues characterized by limited developer activity and faster resolution times. These typically involve well-scoped technical problems such as compilation errors and tool feature requests. In contrast, it performs less effectively on complex issues that demand deep project-specific understanding, such as system-level code debugging and refactoring. The most common deficiencies in unhelpful ChatGPT responses include incorrect information and lack of comprehensiveness. Our findings have wide implications including guiding developers on effective interaction strategies for issue resolution, informing the development of tools or frameworks to support optimal prompt design, and providing insights on fine-tuning LLMs for issue resolution tasks. Ramtin Ehsani, Sakshi Pathak, Esteban Parra, Sonia Haiduc, Preetha Chatterjee |
Empir. Softw. Eng. | 3 |
| 2025 | Towards Implementing and Evaluating AI-Assisted Pull Requests in Software Engineering EducationabstractPull requests allow developers to suggest and review codebase changes collaboratively. This process is standard for maintaining code quality and following best practices. The recent emergence of Large Language Models like ChatGPT and GitHub Copilot has shown great potential in improving coding efficiency and accuracy in software engineering. This paper outlines a study design to explore the integration of an AI tool to streamline PR reviews in a software engineering course. By incorporating the pr-agent into the curriculum, the study aims to evaluate its impact on students' coding skills, understanding of PR processes, and overall learning experience. The evaluation strategy includes collecting quantitative and qualitative feedback from students to assess the effectiveness of the tool. The results will offer insight into the feasibility and benefits of integrating AI tools in software engineering education. Esteban Parra, Sophia Willingham |
CSEE&T | 1 |
| 2025 | Design of An Eye-Tracking Study Towards Assessing the Impact of Generative AI Use on Code Summarization
Suad Mohamed, Najma Ismail, Kimberly Amaya Hernandez, Abdullah Parvin, Michael Oliver, Esteban Parra |
ETRA | 6 |
| 2025 | AI-Powered Commit Explorer (APCE)abstractCommit messages in a version control system provide valuable information for developers regarding code changes in software systems. Commit messages can be the only source of information left for future developers describing what was changed and why. However, writing high-quality commit messages is often neglected in practice. Large Language Model (LLM) generated commit messages have emerged as a way to mitigate this issue. We introduce the AI-Powered Commit Explorer (APCE), a tool to support developers and researchers in the use and study of LLM-generated commit messages. APCE gives researchers the option to store different prompts for LLMs and provides an additional evaluation prompt that can further enhance the commit message provided by LLMs. APCE also provides researchers with a straightforward mechanism for automated and human evaluation of LLM-generated messages. Demo link https://youtu.be/zYrJ9s6sZvo Yousab Grees, Polina Iaremchuk, Ramtin Ehsani, Esteban Parra, Preetha Chatterjee, Sonia Haiduc |
ICSME | 4 |
| 2025 | Hierarchical Knowledge Injection for Improving LLM-based Program Repair
Ramtin Ehsani, Esteban Parra, Sonia Haiduc, Preetha Chatterjee |
ASE | 2 |
| 2025 | Cascading Effects: Analyzing Project Failure Impact in the Maven Central EcosystemabstractThis study examines failure propagation within the Maven Central ecosystem, a critical software dependency repository, through a comprehensive analysis of dependency networks using the Goblin framework. Our dual-sampling methodology, investigating both top dependencies and random libraries, revealed two distinct failure propagation patterns that pose significant risks to ecosystem stability. Core infrastructure failures, particularly evident in cases like the AWS SDK family, can create immediate and widespread disruption. These libraries have many direct dependents - averaging 20,402 projects that would break immediately if the library fails. Furthermore, these core libraries themselves have extensive dependencies, with the AWS SDK family depending on 377 other libraries itself. The impact of failures spreads deeply through the ecosystem, with dependency chains reaching an average of 90.80 levels. Our analysis of peripheral projects reveals their significant cascading effects, with higher average dependency depths of 54.25 levels and chain lengths extending to 116.74 levels, as exemplified by cases like org.apache.camel:camel-swagger-java, which demonstrated a maximum chain length of 647 levels. Our findings highlight specific vulnerabilities in dependency network structures, showing that ecosystem resilience requires both protecting core infrastructure and managing dependency complexity. Mina Shehata, Saidmakhmud Makhkamjonoov, Mahad Syed, Esteban Parra |
MSR | 4 |
| 2025 | Visualizing Bug Lifecycles Within DebianabstractAs software systems evolve, users and developers are responsible for identifying and reporting bug fixes. Addressing these issues is vital to ensuring the health, security, and performance of the system, as well as providing a positive user experience. Analyzing bug resolution times within large-scale open source systems such as Linux is important to understand the overall health of the project, long-term maintenance, and different factors affecting bug resolution. In this paper, we perform an analysis of resolved bugs within the Debian distribution of Linux. We collected a data set of 466 bugs from the Ultimate Debian Database (UDD) for this project.In this paper, we present a collection of interactive visualizations of bug fix data in the Debian distribution of Linux. Namely, (1) bar charts showing distributions of bug fix durations, (2) a line graph of reported and resolved bugs per year, (3) a box plot relating bug severity with resolution time, (4) a bar chart identifying the most frequently affected packages, and (5) a scatter plot tracking how fix durations have changed over time. These visualizations reveal patterns in bug management and resolution, providing insight into long-term maintenance strategies and bug triage response times. These tools can be used to inform decisions related to real-world bug processing behavior in the Linux kernel environment and other large open-source projects. Challenge Demo Video: https://youtu.be/3rL0YWmiRl0 Christina Clements, Esteban Parra |
VISSOFT | 2 |
| 2025 | Interactive 3D Graph Visualization of Linux Kernel SubsystemsabstractThe Linux Kernel consists of thousands of interrelated subsystems that can be challenging to grasp for developers or curious users. We are presenting an interactive 3D graph visualization tool built in Unity that facilitates the comprehension of the cloned Linux kernel Git repository. The visualization creates an undirected graph representing files or folders of desired subsystems consisting of nodes and edges as connections between them. It includes a canvas with slider filters, where the minimum and maximum years allow for more time-specific visualization, while the intensity represents the total lines of code changed in a diff for the given time frame. The user can also navigate in different directions, move closer to the main file or any other area, as well as move and drag the visualization of the subsystem nodes for a better overview of the subsystem network. The current implementation uses a file-writing script which transfers cleaned data from the cloned respiratory, generating a file in the Unity project folder, later used for the file-reading script for Unity graph rendering. The 3D graph representation can be helpful for newcomers, developers, and researchers seeking to explore and learn the complex structure of Linux Kernel subsystems and draw insights about the commit activity on a chosen time frame. Demo link - https://youtu.be/pL97TxFhsq8. Polina Iaremchuk, Esteban Parra |
VISSOFT | 2 |
| 2025 | Visualizing Linux Kernel Contributor Networks in JavaabstractModern software systems have grown increasingly complex, necessitating the collaborative efforts of diverse teams comprised of multiple developers. The combined efforts and collaboration of distributed teams create underlying networks within which these teams operate. This project presents a Java-based visualization tool that analyzes and depicts collaborative patterns within the Linux kernel development community. We extract contributor information from the Linux kernel repository to construct a network graph where nodes represent individual contributors and edges indicate shared contributions on the same commits. By showing which contributors tend to work together, the tool helps identify likely points of contact for bug triage, code maintenance, and understanding subsystem ownership. It also lays the groundwork for a deeper analysis of team structure and collaboration patterns. Ella G. McDevitt, Esteban Parra |
VISSOFT | 2 |
| 2025 | Visualizing Software Evolution in Linux: A Hierarchical Graph-Based Heat MapabstractLinux is a large open-source operating system. Its size makes it difficult for developers to fully grasp the system as a whole. Visualizations of the Linux kernel can provide developers with better program comprehension and understanding of evolution processes. This paper presents a visualization of Linux commit activity using a hierarchical, heat-map based graph. Using a large data set of commit data from Zenodo, we model the Linux structure as a directed hierarchy. Each node in the graph represents a subsystem, and edges show a parent-child relationship within the Linux architecture. The graph is rendered radially in Unity with a color gradient applied to nodes to indicate the volume of commits made to each subsystem. The result is an interactive, intuitive view of development hot-spots in the Linux kernel, aimed at supporting further software evolution analysis. Challenge video link: https://youtu.be/5rNFFZl16v4 Adrian Volpe, Esteban Parra |
VISSOFT | 2 |
| 2024 | Chatting with AI: Deciphering Developer Conversations with ChatGPTabstractLarge Language Models (LLMs) have been widely adopted and are becoming ubiquitous and integral to software development. However, we have little knowledge as to how these tools are being used by software developers beyond anecdotal evidence and word-of-mouth reports. In this work, we present a study toward understanding how developers engage with and utilize LLMs by reporting the results of an empirical study identifying patterns in the conversation that developers have with LLMs. We identified a total of 19 topics describing the purpose of the developers in their conversations with LLMs. Our findings reveal that developers use LLMs to facilitate various aspects of their software development processes (e.g., information-seeking about programming languages and frameworks and soliciting high-level design recommendations) to a similar extent to which they use them for non-development purposes such as writing assistance, general purpose queries, and conducting Turing tests to assess the intrinsic capabilities of the models. This work not only sheds light on the diverse applications of LLMs in software development but also underscores their emerging role as critical tools in enhancing developer productivity and creativity as we move closer to widespread AI-assisted software development. Suad Mohamed, Abdullah Parvin, Esteban Parra |
MSR | 3 |
| 2024 | Creating UML Class Diagrams with General-Purpose LLMsabstractGeneral-purpose large language models (LLMs) have become a versatile tool in software development and maintenance. These models offer support in tasks such as under-standing, writing, and summarizing code. While these LLMs can generate code quickly their use in software modeling is relatively under-explored. This research aims to study ChatGPT's ability to generate class diagrams from source code. To do this, we engineered a prompt that takes in source code to create a UML class diagram for that system. We used this prompt to create class diagrams for 40 systems and assessed the diagrams on their correctness and structure. Our results show that ChatGPT creates class diagrams that accurately capture 90% of the classes and their attributes and 66 % of the associations. While ChatGPT performed nearly flawlessly for smaller projects, the diagrams for larger projects had more issues. We conclude that ChatGPT is best utilized as a complementary tool rather than the sole resource for software modeling and maintenance. Mina Shehata, Blaire Lepore, Hailey Cummings, Esteban Parra |
VISSOFT | 4 |
| 2023 | A Machine Learning Approach to Convert Pseudo-Code to Domain-Specific Programming LanguageabstractThe California Department of Energy oversees a set of standards and rules that new constructions (i.e., buildings) must adhere to within the State of California (the California Energy Code). Energy Code Ace (ECA) is a website that helps users plan their future construction in California and see if the plans for their structure comply with the California Energy Code. However, there are thousands of details and sections in the California Energy Code that need to be translated into questions and fields on the website for the user to fill out. To do this, the Department of Energy provides a web development team at Binary Evolution with XML schematics that contain details of each field and pseudo-code of the logic that fields should follow compared to the other fields on the ECA website. A very large portion of the developers’ job involves reading the pseudo-code and writing real code in a language developed by Binary Evolution. Reading and handwriting the code as described in the schematics can take up to hours of a developer’s time.This paper presents ECAi, a new approach for automatically generating executable code from pseudo-code schematics. ECAi uses a Decision Tree Classifier trained on 23,864 lines of pseudo-code and had to classify a line of pseudo-code into one of 13 categories achieving an 89.87% F1-score. Leveraging the classifier category, ECAi then calls a code-writing algorithm to parse the line of pseudo-code and writes the executable code that corresponds with the given pseudo-code line. The implementation of a prototype of the system shows a statistically significant (p=0.0116) decrease in the time it takes a developer to obtain an executable piece of code. ECAi has saved countless hours of development work during the maintenance workflow of ECA. Jacob Neal, Shane Rogers, Esteban Parra |
ICSME | 3 |
| 2023 | Developers and Modern Communication MediumsabstractAs part of their daily work, software developers use information that has been traditionally sourced from code, version control repositories, colleagues, emails, and online documentation. However, recently developers have started using information from modern communication mediums like online forums, videos, and chat platforms. The rising usage of these platforms, fueled by online learning and remote work, offers vast benefits but also necessitates significant effort to locate task-specific knowledge.Modern communication mediums such as online videos and instant messaging platforms have become more prevalent among developers than ever before. Although these mediums provide a large array of benefits to developers, they can also require a significant amount of effort spent searching and browsing through the information they contain to find the knowledge that is relevant to the developers’ current tasks and needs.My work aims to design mechanisms and systems that assist developers in extracting information from videos and instant messaging. In particular, i) use information retrieval algorithms to produce meaningful tags for software development videos and ii) present evidence on how developers interact with instant messaging platforms. Esteban Parra |
ICSME | 1 |
| 2023 | Extended Abstract of A Comparative Study and Analysis of Developer Communications on Slack and GitterabstractSoftware developers are often using instant messaging platforms to communicate with each other and other stakeholders. Among these platforms, Gitter has emerged as a popular choice and the messages it contains can reveal important information to researchers studying open-source software systems. Uncovering what developers are communicating about through Gitter is an essential first step towards successfully understanding and leveraging this information. This paper builds upon our previously published paper (Parra et al. 2020), which introduced GitterCom for the first time and presented a study of the messages it contains with the goal of observing how developers and other stakeholders communicate about software using Gitter in the context of Gitter communities dedicated to the active development of open source software systems on GitHub. Esteban Parra, Mohammad Alahmadi 0001, Ashley Ellis, Sonia Haiduc |
SANER | 1 |
| 2022 | A comparative study and analysis of developer communications on Slack and GitterabstractSoftware developers are often using instant messaging platforms to communicate with each other and other stakeholders. Among these platforms, Gitter has emerged as a popular choice and the messages it contains can reveal important information to researchers studying open source software systems. Uncovering what developers are communicating about through Gitter is an essential first step towards successfully understanding and leveraging this information. In this paper, we first describe the largest manually labeled and curated dataset of Gitter developer messages, named GitterCom, obtained by manually analyzing and labeling 10,000 Gitter messages in 10 software projects. We then present a qualitative study to understand the extent to which the categories identified in previous work by Lin et al. (2016) found on Slack through surveys are applicable to developer messages exchanged on Gitter. Further, in an effort to automate the labeling process, we investigate the accuracy of 9 traditional machine learning and deep learning algorithms in predicting the intent of Gitter messages. We found that Decision Trees and Random Forest performed the best, achieving an accuracy of 88%, which is very promising for this multi-class classification task. Finally, we discuss the potential directions for future research enabled by labeled Gitter datasets such as GitterCom. Esteban Parra, Mohammad Alahmadi 0001, Ashley Ellis, Sonia Haiduc |
Empir. Softw. Eng. | 1 |
| 2020 | GitterCom: A Dataset of Open Source Developer Communications in GitterabstractTeam communication is essential for the development of modern software systems. For distributed software development teams, such as those found in many open source projects, this communication usually takes place using electronic tools. Among these, modern chat platforms such as Gitter are becoming the de facto choice for many software projects due to their advanced features geared towards software development and effective team communication. Gitter channels contain numerous messages exchanged by developers regarding the state of the project, issues and features of the system, team logistics, etc. These messages can contain important information to researchers studying open source software systems, developers new to a particular project and trying to get familiar with the software, etc. Therefore, uncovering what developers are communicating about through Gitter is an essential first step towards successfully understanding and leveraging this information. Esteban Parra, Ashley Ellis, Sonia Haiduc |
MSR | 1 |
| 2020 | UIScreens: extracting user interface screens from mobile programming video tutorialsabstractMobile apps are one of the most widely used types of software systems in existence today and more programmers and students learn how to develop them everyday. One of the most popular resources for learning mobile programming are videos hosted on social platforms such as YouTube. While useful, this type of resource has also its limitations, especially when developers are looking for user interface (UI) designs for mobile applications, since these are hard to search for and locate in videos. We propose UIScreens, a web-based analysis and search engine that analyzes the visual contents of mobile programming video tutorials, then identifies and extracts the UI screens displayed in the videos. Our tool offers features such as searching for UI screens in videos, displaying an overview of the UI screens identified in a video under each search result, and navigating to the part of a video where a particular UI screen is being displayed and discussed. In a user study, participants agreed that UIScreens is usable and useful to quickly skim through videos, while the UI screens it extracts can help developers further determine the relevance of videos to a search topic. Mohammad Alahmadi 0001, Ahmad Tayeb, Abdulkarim Khormi, Esteban Parra, Sonia Haiduc |
ESEC/SIGSOFT FSE | 4 |
| 2020 | On the relationship between bug reports and queries for text retrieval-based bug localization
Chris Mills, Esteban Parra, Jevgenija Pantiuchina, Gabriele Bavota, Sonia Haiduc |
Empir. Softw. Eng. | 2 |
| 2018 | Are Bug Reports Enough for Text Retrieval-Based Bug Localization?abstractText Retrieval (TR) has been widely used to support many software engineering tasks, including bug localization (i.e., the activity of localizing buggy code starting from a bug report). Many studies show TR's effectiveness in lowering the manual effort required to perform this maintenance task; however, the actual usefulness of TR-based bug localization has been questioned in recent studies. These studies discuss (i) potential biases in the experimental design usually adopted to evaluate TRbased bug localization techniques and (ii) their poor performance in the scenario when they are needed most: when the bug report, which serves as the de facto query in most studies, does not contain localization hints (e.g., code snippets, method names, etc.) Fundamentally, these studies raise the question: do bug reports provide sufficient information to perform TR-based localization? In this work, we approach that question from two perspectives. First, we investigate potential biases in the evaluation of TR-based approaches which artificially boost the performance of these techniques, making them appear more successful than they are. Second, we analyze bug report text with and without localization hints using a genetic algorithm to derive a near-optimal query that provides insight into the potential of that bug report for use in TR-based localization. Through this analysis we show that in most cases the bug report vocabulary (i.e., the terms contained in the bug title and description) is all we need to formulate effective queries, making TR-based bug localization successful without supplementary query expansion. Most notably, this also holds when localization hints are completely absent from the bug report. In fact, our results suggest that the next major step in improving TR-based bug localization is the ability to formulate these near-optimal queries. Chris Mills, Jevgenija Pantiuchina, Esteban Parra, Gabriele Bavota, Sonia Haiduc |
ICSME | 3 |
| 2018 | Automatic tag recommendation for software development video tutorialsabstractSoftware development video tutorials are emerging as a new resource for developers to support their information needs. However, when trying to find the right video to watch for a task at hand, developers have little information at their disposal to quickly decide if they found the right video or not. This can lead to missing the best tutorials or wasting time watching irrelevant ones. Esteban Parra, Javier Escobar-Avila, Sonia Haiduc |
ICPC | 1 |
| 2013 | An empirical study assessing the effect of seeit 3D on comprehensionabstractA study to assess the effect of SeeIT 3D, a software visualization tool is presented. Six different tasks in three different task categories are assessed in the context of a large open-source system. Ninety-seven subjects were recruited from three different universities to participate in the study. Two methods of data collection: traditional questionnaires and an eye-tracker were used. The main goal was to determine the impact and added benefit of SeeIT 3D while performing typical software tasks within the Eclipse IDE. Results indicate that SeeIT 3D performs significantly better in one task category namely overview tasks but takes significantly longer when completing bug fixing tasks. Scores obtained by the subjects in the SeeIT 3D group are 13% better and 45% faster for overview tasks. Bonita Sharif, Grace Jetty, Jairo Aponte, Esteban Parra |
VISSOFT | 4 |