Souti Chattopadhyay

dblp:208/7391 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0003-1644-7344ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Trust Dynamics in AI-Assisted Development: Definitions, Factors, and Implications
abstract
Software developers increasingly rely on AI code generation utilities. To ensure that “good” code is accepted into the code base and “bad” code is rejected, developers must know when to trust an AI suggestion. Understanding how developers build this intuition is crucial to enhancing developer-AI collaborative programming. In this paper, we seek to understand how developers (1) define and (2) evaluate the trustworthiness of a code suggestion and (3) how trust evolves when using AI code assistants. To answer these questions, we conducted a mixed method study consisting of an in-depth exploratory survey with (n=29) developers followed by an observation study (n=10). We found that comprehensibility and perceived correctness were the most frequently used factors to evaluate code suggestion trustworthiness. However, the gap in developers' definition and evaluation of trust points to a lack of support for evaluating trustworthy code in real-time. We also found that developers often alter their trust decisions, keeping only 52% of original suggestions. Based on these findings, we extracted four guidelines to enhance developer-AI interactions. We validated the guidelines through a survey with (n=7) domain experts and survey members (n=8). We discuss the validated guidelines, how to apply them, and tools to help adopt them.
Sadra Sabouri, Philipp Eibl, Morteza Ziyadi, Nenad Medvidovic, Lars Lindemann, Souti Chattopadhyay
ICSE7
2025 Code Today, Deadline Tomorrow: Procrastination Among Software Developers
abstract
Procrastination, the action of delaying or postponing something, is a well-known phenomenon that is relatable to all. While it has been studied in academic settings, little is known about why software developers procrastinate. How does it affect their work? How can developers manage procrastination? This paper presents the first investigation of procrastination among developers. We conduct an interview study with (n=15) developers across different industries to understand the process of procrastination. Using qualitative coding, we report the positive and negative effects of procrastination and factors that triggered procrastination, as perceived by participants. We validate our findings using member checking. Our results reveal 14 negative effects of procrastination on developer productivity. However, participants also reported eight positive effects, four impacting their satisfaction. We also found that participants reported three categories of factors that trigger procrastination: task-related, personal, and external. Finally, we present 19 techniques reported by our participants and studies in other domains that can help developers mitigate the impacts of procrastination. These techniques focus on raising awareness and task focus, help with task planning, and provide pathways to generate team support as a mitigation means. Based on these findings, we discuss interventions for developers and recommendations for tool building to reduce procrastination. Our paper shows that procrastination has unique effects and factors among developers compared to other populations.
Zeinabsadat Saghi, Thomas Zimmermann 0001, Souti Chattopadhyay
ICSE3
2025 Beyond the Page: Enriching Academic Paper Reading with Social Media Discussions
abstract
Figure 1: Surf enriches the paper reading experience by connecting research papers with related social media discussions.The interface displays the paper on the left side, with the right panels presenting organized threads of peer discussions around the paper on social media.Surf enables readers to fluidly navigate between paper content and social discourse, allowing them to develop deeper and more critical understanding without increasing cognitive overhead.
Run Huang, Anna Katherine Zhao, Zeinabsadat Saghi, Sadra Sabouri, Souti Chattopadhyay
UIST5
2025 Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
abstract
Recent AI code assistants have significantly improved their ability to process more complex contexts and generate entire codebases based on a textual description, compared to the popular snippet-level generation. These codebase AI assistants (CBAs) can also extend or adapt codebases, allowing users to focus on higher-level design and deployment decisions. While prior work has extensively studied the impact of snippetlevel code generation, this new class of codebase generation models is relatively unexplored. Despite initial anecdotal reports of excitement about these agents, they remain less frequently adopted compared to snippet-level code assistants. To utilize CBAs better, we need to understand how developers interact with CBAs, and how and why CBAs falls short of developers’ needs. In this paper, we explored these gaps through a counterbalanced user study and interview with ($\mathbf{n}=16$) students and developers working on coding tasks with CBAs. We found that participants varied the information in their prompts, like problem description (48% prompts), required functionality (98% prompts), code structure (48% prompts), and their prompt writing process. Despite various strategies, the overall satisfaction score with generated codebases remained low (mean $=2.8$, median $=3$, on a scale of one to five). Participants mentioned functionality as the most common factor for dissatisfaction (77% instances), alongside poor code quality (42% instances) and communication issues (25% instances). We delve deeper into participants’ dissatisfaction to identify six underlying challenges that participants faced when using CBAs, and extracted five barriers to incorporating CBAs into their workflows. Finally, we surveyed 21 commercial CBAs to compare their capabilities with participant challenges, and present design opportunities for more efficient and useful CBAs.
Philipp Eibl, Sadra Sabouri, Souti Chattopadhyay
VL/HCC3
2024 Generating Function Names to Improve Comprehension of Synthesized Programs
abstract
The hope of allowing programmers to more freely express themselves has led to a proliferation of program synthesis techniques. These tools automatically derive implementations from high-level specifications of user intent. These specifications may take the form of logical formulas, demonstrations, or input-output examples. Synthesizers guarantee that when synthesis is successful, the implementation satisfies the specification. However, they provide no additional information regarding how the implementation works or the manner in which the specification is realized. As a result, they remain algorithmic black boxes which are prone to producing unidiomatic code with procedurally generated identifier names, like $x 1, x 2$, etc. As a result, complicated implementations produced by modern program synthesizers are becoming increasingly hard to understand. One solution to this comprehensibility problem is to produce meaningful identifier names for its variables, functions, etc. While large language models (LLMs) suggest a simple way to obtain human-readable names, our experiments reveal that LLMs frequently produce nonsensical or misleading names when applied to code emitted by program synthesizers. In this paper, we develop an approach to reliably augment the implementation with explanatory names: We recover finegrained input-output data from the synthesis algorithm to enhance the prompt supplied to the LLM and use a combination of a program verifier and a second language model to validate the proposed names before presenting them to the user. Together, these techniques improve the accuracy of the proposed names from $\mathbf{2 4 \%}$ to $\mathbf{7 9 \%}$. A two-phase user study indicates that users significantly prefer the names produced by our technique, and that the proposed names greatly help users in understanding synthesized implementations.
Amirmohammad Nazari, Swabha Swayamdipta, Souti Chattopadhyay, Mukund Raghothaman
VL/HCC3
2021 Reel life vs. real life: how software developers share their daily life through vlogs
abstract
Software developers are turning to vlogs (video blogs) to share what a day is like to walk in their shoes. Through these vlogs developers share a rich perspective of their technical work as well their personal lives. However, does the type of activities portrayed in vlogs differ from activities developers in the industry perform? Would developers at a software company prefer to show activities to different extents if they were asked to share about their day through vlogs? To answer these questions, we analyzed 130 vlogs by software developers on YouTube and conducted a survey with 335 software developers at a large software company. We found that although vlogs present traditional development activities such as coding and code peripheral activities (11%), they also prominently feature wellness and lifestyle related activities (47.3%) that have not been reflected in previous software engineering literature. We also found that developers at the software company were inclined to share more non-coding tasks (e.g., personal projects, time spent with family and friends, and health) when asked to create a mock-up vlog to promote diversity. These findings demonstrate a shift in our understanding of how software developers are spending their time and find valuable to share publicly. We discuss how vlogs provide a more complete perspective of software development work and serve as a valuable source of data for empirical research.
Souti Chattopadhyay, Thomas Zimmermann 0001, Denae Ford
ESEC/SIGSOFT FSE1
2021 Developers Who Vlog: Dismantling Stereotypes through Community and Identity
abstract
Developers are more than "nerds behind computers all day", they lead a normal life, and not all take the traditional path to learn programming. However, the public still sees software development as a profession for "math wizards". To learn more about this special type of knowledge worker from their first-person perspective, we conducted three studies to learn how developers describe a day in their life through vlogs on YouTube and how these vlogs were received by the broader community. We first interviewed 16 developers who vlogged to identify their motivations for creating this content and their intention behind what they chose to portray. Second, we analyzed 130 vlogs (video blogs) to understand the range of the content conveyed through videos. Third, we analyzed 1176 comments from the 130 vlogs to understand the impact the vlogs have on the audience. We found that developers were motivated to promote and build a diverse community, by sharing different aspects of life that define their identity, and by creating awareness about learning and career opportunities in computing. They used vlogs to share a variety of how software developers work and live---showcasing often unseen experiences, including intimate moments from their personal life. From our comment analysis, we found that the vlogs were valuable to the audience to find information and seek advice. Commenters sought opportunities to connect with others over shared triumphs and trials they faced that were also shown in the vlogs. As a central theme, we found that developers use vlogs to challenge the misconceptions and stereotypes around their identity, work-life, and well-being. These social stigmas are obstacles to an inclusive and accepting community and can deter people from choosing software development as a career. We also discuss the implications of using vlogs to support developers, researchers, and beyond.
Souti Chattopadhyay, Denae Ford, Thomas Zimmermann 0001
Proc. ACM Hum. Comput. Interact.1
2020 What's Wrong with Computational Notebooks? Pain Points, Needs, and Design Opportunities
abstract
Computational notebooks - such as Azure, Databricks, and Jupyter - are a popular, interactive paradigm for data scientists to author code, analyze data, and interleave visualizations, all within a single document. Nevertheless, as data scientists incorporate more of their activities into notebooks, they encounter unexpected difficulties, or pain points, that impact their productivity and disrupt their workflow. Through a systematic, mixed-methods study using semi-structured interviews (n=20) and survey (n=156) with data scientists, we catalog nine pain points when working with notebooks. Our findings suggest that data scientists face numerous pain points throughout the entire workflow - from setting up notebooks to deploying to production - across many notebook environments. Our data scientists report essential notebook requirements, such as supporting data exploration and visualization. The results of our study inform and inspire the design of computational notebooks.
Souti Chattopadhyay, Ishita Prasad, Austin Z. Henley, Anita Sarma, Titus Barik
CHI1
2020 A tale from the trenches: cognitive biases and software development
abstract
Cognitive biases are hard-wired behaviors that influence developer actions and can set them on an incorrect course of action, necessitating backtracking. While researchers have found that cognitive biases occur in development tasks in controlled lab studies, we still don't know how these biases affect developers' everyday behavior. Without such an understanding, development tools and practices remain inadequate. To close this gap, we conducted a 2-part field study to examine the extent to which cognitive biases occur, the consequences of these biases on developer behavior, and the practices and tools that developers use to deal with these biases. About 70% of observed actions that were reversed were associated with at least one cognitive bias. Further, even though developers recognized that biases frequently occur, they routinely are forced to deal with such issues with ad hoc processes and sub-optimal tool support. As one participant (IP12) lamented: There is no salvation!
Souti Chattopadhyay, Nicholas Nelson 0002, Audrey Au, Natalia Morales, Christopher A. Sanchez, Rahul Pandita, Anita Sarma
ICSE1
2020 Supporting Code Comprehension via Annotations: Right Information at the Right Time and Place
abstract
Code comprehension, especially understanding relationships across project elements (code, documentation, etc.), is non-trivial when information is spread across different interfaces and tools. Bringing the right amount of information, to the place where it is relevant and when it is needed can help reduce the costs of seeking information and creating mental models of the code relationships. While non-traditional IDEs have tried to mitigate these costs by allowing users to spatially place relevant information together, thus far, no study has examined the effects of these non-traditional interactions on code comprehension. Here, we present an empirical study to investigate how the right information at the right time and right place allows users-especially newcomers-to reduce the costs of code comprehension. We use a non-traditional IDE, called Synectic, and implement link-able annotations which provide affordances for the accuracy, time, and space dimensions. We conducted a between-subjects user study of 22 newcomers performing code comprehension tasks using either Synectic or a traditional IDE, Eclipse. We found that having the right information at the right time and place leads to increased accuracy and reduced cognitive load during code comprehension tasks, without sacrificing the usability of developer tools.
Marjan Adeli, Nicholas Nelson 0002, Souti Chattopadhyay, Hayden Coffey, Austin Z. Henley, Anita Sarma
VL/HCC3
2020 Mental Models of Mere Mortals with Explanations of Reinforcement Learning
abstract
How should reinforcement learning (RL) agents explain themselves to humans not trained in AI? To gain insights into this question, we conducted a 124-participant, four-treatment experiment to compare participants’ mental models of an RL agent in the context of a simple Real-Time Strategy (RTS) game. The four treatments isolated two types of explanations vs. neither vs. both together. The two types of explanations were as follows: (1) saliency maps (an “Input Intelligibility Type” that explains the AI’s focus of attention) and (2) reward-decomposition bars (an “Output Intelligibility Type” that explains the AI’s predictions of future types of rewards). Our results show that a combined explanation that included saliency and reward bars was needed to achieve a statistically significant difference in participants’ mental model scores over the no-explanation treatment. However, this combined explanation was far from a panacea: It exacted disproportionately high cognitive loads from the participants who received the combined explanation. Further, in some situations, participants who saw both explanations predicted the agent’s next action worse than all other treatments’ participants.
Andrew Anderson 0002, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Matthew L. Olson, Alan Fern, Margaret M. Burnett
ACM Trans. Interact. Intell. Syst.7
2019 Latent patterns in activities: a field study of how developers manage context
abstract
In order to build efficient tools that support complex programming tasks, it is imperative that we understand how developers program. We know that developers create a context around their programming task by gathering relevant information. We also know that developers decompose their tasks recursively into smaller units. However, important gaps exist in our knowledge about: (1) the role that context plays in supporting smaller units of tasks, (2) the relationship that exists among these smaller units, and (3) how context flows across them. The goal of this research is to gain a better understanding of how developers structure their tasks and manage context through a field study of ten professional developers in an industrial setting. Our analysis reveals that developers decompose their tasks into smaller units with distinct goals, that specific patterns exist in how they sequence these smaller units, and that developers may maintain context between those smaller units with related goals.
Souti Chattopadhyay, Nicholas Nelson 0002, Yenifer Ramirez Gonzalez, Annel Amelia Leon, Rahul Pandita, Anita Sarma
ICSE1
2019 Explaining Reinforcement Learning to Mere Mortals: An Empirical Study
abstract
We present a user study to investigate the impact of explanations on non-experts? understanding of reinforcement learning (RL) agents. We investigate both a common RL visualization, saliency maps (the focus of attention), and a more recent explanation type, reward-decomposition bars (predictions of future types of rewards). We designed a 124 participant, four-treatment experiment to compare participants? mental models of an RL agent in a simple Real-Time Strategy (RTS) game. Our results show that the combination of both saliency and reward bars were needed to achieve a statistically significant improvement in mental model score over the control. In addition, our qualitative analysis of the data reveals a number of effects for further study.
Andrew Anderson 0002, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Alan Fern, Margaret M. Burnett
IJCAI7
2017 Context in exploratory programming: Towards a theoretical framework
abstract
Creativity theory states good designs are achieved by having a multitude of these designs [1]. Exploratory Programming is the process of trying out designs while writing software. Programmers have to evaluate these alternative implementations in order to implement new ideas [2]. These alternatives often have multiple objectives which might prompt a programmer to work towards multiple goals in episodes. Episodes are distinct periods when a programmer works towards a certain goal. These episodes may be interleaved where programmers compare different episodes [3]. However, there is little research which focuses on how programmers abstract meaningful and appropriate information from their exploration of alternatives to integrate into their current work. We refer to these meaningful abstractions as context.
Souti Chattopadhyay
VL/HCC1
2017 What makes a task difficult? An empirical study of perceptions of task difficulty
abstract
Estimating the difficulty of tasks is imperative for project planning, task assignment, and cost calculation. However, little is known about how and for what purpose software practitioners estimate task difficulty in their day-to-day work. In this paper, we interviewed 15 professionals to understand their needs and perceptions when estimating task difficulty. We find that practitioners do estimate the difficulty of tasks for scheduling and prioritizing their work. Additionally, performing such estimation requires more than one metric, and across more than one domain (i.e. code metrics, process metrics, and task metrics). The use of metrics that encapsulate different aspects of a task allows developers to gain a holistic view of the task and its potential difficulty.
Rafael Leano, Souti Chattopadhyay, Anita Sarma
VL/HCC2