David Gros 0001

dblp:276/0345-1 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0001-4212-2204ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
abstract
The rise of large language models for code has reshaped software development. Autonomous coding agents, able to create branches, open pull requests, and perform code reviews, now actively contribute to real-world projects. Their growing role offers a unique and timely opportunity to investigate AI-driven contributions and their effects on code quality, team dynamics, and software maintainability. In this work, we construct a novel dataset of approximately 110,000 open-source pull requests, including associated commits, comments, reviews, issues, and file changes, collectively representing millions of lines of source code. We compare five popular coding agents, including OpenAI Codex, Claude Code, GitHub Copilot, Google Jules, and Devin, examining how their usage differs in various development aspects such as merge frequency, edited file types, and developer interaction signals, including comments and reviews. Furthermore, we emphasize that code authoring and review are only a small part of the larger software engineering process, as the resulting code must also be maintained and updated over time. Hence, we offer several longitudinal estimates of survival and churn rates for agent-generated versus human-authored code. Ultimately, our findings indicate an increasing agent activity in open-source projects, although their contributions are associated with more churn over time compared to human-authored code.
Razvan Mihai Popescu, David Gros 0001, Andrei Botocan, Rahul Pandita, Premkumar T. Devanbu, Maliheh Izadi
MSR2
2025 Calibration and Correctness of Language Models for Code
abstract
Machine learning models are widely used, but can also often be wrong. Users would benefit from a reliable indication of whether a given output from a given model should be trusted, so a rational decision can be made whether to use the output or not. For example, outputs can be associated with a confidence measure; if this confidence measure is strongly associated with likelihood of correctness, then the model is said to be well-calibrated. A well-calibrated confidence measure can serve as a basis for rational, graduated decision-making on how much review and care is needed when using generated code. Calibration has so far been studied in mostly non-generative (e.g., classification) settings, especially in software engineering. However, generated code can quite often be wrong: Given generated code, developers must decide whether to use directly, use after varying intensity of careful review, or discard model-generated code. Thus, calibration is vital in generative settings. We make several contributions. We develop a framework for evaluating the calibration of code-generating models. We consider several tasks, correctness criteria, datasets, and approaches, and find that, by and large, generative code models we test are not well-calibrated out of the box. We then show how calibration can be improved using standard methods, such as Platt scaling. Since Platt scaling relies on the prior availability of correctness data, we evaluate the applicability and generalizability of Platt scaling in software engineering, discuss settings where it has good potential for practical use, and settings where it does not. Our contributions will lead to better-calibrated decision-making in the current use of code generated by language models, and offers a framework for future research to further improve calibration methods for generative models in software engineering.
Claudio Spiess, David Gros 0001, Kunal Suresh Pai, Michael Pradel, Md. Rafiqul Islam Rabin, Mohammad Amin Alipour, Susmit Jha, Premkumar T. Devanbu, Toufique Ahmed
ICSE2
2022 Robots-Dont-Cry: Understanding Falsely Anthropomorphic Utterances in Dialog Systems
abstract
Dialog systems are often designed or trained to output human-like responses.However, some responses may be impossible for a machine to truthfully say (e.g."that movie made me cry").Highly anthropomorphic responses might make users uncomfortable or implicitly deceive them into thinking they are interacting with a human.We collect human ratings on the feasibility of approximately 900 two-turn dialogs sampled from 9 diverse data sources.Ratings are for two hypothetical machine embodiments: a futuristic humanoid robot and a digital assistant.We find that for some data-sources commonly used to train dialog systems, 20-30% of utterances are not viewed as possible for a machine.Rating is marginally affected by machine embodiment.We explore qualitative and quantitative reasons for these ratings.Finally, we build classifiers and explore how modeling configuration might affect output permissibly, and discuss implications for building less falsely anthropomorphic dialog systems.
David Gros 0001, Yu Li 0013, Zhou Yu 0005
EMNLP1
2021 The R-U-A-Robot Dataset: Helping Avoid Chatbot Deception by Detecting User Questions About Human or Non-Human Identity
abstract
David Gros, Yu Li, Zhou Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
David Gros 0001, Yu Li 0013, Zhou Yu 0005
ACL/IJCNLP (1)1
2020 Code to Comment "Translation": Data, Metrics, Baselining & Evaluation
abstract
The relationship of comments to code, and in particular, the task of generating useful comments given the code, has long been of interest. The earliest approaches have been based on strong syntactic theories of comment-structures, and relied on textual templates. More recently, researchers have applied deep-learning methods to this task---specifically, trainable generative translation models which are known to work very well for Natural Language translation (e.g., from German to English). We carefully examine the underlying assumption here: that the task of generating comments sufficiently resembles the task of translating between natural languages, and so similar models and evaluation metrics could be used. We analyze several recent code-comment datasets for this task: CodeNN, DeepCom, FunCom, and DocString. We compare them with WMT19, a standard dataset frequently used to train state-of-the-art natural language translators. We found some interesting differences between the code-comment data and the WMT19 natural language data. Next, we describe and conduct some studies to calibrate BLEU (which is commonly used as a measure of comment quality). using "affinity pairs" of methods, from different projects, in the same project, in the same class, etc; Our study suggests that the current performance on some datasets might need to be improved substantially. We also argue that fairly naive information retrieval (IR) methods do well enough at this task to be considered a reasonable baseline. Finally, we make some suggestions on how our findings might be used in future research in this area.
David Gros 0001, Hariharan Sezhiyan, Premkumar T. Devanbu, Zhou Yu 0005
ASE1