VLDB 2026 Research / reviewers in the wild / expert
Sydney Nguyen
dblp:326/8051
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0002-7053-699XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 44% Empirical software engineering · 44% Programming languages and type systems · 13% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering
benchmarking |
0.7 | 1 | 2023 | MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation · IEEE Trans. Software Eng. 2023 |
Program synthesis and code generation
code generation with language models |
0.7 | 1 | 2023 | MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation · IEEE Trans. Software Eng. 2023 |
Computing education
programming education |
0.2 | 1 | 2024 | How Beginning Programmers and Code LLMs (Mis)read Each Other · CHI 2024 |
Programming languages and type systems
programming paradigms |
0.2 | 1 | 2023 | MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation · IEEE Trans. Software Eng. 2023 |
Methods — techniques the papers use, named apart from their topics
mixed-methods evaluation · 1.5controlled study · 1.5large language model evaluation · 0.7benchmark translation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | How Beginning Programmers and Code LLMs (Mis)read Each OtherabstractGenerative AI models, specifically large language models (LLMs), have made strides towards the long-standing goal of text-to-code generation. This progress has invited numerous studies of user interaction. However, less is known about the struggles and strategies of non-experts, for whom each step of the text-to-code problem presents challenges: describing their intent in natural language, evaluating the correctness of generated code, and editing prompts when the generated code is incorrect. This paper presents a large-scale controlled study of how 120 beginning coders across three academic institutions approach writing and editing prompts. A novel experimental design allows us to target specific steps in the text-to-code process and reveals that beginners struggle with writing and editing prompts, even for problems at their skill level and when correctness is automatically determined. Our mixed-methods evaluation provides insight into student processes and perceptions with key implications for non-expert Code LLM use within and outside of education. Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, Molly Q. Feldman |
CHI | 1 |
| 2023 | MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code GenerationabstractLarge language models have demonstrated the ability to generate both natural language and programming language text. Although contemporary code generation models are trained on corpora with several programming languages, they are tested using benchmarks that are typically monolingual. The most widely used code generation benchmarks only target Python, so there is little quantitative evidence of how code generation models perform on other programming languages. We propose MultiPL-E, a system for translating unit test-driven code generation benchmarks to new languages. We create the first massively multilingual code generation benchmark by using MultiPL-E to translate two popular Python code generation benchmarks to 18 additional programming languages. We use MultiPL-E to extend the HumanEval benchmark [1] and MBPP benchmark [2] to 18 languages that encompass a range of programming paradigms and popularity. Using these new parallel benchmarks, we evaluate the multi-language performance of three state-of-the-art code generation models: Codex [1], CodeGen [3]and InCoder [4]. We find that Codex matches or even exceeds its performance on Python for several other languages. The range of programming languages represented in MultiPL-E allow us to explore the impact of language frequency and language features on model performance. Finally, the MultiPL-E approach of compiling code generation benchmarks to new programming languages is both scalable and extensible, making it straightforward to evaluate new models, benchmarks, and languages. Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q. Feldman, Arjun Guha, Michael Greenberg 0002, Abhinav Jangda |
IEEE Trans. Software Eng. | 4 |