VLDB 2026 Research / reviewers in the wild / expert
Celine Lee
dblp:259/9736
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 70% Compilers and program optimization · 22% Software testing · 8% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
code generation |
0.9 | 1 | 2025 | Commit0: Library Generation from Scratch · ICLR 2025 |
Program synthesis and code generation › generative programming
library generation |
0.9 | 1 | 2025 | Commit0: Library Generation from Scratch · ICLR 2025 |
Program synthesis and code generation
code translation |
0.8 | 1 | 2024 | Guess & Sketch: Language Model Guided Transpilation · ICLR 2024 |
Program synthesis and code generation › neural program synthesis
neurosymbolic program synthesis |
0.8 | 1 | 2024 | Guess & Sketch: Language Model Guided Transpilation · ICLR 2024 |
Compilers and program optimization › program transformation
source-to-source transformation |
0.8 | 1 | 2024 | Guess & Sketch: Language Model Guided Transpilation · ICLR 2024 |
Software testing
unit testing |
0.3 | 1 | 2025 | Commit0: Library Generation from Scratch · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
static analysis · 1.7interactive feedback · 1.7symbolic solver · 0.8neurosymbolic approach · 0.8language model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Commit0: Library Generation from ScratchabstractWith the goal of benchmarking generative systems beyond expert software development ability, we introduce Commit0, a benchmark that challenges AI agents to write libraries from scratch. Agents are provided with a specification document outlining the library’s API as well as a suite of interactive unit tests, with the goal of producing an implementation of this API accordingly. The implementation is validated through running these unit tests. As a benchmark, Commit0 is designed to move beyond static one-shot code generation towards agents that must process long-form natural language specifications, adapt to multi-stage feedback, and generate code with complex dependencies. Commit0 also offers an interactive environment where models receive static analysis and execution feedback on the code they generate. Our experiments demonstrate that while current agents can pass some unit tests, none can yet fully reproduce full libraries. Results also show that interactive feedback is quite useful for models to generate code that passes more unit tests, validating the benchmarks that facilitate its use. We publicly release the benchmark, the interactive environment, and the leaderboard. Celine Lee, Justin T. Chiu, Claire Cardie, Matthias Gallé, Alexander M. Rush |
ICLR | 3 |
| 2024 | Guess & Sketch: Language Model Guided TranspilationabstractMaintaining legacy software requires many software and systems engineering hours. Assembly code programs, which demand low-level control over the computer machine state and have no variable names, are particularly difficult for humans to analyze.
Existing conventional program translators guarantee correctness, but are hand-engineered for the source and target programming languages in question. Learned transpilation, i.e. automatic translation of code, offers an alternative to manual re-writing and engineering efforts. Automated symbolic program translation approaches guarantee correctness but struggle to scale to longer programs due to the exponentially large search space. Their rigid rule-based systems also limit their expressivity, so they can only reason about a reduced space of programs. Probabilistic neural language models (LMs) produce plausible outputs for every input, but do so at the cost of guaranteed correctness. In this work, we leverage the strengths of LMs and symbolic solvers in a neurosymbolic approach to learned transpilation for assembly code. Assembly code is an appropriate setting for a neurosymbolic approach, since assembly code can be divided into shorter non-branching basic blocks amenable to the use of symbolic methods. Guess & Sketch extracts alignment and confidence information from features of the LM then passes it to a symbolic solver to resolve semantic equivalence of the transpilation input and output. We test Guess & Sketch on three different test sets of assembly transpilation tasks, varying in difficulty, and show that it successfully transpiles 57.6% more examples than GPT-4 and 39.6% more examples than an engineered transpiler. We also share a training and evaluation dataset for this task. Celine Lee, Abdulrahman Mahmoud, Michal Kurek, Simone Campanoni, David Brooks 0001, Stephen Chong, Gu-Yeon Wei, Alexander M. Rush |
ICLR | 1 |