VLDB 2026 Research / reviewers in the wild / expert
Chengxi Li 0011
dblp:334/1846
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 77% Empirical software engineering · 23% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering
benchmarking |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation › code generation evaluation
code generation benchmark |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation
code generation evaluation |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation › code generation with language models
data science code generation |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation
code generation with language models |
0.2 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
surface-form constraint checking · 0.7functional correctness testing · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationabstractWe introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io. Yuhang Lai, Chengxi Li 0011, Ruiqi Zhong, Luke Zettlemoyer, Scott Yih, Daniel Fried, Sida I. Wang, Tao Yu 0009 |
ICML | 2 |