VLDB 2026 Research / reviewers in the wild / expert
Yuhang Lai
dblp:334/2231
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 77% Empirical software engineering · 23% | |
| Artificial intelligence
1 paper |
Reinforcement learning · 50% Language models and text generation · 50% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | HAF-RM: A Hybrid Alignment Framework for Reward Model Training · ACL (1) 2025 |
Machine learning › Reinforcement learning › reward learning
reward model training |
0.9 | 1 | 2025 | HAF-RM: A Hybrid Alignment Framework for Reward Model Training · ACL (1) 2025 |
Empirical software engineering
benchmarking |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation › code generation evaluation
code generation benchmark |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation
code generation evaluation |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation › code generation with language models
data science code generation |
0.7 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Program synthesis and code generation
code generation with language models |
0.2 | 1 | 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
reward modeling · 0.9hybrid alignment · 0.9surface-form constraint checking · 0.7functional correctness testing · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HAF-RM: A Hybrid Alignment Framework for Reward Model TrainingabstractShujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang, Shengbin Yue, Zengfeng Huang, Xuanjing Huang, Zhongyu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shujun Liu, Yuhang Lai, Siyuan Wang 0025, Shengbin Yue, Zengfeng Huang, Xuanjing Huang 0001, Zhongyu Wei |
ACL (1) | 3 |
| 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationabstractWe introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io. Yuhang Lai, Chengxi Li 0011, Ruiqi Zhong, Luke Zettlemoyer, Scott Yih, Daniel Fried, Sida I. Wang, Tao Yu 0009 |
ICML | 1 |