VLDB 2026 Research / reviewers in the wild / expert
James Aung
dblp:383/8457
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 75% Efficient and distributed learning · 25% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | PaperBench: Evaluating AI's Ability to Replicate AI Research · ICML 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM judge |
0.9 | 1 | 2025 | PaperBench: Evaluating AI's Ability to Replicate AI Research · ICML 2025 |
Computing education › machine learning research practices
AI research reproducibility |
0.9 | 1 | 2025 | PaperBench: Evaluating AI's Ability to Replicate AI Research · ICML 2025 |
Program synthesis and code generation
code generation evaluation |
0.3 | 1 | 2025 | PaperBench: Evaluating AI's Ability to Replicate AI Research · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
rubric-based grading · 2.6large language model evaluation · 0.9agent scaffolding · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringabstractWe introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test real-world ML engineering skills such as training models, preparing datasets, and running experiments. We establish human baselines for each competition using Kaggle's publicly available leaderboards. We use open-source agent scaffolds to evaluate several frontier language models on our benchmark, finding that the best-performing setup — OpenAI's o1-preview with AIDE scaffolding — achieves at least the level of a Kaggle bronze medal in 16.9% of competitions. In addition to our main results, we investigate various forms of resource-scaling for AI agents and the impact of contamination from pre-training. We open-source our benchmark code https://github.com/openai/mle-bench to facilitate future research in understanding the ML engineering capabilities of AI agents. Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Aleksander Madry, Lilian Weng |
ICLR | 4 |
| 2025 | PaperBench: Evaluating AI's Ability to Replicate AI ResearchabstractWe introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research. Agents must replicate 20 ICML 2024 Spotlight and Oral papers from scratch, including understanding paper contributions, developing a codebase, and successfully executing experiments. For objective evaluation, we develop rubrics that hierarchically decompose each replication task into smaller sub-tasks with clear grading criteria. In total, PaperBench contains 8,316 individually gradable tasks. Rubrics are co-developed with the author(s) of each ICML paper for accuracy and realism. To enable scalable evaluation, we also develop an LLM-based judge to automatically grade replication attempts against rubrics, and assess our judge’s performance by creating a separate benchmark for judges. We evaluate several frontier models on PaperBench, finding that the best-performing tested agent, Claude 3.5 Sonnet (New) with open-source scaffolding, achieves an average replication score of 21.0%. Finally, we recruit top ML PhDs to attempt a subset of PaperBench, finding that models do not yet outperform the human baseline. We open-source our code (https://github.com/openai/preparedness) to facilitate future research in understanding the AI engineering capabilities of AI agents. Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, Tejal Patwardhan |
ICML | 4 |