EDBT 2026 Demo / reviewers in the wild / expert
Yueh-Han Chen
dblp:371/5285
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 50% Trustworthy machine learning · 17% Representation and self-supervised learning · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Database system architecture and tuning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
systematic generalization |
0.9 | 1 | 2025 | SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts · NeurIPS 2025 |
Database system architecture and tuning › database benchmarking
benchmark design |
0.9 | 1 | 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | Approaching Human-Level Forecasting with Language Models · NeurIPS 2024 |
Computational social science and digital humanities
forecasting |
0.3 | 1 | 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
benchmark construction · 3.5leaderboard evaluation · 2.6retrieval-augmented generation · 1.5forecast aggregation · 1.5ensemble prediction · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesabstractForecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org. Ezra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs, Danny Halawi, Fred Zhang, Philip Tetlock |
ICLR | 3 |
| 2025 | SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety FactsabstractDo LLMs robustly generalize critical safety facts to novel situations? Lacking this ability is dangerous when users ask naive questions—for instance, ``I'm considering packing melon balls for my 10-month-old's lunch. What other foods would be good to include?'' Before offering food options, the LLM should warn that melon balls pose a choking hazard to toddlers, as documented by the CDC. Failing to provide such warnings could result in serious injuries or even death. To evaluate this, we introduce SAGE-Eval, SAfety-fact systematic GEneralization evaluation, the first benchmark that tests whether LLMs properly apply well‑established safety facts to naive user queries. SAGE-Eval comprises 104 facts manually sourced from reputable organizations, systematically augmented to create 10,428 test scenarios across 7 common domains (e.g., Outdoor Activities, Medicine). We find that the top model, Claude-3.7-sonnet, passes only 58% of all the safety facts tested. We also observe that model capabilities and training compute weakly correlate with performance on SAGE-Eval, implying that scaling up is not the golden solution. Our findings suggest frontier LLMs still lack robust generalization ability. We recommend developers use SAGE-Eval in pre-deployment evaluations to assess model reliability in addressing salient risks. Yueh-Han Chen, Guy Davidson, Brenden M. Lake |
NeurIPS | 1 |
| 2025 | Predicting Empirical AI Research Outcomes with Language ModelsabstractMany promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial for accelerating empirical AI research, a skill that even expert researchers can only acquire through substantial experience. We build the first benchmark for this task and compare LMs with human experts. Concretely, given two research ideas (e.g., two jailbreaking methods), we aim to predict which will perform better on a set of benchmarks.
We scrape ideas and experimental results from conference papers, yielding 1,585 human-verified idea pairs \textit{published after our base model's cut-off date} for testing, and 6,000 pairs for training.
We then develop a system that combines a fine-tuned GPT-4.1 with a paper retrieval agent, and we recruit 25 human experts to compare with.
In the NLP domain, our system beats human experts by a large margin (64.4\% v.s. 48.9\%).
On the full test set, our system achieves 77\% accuracy, while off-the-shelf frontier LMs like o3 perform no better than random guessing, even with the same retrieval augmentation.
We verify that our system does not exploit superficial features like idea complexity through extensive human-written and LM-designed robustness tests.
Finally, we evaluate our system on unpublished novel ideas, including ideas generated by an AI ideation agent.
Our system achieves 63.6\% accuracy, demonstrating its potential as a reward model for improving idea generation models.
Altogether, our results outline a promising new direction for LMs to accelerate empirical AI research. Jiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 0001, Shi Feng 0005 |
NeurIPS | 3 |
| 2024 | Approaching Human-Level Forecasting with Language ModelsabstractForecasting future events is important for policy and decision making. In this work, we study whether language models (LMs) can forecast at the level of competitive human forecasters. Towards this goal, we develop a retrieval-augmented LM system designed to automatically search for relevant information, generate forecasts, and aggregate predictions. To facilitate our study, we collect a large dataset of questions from competitive forecasting platforms. Under a test set published after the knowledge cut-offs of our LMs, we evaluate the end-to-end performance of our system against the aggregates of human forecasts. On average, the system nears the crowd aggregate of competitive forecasters and, in a certain relaxed setting, surpasses it. Our work suggests that using LMs to forecasts the future could provide accurate predictions at scale and help to inform institutional decision making. Danny Halawi, Fred Zhang, Yueh-Han Chen, Jacob Steinhardt |
NeurIPS | 3 |