EDBT 2026 Demo / reviewers in the wild / expert
Zachary Jacobs
dblp:199/4250
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Database system architecture and tuning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities · ICLR 2025 |
Database system architecture and tuning › database benchmarking
benchmark design |
0.9 | 1 | 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities · ICLR 2025 |
Computational social science and digital humanities
forecasting |
0.3 | 1 | 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
leaderboard evaluation · 2.6benchmark construction · 2.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesabstractForecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org. Ezra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs, Danny Halawi, Fred Zhang, Philip Tetlock |
ICLR | 4 |