VLDB 2026 Research / reviewers in the wild / expert
Ezra Karger
dblp:348/9676
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0003-8035-8239ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Theory of computation · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesabstractForecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org. Ezra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs, Danny Halawi, Fred Zhang, Philip Tetlock |
ICLR | 1 |
| 2025 | Self-Resolving Prediction Markets for Unverifiable OutcomesabstractPrediction markets elicit and aggregate beliefs by paying agents based on how close their predictions are to a verifiable future outcome. However, outcomes of many important questions are difficult to verify or unverifiable, in that the ground truth may be hard or impossible to access. We present a novel incentive-compatible prediction market mechanism to elicit and efficiently aggregate information from a pool of agents without observing the outcome, by paying agents the negative cross-entropy between their prediction and that of a carefully chosen reference agent. Our key insight is that a reference agent with access to more information can serve as a reasonable proxy for the ground truth. We use this insight to propose self-resolving prediction markets that terminate with some probability after every report and pay all but a few agents based on the final prediction. The final agent is chosen as the reference agent since they observe the full history of market forecasts, and thus have more information by design. We show that it is a perfect Bayesian equilibrium (PBE) for all agents to report truthfully in our mechanism and to believe that all other agents report truthfully. Although primarily of interest for unverifiable outcomes, this design is also applicable for verifiable outcomes. Siddarth Srinivasan, Ezra Karger, Yiling Chen 0001 |
EC | 2 |
| 2025 | AI-Augmented Predictions: LLM Assistants Improve Human Forecasting AccuracyabstractLarge language models (LLMs) match and sometimes exceed human performance in many domains. This study explores the potential of LLMs to augment human judgment in a forecasting task. We evaluate the effect on human forecasters of two LLM assistants: one designed to provide high-quality (“superforecasting”) advice, and the other designed to be overconfident and base-rate neglecting, thus providing noisy forecasting advice. We compare participants using these assistants to a control group that received a less advanced model that did not provide numerical predictions or engage in explicit discussion of predictions. Participants ( N \(=\) 991) answered a set of six forecasting questions and had the option to consult their assigned LLM assistant throughout. Our preregistered analyses show that interacting with each of our frontier LLM assistants significantly enhances prediction accuracy by between 24% and 28% compared to the control group. Exploratory analyses showed a pronounced outlier effect in one forecasting item, without which we find that the superforecasting assistant increased accuracy by 41%, compared with 29% for the noisy assistant. We further examine whether LLM forecasting augmentation disproportionately benefits less skilled forecasters, degrades the wisdom-of-the-crowd by reducing prediction diversity, or varies in effectiveness with question difficulty. Our data do not consistently support these hypotheses. Our results suggest that access to a frontier LLM assistant, even a noisy one, can be a helpful decision aid in cognitively demanding tasks compared to a less powerful model that does not provide specific forecasting advice. However, the effects of outliers suggest that further research into the robustness of this pattern is needed. Philipp Schoenegger, Peter S. Park, Ezra Karger, Sean Trott, Philip Tetlock |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2024 | Full Accuracy Scoring Accelerates the Discovery of Skilled ForecastersabstractReliable detection of skilled forecasters is slow and resource-intensive. The gold-standard approach relies on proper scores and requires forecasters to answer dozens of questions, which may take months or years to resolve. To accelerate skill identification, we propose the Full Accuracy Score (FAS). FAS combines the strengths of objective ground-truth proper scores with the early availability of proper proxy scores that measure the distance between individual and consensus estimates (Witkowski et al. 2017). FAS treats the two inputs as complements, using ground-truth scores on resolved questions and proxy scores on questions with yet unknown answers. The proxy component acts as a running tally of individual performance and helps to complete the evolving picture of relative skill. Pavel Atanasov, Ezra Karger, Philip Tetlock |
EC | 2 |