Philip Tetlock

dblp:126/6307 · also Philip E. Tetlock · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Theory of computation · 3 · 2 since 2021
YearPublicationVenuePosition
2025 ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
abstract
Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org.
Ezra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs, Danny Halawi, Fred Zhang, Philip Tetlock
ICLR7
2025 AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
abstract
Large language models (LLMs) match and sometimes exceed human performance in many domains. This study explores the potential of LLMs to augment human judgment in a forecasting task. We evaluate the effect on human forecasters of two LLM assistants: one designed to provide high-quality (“superforecasting”) advice, and the other designed to be overconfident and base-rate neglecting, thus providing noisy forecasting advice. We compare participants using these assistants to a control group that received a less advanced model that did not provide numerical predictions or engage in explicit discussion of predictions. Participants ( N \(=\) 991) answered a set of six forecasting questions and had the option to consult their assigned LLM assistant throughout. Our preregistered analyses show that interacting with each of our frontier LLM assistants significantly enhances prediction accuracy by between 24% and 28% compared to the control group. Exploratory analyses showed a pronounced outlier effect in one forecasting item, without which we find that the superforecasting assistant increased accuracy by 41%, compared with 29% for the noisy assistant. We further examine whether LLM forecasting augmentation disproportionately benefits less skilled forecasters, degrades the wisdom-of-the-crowd by reducing prediction diversity, or varies in effectiveness with question difficulty. Our data do not consistently support these hypotheses. Our results suggest that access to a frontier LLM assistant, even a noisy one, can be a helpful decision aid in cognitively demanding tasks compared to a less powerful model that does not provide specific forecasting advice. However, the effects of outliers suggest that further research into the robustness of this pattern is needed.
Philipp Schoenegger, Peter S. Park, Ezra Karger, Sean Trott, Philip Tetlock
ACM Trans. Interact. Intell. Syst.5
2024 Full Accuracy Scoring Accelerates the Discovery of Skilled Forecasters
abstract
Reliable detection of skilled forecasters is slow and resource-intensive. The gold-standard approach relies on proper scores and requires forecasters to answer dozens of questions, which may take months or years to resolve. To accelerate skill identification, we propose the Full Accuracy Score (FAS). FAS combines the strengths of objective ground-truth proper scores with the early availability of proper proxy scores that measure the distance between individual and consensus estimates (Witkowski et al. 2017). FAS treats the two inputs as complements, using ground-truth scores on resolved questions and proxy scores on questions with yet unknown answers. The proxy component acts as a running tally of individual performance and helps to complete the evolving picture of relative skill.
Pavel Atanasov, Ezra Karger, Philip Tetlock
EC3
2022 Crowd Prediction Systems: Markets, Polls, and Elite Forecasters
abstract
No abstract available.
Pavel Atanasov, Jens Witkowski, Barbara A. Mellers, Philip Tetlock
EC4
2020 Small Steps to Accuracy: Incremental Belief Updaters Are Better Forecasters
abstract
Laboratory research has shown that both underreaction and overreaction to new information pose threats to forecasting accuracy. This article explores how real-world forecasters who vary in skill attempt to balance these threats. We distinguish among three aspects of updating: frequency, magnitude, and confirmation propensity. Drawing on data from a four-year forecasting tournament that elicited over 400,000 probabilistic predictions on almost 500 geopolitical questions, we found that the most accurate forecasters made frequent, small updates, while low-skill forecasters were prone to confirm initial judgments or make infrequent, large revisions. High-frequency updaters scored higher on crystallized intelligence and open-mindedness, accessed more information, and improved over time. Small-increment updaters had higher fluid intelligence scores, and derived their advantage from initial forecasts. Update magnitude mediated the causal effect of training on accuracy. Frequent, small revisions provided reliable and valid signals of skill. These updating patterns can help organizations identify talent for managing uncertain prospects.
Pavel Atanasov, Jens Witkowski, Lyle H. Ungar, Barbara A. Mellers, Philip Tetlock
EC5
2017 Assessing Objective Recommendation Quality through Political Forecasting
abstract
Recommendations are often rated for their subjective quality, but few researchers have studied quality in terms of objective utility.We explore quality assessment with respect to both subjective (i.e.users' ratings) and objective (i.e., did it influence?did it improve decisions?)metrics in a massive online geopolitical forecasting system, ultimately comparing linguistic characteristics of each quality metric.Using a variety of features, we predict all types of quality with better accuracy than the simple yet strong baseline of recommendation length.For example, more complex sentence constructions, as evidenced by subordinate conjunctions, are characteristic of recommendations leading to objective improvements in forecasting.Our analyses also reveal rater biases; for example, forecasters are subjectively biased in favor of recommendations mentioning business deals and material things, even though such recommendations do not indeed prove any more useful objectively.
H. Andrew Schwartz, Masoud Rouhizadeh, Michael Bishop, Philip Tetlock, Barbara A. Mellers, Lyle H. Ungar
EMNLP4