Skip to content
Longterm Wiki

ForecastBench: Dynamic LLM Forecasting Benchmark

web
forecastbench.org·forecastbench.org

ForecastBench is a dynamic benchmark measuring LLM forecasting accuracy against human baselines, relevant to AI safety as forecasting ability serves as a proxy for general intelligence and helps track AI capability progress toward and beyond human-level performance.

Metadata

Importance: 62/100tool pagetool

Summary

ForecastBench is a contamination-free benchmark that evaluates LLM forecasting accuracy against human comparison groups, including superforecasters. It maintains both a baseline leaderboard (no tools) and a tournament leaderboard (with scaffolding/tools), and projects when LLMs will reach superforecaster-level performance.

Key Points

  • Dynamic, contamination-free benchmark preventing LLMs from training on benchmark questions, ensuring valid capability measurement.
  • Compares LLM forecasting performance against human baselines including superforecasters as a proxy for general intelligence.
  • Dual leaderboards: baseline (raw model performance) and tournament (with tool use, fine-tuning, ensembling).
  • Tracks historical progress in LLM forecasting capabilities and projects date of LLM-superforecaster parity.
  • Open to public submissions, enabling broad participation in capability evaluation.

Cited by 2 pages

PageTypeQuality
Forecasting Research Institute (FRI)Organization55.0
ForecastBenchProject53.0

1 FactBase fact citing this source

EntityPropertyValueAs Of
ForecastBenchFounded DateSep 2024

Cached Content Preview

HTTP 200Fetched Aug 2, 20260 KB
Tournament leaderboard
Resource ID: kb-c808dd961e2e3c1d | Stable ID: sid_WNgUtLr8jP