New Benchmark Evaluates Strategic Thinking in AI Forecast Agents
Researchers have developed "Bench to the Future 2" (BTF-2), a new benchmark to evaluate AI agents' capacity for strategic reasoning in forecasting.

What happened?
A new benchmark called "Bench to the Future 2" (BTF-2) has been developed to assess the strategic reasoning capabilities of AI agents in forecasting. The benchmark consists of 1,417 "pastcasting" questions, where agents research and forecast offline using a fixed research corpus of 15 million documents. This generates complete reasoning traces, allowing for a detailed analysis of the AI agents' processes.
Key facts
| Benchmarknamn | Bench to the Future 2 (BTF-2) |
|---|---|
| Antal frågor | 1 417 |
| Dokumentkorpus | 15 miljoner dokument |
| Möjlig noggrannhetsdifferens | 0.004 Brier score |
”Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to the Future 2 (BTF-2), 1,417 pastcasting questions with a frozen 15M-document research corpus in which agents reproducibly researc”
”BTF-2 detects accuracy differences of 0.004 Brier score, and can distinguish differential agent strengths in research vs. judgment.”
”Expert human forecasters found the dominant strategic reasoning failures of frontier agents are in assessing political and business leaders' incentives, judging their likelihood to follow through on st”
Why it matters
Traditional forecasting benchmarks focus primarily on accuracy and provide limited insight into why certain forecasters perform better than others. BTF-2 enables differentiation between agents' strengths in research versus judgment, and identifies strategic deficiencies such as blind spots and the handling of "black swan" events. This provides a deeper understanding of AI decision-making.
Who is affected?
Researchers and developers in AI forecasting are directly affected, as BTF-2 offers a tool to test and improve strategic reasoning capabilities. Organisations using AI for forecasting can benefit from an increased understanding of AI systems' limitations and strengths. Human forecasting experts also gain insights into the comparative performance of AI.
What else you should know
Expert assessments show that the primary strategic deficiencies in current AI agents lie in the analysis of political and business leaders' incentives and their likelihood of following through on plans.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vad är 'pastcasting'?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "New Benchmark Evaluates Strategic Thinking in AI Forecast Ag"