The test
So we tested Torchcast under constraint. We submitted three new configurations using materially smaller and less capable models. We called them Nano.
Among systems entering the same ForecastBench round, they finished:
🥇 carb-bomb-nano
🥈 wyrm-warlord-nano
🥉 skipper-morgan-nano
On the broader dataset leaderboard, which combines systems evaluated across different rounds and question cohorts, the three Nanos currently rank #13, #17, and #19 overall. Two score above the median human Superforecaster. The third is only slightly below it.
The Nano systems did not replace our flagship systems at the top. They showed something different: Torchcast's forecasting capability survives even when the underlying models become substantially weaker.
How we train
That reflects how we think about training. We train our forecasting models, and the system around them, not to produce the most persuasive explanation, but to make better probabilistic updates. No individual model is treated as an oracle. Torchcast learns which evidence is genuinely new, which signals repeat the same underlying information, how to weigh competing judgments, and when new evidence should or should not move a forecast.
The objective is not to sound certain. The objective is to be uncertain correctly.
The ceiling and the floor
Foundation model intelligence still matters. Our strongest systems show how far we can push forecasting performance. But institutional forecasting must eventually operate continuously across thousands, or even millions, of changing questions. It cannot depend on using the largest and most expensive model for every update.
Our flagship systems show the ceiling. Nano shows that the judgment survives.
If AI is going to support real-world decisions, it must learn not only how to reason, but how certain to be.
