At the end of last year, several large language models were asked to predict the final standings of the football World Cup. The tested systems included Grok, ChatGPT, Gemini, Copilot and DeepSeek. Their forecasts varied, and among them Microsoft Copilot proved to be the most accurate in this case.
What the AIs predicted
- Grok: Argentina — France — Spain
- ChatGPT: Spain — France — Brazil
- Gemini: France — Brazil — Spain
- Copilot: Spain — Argentina — France
- DeepSeek: Argentina — France — England
The Sunday final was contested between Spain and Argentina, with Spain winning the tournament. Only Copilot had forecast both finalists and the champion correctly. The bronze medal, however, went to England rather than France, so some models still missed parts of the podium.
Unexpected results the AIs did not foresee
Brazil failed to reach the semifinals after Norway beat Brazil 2–1, eliminating the South American team from the top four. Both ChatGPT and Gemini had predicted a stronger finish for Brazil, but reality differed.
Grok and DeepSeek produced predictions that were far from the final outcome; the article notes these models effectively misled fans by overrating Argentina in their placings. Copilot’s performance is singled out because it has made clear errors in earlier forecasts — for example, treating a so-called "cat tax" as real in a previous test.
Why this matters
The experiment illustrates that general-purpose language models show widely varying accuracy when asked to predict sports outcomes. None of the models should be treated as an infallible forecasting tool; their outputs depend on how they process information and the internal heuristics they use.
The AIs’ trials are not over: upcoming events such as the Formula 1 season and the United States midterm elections will provide further opportunities to compare model performance.



