In a research paper published on the American website arXiv for studies and research, researchers proposed a new test to measure the ability of language models and deep search agents to predict football matches before they start, using the 2026 World Cup as the first applied experiment.








The research, translated by “Lebanon 24,” says that predicting the result of a football match before the starting whistle does not depend only on knowledge of previous results, but rather requires the use of variable information and providing a clear prediction before the result becomes known.

Therefore, the researchers introduced a new benchmark called “WorldCupArena”, which is a dynamic testbed dedicated to language models and deep search agents. The test, in its first version, is based on the 2026 World Cup matches, with the possibility of reusing the same approach in other tournaments and leagues in the future.

According to the research, before each match the model either receives a standardized information package, or is asked to search for information on its own. After that, he predicts the score and the numerical score, the players expected to excel, possible events, match statistics, in addition to the course of the tournament and its final result.

After the match ends, these predictions are compared to the official score. The test not only measures the accuracy of the winner’s prediction or the exact result, but also uses a criterion that awards some points when the predicted result is close to the actual result, even if it is not an exact match.

The evaluation included 104 matches and 13 different systems. The researchers found that models may achieve similar percentages in predicting the winner, but they differ more clearly when testing details, such as the numerical score, players, events, and statistics.

The research indicates that the best system achieved only limited superiority over betting markets and fans’ expectations in predicting the winner and the exact result, but it showed clearer superiority in predicting results close to reality.

The researchers also concluded that adding online research does not always help improve match predictions, and that some models fail when they tend to select the same stronger candidate. The importance of “WorldCupArena” is that it allows future models to be evaluated on matches that have not yet been played, rather than testing them on results that are already known.