1. Overview
1. Overview
If you're given only the first 10 minutes of a football match, can you predict the eventual result? Predicting the exact scoreline isn't realistic for such a low-scoring, high-variance sport, so the target here is the 3-way result (Home win / Draw / Away win), always measured against a naive baseline ("always predict the most common result in the training data") so "better than random guessing" actually means something.
Three successive model variants were built and tested:
◼
Model A –factoring in 5 in-game stats (shots, xG, completed passes, possession share, fouls) from the first 10 minutes of 230 matches (World Cup 2018 & 2022, Euro 2020 & 2024).
◼
Model B – the same 5 stats, but the sample grown to 314 matches by adding AFCON 2023 and Copa America 2024 – testing whether more data helps.
◼
Model C – Model B’s data plus one external feature: each team’s FIFA World Ranking points at the time of the match, joined from a separate historical ranking dataset by team name and nearest date before kick-off – testing whether external team-strength data helps.
To remove any discrepancy, all data tested was taken from top-flight international tournaments (World Cup, AFCON, Euros, Copa America) within the last 8 years only.
2. Real Results
2. Real Results
These numbers come from actually running the classifier below:
"Model" | "Matches" | "Accuracy" | "Baseline" |
"A (original)" | "230" | "37.0%" | "28.3%" |
"B (expanded)" | "314" | "49.2%" | "38.1%" |
"C (+ FIFA ranking)" | "307" | "49.2%" | "39.3%" |
Headline finding: adding more real match data (A B) gave a large, genuine accuracy jump, as expected. Surprisingly though, adding the FIFA World Ranking feature on top of that (B C) made essentially no further difference – in fact, it led to a marginal decrease in the relative accuracy of our model. Possible explanation: team quality already shows up indirectly in the first 10 minutes of play (a stronger team tends to have more of the ball and more shots early on anyway), so a separate "how good is this team on paper" feature is largely redundant information the model already had. Also quite likely that there simply isn't enough data to significantly reflect an increase in the accuracy percentage between Model B and Model C.
3. Wolfram Language
3. Wolfram Language
3.1 Load the cached features
3.1 Load the cached features
In[]:=
SetDirectory[NotebookDirectory[]];restoreMissing[assoc_]:=If[KeyExistsQ[assoc,"rankPointsDiff"]&&assoc["rankPointsDiff"]===Null,Append[assoc,"rankPointsDiff"->Missing["NoRanking"]],assoc];allFeatures=restoreMissing/@(Association/@Import["features_cache_v2.json","RawJSON"]);Length[allFeatures]
Out[]=
0
3.2 Train and evaluate a model
3.2 Train and evaluate a model
trainingRules=(KeyDrop[#,"label"]->#["label"])&/@allFeatures;SeedRandom[42];shuffled=RandomSample[trainingRules];testSet=shuffled[[1;;63]];trainSet=shuffled[[64;;]];classifier=Classify[trainSet];ClassifierMeasurements[classifier,testSet]["Accuracy"]
0.492063
That's the entire model-training step: `Classify[trainSet]` picks the method, trains it, and `ClassifierMeasurements` evaluates it – no algorithm was named anywhere above.
3.3 Visual analysis
3.3 Visual analysis
BarChart[Counts[allFeatures[[All,"label"]]]//Values,ChartLabels->Counts[allFeatures[[All,"label"]]]//Keys,PlotLabel->"Match outcome distribution",ImageSize->500]
Outcome split across all 314 matches – home advantage is visible, despite the fact that there should be no such advantage due to these games being played at neutral venues (wherever the host country of the tournament is). Draws are the smallest group, which foreshadows why they're the hardest class to predict below.
Chart 2: accuracy vs baseline across all three model variants - see Section 2 table
BarChart{{0.369565,0.282609},{0.492063,0.380952},{0.491803,0.393443}},
Chart 3: Model C’s confusion matrix, as a grouped bar chart
BarChart[{{13,1,8},{4,1,10},{6,2,16}},ChartLabels->{("Actual "<>#)&/@{A,D,H},("Predicted "<>#)&/@{A,D,H}},PlotLabel->"Model C: predictions per actual outcome"]
Observation: draws are extremely hard for the model to accurately predict. Across every model variant tried, the classifier gets very few actual draws right (see the "Actual D" group above) – a well-known, real phenomenon in football analytics, not a bug in this pipeline. Draws don't have a distinct statistical "signature" the way dominant or losing performances do, so this outcome is rarely predicted. Since a draw is sandwiched between two more-locally probable neighbors, the classifier hardly tends to predict it. This could be something to tweak in future versions of the model.
4. Model Developments
4. Model Developments
◼
Tried and worked: growing the sample from 230 to 314 matches (added AFCON 2023, Copa America 2024). Accuracy rose from 37.0% to 49.2% – the single biggest improvement in the whole project, and it came from more data, not smarter modelling.
◼
Tried and didn't help (on its own): joining each team's FIFA World Ranking points at the time of the match. Despite being one of the most respected pre-match strength signals in real football analytics, it added no measurable accuracy here (49.2% 49.2%) once the in-game stats and larger sample were already in place.
◼
Considered and deliberately skipped: individual players' FIFA video-game OVR ratings (e.g. from sofifa.com). This would require matching StatsBomb's player names against a completely different database's naming conventions (nicknames, accents, transliteration) – a large, error-prone task for the time available, but a reasonable future project extension.
6. Advantages / Disadvantages of each language
6. Advantages / Disadvantages of each language
Advantages of WL over Python
Advantages of WL over Python
◼
Built in, specialised functions; no imported libraries needed
◼
Fewer lines of code overall (376 vs 337)
◼
Classify[] was extremely helpful specifically for this project
◼
Much easier to create visuals (as seen above)
◼
Use of WolframAlpha or Free Form Input very useful
◼
Symbolic language
◼
Notebooks: good mixed text interface
◼
Powerful
Advantages of Python over WL
Advantages of Python over WL
◼
Open source / community based + general-purpose
◼
Free
◼
Larger community for troubleshooting problems etc.
◼
More lines means easier to follow step by step logic?
◼
Massive ecosystem of libraries
Overall, Wolfram Language was the better choice for this project specifically.
7. Talking Points + Recommendations to 6th Form Students & Teachers
7. Talking Points + Recommendations to 6th Form Students & Teachers
◼
Use WL's Classify/Predict tools if without ML knowledge.
◼
Great for an efficient first investigation.
◼
Also good for any maths-related pursuits
◼
Make use of WL's specialised functions for specific tasks
◼
Use python for its versatility and for basic coding learning
◼
Despite WL's strengths, it will inevitably be hard to shift coding in a school context away from Python (since it is the most common language at GCSE/A level etc).
◼
CBM (and ultimately a new A-level) is a really good way to properly introduce WL into education and make students and teachers properly aware of it.
◼
For teachers: importance of WL in any non-CS subject (e.g. A-level maths): perfect for adding real computation to a lesson using WL's functions without students requiring much coding knowledge
◼
Versatility of notebooks
◼
messy data: Human review needed despite Claude Code: e.g. Côte d'Ivoire vs Ivory Coast
8. Model limitations
8. Model limitations
◼
Even the expanded dataset is only 314 matches, ~60 in each test set: each match is worth roughly 1.6-2.2 percentage points of accuracy - not enough for such a low-scoring game
◼
Only international tournament matches (World Cups, Euros, AFCON, Copa America) – results may not generalise to domestic league football with different pacing or game standards
◼
7 of 314 matches couldn't be matched to a FIFA ranking row (older/renamed team names) and therefore were discarded for Model C
◼
First-10-minutes stats are a weak signal by nature, but beating the baseline by roughly 9-11 points is evidence of some signal. Not a usable prediction tool, however.
CITE THIS NOTEBOOK
CITE THIS NOTEBOOK
Predicting football match outcomes from the first 10 minutes: Wolfram language vs Python
by Archer Briggs
Wolfram Community, STAFF PICKS, August 31, 2026
https://community.wolfram.com/groups/-/m/t/3789788
by Archer Briggs
Wolfram Community, STAFF PICKS, August 31, 2026
https://community.wolfram.com/groups/-/m/t/3789788