Weekly review

Week of 31 August

31 August 20266 September 2026 · 1,749 words

48.4%
called correctly
244
predictions published before kick-off
23
competitions
7
match days

Between 31 August and 6 September we had published a prediction before kick-off for 244 fixtures, across 7 days and 23 competitions. We called 118 of them correctly, 48.4%.

Does our confidence mean anything

Taking every call in the period and sorting it by how confident we were: we said under 40% on 44 fixtures, averaging 38.5%, and were right 43.2% of the time; we said 40 to 50% on 88 fixtures, averaging 45.2%, and were right 48.9% of the time; we said 50 to 60% on 60 fixtures, averaging 54.3%, and were right 43.3% of the time; we said 60 to 70% on 39 fixtures, averaging 64.5%, and were right 53.8% of the time. The band that misses is 50 to 60%: we overstated it by 10.9 points across 60 fixtures. That is the number to hold us to, and it is the one we would most like to see shrink.

The outcome we almost never call

Of 244 fixtures, 73 ended level — 29.9% of them. We called a draw in none of them. Not rarely: never, and that is worth explaining rather than apologising for. The three probabilities on each card come from simulating the two sides' scoring rates independently and counting how often each outcome appears. That method has a built-in ceiling on the draw: across this period the highest draw probability we published was 27.5%, and the nearest a draw ever came to leading its fixture was 9.7 points short. For the draw to top a card, two teams would have to be so exactly matched that neither win is likelier than the tie — which is rarer than a tie itself. So the headline call cannot name the outcome that happens a quarter of the time, and roughly a quarter of our misses are lost before kick-off. Independence is also the part we would change first. Real football scores are not independent at the low end: 0-0 and 1-1 happen more often than two separate rates predict, which is why the standard correction in the literature exists and why our draw numbers are likely a little low across the board. Until that is fixed, the probability we print for the draw is the figure worth reading on a tight fixture — not the call above it.

The fixtures we called close

On 44 fixtures no outcome reached 40% — the model divided them close to three ways and had, by its own account, little to say. 15 of those ended level, 34.1%, against 29.9% across the period as a whole. That difference matters more than it looks. It means the probabilities do rise when a draw is coming, even though the call printed above them never names one. The information is in the forecast; it is the headline that throws it away. If you read one number on a fixture we have split three ways, read the draw percentage rather than the team we listed first. We were right on 43.2% of them, which is roughly what guessing between two of three outcomes would give you, and we would rather publish a fixture like that flagged as undecidable than manufacture a lean.

Home, away, and which we get right

We backed the home side 184 times and the visiting side 60, so 75% of our calls went to the home team. They landed 44.0% of the time. The away calls landed 61.7%. That gap runs the wrong way for the shape of our output: the calls we make least often are the ones that come off most. The reading we would defend is not that away teams are undervalued in general, but that we only name the visitor when the case is strong, and are readier to give the home side the benefit of a close one. The fix is not to back more away teams; it is to be less willing to call a fixture that is genuinely tight.

When we disagreed with the bookmakers

On 145 fixtures we had both our own number and a bookmaker price for the outcome we called. We were at least 10 points above the market on 25 of them — a real disagreement, several times larger than the margin built into the odds — and those calls came in 40.0% of the time. On the 60 fixtures where our number sat within 5 points of the price, we were right 50.0%. Read that in the least flattering direction, because it is the correct one: when we part company with the market, the market is usually the one that was right. A disagreement is where a model claims to know better, and over this period ours did not earn the claim. It is a reason to treat our boldest departures as the weakest part of the output rather than the most interesting, and we would rather print that here than let it sit unmeasured.

Against the bookmakers' favourite

The benchmark worth measuring against is not chance, it is the shortest price — anyone can read the odds for nothing. On the 145 fixtures where we had all three prices, backing the bookmakers' favourite every time would have been right 72 times, 49.7%. We were right 69 times, 47.6%. We named the same outcome as the market on 126 of them, 87%. So over this period the model came out 2.1 points BELOW simply following the favourite. That is the least comfortable number on this page and it is the one we would want a reader to see first. It does not make the probabilities worthless — a calibrated number tells you something a binary pick never does, and the market has no obligation to explain itself — but anyone here for the pick alone should know that the free version of it is currently doing better.

The goals we expected

Underneath the win probabilities sit two expected-goals figures, and they are the part of the model that is easiest to check and hardest to fudge. Across 244 fixtures we expected 1.53 goals for the home side and 1.19 for the away side, a total of 2.72. What actually happened was 1.55 and 1.39, a total of 2.93. The away figure is the one that drifts: we were under by 0.20 of a goal per match. It is small in absolute terms and it is systematic, which is the more interesting property — a bias that holds across hundreds of fixtures is correctable in a way that a wide spread is not. Rounding each expectation to a whole number reproduced the exact scoreline 19 times, 7.8%. We mention that as a curiosity rather than a claim: nobody should be predicting exact scorelines, and a model that appeared good at it would be a model that had been overfitted.

The margin flag, tested

Each card carries a margin flag, meant to warn that a fixture could finish lopsidedly. Over 244 fixtures, 44 ended with a gap of three goals or more — 18.0%. The flag was raised on 233 of them, 95% of everything we published. That is the finding, and it is not a good one: a warning attached to 95% of fixtures is not a warning, it is a decoration. When raised it was followed by a three-goal margin 17.2% of the time. We are leaving the flag visible and saying plainly that it is not currently carrying information. It will either be recalibrated or removed, and whichever happens will be reported here.

The calls we backed hardest

13 calls were published at 70% or above. 9 of them came in, 69.2%, against an average stated confidence of 77.0%. The most expensive of them was Telstar to beat SC Cambuur at 79.7%, which finished 2–2 — a draw, which is the outcome we structurally cannot call.

The call we would point to

The call we are most pleased with is not the most confident one, it is the one furthest from the price. PSV Eindhoven to beat Ajax Amsterdam in Eredivisie: we published 64.7% where the odds implied about 41.7%, a gap of 23.0 points, and it finished 1–3. 10 calls in this period were that far above the market and came in. We list one rather than all of them because a page of the ones that worked, with the ones that did not left out, is how a track record gets quietly turned into advertising — the figure for how that whole group did is in the section above, and it is not flattering.

Where we were most wrong

Telstar against SC Cambuur, in Eredivisie, finished 2–2. We had published Telstar to beat SC Cambuur at 79.7%, the largest confidence we placed on anything that did not happen in this period. It is worth noting what our five worst misses have in common: 3 of them ended level. The model's most confident errors are not upsets, they are draws. It stays on the record. A site that reports only its hits is not reporting a record, it is reporting a selection.

Against last week

Against the week before: 215 settled fixtures then at 45.1%, 244 now at 48.4%, a move of +3.2 points. More useful than the headline: in the 50 to 70% range, where most of our calls sit, we were 5.2 points below our stated confidence last week and 10.8 points below it this one. That gap widened, which matters more than the headline percentage moving the other way.

Where we published, and how it went

We published across 23 competitions. The five busiest were League Two (24), Championship (24), League One (23), Liga Profesional Argentina (13), Segunda Division (12). Among the 3 with at least 20 settled fixtures, we read League Two best — 50.0% across 24 matches — and League One worst, at 43.5% across 23.

How to read all this

One week of 244 fixtures settles very little on its own. A percentage point either way on a sample this size is noise, and we would rather say so than dress it up. What these reviews are for is the shape that emerges when several of them sit side by side: whether a stated 65% keeps landing near 65%, whether the gap between our number and the market's keeps costing us, whether a competition we read badly this week is still being read badly in three. Those are answerable, and they are answerable in public, which is the only reason to publish the bad figures alongside the good ones.

Every figure here comes from predictions published before the match day — never from a re-run after the results were known. Results come from public fixture feeds. Every period we have reviewed → How the model is measured →

Week of 31 August — football prediction review | Aleavici