Monthly review

August 2026

1 August 202631 August 2026 · 1,982 words

48.8%
called correctly
693
predictions published before kick-off
25
competitions
30
match days

Over August 2026 we had published a prediction before kick-off for 693 fixtures, across 30 match days and 25 competitions. We called 338 of them correctly, 48.8%.

Does our confidence mean anything

Taking every call in the period and sorting it by how confident we were: we said under 40% on 124 fixtures, averaging 38.2%, and were right 42.7% of the time; we said 40 to 50% on 264 fixtures, averaging 44.4%, and were right 42.0% of the time; we said 50 to 60% on 161 fixtures, averaging 54.3%, and were right 47.2% of the time; we said 60 to 70% on 87 fixtures, averaging 64.5%, and were right 63.2% of the time; we said 70% and above on 57 fixtures, averaging 76.1%, and were right 75.4% of the time. The band that misses is 50 to 60%: we overstated it by 7.1 points across 161 fixtures. That is the number to hold us to, and it is the one we would most like to see shrink. At the other end, 70% and above came out within 0.6 points of what we claimed, on 57 fixtures — there, the figure on the card means what it says.

The outcome we almost never call

Of 693 fixtures, 181 ended level — 26.1% of them. We called a draw in none of them. Not rarely: never, and that is worth explaining rather than apologising for. The three probabilities on each card come from simulating the two sides' scoring rates independently and counting how often each outcome appears. That method has a built-in ceiling on the draw: across this period the highest draw probability we published was 29.3%, and the nearest a draw ever came to leading its fixture was 7.0 points short. For the draw to top a card, two teams would have to be so exactly matched that neither win is likelier than the tie — which is rarer than a tie itself. So the headline call cannot name the outcome that happens a quarter of the time, and roughly a quarter of our misses are lost before kick-off. Independence is also the part we would change first. Real football scores are not independent at the low end: 0-0 and 1-1 happen more often than two separate rates predict, which is why the standard correction in the literature exists and why our draw numbers are likely a little low across the board. Until that is fixed, the probability we print for the draw is the figure worth reading on a tight fixture — not the call above it.

The fixtures we called close

On 124 fixtures no outcome reached 40% — the model divided them close to three ways and had, by its own account, little to say. 40 of those ended level, 32.3%, against 26.1% across the period as a whole. That difference matters more than it looks. It means the probabilities do rise when a draw is coming, even though the call printed above them never names one. The information is in the forecast; it is the headline that throws it away. If you read one number on a fixture we have split three ways, read the draw percentage rather than the team we listed first. We were right on 42.7% of them, which is roughly what guessing between two of three outcomes would give you, and we would rather publish a fixture like that flagged as undecidable than manufacture a lean.

Home, away, and which we get right

We backed the home side 517 times and the visiting side 176, so 75% of our calls went to the home team. They landed 46.6% of the time. The away calls landed 55.1%. That gap runs the wrong way for the shape of our output: the calls we make least often are the ones that come off most. The reading we would defend is not that away teams are undervalued in general, but that we only name the visitor when the case is strong, and are readier to give the home side the benefit of a close one. The fix is not to back more away teams; it is to be less willing to call a fixture that is genuinely tight.

When we disagreed with the bookmakers

On 622 fixtures we had both our own number and a bookmaker price for the outcome we called. We were at least 10 points above the market on 83 of them — a real disagreement, several times larger than the margin built into the odds — and those calls came in 30.1% of the time. On the 242 fixtures where our number sat within 5 points of the price, we were right 50.8%. Read that in the least flattering direction, because it is the correct one: when we part company with the market, the market is usually the one that was right. A disagreement is where a model claims to know better, and over this period ours did not earn the claim. It is a reason to treat our boldest departures as the weakest part of the output rather than the most interesting, and we would rather print that here than let it sit unmeasured.

Against the bookmakers' favourite

The benchmark worth measuring against is not chance, it is the shortest price — anyone can read the odds for nothing. On the 622 fixtures where we had all three prices, backing the bookmakers' favourite every time would have been right 297 times, 47.7%. We were right 301 times, 48.4%. We named the same outcome as the market on 518 of them, 83%. The two are within a couple of points, which is the honest summary: on the headline pick, we are matching the market rather than beating it.

The goals we expected

Underneath the win probabilities sit two expected-goals figures, and they are the part of the model that is easiest to check and hardest to fudge. Across 693 fixtures we expected 1.52 goals for the home side and 1.21 for the away side, a total of 2.73. What actually happened was 1.53 and 1.39, a total of 2.91. The away figure is the one that drifts: we were under by 0.18 of a goal per match. It is small in absolute terms and it is systematic, which is the more interesting property — a bias that holds across hundreds of fixtures is correctable in a way that a wide spread is not. Rounding each expectation to a whole number reproduced the exact scoreline 69 times, 10.0%. We mention that as a curiosity rather than a claim: nobody should be predicting exact scorelines, and a model that appeared good at it would be a model that had been overfitted.

The margin flag, tested

Each card carries a margin flag, meant to warn that a fixture could finish lopsidedly. Over 693 fixtures, 120 ended with a gap of three goals or more — 17.3%. The flag was raised on 662 of them, 96% of everything we published. That is the finding, and it is not a good one: a warning attached to 96% of fixtures is not a warning, it is a decoration. When raised it was followed by a three-goal margin 15.9% of the time, and on the 31 fixtures where it was NOT raised the margin was that wide 48.4% of the time — the wrong way round, on a sample too small to be certain but too stark to ignore. We are leaving the flag visible and saying plainly that it is not currently carrying information. It will either be recalibrated or removed, and whichever happens will be reported here.

The calls we backed hardest

57 calls were published at 70% or above. 43 of them came in, 75.4%, against an average stated confidence of 76.1%. The most expensive of them was Feyenoord Rotterdam to beat ADO Den Haag at 87.1%, which finished 2–2 — a draw, which is the outcome we structurally cannot call.

The call we would point to

The call we are most pleased with is not the most confident one, it is the one furthest from the price. Aris to beat Kalamata in Greek Super League: we published 91.3% where the odds implied about 45.5%, a gap of 45.8 points, and it finished 2–3. 25 calls in this period were that far above the market and came in. We list one rather than all of them because a page of the ones that worked, with the ones that did not left out, is how a track record gets quietly turned into advertising — the figure for how that whole group did is in the section above, and it is not flattering.

Where we were most wrong

Feyenoord Rotterdam against ADO Den Haag, in Eredivisie, finished 2–2. We had published Feyenoord Rotterdam to beat ADO Den Haag at 87.1%, the largest confidence we placed on anything that did not happen in this period. It is worth noting what our five worst misses have in common: 4 of them ended level. The model's most confident errors are not upsets, they are draws. It stays on the record. A site that reports only its hits is not reporting a record, it is reporting a selection.

The best and worst week

Split the period by week and the spread is wide: our best week was the week of 10 August, 74 of 142 (52.1%); the worst was the week of 24 August, 97 of 215 (45.1%). Across 4 comparable weeks the range is 45.1% to 52.1%.

Week by week

In order, the weeks ran: week of 3 August, 51 of 101 (50.5%); week of 10 August, 74 of 142 (52.1%); week of 17 August, 97 of 200 (48.5%); week of 24 August, 97 of 215 (45.1%). The spread between the best and worst week is 7.0 points, with no direction to it. That is what a stable model looks like at this volume: the weekly figure moves around, the monthly one much less.

Against last month

Against the month before: 50 settled fixtures then at 50.0%, 693 now at 48.8%, a move of -1.2 points. Volume changed more than accuracy did — up 1286% — so the two figures are not measuring quite the same thing: a month with more fixtures is usually a month with more of the competitions we read least well. The headline difference itself is small enough that we would not call it a change at all.

Where we published, and how it went

We published across 25 competitions. The five busiest were EFL Cup (60), Liga Profesional Argentina (56), MLS (38), Brasileirao (37), Championship (36). Among the 18 with at least 20 settled fixtures, we read DFB-Pokal best — 68.2% across 22 matches — and MLS worst, at 34.2% across 38. We would not yet call that a real difference between leagues. At these sample sizes the gap between the best and worst competition is about what you would get by splitting one league in half at random, and the leagues we cover least are exactly the ones where the rating model has the least history to work from. It is worth watching across several months; it is not worth acting on after one.

How to read all this

One month of 693 fixtures settles very little on its own. A percentage point either way on a sample this size is noise, and we would rather say so than dress it up. What these reviews are for is the shape that emerges when several of them sit side by side: whether a stated 65% keeps landing near 65%, whether the gap between our number and the market's keeps costing us, whether a competition we read badly this month is still being read badly in three. Those are answerable, and they are answerable in public, which is the only reason to publish the bad figures alongside the good ones.

Every figure here comes from predictions published before the match day — never from a re-run after the results were known. Results come from public fixture feeds. Every period we have reviewed → How the model is measured →

August 2026 — football prediction review | Aleavici