GLOBE SPORT LIVE
Olympics FIFA ICC Tournaments Tennis Athletics Basketball Racing World Rankings
AboutContactPrivacy Policy

Prediction models are graded on the wrong thing entirely

A forecast that says sixty per cent and is wrong has not failed, and the way sport judges its models keeps missing that.

Prediction models are graded on the wrong thing entirely

A model gives the away side a thirty per cent chance. The away side wins. By Monday morning somebody has written that the model was wrong, and by the end of the season the model has a reputation, and the reputation is built entirely from the outcomes that felt surprising.

That is not how you assess a probabilistic forecast. A thirty per cent call is supposed to come true roughly three times in ten. If every thirty per cent shot lost, the model would be badly broken, just in the other direction, and nobody would notice because it would look brilliant every weekend.

The measure that matters is calibration. Take every match where a model said thirty per cent, count how many of them happened, and see whether the answer is close to thirty. Do it for each band. A well-calibrated model is boring to read and enormously valuable, because its numbers mean what they say, and you can act on them without translation.

Accuracy, the hit rate everybody quotes, actively rewards cowardice. A model that predicts the favourite every time in a sport where favourites win two thirds of the time posts a two-thirds accuracy and has told you nothing you did not know. It has no information in it. It is a coin weighted by the league table.

Sport is worse at this than most fields, and I think the reason is cultural. Sport is a results business, and a results business hates the sentence it was a good decision that did not work. The same reasoning failure runs through selection, through in-game choices, through sackings. A team that creates chances and loses is described as having played badly.

There is a practical consequence for anybody building these things. If your audience grades you on individual outcomes, you will drift towards confident predictions, because confident predictions look better when they land and no worse when they miss. The incentive points away from honesty.

Publish the calibration curve alongside the forecast. Nobody will read it. It is still the only thing on the page that tells you whether to believe the rest.

📊 Community Poll

Should helmets be mandatory in all amateur sports?