A forecast that gave the favorite a strong chance and watched it lose is not automatically wrong. Assessing prediction quality requires a sample large enough for probability to mean something.
One result carries almost no information
If a forecast says a team wins seven times in ten, then losing is an expected outcome roughly three times in ten. Observing that loss once tells you nothing about the model.
This is the central misunderstanding in most arguments about predictions. The critic treats a probability as a claim about this match, while the forecaster meant it as a claim about a class of matches.
Only a run of results reveals whether the stated confidence was honest, because only a run lets you count how often the unlikely thing actually happened.
Calibration is the property that matters
A well-built forecast is calibrated: across all the matches where it claimed a seven-in-ten chance, close to seven in ten should have gone that way.
Calibration is separate from being bold. A forecaster who says every match is a coin flip is perfectly calibrated and completely useless, because the predictions carry no discrimination.
Good forecasting needs both properties together, and a season is roughly the minimum span over which either one can be measured with any confidence.
Accuracy counts are a misleading scoreboard
Counting how many results a pundit called correctly rewards picking favorites in every fixture, which any spreadsheet can do without insight.
It also penalizes a forecaster who correctly identified that an upset was unusually likely, since the upset still probably did not happen.
Scoring rules used seriously by forecasters instead reward stating a probability close to what eventually occurred, so confident wrong calls cost more than cautious ones.
Seasons change underneath the model
The complication is that a soccer season is not a stable experiment. Squads change in the winter window, coaches are dismissed, and a team in spring may not resemble itself in autumn.
That means the sample is never fully clean, and a forecaster reviewing a season has to separate genuine model error from real changes in the teams being modeled.
Most honest reviews therefore split the season into segments and check whether errors cluster around particular events, rather than reporting one number for the year.
What a reader can do with this
The useful test of any prediction source is whether it publishes its record in full, including the calls that failed, and whether it states probabilities rather than picks.
A source that only remembers its correct calls is not making forecasts. It is telling stories after the fact, which reads similarly and carries none of the same information.