Why a high hit rate proves nothing
A hit rate is the easiest number a forecaster can publish and the easiest one to make look good. It is also the number least likely to tell you whether the forecasts were worth reading.
Four ways to reach a 90% hit rate without being useful
Forecast what was going to happen anyway
The cheapest accuracy comes from calling things with a high base rate. Predict that a major assistant will ship some product change this quarter, or that AI referral volume will keep climbing, and you will be right nearly every time while telling nobody anything they could act on. A record full of near certainties scores beautifully and carries no information. The worth of a call is not whether it landed, it is how far it moved your belief from what you already assumed.
Leave the resolution criteria loose
"Significant change in AI referral traffic" can be graded a hit or a miss after the fact, depending on which way the grader wants it to go. A criterion written at grading time is not a criterion, it is a description of what happened. The test that separates the two is simple: was the exact condition that would count as a hit, on which surface, inside which window, written down and published before the window opened.
Count only the calls you like
A forecast quietly withdrawn, edited, or dropped from the archive never appears in the denominator. Nothing has to be deleted for this to happen: a call reworded a week in, or reissued with a new date, has effectively been ungraded. This is why an archive that is never rewritten matters more than a score that looks good, and why a forecaster who cannot show you the calls they got wrong is showing you a marketing asset.
Score too few calls
Ten graded calls cannot separate a skilled forecaster from a lucky one, because the uncertainty around a hit rate computed from ten results is wide enough to contain both. Early accuracy figures on any young forecasting record, including this one, are not evidence, and treating them as proof is the error the record exists to prevent.
Ask for calibration instead
Meteorology settled this argument decades ago, and the answer is not a hit rate. It is calibration: across the whole record, do the calls issued at 70% happen about seven times in ten. A forecaster can be calibrated and still be useless, if every call sits at the base rate, so calibration is read alongside resolution, which asks whether the probabilities are spread out enough to distinguish one week from another. Useful forecasting needs both.
The standard instrument is the Brier score, proposed by Glenn W. Brier in Monthly Weather Review in 1950. It is the mean squared difference between the stated probability and the outcome recorded as one or zero, so lower is better and zero is perfect. It is a proper scoring rule, which is the property that matters here: it penalises overconfidence and permanent hedging at 50% alike, so a forecaster cannot improve their score by shading numbers toward whatever is safest. A hit rate can be improved that way, which is a large part of why it is the number most often published.
Three questions to ask any forecaster, including this one
- Was the probability published with the call, or attached afterwards. A prediction with no number on it cannot be scored at all.
- Is the resolution criterion specific enough that two people reading it in advance would grade the outcome the same way.
- Does the archive contain calls that failed, at the same size and in the same place as the ones that landed.
Where this leaves our own record
AIOWeather is in BETA and the graded record is short, so it cannot yet support a headline accuracy claim, and one should not be believed from us or from anyone else at this stage. What the accuracy page holds is the individual calls with their dates, stated probabilities, impact, confidence labels and grades, kept whether they landed or not.
This is the limit of the accountability layer, stated plainly. A public archive proves that nothing was removed. It does not prove the calls were hard, and it cannot on its own tell you whether a run of hits was skill. Only a record long enough to check calibration against can do that, and that takes time no amount of publishing shortens.
AI Visibility Forecasting is worth exactly what its record can be checked against. Read the calls, not the percentage. Why a forecast misses covers what the individual failures mean, what a probability means covers the number itself, and the methodology defines how a call is built and graded.