What a forecast accuracy score cannot prove
One percentage is the easiest number for a forecaster to publish and the easiest one to inflate. A hit rate can rise while every forecast behind it gets less useful, and nothing about the number itself would show you that.
This is the page where the method is weakest
Publishing a graded record is the strongest thing AI Visibility Forecasting does, and the summary number on top of that record is the weakest. It is weak in a specific and fixable way, so it is worth stating plainly rather than leaving for someone else to notice.
Three moves raise a hit rate without improving a single forecast. None of them is fraud. All of them are ordinary, and two of them look like good practice.
Forecast things that have already happened
A confirmed change, written up as a forecast, is a hit at the moment it is published. Fill a record with those and the score climbs while the useful output falls to nothing: the reader who wanted to know what to prepare for is being told what they already read last week. This is why every call here carries a confidence label, and why a confirmed change and a reasonable forecast are never presented as the same act. A score that mixes them is measuring how much hindsight was published.
Forecast a market that mostly does not move
In weather, saying tomorrow will resemble today is right most days in most places, and it requires no skill at all. Digital discovery behaves the same way. Most weeks nothing material changes, which is why the most common recommendation on this site is no action, and why that is intentional rather than an omission.
It follows that a forecaster who predicts stability every week will score well and be useless in the single week that matters. An accuracy number only means something next to what the naive assumption would have scored over the same period. Beating zero is not skill. Beating "assume nothing changes" is, and it is a much harder bar.
Choose the denominator
Withholding a forecast when a source has gone stale is correct behaviour, and it also removes the hardest calls from the record. Grading a murky outcome as inconclusive is honest, and it also lets that call fall out of the total. Neither is wrong on its own. Together they mean a percentage describes the calls a forecaster chose to grade, not the questions a reader had.
Hit rate and calibration are different tests, and only one survives averaging
A call published at 30% is supposed to fail roughly seven times in ten. Fold those into an overall hit rate and correct behaviour reads as incompetence, and the rational response is to stop publishing anything below 80%, which deletes exactly the forecasts that carried information.
Calibration asks a different question, band by band: of everything issued at 60%, did about six in ten occur? That is checkable, it is the property worth having, and one blended percentage cannot express it. It is also slow, because each band needs many calls before it says anything at all. Which leads to the uncomfortable part.
Where this leaves our own number
AIOWeather is in BETA and the graded record is short. A short record cannot separate a good forecaster from a lucky one: across a few dozen calls the plausible range around any hit rate is wide enough to contain both, and no amount of presentation narrows it. Reporting a confident accuracy figure now would be the exact error this page describes.
So the accuracy page is an archive first and a score second. What it is worth today is that every call is timestamped, keeps the probability and confidence label it was issued with, and is still there after grading. What it will be worth in a year is calibration.
Six questions to ask any forecaster, including this one
- Was the call published before its window opened, with a timestamp you did not have to take on trust?
- Does the archive still contain the withdrawn, ungraded and inconclusive calls, or only the ones that were scored?
- Is a probability attached to each call, and are results broken out by probability band instead of averaged into one figure?
- What would "nothing changes" have scored over the same period?
- Are grades ever revised after publication, and has a forecast ever been deleted?
- Are there weeks with no forecast at all, or does something appear on every schedule slot?
A record that answers those honestly cannot be compressed into a percentage, which is the point of asking. Why a forecast misses covers what an individual wrong call means, the methodology defines how a call is built and graded, and both sit under AI Visibility Forecasting.