Why a high accuracy score can be worthless
A forecast record that reads 90% correct sounds like proof of skill. It can be produced without forecasting anything at all, and telling the two apart takes about a minute. This page is the one that says where our own scoreboard fails.
The score you can get without making a single call
Rare events break percent correct. If the thing being forecast happens in one week out of twenty, then a forecaster who says "nothing will change" every single week finishes the year at 95% and has never once been useful to anybody. The score is not measuring judgement. It is measuring how rare the event is.
That is uncomfortably relevant here, because the events on this site are rare by construction. A platform materially changing how it surfaces and cites sources is not a weekly occurrence. Most weeks genuinely are quiet, and a long run of quiet weeks correctly called quiet is close to free.
Meteorology named this problem in 1884
John Park Finley ran an experimental tornado forecasting programme for the US Army Signal Service and published his verification results in the American Meteorological Journal in 1884, reporting his forecasts as roughly 96.6% correct. Within months, critics pointed out that a forecaster who simply said "no tornado" every time, everywhere, would have scored about 98.2% on the same data. The person who tried lost to the person who declined to try.
The episode is documented in Allan Murphy's paper "The Finley Affair: A Signal Event in the History of Forecast Verification" (Weather and Forecasting, 1996), and it is the reason verification moved to skill scores: measures that compare a forecast against a naive baseline rather than against the calendar. A skill score of zero means you did exactly as well as saying nothing at all. That is the number carrying information. Percent correct is not.
Four ways an accuracy record flatters itself
Base rate. The largest of the four, and the one above. Ask what a "nothing changes" forecaster would have scored over the same period. If the honest answer is close to the published figure, the record has told you nothing about the method.
Selection. A forecaster who publishes only the calls it is nearly certain of builds a beautiful record and adds nothing to anyone's decision. High accuracy and low usefulness are frequently the same behaviour viewed from two angles.
Ungradable wording. "Continued volatility expected" cannot be wrong. If no state of the world would falsify a statement, it is not a forecast, and it has no business inside a denominator that gets reported as accuracy.
Retroactive movement. Rewording a call after the outcome is known, quietly widening a window, or letting a miss fall out of the archive. This is the only one of the four that is dishonest rather than merely weak, and it is also the easiest to hide.
Four questions to ask of any forecast record
They apply to this one, and they apply to any vendor showing you a hit rate.
What would doing nothing have scored? If that is not stated, work it out from the number of weeks in which anything actually changed.
How many calls are in the denominator, and was anything excluded? A percentage with no count behind it is a decoration. Twelve calls at 92% and two hundred calls at 74% are not comparable, and the second is the more informative record.
Can you still read the failures? If the misses are not there to read, the hits are unverifiable, because you cannot see what was left out.
Could the calls have gone the other way? A record made entirely of calls above 90% is a record of things that were already visible. See what a probability actually means for why the useful range sits lower.
Where this leaves our own scoreboard
Stated plainly, because a page about misleading records that exempts itself would be the thing it is warning about: the accuracy record here is young, its denominator is small, and it is accumulating during a period in which most weeks have been quiet. A small number of graded calls made across a calm stretch cannot yet separate skill from season. It will become able to, and it is not able to yet.
So the record is published with its counts visible rather than as a single headline percentage, and misses stay in the archive permanently. Neither of those makes a young record mature. They make it possible for you to see that it is young, which is the part a flattering scoreboard removes.
What we hold, and what we do not
Every graded call keeps its original wording, its window and its probability. Nothing is reworded after the outcome, nothing is withdrawn for scoring badly, and a quiet week is published as a quiet week rather than dressed as a correct prediction of calm.
What we do not do is lead with a percent correct figure and let it stand as evidence of judgement. It is not, on its own, evidence of anything. This page sits under AI Visibility Forecasting; why a forecast misses covers the individual failed call, and when to ignore a forecast covers what to do with a record you have decided to trust.