Every verdict is frozen at T-24h and scored against what actually happened. Nothing is revised after the fact. The whole record is below, including the misses, because a forecast you can't check is marketing.
When we say 80%, does it happen 80% of the time? Predicted versus actual, bucketed by score.
Verdicts come from an independent weather model scoring wind, gusts, cloud cover, precipitation
probability and convective energy (CAPE) at the exact pad, at the launch hour. The verdict is
frozen 24 hours before the window opens and never touched again. After the window passes, the
launch is resolved as on time or delayed.
CAUTION is scored as neither right nor wrong. It asserts uncertainty rather than a
direction, so counting it either way would flatter the number. Brier score measures probabilistic accuracy. Lower is better, 0.25 is what you get by guessing.
These assessments are indicative and independent. They are not official launch commit
criteria, and they are not a substitute for the range's call.