Forecasts
One testable hypothesis: a meme that has started spreading will cross into reality — a confirmed link to a news trend will appear within 7 days. The criteria were fixed in advance (docs/forecast-contract.md); misses are public and are never deleted.
Hypotheses are written in Russian (the original language); no translation is available yet — the original is shown.
Brier score
0.2638 (0.255–0.2736)
Calibration: predicted probability vs observed rate
Resolved forecasts are split into equal-size groups by predicted probability. For a calibrated model, “predicted” ≈ “observed” in every row; the interval is a 95% Wilson CI (few outcomes → wide, and that is honest).
| Predicted | Observed | Outcomes | 95% CI |
|---|
| 35% | 25% | 689 | 22–28% |
| 50% | 20% | 688 | 17–23% |
Contract #2: inflection → importance
A sharply accelerating trend of local importance will reach national importance (4/5) within 48 hours. Statistics cover the current contract version (v1.1, from 17 Jul 19:51 UTC — the moment the new trigger was actually deployed). Forecasts produced by the previous trigger are excluded from this board: some were voided with a public reason, the rest are excluded as “created by the previous logic”. Clarified 26 Jul: the previous boundary sat 12 hours before the deploy, letting 18 old-trigger forecasts into the board — all with a “hit” outcome (a negative outcome could not occur in that window), inflating the hit rate. The forecasts themselves remain public.
Brier score
0.092 (0.0859–0.0984)
⚠ The outcome-labelling rule has changed: a “hit” now counts only on evidence FROM the forecast window (the previous rule accepted an importance observation at any time, including after the deadline). The Brier and skill above are computed only over the 5758 rows under the new rule; another 1127 resolved forecasts are labelled under the previous rule and are not included — they remain public in the list. The two generations cannot be merged into one series: they answer different questions.
Contract #3: security → patch within 7 days
For a HIGH/CRITICAL vulnerability from the tech issue that is known to OSV but not yet fixed, a fixed version will appear within 7 days of the advisory being published. Resolution comes from api.osv.dev (an external source, not our own data); the sample is biased towards open-source ecosystems covered by OSV. Calibration is preliminary — the base rate will be refined by a backtest over OSV history.
Brier score
0.25 (0.25–0.25)
⏳ openinflection → importanceconfidence 10% · base ratedeadline 19 Sept, 10:22 UTC
⏳ openinflection → importanceconfidence 10% · base ratedeadline 19 Sept, 19:22 UTC
Confidence is a contract constant, not a model: the same number on every forecast of a given type. For “meme → reality” it comes from a backtest; for “inflection → importance” a historical backtest is impossible in principle (the labelling rule cannot be reconstructed retroactively), so the number was chosen deliberately uninformed; for “patch within 7 days” it is interim until a backtest over OSV history. While the stake is constant, skill against the base rate cannot be positive on any data — that is a property of arithmetic, not of the pipeline. A model will be published only once it beats the base on a held-out sample.
Voided forecasts (∅) count as neither hits nor misses, but they mean different things: for “meme → reality” — a data failure (the cluster was hidden by moderation); for “inflection → importance” — the trend disappearing, that is, informative censoring (the rows that vanish are not random); for “patch within 7 days” — a gap in check coverage. Lumping them together as “technical noise” would hide the bias.