Does another week
make a better forecast?
A frozen preseason model and a weekly updating challenger, compared on the same 784 games from the 2025 season. Inspect the errors, the probabilities and the training behind each prediction.
0.60 points less margin error.
The weekly model’s MAE was 13.64, compared with 14.24 for the preseason model. Winner accuracy was 70.2% versus 65.4%.
This is an experiment, not a live betting record. It uses currently published historical data and does not replace the 2026 production model. Rosters, injuries and market prices are absent.
Weekly minus preseason MAE: -0.60 points. Approximate 95% week-block bootstrap range: -0.88 to -0.30. Resamples 22 UTC weeks; repeated teams across weeks can still be dependent.
Does the weekly update travel?
Each row calibrates on the prior season, freezes that probability mapping, and scores the following season. Keeping transitions separate shows whether a result repeats beyond one schedule.
| Test season | Calibrated on | Games | Preseason MAE | Weekly MAE | Weekly winner % | Weekly fits |
|---|---|---|---|---|---|---|
| 2024 | 2023 | 787 | 14.64 | 13.90 | 66.3% | 22 |
| 2025 | 2024 | 784 | 14.24 | 13.64 | 70.2% | 22 |
Margin MAE is points. Weekly fits are dated snapshots; the table is retrospective evidence, not a guarantee of future or betting edge.
Move forward. Never peek ahead.
Establish the field
Fit the score model on 2022–23. Freeze program membership from that fit for the 2024 calibration replay. The 2025 comparison freezes its field from the 2022–24 fit.
Calibrate probabilities
Generate weekly predictions, then use 787 games to fit the challenger’s probability mapping and 80% margin range. The preseason model has its own calibration.
Replay the next season
Freeze calibration. Each Monday, refit using earlier completed games whose starts precede Sunday 00:00 UTC. Compare with a preseason fit that stays fixed all year.
The 24-hour start buffer reduces overlap with unfinished games; it is not proof of historical data availability. Source corrections are not rolled back. Earlier 2025 results may enter later weekly fits, but never their own forecast.
Where does the difference appear?
Loading the historical comparison…
Account for the missing games.
Both methods exclude the same out-of-field games. The source is not a certified FBS membership list. These counts describe this source edition.
Use the prospective ledger to distinguish forecasts actually registered before games from historical experiments.
Reproduce the comparison.
Download game predictions, all weekly coefficients and training-game IDs, the calibration sample, and the file hash manifest.
Experiment edition: Sep 12, 2026. Dataset edition: Sep 12, 2026.
e15877c196505e6c92dc951c0e27c6a4d93e7b34178f1a9bcec8944e59a51b95
What this experiment cannot establish.
Retrospective experiment designed after the evaluation season. Historical source revisions and actual result-availability times are not reconstructed.
The 2023–24 and 2024–25 rolling transitions use prior-season calibration and independent following-season tests; calibration rows are not independent test performance.
Evaluation-season results enter later weekly fits only after the cutoff buffer, but never their own prediction.
The 24-hour start buffer is conservative scheduling, not proof that a historical result had been reported. Calendar weeks use UTC, not source week numbers.
Teams are frozen from the previous-season fit to keep both methods on the same field. New or unseen teams are excluded; performance on them is unknown.
Both methods use scores, program identity and venue only. No roster, recruiting, injury, advanced efficiency or market features are included.
The approximate interval resamples whole calendar weeks. Repeated teams and schedules can remain dependent across those blocks; it is not proof of future improvement or market advantage.
No production forecasts, model registrations, odds observations or prospective ledger results are changed by this experiment.
Temporal evaluation and calibration references: scikit-learn’s time-series evaluation guidance and probability calibration documentation. Source data: SportsDataverse, labeled CC BY 4.0 by its publisher. Source receipts and download URLs are in the summary. Read the production model notebook →