Does another week
make a better forecast?
A frozen preseason model and a weekly updating challenger, compared on the same 5,734 games from 2025–26. Inspect the errors, the probabilities and the training behind each prediction.
0.66 points less margin error.
The weekly model’s MAE was 9.61, compared with 10.26 for the preseason model. Winner accuracy was 70.0% versus 67.7%.
This is an experiment, not a live betting record. It uses currently published historical data and does not replace the 2026–27 preseason model. Rosters, injuries and market prices are absent.
Weekly minus preseason MAE: -0.66 points. Approximate 95% week-block bootstrap range: -0.80 to -0.50. Resamples 23 UTC weeks; repeated teams across weeks can still be dependent.
Does the update travel?
Each row calibrates on the prior season, freezes that mapping, and scores the following season. The 2024, 2025 and 2026 tests stay separate so a strong year cannot hide a weak transition.
| Test season | Calibrated on | Games | Preseason MAE | Weekly MAE | Weekly winner % | Weekly fits |
|---|---|---|---|---|---|---|
| 2023–24 | 2022–23 | 5,695 | 10.22 | 9.40 | 70.6% | 23 |
| 2024–25 | 2023–24 | 5,701 | 10.10 | 9.41 | 71.1% | 23 |
| 2025–26 | 2024–25 | 5,734 | 10.26 | 9.61 | 70.0% | 23 |
Margin MAE is points. Weekly fits are Monday snapshots; they are evidence of temporal replay, not a guarantee of future edge.
The 2023–24 row replays cached 2022 schedule and team-box releases as its prior-season training layer; it does not change production D1 data or current forecasts.
Transition index: all transitions ↗ · 2023–24 evidence ↗ · 2024–25 evidence ↗ · 2025–26 evidence ↗
Does continuity explain the next season?
This separate ridge challenger combines prior net efficiency with exact-ID roster continuity, represented prior minutes and listed-player counts. It is evaluated chronologically and stays outside the production forecast until more dated transitions are available.
| Test season | Training seasons | Teams | MAE | Baseline MAE | Improvement |
|---|---|---|---|---|---|
| 2024–25 | 2023–24 | 286 | 5.61 | 6.48 | 0.87 pts |
| 2025–26 | 2023–24, 2024–25 | 289 | 5.87 | 6.55 | 0.68 pts |
Historical transitions use the NCAA roster release and have no publisher Box BPM, so their scores are not directly comparable with the current ESPN-derived production challenger. 345 teams have current roster features; the 2026–27 scenario is a research prompt and does not change win probabilities, uncertainty or ledger registrations.
Roster releases are source snapshots without a verified pre-season publication clock. The challenger is research-only and does not replace the primary forecast or enter the prospective ledger. Only two historical roster transitions are available for fitting; the held-out evaluation is one season and is not a guarantee of future performance. Roster listings do not establish eligibility, availability, transfer reason, injury status or depth-chart role. Publisher Box BPM is unavailable for some source IDs; rows without prior BPM coverage are withheld from the challenger rather than imputed. The scenario changes the primary margin by the learned team-strength delta but does not recalibrate win probability or uncertainty.
Move forward. Never peek ahead.
Establish the field
Fit 2023–24 efficiency and tempo. Freeze program membership before the next season; ten games in the latest fitting year are required.
Calibrate probabilities
Generate weekly predictions, then use 5,701 games to fit the challenger’s probability mapping and 80% margin range. The preseason model has its own calibration.
Replay the next season
Freeze calibration. Each Monday, refit using earlier completed games whose starts precede Sunday 00:00 UTC. Compare with a preseason fit that stays fixed all year.
The 24-hour start buffer reduces overlap with unfinished games; it is not proof of historical data availability. Source corrections are not rolled back. Earlier 2025–26 results may enter later weekly fits, but never their own forecast.
Where does the difference appear?
Loading the historical comparison…
Account for the missing games.
Both methods exclude the same out-of-field games. The source is not a certified Division I membership list. These counts describe this source edition.
Use the prospective ledger to distinguish forecasts actually registered before games from historical experiments.
Reproduce the comparison.
Download game predictions, all weekly coefficients and training-game IDs, the calibration sample, and the file hash manifest.
Experiment edition: Sep 12, 2026. Dataset edition: Sep 12, 2026.
87187136f46b007234115ef7e60d4f4f7a2cf3682d817faec3fb72dbb7d39a4d
What this experiment cannot establish.
Retrospective replay using current source releases; historical revisions and availability timestamps are not reconstructed.
Weekly fits include only completed records with starts before Monday 00:00 UTC minus 24 hours; exact historical final-publication times are unavailable.
The 2023 calibration transition uses the cached 2022 schedule/team-box releases only for replay; it is retained outside the production D1 warehouse and is not a current forecast input.
2023–25 rolling predictions calibrate the challenger; those calibration results are not independent test performance.
2025–26 games enter later weekly fits only after the cutoff buffer. No game enters its own prediction or any earlier week's fit.
The preseason team's field is frozen before each season. New programs outside it are excluded from both methods.
No roster, availability, injury, recruiting or bookmaker inputs. This experiment does not replace live preseason forecasts or enter the prospective ledger.
The week-block bootstrap describes sampling variation within this one season. It is not a guarantee across future seasons or protection against shared-team dependence between weeks.
Fixed penalties, yearly weights and update cadence; no parameter search was performed. This is a new exploratory comparison on a season already used for the published baseline evaluation.
Temporal evaluation and calibration references: scikit-learn’s time-series evaluation guidance and probability calibration documentation. Source data: SportsDataverse, labeled CC BY 4.0 by its publisher. Source receipts and download URLs are in the summary. Read the production model notebook →