Methodology
Disclosing how the model works and how large the sample is matters more than any performance number on this site. Here is the whole shape of it.
How a slate gets processed
The model compares each day's matchups against a library of historical situational patterns — pitching profile, recent run differential, platoon splits, travel and rest, park factors. A forecast is published only when enough patterns match with sufficient historical strength and no danger signal fires. Everything else is logged as informational and is not published.
The threshold is fixed before the slate is seen. It is not adjusted to produce a certain number of forecasts per day, which is why some days produce three and some produce none.
Grades
Everything that clears threshold is published with a grade describing how densely the patterns matched. The grade describes the model’s certainty. It is not an instruction about money and should not be read as one — and as the record shows, a higher grade has not reliably meant a better result.
A++Densest pattern match
The densest pattern matches. Rarest by design — and over the current window, not the best-performing grade. See the record.
A+Strong pattern match
A strong match with substantial historical support behind it.
AClear pattern match
A clear match that meets the bar without the density of the grades above it.
Why games get rejected
- No pattern match
- Not enough historical situations resemble this matchup closely enough to say anything.
- Danger signal
- A pattern matched, but a countervailing factor fired — a profile the model has been wrong about before.
- Below threshold
- Patterns matched, but the combined historical strength did not reach the fixed cutoff.
Sample size and limitations
This record covers 99 published forecasts across 34 days. That is still a small sample and should be read as one.
- Ninety-nine forecasts over five weeks is a short window. Runs inside it have ranged from 12-0 across four days to 3-11 across seven — neither streak is the model's true rate, and neither is the headline number, exactly.
- The grade ranking has not been stable: the densest-pattern grade was the worst tier over one stretch of this window and the best over another. At these sample sizes the ordering between grades is mostly noise — read every rate with the count next to it.
- Games the model logged as informational rather than publishing ran 173-151 over the same window, close to a coin flip. The published forecasts are the product; the rest is the model declining to say anything useful.
- The model forecasts game outcomes. It does not price markets and does not account for line movement or availability.
- Late scratches, weather, and lineup changes after publication are not re-modeled.
Past performance does not predict future results. Model accuracy varies by sample and by period, and any published record reflects a limited window.