The Record · Public audit
Most sports prediction sites only publish the picks they got right. We publish every single prediction, locked before first pitch, and every outcome. The good ones. The bad ones. All of it.
Results update nightly · Yesterday's games post by midnight ET
Best Bet win rate by direction
slate baseline 57.5%
193-130 · pred 60.0% · cal gap +0.3pp
60%
+2.3pp vs baseline
1-2 · pred 49.5% · cal gap +16.2pp
33%
-9.2pp vs baseline
◆ Closing Line Value
COMING ONLINECLV tracking is coming online.
The most defensible proof a betting model is +EV is closing-line value - the difference between our pick probability at lock time and where the market settled by first pitch. We're instrumenting captures from FanDuel + DraftKings at T-10min from first pitch. The framework is live; the first captures land on the next slate that the book-line scraper covers. We publish positive AND negative CLV - same rule as the rest of the audit.
Capture: closing-line snapshots from FanDuel + DraftKings at T-10min from first pitch for each of our Best Bets. We capture both Yes and No sides so we can de-vig.
De-vigging: Quoted prob / (sum-of-both-sides quoted prob). Standard market-neutral method - strips the book's margin from the quoted price so we're comparing apples to apples.
◆ Track Record · 21,073 predictions · 88 days
2026-04-09 → 2026-07-28
◆ All picks
57.5%
12,117 / 21,073 picks hit
Each dot is one out of every 100 predictions. Filled = the pick hit. Hollow = it didn't. Every prediction is locked before first pitch and matched to the box score after the game resolves - wins and losses both. Nothing is removed or edited.
Average predicted probability: 53.85% vs actual hit rate: 57.50%. The model is conservative - predictions came in below actual outcomes.
Calibration error (5-bin): 3.84 pp - sample-weighted mean of |bin predicted − bin actual| across the 5 quintile buckets. 0 pp = perfectly calibrated; under 1 pp is excellent.
Brier score (Hit): 0.2489 - mean of (predicted − outcome)². Lower is better. 0 = oracle; 0.25 = coin flip; under 0.24 means the probabilities carry information.
Per-stat Brier: HR 0.1024 (n=21,073), K 0.2142 (n=12,235). HR baseline is ~0.029 (3% league HR rate squared); K baseline ~0.18. Lower beats baseline.
Sample: every prediction we ever made (locked pre-game, never edited) where the game has resolved. n = 21,073 across 88 days. Best Bets = 326 curated picks.
Why a dot grid instead of a calibration plot: research across NYT, FT, 538, Polymarket, Whoop, ESPN BET converged on a single pattern - one big number, one iconic shape, one comparison. A 100-dot grid renders the percentage literally; you can count to verify. The calibration plot lives below in the Edge Audit panel and on the dedicated /calibration page for sharp readers.
◆ Calibration proof
Across 21,073 graded picks, locked before first pitch and matched to the official box score. A 56% pick is supposed to lose 44% of the time. Here is the receipt.
◆ Edge Audit · 88 days
2026-04-09 → 2026-07-28
+2.01 ppBest Bet lift over slate baseline
Across 326 curated picks, our Best Bet hit rate is 59.5% vs a slate-wide baseline of 57.5% (n = 21073). Statistical confidence: <90% (z = 0.74).
Top-5 by calibrated hit probability with CI-width tiebreak. Persisted server-side at lock time.
Hits / N
194 / 326Rate
59.5%Δ vs slate
+2.01 ppPure model output: top-5 by calibrated hitProb. No additional curation.
Hits / N
278 / 440Rate
63.2%Δ vs slate
+5.69 pp0-100 composite advantage score. Noisier than hitProb - sorting by it loses signal.
Hits / N
283 / 440Rate
64.3%Δ vs slate
+6.82 ppNaive baseline - qualified hitters ranked by 2026 OBP. Tests whether our model beats a one-stat dumb sort.
Hits / N
291 / 440Rate
66.1%Δ vs slate
+8.64 ppEvery prediction made - n = 21073. The actuarial baseline.
Hits / N
12116 / 21073Rate
57.5%Δ vs slate
—Sanity check - 200 random samples/day, mulberry32 seeded.
Hits / N
50334 / 88000Rate
57.2%Δ vs slate
-0.30 ppSample: every prediction we ever made (locked pre-game, never edited) where the game has resolved. n = 21073 predictions across 88 days of MLB action. Best Bets = 326 curated picks (~5/day).
Lift definition: our hit rate minus the slate baseline hit rate. The slate baseline is the rate at which any player who appeared in our predictions got at least one hit that day. It IS already filtered to confirmed lineups, so it's already a high bar.
Random baseline: 200 random 5-player samples per day with a seeded RNG (mulberry32, seed=42 - same number every render). Matches the slate rate to within 0.2pp, which is the sanity check that we're not silently sampling the head.
Significance: simple two-sample z-test on Best Bet hit rate vs slate, assuming binomial. z = 0.74, p < 0.01 when |z| > 2.58. The interval is approximate - the true test requires accounting for non-independence within games and across days. Take the number as directional, not publication-grade.
What this audit does NOT include yet: closing-line value (CLV), which would require pulling book lines daily. It is in the queue. Until it ships, lift over the slate (and the season-OBP baseline shown above) are the cleanest defensible edge metrics.
Hit Accuracy
↑ 29.8pp vs MLB avg
21,073 predictions · 0 removed
vs league baseline
Cumulative hit rate · every day
Apr 28 to Jul 26
Each point is the running accuracy through that date, settling toward the true rate as the sample grows.
Tier Performance
Hits · n = 21073Top-5 picks per day, curated by calibrated hit probability with tight CI.
95% CI [54.4, 65.0]· Wilson interval
95% CI [56.9, 64.3]
95% CI [56.6, 62.5]
95% CI [58.2, 62.5]
95% CI [56.2, 57.6]
Lift = actual rate − predicted rate. Positive lift means we under-predicted that tier; negative means we over-predicted. Calibrated picks should show near-zero lift across tiers.
7/14/30-day windows within 3pp - consistent, not streaky
Best Verified Calls
Highest-confidence hits - locked before first pitch, published win or lose.
Every Day Since Launch
Click any day · see every call we made
Nailed It
highest confidence published calls, last 30 daysMissed Calls
same 30-day window, wrong - we publish everythingAccuracy Over Time
Running average drawing in real time
57.5%
Full Prediction Log
n = 21,073Every record locked before first pitch · nothing removed
◆ Pro Feature
Full History Unlocked on Pro
7-day free trial · $40/mo after · one hit pays the month
Full prediction history, all three metrics, player-level trends, and a daily AI brief before games start.
Get Pro →7-day free trial · cancel anytime
◆ Best Bets · top 5/day
59.8%
195 / 326 picks hit
Edge measures pct beat over picking that side blind. Cal gap is (predicted - actual): positive = model overconfident, negative = model underclaims. Tested live since 2026-05-18 - fades reserved 1 slot per slate even when over edges rank higher.
CLV per pick: our locked probability minus the de-vigged book probability, expressed in percentage points. Positive = we priced the pick more aggressively than the market settled. Negative = the market thought we were too aggressive.
Publication threshold: 500 resolved captures. Below that the average CLV is dominated by variance - a run of cold matchups could swing the number 0.5pp+ without changing the underlying signal. We publish progress but not the average until the sample clears the threshold.
What this audit does NOT replace: the EdgeAuditPanel above (lift over slate baseline). CLV is the market test; the slate-baseline lift is the internal test. Both are real, both are published unedited, neither is a substitute for the other.
◆ Pitcher track record
29-day window · 2026-06-29 to 2026-07-27
618 resolved · 14 pending
52.9%
95% CI [49-57%] · n=618
34.6%
95% CI [31-38%] · n=618
-0.19
actual - proj
-0.26
actual - proj
| Date | Starts | K Over | QS Rate | K Bias |
|---|---|---|---|---|
| 2026-07-26 | 29 | 55% (16/29) | 21% (6/29) | -0.15 |
| 2026-07-25 | 28 | 61% (17/28) | 61% (17/28) | +0.31 |
| 2026-07-24 | 27 | 44% (12/27) | 33% (9/27) | -0.27 |
| 2026-07-23 | 10 | 30% (3/10) | 60% (6/10) | -0.05 |
| 2026-07-22 | 33 | 64% (21/33) | 36% (12/33) | +0.07 |
| 2026-07-21 | 25 | 60% (15/25) | 36% (9/25) | -0.53 |
| 2026-07-20 | 29 | 52% (15/29) | 34% (10/29) | -0.44 |
| 2026-07-19 | 30 | 50% (15/30) | 40% (12/30) | -0.21 |
| 2026-07-18 | 26 | 77% (20/26) | 35% (9/26) | +0.73 |
| 2026-07-17 | 25 | 56% (14/25) | 44% (11/25) | +0.49 |
| 2026-07-16 | 2 | 100% (2/2) | 50% (1/2) | +1.32 |
| 2026-07-14 | 2 | 0% (0/2) | 0% (0/2) | -4.51 |
| 2026-07-12 | 29 | 52% (15/29) | 21% (6/29) | -0.38 |
| 2026-07-11 | 29 | 52% (15/29) | 34% (10/29) | -0.50 |
Per-stat Brier (lower = better)
Phase 2c calibration (per-stat Platt scaling) lights up after 30+ days of resolved outcomes