We Tested Our Models Against Data They've Never Seen. Here's What Happened.
Out-of-sample validation: the gold standard for proving a sports betting model actually works. We ran 5 models against 2025-2026 data they'd never seen. One passed. Here's the full breakdown.
We Tested Our Models Against Data They've Never Seen. Here's What Happened.
July 29, 2026 · 10 min read · Signal Syndicate Research
Key insight: One model passed the hardest test in sports modeling. Two showed edge from line shopping. Two failed — and we're publishing the full results.
What Is Out-of-Sample Validation?
Every sports betting model can look good on its training data. The real question is: does it work on data it's never seen?
Out-of-sample (OOS) validation is the answer. You train a model on historical data (say, 2024), then test it on a completely separate period (2025-2026). If the model's performance holds up, you have genuine signal. If it collapses, you had overfitting — noise that looked like a pattern.
Most sports betting sites never publish OOS results. The ones that do typically use flat -110 odds assumptions that don't reflect real market conditions. We used actual market odds for every single bet.
Why This Matters
The sports betting industry is built on marketing, not transparency. Twitter touts cite 70% win rates on 15-bet samples. Picks sites show backtests on training data. Nobody runs the hard test — and then publishes the results, including the failures.
This is the transparency standard that separates research from promotion.
Our Methodology
Training data: 2024 MLB season
OOS test data: 2025-2026 MLB seasons
Odds: Real market odds from multiple sportsbooks — no flat -110 assumptions
Edge filter: Kelly criterion > 0% at actual odds
Total bets tested: 5 models, thousands of unique plays
Every pick specifies where to bet and at what odds. Every result is tracked with actual settlement data, not theoretical grades.
The Results
| Model | Status | OOS Bets | OOS Win Rate | OOS ROI | Verdict |
|---|---|---|---|---|---|
| **Run Lines** | 🟢 FLAGSHIP | 620 | **60.8%** | **+4.97%** | Genuine predictive signal |
| **Hits Props OVER** | 🟡 ACTIVE (CLV) | 1,391 | 50.0% | +3.85% | Line shopping edge |
| **MLB Totals UNDER** | 🔴 SHADOW | 1,010 | 49.0% | -4.43% | Overfit — retune in progress |
| **Hits Props UNDER** | 🔵 SHADOW | 0 | — | — | Counterfactual only |
| **F5 Innings UNDER** | 🔵 SHADOW | 86 | — | +13.21% | Insufficient data |
Run Lines: The Flagship Finding
OOS result: 60.8% win rate, +4.97% ROI, 620 bets
Run Lines — road underdogs getting +1.5 runs, filtered by Kelly criterion and edge threshold — is the only model that demonstrated genuine predictive power. The 60.8% win rate held across 620 out-of-sample bets at 87 unique odds values ranging from -205 to +184.
This is not a fluke. The model outperformed its in-sample training results (+1.77% ROI) on OOS data, which is the hallmark of a robust signal.
Why it works: The model identifies situations where the market systematically underprices road underdogs in specific run-line contexts. The Kelly criterion filters out low-confidence plays, and the edge threshold ensures only the strongest signals are promoted.
Hits Props OVER: The CLV Edge
OOS result: 50% win rate, +3.85% ROI, 1,391 bets
A 50% win rate is a coin flip. So how does it generate profit?
Closing Line Value (CLV). By consistently beating the market's closing price, the model captures better odds than the final market price. Winning bets are at longer odds, producing net profit even at 50% accuracy.
This is line shopping as a legitimate strategy — not prediction, but execution. The edge comes from identifying props where the opening line offers value relative to where the market settles.
MLB Totals UNDER: The Overfit Finding
OOS result: 49.0% win rate, -4.43% ROI, 1,010 bets
This is the honest result. The model showed +4.52% ROI on training data and -4.43% on OOS. That's a textbook overfit: the model learned noise in the training period that didn't generalize.
Why it failed: The under-9.0 totals edge filter was tuned to 2024 market conditions. When run environments shifted in 2025-2026, the filter's assumptions broke down. The model is now in shadow retune with OOS data as the true test set.
Hits Props UNDER: Counterfactual
OOS result: Insufficient live data
The UNDER version of the hits props model has a strong backtest (+66.44% ROI on 880 training bets) but zero live OOS data. The counterfactual backtest exists — we're accumulating live data in shadow mode.
F5 Innings UNDER: Research
OOS result: 86 bets, +13.21% ROI
86 bets is insufficient for conclusions. The model shows promise but needs more data. Shadow monitoring continues with a promotion threshold of 200+ bets.
What This Means
Run Lines is our flagship model. The 60.8% OOS win rate is the strongest signal we've found. It's the model we're prioritizing for expansion to NFL, NBA, and NHL.
Hits Props OVER is an active CLV play. The profit comes from line shopping, not prediction. Valuable, but different.
MLB Totals UNDER is in shadow retune. The overfit finding is disappointing but honest. We're retuning with OOS data.
Transparency is the standard. We publish wins and losses. We publish the models that work and the ones that don't. This is the only way to build trust in an industry built on hype.
What's Next
- NFL expansion: Run Lines methodology transfer to NFL spreads, OOS validation by Week 1
- MLB Totals retune: Shadow retune with OOS data as the test set
- Continuous validation: Every promoted model will maintain OOS tracking
- CLV pipeline: Expanding CLV tracking to all models, not just Hits Props
Educational intelligence only · Estimates only · Not betting advice · Past results ≠ future performance
Read the full research report with complete methodology, result tables, and retune plans →
View live model performance with 3-column validation grid →
Blog posts are public education. The app has Research, signals, and Ask Signal.
Open the App Read the MethodologyAll figures are estimates. Past analysis is not a guarantee of future results. Not betting advice.