Why We Didn't Promote a Winning-Looking MLB Totals Model (P43)

Why We Didn't Promote a Winning-Looking MLB Totals Model (P43)

> Signal Syndicate Research Library · Model Promotion Review · Educational intelligence · Not betting advice

Phase 43 (P43) is a case study in why a strong-looking backtest is not a promotion. The gate that stopped it is 75 forward bets before promotion. Signal evaluated MLB totals across a broad lane and a narrower summer UNDER sub-lane. The broad universe flattened to break-even. A sub-lane looked exceptional in 2025 but failed 2024 replication. Our multi-AI courtroom issued MONITOR — not promote — with only three forward shadow bets at review time.

This is intentional process. Picks services market hot slices. A sports intelligence platform publishes the validation story.


Executive Summary

Lane 2025 backtest (est.) Verdict
**P43 broad totals** n=502, ROI +0.01% No durable edge — do not promote
**Summer UNDER 7.0–8.5 (Jun–Sep)** n=186, ROI +15.98% Discovery hypothesis — monitor only
**Same sub-lane 2024 as-of** n=171, ROI −1.76% Replication failed
**2026 shadow OOS** n=3, ROI +27.27% Insufficient sample

Court majority (Phase 72): `MONITOR` · `PROMOTE: NO` · `OOS_READ: INSUFFICIENT` · 72 bets to promotion minimum.

Note: One AI panelist (Llama) argued RETIRE on the sub-lane; majority verdict is MONITOR. The lane is not retired — it remains in shadow accumulation. K props (separate lane) was retired; do not conflate the two decisions.


What We Tested

Policy: Frozen P43 (`P43_a0.69_sc0.85_floor_1.05`) — no retuning, no promote-without-gates.

Broad lane: Full P43 bet universe on 2025 validation.

Candidate sub-lane (pre-registered):

Rule Value
Side UNDER only
Total line 7.0 – 8.5 inclusive
Months Jun, Jul, Aug, Sep
Projection path `advisory_projection` / `mlb_model_v3_asof`

Quant metrics were computed in Python. Interpretation and promotion gates ran through the multi-AI courtroom panel (DeepSeek, Qwen, Llama, R1-7b) with founder alignment.


What We Found

Broad lane — expanded universe flattened

When we evaluated the full P43 2025 sample (502 bets), ROI landed at approximately +0.01% — effectively break-even. A model that looks promising on a narrow slice can disappear when the bet universe expands. That is the core P43 lesson.

Sub-lane — strong 2025, weak 2024 replication

The summer UNDER mid-line slice produced stronger 2025 backtest estimates:

2024 as-of comparable slice: n=171, ROI −1.76% — replication failure. Discovery in one year without prior-year support is an overfit warning, not a promotion case.

2026 forward shadow (partial)

At Phase 72 review:

Three bets is not a sample. Positive early OOS does not override insufficient n, missing CLV history, or failed replication.


Why It Matters

1. Winning-looking ≠ validated. A +16% ROI slice can coexist with a flat broad lane and a negative prior year.
2. Process over marketing. Signal withheld promotion despite headline-friendly metrics — that is the product moat.
3. Shadow mode protects users. `do_not_apply=true` until court and sample gates pass. Users never see unapproved lanes as picks.
4. Transparency builds trust. Publishing non-promotion decisions is as important as publishing wins.


Courtroom Verdict

Field Majority
LANE_VERDICT **MONITOR**
PROMOTE **NO**
OOS_READ **INSUFFICIENT**
CREDITS **HOLD**

Panel disagreement: Qwen → MONITOR; Llama → RETIRE (sub-lane); majority rules MONITOR for the summer UNDER shadow lane.

What we did not do: Promote broad P43. Promote the sub-lane. Expose either lane to users as live picks.

Separate decision: K props May–Jul lane was RETIRED after OOS failure — a different lane, different verdict. See `why-we-retired-k-props-may-jul` in this library.


What We Did Next


What To Monitor


Appendix — Sub-Lane Breakdown (Historical Estimates)

All figures below are backtest estimates. Past performance ≠ future results.

Monthly (2025 lane)

Month n Win% ROI
Jun 50 70.0 +33.63%
Jul 40 60.0 +14.54%
Aug 41 56.1 +7.09%
Sep 55 56.4 +7.60%

Line (2025 lane)

Line n Win% ROI
7.0 24 70.8 +35.22%
7.5 60 61.7 +17.72%
8.5 59 61.0 +16.48%
8.0 42 52.4 ~0.0%
8.2 1 100.0 +90.9%

Edge bucket (2025 lane)

Edge n Win% ROI
1.5–2.0 42 73.8 +40.9%
1.0–1.5 38 65.8 +25.59%
2.0+ 106 53.8 +2.65%

Cherry-pick audit: sub-lane filters were fixed before the validation run (pre-registered from postmortem hypothesis). Lane ROI exceeded major excluded toxic buckets (OVER, Apr/May, 9+ totals, thin edge) — profile-based separation, not post-hoc optimization.


Related Assets


Research only · estimates only · not betting advice · past backtest ≠ future performance · shadow lanes are not user-facing picks.


Signal Syndicate Research Footer

Field Value
**Methodology** Frozen P43 policy · holdout + 2024/2025 replication · multi-AI courtroom (Phase 72)
**Data Source** Signal MLB totals research layer · historical odds estimates
**Sample Size** Broad n=502 +0.01% ROI; sub-lane n=186 +15.98% (2025 est.); 2024 n=171 −1.76%; 2026 OOS n=3
**Date Range** 2024–2025 backtest; 2026 forward shadow partial
**Research Date** 2026-06-10
**Validation Notes** MONITOR · PROMOTE: NO · OOS INSUFFICIENT · 72 bets to 75-bet gate
**Related Signal Research** why-we-retired-k-props-may-jul · mlb-totals-validation-transparency

Signal Syndicate Research found these results through governed validation — not picks marketing.