Same Accuracy, Half the Reward: Watching a Betting Market Get Efficient
Same Accuracy, Half the Reward: Watching a Betting Market Get Efficient
The Signal Syndicate August 2026 research brief. Two audits, one corrected ledger, and a live demonstration of something textbooks describe in the abstract: a betting market deciding our edge was worth less than we thought — while our model kept proving it could hit the target.
First, the correction
We begin with the uncomfortable one, because a research journal that only publishes good news is a marketing blog.
On August 31 we discovered that the program which scores our run-line results had been grading them backwards for six weeks. Our flagship lane wagers only on road underdogs at +1.5 — but the grader inferred each bet's side from the sign of the odds and read the winning margin team-relative. For a dog lane, that inverts results in both directions. Sunday's card, published as 0-3 and minus three units, was actually 3-0 and plus two. In total, 86 of 133 graded run-line outcomes were wrong on the public ledger.¹
The fix was not a silent database edit. Every corrected row carries provenance, the original 738-row ledger was snapshotted before any write, and an independent auditor — a program that re-grades the ledger straight from box scores, deliberately not reusing the production code — re-checked 379 rows across 14 days and found zero mismatches. The corrected record: across the 133 run-line signals graded since the lane went live on July 22, 86 outcomes had been wrong, and the repaired book reads 85-48, plus 28.04 units — a 63.9% win rate that lands on the exact number our pre-launch backtest predicted for this lane's price bucket. (Lifetime across all 140 graded signals, including the pre-launch window: 90-49-1, plus 31.58 units.) The model performed as validated. The scoreboard lied, and we published the confession the same day.
The golden finding: the model held. The price didn't.
Our hits-props OVER lane wagers that a batter records at least one hit. It has been running all August, 181 graded signals, every one a 0.5-line.
Here is the split that matters. Through August 20, the lane won 58.8% of its bets and returned +4.17% ROI. From August 21 through month's end, it won 58.2% — statistically the same accuracy, 79 fresh bets later — and returned −7.77%.²
The win rate barely moved. What collapsed was the payout. Early in the month, an average win paid +0.771 units; the median win paid +0.909, meaning the typical entry price was around −110: risk $1.10 to win $1.00, needing a 52.4% hit rate to break even, which 58.8% clears. By late August the average win paid +0.584 and the median +0.500 — prices around −200. Risk $2 to win $1. Break-even: 66.7%. Same accuracy, half the reward, and suddenly the wrong side of break-even.
Nobody broke our model. The market repriced the trade. Batter-hit over/unders at half-a-hit are a public, well-fed market, and the books can see the same thing everyone sees: these players get hits at a remarkably steady rate. When a market correctly prices an outcome, the payout falls until the edge is gone. We watched it happen on our own ledger in real time — which is, we think, the most interesting result of our summer. The lesson generalizes: an accurate forecast is not a profitable bet unless the price leaves room between accuracy and payout.
The corroborating evidence is closing line value: on the 22 August entries with a stamped closing price,³ the median entry ran −4.1% against the close, and 77% were priced worse than the final market number. When you consistently get a worse price than the close, you are not beating the market to information — the market is telling you your price is the stale one.
The 68.8% that is not edge
Our hits UNDER lane — shadow-graded, never on the board⁴ — produced the month's best illusion: a 68.8% win rate across 366 signals… and a −1.17% return. How? The lane lays short-priced favorites: the average entry sits near −236, which demands 69.7% just to break even. We won 68.8%. The 95% confidence interval on that win rate spans 63.8 to 73.4 — it straddles break-even. This is a coin wearing a highlighter's costume, and a clean illustration of why win rate misleads.
The band map is where the actual information lives.⁵ Sorted by the price we were offered: at −250 to −299 (190 bets), ROI ran −4.5%. At −200 to −249 (121 bets), +0.4% — exactly the break-even it should be. And at −150 to −199 — just 20 bets — the lane won 70.0% against a 63.5% break-even: +9.9% ROI. One row of the table is not a discovery; n=20 is a hypothesis. But it converts a vague question ("is UNDER good?") into a testable one ("do we ever get offered those prices, and from whom?"). That is September's job: instrument the price capture, and let the band map fill in.
One more portfolio truth: the OVER and UNDER daily results correlate at +0.135. Two models, one bet — both are wagers that someone gets a hit. Diversification claims should be read off correlation matrices, not model counts.
What this data cannot say
We hold 181 OVER bets and 366 UNDER bets. To statistically resolve a genuine two-point-percentage edge at these prices requires roughly 4,000 graded bets; a one-point edge, about 16,500. Every finding above is directional, with sample sizes printed next to it on purpose. Nothing here justifies increasing a stake anywhere — our sizing guide is a quarter-Kelly on the lower bound of a confidence interval, which at these numbers means, honestly, near zero. Anyone publishing certainty from three weeks of betting data is selling something. We are not.
September's tests
The NFL pipeline ran its full shadow season in August: 2,400+ player-projection rows, forward odds captured, a grading path that runs end to end. The promotion gate is unchanged — 100 regular-season graded signals before any lane is discussed for the board, and the ledger will say whatever it says. Week 1 is a pre-registered experiment, not a prediction.
What we continue to do
None of this report is a finish line; it is a checkpoint on machinery that keeps running whether or not anyone reads about it.
- CLV, start to finish. We timestamp the price at the moment our engine identifies a signal, and since late August we capture the market again before first pitch, so "close" is measured honestly, not approximated. The 22 stamped entries in this report are the beginning of that archive — coverage grows every day, and the gap between the two prices is the process metric this report keeps quoting.
- Data collection for quarterly out-of-sample testing. Every slate — games we bet, games we skipped, prices that moved and prices that held — goes into an archive that our models never train on mid-quarter. When the quarter closes, that archive becomes the next honest test, and the answer might be "the edge is gone," which is a result, not a failure.
- Daily projection monitoring. Each player's nightly projection is logged against the actual stat line — hits, and the context around them — so drift in accuracy or bias (like the +0.29→+0.37 over-projection above) is caught in days, not seasons.
- Context surveillance. Weather shifts, lineup scratches, umpire assignments, betting-flow sentiment: we log what changed between our morning read and first pitch, and we test whether those changes moved our projections or moved the market's line toward or away from us. Some of it will matter; most of it won't. We are measuring which, not assuming — the same discipline that caught our own scoreboard this month.
Method note. All figures come from our append-only settled ledger through August 30, re-extracted at publication time; corrections are stamped with provenance and published, not patched quietly. A small number of closing-line values carry extreme tails under audit and are excluded from headline claims. Educational research only — not betting advice, and never a recommendation. The ledger is public; the scoreboard updates daily.
Sources & Definitions
- ¹ The correction — 86 of 133 run-line outcomes graded 7/22–8/30 were re-scored on 2026-09-01 after box-score ground-truthing; every corrected row carries a provenance stamp and the pre-correction ledger is snapshotted in full. The independent re-grader re-checks the last 14 days before every morning report ships.
- ² Period split — 102 signals (8/12–8/20) vs 79 (8/21–8/30); the 0.6-point win-rate difference is well inside sampling error for these n. Payout figures are per-unit win payouts, which encode the entry price directly.
- ³ CLV (closing line value) — the price we were offered at signal identification versus the closing price before first pitch; a process metric, not profit. n=22 stamped August entries; extreme-tail values are under data-quality review and excluded from headlines.
- ⁴ Shadow lanes — tracked and graded publicly on the scoreboard but excluded from the live portfolio headline; promotion requires backtest, holdout, and forward-shadow gates.
- ⁵ Band map — bets grouped by the entry odds captured on the published card; break-even win% at a given price is |odds| ÷ (|odds| + 100) — for −236 that is 70.2%. The 69.7% headline is the average of each bet's own break-even across the lane, which differs slightly from the break-even at the average price.
- ⁶ Sample-size math — two-sided α=0.05, power 0.80, at p≈0.69; the number of graded bets required to distinguish a real edge of the stated size from noise.
Continue reading: yesterday's daily analysis · We Published a Wrong Number for Six Weeks · What closing line value actually means · Why win rate misleads
Data: Signal Platform · Edited by: Signal Desk