Skip to content
Research
11 min read

Proving a backtest is not cheating

Look-ahead bias has no visual signature, and the version that matters inflates a result by less than two percentage points. There is a five-line test that catches it anyway, and a way to check the test itself works.

Universe
1,500 liquid US names, point-in-time
Test period
2024–2026

Look-ahead bias does not look like anything. That is the whole problem with it. A backtest that quietly used tomorrow's information produces the same kind of equity curve as one that did not. It rises, it has drawdowns, it has a plausible number of trades. There is no visual signature, and the line of code that causes it usually reads as sensible.

The usual defences are weaker than they feel. Reading the code finds the bug you already suspect and misses the one you do not. A high Sharpe is a hint at best, because plenty of honest strategies score well and plenty of contaminated ones do not. And "the framework prevents it" is a claim about the framework, not about what you built on top of it.

There is a direct test. This article is that test, the strategy I ran it on, and what happened when I deliberately broke the same strategy to check the test actually works.

What look-ahead bias actually is

A backtest simulates decisions in order. On each day it decides using what was knowable then, and the simulation moves forward. Look-ahead bias is any place where information from later leaks into an earlier decision.

The obvious version is easy to avoid: computing a signal from today's closing price and then buying at today's open. Most backtesters make that hard.

The dangerous version does not touch prices at all. It hides in the questions you answer before the simulation starts:

  • Which instruments am I allowed to trade? If that list was drawn using data from the end of the period, every decision in the simulation is conditioned on the future.
  • Which names count as liquid enough? Same problem, one level down.
  • Which companies still exist? A price file assembled today contains the survivors. Trading it back in 2020 means never holding anything that went to zero.

These are not signal bugs. They are setup decisions, made once, in a different part of the code from the strategy, often months earlier. Reviewing the strategy will never find them, because the strategy is not where they are.

The strategy I tested it on

The subject needs to be a real strategy or the test proves nothing about real strategies. Everything below runs the same one.

Cross-sectional momentum. Rank a universe of stocks by how much each has risen recently, hold the strongest, re-rank periodically. The bet is that relative strength persists for a few months.

On each rebalance day, in order:

  1. Take the eligible universe. The 1,500 most liquid US names, reconstructed as of the decision date rather than as of today. Companies that later delisted are present until the day they stop trading.
  2. Score each name by its trailing six-month return, ignoring the most recent month. That gap is the skip, and it exists because names that have just spiked tend to give some of it back, making the last month the least reliable part of the signal.
  3. Hold the top 20, weighted by inverse volatility, so a name that moves twice as much gets half the money.
  4. Sell with hysteresis. A holding is not sold when it drops out of the top 20, but when it falls out of the top 40. Without that buffer a name oscillating around the boundary is bought and sold every month and pays the spread each time.
  5. Rebalance monthly, filling at the next day's open, charging 0.25% per fill.

The probe window runs from 2024-01-02 to 2026-07-31, with the midpoint at 2025-06-30.

None of that matters much to the test. What matters is step 1, because that is where the leak will go.

How the universe gets chosen, and how it leaks

"The 1,500 most liquid US names" is not a fixed list. Liquidity has to be measured, and measuring requires a window.

The honest version fixes that window at the start of the test. Eligibility is decided using only data available then, and the same 1,500 names are used no matter how much data is loaded afterwards.

The broken version lets the window follow whatever data happens to be loaded. If you load prices through 2026, liquidity is ranked through 2026, and the book trades the names that turned out to be liquid later.

Written plainly, that is: pick the names that turned out to be heavily traded, then go back and trade them in 2024. It is one keyword argument. It reads as completely reasonable in code. It is the most common look-ahead in cross-sectional equity work.

The test

Pick a date in the middle of the backtest. Call it the midpoint.

  1. Load prices only up to the midpoint, and run the strategy from its start to the midpoint.
  2. Load all the prices, and run the same strategy from the same start to the end of the data.
  3. Compare the two equity curves over the days they share.

The first run is physically incapable of using anything after the midpoint, because that data is not in the process. The second has all of it. If any decision before the midpoint touched the future, the two disagree somewhere in the overlap.

In outline, and this is the whole thing:

MID = "2025-06-30"

# 1. a world that ends at the midpoint
prices_a = load_prices(start=START, end=MID)
curve_a  = backtest(strategy, prices_a, start=START, end=MID)

# 2. a world that knows everything
prices_b = load_prices(start=START, end=TODAY)
curve_b  = backtest(strategy, prices_b, start=START, end=TODAY)

# 3. the days they share must agree
shared = sorted(set(curve_a) & set(curve_b))
diffs  = [d for d in shared if abs(curve_a[d] - curve_b[d]) > TOL]

assert not diffs, f"first divergence at {diffs[0]}"

The assertion is the test. You do not need to know what the bug is, or to suspect there is one, which is the point: the failure modes that survive review are exactly the ones nobody thought to look for.

Running it on the honest version

The two runs share 374 trading days. On 4 of them the portfolio values are not bit identical.

Before that sounds like a failure, here is the size of those four disagreements: one millionth of a dollar each, on a portfolio worth about $36,598. Three parts in a billion. That is floating-point arithmetic accumulating in a different order, not information moving backwards through time. Every other day matches exactly, and on the last shared day both runs report the same value to the cent: $36,598.

The universes are identical too, differing by 0 names out of 1500.

So the strategy passes. Its past does not depend on its future.

A test that cannot fail proves nothing

That result is worth very little on its own. If the probe would have said "pass" no matter what, it has told me nothing about the strategy and quite a lot about my willingness to believe good news.

So I broke the same strategy in the way described above, changed nothing else, and ran the identical probe.

It is caught on 2024-01-02, the first day of the test, and it never recovers. On the last shared day the two runs of the same backtest over the same dates report $37,008 and $46,667. The same strategy, the same period, $9,659 apart, differing only in how much of the future was sitting on disk.

Frozen, data to midpoint1.48×Frozen, all data1.48×Leaky, data to midpoint1.49×Leaky, all data1.89×
1×1.2×1.4×1.6×1.8×20242025
Fig. 1The same 374 days, run twice each. Only three lines are visible because the two frozen-selection runs are exactly on top of each other. The two leaky runs are not.
Selection frozen0.00%Selection leaks26.10%
0.0%10.0%20.0%20242025
Fig. 2The same thing as a single number per day: how much each backtest's own past moves when the future is added back. Zero means the two runs agree.

By the midpoint the leaky backtest's own past is worth 26.1% more when the future is available, having peaked at 28.62%.

Which names it swapped in

The mechanism is easiest to see in the universe. With selection leaking, the two runs differ by 220 tickers, roughly one name in seven.

They are not random. Names that appear only when the future is visible include AAOI, ACHR, ADT, ALM. These are companies whose trading volume exploded in 2025 and 2026. In January 2024 they were not among the 1,500 most liquid US stocks, and a book running then could not have been holding them.

The names they displaced are unremarkable in the other direction: ACAD, ADNT, AGL, AL, ordinary mid-caps that were liquid in 2024 and drifted out of the top 1,500 later.

So the leak is not a subtle numerical artefact. It is a book that has quietly swapped a seventh of its opportunity set for the stocks that were about to become interesting.

The part that should worry you

Now look at what all that did to the headline numbers.

TrainHoldout
CAGRMax DDMARTradesCAGRMax DDMARTrades
Frozenselection fixed at the start29.9%-35.3%0.8524844.1%-35.3%1.25410
Leakyselection follows the data30.9%-36.7%0.8422045.8%-38.9%1.18354

Run to the end with all the data, the leaky version reports 45.78% against the honest 44.13%. One and a half percentage points.

Nobody reviewing that flinches. It is not a suspiciously perfect equity curve or an impossible Sharpe. Its drawdown is slightly worse, its MAR slightly lower. It looks exactly like a slightly different idea, which is what makes it dangerous. A bias that inflates a backtest by 200% announces itself. This one does not, and no amount of staring at the final table separates it from a real improvement.

The probe separates them in one run, without being told what to look for.

There is a shorter way to say what the leaky version is doing: its past changes. Run that backtest in June 2025 and the first eighteen months earn 30.89%. Run the identical code today, unchanged, and the same eighteen months earn 26.1% more. History should not move. When it does, the backtest is reading something it should not have.

What the probe cannot catch

It is one test, not an audit, and it detects exactly one thing: whether decisions before the midpoint depend on data after it. Three failures it will pass happily.

  1. Bias baked into the data. If the price file already dropped delisted companies, both runs are equally wrong and agree perfectly. The probe checks for time travel, not for a bad dataset. Survivorship has to be answered separately, by going and looking for the corpses.
  2. Look-ahead inside the truncated window. A signal that uses the closing price of the bar it trades on is contaminated, but both runs are contaminated identically and agree. Moving the midpoint catches different ranges; testing the fill timing directly is better.
  3. Everything that is not look-ahead. Overfitting, unrealistic fills, costs that are too low, a universe you could never have traded at size. This test says nothing about any of them.

A pass means "no future data crossed this boundary." It does not mean the backtest is right. Those are different claims and they are worth keeping apart.

Running it on your own work

Nothing here depends on a particular backtester. You need two things: the ability to limit the data your engine loads, and a per-date equity series out of it.

Three practical notes from doing it.

Pick the midpoint where the strategy is active. A midpoint in a stretch where the book is flat tests nothing, because there are no decisions to contaminate. Halfway through, with positions on both sides, is a reasonable default.

Run the two loads in separate processes if the data is large. Two full price sets will not coexist in memory, and the failure mode is the process being killed rather than the test reporting anything useful.

Compare with a tolerance, then look at what fails it. Bit-equality is the wrong bar. Differences at the tenth decimal place are floating-point accumulation; a difference in the second decimal place on day one is a bug. The gap between those two is large enough that you will not have to agonise over which you are looking at.

I run this before publishing any result now. It is cheap, it is mechanical, and it is the difference between believing a backtest and having checked one.

Written by

Koray GocmenFounder

Builds and runs the systematic strategies behind this research.

Full source

The code, the data and the parameter files behind this article.

All research