Skip to content
MarketClueLearn

Why Backtests Lie

Intermediate12 min readLesson 8 of 8

4 steps · one page

In short

A backtest applies a trading rule to historical data and reports what it would have produced. It is a legitimate and necessary tool, and almost every backtest a member of the public ever sees is worthless as evidence.

Reading-order note and boundary. Numbered #8 and drafted last, closing the pillar, so it can draw on every article before it. Boundary: #7 covers distortions already present in historical data; this article covers what goes wrong when someone tests a strategy against that data. Backtesting as a quantitative technique — how to construct one properly — belongs to Pillar 25. This article is about why the results a reader is shown should be discounted heavily.

Those statements are compatible, and the reason they are compatible is the whole content of this article: the problem is rarely the arithmetic, which is usually correct. The problem is the selection process that determined which backtest you were shown.

The central problem: enough tries guarantee a good result

Search a large enough space of rules and you will find one that performed spectacularly, whether or not anything real is happening. This is not a subtle statistical caveat — it is the dominant effect, and it is large enough to swamp everything else in the article. The mechanism is simple. Any measurement taken from a finite sample carries error. Try one rule and you get one draw from that error. Try a thousand and take the best, and you have selected the most extreme draw from a thousand — which will look impressive even when every rule tested is worthless. The arithmetic is stark. A performance ratio estimated over twenty years of data carries a standard error of roughly 0.22. Take the best of ten completely worthless strategies and its apparent ratio is about 0.30; the best of a hundred, about 0.53; the best of a thousand, about 0.70; the best of ten thousand, about 0.84. For scale: the canonical figure for the market itself in this pillar is 0.31. So a researcher who tries a thousand rules containing no genuine edge whatsoever should expect to produce a headline result more than twice as good as the market — from noise alone, with honest arithmetic, on real data. The same effect in the language of significance testing: run a thousand tests at a one-in-twenty threshold and about fifty will pass by chance; tighten to one-in-a-hundred and about ten still will. The decisive question about any backtest is therefore not "what did it return?" but "how many variations were tried before this one was shown to me?" — and that number is almost never disclosed, is often not recorded, and is sometimes not even known to the person who ran the search.

Four more ways the number inflates

Overfitting. Every adjustable parameter — a threshold, a lookback window, a rebalancing frequency, an exclusion rule — is another opportunity to fit the noise in the sample rather than the signal. A rule with enough parameters can describe any history perfectly and predict nothing, and the tell is fragility: if performance collapses when a parameter moves slightly, the parameter was fitted to noise. An out-of-sample test is the standard defence and a weaker one than it appears, because if the out-of-sample period is examined repeatedly and the rule revised in response, it has quietly become in-sample.

Omitted frictions. Backtests routinely understate what trading costs. Bid-ask spreads, market impact for size, borrowing costs for short positions, and the simple fact that the tested price may not have been transactable all reduce the result — and the damage scales with turnover, which is precisely what elaborate strategies have. At 100% annual turnover and a 0.20% round-trip cost the drag is 0.20 points a year; at 300% turnover, 0.60 points; at 500% turnover with 0.25% costs, 1.25 points. Over twenty years on $10,000, that last figure removes 20.6% of the outcome — and it is a cost the backtest may simply not have modelled.

Inherited data problems. Everything #7 establishes applies here too: a backtest run on today's index constituents, or on restated financials, or on a database that dropped its failures, produces a flattered result no matter how carefully the rule was specified. A perfectly designed test on contaminated data returns a contaminated answer.

Regime dependence and publication selection. Every test covers a period with particular conditions — a rate environment, an inflation regime, a set of crises — and a rule tuned to those conditions may be describing the period rather than the market. And finally: you only ever see the backtests that worked. Failed tests are not published, not marketed, and rarely even retained, so the population of visible backtests is selected on exactly the quantity that makes them unreliable — the same structural distortion #7 identified in fund records, operating one level up.

Worked example

Worked example

Worked example (illustrative; fictional). A researcher builds a rule for selecting shares. The search. She tries 1,000 variations — different signals, thresholds, holding periods, and universe filters — and keeps the best. Its apparent performance ratio over twenty years of data is 0.70, against 0.31 for the market. It looks like a discovery. It is exactly what a thousand worthless rules would have produced by chance. The refinement. She improves it: excluding one sector and shifting a lookback window lifts the ratio further. Both changes were chosen because they improved the result on this data, which means both are fitted to this sample's noise. The costs. The rule turns the portfolio over 500% a year. At a 0.25% round-trip cost that is 1.25 points annually — over twenty years on $10,000, 20.6% of the final outcome, none of it in the backtest. The data. The universe was "companies in the index today," so the test never saw the failures. The result. Every calculation is arithmetically correct, no dishonesty occurred at any point, and the strategy has no demonstrated edge whatsoever. Published as a chart with a rising line, it is indistinguishable from a genuine finding — and the information needed to tell the difference, the number of variations tried, appears nowhere on the chart. (All figures illustrative and independently verified; the standard error is 1 ÷ √20 = 0.224, the expected best-of-N draw uses the standard extreme-value approximation √(2 ln N) − (ln ln N + ln 4π) ÷ (2√(2 ln N)), and the 20.6% is the shortfall of a 9.0% return dragged to 7.75% over twenty years.)

Frequently asked

8 questions

Are backtests useless?

As a research tool, no — they're necessary, and a strategy that fails its backtest has learned something real. As evidence presented to a reader, close to useless, because you can't see the selection process that produced the one you were shown.

What's the single most important question to ask about a backtest?

How many variations were tried before this one was shown to me? That number determines almost everything about how much the result means — and it's almost never disclosed, often not recorded, and sometimes unknown even to the person who ran the search.

How much can pure chance produce?

More than most people expect. With twenty years of data, the best of a thousand entirely worthless strategies should show a performance ratio around 0.70 — against roughly 0.31 for the market itself on this pillar's canonical figures. That's honest arithmetic on real data producing a result twice as good as the market from nothing.

What is overfitting?

Fitting a rule to the noise in a sample rather than to any real pattern. Every adjustable parameter is another chance to do it. The tell is fragility: if the result collapses when a threshold or window shifts slightly, the parameter was fitted to noise.

Doesn't out-of-sample testing solve this?

It helps and it's weaker than it looks. If the out-of-sample period is examined repeatedly and the rule adjusted in response, it has quietly become in-sample — and that's the normal way research proceeds rather than an abuse.

Why do trading costs matter so much here?

Because damage scales with turnover, and elaborate strategies trade a lot. At 500% annual turnover with 0.25% round-trip costs, the drag is 1.25 points a year — over twenty years, about a fifth of the outcome. Backtests frequently omit spreads, market impact, and borrowing costs entirely.

If a backtest used clean data and one rule, is it reliable?

Considerably more informative — though still limited by the period it covers, since a rule can describe a particular rate or inflation regime rather than the market. The problem was never a well-specified single test; it's that the tests reaching an audience have been selected from many.

Why does MarketClue not publish backtests?

Because a backtested figure carries no information about the search that produced it, and a reader can't recover that number. Showing one would present something unfalsifiable as evidence. MarketClue shows what happened to real assets over stated periods, labelled as history.

References

Educational and informational only — not investment advice, a recommendation, or an offer to buy or sell any security. Investing involves risk, including the possible loss of principal. Worked examples use fictional companies and figures.