Back to blog

Seven ways a backtest lies to you

22 Aug 20265 min readPocketX Research Desk

Here is the uncomfortable arithmetic of backtesting.

Nobody deploys a strategy whose historical results looked poor. So every strategy that has ever failed in live trading had a backtest that looked acceptable.

A good backtest is therefore a necessary condition and almost no evidence at all. What separates a useful result from a misleading one is knowing the specific ways the number gets inflated.

What a backtest actually proves covers the conditional nature of the result. This article is about the seven concrete mechanisms by which the number ends up wrong.

1. You tested until something worked

The most common failure, and the hardest to see from the inside.

You test a rule. Results are unremarkable. You adjust a parameter. Better. You change the timeframe. Better again. You add a filter. Now it looks good.

Nothing dishonest happened at any step. But you have run twenty variants and selected the best, and with twenty attempts on any dataset, something will look good by chance. What you have found is the luckiest variant, and luck does not persist.

The tell: a strategy whose performance collapses when a parameter shifts slightly. A 20-period average that works and a 22-period one that does not means you found a quirk, not a mechanism.

The defence: decide the rule before testing it. Write the parameters down first. If you must explore, treat every additional test as evidence against the eventual result — and know that you have.

2. The exit was chosen after the fact

You test an entry rule with a simple exit. Results are mediocre. You try a wider stop. Better. A target instead. Better again.

You have now fitted the exit to a known past. In live trading the future does not come with the exit already optimised for it.

The exit governs most of the outcome, which makes it the most damaging thing to fit. Write it at the same time as the entry, before any results exist. That is the discipline described in turning a hunch into a testable rule, and this is why it matters.

3. Too few occurrences to mean anything

A strategy with eight historical trades, six of them profitable, tells you approximately nothing. Eight is well inside the range where coin-flipping produces the same picture.

The instinct on seeing a small sample is to relax the rule so it fires more often — but that produces a different strategy, not more evidence for the original one.

The honest response to too few occurrences is to conclude you cannot judge it, and either test a longer period or accept the uncertainty explicitly. Not to loosen the rule until the sample looks respectable.

4. The period was one regime wearing a disguise

A trend-following rule tested across a sustained bull market will look excellent. It has not been shown to work; it has been shown to work in a trend, which is what it was built to do.

The question that matters is what happens in the conditions it was not designed for. A trend rule in a range-bound market. A mean-reversion rule during a sustained directional move.

Look at the distribution across the period, not the total. A strategy that made all its money in one favourable stretch and drifted the rest of the time is a bet on that stretch recurring — which may be reasonable, but should be a conscious decision rather than an accident hidden in a summary number.

5. Costs were understated

Every trade costs brokerage, taxes and spread. A backtest that models these thinly overstates results, and the error scales with trade count.

A strategy making two trades a month can absorb a modest cost error. One making forty cannot — at that frequency, costs are frequently the difference between a working strategy and a losing one.

This is where high-frequency rules quietly die. The gross result looks strong. Net of realistic costs, the edge is gone.

Read the cost and fill model in the backtest assumptions before reading the returns. If the strategy only works under optimistic fills, it does not work.

6. The data was not what you assumed

Data quality problems are dull and expensive.

  • Missing bars mean conditions that never got evaluated.
  • Corporate actions — splits, bonuses — create apparent price moves that were not real. A rule triggering on a split looks like a signal and is an artefact.
  • Instruments that no longer exist may be absent, which quietly removes the worst outcomes from the sample.
  • Contracts that expired are a specific issue for derivatives, where the instrument itself has a limited life.

The backtest states its data-quality assumptions explicitly, and they are worth reading first — because they define what the result is a statement about. The closed-bar rule is part of this: conditions evaluate when a bar closes, not the moment price touches a level intraday. That cuts both ways, and it is honest, but it means the backtest cannot represent a strategy that reacts intrabar.

7. Live conditions were not modelled

The last gap is between any backtest and the market.

  • Slippage. You will not always fill at the modelled price. Illiquidity and fast markets make this worse precisely when it hurts most.
  • Liquidity. A backtest happily fills orders at any size. The market does not. See ETF liquidity for how visible depth constrains real orders.
  • You. The backtest never skipped a signal because it seemed wrong, never exited early after two losses, never sized up after three wins. It executed every trade mechanically. You will not.

That third one is the largest and least discussed. Most strategies fail not because the logic was wrong but because the person running it stopped following it during the drawdown the backtest showed and they did not believe would feel like that.

Reading a result honestly

Work through the list in order:

  1. How many variants did I test? Count truthfully.
  2. Was the exit written before I saw results?
  3. How many occurrences? Under thirty, treat as indicative only.
  4. What regime was the period? Does it contain conditions that disfavour this rule?
  5. What costs and fills were assumed?
  6. What do the data-quality notes say?
  7. Would I follow this through its worst drawdown?

A strategy that survives all seven is worth something. One that survives none may still be profitable — but you do not have evidence of it, and you should not size as though you do.

What to do with a result that passes

Keep expectations proportionate. A backtest is a conditional statement about a specific past under stated assumptions. It is not a forecast.

Remember what deployment actually means here. Arming a cash strategy creates alerts only — it evaluates closed bars and notifies. Live strategy execution is excluded from the first release. Nothing in the workspace trades on your behalf, which means the backtest is informing your decisions rather than authorising a machine to act. See smart alerts and the execution boundary.

Build and test in the strategy workspace. And treat the backtest for what it is — a filter that eliminates bad ideas, not a mechanism that identifies good ones.

This is research and commentary, not personalised investment advice. Markets carry risk; past performance does not guarantee future results.

PocketX - powered by CapitalBridge

PocketX is a CapitalBridge product. Trading, demat and settlement services are provided by our broking partner, ATS Share Brokers Private Limited.