Why Backtests Look Great and Live Results Do Not
You backtested a moving-average crossover on one stock, five years of daily bars, 2019 through 2023. The equity curve came out clean, the win rate sat at 68%, and on a $5,000 position the thing projected somewhere near $950 a year. You went live in January. Three months and twenty-five trades later the account is down $180.
Nothing broke. The gap between those two results is mostly built into how the backtest was constructed, and it comes from two places: the strategy was selected for fitting the past, and the past was free to trade.
The one variation that fit the noise
Go back to how that 68% number came into existence. You didn’t test one rule. You tested a fast average of 5, 8, 10, 15, or 20 days against a slow average of 30, 50, or 100, with three stop-loss distances, and you looked at the results table and picked the row at the top. Call it fifty variations. Forty-nine of them looked ordinary or bad. One looked excellent.

That row is not a discovery. It is the maximum of fifty noisy measurements, and the maximum of a set of noisy measurements is biased high by construction. Some of what makes the winning row look good is real structure in the data. Some of it is the particular sequence of gaps and reversals that stock happened to print between 2019 and 2023, which will not repeat. The search doesn’t tell you which portion is which. It just hands you the number.
The coin version makes it obvious. Flip fifty fair coins fifty times each, then keep the coin with the most heads. It will beat 50% comfortably, and you will have learned nothing about that coin. You could write up its “edge.” You could even give it a Sharpe ratio.
This is overfitting, and the damage scales with two things: how many variations you tried, and how few trades each test produced. Fifty variations judged on 100 trades apiece is a wide search over thin evidence. One pre-specified rule judged on 2,000 trades is a narrow search over thick evidence. The first result deserves far more suspicion than the second, even when both print the same win rate.
The stocks that aren’t in your data
There’s a second way the past gets flattered, and it hides in the data file rather than the code.

Suppose you widen the test from one stock to the S&P 500, and you build the universe by pulling the index’s current constituent list and running it backward. Every company in that file survived to be on the list today. The ones that went bankrupt, got delisted, or were bought out at a discount are simply absent — not marked as losses, just gone. Your strategy never had the chance to hold them into the ground, because your dataset never offered them.
Someone trading in real time in 2019 did not have that list. They had the 2019 list, which included names that no longer exist. Survivorship bias doesn’t add a fixed penalty you can subtract later; it changes which trades your strategy was ever asked to take. A rule that buys dips looks very different when the dips that never recovered are in the sample.
The single-stock backtest above dodges this one, since the stock is still trading. Broaden the universe and it comes straight back.
What the spreadsheet doesn’t charge you for
Backtests fill you at the price on the chart. Markets don’t.

Three costs go missing, and they’re worth separating because they scale differently. Slippage is the gap between the price that triggered your signal and the price you actually got. Your rule fires on the close; you’re filled on the next open, or you’re filled at 9:31 four cents higher than where you clicked. Commissions and the spread are the toll for the round trip — even a zero-commission broker routes you across a bid-ask spread that a backtest built on closing prices has never seen, because a closing print is a single number and a real market is two. Market impact is your own order pushing the price against you as it executes, which for a 100-share retail order in a liquid name is close to nothing and for a 50,000-share order in a thin one is the whole story.
Frequency decides how much this matters. A strategy that holds for eighteen months pays these costs twice and can absorb a lot of them. A strategy that turns over a hundred times a year pays them a hundred times, and the arithmetic gets ugly fast.
Running the numbers on the same strategy
State the assumptions plainly, because every figure below depends on them. The strategy takes 100 round trips a year. Position size is 100 shares of a stock trading near $50, so about $5,000 at risk per trade. It wins 55% of the time. The average winner makes $0.50 a share, or $50. The average loser costs $0.40 a share, or $40. All of it is hypothetical, built to show the mechanics rather than to describe any real system.
Gross expected value per trade is (0.55 × $50) − (0.45 × $40) = $27.50 − $18.00 = $9.50. Over 100 trades, $950 a year. That’s the number on the backtest report.
Now charge for reality. Commission and fees, $1.00 per round trip. Slippage of two cents a share on entry and two cents on exit, which on 100 shares is $2.00 each way, or $4.00 round trip. Two cents is not a worst case; it’s an ordinary retail fill in a liquid name that isn’t quoted a penny wide.
| Item | Backtest, no frictions | Same strategy, with frictions |
|---|---|---|
| Trades per year | 100 | 100 |
| Win rate | 55% | 55% |
| Average win | $50.00 | $50.00 |
| Average loss | $40.00 | $40.00 |
| Gross EV per trade | $9.50 | $9.50 |
| Commission, round trip | $0.00 | $1.00 |
| Slippage, round trip | $0.00 | $4.00 |
| Net EV per trade | $9.50 | $4.50 |
| Expected year | $950.00 | $450.00 |
Five dollars of friction against a $9.50 edge removes 53% of the expected profit. The win rate didn’t move. The average win didn’t move. The strategy is identical in every respect except that it now pays to trade.
Notice what that means about the edge’s margin for error. The whole thing lives in a band $9.50 wide, and costs eat the bottom $5.00 of it. Anything that shaves another two or three dollars off the gross — a wider spread on volatile days, a fill on the open instead of the close — takes the rest.
When both problems stack
The two failures multiply, and this is where the live account ends up at −$180 instead of merely underperforming.
That $9.50 gross edge was the best of fifty tested variations. Suppose, purely as an illustration, that half of it was selection artifact and half was real. The honest gross edge is then $4.75 per trade. Subtract the same $5.00 of commission and slippage and you get −$0.25 per trade, or about −$25 over a hundred trades, before you account for the fact that a hundred trades is nowhere near enough for the average to show up cleanly.
A backtest that said $950 becomes a live strategy that grinds slowly negative and swings a few hundred dollars either way on noise. Nobody lied. No line of code was wrong. The report answered the question “how would this rule have performed if trading were free and I had known in advance which of fifty variations to pick,” and that is not the question anyone was asking.
What this does not tell you
It does not tell you whether your specific strategy is overfit. The 68% figure above is a made-up illustration, not a claim about crossovers, indicators, or any system in particular. Diagnosing your own requires out-of-sample testing on data the rule never touched during development, and even that isn’t conclusive.
It does not capture regime change. A rule fitted entirely to a calm stretch can behave completely differently once volatility rises, and that failure has nothing to do with parameter search. Backtests describe the past. They are not forecasts, and no amount of statistical hygiene converts one into the other.
It says nothing about your execution. Two people running byte-identical rules will get different live results depending on order type, time of day, whether they use limits or market orders, and whether they take the signal on the third losing trade in a row or quietly skip it.
And it ignores position sizing and risk of ruin entirely. Positive expected value per trade is compatible with a drawdown deep enough that you stop trading before the long-run average has any chance to arrive. Expected value is a statement about an infinite sequence. You are trading a finite one, with a finite tolerance for pain.
Finally, the $4.00 slippage figure is an assumption, not a measurement. Your real number depends on the instrument, order size, time of day, and broker routing. The honest way to get it is to compare your own fills against the price that triggered each signal, over a few dozen trades, and use what you find.
FAQ
Is a backtest useless if it doesn’t include trading costs?
Incomplete, not useless. A cost-free backtest still tells you something about whether the underlying logic has any structure to it. What it cannot tell you is whether that structure survives friction. Read it as an upper bound on performance, never as an estimate.
How much data do I need before a backtest means anything?
There’s no universal number, and the honest answer depends on the ratio between trades and tuned parameters. Twenty trades with five tuned parameters tells you close to nothing. Two thousand trades on one fixed rule tells you a great deal. Every extra parameter you tune is another chance to fit noise, so the data requirement grows with the size of your search, not just with the strategy.
What is out-of-sample testing and why does it matter?
You hold back a chunk of data, develop the strategy using only the rest, then run it once on the untouched portion. If the results collapse there, the original number was fit to noise. It’s the most direct check available against overfitting, though it can be gamed — if you keep adjusting the rule and re-testing on the same held-out data, that data is no longer out of sample.
Does a higher backtested win rate mean a better strategy?
No. Win rate ignores the size of wins against losses entirely. A strategy winning 80% of the time and giving it all back on the occasional disaster can have a worse expected value than one winning 45% with tight losses. Net expected value per trade, after costs, is the figure that matters.
Can a strategy pass out-of-sample testing and still fail live?
Yes, routinely. Out-of-sample testing checks for overfitting to a particular historical dataset. It cannot recover costs that were never modeled, it cannot anticipate a shift in market conditions after the test window ends, and it says nothing about whether you’ll follow the rules once real money is moving. Passing is a necessary step, not a clearance.
Are there rules about advertising backtested performance?
Registered investment advisers in the US fall under the SEC’s Marketing Rule, which governs how hypothetical and backtested performance may be presented — precisely because these numbers are so easy to dress up. If you want the regulators’ own material rather than a secondhand summary, the SEC’s investor education site and FINRA’s investor resources are the primary sources.
What to look at next
Three questions cut through most backtest presentations, whether the backtest is yours or someone else’s. How many variations were tested before this one was chosen? Were commissions, spread, and slippage charged, and at what assumed rate? Did any of the result hold up on data the rule never saw during development?
If the honest answer to the first is “I don’t remember,” treat the result as the maximum of an unknown number of tries. On where slippage and market impact actually come from — the mechanics of order books, quoting, and execution — the Bank for International Settlements publishes market microstructure research that’s useful background before you pick a number to charge yourself.
This article is general information, not financial advice. See our disclaimer.
Sources
Related articles
- What Overnight Financing Charges Cost on a Leveraged Position A worked breakdown of how overnight financing fees on leveraged trades accumulate daily and erode returns over time.
- Stop Losses: The Trade-Off Between Protection and Whipsaw How stop orders actually fill, why tight stops get triggered by normal noise, and the math behind picking a stop distance.
- Position Sizing: The Part Beginners Skip and Regret Learn how to calculate trade size from risk per trade instead of gut feeling, with a worked example and its limits.