Guide · Testing and backtesting

How to avoid overfitting a trading strategy

An overfitted strategy describes the past perfectly and the future badly. By the end you will know the warning signs, the tests that expose curve fitting, and what each test can and cannot prove.

Based on our team’s research and live testing since 2018.

Overfitting, also called curve fitting, happens when a strategy’s rules and settings are tuned so closely to past data that they capture its noise instead of a pattern that repeats. The backtest looks excellent. The live account does not. It is one of the most common reasons an EA works in the backtest but not live.

No single test proves a strategy is not overfitted. A set of tests, each catching a different kind of fitting, gets you close. These are the ones we use in our backtesting and optimisation service.

What are the red flags of an overfitted strategy?

  • Results that look too good on a small sample: a very high profit factor or win rate from a few dozen trades.
  • Many rules and settings, especially filters with no clear reason, such as skipping one weekday or one specific hour.
  • Oddly precise best settings (a 37-period average, a 23-pip stop) whose neighbours perform badly.
  • A near-straight equity curve in the optimised period, followed by a flat or falling one afterwards.
  • Fragility: results collapse with slightly higher costs, a one-bar delay, or a related symbol.
  • Rules added after studying losing trades, one exception at a time.

Why do fewer parameters help?

Each parameter you optimise gives the strategy one more way to bend around the past. With enough of them, it can fit almost any history, including random data. Fewer parameters leave less room to fit noise, so whatever performance remains is more likely to be real.

Practical habits:

  • Fix what you can from reasoning, not optimisation. A session filter can follow the market’s opening hours instead of the best-scoring hour.
  • Optimise in coarse steps over sensible ranges.
  • Remove one rule at a time and retest. If the strategy collapses when a single filter is removed, that filter may be fitting specific trades.

Should you look for peaks or plateaus?

Plateaus. When you scan one setting across a range, a robust strategy shows a broad area where neighbouring values also work. A fitted one shows a single spike. The table shows two strategies scanned across moving-average periods, with illustrative profit factors:

MA period20304050607080
Strategy A (peak)1.020.981.051.621.010.971.03
Strategy B (plateau)1.121.211.261.291.271.221.15

Strategy A’s best number is higher, but nothing around it works, so the 50 setting probably matched the noise in this one sample. Strategy B is less exciting and far more believable. Choose a value from the middle of a plateau, not the highest point of a spike. With two settings, a heat map of scores shows the same thing in two dimensions.

How should you use out-of-sample data?

Split the history before you start. Optimise on the first part. Test the chosen settings once on the part the optimiser never saw. If you go back and adjust after seeing the out-of-sample result, that data has become in-sample, and you need new unseen data.

A single split can be lucky. Walk-forward optimisation repeats the split across many windows and is a stronger test. Keep a final stretch of data out of every experiment, for one last check before going live.

What does a Monte Carlo test actually test?

The most common version reshuffles the order of your backtest trades thousands of times and records the drawdown of each shuffled sequence. Here is a minimal version in Python:

monte_carlo.py22 lines
1import random2 3def max_drawdown(trades, start=10_000.0):4    """Largest fall from a peak, as a fraction of that peak."""5    equity = peak = start6    worst = 0.07    for result in trades:8        equity += result9        peak = max(peak, equity)10        worst = max(worst, (peak - equity) / peak)11    return worst12 13def reshuffle_drawdowns(trades, runs=5_000, seed=1):14    """Same trades, random order: returns the median and 95th percentile drawdown."""15    rng = random.Random(seed)16    results = []17    for _ in range(runs):18        sample = trades[:]19        rng.shuffle(sample)20        results.append(max_drawdown(sample))21    results.sort()22    return results[int(runs * 0.5)], results[int(runs * 0.95)]

What it tells you: how bad the drawdown and the losing streaks could have been with the same trades in a different order. The single backtest sequence is only one of many possible orders, and often not the worst. The 95th-percentile drawdown is a more cautious figure for sizing risk and for checking prop-firm limits.

What it does not tell you: whether the edge is real. Reshuffling keeps exactly the same trades, so the total profit never changes. If the trades came from fitted settings, every shuffle inherits the fitting. It also cannot show market conditions that are missing from the test. Variants that resample trades with replacement, or randomly skip some trades, vary the total as well, but they still only reuse the trades you already have.

How much cost stress should a strategy survive?

Costs are where fitted edges usually die first. Rerun the test with higher spreads, commission and slippage, and with a delay on execution. Illustrative numbers for a strategy whose average trade makes $10 before costs:

ScenarioCosts per tradeNet average trade
Base costs$4$6
Costs × 1.5$6$4
Costs × 2$8$2
Costs × 2.5$10$0

This strategy loses two-thirds of its edge if costs double, and all of it at 2.5 times. Live costs are often higher than backtest costs, around news, at the daily rollover and on slower connections. A strategy with a thin margin over its costs is fragile even if it is not overfitted. On MT5, run these tests on real ticks with a random execution delay, as described in our tick-data backtesting guide.

Why does testing many ideas lead to overfitting?

Every idea you test is another lottery ticket. Try a hundred variants and a few will look good by chance, even if none has an edge. Choosing the best of them is a form of overfitting, even if each single backtest was done correctly.

  • Keep a log of every variant you tried, not only the survivors.
  • Demand stronger evidence after a wide search: more trades, a higher statistical bar, and an untouched final test. Our guide to how many trades a backtest needs explains the numbers.
  • Prefer ideas you can explain before testing them. A reason in advance is worth more than a pattern found afterwards.

Checklist

  • Few optimised parameters, coarse steps, sensible ranges
  • Settings chosen from a plateau, not a spike
  • Out-of-sample or walk-forward results reviewed, with a final untouched period
  • Monte Carlo drawdowns used for risk sizing, not as proof of the edge
  • Edge survives higher costs and execution delay
  • Every tested variant logged, with the bar raised to match
Illustrative numbers. These tests reduce the risk of trading a fitted strategy; none of them can guarantee future results. Past performance does not indicate future results. Not financial advice.

Quick answers

What is the difference between overfitting and curve fitting?

In trading they usually mean the same thing: settings or rules tuned so closely to past data that they capture noise rather than a repeatable pattern.

Is any optimisation overfitting?

No. Choosing sensible settings from data is normal. It becomes overfitting when the settings only work at one exact point, or only on the data they were tuned on.

Can an overfitted strategy still make money live?

For a while, by luck. Over time it tends to behave like a strategy with no edge, minus costs. That is why out-of-sample results matter more than in-sample ones.

How much worse should out-of-sample results be?

Some drop is normal, because in-sample results always include some fitting. A collapse to a loss, or results that depend on one lucky window, are the warning signs.

Do machine learning strategies overfit more easily?

Usually yes, because they have many more free parameters. The same defences apply, with stricter separation between training, validation and final test data.

Written by
Rohit B., Vertex Algorithms

Co-founder and lead developer. Building MT4/MT5 Expert Advisors, Python trading bots and Pine Script since 2018, with trading analysis by co-founders Harshal K. and Bilal M. About the founders →

Think your strategy might be curve fitted?

We retest it out of sample, with walk-forward windows, cost stress tests and Monte Carlo runs, and I tell you plainly whether the edge survives.

WhatsApp