Guide · Testing and backtesting

Walk-forward optimisation explained, with a worked example

Walk-forward optimisation tests whether a strategy’s settings keep working on data they were not tuned on, again and again. By the end you will know the procedure, the two window styles, how to read the result, and how it relates to MT5’s forward option.

Based on our team’s research and live testing since 2018.

A normal optimisation finds the settings that worked best on the past. That is not the same as settings that will work next. Walk-forward optimisation (WFO) checks the difference by repeating a simple step many times: tune the strategy on one stretch of data, then test it on the stretch that comes after, which it has never seen.

It is one of the main tests we use in our backtesting and optimisation work, because it asks the question a trader actually faces: if I had optimised this strategy back then, what would have happened next?

How does walk-forward optimisation work?

  1. Optimise the strategy on an in-sample window, for example 12 months.
  2. Pick one set of settings using a rule you fixed before you started, such as the best custom criterion score.
  3. Test those settings, unchanged, on the next out-of-sample window, for example the following 3 months.
  4. Move forward by the length of the out-of-sample window and repeat.
  5. Join all the out-of-sample results into one equity curve. That curve is the result you judge.

Every trade in the final curve was made with settings chosen before that trade’s data was used. That is what makes it more honest than a single optimisation over the whole history.

Why not just optimise on all the data?

Because an optimiser always finds something. Given enough parameter combinations, some will fit the noise in your history, and the best-looking result is often the most fitted one. A single optimisation cannot tell you whether the strategy has an edge or the optimiser has found the past. WFO can, because the out-of-sample windows are the optimiser’s exam paper, and it never sees the questions in advance. More on this in how to avoid overfitting a trading strategy.

Anchored or rolling: which should you use?

StyleIn-sample windowGood forWatch out for
RollingFixed length that slides forward, dropping the oldest dataMarkets whose behaviour changes; strategies that need recent tuningShort windows give few trades and jumpy settings
AnchoredAlways starts at the same date and grows each stepSlow strategies that need every trade they can getOld data keeps its full weight, so the strategy adapts slowly

Neither is correct in general. Choose based on how the strategy is meant to be used, and decide before you look at results.

What does a walk-forward test look like? A worked example

A trend strategy with two optimised settings, tested from January 2021 to December 2022 with rolling windows: 12 months in-sample, 3 months out-of-sample, moving forward 3 months each step. The returns are average monthly returns, so windows of different lengths can be compared. All numbers are illustrative, not the result of a real test.

WindowIn-sampleOut-of-sampleChosen MA periodIn-sample per monthOut-of-sample per month
1Jan–Dec 2021Jan–Mar 202240+1.6%+1.0%
2Apr 2021–Mar 2022Apr–Jun 202245+1.4%+0.8%
3Jul 2021–Jun 2022Jul–Sep 202280+1.8%−0.2%
4Oct 2021–Sep 2022Oct–Dec 202245+1.5%+1.2%
Average+1.58%+0.70%

How we would read this:

  • Out-of-sample kept a bit under half of the in-sample return (0.70 ÷ 1.58 ≈ 44%). Some drop is normal: in-sample results always include some fitting.
  • Window 3 is the warning. The optimiser jumped from 45 to 80, produced its best in-sample number, and lost out of sample. Unstable settings like this often mean the optimiser chased noise in that window.
  • Four windows is a thin sample. Before trusting it we would want more windows and enough trades in each one. See how many trades a reliable backtest needs.

If you want to set up the windows in a script, this generates them for either style:

wfo_windows.py19 lines
1from datetime import date2 3def add_months(d, months):4    y, m = divmod(d.month - 1 + months, 12)5    return date(d.year + y, m + 1, 1)6 7def wfo_windows(start, end, is_months=12, oos_months=3, anchored=False):8    """Yield (in-sample start, out-of-sample start, out-of-sample end).9    End dates are exclusive. Rolling by default; anchored keeps the start fixed."""10    is_start = start11    oos_start = add_months(start, is_months)12    while add_months(oos_start, oos_months) <= end:13        yield is_start, oos_start, add_months(oos_start, oos_months)14        oos_start = add_months(oos_start, oos_months)15        if not anchored:16            is_start = add_months(is_start, oos_months)17 18for window in wfo_windows(date(2021, 1, 1), date(2023, 1, 1)):19    print(window)

What is walk-forward efficiency?

Walk-forward efficiency (WFE) is the out-of-sample performance divided by the in-sample performance, after putting both on the same time scale (per month or per year). The method is usually credited to Robert Pardo’s work on trading system testing. In the example above, WFE is about 44%.

  • A figure of about 50% or more is a commonly quoted rule of thumb for a reasonable result. It is a convention, not a statistical test.
  • Far above 100% is not a reason to celebrate. It usually means the out-of-sample period happened to suit the strategy, which is luck, not robustness.
  • A good WFE with a losing in-sample result means nothing. Look at the absolute numbers, the drawdowns and the number of trades as well.

Implementations differ: some use net profit, others a risk-adjusted measure. State which one you used, and compare like with like.

How does MT5’s forward option relate to walk-forward?

The MT5 Strategy Tester has a Forward setting: half, one third, one quarter, or a custom start date. The tester optimises on the first part of the period, then runs the best passes on the forward part. According to the MetaTrader 5 help, that is the best 10% of passes in a full (slow) search, or the best 25% with the genetic algorithm. You compare them on the Optimization Results and Forward Results tabs.

That is one in-sample/out-of-sample split, not a full walk-forward. It is a useful first filter, and we use it on every optimisation. But one split can be lucky or unlucky, and it never shows how the settings change over time. To run a real WFO in MT5, you repeat the optimisation for each window’s dates, apply your fixed selection rule, and test the chosen settings on the next window. That can be done by hand, by script, or with third-party tools.

How long should the windows be?

  • In-sample: long enough to contain plenty of trades for every combination you optimise, and more than one type of market.
  • Out-of-sample: long enough for a meaningful number of trades. A common choice is a quarter to a third of the in-sample length.
  • Re-optimisation schedule: the out-of-sample length is also how often you would re-optimise live. Pick something you will actually do.

What are the common mistakes?

  • Trying many window designs and reporting the one that passed. That turns WFO into one more optimisation.
  • Choosing settings by eye in each window instead of with a fixed rule.
  • Leaving costs out of either window. Use real ticks and your broker’s spreads and commission, as in our tick-data backtesting guide.
  • Ignoring drawdown. Check the joined out-of-sample curve against your risk limits. Prop-firm traders can test it against their firm’s floor with the drawdown calculator.

Checklist

  • Window style, lengths and selection rule fixed before the first run
  • Few optimised parameters, coarse steps
  • Real ticks and full costs in every window
  • Out-of-sample windows joined and judged as one curve
  • WFE, drawdown, trade count and parameter stability all reviewed
  • MT5’s forward split used as a first filter, not as the whole test
Illustrative numbers. A walk-forward test reduces the risk of trusting fitted settings; it does not predict future results. Past performance does not indicate future results. Not financial advice.

Quick answers

Is walk-forward testing the same as forward testing on a demo?

No. Walk-forward uses historical data in repeated optimise-then-test steps. Forward testing on a demo runs the finished EA on new live prices. Both are useful, and a demo forward test usually comes after a walk-forward test.

How many walk-forward windows do I need?

Enough that one lucky or unlucky window does not decide the verdict. Several windows, each with a meaningful number of trades, are more informative than two or three.

Which settings do I trade after a walk-forward test?

Usually the settings chosen by the most recent in-sample window, re-optimised on the same schedule as the test. If you trade fixed settings instead, test those fixed settings across all windows.

Can walk-forward optimisation be overfitted?

Yes. If you try many window lengths, ranges and criteria until the result looks good, you have optimised the walk-forward test itself. Choose the design first and run it once.

Does MetaTrader 4 support walk-forward testing?

Not natively. You can run the steps by hand with date ranges in the MT4 tester, but MT5’s tester is faster, has a built-in forward split and supports real tick data.

Written by
Rohit B., Vertex Algorithms

Co-founder and lead developer. Building MT4/MT5 Expert Advisors, Python trading bots and Pine Script since 2018, with trading analysis by co-founders Harshal K. and Bilal M. About the founders →

Want a walk-forward test of your EA?

We run rolling walk-forward tests on real tick data with your broker’s costs, and I send the out-of-sample results, the parameter history and a plain-English verdict.

WhatsApp