How many trades do you need for a reliable backtest?
There is no magic number, but you can work out a sensible one for your strategy. By the end you will know why 30 trades proves little, how the size of your edge sets the sample you need, and why testing many variants raises the bar.
Based on our team’s research and live testing since 2018.
Traders often ask us for a single number: how many trades before a backtest means something? The honest answer is “it depends on the edge”, which is not much use on its own. This guide gives you the heuristics people use, explains where they come from, and shows a short calculation you can run on your own trade list.
Everything here is a heuristic, not a law. Markets change, trades are not perfectly independent, and no sample size turns a backtest into a promise.
Why does the number of trades matter?
A backtest gives you an average result per trade. That average is an estimate, and every estimate contains noise. With few trades, a short run of lucky wins can make a strategy with no edge look good. With more trades, luck has more chances to cancel out.
One fact from statistics drives everything else: the uncertainty around an average shrinks with the square root of the number of trades. Four times as many trades only halves the uncertainty. That is why going from 30 to 100 trades matters a lot, and going from 1,000 to 1,100 barely matters.
Is 30 trades enough for a backtest?
You will often see 30 quoted as a minimum. It comes from general statistics teaching, where averages of around 30 observations start to behave predictably. For trading, 30 is a weak floor, not a target:
- Trading results have fat tails. Two or three large wins or losses can dominate 30 trades.
- Thirty trades often come from one market regime, such as a single trend or a single quiet range.
- At 30 trades, only a very large edge can be told apart from luck.
Treat a strategy with fewer than 30 trades as untested. Between 30 and 100, treat it as an idea that deserves more testing.
What is a practical target?
A common practical target is 100 to 200 trades or more, spread across different market conditions. It is the reason the custom optimisation criterion in our tick-data backtesting guide gives a score of zero to any run with fewer than 100 trades, and it is the first thing we check when a client sends a backtest to our backtesting and optimisation service.
But the right number for your strategy depends on two things: how large the average result per trade is, and how much results vary from one trade to the next.
How do you work out the number for your strategy?
A simple check is the t-statistic of the average trade: the average result divided by its standard error. A t-statistic of about 2 is the conventional bar for one test that was planned in advance. Below that, luck explains the result easily. Here is the calculation in Python. It works with results in money, pips or R (multiples of the amount risked per trade), as long as costs are included.
1import math2import statistics3 4def edge_check(results):5 """results: one number per closed trade (money, pips or R), costs included."""6 n = len(results)7 mean = statistics.mean(results)8 sd = statistics.stdev(results)9 t_stat = mean / (sd / math.sqrt(n))10 return n, mean, sd, t_stat11 12def trades_needed(mean, sd, t_target=2.0):13 """Rough number of trades before an edge of this size reaches t_target."""14 return math.ceil((t_target * sd / mean) ** 2)15 16# Illustrative: average +0.1R per trade, standard deviation 1.0R17print(trades_needed(0.1, 1.0)) # 400Rearranging the same formula tells you roughly how many trades an edge of a given size needs to reach that bar. The table uses illustrative numbers, measured in R:
| Average trade | Standard deviation | Trades for t ≈ 2 | What it means |
|---|---|---|---|
| +0.50R | 1.2R | 24 | A large edge shows up quickly |
| +0.30R | 1.2R | 64 | A solid edge still needs dozens of trades |
| +0.10R | 1.0R | 400 | A small edge needs hundreds |
| +0.05R | 1.0R | 1,600 | A very small edge needs thousands, and costs can erase it |
The pattern matters more than the exact figures: halve the edge and you need four times the trades. Many real strategies sit at the small-edge end of this table, which is why a few dozen trades rarely settle anything.
Does trade frequency change the answer?
The statistics count trades, not years. But frequency changes what those trades represent.
- Frequent strategies such as scalpers and intraday systems collect hundreds of trades quickly. That sample may still come from a few months of one market regime, and costs weigh more on each trade. Test across several years anyway.
- Slow strategies such as swing systems on H4 or daily charts may produce only a few dozen trades a year. Reaching 200 trades can take many years of data, or running the same rules on several related symbols.
So ask two questions, not one: are there enough trades, and do they cover enough different conditions? Look for trends, ranges, high and low volatility, and news-heavy periods in the test.
How many trades per optimised parameter?
Every setting you optimise gives the strategy another way to fit the past. With five optimised settings and 150 trades, the optimiser has a lot of freedom to find a combination that happened to match the noise in those 150 trades.
There is no agreed ratio. The working principle is simple: the more parameters you optimise, the more trades you need, and the more unseen data you should hold back for checking. In practice we would rather cut parameters than chase more data. A strategy with two optimised settings and 300 trades is far easier to trust than one with eight settings and the same 300 trades.
What is the multiple-testing trap?
The t-statistic above assumes you tested one idea. Most strategy development tests dozens or thousands of variants: different indicators, periods, filters and symbols. Test enough variants and some will pass a t ≈ 2 bar by luck alone. Pick the best of them and its backtest is biased upward, however many trades it has.
- Keep a count of every variant you tried, including the ones you threw away.
- Raise the bar after a wide search. Some academic work on multiple testing argues for a t-statistic of about 3 rather than 2, and formal corrections such as the deflated Sharpe ratio exist for this problem.
- Hold back data that no variant ever saw, and test only your final choice on it, once.
We cover this in more depth in how to avoid overfitting a trading strategy, and rolling out-of-sample tests in walk-forward optimisation explained.
What does a large sample still not tell you?
- It does not prove the edge will continue. Markets change, and a past edge can fade.
- It does not fix a flawed test. Missing costs, code that reads the unfinished bar, or a repainting indicator give you a large sample of trades that could never have happened.
- Trades are not fully independent. Overlapping positions and clustered losing streaks make the true uncertainty larger than the formula suggests.
Checklist
- Fewer than 30 trades: untested. 100 to 200 or more: a reasonable start, if spread across different conditions
- t-statistic of the average trade calculated, with all costs included
- Small edge? Plan for hundreds or thousands of trades
- Few optimised parameters, and a written count of every variant tried
- An out-of-sample period, checked once at the end
- Risk per trade sized from the worst tested losing streak, for example with the position size calculator
Quick answers
Is 50 trades enough to trust a strategy?
Usually not. Fifty trades can only separate a large edge from luck. Treat it as a promising first look, then test over more data, more market conditions or related symbols.
Does a longer backtest period always help?
More trades help, but very old data may come from a market that behaved differently. Aim for enough trades across several regimes, and check that the recent part of the test still holds up.
Can I combine trades from several symbols?
Yes, if the same rules run unchanged on each symbol. That raises the sample and shows whether the idea is general. Results on correlated symbols are not fully independent, so do not count them as entirely new evidence.
Should I count trades or winning days?
Count the unit the strategy makes decisions on, usually closed trades. For strategies that hold many overlapping positions, daily returns are often the more honest unit.
Does a high win rate reduce the trades I need?
Not on its own. What matters is the average result per trade compared with how much results vary. A high win rate with rare large losses can need more trades, not fewer.