What is an out-of-sample test — and does the filter that won on Bitcoin still work on ETH and SOL?
Two days ago we tested nine trend filters on the same Bitcoin breakout rule. One stood out: the 50-day filter made +19.4R where the plain rule made +1.2R. We also said, on the same page, not to trust it yet — it was the best of nine tries, on one coin. This page does the obvious next thing. We froze every setting, changed only the coin, and ran it on Ethereum and Solana. Here is what a result looks like when it meets data it has never seen.

KEY TAKEAWAYS
- An out-of-sample test is the cheapest honesty check in trading: build the rule on one set of data, then run it once, unchanged, on data you have not looked at.
- The 50-day filter improved the result on all three coins. But it improved the average trade by +0.26R on BTC, only +0.04R on ETH and +0.29R on SOL — and on ETH the rule still lost money.
- On ETH a filter that picks trades at random would have done as well 35% of the time. Cutting trades from a losing rule shrinks the loss automatically; that is not the same as the filter knowing something.
- Pooling ETH and SOL: 83 filtered trades made +3.5R, against −10.6R for 153 unfiltered ones. Random filters matched that 14% of the time. Encouraging, not convincing.
- The test also exposed a setting nobody questioned: the “skip stops wider than 6%” rule, harmless on BTC, threw away 80% of SOL’s breakout candles.
What is an out-of-sample test, in one paragraph?
An out-of-sample test checks a trading rule on data that played no part in creating it. The data you used to design, tweak and choose the rule is the in-sample data; everything else is out-of-sample. That can be a later stretch of the same chart (hold back the last year, look at it once at the end) or a different market (build on BTC, check on ETH). The one condition is that nothing changes between the two runs — same entry, same stop, same target, same filter, same numbers.
Why bother? Because any rule you tuned will look better on the data you tuned it on. Some of that is a real pattern in the market, and some of it is the rule bending to fit noise that will never repeat. The in-sample result cannot tell you which part is which. The out-of-sample result can, roughly: the real pattern tends to survive, the fitted noise does not.
What exactly did we carry over from Bitcoin?
Everything, with no adjustment. That is the hard part of the method: the temptation to “fix” one setting for the new coin is strong, and the moment you give in, the test is no longer out-of-sample.
- Signal: a 4-hour candle closes above the highest high of the previous 20 candles, after the candle before it had not.
- Stop and target: stop at the lowest low of the last 10 candles; target 2R; trade closed at market after 60 candles (10 days). Stops under 0.3% or over 6% of price are skipped.
- Filter: long only when the last completed daily close is above the 50-day simple moving average.
- Period: signals from 1 January 2022 to 30 September 2026 on Binance spot candles; one trade at a time.
We also re-ran Bitcoin with the new code first. It reproduced the original numbers exactly (122 trades and +1.2R without the filter, 72 trades and +19.4R with it), so any difference below comes from the coin, not the code.
What happened on ETH?
The filter made a losing rule lose less. It did not make it a winning rule. Ethereum is the harsher test of the two, because the plain breakout rule had nothing to work with there:
| ETH/USDT 4H, Jan 2022 – Sep 2026 | No filter | Above 50-day MA |
|---|---|---|
| Trades | 100 | 57 |
| Hit the 2R target | 25 (25.0%) | 15 (26.3%) |
| Total result | −14.4R | −5.8R |
| Average per trade | −0.14R | −0.10R |
| Worst peak-to-trough run | −24.7R | −12.1R |
Look at the second-to-last row before the third. The total improved by 8.6R, but most of that is arithmetic: a rule that loses on average loses less if it trades less. Per trade, the filter moved ETH from −0.14R to −0.10R. With a standard error of about 0.18R on 57 trades, that difference is invisible.
Two more checks point the same way:
- Veto audit. Of the 100 trades the plain rule took, the filter would have allowed 51 (net −5.3R) and refused 49 (net −9.1R). The refused pile was a bit worse, which is the right direction — but both piles lost.
- Random filter. We drew 57 trades at random from the plain rule’s 100, twenty thousand times. 35.4% of those random “filters” did as well as −5.8R or better. A filter that knew nothing would match it about one time in three.
The year that hurt most was 2024. With the filter on, ETH took 17 breakouts above its 50-day average that year and only 4 reached the target, for −5.0R. The screenshot shows why: twice, in March and in late May, Ethereum broke out well above a rising 50-day line — textbook “with the trend” — and both times the move stalled within days.

The filter did its job in the narrow sense — it only allowed longs in an uptrend. The problem is that on ETH, in this period, breakouts in an uptrend were not much better than breakouts in general. A filter cannot rescue a signal that has no edge to begin with. It can only choose among the signal’s trades.
What happened on SOL?
Here the filter looked good again — on too few trades to lean on:
| SOL/USDT 4H, Jan 2022 – Sep 2026 | No filter | Above 50-day MA |
|---|---|---|
| Trades | 53 | 26 |
| Hit the 2R target | 18 (34.0%) | 11 (42.3%) |
| Total result | +3.8R | +9.3R |
| Average per trade | +0.07R | +0.36R |
| Worst peak-to-trough run | −8.7R | −3.7R |
This is the pattern you would hope for from a real filter: fewer trades, a higher hit rate, a smaller drawdown, a better average. The veto audit agrees — the 25 allowed trades netted +8.0R, the 28 refused trades −4.2R. And unlike on Bitcoin, the neighbouring settings behave: the 100-day version made +9.4R, the 150-day +10.0R, the 200-day +7.0R. A smooth result across settings is what a real effect tends to look like.
The catch is size. Twenty-six trades over almost five years. Random draws of 26 trades from the 53 matched +9.3R 7.8% of the time, and the average of +0.36R per trade has a standard error of about 0.29R. One or two different outcomes would change the story. In 2026 alone the filtered rule took 10 SOL trades and lost 1.7R.
So did the filter pass?
Half. Here is the scorecard, built from the same checks we used on the Bitcoin page:
| Question | ETH | SOL |
|---|---|---|
| Does the filter improve the average trade? | Barely (+0.04R) | Yes (+0.29R) |
| Is the filtered rule profitable? | No (−5.8R) | Yes (+9.3R) |
| Are refused trades worse than allowed ones? | Slightly | Yes |
| Would a random filter do as well? | Often (35%) | Sometimes (8%) |
| Enough trades to judge? | Barely (57) | No (26) |
Putting the two new coins together gives the cleanest single number: 83 filtered trades made +3.5R (31.3% hit the target), while the 153 unfiltered trades lost 10.6R. Random selections of 83 trades from those 153 did as well 13.9% of the time.
So the honest reading is: the direction carried over, the size did not. On Bitcoin the filter added a quarter of an R to every trade. On the two new coins together it added about a tenth (from −0.07R to +0.04R per trade), with a margin of error as big as the effect. That is exactly the shrinkage you should expect from a setting that was the best of nine. The in-sample winner almost always comes back smaller; the question is only whether anything is left. Here, something might be. Not enough to size up on.
What hidden assumption did the test expose?
A setting we never thought of as a choice. The rule skips any breakout whose stop is more than 6% below the entry, to avoid trades where the stop is unreasonably wide. On Bitcoin that limit rarely bites. On the other coins it decides most of the outcome:
| Breakout candles, Jan 2022 – Sep 2026 | Stop within 0.3–6% | Stop wider than 6% (skipped) |
|---|---|---|
| ETH/USDT 4H | 171 | 204 (54%) |
| SOL/USDT 4H | 78 | 304 (80%) |
That is why SOL produced only 53 trades. The “same rule” on Solana is really a rule that only trades Solana’s quietest breakouts. Whether that is good or bad, it is not the rule we thought we were testing. This is one of the most useful things an out-of-sample run does: it shows you which of your numbers were secretly fitted to the first market. A fixed percentage tuned to Bitcoin’s volatility is one of them; a version based on each coin’s own average range would travel better — but that is a new rule, and it needs its own test.
How do you run an out-of-sample test on your own rule?
Decide what “passing” means before you look, run it once, and accept the answer. A routine that works with a spreadsheet:
- Freeze the rule in one sentence. Entry, stop, target, time limit, filter, every number. If you cannot write it down, you cannot test it.
- Choose the out-of-sample data before you tune anything. The last 12–24 months, or a second market. Do not open that chart while you are building.
- Write the pass test in advance. For example: “average trade above +0.1R and refused trades worse than allowed trades.” Use the scorecard above as a template.
- Run it once. If it fails and you change the rule, the data you just used is now in-sample. You need fresh data for the next attempt.
- Expect shrinkage. Plan your position size on the out-of-sample result, not the in-sample one. Our Kelly criterion page shows how much an overstated edge costs you.
- Practise on the new market blind. Numbers tell you whether a rule works; replay tells you whether you can follow it on an unfamiliar chart.

In replay, switch the coin to one you did not build the rule on, apply the filter from the higher-timeframe panel, and step forward bar by bar. Log each decision. Twenty logged decisions will not prove anything statistically, but they will show you quickly whether the rule even makes sense on a market that moves differently.
What an out-of-sample test is NOT
- It is not a second optimisation. Re-tuning the setting for ETH and reporting the best one is an in-sample result on ETH. On ETH every length we tried lost money; the best (50-day) lost the least.
- It is not proof. Passing one out-of-sample test with 26 or 57 trades moves the odds; it does not settle them.
- It is not repeatable on the same data. Once you have looked at the held-back data and changed something because of it, it stops being held back. Most people burn their out-of-sample data in the first evening.
- It is not only for code. A discretionary trader can do the same thing: write the checklist, then apply it to a coin or a year you have not studied, and journal the results.
Where this reasoning breaks down
- Correlated markets. ETH and SOL move with Bitcoin much of the time, so they are not fully independent tests. A different asset class would be a tougher one.
- Small samples. 26 to 100 trades per run. Differences of a few R are well inside the noise, as the random-filter checks show.
- The 6% stop limit changes the rule per coin (section 6). The fair reading is “this rule as written”, not “breakouts on SOL in general”.
- Same period. We changed the market but not the years. A time split (choose on 2022–2024, judge on 2025–2026) on ETH gives the same verdict: the 50-day was the best length in 2022–2024 at −3.5R and lost 2.4R on 20 trades afterwards. On SOL the best length in 2022–2024 was the 150-day, and it made +1.0R on 11 trades afterwards.
- No costs, long only. Fees and slippage would cut every total by roughly 1–3R; shorts were not tested.
- Random-filter checks are approximate. They sample from the plain rule’s trades and do not re-run the one-trade-at-a-time logic.
FAQ
What is an out-of-sample test in trading?
It is a test of a trading rule on data that was not used to build or tune it, such as a later time period or a different market, with every setting left unchanged. It shows how much of a backtest result was a real pattern and how much was the rule fitting past noise.
How much data should I keep out-of-sample?
Enough to produce a meaningful number of trades, usually at least 30 and ideally 100 or more. Many traders hold back the most recent 20 to 30 percent of their history, or test on a second market. The key is to decide before you start tuning and not to look at that data until the end.
Why do out-of-sample results usually look worse?
Because the in-sample result includes some luck that the tuning process captured. When you pick the best of several settings, you pick partly the one that happened to fit the past noise best. That part does not repeat. In our test the 50-day filter added about 0.26R per trade on Bitcoin, where it was chosen, and about 0.11R per trade on ETH and SOL combined.
Does a 50-day moving average filter work on Ethereum?
Not on its own in our test. On every 20-bar breakout on the ETH/USDT 4-hour chart from January 2022 to September 2026, the filter cut the trades from 100 to 57 and the loss from 14.4R to 5.8R, but the filtered rule still lost money and a random filter did as well 35% of the time.
Can I re-test after an out-of-sample failure?
Yes, but not on the same data. Once you have changed the rule because of what the out-of-sample data showed, that data has become in-sample. Use a new period or a new market for the next test, or wait and collect live results in a journal.