Out-of-sample testing
Out-of-sample testing checks a rule on data that was not used to design or tune it. Only the first look at that data counts as a clean test.
How Out-of-sample testing is calculated
Set a block of history aside before building the rule, often the most recent part, and write down the rule. Run it once on the held-back block. If you change the rule after seeing the result, that block has become part of the design and a new one is needed.
How it is read
A result that holds on held-back data is better evidence than one fitted on the data used to build it. A single held-back window is still one sample.
Common mistakes
- Peeking at the held-back data and then adjusting.
- Holding back too little data to produce enough trades.
- Choosing the held-back period after seeing which one looks good.
What Tickfloor tested
Tickfloor's research desk has backtested 678 strategies, net of modelled trading costs. After correcting for the 761 tests run, 0 passed. That does not prove no strategy works. These are historical diagnostics, not validation under Testing Standard v2: the backtest harness predates that standard and has not been re-run to meet it.
Lessons that cover it
- Testing on prices the strategy never saw
- Split-half, concentration and regimes: is it a few lucky trades?
- Momentum: +20,310% in discovery, a ruin path in reality
Related concepts
General information only. It doesn't consider your objectives, finances or needs. Tickfloor holds no financial services licence and never places trades.