Multiple testing
Multiple testing is the problem that the more rules you try, the more likely one looks good by chance. A result has to be judged against how many were tried.
How Multiple testing is calculated
If you run 20 independent tests at a 5 percent significance level, you expect about one false positive. Corrections raise the bar: Bonferroni divides the threshold by the number of tests, and the Benjamini-Hochberg method controls the false discovery rate. The Deflated Sharpe Ratio adjusts a Sharpe ratio for the number of trials.
How it is read
A result that was the best of many tries needs a stronger test than one specified in advance.
Common mistakes
- Counting only the tests you kept.
- Treating small changes to parameters as one test. Each is a try.
- Reading an unadjusted p-value from the best of many.
Test it yourself
The Lab counts attempts on each experiment and shows the count on the receipt, so the number of tries stays visible. The Lab charges trading costs on every trade, fills on the next day's open, and shows how many attempts you have made.
What Tickfloor tested
Tickfloor's research desk has backtested 678 strategies, net of modelled trading costs. After correcting for the 761 tests run, 0 passed. That does not prove no strategy works. These are historical diagnostics, not validation under Testing Standard v2: the backtest harness predates that standard and has not been re-run to meet it.
Lessons that cover it
- Why testing many ideas finds lucky ones
- Indicators on price bars: 6,428 tests and the ADA trap
- Insider buying clusters: market beta in a costume
- What trading our own signals would have done
Related concepts
General information only. It doesn't consider your objectives, finances or needs. Tickfloor holds no financial services licence and never places trades.