Pairs Trading: does it work?
Pairs trading picks two assets that usually move together and bets on the gap closing when it widens. Statistical arbitrage is the same idea run across many assets at once.
The rule we tested. Four versions: long-leg-only divergence on US stocks, fixed same-industry pairs reverting from a 60-day z-score, re-convergence after a correlation breakdown on ETF pairs, and PCA-based statistical arbitrage across crypto majors.
- Pairs-trading divergence, long leg only (wide US stocks) (171 US stocks (AAPL, NVDA, AVGO, ORCL, …)), 2016-09-15 to 2025-03-14, costs 0.05% per side: −2.2 pts/yr vs the benchmark across 91 changes to the portfolio holdings (not completed trades), unadjusted p=1.000, BH-adjusted q=1.000: did not beat the benchmark.
- Fixed same-industry pairs, 60-day z-score reversion (40 US stocks (AAPL, NVDA, AMZN, GOOGL, …)), 2016-09-15 to 2025-03-14, costs 0.05% per side: −19.5 pts/yr vs the benchmark across 567 changes to the portfolio holdings (not completed trades), unadjusted p=1.000, BH-adjusted q=1.000: did not beat the benchmark.
- Correlation-breakdown re-convergence on fixed ETF pairs (15 ETFs (GLD, SLV, TLT, IEF, …)), 2005-01-04 to 2025-03-14, costs 0.05% per side: −7.1 pts/yr vs the benchmark across 10 changes to the portfolio holdings (not completed trades), unadjusted p=1.000, BH-adjusted q=1.000: did not beat the benchmark.
- PCA-based statistical arbitrage across crypto majors (18 crypto pairs (BTCUSDT, BNBUSDT, XRPUSDT, ADAUSDT, …)), 2017-08-18 to 2025-03-14, costs 0.10%–0.17% per side: −89.0 pts/yr vs the benchmark across 735 changes to the portfolio holdings (not completed trades), unadjusted p=1.000, BH-adjusted q=1.000: did not beat the benchmark.
These versions did not pass these tests. US stocks −2.2 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000); US stocks −19.5 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000); ETFs −7.1 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000); crypto pairs −89.0 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000). None cleared Tickfloor's Benjamini-Hochberg q<0.10 bar once weighed against every other rule tested alongside it. This describes the tested implementations, the costs and the benchmark stated here, not every version of the rule. These are historical diagnostics, not validation under Testing Standard v2: the backtest harness predates that standard and has not been re-run to meet it.
This is one line in a wider check: Tickfloor's research desk has run 517 strategies, most of them against an equal-weight benchmark of the same assets rebalanced daily that pays no costs (the strategies pay theirs), and after correcting for how many were tested (Benjamini-Hochberg, 598 tests), 0 passed. That does not prove no strategy works, and it says nothing about untested versions of a rule. These are historical diagnostics, not validation under Testing Standard v2: the backtest harness predates that standard and has not been re-run to meet it.
General information only. It doesn't consider your objectives, finances or needs. Tickfloor holds no financial services licence and never places trades.
Does pairs trading work?
These versions did not pass these tests. US stocks −2.2 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000); US stocks −19.5 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000); ETFs −7.1 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000); crypto pairs −89.0 pts/yr vs the benchmark (unadjusted p=1.000, BH-adjusted q=1.000). None cleared Tickfloor's Benjamini-Hochberg q<0.10 bar once weighed against every other rule tested alongside it. This describes the tested implementations, the costs and the benchmark stated here, not every version of the rule. These are historical diagnostics, not validation under Testing Standard v2: the backtest harness predates that standard and has not been re-run to meet it.
What exact rule did Tickfloor test?
Four versions: long-leg-only divergence on US stocks, fixed same-industry pairs reverting from a 60-day z-score, re-convergence after a correlation breakdown on ETF pairs, and PCA-based statistical arbitrage across crypto majors.
Is this financial advice?
No. This measures a publicly claimed strategy rule, not a recommendation. General information only, not personal advice.
See the full numbers on the Mean reversion, Machine learning family pages, or the full method and every result.