# Backtests Are the Most Optimistic Estimates a Strategy Will Get

> 208 of 212 published stock-return predictors, 1926 to 2024, compared using reconstructed backtest, forward and post-publication returns, not verified live trading. Pooled returns fell 29.4% in the forward window and 55.3% after publication. A 50/50 blend cut forecast error by 14.5%. 13 research pages, 12 direct sources.

Published: 2026-09-28
Publisher: BlackRidge (https://blckridge.com/)
Canonical: https://blckridge.com/research/forward-validation-20260928/
PDF: https://blckridge.com/research/forward-validation-20260928/forward-validation-20260928-en.pdf

---

# Backtests Are the Most Optimistic Estimates a Strategy Will Get

Of 212 published stock-return predictors, 208 enter the main comparison after exclusions. Historical returns are reconstructed for backtest, forward and post-publication periods. These are not verified live trading results.

## Published predictors lost 29% of their backtest return in the forward window and 55% after publication

The backtest is the high-water mark of a strategy's evidence [04] . The decline starts before publication in these reconstructed returns. A forward window (sample end to publication, median 48 months) provides a useful check, although authors may have seen those years. Adding it to the backtest cut forecast error by about a seventh on average, a gain coming entirely from predictors published after 2006 (down 32.6%, with no improvement for older predictors) [01] .

## The edge peaks in the last five backtest years and loses a third in the first five years after

The peak just before the sample ends is what selection looks like. Researchers stop the sample where the result is strongest. Sometimes the effect was simply strongest then [01] [02] .

Average return of 208 predictors as a percentage of their in-sample mean, from 10 years before to 20 years after the original sample ends [01] [02] . Years from original sample end

## The strongest backtests lost the most and ended level with the middle

Investors often treat a massive in-sample t-statistic as proof of a durable edge. The data tell a different story. Across 208 predictors, the in-sample t-statistic explains about 3% (R-squared 0.03) of the variation in later returns [01] .

- Group
- In-sample t range
- Backtest return a month
- After publication a month
- Share kept
- Weakest fifth
- 1.0–2.3
- 0.42%
- −0.05%
- −11.9%
- Second fifth
- 2.3–2.9
- 0.50%
- 0.34%
- 67.4%
- Middle fifth
- 2.9–3.7
- 0.73%
- 0.43%
- 59.6%
- Fourth fifth
- 3.7–5.3
- 0.79%
- 0.39%
- 49.3%
- Strongest fifth
- 5.4–14.2
- 1.07%
- 0.40%
- 37.3%

## Nine in ten predictors cleared t = 2 in the backtest, fewer than three in ten after publication

Most backtests clear traditional significance hurdles easily. Shorter track records after publication explain only a small fraction of the subsequent drop [01] . Harvey, Liu and Zhu have already argued that any new factor needs to clear a t-statistic of 3.0 [05] .

- t-statistic band
- in sample
- after publication
- below 0
- 0.0%
- 15.0%
- 0 to 1
- 0.5%
- 29.6%
- 1 to 2
- 10.1%
- 26.2%
- 2 to 3
- 30.8%
- 17.5%
- above 3
- 58.7%
- 11.7%

## The best of 100 useless ten-year backtests shows a Sharpe ratio of 0.8 on average

Every variant a quantitative team tries acts as a lottery ticket. The winning ticket's reported Sharpe ratio demands harsh judgment against the sheer number of tickets bought. Selection makes the best of many zero-skill results look like skill; it does not isolate forecasting ability [06] .

Expected maximum Sharpe ratio of N independent zero-skill strategies, annualised, using the deflated Sharpe ratio formula [06] . Number of strategies tried (log scale)

## A Sharpe ratio of 0.5 needs 11 years of evidence, 17 if it crashes like momentum

A six-month paper-trading period says almost nothing about a modest strategy [07] . Sample size dictates confidence. The shape of returns matters too: momentum-like skew stretches the requirement from 11 to 17 years [07] [03] . Any strategy carrying crash risk requires a massive observation window just to rule out luck.

Minimum track record length, years, 95% one-sided, monthly data. Momentum skewness and kurtosis from 1927-2026 data [07] [03] . Annualised Sharpe ratio

## Dropping micro caps and weighting by size removes a quarter to a third of the backtest before any cost

Statistical overfitting only explains part of the performance gap. Academic backtests routinely assume equal weights across universes thick with tiny, illiquid equities that no institutional fund could efficiently trade [01] . The problem worsens with complexity, as bank-promoted alternative beta strategies suffer massive degradation when moved from paper to production [08] .

## Every classic factor earned less after publication. Value lost money for 14 years.

Even the best-known factors kept only a fraction of their backtest returns after decades of scrutiny [03] . Value earned −5.7% a year from 2007 to 2020, then surged to +7.9% a year from 2021 through August 2026 [03] . A forward test that covers one regime can mislead you in either direction.

Rolling ten-year annualised return, %, June and December points, 1973-August 2026 [03] .

- Factor
- Published
- Original sample
- In sample, % a year
- After publication, % a year
- Worst decline after
- Value (HML)
- 1980
- 1962–1976
- 7.0%
- 2.7%
- −57.8%
- Size (SMB)
- 1981
- 1926–1975
- 1.7%
- −0.4%
- −54.9%
- Momentum
- 1993
- 1964–1989
- 9.3%
- 2.9%
- −57.8%
- Profitability (RMW)
- 2006
- 1977–2003
- 4.2%
- 2.8%
- −25.9%
- Investment (CMA)
- 2008
- 1968–2003
- 5.4%
- 0.0%
- −27.6%

## Adding a four-year forward window cut forecast error by 14.5%, all of it among papers published after 2006

The forward window alone offers no more predictive power than the original backtest. Averaging the two produces a clear improvement [01] . R-squared rises from 0.19 to 0.26, and a flat or negative forward window is a strong warning (19.3% of the backtest kept, against 51.1%).

## Momentum gave back 17.2% in July and August 2026, a two-month loss seen in 1.3% of windows since 1927

A regime turn is when a backtest calibrated on the previous regime is found out. The gap between the trailing year in June and in August shows how fast. French's momentum factor recorded exactly this shift [03] .

trailing 12-month return, %, monthly, January 2016 to August 2026, [03]

## Forward data reduces forecast error and exposes backtest decay

We examined published stock predictors to measure performance on data outside the original sample. Returns in the post-publication period earned 55.3% less than the backtest. The decay begins immediately.

- Measure
- What it shows
- Value
- Forward window relative return
- Return shortfall compared to the original sample
- −29.4%
- Post-publication relative return
- Return shortfall after the research appears in print
- −55.3%
- Strongest fifth retention
- Share of performance maintained by the best backtests
- 37.3%
- Random backtest Sharpe
- Expected ratio from multiple attempts with zero skill
- 0.80
- Years to confirm Sharpe
- Minimum track record needed at standard confidence levels
- 11.0 y
- Change in forecast error
- Improvement in predicting post-publication returns using forward data
- −14.5%

Before allocating capital to a systematic strategy, demand the number of variants tried and a frozen rule. Evaluate these alongside years of forward data judged together with the backtest.

## Selected sources

This document is research, not investment advice. Returns appear gross of costs and taxes. Past decay does not place a limit on future decay.

## Keep up with the research.

Choose optional research updates, or learn how the account structure works.

## Independent research, introductions clearly defined.

BlackRidge is an independent research bureau introducing private investors to quantitative traders through multiple strategy providers. We publish quantitative research. Investors access strategies through PAMM accounts at the broker. We are not a fund or broker and never hold client money.

Your funds stay at the broker. The account is opened in your own name. BlackRidge never receives, holds, or has withdrawal rights over client capital.

No fee on your profits. A one-time access payment of 6.7% of the agreed trading level, paid to the selected strategy provider. No management fee, no performance fee, no profit share, in any year.

You fund the risk, not the exposure. Allocations are notionally funded: you agree a trading level and fund the margin and drawdown allowance behind it. Trading losses can exceed the deposit without applicable negative balance protection; the separate 6.7% access payment is nonrefundable.

blckridge.com/research : published research. Notional funding and PAMM accounts explained in full, and answers to the questions this raises.

We use AI models to gather and aggregate source material and to help prepare each report. Read it with its cited sources, sample, methods and limitations, and send corrections through research methodology and corrections .

The client pays the selected strategy provider a one-time access payment of 6.7% of the agreed trading level. BlackRidge receives an introduction fee from that provider and does not receive broker compensation. About BlackRidge .

Minimum funded capital $25,000. The strategy, its live records, the broker, the selected strategy provider and the full commercial terms are presented on an introductory call and confirmed in writing before any payment or deposit.

This report is published for information only. It is not investment advice and does not take account of your circumstances. Trading leveraged instruments including CFDs carries substantial risk and is not suitable for all investors; you may lose the capital you fund and, depending on your broker's terms, may owe more than your deposit. Findings in this report may combine third-party evidence, historical calculations and illustrative scenarios; they are not verified live trading results of any strategy provider. Past performance does not indicate future results.

## Sources

- [01] [Open Source Asset Pricing data, October 2025 release](https://www.openassetpricing.com/data/). We pulled the monthly long-short returns of 212 predictors. The dataset provides original sample years, publication dates and liquidity-screened versions.
- [02] [Chen and Zimmermann, Open Source Cross-Sectional Asset Pricing](https://www.federalreserve.gov/econres/feds/files/2021-037pap.pdf). This paper outlines the methodology behind the replication.
- [03] [Kenneth French Data Library](https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html). This repository supplied the returns for value, size, momentum, profitability and investment factors through August 2026.
- [04] [McLean and Pontiff, Does Academic Research Destroy Stock Return Predictability? (Journal of Finance 2016)](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2156623). They documented the original 26% out-of-sample decay and 58% post-publication decay.
- [05] [Harvey, Liu and Zhu, ...and the Cross-Section of Expected Returns (Review of Financial Studies 2016)](https://www.nber.org/papers/w20592). The authors argue a newly discovered factor requires a t-statistic above 3.0.
- [06] [Bailey and Lopez de Prado, The Deflated Sharpe Ratio (Journal of Portfolio Management 2014)](https://www.davidhbailey.com/dhbpapers/deflated-sharpe.pdf). Their math yields the expected maximum Sharpe ratio from N trials with zero true skill.
- [07] [Bailey and Lopez de Prado, The Sharpe Ratio Efficient Frontier (Journal of Risk 2012)](https://www.davidhbailey.com/dhbpapers/sharpe-frontier.pdf). We used their formula for calculating the minimum track record length.
- [08] [Suhonen, Lennkh and Perez, Quantifying Backtest Overfitting in Alternative Beta Strategies (Journal of Portfolio Management 2017)](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2757113). They found a 73% median decay in live Sharpe ratios across 215 bank-promoted strategies.
- [09] [Wiecki, Campbell, Lent and Stauth, All That Glitters Is Not Gold (Journal of Investing 2016)](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2745220). Testing 888 trading algorithms revealed a backtest Sharpe ratio predicts out-of-sample results with an R-squared below 0.025.
- [10] [Jacobs and Mueller, Anomalies across the globe: Once public, no longer existent? (Journal of Financial Economics 2020)](https://www.sciencedirect.com/science/article/pii/S0304405X19301618). Looking at 241 anomalies in 39 markets showed the US is the only market with a reliable post-publication decline.
- [11] [SEC press release 2023-173 (September 2023)](https://www.sec.gov/newsroom/press-releases/2023-173). The regulator charged nine advisers a combined $850,000 for advertising hypothetical performance without required policies.
- [12] [SEC Division of Investment Management, Marketing Compliance FAQ](https://www.sec.gov/rules-regulations/staff-guidance/division-investment-management-frequently-asked-questions/marketing-compliance-frequently-asked-questions). Staff published their latest interpretations on 15 January 2026.

---

## Citation context

Of 212 published stock-return predictors, 208 enter the main comparison. Pooled returns fell 29.4% in the forward window and 55.3% after publication. A 50/50 blend of backtest and forward return cut forecast error by 14.5%.

Sample and method: historical returns reconstructed for backtest, forward and post-publication periods, 1926 to 2024.

Limits: these are not verified live trading results. A forward window after the original sample is a useful check, although authors may have seen those years.

Primary input: [Open Source Asset Pricing data, October 2025 release](https://www.openassetpricing.com/data/).

Stable permalink: [https://blckridge.com/research/forward-validation-20260928/#citation-context](https://blckridge.com/research/forward-validation-20260928/#citation-context).

## Read next

- [Artificial Intelligence in Quant Trading](/research/artificial-intelligence-quant-trading-2026/). For the practical AI workflow behind quantitative strategies, including evaluation and governance beyond return evidence.

