BlackRidge
Systematic Strategy Evaluation

Backtests Are the Most Optimistic Estimates a Strategy Will Get

Of 212 published stock-return predictors, 208 enter the main comparison after exclusions. Historical returns are reconstructed for backtest, forward and post-publication periods. These are not verified live trading results.

September 2026
BlackRidge Research
Forward Validation | BlackRidge Research01
BlackRidge
The Overfitting Gap
02 / 13

Published predictors lost 29% of their backtest return in the forward window and 55% after publication

The backtest is the high-water mark of a strategy's evidence [04]. The decline starts before publication in these reconstructed returns. A forward window (sample end to publication, median 48 months) provides a useful check, although authors may have seen those years. Adding it to the backtest cut forecast error by about a seventh on average, a gain coming entirely from predictors published after 2006 (down 32.6%, with no improvement for older predictors) [01].

Forward window decay
−29.4%
Pooled return drop before publication.
Post-publication decay
−55.3%
Reconstructed post-publication return drop.
Significant predictors remaining
29.1%
Share retaining a t-statistic above two after publication.
Change in forecast error
−14.5%
Gain from judging the backtest and forward window together.
Backtest (in sample)
100.0%
Forward window
70.6%
After publication
44.7%
Average monthly long-short return for three periods (pooled, each predictor scaled by its own backtest average, 208 predictors, 1926-2024) [01].
Observation
More than half of the decay happens before investors can trade on the paper, making it the closest available measure of backtest optimism and an upper bound [04]. Pooled returns dropped 29.4% in the forward window and 55.3% after publication. Simple averages show a steeper decline: 0.70% a month in sample fell to 0.42% in the forward window and 0.30% after publication, drops of 40.5% and 56.9% respectively [01].
Interpretation
The performance drop starts before publication. Strong simulated results offer no safety margin. The strongest backtests lost the most out of sample and ended up matching the middle of the pack after publication.
Implication
Do not discard the backtest mean to rely solely on the forward period, as it predicts reconstructed post-publication returns no better than the original sample (R-squared 0.17 versus 0.19). The forecasting advantage comes strictly from judging both periods together (0.26).
Limitations
These returns are rebuilt on today's data, and the forward window is short (median 48 months against 336 in sample). Authors may have seen these forward years while writing, making this an upper bound on true overfitting. Finally, the reliable post-publication decline is a US phenomenon that may not travel globally [10].
Forward Validation | BlackRidge Research02
BlackRidge
Decay in Event Time
03 / 13

The edge peaks in the last five backtest years and loses a third in the first five years after

The peak just before the sample ends is what selection looks like. Researchers stop the sample where the result is strongest. Sometimes the effect was simply strongest then [01] [02].

Return as % of backtest
PredictorsBacktest average
50%100%−10−50+5+10+15+20
Average return of 208 predictors as a percentage of their in-sample mean, from 10 years before to 20 years after the original sample ends [01] [02]. Years from original sample end
Last five sample years
114.6%
Return relative to the full backtest average.
First five years after
68.4%
Average edge kept immediately after the sample.
Years six to ten after
48.5%
Continued deterioration as the strategy ages.
Interpretation
A strategy peaks right before its history ends. The subsequent drop exposes the cost of fitting rules to a known past. Live markets offer no such luxury.
Limitations
Four of the 212 predictors were excluded. Calendar years define the stage splits. Jacobs and Mueller find the decline reliably only in the US [10].
Forward Validation | BlackRidge Research03
BlackRidge
IN-SAMPLE SIGNIFICANCE
04 / 13

The strongest backtests lost the most and ended level with the middle

Investors often treat a massive in-sample t-statistic as proof of a durable edge. The data tell a different story. Across 208 predictors, the in-sample t-statistic explains about 3% (R-squared 0.03) of the variation in later returns [01].

Weakest fifth
−11.9%
Second fifth
67.4%
Middle fifth
59.6%
Fourth fifth
49.3%
Strongest fifth
37.3%
Share of in-sample return kept after publication, %, by quintile of in-sample t-statistic, 208 predictors [01].
GroupIn-sample t rangeBacktest return a monthAfter publication a monthShare kept
Weakest fifth1.0–2.30.42%−0.05%−11.9%
Second fifth2.3–2.90.50%0.34%67.4%
Middle fifth2.9–3.70.73%0.43%59.6%
Fourth fifth3.7–5.30.79%0.39%49.3%
Strongest fifth5.4–14.21.07%0.40%37.3%
Interpretation
Beyond moderate significance, a higher t-statistic buys no extra post-publication return: the strongest fifth ended level with the middle (0.40% against 0.43% a month), while only the weakest fifth lost money.
Implication
Stop paying a premium for paper significance. You should rank competing models by their out-of-sample decay instead.
Limitations
The weakest group turned negative partly because a t-statistic below 2.0 was never a strong claim. These are just averages. Those pooled numbers hide wide dispersion among the individual predictors inside each group.
Forward Validation | BlackRidge Research04
BlackRidge
Significance
05 / 13

Nine in ten predictors cleared t = 2 in the backtest, fewer than three in ten after publication

Most backtests clear traditional significance hurdles easily. Shorter track records after publication explain only a small fraction of the subsequent drop [01]. Harvey, Liu and Zhu have already argued that any new factor needs to clear a t-statistic of 3.0 [05].

t above 2 in sample
89.4%
t above 2 with no decay
74.8%
t above 2 after publication
29.1%
t above 3 in sample
58.7%
t above 3 with no decay
49.0%
t above 3 after publication
11.7%
Share of predictors, %, 208 in sample and 206 after publication [01].
t-statistic bandin sampleafter publication
below 00.0%15.0%
0 to 10.5%29.6%
1 to 210.1%26.2%
2 to 330.8%17.5%
above 358.7%11.7%
Observation
Backtested strategies project high confidence, with 89.4% clearing a t-statistic of 2 and 58.7% clearing t = 3. Reality sets in after publication when those figures collapse to 29.1% and 11.7% [01]. Fully 15.0% of predictors recorded a negative average return after publication [01].
Interpretation
The original results still contain information: across 206 predictors the in-sample mean explains an R-squared of 0.19 of post-publication returns [01]. Yet the collapse highlights that historical t-statistics dramatically overstate the actual signal. Initial certainty rarely survives contact with unseen data.
Limitations
The no-decay baseline assumes the return-to-risk ratio stays completely constant. Post-publication t-statistics are lower partly because arbitrage took the return, not only because the original authors overfit the data [04].
Forward Validation | BlackRidge Research05
BlackRidge
The Mathematics of Luck
06 / 13

The best of 100 useless ten-year backtests shows a Sharpe ratio of 0.8 on average

Every variant a quantitative team tries acts as a lottery ticket. The winning ticket's reported Sharpe ratio demands harsh judgment against the sheer number of tickets bought. Selection makes the best of many zero-skill results look like skill; it does not isolate forecasting ability [06].

Expected maximum Sharpe ratio
5-year backtest10-year backtest20-year backtest
00.511.51101001 000
Expected maximum Sharpe ratio of N independent zero-skill strategies, annualised, using the deflated Sharpe ratio formula [06]. Number of strategies tried (log scale)
Ten-year backtests
0.80
Expected outcome from one hundred independent zero-skill trials.
Five-year, 1,000 trials
1.46
Average top result when searching across one thousand variations.
Twenty-year backtests
0.57
Apparent performance generated entirely by luck across one hundred attempts.
Interpretation
Finding a good backtest tells you nothing if you ignore how many bad ones hit the cutting room floor. A high Sharpe ratio often measures the computing power spent searching for it, rather than any structural edge in the market. The shorter the sample period, the easier the noise masquerades as alpha.
Implication
Investment committees should discount presented performance by the estimated number of discarded variations. Demand a substantially longer track record when the underlying strategy relies on aggressive computational searches.
Limitations
Tested variations are rarely completely independent. The effective number of trials is smaller than the raw count, and estimating that true effective number remains the hardest part of the evaluation.
Forward Validation | BlackRidge Research06
BlackRidge
Track record length
07 / 13

A Sharpe ratio of 0.5 needs 11 years of evidence, 17 if it crashes like momentum

A six-month paper-trading period says almost nothing about a modest strategy [07]. Sample size dictates confidence. The shape of returns matters too: momentum-like skew stretches the requirement from 11 to 17 years [07] [03]. Any strategy carrying crash risk requires a massive observation window just to rule out luck.

Years of monthly returns needed (95% confidence)
Normal returnsMomentum-shaped returns
1310320.250.400.500.600.751.001.251.502.00
Minimum track record length, years, 95% one-sided, monthly data. Momentum skewness and kurtosis from 1927-2026 data [07] [03]. Annualised Sharpe ratio
Normal returns at 0.5 Sharpe
11.0 y
Years required to reject zero skill at 95% confidence.
Momentum shape at 0.5 Sharpe
17.3 y
Years required when returns carry fat left tails.
Momentum's long-run Sharpe (0.45)
20.2 y
Years required to reject zero skill at momentum's own long-run Sharpe ratio.
Interpretation
A standard two-year forward test offers less security than the industry assumes. With normal returns, a two-year record at 95% confidence is enough only for a true Sharpe ratio of about 1.25 or higher. Reaching that 1.25 mark takes 1.9 years. A 1.0 Sharpe needs 2.9 years.
Implication
Price the shape of returns into any evaluation of a live track record. Short histories offer zero protection against strategies with left-tail risk.
Limitations
This test only rejects zero skill, leaving the actual backtest level unconfirmed. High-frequency strategies run many independent bets and can shorten the time requirement. They never bypass it entirely.
Forward Validation | BlackRidge Research07
BlackRidge
Implementation
08 / 13

Dropping micro caps and weighting by size removes a quarter to a third of the backtest before any cost

Statistical overfitting only explains part of the performance gap. Academic backtests routinely assume equal weights across universes thick with tiny, illiquid equities that no institutional fund could efficiently trade [01]. The problem worsens with complexity, as bank-promoted alternative beta strategies suffer massive degradation when moved from paper to production [08].

Original backtest
0.7%
Excluding micro caps
0.5%
Value-weighted
0.5%
Post-publication original
0.3%
Post-publication value-weighted
0.1%
Average long-short return, % a month, before costs, same predictors [01].
Excluding micro caps
−26.2%
Return impact of dropping stocks below the NYSE 20th size percentile.
Value weighting
−34.6%
Return impact of weighting by market capitalization rather than equally.
Live Sharpe deterioration
−73%
Median drop for bank-promoted alternative beta strategies compared to their backtests [08].
Observation
Data from 888 trading algorithms reveals the backtest Sharpe ratio predicts out of sample performance with an R-squared below 0.025 [09]. A simulated track record explains practically nothing about future success.
Interpretation
Paper portfolios ignore market friction and capital capacity. Adjusting the rules to reflect realistic institutional mandates destroys a massive portion of the theoretical returns. You are pricing in execution constraints that the original authors simply ignored.
Limitations
These adjustments capture only structural portfolio rules, as transaction costs are not deducted anywhere in the primary dataset. Trading costs would further reduce these reconstructed returns. The alternative beta data also suffers from self selection, reflecting only the strategies banks chose to actively promote.
Forward Validation | BlackRidge Research08
BlackRidge
Famous Factors
09 / 13

Every classic factor earned less after publication. Value lost money for 14 years.

Even the best-known factors kept only a fraction of their backtest returns after decades of scrutiny [03]. Value earned −5.7% a year from 2007 to 2020, then surged to +7.9% a year from 2021 through August 2026 [03]. A forward test that covers one regime can mislead you in either direction.

Rolling 10-year annualised return (%)
Value (HML)MomentumProfitability (RMW)
−5%0%5%10%15%19801990200020102020
Rolling ten-year annualised return, %, June and December points, 1973-August 2026 [03].
FactorPublishedOriginal sampleIn sample, % a yearAfter publication, % a yearWorst decline after
Value (HML)19801962–19767.0%2.7%−57.8%
Size (SMB)19811926–19751.7%−0.4%−54.9%
Momentum19931964–19899.3%2.9%−57.8%
Profitability (RMW)20061977–20034.2%2.8%−25.9%
Investment (CMA)20081968–20035.4%0.0%−27.6%
Interpretation
A backtest provides an optimistic baseline that needs a haircut. You cannot dismiss it entirely. Across 206 predictors the in-sample mean still explains 0.19 of the variation in post-publication returns [01].
Limitations
French's factor definitions differ from the original papers. Returns are before costs.
Forward Validation | BlackRidge Research09
BlackRidge
Forward Validation
10 / 13

Adding a four-year forward window cut forecast error by 14.5%, all of it among papers published after 2006

The forward window alone offers no more predictive power than the original backtest. Averaging the two produces a clear improvement [01]. R-squared rises from 0.19 to 0.26, and a flat or negative forward window is a strong warning (19.3% of the backtest kept, against 51.1%).

Backtest only
0.19
Forward window only
0.17
Average of both
0.26
Share of variation in post-publication monthly return explained (R-squared), 206 predictors [01].
Change in forecast error
−14.5%
Drop in mean absolute forecast error when using the combined average.
Negative or flat forward window
19.3%
Share of the backtest return kept after publication by the 55 flagged predictors.
Positive forward window
51.1%
Share of the backtest return kept after publication by predictors that earned a profit.
Observation
We measure forecast error as the mean absolute error of the post-publication monthly mean in percentage points a month. Using a fixed 50/50 average confirms the gain beats mere shrinkage: the best hindsight shrinkage leaves the error essentially unchanged at 0.49, while the forward window drops it to 0.42.
Interpretation
Backtests simply measure what fit historical noise. The forward window tests whether the exact same rules survived their first contact with unseen data, stripping away the benefit of hindsight.
Implication
Freeze your rules to evaluate the backtest and forward window together. Treat a flat or negative forward window as a red flag that needs an explanation, not a reason to discard the strategy outright.
Limitations
The forward window was never a clean blind test, and an R-squared of 0.26 leaves three quarters of the variation unexplained. The error fell 32.6% for 81 predictors published after 2006 but 0.0% for 125 older ones, while 17 of the 55 flagged predictors still kept half their backtest.
Forward Validation | BlackRidge Research10
BlackRidge
Regime shifts
11 / 13

Momentum gave back 17.2% in July and August 2026, a two-month loss seen in 1.3% of windows since 1927

A regime turn is when a backtest calibrated on the previous regime is found out. The gap between the trailing year in June and in August shows how fast. French's momentum factor recorded exactly this shift [03].

Trailing 12-month return (%)
−20%−10%0%10%20%201620182020202220242026
trailing 12-month return, %, monthly, January 2016 to August 2026, [03]
Momentum decline
−17.2%
Combined loss for the factor in July and August 2026.
Historical rank
15 / 1,195
Worst loss rank out of 1,195 two-month windows since 1927.
Profitability return
+11.1%
Profitability (RMW) in July 2026, its third-best month out of 758 since 1963.
Interpretation
A factor can give back most of a year's gain in two months. A forward window shorter than one regime would never show this. Evaluators need long out-of-sample data to see how strategies behave when the environment changes.
Implication
A forward test should span more than one market environment. On the regulatory side, the SEC charged nine advisers $850,000 in September 2023 over hypothetical performance [11], and SEC staff keeps updating the Marketing Rule FAQ (latest 15 January 2026) [12].
Forward Validation | BlackRidge Research11
BlackRidge
Forward Validation
12 / 13

Forward data reduces forecast error and exposes backtest decay

We examined published stock predictors to measure performance on data outside the original sample. Returns in the post-publication period earned 55.3% less than the backtest. The decay begins immediately.

01
When evaluated on unchanged rules, the forward window earned 29.4% less than the original backtest.
02
The strongest fifth kept 37.3%.
03
The best of 100 zero-skill ten-year backtests shows 0.80 on average [06].
04
A Sharpe ratio of 0.5 needs 11.0 years of monthly data to be told apart from zero at 95% confidence [07]. Short records mislead.
05
Judged together with a forward window, forecast error fell 14.5% on average [01]. The benefit is not uniform across cohorts. All of it occurred among the 81 predictors published after 2006.
Forward Validation Summary
MeasureWhat it showsValue
Forward window relative returnReturn shortfall compared to the original sample−29.4%
Post-publication relative returnReturn shortfall after the research appears in print−55.3%
Strongest fifth retentionShare of performance maintained by the best backtests37.3%
Random backtest SharpeExpected ratio from multiple attempts with zero skill0.80
Years to confirm SharpeMinimum track record needed at standard confidence levels11.0 y
Change in forecast errorImprovement in predicting post-publication returns using forward data−14.5%

Before allocating capital to a systematic strategy, demand the number of variants tried and a frozen rule. Evaluate these alongside years of forward data judged together with the backtest.

Forward Validation | BlackRidge Research12
BlackRidge
Appendix
13 / 13

Selected sources

[03]
Kenneth French Data Library
https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html
This repository supplied the returns for value, size, momentum, profitability and investment factors through August 2026.
[11]
SEC press release 2023-173 (September 2023)
https://www.sec.gov/newsroom/press-releases/2023-173
The regulator charged nine advisers a combined $850,000 for advertising hypothetical performance without required policies.

This document is research, not investment advice. Returns appear gross of costs and taxes. Past decay does not place a limit on future decay.

Forward Validation | BlackRidge Research13

Citation context

Of 212 published stock-return predictors, 208 enter the main comparison. Pooled returns fell 29.4% in the forward window and 55.3% after publication. A 50/50 blend of backtest and forward return cut forecast error by 14.5%.

Sample and method: historical returns reconstructed for backtest, forward and post-publication periods, 1926 to 2024.

Limits: these are not verified live trading results. A forward window after the original sample is a useful check, although authors may have seen those years.

Primary input: Open Source Asset Pricing data, October 2025 release.

Stable permalink: https://blckridge.com/research/forward-validation-20260928/#citation-context.

Cite this report

When you use a result that another publication established, cite that original work; it is linked in Sources. Cite this report for our synthesis, explanation or an identified recalculation. No link is required in return.

Author
BlackRidge
Published
28 September 2026
Stable link
https://blckridge.com/research/forward-validation-20260928/

Download BibTeXDownload referenceCitation and reuse

Keep up with the research.

Choose optional research updates, or learn how the account structure works.

Subscribe to research updatesNext: Start here
BlackRidge
About this research
—
—
About BlackRidge

Independent research,
introductions clearly defined.

BlackRidge is an independent research bureau introducing private investors to quantitative traders through multiple strategy providers. We publish quantitative research. Investors access strategies through PAMM accounts at the broker. We are not a fund or broker and never hold client money.

01

Your funds stay at the broker. The account is opened in your own name. BlackRidge never receives, holds, or has withdrawal rights over client capital.

02

No fee on your profits. A one-time access payment of 6.7% of the agreed trading level, paid to the selected strategy provider. No management fee, no performance fee, no profit share, in any year.

03

You fund the risk, not the exposure. Allocations are notionally funded: you agree a trading level and fund the margin and drawdown allowance behind it. Trading losses can exceed the deposit without applicable negative balance protection; the separate 6.7% access payment is nonrefundable.

Continue reading

blckridge.com/research: published research. Notional funding and PAMM accounts explained in full, and answers to the questions this raises.

We use AI models to gather and aggregate source material and to help prepare each report. Read it with its cited sources, sample, methods and limitations, and send corrections through research methodology and corrections.

The client pays the selected strategy provider a one-time access payment of 6.7% of the agreed trading level. BlackRidge receives an introduction fee from that provider and does not receive broker compensation. About BlackRidge.

Minimum funded capital $25,000. The strategy, its live records, the broker, the selected strategy provider and the full commercial terms are presented on an introductory call and confirmed in writing before any payment or deposit.

This report is published for information only. It is not investment advice and does not take account of your circumstances. Trading leveraged instruments including CFDs carries substantial risk and is not suitable for all investors; you may lose the capital you fund and, depending on your broker's terms, may owe more than your deposit. Findings in this report may combine third-party evidence, historical calculations and illustrative scenarios; they are not verified live trading results of any strategy provider. Past performance does not indicate future results.