Chapter 5 of 5 · 11 min

Is a Sharpe ratio significant? t = Sharpe × √years

A Sharpe ratio is a mean over a volatility, so testing it is testing the mean, and the t-statistic is the annual Sharpe times the square root of the years. That one line explains why track records are so hard to judge.

By the end of this chapter you can
  • Turn a Sharpe ratio and a track record into a t-statistic
  • Solve for the years a Sharpe needs, or the Sharpe a track record needs, to clear a hurdle
  • Annualize a daily Sharpe with √252 and say why sampling more often does not help
  • Count the false positives a search over many strategies produces
1

The intuition

A manager shows three years with an annual Sharpe of 1.0. The t-statistic on the mean excess return is the Sharpe times the square root of the years: 1.0 × √3 = 1.73, short of the usual 1.96. A genuinely good strategy with a Sharpe of one needs about four years before the conventional test believes it; a Sharpe of 0.5 needs fifteen. Significance is a property of the evidence, not of the strategy.

And the evidence is easy to manufacture. A researcher who tries twenty variants with no edge at all has a 40% chance that at least one clears 1.96 by luck, and should expect half of one to. That is why people who study factor discovery argue for a hurdle of 3 rather than 1.96, and why the first question about any backtest is how many were run.

The key idea

t ≈ SR_annual × √years. Years needed for a hurdle t*: (t* ÷ SR)². Minimum Sharpe for t* over Y years: t* ÷ √Y. SR_annual = SR_daily × √252. P(at least one of N no-edge strategies shows t > 1.96) = 1 − 0.975^N; expected number = 0.025 × N.

2

Why it works

  • The conventions here: independent, normally distributed returns with constant mean and volatility. The Sharpe's own estimation error and fat tails are ignored. 'Significant' means t above 1.96, the conventional 5% two-sided bar, or 3.0, the stricter bar proposed for data-mined strategies. No-edge strategies each clear 1.96 by luck with probability 2.5%, one tail.
  • Why t = SR × √years. t is the mean over its standard error, mean ÷ (σ ÷ √n). With annual figures, mean ÷ σ is the Sharpe and n is the years.
  • Frequency does not help. Daily data has 252 times the observations, but the daily Sharpe is the annual one over √252, and the two cancel. Only calendar time adds information about the mean.
  • Leverage does not help either. Doubling the position doubles the mean and the volatility: same Sharpe, same t. Leverage changes the size of the bet, not the strength of the evidence.
  • Run it backwards. Years for a hurdle: (t* ÷ SR)². The Sharpe a track record needs: t* ÷ √years. A five-year record needs a Sharpe of 0.88 to clear 1.96.
  • Multiple testing. Each of N independent no-edge strategies clears the bar with probability 2.5%. P(at least one) = 1 − 0.975^N, and the expected count is 0.025 N. A Bonferroni correction tests each at 5% ÷ N; either way, report N.
Annual Sharpe 1.0; a 3-year track record; a search over 20 variants
t-statistic: 1.0 × √31.73, below 1.96
Years needed for 1.96 at a Sharpe of 1.0: (1.96 ÷ 1.0)²3.84 years
Minimum Sharpe over 5 years for 1.96: 1.96 ÷ √50.88
Daily Sharpe of 0.0630 annualized: × √2521.00
20 no-edge variants: P(at least one t > 1.96) = 1 − 0.975²⁰39.7%
Expected false positives: 20 × 2.5%0.5

At the stricter hurdle of 3.0 the Sharpe-1.0 strategy needs 9 years, and a five-year record needs a Sharpe of 1.34.

3

The formulas

t ≈ SR_annual × √years

The mean over its standard error, in annual units.

Years needed = (t* ÷ SR)²; minimum SR = t* ÷ √years

The same line, solved for the years or for the Sharpe.

SR_annual = SR_daily × √252

Mean scales with 252, volatility with √252.

P(at least one false positive in N) = 1 − 0.975^N; expected = 0.025 N

Independent no-edge strategies, 2.5% each, one tail.

4

Worked example

Sharpe times the square root of the years. The follow-up asks why the years matter but the number of daily observations does not.

Drawing the numbers…
5

See it move

Same manager. Change the Sharpe ratio, the years of track record, the hurdle, and how many variants a researcher tried.

Drawing the numbers…
Try this
  • Raise the Sharpe. The t-statistic rises and the years needed fall as the square.
  • Add years. The t-statistic rises with the root, and the minimum Sharpe the record needs falls.
  • Switch to the stricter hurdle. The years needed and the minimum Sharpe both rise; the t-statistic itself does not change.
  • Try more variants. The chance of a false positive and the expected count both rise.
6

Run it backwards

A strategy has a stated true Sharpe. How many years before its t-statistic reaches the hurdle?

Drawing the numbers…

Solve SR × √years = t* for the years: (t* ÷ SR)². A Sharpe of 0.5 needs four times the years a Sharpe of 1.0 does.

The follow-up is why the stricter hurdle exists: most strategies that reach a desk are survivors of many that were tried.

7

Traps

Believing more frequent data makes a Sharpe more significant.
Daily Sharpe is annual ÷ √252 and the observations are × 252; they cancel. Only calendar time counts.
Asking for the volatility to judge a Sharpe.
The Sharpe already divides by it. Leverage changes the bet, not the evidence.
Treating any Sharpe above one as significant.
t = SR × √years. A Sharpe of one needs about four years at 1.96.
Ignoring how many strategies were tried.
Twenty no-edge variants give a 40% chance of a fluke at 1.96. Raise the bar with the number of trials, and report the number.
Taking t as proof under any return distribution.
The line assumes independent normal returns. Fat tails and autocorrelation make the evidence weaker than t says.
8

Say it in the interview

The interviewer asks

A candidate shows a Sharpe of 1.5 over two years of live trading. Convinced?

Say yours out loud first, then compare.
9

Check yourself

5 fresh questions, with new numbers. Answer each one correctly to finish the chapter. Get one wrong and you will see the full working, then you can try it again with new numbers.

Answers within 1% are marked right. Type the number; $, %, x and M are fine. First tries count toward Learned: the topic is Learned once every chapter is done and 75% of first tries were right.

0 of 5
Drawing your questions…
Remember
  • t ≈ Sharpe × √years. A Sharpe of one needs about four years at 1.96.
  • Years needed = (t* ÷ SR)²; minimum Sharpe = t* ÷ √years.
  • Sampling more often and levering up do nothing for significance.
  • P(a fluke in N tries) = 1 − 0.975^N. Ask how many were tried.