Skip to content
MarketClueLearn

Quantitative Analysis: The Concept and Its Hazards

Advanced13 min readLesson 14 of 19

10 steps · one page

In short

Quantitative analysis means replacing case-by-case judgement with rules stated in advance and applied uniformly.

Scope. This article explains what quantitative analysis is as an approach and what statistical hazards are specific to applying it to markets. It does not teach strategy construction. No factor, rule, model specification or parameter is offered, no backtested result appears, and nothing here is a recipe. It is the counterpart to Pillar 22's article on multiple testing, which states the underlying statistical argument in general form.

A quantitative analyst specifies what will be measured, over what universe, on what data, and then applies that specification everywhere without exception — including to the cases where it produces an uncomfortable answer.

That definition contains the approach's real contribution and its central hazard at the same time. The contribution is that a rule stated in advance can be checked, which a judgement cannot. The hazard is that a rule can be discovered by searching rather than reasoned to, and a rule found by searching looks identical to one that was reasoned to.

The term covers three activities with very different standing

Measurement and risk — computing exposures, volatilities, correlations and concentrations across a portfolio. This is largely descriptive statistics applied carefully, and it is the best-founded part of the field.

Execution and cost — modelling how orders move prices, how much trading costs, and how to schedule it. This is an engineering problem with observable feedback, and practitioners can tell relatively quickly whether they are right.

Return prediction — identifying characteristics associated with future returns. This is the contested part, it is what most people mean by the word, and almost everything difficult in this article is about it.

Conflating the three flatters the last by association with the first two. A firm can be excellent at measuring risk and no better than chance at predicting returns, and the two capabilities have almost nothing to do with each other.

What the approach genuinely does better

Consistency. The same criteria are applied to the four-hundredth company as to the first, without fatigue, boredom or a prior opinion about the management.

Breadth. A rule can examine a universe no analyst could read through.

Auditability. The decision can be reconstructed afterwards. Why a discretionary decision was made is often not recoverable even by the person who made it.

Falsifiability, which is the real one. A rule stated precisely can be shown to be wrong. A judgement expressed as the shares look attractive here cannot be, because it has no content that could fail. The discipline of stating a claim precisely enough to be refuted is the approach's actual intellectual contribution, and it survives every criticism below.

Hazard one: fit arrives without any relationship being present

Adding explanatory variables improves in-sample fit mechanically, whether or not those variables have anything to do with the outcome. The expected in-sample R-squared from fitting k predictors to n observations of pure noise is approximately k divided by (n − 1) — and the fit is entirely spurious by construction.

The table below fits random predictors to a random target over 250 observations, then applies the fitted model to fresh data.

Predictors fitted (250 observations)In-sample R-squaredTheoretical k / (n − 1)Out-of-sample R-squared
50.0190.020−0.03
100.0390.040−0.05
200.0820.080−0.10
500.2000.201−0.26
1000.4040.402−0.69

Worked example — a 40% fit from nothing at all. One hundred meaningless predictors on 250 observations of noise produce an in-sample R-squared of 0.40, matching the theoretical value almost exactly. Applied to fresh data from the identical process, the same model produces an R-squared of −0.69 — meaning it is substantially worse than having used the sample mean. The in-sample figure is not evidence that has been overstated; it is not evidence at all. It is a deterministic consequence of the number of parameters relative to the number of observations, and it would appear identically if the predictors were the results of coin flips.

Hazard two: the best of many is not the best, it is the luckiest

This is the multiple-testing problem, and in a quantitative setting it operates at industrial scale. If N candidate strategies are evaluated and the best is selected, the test statistic of the winner is the maximum of N draws, and the maximum of many draws is large even when every individual draw is meaningless.

Candidates evaluatedExpected t-statistic of the best one, on pure noiseShare of runs where the best exceeds t = 2.0
10.022.7%
101.5219.0%
1002.5189.3%
1,0003.24100.0%
10,0003.86100.0%
100,0004.39100.0%

Worked example — where the conventional significance bar goes to die. Search 1,000 candidates on data with no structure in it and the best one clears the conventional t = 2.0 threshold in 100% of runs, with an expected t-statistic of 3.24. The convergence with the published literature is the striking part. A 2016 study in the Review of Financial Studies catalogued 316 factors proposed to explain the cross-section of returns, built a multiple-testing framework around that search, and concluded that a newly proposed factor should have to clear a t-statistic above 3.0 rather than the customary 2.0 — and that most claimed findings in financial economics are likely to be false. An arbitrary simulation of a thousand noise trials lands at 3.24; a careful accounting of the actual published search lands at 3.0. They are measuring the same thing from opposite directions. (Simulated values; the expected maximum of N standard normal draws is a known quantity and the figures reproduce it to the second decimal.)

Hazard three: the record is far shorter than it looks

A track record is a sample, and its length governs what can be concluded from it. The standard error of an annualised Sharpe ratio is approximately the reciprocal of the square root of the number of years observed.

Length of recordApproximate standard error95% interval around an observed Sharpe of 1.0
3 years0.58−0.13 to 2.13
5 years0.450.12 to 1.88
10 years0.320.38 to 1.62
20 years0.220.56 to 1.44
Worked example

Worked example

Worked example — five years settles almost nothing. A five-year record showing a Sharpe ratio of 1.0 is consistent with a true value anywhere from roughly 0.1 to 1.9. The lower end is close to indistinguishable from a passive holding; the upper end would be exceptional. Even twenty years leaves an interval from 0.56 to 1.44. This is not a criticism of any particular record — it is the reason why the phrase a strong five-year track record carries less information than it appears to, and why the length of a sample deserves as much attention as its result.

Hazard four: the subject changes when it is studied

A physical constant does not react to being measured. A market does. When a characteristic associated with returns is published, capital moves toward it, and the movement of capital is precisely what would remove the effect. The literature documents this decay after publication, which means a finding can be entirely genuine at the time of discovery and worthless by the time it is widely known — without anyone having done anything wrong.

The consequence for evaluating any quantitative claim is severe. An out-of-sample test on data from before publication is not really out of sample in the sense that matters, because the world had not yet adapted. The only fully honest test runs forward from the moment the claim was made, and that test takes as long as the horizon of the claim.

Hazard five: the data is not what it appears to be

Survivorship. A universe assembled today contains the companies that still exist. Testing a rule on it asks how the rule would have done on the survivors, which is not a question anyone needed answered.

Look-ahead. Financial statement data is dated to the period it describes, not the date it became public. Using it on the earlier date embeds knowledge nobody had.

Restatement. Databases hold the current version of a figure, not the version originally published. A rule tested on restated figures was tested on information that did not exist at the time.

Point-in-time data solves all three and is expensive, which is why it is frequently not used.

Worked example

Worked example

What careful practice looks like, described rather than prescribed. The disciplines that address the hazards above are well established: state the hypothesis and its economic rationale before looking at the data, since a mechanism proposed afterwards is a description of the result rather than a test of it; count and report every specification examined, not only the surviving one, because the count is what determines the appropriate bar; hold out data and touch it once, given that a holdout consulted repeatedly is simply more in-sample data; and prefer few parameters, because the first table in this article shows what parameters buy in fit and what they cost out of sample. This is a description of how the field polices itself, not an instruction set. MarketClue does not construct strategies and does not teach their construction.

The honest position

Quantitative analysis is not a superior method or an inferior one; it is a method whose failures are legible. A discretionary analyst who has been wrong for ten years can usually explain why the ten years were unrepresentative. A quantitative claim written down in advance cannot. That legibility is why the criticisms in this article exist in such detail — the field generated them itself, about itself, in its own journals. No comparable literature exists auditing the reliability of judgement, not because judgement is more reliable but because it is harder to hold still long enough to be examined.

Frequently asked

8 questions

What is quantitative analysis in investing?

Replacing case-by-case judgement with rules stated in advance and applied uniformly across a universe. The term covers risk measurement, execution modelling and return prediction, which have very different standing — the first two are far better founded than the third.

What is overfitting, concretely?

Fitting a model to noise. One hundred meaningless predictors on 250 observations produce an in-sample R-squared of about 0.40 and an out-of-sample R-squared of about −0.69. The in-sample fit is a deterministic function of the number of parameters, not evidence of a relationship.

Why does testing many strategies undermine the result?

Because the winner is the maximum of many draws. Searching 1,000 candidates on structureless data produces a best t-statistic above the conventional 2.0 bar in 100% of runs, with an expected value of 3.24. The bar was designed for a single test.

What is the factor zoo?

Shorthand for the accumulation of published characteristics claimed to explain returns. A 2016 study catalogued 316 of them, argued that the sheer size of the search means the usual significance bar is inappropriate, and proposed a t-statistic hurdle above 3.0 instead of 2.0.

How long a track record is needed to conclude anything?

Longer than is usually available. A five-year record showing a Sharpe ratio of 1.0 is consistent with a true value from roughly 0.1 to 1.9; twenty years narrows it only to 0.56 to 1.44.

Why does a finding stop working after it is published?

Because capital moves toward it, and that movement is what removes the effect. A result can be genuine at discovery and worthless once widely known, with no error by anyone. It also means pre-publication out-of-sample tests are not out of sample in the way that matters.

What are survivorship and look-ahead bias?

Survivorship: a universe assembled today contains only the companies that still exist, so a test on it measures performance among survivors. Look-ahead: using financial data on the date it describes rather than the date it became public, which embeds knowledge nobody had. Point-in-time data avoids both and costs more.

Is quantitative analysis better than judgement?

It is not better or worse — its failures are simply legible. A rule written in advance can be shown to be wrong; a judgement usually cannot. That is why the field has produced such a detailed literature criticising itself, and why no comparable audit exists for discretionary decisions.

References

Educational and informational only — not investment advice, a recommendation, or an offer to buy or sell any security. Investing involves risk, including the possible loss of principal. Worked examples use fictional companies and figures.