Benchmarking and Measuring Performance
6 steps · one page
In short
A return means nothing on its own. It means something compared with what else was available, and choosing the comparison is most of the work.
Scope. MarketClue selects no benchmark on a reader's behalf, publishes no performance rating, and does not describe any result as good or bad. This article explains what a benchmark has to be to mean anything, and why two correct measures of the same portfolio over the same period can differ by more than twenty percentage points a year. All figures are worked illustrations and are not forecasts.
What a benchmark has to be
Five properties, and a benchmark failing any of them will produce a verdict rather than a measurement.
Specified in advance. A benchmark chosen after the period is a benchmark chosen because of the answer it gives.
Investable. If the comparison could not actually have been held — because it excludes costs, or contains things that cannot be bought — then beating it establishes nothing.
Measurable and unambiguous. Its constituents and weights must be knowable, so that the comparison can be reconstructed by someone else.
Appropriate. It has to reflect the same kind of exposure. Comparing a portfolio to a benchmark carrying different risk measures the difference in risk and calls it skill.
Agreed in advance by whoever is being measured. Otherwise the choice becomes an argument conducted after the fact.
Worked example
Why the choice does so much work. Every portfolio outperforms some benchmark and underperforms another. The selection is therefore not a technical preliminary to the measurement — in many cases it is the measurement, and the person choosing has a great deal of latitude before anyone has computed anything. This is the same structural problem as the confidence level in risk measurement and the construction choices in the cost of capital: a decision made before the arithmetic determines the arithmetic's conclusion, and it is invisible in the result.
Two correct returns for the same portfolio
This is the finding most worth taking from the article, because both numbers are right and they answer different questions.
Time-weighted return measures what the investments did, removing the effect of money going in and out. Money-weighted return — the internal rate of return — measures what happened to the money, including the timing of contributions.
Consider a portfolio over two years: $10,000 invested, up 50% in year one to $15,000; then $85,000 added, taking it to $100,000; then down 20% in year two, ending at $80,000.
| Measure | Result | What it answers |
|---|---|---|
| Time-weighted | +9.54% a year | How did the investments perform? |
| Money-weighted | −14.49% a year | What happened to the money? |
Worked example — the same two years, twenty-four percentage points apart. Time-weighted, the portfolio returned +9.54% a year. Money-weighted, it returned −14.49% a year. The gap is 24.0 percentage points annually, and neither figure is wrong. The investments genuinely did well: up 50% then down 20% compounds to +20% over two years. The money genuinely did badly: most of it arrived immediately before the decline, so the large balance experienced the loss and the small balance experienced the gain. Which measure is appropriate depends entirely on who controlled the timing. A manager who did not choose when contributions arrived should be measured time-weighted; the investor whose money it is experienced the money-weighted result and has $80,000. The number quoted in any performance discussion should always be identified as one or the other, and a figure presented without that label is not interpretable.
Tracking error, and what it does and does not say
Tracking error is the variability of the difference between a portfolio and its benchmark. It measures how closely one follows the other, and it is symmetric — it treats outperformance and underperformance identically.
| Tracking error | Roughly two-thirds of periods land within | Roughly 95% within |
|---|---|---|
| 0.5% | 0.5 pp of the benchmark | 1.0 pp |
| 2.0% | 2.0 pp | 4.0 pp |
| 5.0% | 5.0 pp | 10.0 pp |
A low tracking error says a portfolio behaves like its benchmark. It says nothing about whether either did well, and a portfolio can track a falling benchmark impeccably.
Three ways performance comparisons mislead
Survivorship. Comparisons against a peer group compare against the funds that still exist. Those that closed are not in the average, and they did not close because they were doing well. This is the same problem Pillar 22 covers, arriving here through the denominator.
Period selection. A start date is a choice, and moving it by a year can reverse a conclusion. A comparison whose start date was chosen after the data existed is not evidence.
Risk mismatch. A portfolio taking more risk than its benchmark should outperform on average, and does so without any skill being involved. Comparing returns without comparing variability is comparing two different things and reporting the difference as performance.
Frequently asked
8 questions
Why does a return need a benchmark?
Because a return means nothing on its own — only compared with what else was available. Every portfolio outperforms some benchmark and underperforms another, so choosing the comparison is most of the work.
What makes a benchmark valid?
Five properties: specified in advance, investable, measurable and unambiguous, appropriate to the same kind of exposure, and agreed by whoever is being measured. Failing any of them produces a verdict rather than a measurement.
What is the difference between time-weighted and money-weighted return?
Time-weighted measures what the investments did, removing the effect of money moving in and out. Money-weighted — the internal rate of return — measures what happened to the money, including contribution timing.
How far apart can they be?
Very far. In the worked example the same two years return +9.54% a year time-weighted and −14.49% money-weighted — a gap of 24.0 percentage points — because most of the money arrived immediately before the decline.
Which one is correct?
Both. Which is appropriate depends on who controlled the timing: a manager who did not choose when contributions arrived should be measured time-weighted, while the investor whose money it is experienced the money-weighted result.
What is tracking error?
The variability of the difference between a portfolio and its benchmark. At 2.0%, roughly two-thirds of periods land within 2.0 percentage points of the benchmark and about 95% within 4.0.
Does low tracking error mean good performance?
No. It means the portfolio behaves like its benchmark and says nothing about whether either did well. A portfolio can track a falling benchmark impeccably.
What are the main ways comparisons mislead?
Survivorship, since peer groups contain only the funds that still exist and closures were not caused by success; period selection, where moving a start date by a year can reverse a conclusion; and risk mismatch, where a portfolio taking more risk outperforms on average with no skill involved.
References
- CFA Institute — Global Investment Performance Standards (GIPS): time-weighted and money-weighted return conventions and benchmark requirements —
- Investor.gov (SEC) — Benchmark —
- Pedersen (2018) — Sharpening the Arithmetic of Active Management, Financial Analysts Journal 74(1) (why beating a benchmark is a zero-sum comparison after costs) —
Educational and informational only — not investment advice, a recommendation, or an offer to buy or sell any security. Investing involves risk, including the possible loss of principal. Worked examples use fictional companies and figures.