Sentiment Analysis
8 steps · one page
In short
Sentiment analysis attempts to measure the mood or positioning of market participants rather than the economics of a business or the geometry of a price.
Scope. This article describes what sentiment measures attempt to capture, how they are built, and why they are unusually hard to evaluate. It gives no readings to watch, no levels and no contrarian rules. No backtested result appears and no sentiment measure is described as reliable or unreliable.
It sits alongside the fundamental and technical traditions as a third thing to look at, and its underlying claim is different from either: not that a company is worth something, and not that price history predicts price, but that the state of other people's expectations is itself measurable and informative.
That claim is coherent and the measurement is genuinely difficult, and this article spends most of its length on why — because the difficulty is not usually explained, and it is what determines how much weight the whole field can bear.
The four families of measure
Surveys. Participants are asked what they expect. Bull-and-bear surveys of individual investors and of newsletter writers, institutional expectation surveys, and broader consumer-confidence measures all sit here. They measure what people say.
Positioning and flow. Options activity such as the ratio of puts to calls, short interest, fund flows, futures positioning by category of participant, and margin balances. They measure what people have done.
Price and volatility derived. Implied-volatility indices, the shape of the volatility surface, and credit spreads. These are prices rather than opinions, and they are the most objective of the four — though what they measure is the cost of insurance, which is not the same thing as fear.
Text. Automated reading of news, regulatory filings, earnings-call language and social media, scored for tone.
Worked example
The distinction that does the most work: saying versus doing. A survey response costs nothing. A position costs money and can be wrong expensively. The two frequently disagree, and when they do, the positioning data is the harder one to fake and the harder one to interpret. A large put position can mean a participant expects a decline, or that they hold the underlying and are insuring it — an identical observation with opposite meanings. Sentiment measures are not interchangeable, and treating a survey reading and a positioning reading as two views of the same quantity is the most common error in reading this material.
Why text-based measures are harder than they look
Scoring language for tone requires deciding which words are negative, and financial language is not ordinary language. A study published in the Journal of Finance in 2011 examined a large sample of annual filings and found that almost three-quarters of the words flagged as negative by a widely used general-purpose dictionary are not negative in a financial context. Words such as liability, cost, depreciation and capital carry no negative charge in a financial statement; they are simply the vocabulary of accounting.
The consequence is that a naive tone score of a financial document is substantially measuring how much accounting is in it. Purpose-built financial word lists exist because of this finding, and the episode is a useful general warning: an automated measure can be precise, reproducible and consistently wrong, and none of those first three properties will reveal the fourth.
The contrarian framing, and what it costs
Sentiment is most often read against the crowd — extreme optimism treated as a warning, extreme pessimism as its opposite. The reasoning is not silly: if participants are already fully committed, the marginal buying power that would push prices higher has been spent.
The difficulty is what that framing does to the claim's testability. Every reading becomes a statement about where the current level sits in a distribution, so the claim depends entirely on the definition of extreme — and extreme is defined by percentile of history, which changes as history accumulates. A reading that was in the top decile of its first ten years may be unremarkable across thirty.
The measurement problem, computed
Sentiment series are highly persistent: a bullish week is followed by another bullish week far more often than not. That persistence means consecutive readings are not independent observations, and the effective sample is much smaller than the raw count.
For a series with first-order autocorrelation ρ, the effective number of independent observations is approximately n(1 − ρ)/(1 + ρ). Twenty years of weekly readings is 1,040 raw observations.
| Autocorrelation of the series | Effective independent observations from 1,040 weekly readings | Independent top-decile episodes |
|---|---|---|
| 0.0 | 1,040 | 104 |
| 0.5 | 347 | 35 |
| 0.8 | 116 | 12 |
| 0.9 | 55 | 6 |
| 0.95 | 27 | 3 |
The second half of the calculation asks how many independent episodes would be needed to establish that outcomes after extreme readings differ from a coin flip.
| If the true hit rate after an extreme reading were | Independent episodes needed to detect it (80% power, 5% level) |
|---|---|
| 55% | 783 |
| 60% | 194 |
| 65% | 85 |
| 70% | 47 |
| 80% | 20 |
Worked example — two decades of data cannot settle the question. A weekly sentiment survey with autocorrelation of 0.8 yields roughly 116 effective observations over twenty years, and about 12 independent extreme episodes. Detecting even a 70% hit rate would require 47 episodes; detecting a 60% hit rate would require 194. With twelve episodes, only an effect of overwhelming size — around 80% — would be detectable at all. This cuts in both directions and that is the point. It means confident claims that a sentiment measure works are not supported by the available data, and equally that confident dismissals are not either. The honest description is that most sentiment claims are, at present sample sizes, not testable within a professional lifetime of data — which is a different and more uncomfortable statement than saying they are wrong. (Effective-sample formula and a two-sided binomial power calculation against a 50% null; both are standard constructions and the figures follow from them exactly.)
What has actually been established
The most influential result in the area is specific and worth stating precisely. A study published in the Journal of Finance in 2007 measured the tone of a daily newspaper column quantitatively and found that high media pessimism predicts downward pressure on prices followed by a reversion toward fundamentals, and that unusually high or unusually low pessimism predicts high trading volume. The author reported these results as consistent with theoretical models of noise and liquidity traders, and as inconsistent with media content being merely a proxy for new information about fundamental values, a proxy for volatility, or a sideshow unrelated to markets.
Worked example
What that result supports, and what it does not. It supports the proposition that measured sentiment is related to market behaviour, that the relationship has a plausible mechanism, and that it is not merely a restatement of news about fundamentals — which is a real finding and rules out the most dismissive position. It does not support the framing sentiment usually arrives in. The effects documented are short-horizon, concern the aggregate market rather than any individual security, include reversion as part of the finding rather than a durable move, and were measured on one specific text source over a defined period. The distance between that and a reader watching a sentiment gauge to decide something is very large.
Two structural problems that no amount of data fixes
Sentiment largely follows returns. People report feeling optimistic after prices have risen. A sentiment series is therefore substantially a lagged transformation of the price series, which places it in the same category as the indicators in the article on moving averages, RSI and MACD — a rearrangement of information already present rather than an addition to it. The burden on any sentiment measure is to show it adds something beyond trailing returns, and that is a much harder demonstration than showing it is correlated with what follows.
And the measure is part of the system it measures. A widely followed sentiment gauge is watched by participants who adjust for it, which changes the behaviour it was built to observe. This is the same reflexivity that closes the technical cluster in Support and Resistance, and it applies with more force here, because sentiment is by definition a measurement of other participants rather than of an asset.
Frequently asked
8 questions
What is sentiment analysis?
An attempt to measure the mood or positioning of market participants rather than the economics of a business or the history of a price. Its underlying claim is that the state of other people's expectations is itself measurable and informative.
What kinds of sentiment measures exist?
Four families: surveys, which measure what people say; positioning and flow data, which measure what people have done; price and volatility derived measures such as implied-volatility indices; and automated tone scoring of news, filings and social media.
Why do surveys and positioning data disagree?
Because a survey answer costs nothing and a position costs money. Positioning is harder to fake but harder to read — a large put position may express an expectation of decline or may be insurance on a holding, and the observation looks identical either way.
What is a contrarian indicator?
A sentiment reading interpreted against the crowd, on the reasoning that fully committed participants have already spent the buying power that would push prices further. The framing is coherent but it makes every claim depend on the definition of extreme, which is a percentile of history and therefore changes as history accumulates.
Why is sentiment so hard to test?
Because the series are highly persistent. Twenty years of weekly readings with autocorrelation of 0.8 gives roughly 116 effective observations and about 12 independent extreme episodes — while detecting even a 70% hit rate would require 47. Most sentiment claims are not testable within a professional lifetime of data, in either direction.
Do automated news sentiment scores work?
They work as measurements of something; the question is what. Nearly three-quarters of the words a widely used general-purpose dictionary flags as negative are not negative in a financial context, so a naive tone score of a filing is partly measuring how much accounting language it contains. Purpose-built financial word lists exist precisely because of this.
Has any sentiment effect been established?
Yes, narrowly. A 2007 Journal of Finance study found that high media pessimism predicts downward price pressure followed by reversion toward fundamentals, and that unusually high or low pessimism predicts high trading volume. The effects are short-horizon, concern the aggregate market rather than individual securities, and include the reversion as part of the finding.
Does MarketClue publish a sentiment indicator?
No. There is no sentiment score, mood index, contrarian flag or social tone reading, and no market condition is characterised as optimistic, complacent or fearful. Those are conclusions about a population of people presented as data about a security.
References
- Tetlock (2007) — Giving Content to Investor Sentiment: The Role of Media in the Stock Market, Journal of Finance 62(3) —
- Loughran and McDonald (2011) — When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks, Journal of Finance 66(1) —
- University of Notre Dame — the resulting financial sentiment word lists —
Educational and informational only — not investment advice, a recommendation, or an offer to buy or sell any security. Investing involves risk, including the possible loss of principal. Worked examples use fictional companies and figures.