The Emperor Has No Benchmark
Alpha is a function of two things. The field keeps pretending it is a function of one.
Somewhere in America right now, a chief financial officer is deciding whether to build a factory. Her analyst has computed the firm’s cost of equity — the return shareholders demand for tying up their money in something as undignified as a cheap, boring, profitable company. Using the standard model, the one taught in every MBA program on Earth, the answer comes back around five percentage points over the risk-free rate. Using a newer version of the same model — same data, same market, same firm — the answer comes back around seven.
Two percentage points of discount rate is not a rounding error. Compounded over the life of a factory, it is the difference between building it and not building it. Between the jobs existing and not existing. And the two numbers do not differ because anyone made a mistake. They differ because of a modeling choice — which definition of “risk” the analyst subtracted — made years earlier by academics the CFO will never meet, for reasons she will never read.
She thinks she is estimating a number. She is choosing one, and calling the choice an estimate.
Last time, I wrote about a trio of papers showing that the deep structural parameters of finance — demand elasticities, the machinery under the hood — are not measurements at all. I ended by promising a third essay about the strangest paper of the lot: one that set out to defend academic finance and ended up indicting it harder than the demolitions. That paper is Gormsen and Jensen’s “Conditional Risk,” and this is that essay.
The claim, stated up front so you can hold me to it: alpha — outperformance, skill, the thing this entire industry is organized around measuring — is a function of two arguments. Alpha of a strategy, against a benchmark. There is no benchmark-free alpha. Yet the field talks, publishes, and sells as if the function had one argument. (Strictly there are more suppressed arguments still — sample, estimator, information set, costs — but the benchmark is the biggest one hiding in plain sight.) Every fight about whether value works, whether momentum is real, whether your fund manager earns his fee, is partly a fight about what was subtracted before the residual got called skill. And the second argument is not supplied by nature. It has to be chosen, defended, and disclosed. The field is diligent about the first, sporadic about the second, and allergic to the third.
The Sixty-Second Edifice
To see why this matters, you need sixty seconds of history. I will be fair to the edifice — fairer than it deserves — because you cannot appreciate the demolition until you have admired the architecture.
In 1952 Harry Markowitz had the genuinely great idea: do not judge an investment by itself; judge it by what it does to your portfolio. Individual wiggles cancel; co-movement does not. Diversification is the closest thing markets offer to a free lunch, and Markowitz turned that observation into mathematics. In 1964 William Sharpe pushed the logic to its equilibrium conclusion: if everyone does the Markowitz math (and agrees on the inputs, and can borrow freely, and trades without friction), everyone holds the same portfolio of risky assets — the market itself — and exactly one quantity determines any asset’s expected return: beta, its co-movement with that market. Risk you could have diversified away pays you nothing, because why would it? This is the Capital Asset Pricing Model. Nobel Prizes followed — Markowitz and Sharpe shared the 1990 award with Merton Miller — and it deserved them, as theory.
In 1968 Michael Jensen did the operational thing: run the regression. Whatever return the benchmark explains is beta. Whatever is left over — the intercept — gets a Greek letter, and the Greek letter gets a priesthood. Alpha. Every alpha you have ever been quoted descends from that intercept.
Then in 1977 Richard Roll pointed out the crack in the foundation, and the field has been stepping over it for fifty years. The CAPM’s “market portfolio” means everything — every stock, every bond, every house, every farm, every piece of human capital, your uncle’s vintage guitars. It is unobservable. Not hard to observe. Unobservable. So every test of the CAPM ever run substitutes a proxy — in the academy a broad CRSP-style composite, in practice the S&P 500 — and is therefore, Roll proved, a joint test of the CAPM and of the proxy’s adequacy as the market portfolio. A rejection cannot tell you which one failed; an acceptance certifies nothing about the true market portfolio either. The profession read the paper, agreed it was devastating, and went right on running the regressions.
Hold Roll’s point in your hand. This essay is what happens when you refuse to put it down. The benchmark enters twice — once as a model (which risk adjustment?) and once as a proxy (which portfolio stands in for “the market”?). We will now watch both choices fail to be facts, in public, one at a time.
Oh — and whether anyone can actually estimate the inputs to Markowitz’s beautiful machine is a separate question, with an answer so funny it gets its own essay later in this series. Patience.
Act One: The Benchmark That Will Not Hold Still
Here is the most-studied number in empirical finance: the alpha of value — the return to buying cheap companies and shorting expensive ones, beyond what the CAPM says the risk deserves. Thousands of papers. Careers, Nobel arguments, trillion-dollar product categories. One number.
Except it is not one number, because betas move. A company’s market sensitivity is not a constant of nature; it drifts as the firm changes leverage and business mix, and as the market changes composition around it. The original CAPM ignores this — it is unconditional, one beta for all seasons. The conditional CAPM lets beta vary with time. Think of it as a camera: the unconditional model shoots a fifty-year film at a single exposure setting; the conditional model has auto-exposure. Same camera. Different photograph.
You would think the field could settle which photograph to use. It has been trying for thirty years. Watch.
In 2006, Lewellen and Nagel published a paper whose title is its thesis: “The Conditional CAPM Does Not Explain Asset-Pricing Anomalies.” Their argument is a plausibility bound, and it is elegant: for time-varying betas to rescue the CAPM, betas would have to swing with the market premium implausibly hard — several times harder than anything in the data. They measured betas directly, in short windows, no state variables required. Betas do move, meaningfully.[^1] But not in the right way, not remotely enough, and momentum’s conditional alpha stays enormous — on the order of one percent per month. Verdict: conditioning cannot save you. The anomalies are real. (Savor that word while it is still legal. An anomaly is a deviation from normal — and the “normal” here is an efficient market, which is to say the anomaly is itself measured against the very thing that does not exist. Even the field’s vocabulary is benchmarked against a fiction. Onward.)
In 2011, Boguth, Carlson, Fisher, and Simutin found the hole. The prior literature — Lewellen-Nagel’s bound included — had focused on betas co-moving with the market’s premium and largely waved off betas co-moving with the market’s volatility, a channel they showed is plausibly two to ten times larger. They also named a second sin, “overconditioning”: estimating conditional risk using information no investor could have had at the time, which can inflate measured alphas by a factor of two and a half. Run the estimation properly, with instruments, and momentum’s alpha drops twenty to forty percent. Not zero. Not one percent a month either. Verdict: the anomalies are smaller than you think, and the size of your alpha depends on your assumptions about what investors knew and when — which is to say, on your benchmark.
That same year Nagel himself, with Ken Singleton, built the statistically optimal version of conditional-model testing — different machinery, aimed at the era’s fashionable consumption-based models — and rejected those too, by yet another route. Better conditioning machinery did not produce an acquittal. It produced a fourth verdict. This is what careful people look like when the object they are chasing does not exist.
Then in 2024, Gormsen and Jensen arrived with the paper that was supposed to be the defense. Instead of estimating betas — rolling windows, state variables, all the machinery the previous combatants fought over — they built a tradable portfolio designed to harvest the return to conditional risk, both flavors, premium-timing and volatility-timing. To build it they still had to estimate the market’s conditional premium and its conditional variance. They did not eliminate estimation. They relocated it — into the benchmark itself, which is where this essay keeps telling you the bodies are buried. Add the factor to the benchmark and see what happens to the alphas.
What happens is a massacre. Value’s alpha in the full US sample, 1964 to 2022: 4.58 percent per year unconditionally, 3.33 percent conditionally — about a quarter of the most famous number in finance was never alpha at all, just compensation for risk the old camera could not see.[^2] Post-1996, it is not a quarter. It is essentially everything: 2.45 percent unconditionally becomes 0.43 percent conditionally, with a t-statistic of 0.15 — statistically indistinguishable from nothing. In France, in Germany, in Sweden, conditional risk accounts for effectively all of value’s alpha. The investment factor gives up about half. Momentum is more stubborn — only about seven percent of its alpha over the long sample, though closer to a quarter in recent decades — which will matter enormously in a moment. And the cost-of-capital gap for a value firm — our CFO’s five-versus-seven — is this paper’s own worked example.
Now, the reflex reading of this literature is: science progressing. Lewellen-Nagel measured crudely, Boguth et al. refined, Gormsen-Jensen perfected; the estimates are converging on the true alpha. That reading is wrong, and seeing why it is wrong is the whole point of this essay.
Gormsen and Jensen did not measure Lewellen and Nagel’s number more accurately. They computed a different number and gave it the same name. Short-window realized betas (measure beta afresh in each short window and watch it move), instrumented realized betas (predict beta using only what investors could have known at the time), and a tradable timing portfolio (fold the estimation into a benchmark-side factor that harvests what time-varying beta earns) are not three approximations of one quantity. They are three definitions — three operational answers to the question “what do we subtract before we call the remainder skill?” And no, these papers are not one identical experiment run under a shared design — the test assets differ, the samples differ, the information sets differ. That is not a defense of the edifice. It is the indictment: there is no shared design because there is no shared object. Take value, the factor this whole essay follows. Lewellen and Nagel, measuring conditional betas directly, find its alpha essentially untouched — conditioning explains almost none of it. Gormsen and Jensen, folding the conditioning into a tradable portfolio instead, find the alpha of the same strategy in the same market cut by a quarter over the long sample and erased entirely after 1996. Intact, or gone, depending on what you subtract. Nothing is converging, because there is nothing to converge to. “Value’s alpha” is not a property of value. It is a property of the pair — value, and the lamp you chose to hold over it. Move the lamp, move the shadow, and the field has spent thirty years publishing increasingly sophisticated shadow measurements while insisting the shadow belongs to the object alone.
Be careful to take the right lesson. Gormsen and Jensen have not discovered the correct benchmark — crowning their conditional model as truth would just be electing a new emperor, and their own full-sample estimate leaves value 3.33 percent of alpha, so the model that kills value post-1996 certifies it pre-1996. The lesson is one level up: adjust for risk a little differently — defensibly differently — and the most-studied number in finance runs from substantial to zero. A number that behaves like that is not a measurement of the strategy. It is a negotiation between the strategy and the model — and only one of the parties gets named in the abstract.
And let me concede the obvious before a referee does it for me: chosen does not mean arbitrary. Coordinates are chosen; the distance between two points survives the choice. Celsius is a choice; the fever is real. A benchmark can be defended — by mandate, by opportunity set, by liabilities, by what the investor could actually have held instead. A risk model, an index mandate, and a feasible alternative are in fact three different jobs, and part of the industry’s confusion is letting one applicant interview for all three positions at once. The scandal is not that finance chooses benchmarks. The scandal is that finance writes the choice into the regression and then deletes it from the prose.
At this point the practitioner in the back of the room — and reader, I have been that practitioner for thirty years — stands up and says: this is why nobody on a trading desk uses the CAPM. Academics can keep their conditional models. My benchmark is the S&P 500. Beat the index, you have alpha; trail it, you do not. Solid ground.
I have some news about the solid ground.
Act Two: The Benchmark That Is People
First, a confession delivered without irony: I love the S&P 500. I have traded it, been measured against it, and watched it embarrass geniuses for three decades. The people who maintain it do an excellent job, and nothing that follows is a criticism of them. It is a criticism of what the rest of us have built on top of them — because what the rest of us have built is a monument to objectivity resting on judgment calls.
Start with what the S&P 500 is not. It is not the 500 largest American companies. It never has been. It is a portfolio of roughly 500 companies selected by a committee — the U.S. Index Committee at S&P Dow Jones Indices, market professionals who meet in private. Do not take my word for it; take S&P’s. The current methodology document — June 2026 edition, publicly posted — says it in one sentence: “Constituent selection is at the discretion of the Index Committee and is based on the eligibility criteria.” There are criteria — a market-cap floor ($22.7 billion at this writing), liquidity, float, and famously a profitability screen: positive GAAP earnings in the most recent quarter and summed over the trailing four. But clearing the criteria makes a company eligible, not included. The criteria are the velvet rope. The committee is the bouncer.
And the turnstile only checks tickets on the way in. The same document: “eligibility criteria are for addition to an index, not for continued membership. As a result, an index constituent that appears to violate criteria for addition to that index is not deleted unless ongoing conditions warrant an index change.” A company that stops qualifying stays in — unless, in the committee’s judgment, conditions warrant. The gate has rules. Membership has judgment. (Whether it also has privileges is another question entirely.)
If that sounds like an active manager’s charter, you have company in that opinion, and the company is the man who ran the committee. David Blitzer, chairman of the Index Committee for two decades, wrote in 2014 — publicly, cheerfully, on S&P’s own blog — that “there are no rigid or absolute rules for the S&P 500.” His post is a small masterpiece of saying the quiet part out loud. He tells the story of 2008: Lehman collapses, Washington rescues AIG — the original terms handed the government a 79.9 percent equity interest[^3] — and AIG’s public float crashes through the index’s fifty-percent floor. The rulebook said: remove it. Removing one of America’s largest financial companies from the index, that week, with every index fund on Earth forced to sell — Blitzer’s own words — “would have sent the markets tumbling yet again.” So, his words again: “the Index Committee quietly set the 50% float rule aside.”
Read that twice. In the most dangerous week in modern financial history, the world’s most important benchmark looked at its own rulebook and decided the rulebook was wrong. And here is the uncomfortable part: they were probably right. It was good judgment. That is precisely the point. A formula cannot decide its own exceptions. People can, and did, and do.
Tesla, 2020, same lesson in slow motion. The profitability screen kept Tesla out while it grew into one of the largest companies in America. In the summer of 2020 Tesla finally cleared the screen — its fourth straight profitable quarter, which is more than the formal test even demands (a positive latest quarter, and a positive sum over the trailing four). Eligible at last, overqualified even, and enormous. In September, the committee passed it over anyway, adding three much smaller companies while the financial press gaped. In December, Tesla went in — at a valuation well over half a trillion dollars, the largest addition in the index’s history by float-adjusted market capitalization, forcing index funds to execute roughly $73 billion of trades in a single day. Eligibility is necessary. Inclusion is a decision. You could watch the gap between those two words trade, live, for three months.
How big is the discretion, in general? Adriana Robertson, a law professor with an empiricist’s temperament, counted. In her 2015-to-2017 sample, the committee’s own criteria left it roughly six hundred eligible companies for five hundred seats — more than a hundred discretionary substitutions available on a typical day, representing about five percent of the index’s then-value. Call it a trillion dollars, sitting in the space between “the rules” and “the portfolio.” And when she compared the actual index to the obvious mechanical alternative — just buy the 500 largest stocks, no committee — she found that on the average day from 1989 through 2017, roughly 140 of the S&P 500’s members, better than a quarter of the list, were not in the mechanical top 500. The index is not a formula’s output with a committee attached for ceremony. It is a committee’s portfolio with a formula attached for comfort.
Robertson’s conclusion deserves to be quoted in the register of dry understatement law reviews are made for: index investing is “passive in name only” — it is delegated management, where the party you have delegated to is not a fund manager but the index creator. Sit with what that does to the industry’s favorite word. Passive investing does not exist. Passive implementation exists — the fund really does follow the index, mechanically and admirably. Selection-free index construction does not exist. The discretion has been delegated upstream. So there are only two kinds of active management: the kind you pay eighty basis points for, and the kind you pay three basis points for because the stock-picking committee in Manhattan works for the index provider instead of for you. The three-basis-point committee has a better track record. But it is a committee.
Now — and this is where I am honor-bound to argue against the most seductive version of my own case — does that mean the committee’s judgment is why the index is so hard to beat? Blitzer’s post gleefully relays that Bill Miller, the Legg Mason legend who beat the market fifteen years running, used to argue exactly this: the S&P 500’s record proves active management works, because the Index Committee are good active managers. It is a wonderful line. I want it to be true. The honest answer is: not established. Robertson herself measured the return gap — her mechanical index tracks the real one at a correlation of 0.9992 — and pointedly did not test whether the committee’s choices earn risk-adjusted anything. The long-run performance gap between the S&P 500 and the mechanical alternatives — Russell 1000, top-500-by-cap — is basis points per year, not a skill signature.[^4] And most of “hard to beat” needs no genius at all: it is Sharpe’s pitiless arithmetic — within any specified market, the average actively managed dollar earns that market’s return minus fees, as a matter of accounting rather than economics. And note what even the field’s one bulletproof theorem demands before it starts: someone must specify the market. The benchmark enters before the proof begins. The committee does not have to be brilliant for you to lose to it after costs; the average manager loses to the aggregate that managers collectively are, and the index sits close enough to that aggregate to collect the trophy.
And no — before the closet indexers uncork anything — none of this rescues the fund industry from SPIVA’s annual beatings. Losing to a portfolio chosen in advance is still losing; the scoreboard stands. What dies is the scoreboard’s philosophy: the pretense that the thing every manager in America is measured against is a neutral, skill-free fact of nature. It is not a benchmark-free null hypothesis. It is somebody’s portfolio. There is no benchmark-free null. There never was.
Act Three: Turtles
So test it, says the empiricist who has read this far. Fine: the committee is a portfolio manager. Measure the committee’s alpha. I admire the spirit. Now do it. Alpha against what? The mechanical top 500? Says who — why is size the right sort criterion rather than liquidity, or float, or profitability, which the committee itself uses? Russell’s rulebook? Russell’s rules were also written by people, and rewritten by people, most recently after the 2020 reconstitution embarrassments. Whichever yardstick you pick to measure the yardstick, someone picked it, and the regress does not terminate. It is the oldest joke in epistemology: turtles all the way down. Except this time the turtles charge licensing fees.
Do not trust me; trust the chairman. In that same 2014 post, Blitzer ran the regress one loop himself, in public, and you can hear the vertigo: “Of course, the rules in a rules-based index can be re-written (and sometimes are). But who will re-write the rules? Are there rule-makers ready to act or does someone need to form a committee?” There is no bottom, and the man at the bottom said so. Robertson found the farcical limit case: in her end-2016 sample, 81 of 571 U.S. equity ETFs — about one in seven — tracked indices created by the sponsor or an affiliate. The ruler, measuring itself, for a fee.
One exit remains, and the theorists will already be reaching for it: forget indices and models, the true benchmark is the stochastic discount factor — the deep object that prices everything, of which all these CAPMs and committees are mere shadows. Correct in principle! But an abstraction is not a yardstick. To make the SDF operational you must commit to preferences, information sets, test assets — choices of exactly the kind the previous essay watched contaminating every “measurement” of market structure, where the field’s own trilemma formalized the trap for demand elasticities. The exit door is real. It opens back into the same room. This is not three separate scandals. It is one scandal with three organs: the discovered predictors do not forecast (Emperor 1), the structural parameters are not measurements (Emperor 2), and the yardstick is a committee, or a contested model, chosen by somebody, all the way down (this essay). The field built a cathedral of quantitative rigor and put a suggestion box at the altar.
What should a working investor do with this? The same thing I told you the trilemma implies, now with the third leg under it. Every published alpha — every backtest, every factor premium, every fund ranking, every SPIVA table, every “our strategy added 200 basis points” — is a joint bet on a strategy and a yardstick, and the second bet is the one nobody tells you you are making. So when someone quotes you an alpha, the first question is not “how big?” It is “against what?” And the second question is “who chose the what, and what were they optimizing when they chose it?” You will be amazed how many careers, product launches, and Nobel-adjacent reputations cannot survive the second question.
And with that, the Emperor trilogy is complete: no alpha, no identification, no benchmark. The predictions do not predict, the measurements do not measure, and the ruler is a committee. Alpha is relational. The field has always known this in the mathematics; it forgets it in the language; it has built institutions on the omission.
Naturally, therefore — in the grand tradition of Douglas Adams, whose Hitchhiker’s trilogy ran to five volumes — the fourth part of the trilogy arrives shortly. Because one wall of the cathedral is still standing: the factors themselves. And there is a loose end even sharper than that. Momentum was the one predictor that survived Emperor 1’s massacre — and it is also the one whose alpha will not sit still: on the order of a percent a month under Lewellen and Nagel, a fifth to two-fifths smaller once Boguth and coauthors instrument it properly, barely dented over the long sample yet down about a quarter in recent decades under Gormsen and Jensen. Something about momentum is being conserved across all of these lamps, and I do not think it is a factor loading. Next time.
Meanwhile, somewhere in America, the CFO picks a discount rate. She builds the factory, or she does not. Twenty years from now, someone will score that decision — against a benchmark that has not been chosen yet, by people who have not been hired yet, using a model that has not been contested yet. She is not measuring the cost of capital. She is choosing a model and reading a number off it — which is all anyone has ever done. And somewhere in Manhattan, a committee meets quietly, exercises excellent judgment, writes nothing down, and remains the closest thing to an objective standard that the entire quantitative edifice of modern finance has ever produced.
May they never retire.
[^1]: Lewellen and Nagel estimate the standard deviations of conditional betas at roughly 0.25 for value, 0.30 for size, and 0.60 for momentum — real variation, just several times too small in the dimension that would rescue the CAPM. Figures are from the working-paper version; treat magnitudes as approximate.
[^2]: Throughout, the Gormsen-Jensen figures are from the published Journal of Financial Economics version (2024). In their own alternative specifications, conditional risk explains as much as half of value’s unconditional alpha even in the long sample — so the quarter quoted above is, if anything, the conservative reading.
[^3]: Blitzer’s 2014 recollection puts the government’s AIG stake at “over 90%”; it eventually got there in later restructurings, but the September 2008 rescue terms specified a 79.9 percent equity interest. The float breach, and the waiver, are not in dispute.
[^4]: Verified directly: over 2004–2026, monthly total returns on the S&P 500 (via SPY) and the Russell 1000 (via IWB) differ by roughly one basis point per year, at a correlation of 0.998. The widest gap among the big mechanical large-cap alternatives — the CRSP US Large Cap index, via VV — is about 24 basis points per year, and most of that is the funds’ expense-ratio difference, not index construction. Whatever the committee is adding, it is invisible at the yardstick’s own resolution. Robertson’s in-paper measurement agrees: a 0.9992 correlation between the actual index and her mechanical top-500, 1989–2017.
This is the third essay in the Emperor series. Read the first, The Emperor Has No Alpha, and the second, The Emperor Has No Identification.
If you enjoy watching cherished stories audited down to the studs, my book, The Science of Free Will, asks an equally uncomfortable question about an equally cherished story. For more markets work, see Trading; the physics-minded will find its cousins in the Physics series. New here? Start with The Paradox of India and the full India series — and for related country-specific work, there’s Albion, Britain’s institutional decline, first spotted in 1979.
For more on the background behind this work, see my first and second interviews with Titans of Tomorrow, and Part 2 of my Algorithmic Advantage Podcast conversation with Simon M.







If there is no alpha, how do I measure my performance?
What is skill, respected sir? Is it, just make money more than others if the amount of risk and capital are assumed to be the same?
Also, I know its a bit overeaching, but could you tell a bit about how to find a statistical edge in our trading, I've been working on it for a while, I code, read papers and do research, but I have yet to find a stable edge that I could execute on.
Thank you