Skip to content
Berk Günberk
Theme

1952 · The Journal of Finance

Portfolio Selection

What it objects to

Markowitz starts from a criticism. The rule in common use at the time was to maximise discounted expected return. The paper shows that this rule has a flaw: it never prefers a diversified portfolio. Whichever security has the highest expected return, the rule says to put everything into it.

Yet diversification is both observed and sensible. Investors hold portfolios, and it would be hard to call them foolish for doing so. Something is missing from the rule: it does not see risk at all.

The E-V rule

Markowitz defines two quantities. The expected return of a portfolio is a weighted sum:

Variance depends not only on the variance of the individual assets but on the covariance between them:

The symbols keep the same meaning throughout the page:

SymbolMeaning
Xᵢthe weight of asset i in the portfolio. The weights sum to 1, and with no short selling none of them can be negative.
μᵢthe expected return of asset i
σᵢⱼthe covariance of assets i and j, that is, how much they move together. When i = j this is the asset's own variance.
Ethe expected return of the portfolio
Vthe variance of the portfolio
μthe vector holding all the μᵢ
Σthe matrix holding all the σᵢⱼ; the covariance matrix

One warning about a collision: the summation sign ∑ in the equations and the name of the covariance matrix Σ are the same Greek letter. The first means "add up" and carries a counter beneath it; the second is the name of a matrix.

The cross terms in the second expression are the paper's real contribution. The investor wants E and avoids V; there is a trade-off between them. The E-V rule accepts that trade-off and says: if some portfolio gives lower V at the same E, there is no reason to choose the other one.

The three-asset geometry

With three assets X₃ = 1 − X₁ − X₂, so every portfolio lands on a plane and the attainable set becomes a triangle. Portfolios with the same expected return lie on lines; those with the same variance lie on concentric ellipses whose centre is the minimum variance point.

The paper invites the reader to set up three cases. The minimum variance point can fall the triangle; under strong correlation it moves and the efficient set settles on the boundary. When two assets share an expected return () the isomean lines run parallel to an edge.

The attainable set, isomean lines, isovariance ellipses and efficient set in the three-asset case.Expected returns range from %5.0 to %12.0; a 3-segment efficient set is drawn. The minimum variance point is (0.31, 0.11), inside the triangle. Layers shown: isomean lines, isovariance ellipses.X₁ = 1X₂ = 1X₃ = 1
Figure parameters
Σ

The attainable set, isomean lines, isovariance ellipses and efficient set in the three-asset case.

The efficient frontier

The same set can be drawn in the language of outcomes: variance on the horizontal axis, expected return on the vertical. The efficient set becomes a curve made of connected parabolic segments; the joints are the points where an asset enters or leaves the portfolio.

The numbered points are the individual assets. That the curve passes to their left is the paper's core claim: the same expected return is available at lower variance. The gain is largest when the assets are ; under the curve moves towards the assets and the benefit of diversification melts away.

The efficient frontier: expected return against variance. Numbered points are the individual assets.The efficient frontier is made of 2 parabolic segments; expected return ranges from %6.7 to %12.0. The minimum variance portfolio carries 37% less variance than the single lowest-variance asset.1230.0140.090V — annual variance%5.0%12.0E — annual expected return
Figure parameters
Σ

The efficient frontier: expected return against variance. Numbered points are the individual assets.

Why diversification has to be of the right kind

The warning on page 89 is still often skipped: diversification is not a matter of count. Sixty securities from the same industry do not count as being as diversified as sixty securities drawn from different industries.

The reason sits in the expression for V above. As a portfolio grows, the weight of the individual variance terms falls while the number of covariance terms grows with the square; past a certain point what determines the variance of the portfolio is no longer the variance of the assets but the covariance between them. Within one industry covariance is high, so each additional holding contributes less and less.

This is why the paper does not simply say "diversify" but "diversify in the right way".

What it does not solve

Throughout the paper μ and Σ are assumed known. Markowitz does not hide this; on the last page he says that the first stage of selection — moving from observation and experience to beliefs — falls outside the study.

In practice both inputs are estimated from a sample, and estimates carry error.

Estimation error

The experiment below splits that error into three sources: what is lost when only the expected returns are wrong, when only the variances are wrong, and when only the covariances are wrong. The measure is cash equivalent loss; a portfolio built from estimated inputs is evaluated under the true parameters.

The result does not weaken the paper's framework so much as mark its edge: at these settings the overwhelming share of the loss comes from μ. How sharp the ratio is depends on risk tolerance — we come back to that below — but the ordering never changes. With a the loss grows; with a it shrinks but does not vanish. The 1/N portfolio, which optimises nothing, is drawn for comparison.

The cost of estimation error: how much error in each input actually costs.The estimation error experiment has not run yet; only the axis is drawn.
Figure parameters
Σ

The cost of estimation error: how much error in each input actually costs.

This experiment is not a literal reproduction of the paper, and does not try to be. Chopra & Ziemba apply a controlled perturbation to the inputs: each one receives the same relative error, which isolates the question "under equal error, which matters most". The error here is genuine sampling error — a sample of T periods is drawn and the sample estimators are used. The question is a different one: "what do I lose if I work with T periods of data". For this page the second question is the right one, because the issue is where the inputs come from in the first place.

The difference shows up in the result. The ratio here carries two effects at once: how sensitive the optimisation is to each input and how precisely each input can be estimated from a sample. The standard error of a mean is σ/√T while that of a variance is of order √(2/(T−1)); the two do not improve at the same rate.

The ratio is also not a universal constant — it is sensitive to risk tolerance. Pull the slider above down to 0.25 and the ordering approaches the frequently quoted 11:2:1; push it to 1.0 and the dominance of μ becomes far sharper. The one thing that does not change is the ordering itself: μ costs more than variance, and variance more than covariance.

Much of the next fifty years of literature is a set of answers to that gap: shrinkage, Black-Litterman, resampling, risk parity.

Where this goes next

Here is where Markowitz leaves things: choosing a portfolio is no longer a matter of intuition but a problem defined between two quantities. The answer to "which portfolio is good" is not one portfolio but a curve, and where you sit on that curve is up to you. Diversification stops being a piece of advice and becomes a consequence of the arithmetic.

But the setup leaves two questions open. The first we saw above: where the inputs come from. The second is subtler — why E and V? That the investor cares about exactly these two quantities is an assumption in Markowitz, not a result.

In the same year someone else was circling the same problem from a completely different direction. A. D. Roy starts not from mean and variance but from avoiding disaster: keep the return from falling below some threshold. Set up that way, the problem arrives somewhere surprisingly close to Markowitz's. Years later Markowitz would write that Roy deserves a share of the credit for founding portfolio theory.

That is the next paper in the series.

In the unconstrained case the Critical Line Algorithm was compared against the Merton (1972) analytic solution on 2,000 randomly generated valid inputs. The largest relative deviation was 3.79e-12, against a tolerance of 1e-9.