Lecture 4
PIMES/UFPE
Lecture 3 built \(\bb = (\bX'\bX)^{-1}\bX'\bY\) out of pure algebra: it exists for any design of full column rank, whether or not the model is true. We now ask what it is good for.
The estimate you compute is one number. The estimator is a function of the sample, and the sample is random — so a different draw gives a different \(\bb\).
We cannot write the whole sampling distribution from A1–A4 alone — that needs the normality assumption A6, which we add at the end of this lecture. But we can already compute its moments, and the first two answer the questions that matter.
\(y_i = \mathbf{x}_i'\bbeta + \varepsilon_i\) with \((y_i, \mathbf{x}_i)\) i.i.d., \(\rank(\bX) = K\) (A2) and \(\E[\beps \mid \bX] = \bzero\) (A3).
Nothing is assumed about \(\Var[\beps \mid \bX]\).
That silence is deliberate, and it is where this lecture departs from the traditional treatment.
Everything in this lecture follows from one substitution. Writing \(\bY = \bX\bbeta + \beps\),
\[\bb = \bbeta + \eqt{a}{(\bX'\bX)^{-1}\bX'}\eqt{b}{\beps}\]
The estimator equals the truth plus a term linear in \(\beps\). Unbiasedness, variance and efficiency are all read off that single term.
Theorem 1 (Unbiasedness of least squares) Under A1–A3, \(\E[\bb \mid \bX] = \bbeta\), and therefore \(\E[\bb] = \bbeta\).
Worked on the board. The skeleton:
Proof. Take conditional expectations in the decomposition above. The matrix \((\bX'\bX)^{-1}\bX'\) is fixed given \(\bX\), so it passes outside, leaving \((\bX'\bX)^{-1}\bX'\E[\beps \mid \bX] = \bzero\) by A3. The unconditional statement follows by iterated expectations. \(\square\)
Notice which assumption did the work: A3 alone. Not homoskedasticity, not normality, not even large samples. Unbiasedness is exactly as strong as the exogeneity assumption behind it — no stronger, and no weaker.
Unbiasedness is a statement about the distribution, not any one estimate. Draw 5,000 samples from a DGP with a known slope \(\beta = 1\) and recompute \(\bb\) in each:
Figure 1: Sampling distribution of the slope over 5,000 samples. The correctly specified OLS estimator (teal) piles up on the true value (dashed): its average across samples is the truth. An estimator that omits a relevant regressor (amber) is centred elsewhere — same data, but biased.
Unbiasedness needed only A3. The most common way A3 fails in practice is not exotic: a variable that belongs in the model is left out. Lecture 2 flagged this as the leading example of an exogeneity failure and promised the derivation here.
The true model is \(\bY = \bX\bbeta + \mathbf{z}\gamma + \beps\) with \(\E[\beps \mid \bX, \mathbf{z}] = \bzero\), but we regress \(\bY\) on \(\bX\) only, leaving out the relevant \(\mathbf{z}\).
Substituting the true model into \(\bb = (\bX'\bX)^{-1}\bX'\bY\),
\[\bb = \bbeta + \underbrace{(\bX'\bX)^{-1}\bX'\mathbf{z}}_{\mathbf{p}_{X\cdot z}}\gamma + (\bX'\bX)^{-1}\bX'\beps,\]
so taking expectations conditional on \(\bX\) and \(\mathbf{z}\),
Theorem 2 (Omitted-variable bias) \[\E[\bb \mid \bX, \mathbf{z}] = \bbeta + \mathbf{p}_{X\cdot z}\,\gamma, \qquad \mathbf{p}_{X\cdot z} = (\bX'\bX)^{-1}\bX'\mathbf{z}.\]
The bias is the regression of the omitted variable on the included ones, scaled by the omitted variable’s own coefficient \(\gamma\). It vanishes iff \(\gamma = 0\) (the variable was irrelevant) or \(\bX'\mathbf{z} = \bzero\) (it is orthogonal to every included regressor).
\(\mathbf{p}_{X\cdot z}\) is itself a vector of regression coefficients — exactly the object Frisch–Waugh–Lovell describes. By the single-extra-regressor corollary of Lecture 3, the bias in a single coefficient \(b_k\) is
\[\E[b_k \mid \bX, \mathbf{z}] = \beta_k + \gamma\,\frac{\Cov(\mathbf{z}, x_k \mid \text{other }x)}{\Var(x_k \mid \text{other }x)}.\]
Example 1 (Ability bias in the returns to education) \[\text{Income} = \beta_0 + \beta_1\,\text{Educ} + \beta_2\,\text{age} + \beta_3\,\text{age}^2 + \beta_4\,\text{Ability} + \varepsilon.\]
Ability is unobserved, so we omit it. Then \[\E[b_1 \mid \bX, \text{Ability}] = \beta_1 + \beta_4\,\frac{\Cov(\text{Ability}, \text{Educ} \mid \text{age}, \text{age}^2)}{\Var(\text{Educ} \mid \text{age}, \text{age}^2)}.\] If more able people earn more (\(\beta_4 > 0\)) and also study more (positive partial covariance), \(b_1\) is biased upward: the return to schooling is overstated because it absorbs the return to ability.
Unbiasedness says the estimates centre on the truth. It says nothing about how far they scatter — and scatter is what standard errors report.
Theorem 3 (Variance of least squares) Let \(\bD = \E[\beps\beps' \mid \bX]\) be the \(n\times n\) conditional variance matrix of the disturbances, with no restriction on its form. Then \[\Var[\bb \mid \bX] = (\bX'\bX)^{-1}\,\bX'\bD\bX\,(\bX'\bX)^{-1}.\]
Proof. \(\bb - \bbeta = (\bX'\bX)^{-1}\bX'\beps\) is linear in \(\beps\); apply \(\Var[\bA\beps] = \bA\,\Var[\beps\mid\bX]\,\bA'\) with \(\bA = (\bX'\bX)^{-1}\bX'\) and \(\Var[\beps\mid\bX] = \bD\). \(\square\)
The expression has a bread–filling–bread shape and is universally called the sandwich formula. It holds under A1–A3 alone — the same assumptions as unbiasedness, and no more.
Everything here conditions on \(\bX\). Averaging over the design gives the unconditional variance, and the law of total variance is unusually clean, because the first moment is constant: \[\Var[\bb] = \E\big[\Var[\bb \mid \bX]\big] + \underbrace{\Var\big[\E[\bb \mid \bX]\big]}_{=\ \bzero} = \E\big[\Var[\bb \mid \bX]\big].\]
The model of Lecture 2 constrained the conditional mean and said nothing about the conditional variance. We now constrain the variance too — and give the restriction the name the rest of the course uses.
Definition 1 (A4 · Conditional homoskedasticity) \[\bD = \E[\beps\beps' \mid \bX] = \sigma^2\bI_n,\] which packs two claims together: homoskedasticity, \(\sigma^2(\mathbf{x}_i) = \sigma^2\) for every \(i\), and no correlation across observations, \(\Cov[\varepsilon_i,\varepsilon_j \mid \bX] = 0\) for \(i \neq j\).
Under A4 the filling \(\bX'\bD\bX\) becomes \(\bX'(\sigma^2\bI_n)\bX = \sigma^2\bX'\bX\), one factor cancels on each side, and what is left is the formula every textbook opens with: \[\Var[\bb \mid \bX] = \sigma^2(\bX'\bX)^{-1}.\]
Important
The familiar formula is not the general result. It is what the sandwich collapses to when the conditional variance happens to be constant. Teaching it first, and the sandwich later as a repair, gets the logic backwards.
Both forms need an estimate. Under A4 it suffices to estimate the scalar \(\sigma^2\).
Theorem 4 (Unbiased estimation of \(\sigma^2\)) Under A1–A4, \(s^2 = \be'\be/(n-K)\) satisfies \(\E[s^2 \mid \bX] = \sigma^2\).
Proof. \(\be = \bM\beps\), so \(\E[\be'\be \mid \bX] = \E[\beps'\bM\beps \mid \bX] = \sigma^2\tr(\bM) = \sigma^2(n-K)\), using \(\rank(\bM) = n-K\) from the properties of \(\bP\) and \(\bM\) in Lecture 3. \(\square\)
The divisor \(n-K\) is not a convention: it is exactly what makes the expectation come out right, and it is the trace of the residual maker from Lecture 3.
When \(\bD\) is unrestricted there is no single \(\sigma^2\) to estimate — there are \(n\) conditional variances and only \(n\) observations. The way out is that the sandwich needs only the \(K \times K\) filling \(\bX'\bD\bX\), not \(\bD\) itself.
Definition 2 (Heteroskedasticity-robust covariance estimator) \[\widehat{\Var}[\bb \mid \bX] = (\bX'\bX)^{-1}\left(\sum_{i=1}^n e_i^2\,\mathbf{x}_i\mathbf{x}_i'\right)(\bX'\bX)^{-1}.\] Each observation contributes its own squared residual as an estimate of its own variance.
The corrections about to appear depend on leverage — the \(i\)-th diagonal of the projection (hat) matrix \(\bP = \bX(\bX'\bX)^{-1}\bX'\) built in Lecture 3.
Definition 3 (Leverage) \[h_{ii} = [\bP]_{ii} = \mathbf{x}_i'(\bX'\bX)^{-1}\mathbf{x}_i \in [0,1], \qquad \sum_{i=1}^n h_{ii} = K.\]
Leverage measures how much observation \(i\) pulls its own fitted value, \(\partial \hat y_i/\partial y_i = h_{ii}\); a high value flags an \(\mathbf{x}_i\) far from the centre of the design.
Key fact: under homoskedasticity \(\E[e_i^2 \mid \bX] = \sigma^2(1-h_{ii})\) — a raw \(e_i^2\) understates \(\sigma^2\), most at high-leverage points.
The estimator above is HC0. Because \(\E[e_i^2] < \E[\varepsilon_i^2]\) — residuals are shrunk by fitting — it is biased downward in small samples, and several corrections exist:
| Correction | |
|---|---|
| HC0 | \(e_i^2\) |
| HC1 | \(\dfrac{n}{n-K}\,e_i^2\) |
| HC2 | \(\dfrac{e_i^2}{1-h_{ii}}\) |
| HC3 | \(\dfrac{e_i^2}{(1-h_{ii})^2}\) |
HC2 and HC3 divide \(e_i^2\) by \((1-h_{ii})\) and by its square — the biggest corrections fall on the high-leverage points, whose residuals shrink most. The exponent on \((1-h_{ii})\) runs \(0, 1, 2\) across HC0, HC2, HC3; HC4 (Cribari-Neto, 2004) makes it adaptive. Stata’s robust is HC1 and R’s vcovHC defaults to HC3, so “robust standard errors” without a label is ambiguous.
Theorem 5 (Gauss-Markov) Under A1–A4, for any estimator \(\tilde{\bbeta} = \bA\bY\) that is linear in \(\bY\) and unbiased conditional on \(\bX\), \[\Var[\tilde{\bbeta} \mid \bX] - \Var[\bb \mid \bX] \;\succeq\; \bzero,\] that is, the difference is positive semi-definite. Least squares is BLUE — the best linear unbiased estimator.
Derived on the board — the longest derivation of the course so far. The skeleton:
Proof. Write \(\bA = (\bX'\bX)^{-1}\bX' + \bD\). Unbiasedness for all \(\bbeta\) forces \(\bD\bX = \bzero\). Computing the variance under A4, the cross terms vanish precisely because of that restriction, leaving \(\sigma^2(\bX'\bX)^{-1} + \sigma^2\bD\bD'\). The second term is positive semi-definite. \(\square\)
The result is narrower than its reputation. Each qualifier is load-bearing:
Warning
Gauss-Markov is the one place in this lecture where A4 genuinely cannot be dropped. Under heteroskedasticity least squares remains unbiased but is no longer efficient — GLS is, and that is Lecture 7.
We obtained \(\bb\) by minimising \(S(\mathbf{b}) = (\bY - \bX\mathbf{b})'(\bY - \bX\mathbf{b})\) — an instance of \[\hat\btheta = \argmin_{\btheta} Q_n(\btheta).\] Estimators defined this way are M-estimators. Maximum likelihood, in Lecture 13, is another: same shape, different criterion. Much of what we proved today will be proved once more, in that generality, for the whole family.
Setting the derivative of \(S(\mathbf{b})\) to zero gives \(\bX'\be = \bzero\) — a moment condition, and exactly the sample analogue of the population \(\E[\mathbf{x}e] = \bzero\) derived in Lecture 2. It pins \(\bb\) down without minimising anything: least squares is equally the estimator that forces the sample orthogonality to hold. Estimators defined by such an estimating equation, \(\tfrac1n\sum_i \boldsymbol{\psi}(W_i,\btheta) = \bzero\), are Z-estimators (the method of moments). OLS is both at once — an M-estimator that minimises a criterion and a Z-estimator that solves a moment condition — because differentiating the criterion is what produces the condition. This is the thread to IV (Lecture 10) and GMM (Lecture 12).
The mean and the variance are all A1–A4 will give. To pin down the whole sampling distribution — every quantile, not just centre and spread — we must say something about the shape of the disturbances. One assumption does it.
Definition 4 (A6 · Conditional normality) \[\beps \mid \bX \;\sim\; \Normal(\bzero,\ \sigma^2\bI_n).\] The disturbances are jointly normal, homoskedastic, and uncorrelated across observations, given the design. This is A4 plus a full distributional shape — strictly stronger than anything assumed so far.
What A6 implies: the entire sampling distribution of \(\bb\) becomes known — exact and for every \(n\), not merely its first two moments (Theorem 6). It is exactly this that Lecture 5 will drop.
Since \(\bb - \bbeta = (\bX'\bX)^{-1}\bX'\beps\) is a linear function of a normal vector, it is normal too — exactly, for every \(n\), not just in the limit.
Theorem 6 (Normal regression) Under A1–A4 and A6, conditional on \(\bX\), \[\bb \;\sim\; \Normal\!\big(\bbeta,\ \sigma^2(\bX'\bX)^{-1}\big).\]
This exact normal law is the finite-sample counterpart of the asymptotic normal of Lecture 5. The inference it powers — the \(\chi^2\) distribution of \(s^2\) and the \(t\)-ratio of a coefficient — is built on it when we turn to testing in Lecture 6.
| Property | Needs | Holds under heteroskedasticity? |
|---|---|---|
| Unbiasedness | A1–A3 | yes |
| Sandwich variance | A1–A3 | yes |
| \(\sigma^2(\bX'\bX)^{-1}\) | + A4 | no |
| \(s^2\) unbiased for \(\sigma^2\) | + A4 | no |
| Efficiency (Gauss-Markov) | + A4 | no |
| Exact normality of \(\bb\) | + A6 | no |
Read the last column. Only efficiency and the simplified formulas depend on A4, and the exact normal distribution of \(\bb\) on the stronger A6; the estimator itself and its variance do not. That is the whole case for treating robust inference as the default rather than the exception.
Everything today was exact and conditional on \(\bX\) — but the distribution, unlike the two moments, came at a price: assumption A6. That price is steep, and it is where the finite-sample story runs out.
| Hansen | Chapter 4 — §4.2–4.5 (mean, variance, Gauss-Markov), §4.8–4.10 (covariance estimation, standard errors) |
| Greene | Chapter 4 — the assumptions, omitted-variable bias, and the Gauss-Markov theorem |