Generalized Least Squares and Clustered Dependence

Lecture 7

Henrique Veras

PIMES/UFPE

Where we are

Where this lecture sits

DGPIdentificationEstimationAsymptoticsInference
Computation

Lectures 4–5 already made heteroskedasticity the default: \(\E[\beps\beps'\mid\bX]=\bD\) is a general positive-definite matrix, OLS is consistent, and the robust (sandwich) standard errors deliver valid inference with no homoskedasticity assumption. So the classical “violation of A4” story is behind us. Two questions remain — and they are what this lecture is about.

Two questions robust OLS does not answer

  • Efficiency — OLS is consistent and, with robust errors, correctly measured. But is it the best we can do? Under a general \(\bD\), Gauss–Markov no longer crowns OLS.
  • Dependence — the sandwich assumes observations are independent. What if they come in groups (clusters), correlated within?

Part I · The generalized model and the cost to OLS

The generalized regression model

One assumption, dropped long ago

The model is the familiar linear one, with the error covariance left general: \[\bY=\bX\bbeta+\beps,\qquad \E[\beps\mid\bX]=\bzero,\qquad \E[\beps\beps'\mid\bX]=\bD,\] with \(\bD\) symmetric positive-definite (an \(n\times n\) matrix). Homoskedasticity, \(\bD=\sigma^2\bI\), is the special case — not the starting point.

Definition 1 (The two leading structures) Heteroskedasticity (this lecture’s focus): errors uncorrelated across observations but with different variances — \(\bD=\operatorname{diag}(\sigma_1^2,\dots,\sigma_n^2)\), i.e. \(\E[\varepsilon_i^2\mid\mathbf{x}_i]=\sigma_i^2(\mathbf{x}_i)\). Dependence: \(\bD\) has non-zero off-diagonal entries — serial correlation (time series) or within-cluster correlation (Part V).

OLS still works — it just is not the best

Nothing about consistency or valid inference breaks under a general \(\bD\):

  • Unbiased and consistent: \(\E[\bb\mid\bX]=\bbeta\) and \(\plim\,\bb=\bbeta\) — neither used A4.
  • Inference is handled: the robust \(\widehat{\Asyvar}[\bb]\) of Lecture 5 is valid without A4.
  • But not efficient: Gauss–Markov needed \(\bD=\sigma^2\bI\). Drop it and OLS loses its BLUE crown.

Part II · Generalized Least Squares

The cost of ignoring D

Every observation weighted equally

A general \(\bD\) can break A4 in two ways — heteroskedasticity (the diagonal of \(\bD\) varies across observations) or autocorrelation (non-zero off-diagonals), or both. OLS stays unbiased and consistent either way, but it pays for ignoring \(\bD\) in precision. Compare the variance OLS assumes with the true one: \[\underbrace{\sigma^2(\bX'\bX)^{-1}}_{\text{under A4}} \qquad\text{vs.}\qquad \underbrace{(\bX'\bX)^{-1}\bX'\bD\bX(\bX'\bX)^{-1}}_{\text{true, general }\bD}.\]

  • Write \(\bX'\bX=\sum_i\mathbf{x}_i\mathbf{x}_i'\): the A4 formula puts the same weight on every observation.
  • But a noisy observation (large \(\sigma_i^2\)) carries less information — efficiency wants to downweight it, with weight \(\propto 1/\sigma_i^2\).
  • OLS discards that weighting. Putting it back is exactly what GLS does.

GLS with known D

Transforming to constant variance

If the trouble is unequal (and correlated) variances, the cure is to rescale the observations so the variances become equal and the correlations vanish — then the classical model, and OLS’s optimality, are back. Since \(\bD\) is positive-definite, there is a matrix \(\mathbf{L}\) with \[\mathbf{L}\,\bD\,\mathbf{L}'=\bI,\qquad\text{equivalently}\qquad \bD^{-1}=\mathbf{L}'\mathbf{L}\] — the matrix analogue of “dividing by the standard deviation” (a Cholesky factor is one such \(\mathbf{L}\)).

On the Board

Pre-multiply the model by \(\mathbf{L}\): with \(\bY_*=\mathbf{L}\bY\), \(\bX_*=\mathbf{L}\bX\), \(\beps_*=\mathbf{L}\beps\), \[\bY_*=\bX_*\bbeta+\beps_*,\qquad \E[\beps_*\beps_*'\mid\bX]=\mathbf{L}\,\bD\,\mathbf{L}'=\bI.\] The transformed errors are homoskedastic and uncorrelated — the classical model holds for \((\bY_*,\bX_*)\), so OLS on the transformed data is efficient. That is GLS.

The GLS estimator

Theorem 1 (Generalized least squares) OLS on the transformed model is \[\hat{\bbeta}_{\text{GLS}} =(\bX_*'\bX_*)^{-1}\bX_*'\bY_* =(\bX'\bD^{-1}\bX)^{-1}\bX'\bD^{-1}\bY.\] Among linear unbiased estimators it has the smallest variance — the generalized Gauss–Markov (Aitken) theorem — with \[\Var[\hat{\bbeta}_{\text{GLS}}\mid\bX]=(\bX'\bD^{-1}\bX)^{-1}.\] OLS is the special case \(\bD=\sigma^2\bI\) (weight matrix \(\bI\) instead of \(\bD^{-1}\)).

Weighted least squares: the heteroskedastic case

Corollary 1 (Weighted least squares) When the violation is pure heteroskedasticity, \(\bD=\operatorname{diag}(\sigma_1^2,\dots,\sigma_n^2)\), the transformation is just dividing each observation by its own \(\sigma_i\) (\(\mathbf{L}=\operatorname{diag}(1/\sigma_1,\dots,1/\sigma_n)\)). GLS then takes the familiar weighted form \[\hat{\bbeta}_{\text{WLS}} =\Big(\textstyle\sum_i \tfrac{1}{\sigma_i^2}\,\mathbf{x}_i\mathbf{x}_i'\Big)^{-1}\Big(\textstyle\sum_i \tfrac{1}{\sigma_i^2}\,\mathbf{x}_i y_i\Big) \ -\ \text{exactly the weighting OLS was missing.}\]

  • Each observation enters with weight \(1/\sigma_i^2\): precise points count more, noisy points less — the weight the previous slides said OLS should have used.
  • A classic special case: variance proportional to a regressor, \(\sigma_i^2=\sigma^2 x_{ik}^2\) — then WLS is simply OLS after dividing \(y_i\) and \(\mathbf{x}_i\) by \(x_{ik}\).

Part III · Feasible GLS

When D is unknown

The skedastic regression

GLS needs \(\bD\), which we never know. But an unrestricted \(\bD\) has \(n\) variances — too many to estimate from \(n\) observations. Model the variance with a few parameters:

Definition 2 (The skedastic regression) Specify \(\sigma_i^2=\sigma^2(\mathbf{z}_i,\boldsymbol{\alpha})\) for a small parameter vector \(\boldsymbol{\alpha}\) and observed \(\mathbf{z}_i\) (often a subset/transform of \(\mathbf{x}_i\)) — for example the linear \(\sigma_i^2=\alpha_0+\mathbf{z}_i'\boldsymbol{\alpha}_1\) or the multiplicative \(\sigma_i^2=\exp(\alpha_0+\mathbf{z}_i'\boldsymbol{\alpha}_1)\). Estimate \(\boldsymbol{\alpha}\) by regressing the squared OLS residuals \(\hat e_i^2\) on \(\mathbf{z}_i\) — the skedastic regression.

The FGLS estimator

Theorem 2 (Feasible GLS) Form \(\widehat{\bD}=\operatorname{diag}(\hat\sigma_1^2,\dots,\hat\sigma_n^2)\) from the skedastic regression and plug it in: \[\hat{\bbeta}_{\text{FGLS}}=(\bX'\widehat{\bD}^{-1}\bX)^{-1}\bX'\widehat{\bD}^{-1}\bY.\] If the skedastic model is correctly specified, estimating \(\boldsymbol{\alpha}\) is asymptotically free: \(\hat{\bbeta}_{\text{FGLS}}\) has the same asymptotic distribution as infeasible GLS.

On the Board

Why “asymptotically free”: the sampling error in \(\hat{\boldsymbol{\alpha}}\) enters \(\hat{\bbeta}_{\text{FGLS}}\) only through \(\widehat{\bD}\pto\bD\), and a \(\sqrt{n}\)-consistent \(\hat{\boldsymbol{\alpha}}\) perturbs \(\hat{\bbeta}\) by a term that vanishes faster than \(1/\sqrt{n}\). So FGLS and GLS share the limit \(\Normal(\bzero,(\bX'\bD^{-1}\bX/n)^{-1})\).

Why not always FGLS?

The efficiency–robustness trade-off

If FGLS is efficient and asymptotically free, why is OLS still the default? Three reasons — this is the heart of the lecture.

  1. It can do worse. In finite samples the estimated weights carry more noise than information; FGLS can have larger variance than OLS.
  2. Negative variances. Fitted \(\hat\sigma_i^2\) can be \(\le 0\), forcing an arbitrary trimming rule.
  3. OLS is robust; FGLS is not. OLS is consistent for the linear projection even if the conditional mean is misspecified. FGLS’s efficiency is bought with the stronger assumption of a correct conditional mean — get it wrong and OLS and FGLS converge to different limits.

Part IV · Testing for heteroskedasticity

The skedastic regression, tested

White and Breusch–Pagan are one test

Theorem 3 (Testing constant variance) Homoskedasticity is \(H_0:\boldsymbol{\alpha}_1=\bzero\) in the skedastic regression \(\hat e_i^2=\alpha_0+\mathbf{z}_i'\boldsymbol{\alpha}_1+\nu_i\) — the variance does not depend on \(\mathbf{z}_i\). Test it with the joint Wald/\(F\) on \(\boldsymbol{\alpha}_1\) (equivalently \(nR^2\dto\chi^2_q\), \(q=\dim\boldsymbol{\alpha}_1\)). The two classic choices differ only in \(\mathbf{z}_i\): White uses the squares and cross-products of \(\mathbf{x}_i\); Breusch–Pagan uses a general \(\mathbf{z}_i\). Under independence they are essentially the same test.

What the test is — and is not — for

Under the Hood—Do not misuse the test

A heteroskedasticity test answers one scientific question: does \(\sigma^2\) depend on the regressors? It is not a tool to choose between OLS and FGLS, nor between classical and robust standard errors — “hypothesis tests are not designed for these purposes.” You use robust errors regardless of the test. Run the test only if whether \(\sigma^2(\mathbf{x})\) varies is itself of economic interest.

Part V · Clustered dependence

When observations come in groups

Clusters: dependence by design

Definition 3 (Clustered sampling) The sample is \(C\) groups (clusters) — firms in industries, students in schools, individuals in villages — \[y_{i,c}=\mathbf{x}_{i,c}'\bbeta+\varepsilon_{i,c},\] with errors correlated within a cluster but independent across clusters — off-diagonal structure in \(\bD\), block by block, from how the sample was drawn. OLS stays unbiased when \(\E[\beps_c\mid\bX_c]=\bzero\), which requires within-cluster interactions to be in the model (a pupil’s outcome must not depend on classmates’ regressors).

  • Within a school, pupils share teachers, neighborhood, shocks — their errors move together.
  • Even the White (heteroskedasticity-robust) errors are wrong here: they assume independence.
  • Ignoring clustering understates standard errors, often badly.

The cluster-robust variance

Theorem 4 (Cluster-robust covariance) With \(\bX_c,\beps_c\) the regressors and errors of cluster \(c\), OLS is unchanged, and \[\Var[\bb\mid\bX]=(\bX'\bX)^{-1}\Big[\textstyle\sum_{c=1}^{C}\bX_c'\,\bD_c\,\bX_c\Big](\bX'\bX)^{-1}, \qquad \bD_c=\E[\beps_c\beps_c'\mid\bX_c].\] It is estimated by the cluster sandwich, summing the cluster score vectors \(\bX_c'\hat{\be}_c\): \[\widehat{\Asyvar}[\bb]=(\bX'\bX)^{-1}\Big[\textstyle\sum_{c=1}^{C}(\bX_c'\hat{\be}_c)(\bX_c'\hat{\be}_c)'\Big](\bX'\bX)^{-1}.\] White’s estimator is the special case of one observation per cluster (\(\bD_c\) scalar).

The catch: asymptotics in the number of clusters

Under the Hood—Few clusters break the sandwich

The cluster-robust variance is consistent as the number of clusters \(C\to\infty\) — not as \(n\to\infty\). Stata’s finite-sample correction multiplies it by \(a_n=\dfrac{n-1}{n-k}\cdot\dfrac{C}{C-1}\) (the \(C/(C-1)\) helps when \(C\) is small; \(C=n\) recovers the i.i.d. factor). But no correction rescues small \(C\): \(C\) is the effective sample size — \(C=50\) clusters is like heteroskedasticity-robust inference with \(n=50\) observations, and unequal cluster sizes make it worse. Treat it as a small-sample problem, not a fixed reference distribution.

At what level to cluster?

The cluster level is a choice

The cluster level is a choice, and it is a genuine trade-off — not “always cluster more”.

  • Too fine (household instead of village): the variance is biased — it omits the very covariances clustering exists to capture — so standard errors are too small and significance is spurious.
  • Too coarse (state instead of village): unbiased, but the variance estimator is noisy — few, very unequal clusters.
  • So “the standard errors changed, therefore cluster” is not a valid rule: the change may be bias reduction (better) or sampling noise (worse). There is no clean answer.

When robust is not enough

The limits of a robust variance

The Four Questions—What robust standard errors do and do not fix
  • They fix the variance — the sandwich (White or cluster) gives valid standard errors under heteroskedasticity or clustering.
  • They do not fix inconsistency — if \(\E[\mathbf{x}\varepsilon]\neq\bzero\) (endogeneity, Lecture 10), no variance formula saves you; the point estimate is wrong.
  • They do not fix dependence you failed to model — the White errors are wrong under clustering; you must match the sandwich to the actual dependence structure.

Seeing it: coverage under heteroskedasticity

Does the interval cover 95%?

Figure 1: Empirical coverage of a nominal 95% confidence interval for the slope, over increasing heteroskedasticity (\(\sigma_i = x_i^{\gamma}\)). The conventional interval \(s^2(\mathbf{X}'\mathbf{X})^{-1}\) under-covers as \(\gamma\) grows; the heteroskedasticity-robust (HC1) interval stays near 0.95.

Summary

What to take away

  • Under a general \(\bD\), OLS is consistent and (with robust errors) correctly measured, but not efficient.
  • GLS is the efficient BLUE when \(\bD\) is known (rescale to constant variance, then OLS); WLS is the heteroskedastic case; FGLS estimates \(\bD\) from a skedastic regression.
  • But: FGLS buys efficiency by assuming the conditional mean is right — OLS + robust is the default.
  • Heteroskedasticity tests rarely change conduct; use robust errors regardless.
  • Clustering is dependence by design: use the cluster sandwich, mind the few-cluster problem.