Lecture 1
PIMES/UFPE
Econometrics is often introduced as the meeting point of economic theory, mathematics and statistics. That is true, and not very useful: it lists the ingredients without saying what is being cooked.
Econometrics is the study of how we learn about unknown parameters from data generated by a probabilistic process.
Not a catalogue of techniques. A few questions, asked repeatedly.
An economist reports that an additional year of schooling raises earnings by 8%.
Three things must be settled before that number means anything — and they are genuinely separate, which is why confusing them is the most common failure in applied work.
How are the estimator and the inference procedure actually carried out?
Not “which command do I type”, but: what sequence of numerical operations produces the number on the screen, and when does that sequence fail?
Important
If a parameter is not identified, collecting more data does not help. Not “helps slowly” — does not help.
The estimator may converge, beautifully and with tight standard errors, to the wrong thing.
Select a stage to see the question it asks and where the course answers it.
Estimation branches. There are two ways to build an estimator, and nearly everything in the course is one or the other. We meet both in the linear model, in Lecture 3, long before either has a name.
Asymptotic theory sits between estimation and inference, not off to the side. It is not a mathematical preliminary to be endured. It is the machinery that converts “here is an estimator” into “here is what we can say about the parameter”.
A data generating process is a complete probabilistic description of how the observed sample was produced: the distribution of the observable variables, and the relationship between observations.
Write \(W_i\) for the observables of unit \(i\) — in regression, \(W_i = (y_i, \mathbf{x}_i)\) — and assume \(\{W_i\}_{i=1}^n\) is drawn from a population distribution \(F\).
We never observe the DGP. It is an assumption we impose about how the data came about — one we choose, and could have chosen differently. That makes it sound optional, and it is not: the DGP is the only thing that gives the central vocabulary of this course any meaning at all.
Bias, consistency, standard error, confidence interval — every one of these is a statement about what would happen across repeated samples from the DGP. Remove the DGP and not one of those words can be defined, let alone computed.
Definition 1 (Estimand, estimator, estimate) The estimand \(\theta\) is the unknown feature of the population — a fixed, non-random functional of \(F\).
An estimator \(\hat\theta_n = \hat\theta(W_1,\dots,W_n)\) is a rule mapping samples to numbers. Because the sample is random, the estimator is a random variable.
An estimate is the number the estimator returns on the sample you have. It is not random.
The single most important idea in the course is that the estimator has a distribution. We see one sample and compute one number, but that number is a draw from a distribution generated by the DGP — and it is that distribution, not the number, which every claim about bias, precision and significance refers to.
Since we only ever observe one draw, the distribution is invisible. Simulation is the one device that can show it to us: fix a DGP, draw from it repeatedly, and watch what the estimator does.
Figure 1: Each grey line is the fitted regression from a different sample of size \(n = 40\) drawn from the same population; the dashed line is the population relationship.
Collapsing the bundle of lines to the slope alone gives the object we will spend Lectures 4 to 6 characterising: the sampling distribution of a single coefficient.
Figure 2: The sampling distribution of \(\hat\beta_1\) from the simulation above, with the estimand marked.
The third is the remarkable one. A single sample carries enough information to estimate the spread of a distribution we can never observe.
New Jersey raised its minimum wage in 1992; Pennsylvania did not. Comparing employment changes in fast-food restaurants across the border, Card and Krueger (1994) found no evidence of employment losses.
The paper is famous for its conclusion. It belongs in the first lecture for a different reason: the estimation is arithmetic. What made it publishable was the case that the comparison isolates the effect of the policy.
Quarter of birth affects how much schooling a person completes, through compulsory attendance laws — but there is no reason it should affect earnings in any other way (Angrist and Krueger 1991). That second claim is what makes the first one useful.
Lectures 10 and 11 are about turning arguments of this shape into estimators, and about how fragile they can be.
Theoretical econometrics develops estimators and establishes their properties. Applied econometrics uses them to answer economic questions.
This course is aimed at the applied economist — and it is a theoretical course, because the applied economist who does not know why an estimator works cannot tell when it stops working.
Not that you can say “I know OLS, IV, logit and panel data.”
That, faced with an empirical question, you can think:
I have a DGP and a parameter of interest. I need to establish whether that parameter is identified, choose an estimation strategy suited to it, work out the properties of the resulting estimator, and construct inference valid under the assumptions I am actually willing to make.
| Hansen | Chapter 1 |
| Greene | Chapter 1 |