9  Difference-in-Differences Fundamentals

9.1 Introduction

Difference-in-differences (diff-in-diff) is one of the most widely used quasi-experimental methods inside and outside of economics, though that was not always the case. The figure below shows a steady rise in the mentions of diff-in-diff, a trajectory mapped by Currie, Kleven, and Zwiers (2020) using the universe of working papers at the National Bureau of Economic Research (NBER) and top economics journals. This figure captures diff-in-diff’s momentum within economics since the early 1980s, when it began being used more frequently to study topics in labor economics like the returns to job training programs (O. Ashenfelter and Card 1985).

The method is much older than the figure below suggests. That figure only shows mentions of it within economics, but the method was not originally developed within economics. It was born in medicine in the 19th century. Not to be too melodramatic, but diff-in-diff has had many lives and many deaths like the Phoenix—a mythical bird that dies only to be reborn in new times and places. Each time diff-in-diff was born, it was brought forth by a researcher attempting to decipher data in a way that was transparent and that could easily be communicated to policymakers. Those two features seem to be common in each part of its life cycle, and through these cycles of death and rebirth, researchers’ understanding of its strengths and limitations, as well as how best to use it, was honed and improved upon, which has contributed to making it what it is today—the preferred method by many for working with longitudinal data conducting causal inference.

Figure 9.1: (Currie, Kleven, and Zwiers 2020) display of growing popularity of diff-in-diff since 1980 using percent of NBER papers and percent of top 5 journals in economics that mention it.

Ignaz Semmelweis in Vienna

The Vienna General Hospital was a teaching hospital offering free childbirth services. In the 1840s, the hospital had two maternity clinics: the First Clinic, staffed by male doctors and medical students, and the Second Clinic, run by female midwives. And, hidden inside these clinics was a troubling, perplexing mystery. Maternal mortality rates in the First Clinic were nearly three times higher than in the Second Clinic, primarily due to postpartum infections. This disparity seemed inexplicable, in part because it was the same hospital. There were differences between the clinics, but the differences could not have caused the disparities—or so said the primitive theories about disease that people believed at the time.

The clinics differed in two ways—how women were assigned to either clinic, and what happened in the delivery rooms of each clinic. First, depending on the day and time that women arrived in labor, they would be directed to either Clinic 1 or 2. This was by law. Years earlier, Vienna had adopted a policy requiring women to be admitted to either the First or Second Clinic depending on their arrival time, which we now know was randomizing women in labor to either clinic. Thus, since no one was choosing their clinic, but rather chance was choosing it for them, the differences in maternal mortality could not be attributed to pre-existing conditions. It had to be something in one of the two clinics, but what, and why?1

Ignaz Semmelweis was a Hungarian physician in the 1840s who served as chief resident at the Vienna General Hospital. Semmelweis, from what I could gather reading his book, The Etiology, Concept, and Prophylaxis of Childbed Fever (Semmelweis 1861), was a meticulous observer and a careful thinker. His attention to every possible detail related to every scrap of information around him that might be relevant to understanding the disease was impressive. Every possible factor would be scrutinized and then systematically eliminated as a potential cause for the disparity in death rates. His deduction happened by chance when one of his colleagues developed symptoms identical to those of the infected mothers and subsequently died. This heartbreaking event appeared to set Semmelweis’s theory into motion. He noted that the hospital, as a teaching institution, was both delivering children for women in labor and a medical school to train future physicians, and part of their training was in anatomy.

Physicians and students in the First Clinic conducted autopsies on cadavers as part of their education and then moved directly to assist in childbirth without having disinfected themselves thoroughly. As the midwives did not partake in the anatomy classes, they did not handle the cadavers. Semmelweis hypothesized that doctors and students were inadvertently transferring “cadaverous particles” to women during labor, causing deadly infections. Although he did not advance an actual germ theory of disease, he correctly intuited that some harmful substance was being transmitted from the cadavers to patients, and though it was invisible, it was nonetheless lethal. And so, both to falsify his theory and address the problem at hand, he implemented an unprecedented new policy: he required all doctors in the First Clinic to wash their hands with a chlorine solution before assisting in childbirth.

The results were staggering. The figure below shows the time series for the two clinics’ maternal mortality rates before and after the handwashing rule was implemented in May 1847.2 Maternal mortality rate in the First Clinic, which had been four times higher than that of the Second Clinic one year before the policy, plummeted to the level of the Second Clinic immediately. This sharp decline confirmed his hypothesis.

Figure 9.2: Time series plot of maternal mortality rates (%) by clinic from 1841 to 1858. Clinic 1 represents physicians and midwives, while Clinic 2 represents midwives only. The dashed vertical line marks May 1847, when handwashing began in Clinic 1.

But, despite the overwhelming success of the intervention, Semmelweis struggled to convince his peers. Colleagues attributed the drop in mortality to unrelated factors, such as improved ventilation, and dismissed the idea that unseen “particles” could carry disease. A strongly entrenched theory about disease transmission at the time made such a pathway impossible, so the evidence Semmelweis presented was not persuasive. Isolated and increasingly agitated, Semmelweis clashed with his colleagues, leading to his alienation from the medical community. Tragically, his mental health deteriorated, and he was eventually committed to a mental institution by a close friend, where he died from injuries sustained during his confinement, his work largely unrecognized and unappreciated.

John Snow’s Grand Experiment

In the 19th century, very little was known about cholera other than it killed its hosts quickly and painfully. Once a person became infected, they would die usually within days from constant vomiting and acute diarrhea leading to severe dehydration. It claimed tens of thousands of lives in London over several epidemic waves in the early to mid-1800s. At the time, the medical community attributed the disease to something called miasma, which was a theory that disease traveled through poisons found in foul-smelling air. Microscopes were still too rudimentary to see microorganisms, so invisible agents in water were unimaginable. What was not unimaginable was the stink. Fueled by London’s industrial revolution and expanding population, the city reeked with odors from the factories, sewage, and waste runoff into the Thames River. Miasma seemed plausible and efforts to clean up the city in direct response to it were probably, ironically, also helpful at improving public health—just not helpful at addressing cholera epidemics.

Diff-in-diff dies in Vienna and is then reborn almost immediately to a London resident named John Snow. Snow was a physician who had been on the frontlines of these outbreaks and was skeptical that miasma was causing the epidemics (Freedman 1991; Johnson 2007; Coleman 2019). Observing puzzling patterns of transmission, he found evidence that didn’t fit the miasma explanation. For example, cholera followed trade routes, seeming to affect sailors who had visited cholera-infected ports but sparing those who hadn’t. Poorer areas with poor sanitation were hardest hit, but some buildings and even adjacent neighborhoods showed dramatically different infection rates. Sometimes a building right next to another would go untouched while the first one suffered considerable casualties. These inconsistencies raised doubts in Snow’s mind about miasma as a comprehensive explanation.

Snow’s alternative hypothesis was original and correct: he posited that cholera was caused by a living organism that entered the body through food and drink, flowed through the alimentary canal, and returned to the water supply through waste, creating a vicious cycle. His ability to test this theory came about through a stroke of luck. By the mid-19th century, two water companies—Lambeth Waterworks Company and Southwark and Vauxhall Waterworks Company—both drew water from the Thames downstream from the city center. Between 1849 and 1854, London required the water utility companies to move their pipes upstream. Lambeth moved its intake pipes upstream, above the city and beyond the polluted sewage discharge points, but by 1854, Southwark and Vauxhall had not yet done so and households were still drawing their water from the river downstream. By the 1854 epidemic, Snow had access to a natural experiment: households served by Lambeth were receiving cleaner water, while those served by Southwark and Vauxhall were drinking contaminated water from the same neighborhoods.

Snow meticulously gathered data on cholera mortality rates, and then went door to door in the neighborhoods served by the two water companies to inquire about which utility company serviced that building. Sometimes the families knew and would share with him that information, and sometimes they didn’t. But because in 1854 the water source for Lambeth was further up the river, the salt content was different, so he would collect water samples from the tap and go back to his office to test it. His concern for data quality is even more impressive than his command over research design.

His analysis showed striking contrasts: cholera mortality rates in Lambeth households saw significant declines from 1849 to 1854 compared to those in Southwark and Vauxhall households. Snow’s landmark manuscript On the Mode of Communication of Cholera (Snow 1855), presented this evidence in a series of tables linking cholera transmission to contaminated water, and these tables and the logic behind them can now be recognized as a precursor to diff-in-diff.

But like Semmelweis, his compelling evidence did not persuade policymakers nor did it appear to persuade people in medicine. Miasma theory was so deeply entrenched that it was just not possible for even compelling evidence to crack it open.3 And the fact that our work can be both ignored and yet important is probably itself one of the most important lessons to be learned from that episode. But we also learn more than the importance of fighting for the sake of fighting, because Snow, like Semmelweis, also left a legacy of being open-minded, deeply attentive to the details, meticulous, and concerned about data quality, and answering one’s own questions using careful research designs.

It would be over a century before diff-in-diff was brought back to life, this time by a labor economist at Princeton in the early to mid-1970s.

9.2 Four Averages and Three Subtractions

In this section, we will introduce diff-in-diff, not in a regression model, but as “four averages and three subtractions.” Once I have done everything I can to explain the logic of diff-in-diff using only four averages and three subtractions, we will move into regression models. But to ensure everyone is on the same page, I will start first with a simple table that explains the logic of diff-in-diff, then introduce potential outcomes notation, then introduce a regression model, and then conduct some analysis with code. My goal is to make the fundamentals of diff-in-diff as accessible as they can be, from as many angles as possible, so that no one gets left behind. This material forms the core of the method, which will carry us through the more technical and complex designs, including covariates and differential timing.

Orley Ashenfelter Goes to Washington

The third birth of diff-in-diff happens when Orley Ashenfelter, a first-generation quantitative labor economist (Card and Farber 2005), made it a foundational tool for contemporary policy evaluation through his own applications of it to studying job training programs in the 1970s and 1980s. And to help us see the connection between diff-in-diff and the broader Mississippi River metaphor, I have updated our “Two Rivers into Causal Inference” graphic to emphasize Ashenfelter, as well as his student and coauthor who is also closely associated with the design, the Nobel Laureate David Card.

TODO Two rivers into causal inference

After completing his PhD at Princeton in 1968, Orley Ashenfelter joined the Office of Evaluation in Washington, DC, where he was tasked with studying job training programs. The role gave him access to a rare longitudinal dataset on labor market outcomes, allowing him to track participants’ earnings over time.

His approach echoed the logic of Semmelweis and Snow: observe two groups—those exposed to a treatment and those not—and track their outcomes over time. But unlike Semmelweis or Snow, Ashenfelter used panel data and estimated fixed effects models with time and thousands of individual controls to study the program’s effect on earnings.

Yet, like Semmelweis and Snow, Ashenfelter wasn’t just trying to convince other researchers. He wanted to communicate results to policymakers—people without the background or patience for technical terms like “regression.” It was in that context, facing the rhetorical demands of DC bureaucrats, that Ashenfelter began shifting his language away from econometric jargon and toward something simpler: difference-in-differences.

Illustrating Diff-in-Diff with a Table

Returning to Snow’s grand experiment, recall that between 1849 and 1854, the Lambeth Waterworks Company relocated its intake pipes upstream in compliance with London policy, while the Southwark and Vauxhall Waterworks Company did not (Johnson 2007). I’ll illustrate this design with Snow’s study using the following table. In Table 9.1, the quasi-experiment is organized as four averages (cholera mortality rates) and three subtractions. My notation will be simple. I’ll use \(Y\) as a measure of the average cholera mortality in each period (1854, or “after,” and 1849, or “before”) for Lambeth (our treatment group) and Southwark and Vauxhall (our comparison group).

Cholera mortality is expressed as the sum of different things on the right side of the equal sign. In 1849, before Lambeth moved its pipes, average cholera mortality in Lambeth will be \(Y = L\), and for Southwark and Vauxhall, it’s \(Y = SV\). In 1854, “after,” we’ll represent changes in cholera mortality for Lambeth as \(L + \mathbf{D + L_t}\). Here, \(\mathbf{L_t}\) is our counterfactual trend—representing what we expect Lambeth’s cholera mortality would have been had they not moved their pipes. The variable \(\mathbf{D}\) stands for the effect of Lambeth’s pipe relocation on cholera mortality, but only for Lambeth, and only afterwards, as that is the only row in which it appears.

Table 9.1: Lambeth and Southwark and Vauxhall, 1849 and 1854
Companies Time Average mortality \(D_1\) \(D_2\)
Lambeth Before \(Y = L\)
After \(Y = L + \mathbf{L_t + D}\) \(\mathbf{L_t + D}\)
\(\mathbf{D} + (\mathbf{L_t} - SV_t)\)
Southwark and Vauxhall Before \(Y = SV\)
After \(Y = SV + SV_t\) \(SV_t\)

Diff-in-diff calculations involve only two steps.

  1. Subtract “after” from “before” for each group, yielding Lambeth’s \(D_1\) and Southwark and Vauxhall’s \(D_1\). These first two subtractions isolate the changes in cholera mortality from 1849 to 1854 for each company, leaving only the trends and the causal effect.

  2. Calculate the third subtraction. Subtract Southwark and Vauxhall’s \(D_1\) from Lambeth’s \(D_1\). This yields the “difference-in-differences” estimate, \(D_2 = \mathbf{D} + (\mathbf{L_t} - SV_t)\).

In the final step, the difference-in-differences gives us not the isolated effect \(\mathbf{D}\) alone but the sum of three terms. The observed difference, \(\mathbf{L_t} - SV_t\), represents any counterfactual trend, so the validity of the diff-in-diff as the true effect depends on whether the trend in cholera mortality for Lambeth, had it not moved its pipe, would match the trend we observed in Southwark and Vauxhall. When that assumption holds, \(D\) can be estimated directly.

Snow’s presentation of results was not as straightforward as our diff-in-diff table here; his tables covered many neighborhoods, requiring a close analysis of several comparisons. However, we can reconstruct a modified version of his Table XII, focusing on his findings for these two water companies, as shown in Table 9.2. Here, population-adjusted cholera mortality is displayed per 10,000 households. In Southwark and Vauxhall, the mortality rate remained high between 1849 and 1854, whereas in Lambeth it dramatically decreased.

Following our diff-in-diff framework of four averages and three subtractions, we observe a reduction in cholera mortality of 97 per 10,000 households, assuming \(\mathbf{L_t} = SV_t\). This number, –97, is Snow’s estimate of the average effect of moving the pipe on cholera mortality and it suggests that cleaner water may have prevented an additional 97 deaths per 10,000 households.

Table 9.2: Average Effect of Lambeth’s Pipe Relocation on Cholera Mortality per 10,000 Households (Modified Snow’s Table XII)
Companies Time Average mortality \(D_1\) \(D_2\)
Lambeth Before 62
After 14 -48
-97
Southwark and Vauxhall Before 565
After 614 49

Using the data in Table 9.2 and applying Orley’s “four averages and three subtractions,” we estimate that by moving its pipe upstream, Lambeth saw 97 fewer cholera deaths per 10,000 households in 1854 than would have occurred had the pipe remained downstream—so long as the counterfactual trend, \(\mathbf{L_t}\), equals the actual comparison group trend, \(SV_t\). The diff-in-diff approach is straightforward yet powerful. It relies on these simple averages and subtractions to yield causal insights, provided that our comparison group accurately represents the counterfactual outcome in the treatment group.

Illustrating Diff-in-Diff with a Regression

For a very long time, I think most people thought that the phrase “difference-in-differences” was a synonym for two-way fixed effects (TWFE). Why did they think that? Well, the truth is quite possibly a surprise to many. The reason that diff-in-diff and TWFE regressions were thought to be the same thing is because there are, in fact, three specific regression formulas that are numerically identical to calculating four averages and three subtractions when working with only two groups and two time periods (Baker et al. 2025).

Equation 9.1 presents the “four averages and three subtractions” equation, or what is more often now called simply the \(2 \times 2\) calculation (Goodman-Bacon 2021; Baker et al. 2025). Starting now, and for the remainder of this chapter and the next chapter, you will see me oscillate between referring to Equation 9.1 as “four averages and three subtractions” and the simple \(2 \times 2\). I like to use both terms interchangeably so that readers can constantly be reminded that at its core, diff-in-diff is nothing more than four averages and three subtractions, to quote Orley Ashenfelter, but as that’s a mouthful, and keeping with the emerging nomenclature in the econometrics of diff-in-diff, I also want you to equate that with the phrase “\(2 \times 2\).” So, let’s now look at Equation 9.1 so we can connect it with regression. \[ \begin{equation} \widehat{\mathbf{\delta}}_{DiD} = \bigg ( \overline{y}_k^{post(k)} - \overline{y}_k^{pre(k)} \bigg ) - \bigg ( \overline{y}_U^{post(k)} - \overline{y}_U^{pre(k)} \bigg ) \label{eq:4averages} \end{equation} \tag{9.1}\] Here, \(k\) represents a collection of units treated at the same time, and \(U\) is a group of units who were not treated in either the preperiod or the postperiod.

Now consider the following regression specification, shown in Equation 9.2: \[ \begin{equation} Y_{ist} = \alpha_0 + \alpha_1 Treat_{is} + \alpha_2 Post_{t} + \mathbf{\delta} (Treat_{is} \times Post_t) + \varepsilon_{ist} \label{eq:ols_did} \end{equation} \tag{9.2}\]

The key question is whether the estimated coefficient \(\widehat{\delta}\) in Equation 9.2 will match \(\widehat{\mathbf{\delta}}_{DiD}\) in Equation 9.1.

To explore this, let’s put both equations to the test using real data. We’ll draw on data from Cheng and Hoekstra (2013), which provides a balanced panel of 50 states from 2000 to 2010. During this period, various states passed “stand your ground” laws allowing lethal force in self-defense beyond the home. But for simplicity, and so it fits with the simple 2 \(\times\) 2 calculations we’re focused on in this chapter, I’ll restrict the sample to only the states that passed the law in 2006 plus the states that never passed it during this period. In other words, I’m going to drop all states from the sample that passed the law in 2005, 2007, 2008, and 2009. I will also look only at the effect of the reforms on homicide, since that is the focus of Cheng and Hoekstra (2013).

equivalence.do

Code
* equivalence.do.  Showing that a particular OLS specification is numerically equivalent to "four averages and three subtractions". Data from Cheng and Hoekstra (2013) castle doctrine study. 

clear 
capture log close

use https://github.com/scunning1975/mixtape/raw/master/castle.dta, clear
xtset sid year

* Keep only states treated in 2006 and those never treated
drop if effyear==2005 | effyear==2007 | effyear==2008 | effyear==2009

* Create dummies for treatment states and post treatment period
drop    post
gen     post = 0
replace post = 1 if year>=2006

gen     treat = 0
replace treat = 1 if effyear==2006

* Four averages
egen y11 = mean(l_homicide) if post==1 & treat==1
egen y10 = mean(l_homicide) if post==0 & treat==1
egen ey11 = max(y11)
egen ey10 = max(y10)

egen y01 = mean(l_homicide) if post==1 & treat==0
egen y00 = mean(l_homicide) if post==0 & treat==0
egen ey01 = max(y01)
egen ey00 = max(y00)

* Diff-in-diff equation as "four averages and three subtractions"
gen did = (ey11 - ey10) - (ey01 - ey00) // (1.798907 - 1.78607 ) - (1.196448 - 1.251846)
sum did // 0.0682359
di (1.798907 - 1.78607 ) - (1.196448 - 1.251846) // 0.068235

* OLS regression model equivalence
xi: reg l_homicide post treat i.post*i.treat, cluster(state) // 0.0682359

equivalence.R

Code
################################################################################
# name: equivalence.R
# description: show that did equation is numerically equivalent to OLS specification
################################################################################

# Load necessary libraries
# install.packages(c("tidyverse", "fixest", "haven"))
library(tidyverse)
library(fixest)
library(haven)

# Load the data
data <- haven::read_dta("https://github.com/scunning1975/mixtape/raw/master/castle.dta")

# Filter the data
data <- data %>%
  filter(!(effyear %in% c(2005, 2007, 2008, 2009)))

# Generate post and treat variables
data <- data %>%
  mutate(
    year = as.numeric(year),
    post = ifelse(year >= 2006, 1, 0),
    treat = ifelse(!is.na(effyear), 1, 0)
  )

# Calculate means for different groups
(mean_values <- data %>%
  group_by(post, treat) %>%
  summarise(mean_l_homicide = mean(l_homicide, na.rm = TRUE)))

# Calculate the DiD manually
(did <- with(mean_values, {
  (mean_l_homicide[post == 1 & treat == 1] - mean_l_homicide[post == 0 & treat == 1]) - 
  (mean_l_homicide[post == 1 & treat == 0] - mean_l_homicide[post == 0 & treat == 0])
}))


# Run the regression 
model <- feols(
  l_homicide ~ i(post) + i(treat) + post:treat, 
  data = data, cluster = ~ state
)

model

Below in Table 9.3 I present the calculation I did in Equation 9.1 as well as the OLS estimate of \(\widehat{\delta}\) from Equation 9.2. And notice—the numbers are identical.

Table 9.3: Estimates of Castle Doctrine Reform on Log Homicides Presented Using Four Averages and Three Subtractions and OLS
Four Averages and Three Subtractions OLS Estimate
DiD Estimate 0.0682359 0.0682359

To see why this particular OLS specification is identical to the four averages and three subtractions calculation in Equation 9.1, let’s calculate the four averages and three subtractions ourselves using a regression equation so we can see for ourselves that the two are the same. The OLS equation again is: \[ \begin{equation} Y_{ist} = \alpha_0 + \alpha_1 Treat_{is} + \alpha_2 Post_{t} + \mathbf{\delta} (Treat_{is} \times Post_t) + \varepsilon_{ist} \end{equation} \tag{9.3}\]

If we estimate this equation using OLS, then we get fitted values for \(\widehat{\alpha_0}\), \(\widehat{\alpha_1}\), and \(\widehat{\delta}\). Below are calculated sample mean log homicides, \(\overline{Y}\), equal to the coefficients estimated with OLS to illustrate what I mean. For each average, you simply sum the fitted values from the previous equation that correspond to that group and time period.

  1. Non-Reform States Pre: \(\overline{Y}_{NR,Pre} = \widehat{\alpha_0}\)

  2. Non-Reform States Post: \(\overline{Y}_{NR,Post} = \widehat{\alpha_0} +\widehat{\alpha_2}\)

  3. Reform States Pre: \(\overline{Y}_{R,Pre} = \widehat{\alpha_0} +\widehat{\alpha_1}\)

  4. Reform States Post: \(\overline{Y}_{NR,Post} = \widehat{\alpha_0} + \widehat{\alpha_1} + \widehat{\alpha_2} + \widehat{\delta}\)

When we plug those values into the diff-in-diff equation, we get this: \[ \begin{eqnarray*} \widehat{\delta}_{DiD} &=& \bigg (\overline{Y}_{NR,Post} - \overline{Y}_{R,Pre} \bigg ) - \bigg ( \overline{Y}_{NR,Post} - \overline{Y}_{NR,Pre} \bigg ) \\ &=& \bigg ( [ \widehat{\alpha_0} +\widehat{\alpha_1} +\widehat{\alpha_2} + \widehat{\delta} ] - [ \widehat{\alpha_0} + \widehat{\alpha_1} ] \bigg ) - \bigg ( [\widehat{\alpha_0} + \widehat{\alpha_2}] - [\widehat{\alpha_0}] \bigg ) \\ &=& \bigg ( \widehat{ \alpha_2} + \widehat{\delta} \bigg ) - \bigg ( \widehat{\alpha_2} \bigg ) \\ &=& \delta \end{eqnarray*} \tag{9.4}\]

The \(\widehat{\delta}\) in Equation 9.2, when estimated with OLS, is the exact same calculation as the four averages and three subtractions method in Equation 9.1. It’s no wonder then, given the equivalence of the two methods, that Orley would’ve chosen the “four averages and three subtractions” to explain the results of his analysis of job training programs to people in DC who did not have economics or statistics backgrounds. It’s not that he was dumbing it down. Rather, it’s that regression coefficient is literally four averages and three subtractions, so that’s what he chose to convey.

I said earlier that there are in fact three regression specifications that calculate “four averages and three subtractions,” but I have just given us only one. What then are the other two? Equations Equation 9.5, Equation 9.5, and Equation 9.5 list all three regression specifications that, curiously enough, are numerically identical to the “four averages and three subtractions” equation from Equation 9.1. \[ \begin{eqnarray} Y_{it} &=& \alpha + \beta_1 \text{Post}_t + \beta_2 \text{Treat}_i + \delta (\text{Post}_t \times \text{Treat}_i) + \varepsilon_{it} \label{eq:did_interaction} \\ Y_{it} &=& \alpha_i + \gamma_t + \delta (\text{Post}_t \times \text{Treat}_i) + \varepsilon_{it} \label{eq:did_twfe} \\ \Delta Y_i &=& \alpha + \delta \text{Treat}_i + \varepsilon_i \label{eq:did_longdiff} \end{eqnarray} \tag{9.5}\] where Equation 9.5 first calculates the difference between the pre- and post-treatment outcomes for each unit, called the “long difference,” which is here represented with \(\Delta Y_i\), and then regresses the long difference onto a treatment dummy. The code for these three regressions is listed below in both Stata and R.

equivalence2.do

Code
* name: equivalence.do
* author: scott cunningham
* description: OLS and Manual are the same

clear 
capture log close

use https://github.com/scunning1975/mixtape/raw/master/castle.dta, clear
xtset sid year

drop if effyear==2005 | effyear==2007 | effyear==2008 | effyear==2009

drop    post
gen     post = 0
replace post = 1 if year>=2006

gen     treat = 0
replace treat = 1 if effyear==2006

keep if year==2006 | year==2005

* Example 1: OLS regression with interactions
reg l_homicide post##treat, cluster(sid)

* Example 2: Twoway fixed effects (state and year fixed effects)
xtreg l_homicide c.treat#c.post i.year, fe vce(cluster sid)

* Example 3: Regress "long difference" onto treatment dummy
preserve
    keep sid year l_homicide prison treat
    reshape wide l_homicide prison, i(sid) j(year)
    gen diff = l_homicide2006 - l_homicide2005
    reg diff treat, vce(cluster sid)
restore

equivalence2.R

Code
# name: equivalence2.R
# author: scott cunningham  
# description: OLS and Manual are the same

# Load required libraries
library(haven)
library(dplyr)
library(fixest)
library(tidyr)

# Clear workspace and load data
rm(list = ls())

# Load the castle dataset
castle <- read_dta("https://github.com/scunning1975/mixtape/raw/master/castle.dta")

# Set up panel structure (equivalent to xtset)
castle <- castle %>%
  arrange(sid, year)

# Drop specific years and create variables
castle <- castle %>%
  filter(!(effyear %in% c(2005, 2007, 2008, 2009))) %>%
  select(-post) %>%
  mutate(
    post = ifelse(year >= 2006, 1, 0),
    treat = ifelse(effyear == 2006, 1, 0)
  ) %>%
  filter(year %in% c(2005, 2006))

# Example 1: OLS regression with interactions
cat("Example 1: OLS regression with interactions\n")
model1 <- feols(l_homicide ~ post * treat, 
                data = castle, 
                cluster = ~sid)
summary(model1)

# Example 2: Twoway fixed effects (state and year fixed effects)
cat("\nExample 2: Twoway fixed effects (state and year fixed effects)\n")
model2 <- feols(l_homicide ~ treat:post + factor(year) | sid, 
                data = castle, 
                cluster = ~sid)
summary(model2)

# Example 3: Regress "long difference" onto treatment dummy
cat("\nExample 3: Regress 'long difference' onto treatment dummy\n")

# Create the difference data (removing prison variable that doesn't exist)
diff_data <- castle %>%
  select(sid, year, l_homicide, treat) %>%
  pivot_wider(
    names_from = year,
    values_from = l_homicide,
    names_prefix = "l_homicide_"
  ) %>%
  mutate(
    diff = l_homicide_2006 - l_homicide_2005
  )

model3 <- feols(diff ~ treat, 
                data = diff_data, 
                cluster = ~sid)
summary(model3)

To summarize, I refer to the manual method of calculating four averages and three subtractions as the \(2 \times 2\). And then I refer to Equation 9.5 as the simple interaction OLS specification, Equation 9.5 as the TWFE OLS specification because it controls for individual and time fixed effects, and Equation 9.5 as the “long difference” OLS specification. All four of them, regardless of which one you choose, gives the exact same number–0.0682359–as seen in Table 9.4 below.

Table 9.4: Estimates of Castle Doctrine Reform on Log Homicides Presented Using Four Averages and Three Subtractions and OLS
\(2 \times 2\) Interaction OLS TWFE Long difference
DiD Estimate 0.0682359 0.0682359 0.0682359 0.0682359

Note: All four methods are numerically identical to one another.

This equivalence reveals something important about both statistical practice and communication. Since these OLS specifications are computationally equivalent to “four averages and three subtractions,” practitioners often prefer regression because it provides access to standard statistical inference tools like standard errors and hypothesis tests.

But Orley’s goal was different: communicating with nonpractitioners. The history of difference-in-differences has always emphasized clear communication—from Ignaz Semmelweis and John Snow onward, the goal was conveying urgent results to those outside the analytical weeds. Orley coined “difference-in-differences” as a truthful but accessible nickname for these regression specifications, avoiding unnecessary OLS jargon while preserving analytical integrity.

This highlights the hidden curriculum of causal inference: successful communication of complex ideas to nonpractitioners. Many of us learn this only after failed attempts to explain fixed effects to journalists or policymakers.4

9.3 Potential Outcomes and Identification

Diff-in-diff is a causal design, not simply four numbers subtracted in a row. What is it that allows one to go, therefore, from four averages and three subtractions to making a claim about causality? What’s the secret? To understand the causal interpretation of diff-in-diff, as we have done in the previous chapters, we have to introduce potential outcomes. Without potential outcomes, at least in the design tradition that this book is mostly focused on, we really can’t talk about causality. But fortunately, we get to keep our four averages and three subtractions when we do it.

No Anticipation Assumption

I am not a huge fan of the nickname of the next assumption as it immediately invokes human reasoning that is actually not what is required by the assumption. The second assumption is commonly called the no anticipation assumption, which implies that to use diff-in-diff, we must as social scientists stop believing that humans can hear of a future event and back out its implications for the present. But before we dive into that sort of thing, I think we should ground ourselves in what no anticipation literally means. All that no anticipation means is that the potential outcome of the treatment group prior to the treatment’s occurrence is \(Y^0\). No anticipation, in other words, is shorthand for “the treatment group is not treated until it is in fact treated."

To help us understand what no anticipation means, I am going to begin with its violation in the hopes that by seeing its violation, we might better understand that the no anticipation is obvious, mundane, and also crucial when designing any diff-in-diff. I’ll start with the diff-in-diff equation and then make a substitution with the switching equation that is different than we did before.

\[ \begin{eqnarray} \widehat{\delta}_{DiD} &=& \bigg ( E[Y_k|Post] - E[Y_k|Pre] \bigg ) - \bigg ( E[Y_U | Post ] - E[ Y_U | Pre] \bigg) \nonumber \\ &=& \bigg ( E[Y^1_k|Post] - \mathbf{E[Y^{\textbf{1}}_k|Pre]} \bigg ) - \bigg ( E[Y^0_U | Post ] - E[ Y^0_U | Pre] \bigg) \end{eqnarray} \tag{9.15}\]

Notice that in the second row of Equation 9.15, when I replaced \(Y\) with potential outcomes using the switching equation, I made \(Y=Y^1\) for the treatment group in both periods. That is what a no anticipation assumption violation means—it means that for your “pretreatment period,” the treatment group was already treated. Why would you ever use as your baseline, in diff-in-diff, an already treated period? That’s not our point yet—our point here is only to understand the consequences of that choice, not to prescribe whether it is good or bad to do so.

So, if you have a no anticipation assumption violation, then what exactly would diff-in-diff identify? Since the calculation is the same, does this equal the plus parallel trends, or is it something else, and if it is something else, what? To answer that question, I need to make a couple of substitutions. I’m going to add in two zeroes based on the differences in counterfactuals so that I better understand what diff-in-diff is identifying in this case. \[ \begin{eqnarray} &=& \bigg ( E[Y^1_k|Post] - \mathbf{E[Y^{\textbf{1}}_k|Pre]} \bigg ) - \bigg ( E[Y^0_U | Post ] - E[ Y^0_U | Pre] \bigg) \nonumber \\ &&+ \mathbf{E[Y^{\textbf{0}}_k|Post]} - \mathbf{E[Y^{\textbf{0}}_k|Post]} \nonumber \\ &&+ \mathbf{E[Y^{\textbf{0}}_k|Pre]} - \mathbf{E[Y^{\textbf{0}}_k|Pre]} \end{eqnarray} \tag{9.16}\]

Now let me move these terms in Equation 9.16 around so that I can better understand what diff-in-diff is identifying if our treatment group was already treated at baseline. \[ \begin{eqnarray} &=& \underbrace{\bigg ( E[Y^1_k|Post] - \mathbf{E[Y^{\textbf{0}}_k|Post]} \bigg )}_{\text{\emph{ATT} at post-treatment for group $k$}} \nonumber \\ &&+ \underbrace{\bigg ( \mathbf{E[Y^{\textbf{0}}_k|Post]} - \mathbf{E[Y^{\textbf{0}}_k|Pre]} \bigg ) - \bigg ( E[Y^0_U | Post ] - E[ Y^0_U | Pre] \bigg )}_{\text{Non-parallel trends bias}} \nonumber \\ && - \underbrace{\bigg ( \mathbf{E[Y^{\textbf{1}}_k|Pre]} - \mathbf{E[Y^{\textbf{0}}_k|Pre]} \bigg )}_{\text{\emph{ATT} at baseline for group $k$}} \label{eq:did_na3} \end{eqnarray} \tag{9.17}\]

And to simplify, I’ll group all of that into this simple expression: \[ \begin{equation} \widehat{\delta}_{DiD} = \mathit{ATT}_k(Post) + PT - \mathit{ATT}_k(Pre) \label{eq:did_na3a} \end{equation} \tag{9.18}\]

Equation 9.17 states that when you estimate diff-in-diff using two periods and two groups, and your treatment group is treated in the pretreatment period, the only way that this can identify the is if 1) parallel trends holds and 2) the for your treatment group at baseline was zero. This means either that the future treatment was a surprise—hence it was not “anticipated"—or that it was known ahead of time, but its treatment effect was zero at baseline. Either way, \(Y=Y^0\) for the treatment group in the preperiod, which is all that is implied by no anticipation.

I think it may be helpful to take this somewhat abstract decomposition from theory to simulated data. This simulation will create 1,000 firms with 25 per state, in 40 states, and over four years. It will then follow those 1,000 firms from 1990 to 1993. The treatment happens only to the treatment group in 1991, and once treated, the firms stay treated. This means that only 1990 is untreated for the treatment group, so if I wanted to estimate the using diff-in-diff, I would need to use 1990 as my baseline.

I generated two types of treatment effects. The first, labeled delta_c in the code, is equal to 10 in 1991, 1992, and 1993. This therefore means that the is 10. These are the constant treatment effects. Then I generated a second type of treatment effects labeled delta_d equal to 10 in 1991, 20 in 1992, and 30 in 1993. This meansthat the is 20 in this case. These are dynamic treatment effects. As Equation 9.17 says that diff-in-diff will be zero if treatment effects are constant, but nonzero if dynamic, we will want to check each to confirm.

na.do

Code
* name: na.do

clear
capture log close
set seed 20200403

* First create the states
set obs 40
gen state = _n

* Finally generate 1000 firms.  These are in each state. So 25 per state.
expand 25
bysort state: gen firms=runiform(0,5)
label variable firms "Unique firm fixed effect per state"

* Second create the years
expand 4
sort state
bysort state firms: gen year = _n
gen n=year

replace year = 1990 if year==1
replace year = 1991 if year==2
replace year = 1992 if year==3
replace year = 1993 if year==4
egen id =group(state firms)

* Treatment group is upper half
gen     group=0
replace group=1 if id >= 500

* Correct start date so that NA is satisfied
gen     post=0  
replace post=1 if year >= 1991

* Incorrect start date so that NA is violated
gen     post_na=0
replace post_na=1 if year>=1992

* Data generating process
gen e   = rnormal(0,1)

* Potential outcomes
gen     y0 = firms + n + e 

* Constant treatment effects
gen     y1_c = y0
replace y1_c = y0 + 10 if year>=1991

* dynamic treatment effects
gen     y1_d = y0 
replace y1_d = y0 + 10 if year==1991
replace y1_d = y0 + 20 if year==1992
replace y1_d = y0 + 30 if year==1993

* Treatment effects
gen     delta_c = y1_c - y0
gen     delta_d = y1_d - y0

su delta_c if year>=1991 & group==1
su delta_d if year>=1991 & group==1

* Treatment period
gen     d = 0
replace d = 1 if year>=1991 & group==1

* Switching equation for constant
gen     y_c = d*y1_c + (1-d)*y0
gen     y_d = d*y1_d + (1-d)*y0

* Aggregate causal parameters
egen att_c = mean(delta_c) if year>=1991 & group==1

egen att_d = mean(delta_d) if year>=1991 & group==1

su att_c att_d

* Correct specification
regress y_c group##post, robust // did = 10
regress y_d group##post, robust // did = 20

* Incorrect specification
regress y_c group##post_na if year>=1991, robust // did = 0
regress y_d group##post_na if year>=1991, robust // did = 15

na.R

Code
# Set seed for reproducibility
library(sandwich)
library(lmtest)

set.seed(20200403)

# Create the base data frame
n_states <- 40
n_firms_per_state <- 25
n_years <- 4

# Create expanded data frame for states and firms
states_rep <- rep(1:n_states, each = n_firms_per_state)
firms_data <- data.frame(
  state = states_rep,
  firms = runif(n_states * n_firms_per_state, 0, 5)
)

# Now expand for years
firms_data <- firms_data[rep(seq_len(nrow(firms_data)), each = n_years), ]
firms_data <- firms_data[order(firms_data$state), ]

# Create year variable
firms_data$year <- rep(1:4, times = nrow(firms_data)/4)
firms_data$n <- firms_data$year

# Replace years with actual dates
firms_data$year <- ifelse(firms_data$year == 1, 1990,
                          ifelse(firms_data$year == 2, 1991,
                                 ifelse(firms_data$year == 3, 1992, 1993)))

# Create unique ID for each firm
firms_data$id <- as.numeric(factor(paste(firms_data$state, firms_data$firms)))

# Treatment group assignment (upper half of IDs)
firms_data$group <- ifelse(firms_data$id >= 500, 1, 0)

# Create post indicators
firms_data$post <- ifelse(firms_data$year >= 1991, 1, 0)
firms_data$post_na <- ifelse(firms_data$year >= 1992, 1, 0)

# Generate error term
firms_data$e <- rnorm(nrow(firms_data), 0, 1)

# Generate potential outcomes
firms_data$y0 <- firms_data$firms + firms_data$n + firms_data$e

# Constant treatment effects
firms_data$y1_c <- firms_data$y0
firms_data$y1_c[firms_data$year >= 1991] <- 
  firms_data$y0[firms_data$year >= 1991] + 10

# Dynamic treatment effects
firms_data$y1_d <- firms_data$y0
firms_data$y1_d[firms_data$year == 1991] <- 
  firms_data$y0[firms_data$year == 1991] + 10
firms_data$y1_d[firms_data$year == 1992] <- 
  firms_data$y0[firms_data$year == 1992] + 20
firms_data$y1_d[firms_data$year == 1993] <- 
  firms_data$y0[firms_data$year == 1993] + 30

# Calculate treatment effects
firms_data$delta_c <- firms_data$y1_c - firms_data$y0
firms_data$delta_d <- firms_data$y1_d - firms_data$y0

# Create treatment indicator
firms_data$d <- ifelse(firms_data$year >= 1991 & firms_data$group == 1, 1, 0)

# Generate observed outcomes
firms_data$y_c <- firms_data$d * firms_data$y1_c + (1 - firms_data$d) * firms_data$y0
firms_data$y_d <- firms_data$d * firms_data$y1_d + (1 - firms_data$d) * firms_data$y0

# Calculate aggregate treatment effects
att_c <- mean(firms_data$delta_c[firms_data$year >= 1991 & firms_data$group == 1], 
              na.rm = TRUE)
att_d <- mean(firms_data$delta_d[firms_data$year >= 1991 & firms_data$group == 1], 
              na.rm = TRUE)

# Print summary of treatment effects
cat("Average Treatment Effects:\n")
cat("Constant ATT:", att_c, "\n")
cat("Dynamic ATT:", att_d, "\n\n")

# Fit regression models
# Correct specification - Constant Treatment Effects
model_c <- lm(y_c ~ factor(group) * factor(post), data = firms_data)
coeftest(model_c, vcov = vcovHC(model_c, type = "HC1"))

# Correct specification - Dynamic Treatment Effects
model_d <- lm(y_d ~ factor(group) * factor(post), data = firms_data)
coeftest(model_d, vcov = vcovHC(model_d, type = "HC1"))

# Incorrect specification - Constant Treatment Effects
model_c_na <- lm(y_c ~ factor(group) * factor(post_na), 
                 data = subset(firms_data, year >= 1991))
coeftest(model_c_na, vcov = vcovHC(model_c_na, type = "HC1"))

# Incorrect specification - Dynamic Treatment Effects
model_d_na <- lm(y_d ~ factor(group) * factor(post_na), 
                 data = subset(firms_data, year >= 1991))
coeftest(model_d_na, vcov = vcovHC(model_d_na, type = "HC1"))

I’ll present my results using the DiD regression specification from Equation 9.2, which I’ll rewrite here so you don’t have to flip back. \[ \begin{eqnarray*} Y_{ist} = \alpha_0 + \alpha_1 Treat_{is} + \alpha_2 Post_{t} + \mathbf{\delta} (Treat_{is} \times Post_t) + \varepsilon_{ist} \end{eqnarray*} \tag{9.19}\]

I then ran four regressions and present that output in Table 9.11. Column 1 uses the data that has constant treatment effects, which means the is 10. And in this specification I used as my baseline 1990, which was not yet treated. The coefficient in column 1 is 9.976, which is approximately correct. In column 2, I used the data that had dynamic treatment effects, which recall, had an of 20. And again, with the correct specification, my estimation was approximately correct again at 19.976.

But then in columns 3 and 4, I chose to intentionally specify the model where I used as my baseline 1991, which as we know was already treated. In 1991, the under constant treatment effects is 10. But in Equation 9.17, we know that if you use as your baseline a period that is already treated, and treatment effects are constant from pre to postperiod, then diff-in-diff under parallel trends will equal zero (i.e., in the post minus in the pre is zero if the is the same in both periods). And if you look closely at column 3, that’s exactly what I find.

In column 4, I ran the same specification, only I used the variable for which dynamic treatment effects is true. The treatment effect in 1991 is 10, and the treatment effect in 1992 is 20 and 30 in 1993, as I said. That means that the in the postperiod, defined as 1992 and 1993, is 25. But since it is 10 in the preperiod, then Equation 9.17 says that the diff-in-diff equation will be 25–10, or 15. And that’s what I found.

Table 9.11: Diff-in-Diff Results With Simulated Data With and Without No Anticipation Violation
(1) (2) (3) (4)
DiD coefficient 9.976 19.976 0.017 15.017
(0.131) (0.266) (0.136) (0.220)
\(ATT\) 10 20 10 20
Specification NA NA NA Violated NA Violated

Now that we see conceptually that diff-in-diff requires no anticipation, what exactly does that mean for practice? It’s very simple—you want to make sure that your baseline period in your diff-in-diff is not treated. If it is treated, then it will attenuate your results as just shown. How do you do that? The main way is to correctly date the treatment timing such that the preperiod was never treated. On the safe side, even if the group was treated only briefly, let that be the start of your first treatment period. Make your baseline completely clean so that you can avoid this problem.

It’s possible that no anticipation can be violated, too, if people are looking forward in time and in response to a future policy, and change their behavior at baseline. If they do, then it is as though they were treated at baseline. All of this is an extension of SUTVA in many ways, only no anticipation is limiting interference from the future to the past in this case. In such instances where you are truly concerned about people or organizations changing behavior (e.g., exiting markets, firing people, hiring people) in response to future policies, then the best bet would probably be to hedge and make the baseline the period before the law change was announced, as opposed to the period before the law change was enforced.

Just note, though, that when you do make a change like that, you are technically also making two other changes. First, the has changed because the is an average over all post-treatment periods, and by rolling back the treatment date, you have necessarily lengthened the post-treatment window to include a period of announcement but not yet enforcement. This is just something you want to be aware of as it changes interpretation.5 And it will change your parallel trends assumption because recall that your parallel trends assumption is always with respect to some fixed baseline, which you have now changed so as to satisfy no anticipation.

None of these are right or wrong in some abstract sense—they just are the case, and you want to be cognizant of them as you make these design decisions so that any decision you made, you made with your eyes wide open and without crossing your fingers. You want it to be that the decisions you make were intentional, and not made for you because of a misunderstanding.

Avoid Already-Treated Controls

The following is not an assumption so much as a strong recommendation. Researchers are strongly advised to not use an already-treated group as a control. And in many ways, a reader reading this might say “obviously I’m not going to use a treated group as a control because I need a control group to be, by definition, not treated.” That’s absolutely true, and you should listen to your gut about that. But the problem is that some statistical models that we will review behind the scenes actually make calculations that use already-treated groups as a control, but without telling us. In other words, some methods have secrets and skeletons in their closets and we need to get a handle now about why those secrets are or are not problematic.

What I want to do here is just walk you through the steps from first principles so that you can see for yourself, outside of any statistical model, why diff-in-diff with an already-treated group as a control is problematic. Let’s go through our series of steps starting with Equation 9.20, which defines the diff-in-diff equation and then makes substitutions to potential outcomes based on whether a unit is treated or not. \[ \begin{eqnarray} \widehat{\delta}_{DiD} &=& \bigg ( E[Y_k|Post] - E[Y_k|Pre] \bigg ) - \bigg ( E[Y_U | Post ] - E[ Y_U | Pre] \bigg) \nonumber \\ &=& \bigg ( E[Y^1_k|Post] - E[Y^0_k|Pre] \bigg ) - \bigg ( \mathbf{E[Y^{\textbf{1}}_U | Post ]} - \mathbf{E[Y^{\textbf{1}}_U | Pre]} \bigg) \nonumber\\ \label{eq:did_at1} \end{eqnarray} \tag{9.20}\]

I put the last two terms in bold to emphasize that they are in fact treated units. I’m going to now add zeroes, but this time I am going to add three zeroes: \[ \begin{eqnarray} &=& \bigg ( E[Y^1_k|Post] - E[Y^0_k|Pre] \bigg ) - \bigg ( \mathbf{E[Y^{\textbf{1}}_U | Post ]} - \mathbf{E[Y^{\textbf{1}}_U | Pre]} \bigg) \nonumber \\ &&+ \mathbf{E[Y^{\textbf{0}}_k|Post]} - \mathbf{E[Y^{\textbf{0}}_k|Post]} \nonumber \\ &&+ \mathbf{E[Y^{\textbf{0}}_U|Post]} - \mathbf{E[Y^{\textbf{0}}_U|Post]} \nonumber \\ &&+ \mathbf{E[Y^{\textbf{0}}_U|Pre]} - \mathbf{E[Y^{\textbf{0}}_U|Pre]} \label{eq:did_at2} \end{eqnarray} \tag{9.21}\]

And we’ll just rearrange that into a form that is more interpretable in terms of the and parallel trends plus any additional biases: \[ \begin{eqnarray} &=& \underbrace{E[Y^1_k|Post] - \mathbf{E[Y^{\textbf{0}}_k|Post]}}_{\text{\emph{ATT}}} \nonumber \\ &&+ \underbrace{\bigg (\mathbf{E[Y^{\textbf{0}}_k|Post]} - E[Y^0_k|Post] \bigg ) - \bigg ( \mathbf{E[Y^{\textbf{0}}_U|Post]} - \mathbf{E[Y^{\textbf{0}}_U|Pre]} \bigg )}_{\text{Non-parallel trends bias}} \label{eq:did_at3}\\ &&- \underbrace{\bigg [ \bigg ( \mathbf{E[Y^{\textbf{1}}_U | Post ]} - \mathbf{E[Y^{\textbf{0}}_U|Post]} \bigg ) - \bigg ( \mathbf{E[Y^{\textbf{1}}_U | Pre]} - \mathbf{E[Y^{\textbf{0}}_U|Pre]} \bigg ) \bigg ]}_{\text{Change in \emph{ATT} for group $U$ from pre to post}} \nonumber \end{eqnarray} \tag{9.22}\]

And now let me just rewrite this out for you to see this expressed without all that notation: \[ \begin{equation} \widehat{\delta}_{DiD} = \mathit{ATT}_k + PT - \Delta \mathit{ATT}_U \end{equation} \tag{9.23}\]

If you use an already-treated group as a control in a diff-in-diff, then even if you have parallel trends, your diff-in-diff will be equal to the minus the change in the from the preperiod to the postperiod for our control group. It’s similar to the bias under a no anticipation (NA) violation in that it is biasing “downward,” but now the problems reverse. If the treatment effects for the control group are constant, then there is no bias, whereas with an NA violation, a constant treatment effect caused the diff-in-diff equation to equal zero. But if the treatment effect in the control group is changing between the pre and postperiods, then the diff-in-diff equation will shave off some of the an amount equal to that change.

9.4 Diff-in-Diff and the Minimum Wage

To motivate this section on the minimum wage, I’ve tweaked our Mississippi River metaphor to emphasize, once again, the Princeton side, as the material in this section turned out to be fairly influential in two ways: to the empirical literature on minimum wages, and secondly, to what appears to have been a turning point in the widespread adoption and acceptance of diff-in-diff as a design-based approach to causal inference.

TODO Two rivers into causal inference

David Card and Alan Krueger were colleagues at Princeton, both deeply embedded in the Princeton Industrial Relations Section, both esteemed and transformative labor economists.6 In late 1991, they learned that New Jersey was raising its minimum wage later in 1992, but that Pennsylvania, a neighboring state, would not be. They began planning the study immediately and decided they would focus on fast food restaurant workers in both states using a survey methodology.7 To create their survey, using the phone book, they collected the names of all the fast food restaurants (not including McDonald’s) in the two states and hired someone to call each of them to collect data on workers’ labor market outcomes like wages and employment. They needed the survey done twice, too—once before and once after the implementation of the law change. So the first round of calls were made in February 1992, which will be the preperiod, and then again in November 1992, which will be the postperiod.

Figure 9.3: Distribution of wages for NJ and PA in November 1992 from (Card and Krueger 1994).

For studies like these, it is recommended that if at all possible, before you look at the effect of some policy on your main outcomes of interest, you first see if the policy had any first stage effect on take-up. This is sometimes called by the labor economics community the policy’s “bite.” In this minimum wage context, it simply means you need to show that when the minimum wage increased, low-wage workers saw their wages increase. If they didn’t, then it’s hard to imagine why any other effect would’ve happened. Card and Krueger showed that, and the figure above is one of the pictures they produced to do so:

The figure above shows the distribution of wages in November 1992 after the minimum wage hike.8 As can be seen, the minimum wage hike was binding, evidenced by the mass of wages at the minimum wage in New Jersey. The interpretation here is that the minimum wage increase had “bite” on New Jersey workers because its intended goal—to raise low wage workers’ wages—happened.

A long-standing source of contention in the field of economics regards the empirical estimates of the minimum wage on employment, not on wages. Theoretically, neoclassical models of competitive labor markets predict declines in employment from increases in the minimum wage, particularly over longer time horizons where the growth rate in employment could be affected (Meer and West 2016). But as we saw in the early chapter on Giffen behavior by Jensen and Miller (2008), the firm’s response to higher wages is just as complex, albeit for different reasons, as households’ responses to higher prices. It’s possible, even within that broader neoclassical tradition, for minimum wages to increase employment, depending on the amount of competition in the area (Robinson 1933).

But recall, also, the ethos of Princeton’s Industrial Relations Section—it did seem like the Section’s hyper focus on realistic empirical estimates over theoretical claims had shaped the entire worldview of the economists there such that the attachment to theoretical claims like that would always require rigorous falsification using high-quality data and valid research designs. So, Card and Krueger collected their own data both before and after for the same fast food restaurants using an original telephone survey, managed attrition between the sampling with a high degree of success, and reported their results in a table like Table 9.12, below.

Table 9.12: Simple DD Using Sample Averages on Full-Time Employment
Dependent variable PA NJ NJ - PA
FTE before 23.3 20.44 -2.89
(1.35) (0.51) (1.44)
FTE after 21.147 21.03 -0.14
(0.94) (0.52) (1.07)
Change in mean FTE -2.16 0.59 2.76
(1.25) (0.54) (1.36)

Standard errors in parentheses.

If employment responded to a minimum wage increase as it would in a highly competitive labor market, we would expect employment to decrease due to higher wage costs. However, in their sample, Card and Krueger (1994) find the opposite—a positive effect of +2.76 in mean full-time-equivalent employment. Importantly, this result isn’t just noise: dividing the coefficient by the standard error gives a t-statistic of about 2, indicating the result is statistically significant at conventional levels, which allows us to reject the null hypothesis of no effect.9

The response to this paper’s finding of a positive employment effect was mixed. Some labor economists weren’t surprised, given ongoing debates about empirical methods in the field. Others found the results implausible because they contradicted standard competitive labor market theory. Regardless of one’s view of the findings, the paper was undeniably influential both for minimum wage research and the broader adoption of diff-in-diff in economics (Currie, Kleven, and Zwiers 2020).

Figure 9.4: DD regression diagram.

9.6 The Importance of Falsifications in Diff-in-Diff

Scientific theories achieve credibility not just by explaining observed phenomena, but by excluding alternative explanations. That means that having alternative explanations that provide precise, falsifiable predictions is an important part of any causal study. This principle was central to philosopher Karl Popper’s view of science. Popper (1959) argued that a theory is scientific only if it can be tested and potentially proven wrong. In Popper’s words: “In so far as a scientific statement speaks about reality, it must be falsifiable; and in so far as it is not falsifiable, it does not speak about reality.” Thus, a good theory doesn’t simply account for what we already know; it makes bold predictions that put your original results at risk for being rejected.

Applying this concept to causal inference, including diff-in-diff, means conducting tests to rule out alternative explanations. In diff-in-diff, falsifications are tests of alternative explanations for your results. What is needed then is two theories—the one being that your treatment caused the results of your diff-in-diff, and the other being that something else did. A falsification in this case would require that the alternative theory make other testable predictions not relevant to the treatment you’re studying, and then testing for those predictions. If, when you test the alternative theory’s predictions on something else, and you fail to reject the null hypothesis, the feasibility of the alternative explanation for your main results is weakened.11

There are roughly two kinds of falsification designs that people employ, which I call “same outcome, alternative groups” and “same group, alternative outcomes.” Let’s review them both now.

Same Outcome, Alternative Groups

The idea behind “Same Outcome, Alternative Groups” falsification tests is to imagine an alternative explanation for your diff-in-diff results that would also impact a group irrelevant to your treatment and then applying your diff-in-diff model to that context. Consider this as an example: You are studying the effect of the minimum wage on employment and therefore focus your attention on low-wage workers whose wages are increasing as a result of the minimum wage increase. You find in this study that increases in the minimum wage reduce employment.

But perhaps there is another explanation and that is that there were state-wide declines in employment due to the area’s faltering economies. As a result of these broader, more systemic problems, many industries and many workers, not just the lowest-wage ones, saw declines in employment. Fortunately, this is testable with a “same outcome, alternative group” falsification. You would simply reestimate your diff-in-diff model using the higher-wage workers’ employment as your outcome of interest. If you find using this falsification declines in employment by a group of people who aren’t paid the minimum wage, it calls into question that the original results could be due to the minimum wage.

Rejecting the null hypothesis on the alternative group does not mean that you can be certain your original findings were spurious because, again, it’s entirely possible that they still were true. The point is that the falsification is a secondary source of evidence that the parallel trends assumption was plausible in the first case, but just like pretrends can be suggestive evidence that parallel trends holds, falsifications like the “same outcome, alternative groups” are also suggestive evidence. In the next section, I will discuss a fabulous example of the “same outcome, alternative group” falsification by Miller, Johnson, and Wherry (2021) involving a placebo group unaffected by Medicaid expansion. But I’ll wait to show you that.

Same Group, Alternative Outcomes

An alternative approach to falsification testing within the diff-in-diff framework involves examining the same group but looking at alternative outcomes that should not be affected by the treatment. If the treatment truly impacts a given outcome, it shouldn’t be influencing unrelated outcomes for the same group. Researchers have applied this reasoning in clever ways to challenge established models and bring more scrutiny to popular empirical approaches.

Imagine a study examining the impact of state increases in cigarette taxes on smoking-related hospitalizations. The primary hypothesis is that the taxes reduce hospitalizations for smoking-related illnesses, such as heart attacks, due to reductions in smoking. But again, maybe this is a spurious finding driven by a different policy that happened at the same time in the same place. Perhaps at the same time that states increased their cigarette taxes, states expanded public insurance that increased preventative healthcare. This expansion in healthcare is thought to be the primary cause ofdeclines in heart attacks, not the cigarette taxes, but if that is true, then the expansion in healthcare should affect other conditions unrelated to smoking, such as appendicitis or fractures. These conditions should not be influenced by cigarette taxes, as they are unrelated to smoking behavior or secondhand smoke exposure, but they should be influenced by more generous public insurance. If the analysis finds no effect of cigarette taxes on hospitalizations for these unrelated conditions, it strengthens the argument that the observed reduction in smoking-related hospitalizations was caused by cigarette taxes and not due to broader confounding trends affecting allhospitalizations.

By contrast, if the analysis reveals a similar reduction in hospitalizations for non-smoking-related conditions, this could indicate that your main results are spurious, driven by broader factors, such as changes in hospital reporting practices or concurrent health policies, rather than the taxes themselves. For this to be done well, though, one must have a credible and compelling alternative theory that makes additional predictions not made by the program you are studying.

I’ll conclude with two real-world examples that I find delightful examples of “same group, alternative outcomes” falsification. One striking example comes from the literature on rational addiction. For those unfamiliar with this literature, Becker and Murphy (1988) proposed that addiction could be rational, and that theoretical framework has, in turn, been a cornerstone theoretical framework for studying public policies aimed at regulating addictive goods. Researchers would typically estimate empirical models derived from the Becker and Murphy (1988) theoretical framework to commodities and activities like alcohol, tobacco, and gambling, often finding evidence that these behaviors align with the rational addiction framework.

In a clever falsification exercise, Auld and Grootendorst (2004) applied the rational addiction model to unrelated commodities—such as milk, eggs, and oranges—that could not plausibly be considered addictive. Surprisingly, the model suggested that milk was one of the most addictive goods studied. This finding cast doubt not so much on the theoretical underpinnings of rational addiction, but on the empirical techniques applied to aggregate data commonly used to support it. If the model is misclassifying milk as addictive, when milk may not be addictive,12 it raises serious questions about the reliability of the same models applied to alcohol, tobacco, etc. using similar data.

Another fascinating example of the “same group, alternative outcomes” falsification was two studies on peer effects and obesity. Estimating peer effects is notoriously difficult, as highlighted by Manski (1993), because peers are chosen, not randomized. Despite these known problems of endogeneity, influential studies found that behaviors like obesity, smoking, and even happiness seemed “contagious” within social networks (Christakis and Fowler 2007). However, Cohen-Cole and Fletcher (2008) used the common empirical peer effect model on a dataset with social network information to examine outcomes that couldn’t possibly be shared with or contagious between peers such as acne, height, and headaches. Yet, even these traits appeared to exhibit “contagion” effects in observational data.

Unfortunately for some of us, we have not been able to become taller by having tall friends. A more likely explanation for why tall people have tall friends is that they all play together on the basketball team.

Again, it does not mean that the original results must be wrong simply because the same empirical model finds evidence for peer effects on alternative outcomes that cannot be influenced by peers, like headaches and height, but it is suggestive that something is wrong with that empirical model if it is.

These two examples show how examining alternative outcomes within the same group can provide a robust falsification test, uncovering potential flaws in empirical designs without challenging the theoretical framework itself.

Inference

Many studies employing diff-in-diff strategies use data spanning multiple years—not just one pre- and one post-treatment period as in Card and Krueger (1994). The variables of interest in these setups often vary only at the group level, such as by state, and outcome variables are typically serially correlated. For instance, in Card and Krueger (1994), employment in each state is likely correlated within the state and also serially correlated over time. Bertrand, Duflo, and Mullainathan (2004) show that conventional standard errors can severely understate the variability of the estimators, leading to biased downward standard errors that are “too small” and consequently over-reject the null hypothesis. To address this, Bertrand, Duflo, and Mullainathan (2004) propose the following solutions: block bootstrapping, aggregation into two periods, and clustering standard errors at the group level. Below is a detailed explanation of each method.

  1. Block bootstrapping: If the block is a state, then block bootstrapping involves resampling states with replacement. Implementing this requires programming to loop through samples and store estimates, with mechanics similar to randomization inference. Readers interested in programming block bootstraps can adapt their approach from these general principles.

  2. Aggregation: This approach avoids the time-series dimension by simplifying to one pre- and one postperiod. For settings with just one untreated group and no differential timing, this is straightforward: calculate averages for the pre- and postperiods and conduct difference-in-differences on these aggregates. When facing differential timing, residualization steps are added: first, regress the outcome on panel unit and time fixed effects, keeping only residuals for treated groups. Then divide the residuals into pre- and postperiods and regress on the post-treatment indicator. This approach does not recover the original point estimate, so it is typically used only when simpler methods are impractical.

  3. Clustering: Clustering the standard errors by group (often the level of treatment) is the most common and straightforward method for correcting standard errors, as it accounts for arbitrary serial correlation in errors within groups over time. For example, in state-level panels, clustering at the state level is standard. This method is widely available in statistical software, making it easy to apply without additional programming.

Inference in panel settings remains complex, especially with few clusters. When cluster numbers are small, clustering alone may lead to inflated false-positive rates. In extreme cases, such as with only one treatment unit, clustering may fail entirely: even techniques like the wild bootstrap can have high false positive rates (Cameron, Gelbach, and Miller 2008; MacKinnon and Webb 2017). For these cases, randomization inference—as shown by Buchmueller, DiNardo, and Valletta (2011)—might be an alternative approach worth pursuing.

9.7 Diff-in-Diff in the Courtroom

There is a difference between your main results from diff-in-diff and evidence, but what is the difference? What do I mean by evidence and what do I mean by results? And what constitutes evidence in a diff-in-diff? Consider this courtroom analogy.

The prosecutor and defense attorney are attempting to persuade a judge and jury of their point of view using evidence, logic, and precedent. This is a zero–sum game because if the prosecution wins, the defense loses and vice versa. But what defines the prosecutor’s success is shifting the jury’s collective beliefs that the defendant is guilty “beyond a reasonable doubt.” And to do that there are three parts to the prosecution’s case, each of which represents something distinct and unique in your own diff-in-diff study.

  1. Assertion of the defendant’s guilt. Note that the assertion of guilt is not evidence of guilt, but rather simply a claim that the defendant is guilty of some offense.

  2. Presentation of evidence. Evidence might include eye witnesses who can put the defendant at the crime scene, the broken glass showing signs of forced entry, and fingerprints found on the smoking gun.

  3. Evidence of a credible motive. Motive would be the jilted lover, or an insurance policy for which the defendant was the sole beneficiary.

Each of these things on the itemized list we know by heart because we’ve seen countless movies and television shows set in the courtroom, or we know people who are lawyers. We know that the claim of guilt and the evidence for guilt are completely different. And we know that a lawyer’s case (point 1) gets stronger the more relevant evidence that is produced (point 2) and the more credible a theory of why the defendant would do it (point 3) is presented.

You are the prosecution in a case entitled “The Causal Effect of \(D\) on \(Y\) Using Diff-in-Diff.” Your diff-in-diff estimates are point 1, the claim of guilt. But these are not your evidence for the claim of guilt and to focus on point 1 to the exclusion of point 2 is a case so weak that readers and the public are justified to reject it. Your diff-in-diff study needs conclusive evidence, not merely a table of regression coefficients with asterisks by the numbers.

In this section, I will extend this courtroom metaphor to suggest the five elements of a strong study using diff-in-diff. The list that I will be discussing is:

  1. Tables and figures showing the policy had bite.

  2. Graphical event studies.

  3. Tables and figures showing reasonable falsifications.

  4. Tables and figures showing Main Results.

  5. Tables, figures, and narrative suggesting plausible mechanisms that explain your results.

To illustrate each of these, I would like to do so in the context of an excellent study involving the expansion of public insurance in the United States in the mid-2010s under then president Barack Obama and his vice president, Joseph Biden.

Medicaid and Mortality

A provocative study by Miller, Johnson, and Wherry (2021) examined the expansion of Medicaid under the Affordable Care Act (ACA). They were primarily interested in the effect that this expansion had on population mortality. Earlier work had cast doubt on Medicaid’s effect on mortality (Finkelstein et al. 2012; Baicker et al. 2013), so revisiting the question with a larger sample size had value.

Like Snow before them, the authors link datasets on deaths with a large-scale federal survey data, thus showing that shoe-leather often goes hand in hand with good design. They use these data to evaluate the causal impact of Medicaid enrollment on mortality using a diff-in-diff design. Their focus is on the near-elderly adults (i.e., adults below but near the age 65) in states with and without the Affordable Care Act Medicaid expansions, and they find a 0.13 percentage point decline in annual mortality, which is a 9.3% reduction over the sample mean, as a result of the ACA expansion. This effect is a result of a reduction in disease-related deaths and gets larger over time. Medicaid, in their estimation, caused a non-trivial number of lives to be saved.

As with many contemporary diff-in-diff designs, Miller, Johnson, and Wherry (2021) evaluate the plausibility of parallel trends with event studies in which they plotted regression coefficients with 95% confidence intervals on their treatment leads and lags. Including leads and lags into the diff-in-diff model allowed the reader to check both the degree to which the post-treatment treatment effects were dynamic, and whether the two groups were comparable on outcome dynamics prior to Medicaid expansion. Models like this one usually follow a form like: \[ Y_{its} = \gamma_s + \lambda_t + \sum_{\tau=-q}^{-1}\gamma_{\tau}D_{s\tau} + \sum_{\tau=0}^m\delta_{\tau}D_{s\tau}+x_{ist}+ \varepsilon_{ist} \tag{9.29}\] Treatment occurs in year 0. You include \(q\) leads or anticipatory effects and \(m\) lags or post-treatment effects.

Miller, Johnson, and Wherry (2021) produce numerous event studies, which when taken together, tell the main parts of the story of their paper. I will focus on five of them. The event study plots are, to me, quite powerful. Let’s look at the first three related to the “bite” of the policy expansion.

State expansion of Medicaid under the Affordable Care Act increased Medicaid eligibility (the figure below), which is not altogether surprising. But it also caused an increase in Medicaid enrollment (the figure below), as well as a reduction in the percent of the population uninsured (the figure below). All three of these are simply showing that the ACA Medicaid expansion had “bite”—people enrolled and became insured who otherwise would not have been insured. The last outcome—uninsured—is helpful because it suggests that the second outcome—enrollment—was not merely people switching from private insurance to public insurance. At least some of it was coming from people switching from uninsured to insured, thus giving them better access to free healthcare, presumably for the first time.

Figure 9.9: Estimated effect of Medicaid expansion on Mediciad eligibility (Miller, Johnson, and Wherry 2021).

The authors also present a “same outcome, alternative group” falsification. Some context about the institutions of public insurance may be useful for readers outside of the United States. The United States has two universal healthcare programs—Medicaid aimed at poor people and Medicare aimed at elderly people aged 65 and older. The expansion of Medicaid would hypothetically only matter for the poor; it would not matter for the elderly population. And since the authors are focused on the “near elderly” population, any alternative explanation for their results that is not Medicaid, but rather common to older people in those expanding states, should probably show up for the elderly population not enrolled in Medicaid.

Figure 9.10: Estimated effect of Medicaid expansion on Mediciad coverage (Miller, Johnson, and Wherry 2021).
Figure 9.11: Estimated effect of Medicaid expansion on uninsured status (Miller, Johnson, and Wherry 2021).

The figure below presents graphical evidence for the placebo elderly population. And using event studies, they show both that the elderly do not enroll in Medicaid when it expands, and they show no change in mortality either. Whatever is going on in the expanding states relative to the non-expanding states, it does not appear to be affecting the mortality of elderly people.

Figure 9.12: Falsification estimates effect of Medicaid expansion on 65 and older Medicaid coverage and mortality (Miller, Johnson, and Wherry 2021).
Figure 9.13: Falsification estimates effect of Medicaid expansion on 65 and older Medicaid coverage and mortality (Miller, Johnson, and Wherry 2021).
Figure 9.14: (Miller, Johnson, and Wherry 2021) estimates of Medicaid expansion’s effects on annual mortality using leads and lags in an event study model.

Next, the authors looked at the mortality rates of the near elderly population, which are their main results. They present event studies for this outcome, just like they had with the others, so that we can investigate the pretrends. The figure above shows some evidence that the expansion of Medicaid may have caused near elderly mortality rates to decline, though the standard errors are large enough that while each one is different from zero, they are not different from another (including the ones pretreatment), and a linear trend cannot be rejected (Rambachan and Roth 2023). But when the figure above is taken alongside the falsifications and the other evidence presented, it is compelling to me and worth considering as plausible evidence that Medicaid’s expansion saved lives.

Recall, though, the last element in our courtroom drama: motive. Why did the defendant kill the victim with the candlestick? Was it revenge? Was it opportunistic murder? Without a plausible motive, the jury and the judge can find it hard to believe beyond a reasonable doubt why someone would commit such a heinous crime. And that same doubt creeps into a diff-in-diff study oftentimes without some evidence for the mechanism driving the main results. The authors show evidence suggesting that the reason the mortality coefficients were so negative was because individuals who got on Medicaid because of the expansion received care that allowed them to get treatment for life-threatening illnesses—a finding that itself has policy implications.

I consider Miller, Johnson, and Wherry (2021) to be in the same category as Card and Krueger (1994)—an important policy study using difference-in-differences that also serves as an excellent research exemplar.

Of course, Miller, Johnson, and Wherry (2021) might still be wrong. We can never know the counterfactual mortality had expansion states not expanded Medicaid. This uncertainty isn’t unique to their study—it’s the fundamental challenge that motivates causal inference. Any difference-in-differences study could be wrong since parallel trends cannot be directly tested without observing the missing counterfactual \(E[Y^0|D-1,Post]\).

What I find compelling about Miller, Johnson, and Wherry (2021) is their layered evidence and careful presentation. They address multiple alternative explanations, leaving skeptical readers with few remaining objections. Like Snow’s cholera study, they combine visualization, narrative, and falsification tests alongside their main results. Their thoughtful communication helps readers understand the analysis without sacrificing rigor—echoing Ashenfelter’s goal when coining “difference-in-differences."

I encourage reading this paper not just to learn about Medicaid and mortality, but to study how the authors construct their argument, what evidence they seek, and how they present it through clear tables and compelling visualizations.

9.9 Concluding Remarks

We covered a lot of ground in this chapter. We learned about the history of diff-in-diff, for instance, and its data requirements. We discussed the parameter that diff-in-diff identifies under three assumptions—no anticipation, SUTVA, and parallel trends—and we discussed the diff-in-diff equation of “four averages and three subtractions.” We also discussed the regression specification that one can use in place of manually calculating the diff-in-diff equation, the role of the event study, bite, and falsifications in diff-in-diff as well as discussing in detail the types of behavior that threaten parallel trends and the types that don’t. And occasionally I offered up my own opinions, which is always dangerous, but I did it anyway because Danger is my middle name.

But we aren’t done with diff-in-diff yet because we now want to know what our options are if parallel trends is not plausible. And what do we do when there is more than just one treatment group, which was the only situation we covered in this chapter. In the next chapter we will build on what we did here in this chapter, and hopefully between the two of them, you’ll have whatever you need when it’s time to stay and when to move on from diff-in-diff altogether.


  1. The situation in the First Clinic was so dire, and its reputation so well-known, that mothers in labor would go to great lengths to avoid it—some even preferred to give birth in the streets and arrive at the hospital only afterward, umbilical cord in hand, rather than risk being admitted to a clinic seen as a place of death.↩︎

  2. These data are publicly available from Wikipedia at https://en.wikipedia.org/wiki/Ignaz_Semmelweis.↩︎

  3. A fable might help bring these and other stories in this chapter to life. A man goes to his doctor and states that he has died. The doctor, puzzled, tells the man that he cannot be dead because he is speaking to him right now, but the patient says, “I understand that, but I am dead, which must mean dead people can speak.” The doctor thinks about it and then asks the man, “Can dead people bleed?” The patient pauses, thinks about it, and says “No, dead people cannot bleed.” The doctor then pulls out a needle and pricks the man’s finger causing a droplet of blood to form. The man stares for a long time at the blood on his hand and whispers to himself, “What do you know? I guess dead men can bleed.” Sometimes we hold beliefs that make it impossible to believe the obvious, even when the obvious is right in front of our noses.↩︎

  4. I once made the rookie mistake of trying to explain triple differences to a journalist over the phone. That was probably a half hour wasted for both of us.↩︎

  5. It probably technically violates the SUTVA requirement that the treatment be the same for all units, too, as now the treatment is “announced but not enforced” and then changes to “announced and enforced,” but I don’t want to be too much of a downer, so I’ll just say it’s up to you.↩︎

  6. Alan Krueger passed away in 2018.↩︎

  7. Economists are not, as a group, known for conducting their own surveys, but interestingly in 1994, Krueger actually published two articles in the American Economic Review using his own original surveys—one being the study on minimum wages with Card and Krueger (1994) that we will discuss, and another being a paper estimating the returns to schooling using a sample of twins he coauthored with Orley (O. C. Ashenfelter and Krueger 1994).↩︎

  8. That’s a cool picture. It’s pretty striking when there’s no white space above the black histogram for NJ at $5.05. You should think carefully about the use of white space in your graphs and what it conveys and what you are trying to convey. Put yourself, always, in the reader’s shoes when making these figures. What do they probably see? Oftentimes you and I see something else entirely because we’ve been working on these projects for what feels like forever, but this is the first and maybe last time they’ll ever see this picture, so it needs to say exactly what you are trying to say.↩︎

  9. A t-statistic above 1.96 implies statistical significance at the 5% level in a two-tailed test.↩︎

  10. I first saw this graphical presentation by the economist Fabian Waldinger, but where I’ve butchered the example, blame me, and where it is helpful, thank him.↩︎

  11. In my opinion, falsifications are more important than event studies in drawing conclusions about the plausibility of parallel trends, but I suspect I am in the minority.↩︎

  12. For whatever it’s worth, I feel like I am addicted to milk. The stomach, much like the heart, loves what it loves.↩︎

  13. This simulation is available online at the free book website but is not shown here to save space.↩︎

  14. Workers differ in characteristics like age and race, but for the purposes of this simulation, these covariates will not influence selection into treatment.↩︎