Finance and Economics Discussion Series: Accessible versions of figures for 2026-045

An Evaluation of Difference-in-Differences Methods Using Placebo Event Studies

Accessible version of figures


Figure 1: Synthetic control estimates without normalization (left) vs. with (right)
Note: The outcome variable is the seasonally adjusted LAUS employment growth rate.
Source: Bureau of Labor Statistics; Census Bureau.

This figure contains two panels, each showing the distribution of synthetic control placebo event study estimates at different horizons. Both charts display time to treatment in months on the x-axis, ranging from -12 to 12, with 0 marking the treatment time (indicated by a vertical dashed line). The y-axis shows estimates and ranges from -4 to 4, with a horizontal solid line at 0. Each chart displays five lines representing different percentiles: 10th percentile (black), 25th percentile (red), 50th percentile (blue), 75th percentile (orange), and 90th percentile (light blue). The left panel shows the distribution for unnormalized estimates, while the right panel shows the distribution of normalized estimates. Both sets of estimates are more tightly concentrated around zero in the pre-treatment period, before widening post-treatment. However, the un-normalized estimates show more spread both before and after treatment.

Return to text.


Figure 2: Mean squared error of estimates by method and seasonal adjustment
Note: The outcome variable is the LAUS employment growth rate, and the predictor for matching estimators is foreign-born share. We exclude generalized synthetic control because of its sensitivity to outliers. From left to right, the estimators shown are coarsened exact matching (CEM), Callaway and Sant’Anna (CSA), D’Chaisemartin and D’Haultfœuille (DCDH), elastic net synthetic control (ENSC), imputation (Imp.), local projections (LP), matrix completion (MC), propensity score matching (PSM), synthetic control (SC), synthetic difference-in-differences (SDID), two-way fixed effects (TWFE), and Wooldridge difference-in-differences (WDID). The salmon bar denotes non-seasonally adjusted (NSA) data, whereas the teal bar denotes seasonal adjusted (SA) data.
Source: Bureau of Labor Statistics.

This is a grouped bar chart comparing mean squared error (MSE) of placebo estimates across estimators, showing separately the MSE for seasonally adjusted versus not seasonally adjusted data. The y-axis measures MSE, ranging from 0.0 to 3.5. The x-axis shows twelve estimators: coarsened exact matching (CEM), Callaway and Sant'Anna (CSA), D'Chaisemartin and D'Haultfoeuille (DCDH), elastic net synthetic control (ENSC), imputation (Imp.), local projections (LP), matrix completion (MC), propensity score matching (PSM), synthetic control (SC), synthetic difference-in-differences (SDID), two-way fixed effects (TWFE), and Wooldridge difference-in-differences (WDID). Each estimator displays two bars: a red-colored bar representing non-seasonally adjusted (NSA) data and a green bar representing seasonally adjusted (SA) data. The seasonally adjusted specification consistently shows lower MSE values than the not seasonally adjusted specification across all estimators. Most NSA bars cluster between approximately 2.1 and 2.4, with ENSC showing a notably higher NSA value of approximately 3.2, the highest in the entire chart. The SA bars generally range from approximately 1.55 to 2.0, with ENSC and MC showing the highest SA values at about 2.7 and 2.0, respectively, while most other estimators display SA values between 1.55 and 1.7.

Return to text.


Figure 3: Mean squared error of estimates by method and covariate specification
Note: The outcome variable is the seasonally adjusted LAUS employment growth rate. From left to right, the estimators shown are coarsened exact matching (CEM), Callaway and Sant’Anna (CSA), local projection (LP), propensity score matching (PSM), synthetic control (SC), and Wooldridge difference-in-differences (WDID). From left to right in each set of bars, each colored bars represents a unique covariate specification:
None (except matching estimators)
One covariate: average age; average income; Black population share; college-educated share; female employment rate; foreign-born share; manufacturing share of employment; poverty rate
Two covariates: average income, average age; college-educated share, female employment rate; foreign-born share, Black population share; poverty rate, female employment rate
Four covariates: average income, college-educated share, Black population share, manufacturing share of employment
Source: Bureau of Labor Statistics; Census Bureau.

This is a grouped bar chart displaying the mean squared error (MSE) for six different estimation methods across multiple covariate specifications. The y-axis measures MSE, ranging from 0.0 to 2.5. The x-axis shows six estimators: coarsened exact matching (CEM), Callaway and Sant'Anna (CSA), local projection (LP), propensity score matching (PSM), synthetic control (SC), and Wooldridge difference-in-differences (WDID). Each estimator has about a dozen bars representing MSE for different covariate specifications, displayed in rainbow colors ranging from orange to magenta. CEM shows the most variation in MSE values, ranging from approximately 1.3 to 2.2, with the magenta bar reaching the highest MSE value in the entire chart at about 2.2. The other estimators display more consistent performance, with MSE values clustering between approximately 0.7 and 1.5.

Return to text.


Figure 4: Normalized MSE of estimates by state (percentile)
Note: We exclude generalized synthetic control and elastic net synthetic control because of their sensitivity to outliers. We obtain normalized MSE by calculating MSE by state within each specification (outcome \(\times\) estimator \(\times\) covariates ), taking the percentage relative to the difference between the maximum and minimum values, and then averaging across specifications by state.
Source: Bureau of Labor Statistics; Census Bureau; Federal Housing Finance Agency.

This is a choropleth map of the United States displaying the normalized mean squared error (MSE) of estimates by state as percentiles. The map includes all 50 states plus the District of Columbia. The states are colored according to the normalized MSE of estimates for that state, ranging from 0 to 100, with purple representing the lowest percentiles (around 0) and yellow representing the highest percentiles (around 100), with intermediate values shown in shades of pink, rose, and orange. States exhibiting low normalized MSE include Pennsylvania, Ohio, Illinois, Texas, and North Carolina; states exhibiting high normalized MSE include Nevada, Hawaii, Alaska, DC, and Arizona.

Return to text.


Figure 5: Average control weight by state (percentile)
Note: Weights are pooled across events, covariates, and outcomes.
Source: Bureau of Labor Statistics; Census Bureau; Federal Housing Finance Agency.

This figure contains four choropleth maps of the United States, arranged in a grid, each displaying average control weights by state expressed as percentiles for different estimation methods. Each map includes all 50 states plus the District of Columbia. All four maps use the same color gradient for weight ranging from 0 to 100, with white representing the lowest percentiles (around 0) and dark blue representing the highest percentiles (around 100). The top left map shows the weights for propensity score matching, which heavily weights Iowa, Nebraska, Texas, Vermont, and Connecticut. The top right map shows the weights for synthetic control, which heavily weights California, Michigan, Georgia, and New Hampshire. The bottom left map shows the weights for elastic net synthetic control, which heavily weights New Mexico, Nevada, Alaska, North Dakota, and Maine. The bottom right map shows the weights for synthetic difference-in-differences, which heavily weights Utah, Nevada, Michigan, Arizona, and Connecticut. The four panels demonstrate substantial variation in which states receive the highest control weights depending on the estimation method employed. No single geographic pattern dominates across all four methods, highlighting how different estimators weight control units differently in the construction of counterfactuals.

Return to text.