Public-data estimates and the case for utility-held records — walking from naive comparisons to quasi-experiments, and why none of it is causal yet
Igor Geyn · June 2026
Full analysis: igorgeyn.com/blog
PG&E's headline figure for 2024 — "up to 76,314 acres avoided," 10,841 buildings, 10,868 residents — comes from 24-hour fire simulations at damage locations. The counterfactual is generated inside a model.
Across ~12.2 million PG&E circuit-days, shut-off days ignite at 15.1× the rate of ordinary days (SCE: 24.5×). Read causally, shutoffs cause fires. They don't — utilities cut power exactly where and when risk peaks.

Restrict to high fire-threat circuits on days a PSPS event was underway somewhere. Three-quarters of PG&E's raw gap evaporates (15.1× → 3.7×); half of SCE's (24.5× → 12.1×).

Compare each circuit only to itself (fixed effects): 2.8× [1.8, 4.3]. Then match each shut-off day to look-alike days in the same storm on wind, humidity, fuel, terrain: the gap persists at +0.0036 ignitions per circuit-day.
| Design | Rate ratio | Risk difference |
|---|---|---|
| Raw difference in means | 15.1× | +0.0048 |
| High-risk circuit-days only | 3.7× | +0.0040 |
| Circuit fixed effects (within-circuit) | 2.8× | +0.0034 |
| Matched within storms (ATT) | 4.4× [1.4, 24.1] | +0.0036 |
A doubly-robust AIPW estimator — lasso propensity and outcome models, cross-fit by circuit — removes as much observed-confounding bias as the measured covariates allow. It lands where simple matching did.

Instead of removing the bias, bound it. Assumption-free Manski bounds run [-0.994, +0.006]. Add one mild assumption — a de-energized line cannot start a fire — and the harmful edge pins at zero.
Re-run the matched design on fires PSPS physically cannot prevent — lightning, arson, vehicle fires. The placebo lands near 1× while the real utility outcome over the same years (2014–2020) sits near 2.8×.

A switching difference-in-differences (PanelMatch, 1,936 matched sets) returns intervals that all span zero — and pre/post lines that cross at the shutoff, the fingerprint of risk-based timing. The rule-change IV fails outright.

Steps 0–4 on a common scale: a 15× gap that collapses, then plateaus in a narrow positive band no design can push to zero. Steps 5–9: bounded-but-wide, clean-but-blind, noisy, falsified.

Three designs deliver treatment variation that doesn't run through the operator's judgment — each gated on a specific piece of utility-held data.
| Design | Identifying variation | Needs from the utility |
|---|---|---|
| Border RD | Adjacent locations, same weather — only one side loses power | Circuit & boundary geometry |
| Fuzzy RD at the FPI cutoff | Circuits just over vs. just under the de-energization threshold | Decision-time FPI values & thresholds |
| Reform-induced DiD, done right | Rule revisions whose timing was policy-driven, not risk-driven | Revision-rationale record |
Igor Geyn · June 2026