How Effective Is PSPS
at Preventing Wildfires?

Public-data estimates and the case for utility-held records — walking from naive comparisons to quasi-experiments, and why none of it is causal yet

The Benefits of PSPS Are Simulated, Not Measured

PG&E's headline figure for 2024 — "up to 76,314 acres avoided," 10,841 buildings, 10,868 residents — comes from 24-hour fire simulations at damage locations. The counterfactual is generated inside a model.

76,314
acres "avoided" in 2024 — simulated
100%
effectiveness assumed by the best published EPSS study
0
published checks against observed ignition outcomes

The Trap: Shut-Off Days Show 15× More Ignitions

Across ~12.2 million PG&E circuit-days, shut-off days ignite at 15.1× the rate of ordinary days (SCE: 24.5×). Read causally, shutoffs cause fires. They don't — utilities cut power exactly where and when risk peaks.

DAG showing confounding: dangerous conditions U drive both the shutoff D and ignitions Y

Step 1: Compare Only Where Shutoffs Were on the Table

Restrict to high fire-threat circuits on days a PSPS event was underway somewhere. Three-quarters of PG&E's raw gap evaporates (15.1× → 3.7×); half of SCE's (24.5× → 12.1×).

Bar chart: raw rate ratio collapses when restricted to comparable high-risk circuit-days at both PG&E and SCE

Steps 2–3: Same Circuit, Same Storm — the Gap Survives

Compare each circuit only to itself (fixed effects): 2.8× [1.8, 4.3]. Then match each shut-off day to look-alike days in the same storm on wind, humidity, fuel, terrain: the gap persists at +0.0036 ignitions per circuit-day.

DesignRate ratioRisk difference
Raw difference in means15.1×+0.0048
High-risk circuit-days only3.7×+0.0040
Circuit fixed effects (within-circuit)2.8×+0.0034
Matched within storms (ATT)4.4× [1.4, 24.1]+0.0036

Step 4: Flexible ML Lands in the Same Band

A doubly-robust AIPW estimator — lasso propensity and outcome models, cross-fit by circuit — removes as much observed-confounding bias as the measured covariates allow. It lands where simple matching did.

AIPW estimate falls inside the same +0.003 to +0.004 risk-difference band as earlier designs

The Bounds Convict the Benchmark

Instead of removing the bias, bound it. Assumption-free Manski bounds run [-0.994, +0.006]. Add one mild assumption — a de-energized line cannot start a fire — and the harmful edge pins at zero.

[-0.994, +0.006]
assumption-free Manski band
≤ 0
effect range under monotone treatment response
+0.004
benchmark estimate — outside the MTR band

Falsification: A Clean Placebo, With a Blind Spot

Re-run the matched design on fires PSPS physically cannot prevent — lightning, arson, vehicle fires. The placebo lands near 1× while the real utility outcome over the same years (2014–2020) sits near 2.8×.

Placebo comparison: non-utility ignitions show no gap while utility ignitions show ~2.8x over the same period

Two Quasi-Experiments, Two Dead Ends

A switching difference-in-differences (PanelMatch, 1,936 matched sets) returns intervals that all span zero — and pre/post lines that cross at the shutoff, the fingerprint of risk-based timing. The rule-change IV fails outright.

Event study: post-shutoff effects positive, pre-shutoff placebos negative, all intervals spanning zero

Every Public-Data Approach Lands in the Same Place

Steps 0–4 on a common scale: a 15× gap that collapses, then plateaus in a narrow positive band no design can push to zero. Steps 5–9: bounded-but-wide, clean-but-blind, noisy, falsified.

Two-panel summary: rate ratio collapsing then plateauing; risk differences clustering in a +0.003 to +0.004 band

What Would Actually Answer the Question

Three designs deliver treatment variation that doesn't run through the operator's judgment — each gated on a specific piece of utility-held data.

DesignIdentifying variationNeeds from the utility
Border RDAdjacent locations, same weather — only one side loses powerCircuit & boundary geometry
Fuzzy RD at the FPI cutoffCircuits just over vs. just under the de-energization thresholdDecision-time FPI values & thresholds
Reform-induced DiD, done rightRule revisions whose timing was policy-driven, not risk-drivenRevision-rationale record

Key Takeaways

  1. PSPS benefits are simulated, not measured. The headline "acres avoided" figures are model counterfactuals; the best published study assumes 100% effectiveness. No one had checked against observed ignitions.
  2. The naive comparison is a trap. Shut-off days show 15× the ignition rate — but utilities de-energize exactly where risk peaks. The raw gap measures the risk, not the policy. Same pattern at SCE (25×).
  3. Adjustment hits a wall at +0.003 to +0.004. Restriction, fixed effects, matching, and doubly-robust ML all land in the same positive band — and under the mild assumption that de-energized lines can't spark, that number has to be bias, not effect.
  4. Utility data could settle it. Border RD, fuzzy RD at the FPI cutoff, and a defensible reform DiD are all feasible — but each needs records only utilities hold: scoped-but-not-shut-off controls, decision-time risk scores, boundary geometry, revision rationale.

Read the Full Analysis

igorgeyn.com/blog

Igor Geyn · June 2026

Igor Geyn · igorgeyn.com