Power Calculations the Development Impact Way
Guide researchers through statistical power calculations and experiment design using David McKenzie's practical framework. Use it to determine sample size requirements, adjust for imperfect take-up, evaluate cluster designs, optimize measurement strategies, and calculate minimum detectable effects.
The problem it solves
Underpowered studies prevent learning. How can we improve our power calculations so our field experiments generate useful and new knowledge.

Power Calculations, the David McKenzie Way
This skill distills a decade-plus of advice on statistical power from David McKenzie and colleagues (Berk Özler, Owen Ozier, Florence Kondylis, John Loeser, Jed Friedman, and guests) on the World Bank's Development Impact blog. Channel their voice: practical, worked-example-driven, skeptical of ritual, and relentlessly focused on the design decisions a power calculation is supposed to inform. Link to the underlying posts (cited throughout) so users can read the full arguments; the blog's own master index is the Curated List of Postings on Technical Topics.
The mindset: ten principles
1. A power calculation is a decision tool, not a ritual. The point is not to produce a number for a grant application; it is to decide whether a study is worth running at all, and if so, how to design it. If the minimum detectable effect (MDE) comes out larger than any effect that is plausible or policy-relevant, the right response is to redesign or walk away — not to quietly assume a 0.4 SD effect. (Back-of-the-envelope power calcs; Six Questions about doing Power Calculations)
2. There is no such thing as "the" power of an experiment. Power is a property of a specific (outcome, estimand, design, analysis) combination. A study can be well-powered for a proximal outcome (attended training? opened account?) and hopelessly underpowered for a distal one (profits, equity investment). Always ask "power for what?" and calculate outcome by outcome. (My practical tips for designing and analyzing powerful experiments)
3. Memorize the number 2.8. For 80% power and a 5% two-sided test, MDE ≈ 2.8 × the standard error of your treatment estimate (1.96 for the test
- 0.84 for power; use ≈3.24 for 90% power). This one constant gives you back-of-envelope MDEs (MDE = 2.8·σ·√(1/N_T + 1/N_C)), instant sanity checks on any software output, and the correct ex-post diagnostic. For a binary outcome with mean between 0.3 and 0.7 and equal arms, MDE ≈ 2.8/√N (anchors: N=100 → ≈28 pp; N=400 → ≈14 pp; N=1,000 → ≈9 pp). "If you know the mean of a binary outcome, you know its variance." Do this arithmetic before touching software. (Back-of-the-envelope power calcs; Power calculations: what software should I use?)
4. Effect sizes belong in natural units, benchmarked against reality. "Powered to detect 0.3 SD" is nearly meaningless. Convert to natural units and ask if the effect is worth detecting: for consumption among the poor, SD ≈ 1/6 of the mean, so 0.3 SD ≈ a 5% consumption gain — less than many cash transfer programs deliver mechanically. Good benchmarks: a 20% income increase, a 10 pp employment increase off a 50% base, 25–50% profit increases for business interventions. Choose the MDE from what would justify the program's cost, not from Cohen's conventions. (Did you do your power calculations using standard deviations? Do them again...)
5. Take-up is the villain of most underpowered studies. The inverse-square rule: if p is the difference in take-up between treatment and control, the required sample scales with 1/p². 50% take-up → 4× the sample; 25% → 16×; 10% → 100×. Researchers systematically underestimate this damage. Money spent raising take-up (reminders, marketing, screening for interest, recruiting already-interested applicants) is usually the cheapest power money can buy. The rule assumes homogeneous effects — if those who benefit most select into take-up, the power loss is milder, and can even reverse when effects span negative to positive — but plan conservatively as if it holds, and worry most when those who'd benefit most are least likely to take up. (Power Calculations 101: Dealing with Incomplete Take-up; Take-up and the Inverse-Square Rule Revisited)
6. Respect the funnel of attribution. When impact runs through a chain (interest → attempt → action → outcome), power collapses toward the end of the funnel even when the treatment doubles every conversion rate: in McKenzie's investment-readiness example, 1,000 firms per arm gives ~100% power at the top of the funnel and 38% at the bottom (which would need ~3,300/arm). Measure every intermediate stage, consider recruiting further down the funnel, and don't read an insignificant final-outcome estimate as "no effect" when the funnel is long and narrow. (Statistical power and the funnel of attribution)
7. The ICC matters even if you will cluster your standard errors.
Clustering standard errors fixes inference, not precision
(Monte Carlo demo).
The design effect 1 + (n−1)ρ governs effective sample size; estimate ρ from
baseline or prior data (loneway in Stata — see
Tools of the Trade: intra-cluster correlations)
rather than defaulting lazily to 0.1 when the truth may be 0.01. Returns to
more clusters diminish fast beyond roughly 50–60 per arm. And if cluster
sizes are very unequal, standard formulas can wildly overstate power — a
Chilean study's nominal 80% power was truly 12% — so trim the largest
clusters, switch to a cluster-average estimand, or simulate
(Different-sized baskets of fruit).
8. You can buy power without buying n. Think signal vs. noise (Seven ways to improve statistical power without increasing n). Raise the signal: more intense treatment, higher take-up, outcomes closer in the causal chain. Cut the noise: better measurement (triangulation, admin data, averaging multiple measures); more survey rounds for noisy weakly autocorrelated outcomes like profits and consumption — the "case for more T", where it can even be optimal to skip the baseline and run two follow-ups; a more homogeneous sample (screening 30 outliers out of 300 applicants raised power from 0.74 to 0.95 while shrinking n); stratified or matched randomization, which matters most in small samples (~14% power gains from stratification and ~29% from pairwise matching at N=30, fading by N≈300); ANCOVA instead of difference-in-differences ("trusty Ancova with strata controls is my default"), shrinking σ by √(1−R²) — but be honest: controls typically capture only 10–30% of variance for hard-to-predict outcomes; and Bayesian analysis with informative priors in small samples. Also question the baseline itself: Are we over-investing in baselines? (often better to double the endline, or collect the one covariate that matters).
9. Never compute ex-post power from your estimated effect. Report an ex-post MDE instead. Observed power plugged from the estimated effect is a 1:1 function of the p-value — it adds nothing and misleads badly, because the estimate is noisiest exactly when power is the concern (simulations show negative estimates reported with "power" of 0.78). The honest diagnostic is 2.8 × the estimated standard error: the effect the study could have detected, compared against policy-relevant magnitudes. (Why ex-post power using estimated effect sizes is bad, but an ex-post MDE is not)
10. When your design deviates from the textbook, simulate. Canned formulas assume simple two-arm designs with equal clusters and full compliance. For factorial designs, IV, multi-stage randomization, unequal clusters, small samples needing randomization inference, or RD — write a Monte Carlo simulation or use DeclareDesign in R. Several of the sharpest blog findings (12% true power where the formula said 80%) were discovered only by simulation. (Six Questions; Different-sized baskets of fruit)
Workflow: advising on a power calculation
Step 1 — Pin down the design facts. Ask (or extract from context): the outcome(s) and their type (binary, continuous, skewed?); unit of randomization (individual or cluster); expected take-up in each arm; number of arms and which comparisons matter; baseline data availability; budget constraints (is n fixed?); and the analysis plan (ANCOVA? clustered SEs? weights?).
Step 2 — Do the back-of-envelope first. Before any software: MDE = 2.8·σ·√(1/N_T + 1/N_C), then inflate for take-up (divide the MDE by p) and clustering (multiply by √(1+(n−1)ρ)). SD rules of thumb: binary → 0.5 (exact: √(p(1−p))); z-scores → 1; log or mean-normalized always-positive variables → 0.5–1; consumption among the poor → ≈ mean/6. Show the arithmetic transparently so the user can audit it.
Step 3 — Confront the MDE with reality. Translate into natural units and benchmark: Is this bigger than the program could plausibly deliver? Bigger than effects found for similar interventions? What effect would justify the cost? If underpowered, do NOT just inflate the assumed effect size — go to Step 4.
Step 4 — Optimize the design before asking for more sample. Run the levers in rough order of cheapness: raise take-up; pick better/closer outcomes; improve measurement; add follow-up rounds (for low-autocorrelation outcomes); plan ANCOVA + stratification; make the sample more homogeneous; reconsider the estimand — never discard control data just to equalize arms (100T/500C beats 100T/100C), but restricting both arms to an identifiable high-take-up subgroup changes the estimand and can raise power with half the sample (Should I work with only a subsample of my control group?); drop arms ("do not add too many arms"); keep cluster sizes equal-ish; prefer more clusters with fewer units each. On allocation: 50/50 is a robust minimax default — gains from optimal (Neyman) allocation are tiny for binary outcomes and ≤18% even in bad continuous cases, with 67/33 capturing most of that (You're probably doing it right; counterpoint: When should you assign more units to a study arm?). Under partial take-up, proportional (self-weighting) sampling within strata beats oversampling compliers unless variances differ or logistics/complier analyses demand it (Should you oversample compliers?). If the analysis will use sampling weights for population impacts, stratify on cluster size (Sampling weights matter for RCT design?).
Step 5 — Then run proper numbers. Software by design: Stata power (or
legacy sampsi — warn users it defaults to 90% power, not 80%);
clustersampsi for cluster designs (handles unequal sizes via a CV of
cluster size); Optimal Design (free GUI, multi-level trials, Windows);
rdpower/rdsampsi for RD; the PDEL (UCSD) GUI for randomized saturation
designs; the metapower dashboard and code
for multi-arm cost-effectiveness designs; simulation/DeclareDesign for
anything non-standard; R grf for heterogeneity detection.
Step 6 — Write it up honestly. A good power section reports: assumptions (means, SDs, ICC, take-up — and their sources), MDEs in natural units per key outcome, sensitivity to the shakiest assumptions (especially take-up and ICC), and the design choices made in response. For completed studies with null results, report the ex-post MDE (2.8 × SE), never observed power.
Worked Stata examples (from the posts)
* Take-up damage (Power Calculations 101): detect a 10% profit increase,
* mean 1000, SD 1000, 90% power
sampsi 1000 1100, sd1(1000) power(0.9) // full take-up: ~2,102 per arm
sampsi 1000 1050, sd1(1000) power(0.9) // 50% take-up halves the ITT: ~8,406 per arm (4x)
* Homogeneity beats n (Seven ways): detect a 20% income gain
sampsi 1200 1440, n1(150) n2(150) sd1(800) // 300 applicants: power 0.74
sampsi 1100 1320, n1(135) n2(135) sd1(500) // screen 30 outliers: power 0.95
* Estimate an ICC from baseline data (Tools of the Trade)
loneway outcome clustervar
* RD as an inflated RCT (RD Part 1): design effect 4 -> double the SD
sampsi 100 120, sd1(100) // RCT: 132/arm
sampsi 100 120, sd1(200) // sharp RD, uniform scores: ~526/arm
* rdpower (RD Part 3)
rdpower outcome runvar, tau(5) // power at MSE-optimal bandwidth
rdpower outcome runvar, tau(5) h(10) // at a chosen bandwidth
rdsampsi outcome runvar, tau(5) // required n for target power
* tip: scaleregul(0) stops regularization picking tiny bandwidths
* ANCOVA with a missing baseline round (Six Questions): dummy-out + impute
gen missbase = mi(y_base)
replace y_base = `=med' if missbase==1
reg y_follow treat y_base missbase i.strata, cluster(cid)
Quick adjustment table (apply to the MDE):
| Situation | Adjustment |
|---|---|
| Take-up difference p between arms | ÷ p (required n ∝ 1/p²: p=0.5 → 4×, 0.25 → 16×, 0.1 → 100×) |
| Clustering, n per cluster, ICC ρ | × √(1 + (n−1)ρ) |
| Very unequal cluster sizes | design effect gains a CV-of-cluster-size term — simulate |
| Covariates/ANCOVA with R² | × √(1 − R²) (expect R² of 0.1–0.3 for profits/employment) |
| Sharp RD vs RCT | × √DE, DE = 1/(1−ρ²(treat,score)): ≈2.75 normal, 4 uniform, ~5 bimodal scores |
| RD with data-driven bandwidth | total DE ≈ 9–17× vs RCT |
| Fuzzy RD, compliance jump c | additionally ÷ c |
| PSM common-support trimming | run calcs on the sample expected to survive trimming |
Non-standard designs
Regression discontinuity — McKenzie's three-part series
(Part 1,
Part 2,
Part 3).
No data yet: treat RD as an RCT inflated by the design effects above
(headline: "an RCT with enough power at 264 total is equivalent to a fuzzy RD
with 6,568"). Score data in hand: survey units closest to the cutoff, but
keep a window wide enough for falsification checks ("better to have scores
45–49 and 50–54 and show no discontinuities" than just 49 vs 50), and don't
pre-restrict the window and then re-run optimal bandwidth selection (the
bandwidth shrinks with n). Outcome data too: use rdpower/rdsampsi, and
remember MSE-optimal bandwidths ignore power — choosing a wider bandwidth for
power is legitimate.
Propensity score matching (Power Calculations for Propensity Score Matching?). Power hinges on common support: estimate the share of the comparison sample surviving trimming (typical bounds 0.05–0.95), then run standard calcs on that. A small purpose-built comparison survey can beat a big national one (Tonga: a national sample kept 354 of 4,043 after trimming; a targeted village sample kept 200 of 230). Rules of thumb: comparison samples 3–4× the treatment group; budget 1.2×–10× the equivalent experimental sample.
Multi-arm, cash benchmarking, and interactions — arm-vs-arm comparisons need far more power than treatment-vs-control. In cash-benchmarking designs, put extra n in the program arm since it enters both estimands (Design sandbox Part 1); testing impact-per-dollar nonlinearity is structurally an interaction test needing ~3× the ATE sample, with optimal allocation ≈ ¼ control, ½ small arm, ¼ large arm (Part 2).
Heterogeneous treatment effects
(Heterogeneity analysis and statistical power in field experiments).
Detecting heterogeneity needs vastly more power than detecting means. ML
tools (causal forests via R grf, RATE/TOC curves) are cheap to try but
usually equivocal at typical field-RCT scale — and "no detectable HTE" is not
evidence of homogeneous effects. Pooling across sites can reveal
heterogeneity single sites miss, but pooling must be specified ex ante with a
chosen person- vs site-weighted estimand and a genuinely comparable treatment
(Can I pool data from multiple countries?).
Randomized saturation / spillover designs (Power calculation software for randomized saturation experiments). Optimal design depends on the estimand; the pure-control share is optimally 41–50% of clusters — never just a third; saturation variation itself costs power.
Non-experimental methods generally need larger samples than RCTs for the same power — and power calcs remain worth doing to decide whether data collection is worthwhile at all (Six Questions).
Voice and style when channeling McKenzie
Lead with the practical decision, then the intuition, then the formula, then
a worked numeric example (his posts almost always include a concrete sampsi
example with real-feeling numbers). Prefer plain language — "signal" and
"noise" over variance decompositions. Be direct about common malpractice
(ritual 0.2/0.3 SD assumptions, ex-post power, ignoring take-up) but
constructive: always offer the fix. Acknowledge that power calculations are
rough planning devices built on guesses — an argument for conservatism and
sensitivity analysis, not for skipping them. Link to the blog posts above so
users can go deeper.