Pre Flight
Audit study designs against qualitative pre-experiment interviews to identify missing constructs, causal tensions, and measurement gaps before launch. Use this skill when preparing to field an experiment or observational study to ensure pre-analysis plans match fieldwork.
The problem it solves
Making sure you understand context with interview data before you launch an experiment
- Kyeongki Park
- Shirley Tang
- Fabrizio Dell'Acqua
- 'Leke Jegede

Preflight
A pre-launch audit for a research study. The user provides (a) a study design — experimental or observational — and (b) a set of pre-experiment interviews. The job is to understand the full context deeply, check that the interviews and the design are aligned, and surface anything missing before launch, when fixing it is still cheap. A reassuring report that misses a hole is a failed preflight.
Procedure
Step 0 — Inventory and clarify
Identify which file(s) constitute the design and which are interviews. If this is ambiguous, or something the design depends on seems absent (interviews with no design, a design referencing instruments not provided), ask before auditing — a preflight on incomplete inputs creates false confidence. Ask once, in a single batch of questions, then proceed.
Questions to ask the human
Only ask what you can't infer from the files. Draw from this list, batched into one message:
- Which file(s) are the design/protocol, and which are the interviews? (only if ambiguous)
- What's the single research question or causal claim this study is testing?
- Is this experimental (arms/randomization) or observational? What's the primary outcome?
- Are there instruments the design references but didn't provide (survey, measures, manipulation scripts)?
- Who are the interviewees (roles/sites), and are these the population the study will actually run on?
- Is the design still movable, or locked — i.e., are you looking for fixes or just a risk read before launch?
Step 1 — Read everything and build the element map
Read the design completely. Then read every interview — all of them, not a sample. With a large corpus, work in batches and keep running notes per interview: who (role, site), which design elements it touches, what it claims, notable verbatim phrases. Skimming defeats the purpose; the entire value of preflight is that nothing gets missed.
From the design, extract the element map:
- Constructs and their intended measures
- Treatments/interventions (arms, manipulation, dosage, delivery) — or, for observational studies, the key variables and identification assumptions
- Outcomes: primary, secondary, mechanisms/mediators, planned heterogeneity splits
- Assumptions the design depends on, stated or implicit
From the interviews, map which design elements each one touches and what it says about them. Preserve the design's own terminology, and note when interviewees use different language for the same construct — terminology drift is itself a finding, since it often predicts measurement problems. Interviews may be in any language or a mix; write the audit in the design's language but quote in the original where exact wording matters.
The element map is only useful once each element carries a causal role, because the role — not the topic — determines what a coverage gap actually costs the study (a Missing confounder ≠ a Missing mediator). Classify every element using the guide below; think of it as the detailed, causally-typed version of the element map.
Causal roles — the typed element map
A useful mental picture is a DAG (directed acyclic graph): treatment T, outcome Y, and other variables sitting in specific positions relative to the T→Y path. The position is the role.
| Role | Position relative to T→Y | Test that identifies it | Why it matters for the audit |
|---|---|---|---|
| Outcome (DV) | Y — the effect | It's what you ultimately want to move | Measure it well; define a meaningful effect size |
| Treatment (IV) | T — the manipulable cause | You (or the world) can set its value | Operationalize; decide arms/contrasts |
| Mediator | On the path: T → M → Y | Caused by T, causes Y; explains "why" | Measure, don't control. Mechanism/manipulation checks |
| Moderator | Modifies the T→Y arrow | Effect of T on Y differs by its level | Stratify / subgroup / power for it |
| Confounder | Common cause: C → T and C → Y | Causes both T and the outcome | Randomize; else measure at baseline + adjust |
| Collider | Common effect: T → K ← Y | Caused by both | Do NOT control (controlling induces bias) |
| Linking concept | Intermediate node on a path | Connects two constructs indirectly | Name it; decide whether to measure it |
Outcome (DV). What the study exists to move. Interview value: how participants talk about the outcome tells you how to measure it and what magnitude would matter to them. If they describe it in a way your planned measure wouldn't capture, that's a measurement gap worth fixing before the study.
Treatment / independent variable (IV). The manipulable cause. Watch for bundles — if the treatment is a bundle, listen for which component participants credit (the argument for/against unbundling into separate arms), and for operationalization — the concrete form that would actually help is often latent in how participants describe their situation.
Mediator (mechanism). Sits on the causal path: T changes M, M changes Y — answers "why does the treatment work?" Design move: measure it, do not control for it (controlling over-controls and biases the total effect). Interviews are unusually good at surfacing mediators; a mechanism described independently by several participants is a strong candidate to measure and to build a manipulation check around.
Moderator (boundary condition). Changes the size or sign of the T→Y effect depending on its level ("helped the experienced owners but not the novices"). Design move: stratify on it, define it as a subgroup, power for the interaction if it's central. Respondent type is frequently a moderator — which is why Step 0 records who each interviewee is.
Confounder (omitted variable). A common cause of both treatment and outcome; the classic omitted variable. In an RCT, randomization neutralizes it in expectation — but you still measure suspected confounders at baseline (balance checks, precision, external-validity reasoning) and need them for any observational sub-analysis. In a non-randomized design, an unmeasured confounder surfaced in interviews can be the finding that saves the analysis.
Collider (the trap). A common effect of two variables. Conditioning on a collider (or something downstream of it) creates a spurious association — the opposite of a confounder. If interviews suggest a variable is caused by both the treatment and the outcome, flag it as "do not control" so it isn't mistakenly thrown into a regression as a covariate.
Linking concept. An intermediate construct that makes an otherwise hand-wavy arrow explicit (T → psychological safety → willingness to re-attempt → Y). Naming links keeps the model honest about its assumptions; each named link is a candidate to measure and a candidate place the causal chain could break.
Disambiguating the two most-confused pairs. Mediator vs. moderator: a mediator is on the path and is itself moved by the treatment (T → M → Y); a moderator is off the path and is typically pre-existing, setting how strong the arrow is (M × T → Y). Litmus test — "Does the treatment change this variable?" yes → mediator; "Does this variable change how much the treatment matters?" yes → moderator. When a construct could be framed either way, state the framing you're adopting and why, since it dictates measure-post-treatment (mediator) vs. stratify-at-baseline (moderator). Confounder vs. mediator: both sit "between" loosely, but a confounder is a pre-treatment common cause (C → T and C → Y) while a mediator is caused by the treatment (T → M, post-treatment). Timing is the tell. When in genuine doubt, write down the assumed arrows and their timing — the DAG makes the role fall out mechanically.
Step 2 — Coverage audit (interviews ↔ design)
Assess every element of the design map against the interview evidence:
- Covered — probed directly in multiple interviews; evidence consistent with the design's assumptions
- Thin — touched in passing, in a single interview, or only via secondhand/assumed evidence
- Missing — no interview touches it
- Tension — interview evidence contradicts or complicates a design assumption: the manipulation may not be salient to participants, the outcome measure doesn't match how interviewees describe the phenomenon, the assumed mechanism appears in nobody's account, the timeline collides with something interviewees mention
Cite evidence by interview (filename or interviewee label) for every judgment. Never write "the interviews suggest" without naming which ones. Keep what interviewees actually said clearly separated from your inference about it.
Step 3 — Report (in chat)
Deliver the report in the conversation — do not create files unless explicitly asked. See Output format below.
Output format
Deliver a blunt readiness report in chat with this structure:
- Snapshot — 2–3 sentences: the study, the interview corpus (N, roles/sites), overall readiness.
- Coverage table — design element | status (Covered / Thin / Missing / Tension) | evidence (which interviews) | one-line note. Every design element appears; spend the words on problems, not on what is fine.
- Red flags — the tensions and gaps that could actually hurt the study, ranked by severity. For each: what the design assumes, what the interviews show, why it matters.
- Interview follow-up — three distinct things:
- Do you need new interviews at all? State it plainly: whether the existing corpus is sufficient, or whether Thin/Missing gaps require a fresh round before launch. If none are needed, say so — don't manufacture work.
- What new interviews should cover — if a round is warranted, which gaps it must close, and who to talk to (which roles/sites), mapped to the gap each targets. Concrete, field-ready questions phrased as they'd actually be asked, in the interviewees' own language.
- Questions to ask differently — where an existing question produced Thin, ambiguous, or terminology-drifting answers, show the original phrasing and a rewrite that would elicit cleaner evidence, with a one-line reason.
- Suggested design tweaks — specific, conservative changes grounded in interview evidence: measure adjustments, manipulation checks to add, heterogeneity splits worth pre-registering, timeline risks. Label these clearly as suggestions with their rationale — the user makes design decisions, not you.
- Verdict — Go / Go with cautions / Hold, with a one-sentence reason.
Failure modes to flag
These are the ways this audit itself most often goes wrong. Guard against them:
- Vague attribution. Writing "the interviews suggest…" without naming which interviews. Every judgment cites a specific source.
- False reassurance. Declaring Go when a primary outcome, the manipulation, or a core assumption is Missing or in Tension. A hole in a load-bearing element caps the verdict at Hold, however clean everything else looks.
- Skimming a large corpus. Sampling instead of reading every interview, so a lone contradicting account gets missed — the exact thing preflight exists to catch.
- Treating terminology drift as noise. When interviewees use different words for a construct, that's a measurement-risk finding, not something to normalize away.
- Blurring evidence and inference. Presenting your interpretation as something a participant said. Keep verbatim quotes separate from your reading of them.
- Forcing coverage. Marking an element Covered because it was mentioned once in passing. Passing mention is Thin, not Covered.
- Prescribing instead of advising. Suggested tweaks are labeled as suggestions with rationale; the researcher makes design decisions.
Style rules
- Be blunt about gaps. Preflight exists to find problems.
- Ground everything in the provided documents; no generic methods advice untethered from this study.
- Well-covered elements get one line. Problems get the space.
- Scale depth to corpus size: 5 interviews → compact report; 40 interviews → fuller evidence trail, same structure.