Pre-flight
Audit study designs against pre-experiment qualitative interviews to identify gaps, misalignments, and unsupported causal assumptions before launching a trial. Use this skill before fielding an experiment or observational study to confirm interviews support planned measures, treatments, and outcomes.
The problem it solves
Making sure you understand context with interview data before you launch an experiment
- Kyeongki Park
- Shirley Tang
- Fabrizio Dell'Acqua
- Leke Jegede

Preflight
A pre-launch audit for a research study. The user provides (a) a study design — experimental or observational — and (b) a set of pre-experiment interviews. The job is to understand the full context deeply, check that the interviews and the design are aligned, and surface anything missing before launch, when fixing it is still cheap. A reassuring report that misses a hole is a failed preflight.
Procedure
Step 0 — Inventory and clarify
Identify which file(s) constitute the design and which are interviews. If this is ambiguous, or something the design depends on seems absent (interviews with no design, a design referencing instruments not provided), ask before auditing — a preflight on incomplete inputs creates false confidence. Ask once, in a single batch of questions, then proceed.
Questions to ask the human
Only ask what you can't infer from the files. Draw from this list, batched into one message:
- Which file(s) are the design/protocol, and which are the interviews? (only if ambiguous)
- What's the single research question or causal claim this study is testing?
- Is this experimental (arms/randomization) or observational? What's the primary outcome?
- Are there instruments the design references but didn't provide (survey, measures, manipulation scripts)?
- Who are the interviewees (roles/sites), and are these the population the study will actually run on?
- Is the design still movable, or locked — i.e., are you looking for fixes or just a risk read before launch?
Step 1 — Read everything and build the element map
Read the design completely. Then read every interview — all of them, not a sample. With a large corpus, work in batches and keep running notes per interview: who (role, site), which design elements it touches, what it claims, notable verbatim phrases. Skimming defeats the purpose; the entire value of preflight is that nothing gets missed.
From the design, extract the element map:
- Constructs and their intended measures
- Treatments/interventions (arms, manipulation, dosage, delivery) — or, for observational studies, the key variables and identification assumptions
- Outcomes: primary, secondary, mechanisms/mediators, planned heterogeneity splits
- Assumptions the design depends on, stated or implicit
From the interviews, map which design elements each one touches and what it says about them. Preserve the design's own terminology, and note when interviewees use different language for the same construct — terminology drift is itself a finding, since it often predicts measurement problems. Interviews may be in any language or a mix; write the audit in the design's language but quote in the original where exact wording matters.
When classifying an element's causal role (outcome, treatment, mediator, moderator,
confounder, collider, linking concept) — and especially when disambiguating the
confusable pairs (mediator vs. moderator, confounder vs. mediator) — consult
references/causal-roles.md. Getting the role right is what makes the coverage audit
meaningful: role, not topic, determines what a gap actually costs the study.
Step 2 — Coverage audit (interviews ↔ design)
Assess every element of the design map against the interview evidence:
- Covered — probed directly in multiple interviews; evidence consistent with the design's assumptions
- Thin — touched in passing, in a single interview, or only via secondhand/assumed evidence
- Missing — no interview touches it
- Tension — interview evidence contradicts or complicates a design assumption: the manipulation may not be salient to participants, the outcome measure doesn't match how interviewees describe the phenomenon, the assumed mechanism appears in nobody's account, the timeline collides with something interviewees mention
Cite evidence by interview (filename or interviewee label) for every judgment. Never write "the interviews suggest" without naming which ones. Keep what interviewees actually said clearly separated from your inference about it.
Step 3 — Report (in chat)
Deliver the report in the conversation — do not create files unless explicitly asked. See Output format below.
Output format
Deliver a blunt readiness report in chat with this structure:
- Snapshot — 2–3 sentences: the study, the interview corpus (N, roles/sites), overall readiness.
- Coverage table — design element | status (Covered / Thin / Missing / Tension) | evidence (which interviews) | one-line note. Every design element appears; spend the words on problems, not on what is fine.
- Red flags — the tensions and gaps that could actually hurt the study, ranked by severity. For each: what the design assumes, what the interviews show, why it matters.
- Follow-up interview questions — concrete, field-ready questions that would close each Thin or Missing gap, mapped to the gap they close. Phrase them as they would actually be asked, in the interviewees' own language.
- Suggested design tweaks — specific, conservative changes grounded in interview evidence: measure adjustments, manipulation checks to add, heterogeneity splits worth pre-registering, timeline risks. Label these clearly as suggestions with their rationale — the user makes design decisions, not you.
- Verdict — Go / Go with cautions / Hold, with a one-sentence reason.
Failure modes to flag
These are the ways this audit itself most often goes wrong. Guard against them:
- Vague attribution. Writing "the interviews suggest…" without naming which interviews. Every judgment cites a specific source.
- False reassurance. Declaring Go when a primary outcome, the manipulation, or a core assumption is Missing or in Tension. A hole in a load-bearing element caps the verdict at Hold, however clean everything else looks.
- Skimming a large corpus. Sampling instead of reading every interview, so a lone contradicting account gets missed — the exact thing preflight exists to catch.
- Treating terminology drift as noise. When interviewees use different words for a construct, that's a measurement-risk finding, not something to normalize away.
- Blurring evidence and inference. Presenting your interpretation as something a participant said. Keep verbatim quotes separate from your reading of them.
- Forcing coverage. Marking an element Covered because it was mentioned once in passing. Passing mention is Thin, not Covered.
- Prescribing instead of advising. Suggested tweaks are labeled as suggestions with rationale; the researcher makes design decisions.
Style rules
- Be blunt about gaps. Preflight exists to find problems.
- Ground everything in the provided documents; no generic methods advice untethered from this study.
- Well-covered elements get one line. Problems get the space.
- Scale depth to corpus size: 5 interviews → compact report; 40 interviews → fuller evidence trail, same structure.