Treatment invertibility stress-test
Audit an experimental design to check whether a proposed treatment uniquely maps to its target theoretical construct. Use before data collection to identify confounding constructs, construct invalidity, and required design fixes or control arms.
The problem it solves
Stress-test an experimental treatment for 1-1 mapping (invertibility) between the manipulation and the theoretical construct it claims to move.
- Brent Goldfarb

Treatment invertibility stress-test
What this skill does and why
A hypothesis lives at the construct level ("status increases advice-taking"). An experiment lives at the treatment level ("we told half the subjects the advisor went to Harvard"). The inference the researcher wants to make runs backward: treatment effect observed, therefore construct did the work.
That backward inference is only licensed if the mapping from construct to treatment is invertible: given the effect, the construct-level explanation is unique. It almost never is. The treatment is one point in a large set of admissible specifications of the construct, and it usually moves other constructs on the way down. The researcher's judgment call at each specification step is exactly where rival explanations enter, and those calls are usually tacit. This skill makes them explicit and stress-tests each one.
The framing to hold throughout: specification is a descent down an abstraction ladder. The construct sits at the top. Each rung down is a choice (context, marker, delivery, dosage, wording). By the bottom rung you have the actual stimulus. Invertibility fails in two ways:
- Fan-out (one-to-many downward). Many admissible treatments specify the same construct. If the effect depends on which one you picked, you learned about your stimulus, not your construct.
- Fan-in (many-to-one upward). Your one treatment is an admissible specification of several constructs. The observed effect inverts to multiple theories.
A design severely tests the hypothesis only if the predicted effect would be improbable were the hypothesis false (Mayo). Every surviving rival construct is a way the effect shows up with the hypothesis false. The audit's job is to enumerate rivals and kill them, or to say plainly that the design cannot.
Phase 1: Interrogate. No audit until the mapping is pinned down.
Do not audit a vague design. If the user has a document, extract what you can from it first; then ask only for what is missing. Otherwise interrogate directly (use AskUserQuestion where available, batched, one round if possible). You need all six items:
- The hypothesis, verbatim, with the construct named. Push on ambiguity: if the hypothesis admits two readings, the audit is over before it starts, say so and get it fixed. Is the claim causal or relational? What is the implied direction and functional form?
- The construct definition the user is committed to. Which literature's version? (Status as deference-earning rank is a different construct from status as quality signal.)
- The exact treatment, both arms, at stimulus level: the actual wording, image, price, confederate script. "We manipulate status" is a construct, not a treatment. Get the materials or the closest draft. The control arm matters as much as the treatment arm: what is held at baseline, and is the control a true zero of the construct or a different point on it?
- The sample and recruitment: who, from where, incentivized how, and what the subject believes the study is about.
- The target setting: the real-world context the hypothesis is actually about (founders pitching VCs, employees taking advice, buyers evaluating products). Not "organizations generally." Name it.
- The outcome measure, specified at the same level of concreteness as the treatment.
If the user cannot supply an item, that is itself a finding. Record it as a red flag, do not fill the gap with a charitable guess.
Phase 2: Build the specification ladder
Before scoring anything, write the ladder out explicitly. Top rung: the construct as defined in item 2. Bottom rung: the stimulus from item 3. Name each intermediate rung and the choice made at it.
Example (status):
- Rung 0: Status (rank conferring deference expectations)
- Rung 1: Context: advice-taking between strangers online
- Rung 2: Status marker: educational pedigree (chosen over title, follower count, endorsement)
- Rung 3: Specific marker: "Harvard MBA" vs. no mention (chosen over Harvard vs. state school)
- Rung 4: Delivery: one line in a profile vignette, subject reads silently
At every rung, list the admissible alternatives the researcher did not pick. This is the fan-out. Then, at the bottom, list every construct for which this same stimulus is an admissible specification. This is the fan-in. ("Harvard MBA" specifies status, but also perceived competence, perceived similarity to the subject, expected income, political priors.)
The ladder is the audit's working object. Everything in Phase 3 refers back to it.
Phase 3: The five-dimension audit
Score each dimension. For each: state the finding in one or two sentences, name the specific rival constructs or failure it implies, and rate severity (LOW / MODERATE / SEVERE). Severity means: how much does this dimension, alone, undermine the inference from effect back to construct?
1. Abstraction distance
Count the rungs between construct and stimulus and the width of the fan at each. More rungs and wider fans mean more researcher judgment embedded in the design and a larger admissible set the single treatment is standing in for. The question to answer: if a colleague re-specified the construct independently, honestly, from the same definition, what is the chance they land on a materially different treatment? If high, a null or an effect is uninformative about the construct without a robustness arm. Distance itself is not damning; undocumented distance is. Flag any rung where the user made a choice they cannot defend as forced.
2. Manipulation depth
How heavy is the intervention? A one-line vignette difference is shallow; a confederate interacting for ten minutes, a real stake, a changed institutional rule is deep. Depth buys realism and detectability but moves more constructs at once: a deep manipulation is a bundle. Enumerate everything the manipulation plausibly moves besides the target construct, including affect, attention, demand (subjects inferring the hypothesis), and salience of the experiment itself. The exclusion-restriction question, asked at design time: does the treatment reach the outcome only through the target construct? List every other path.
3. Behavior vs. composition
Does the treatment change what agents do, or which agents show up and act? Selection masquerades as treatment constantly: an incentive manipulation that changes who persists to the outcome measure, a framing that induces differential attrition, a recruitment message that draws a different pool across arms. Ask: is the set of subjects contributing outcome data identical in composition across arms, and would the marginal subject induced in or out by the treatment differ on anything correlated with the outcome? If the answer is "the treatment works partly by changing who acts," the hypothesis being tested is not the one stated.
4. Type vs. expected behavior
When the manipulation targets perceptions (most vignette and signaling designs do), pin down what perception it moves: the target's type (this person is competent, high-quality, trustworthy) or expected behavior (this person will be listened to, will get resources, is owed deference)? These are different theoretical claims with different rivals and different downstream predictions, and most status, legitimacy, and signaling manipulations move both. If the hypothesis is about one and the treatment moves both, the design needs either a discriminating arm or a discriminating measure. Ask what pattern of results would distinguish them; if no pattern in this design can, say so.
5. Setting correspondence
Compare the experimental setting to the target setting from interrogation item 5, rung by rung. Does the ladder that exists in the experiment exist in the field? Do agents in the target setting encounter this construct through anything like this marker, at this dosage, with these stakes, with this information set? A clean lab effect on a ladder that has no field analog answers a different question than the hypothesis asked. Also run it in reverse: in the field, the construct arrives bundled with correlates (status arrives with resources, networks, track records); the lab isolates the construct precisely by cutting those bundles, so state which bundle-cuts change the meaning of the construct itself.
Phase 4: The invertibility test
The summary move. Instruct the user to imagine the study ran and the predicted effect appeared, clean and significant. Then write down every construct-level story consistent with that result, drawing on the fan-in list and every rival surfaced in Phase 3. Number them.
One surviving story: the mapping is invertible; the design severely tests the hypothesis. Two or more: it does not, and the audit must say which arms, measures, or design changes would kill each rival:
- Rival-killing arm: a condition that moves the rival without the construct, or vice versa (Harvard MBA vs. equally-competent-but-low-pedigree profile separates status from competence).
- Discriminant manipulation check: measure the rival construct too, not just the target. A manipulation check that only confirms the target moved cannot rule out that the rival moved with it.
- Dose or marker variation: a second specification from the admissible set. If the effect survives re-specification, the fan-out worry shrinks.
- Composition lock: measure the outcome on the full assigned sample, or show attrition/participation balance, when dimension 3 flagged selection.
Distinguish fixable from fatal. A rival killable with one added arm is a design note. A rival inherent to the setting (the construct cannot be unbundled from it in any admissible specification) is a scope limit the paper must own, and the honest move is to weaken the hypothesis to what the design can actually test.
Output format
Produce a working document, not a memo to a stranger. The user is a skeptical academic auditing their own design pre-launch; write like a sharp colleague at a whiteboard. Follow this structure exactly:
Invertibility audit: [construct] via [treatment shorthand]
The mapping under test
Hypothesis (verbatim), construct definition, treatment (both arms), sample, target setting, outcome. One line each. Flag anything the user could not supply.
Specification ladder
The rungs, the choice at each, the admissible alternatives not taken. Then the fan-in list: every construct this stimulus admissibly specifies.
Audit
The five dimensions, each: finding, named rivals, severity rating.
Invertibility test
The numbered rival stories. For each: the fix (arm, measure, variation, lock) or the word "fatal" with the scope limit it implies.
Verdict
- GREEN: One surviving story, or all rivals killed by the existing design. Proceed.
- YELLOW: Rivals survive but each has a concrete fix. List the fixes as a punch list, ordered by cost.
- RED: A rival is fatal in this design, or the hypothesis/construct/treatment was too underspecified to audit. State what has to change before the audit can go green, which may be the hypothesis itself.
One paragraph after the color: the two or three things that matter most, in plain prose. No summary beyond that.
Style constraints for the output
The user maintains a strict writing style file. In the audit document: short paragraphs, active voice, contractions fine, no em dashes (use commas, colons, parentheses). Never use these words: delve, robust, leverage, crucial, pivotal, foster, showcase, enhance, holistic, seamless, underscore, highlight (as verb), landscape (abstract), meticulous, innovative, transformative. No negative-parallelism constructions ("This isn't X, it's Y"): state the affirmative claim first and let contrast follow if needed. No praise of the design, no reassurance, no "great question." If the design is weak, the verdict says so plainly. Concrete beats abstract everywhere: name the actual stimulus wording in findings, not "the manipulation."
What this skill is not
Not a referee report (that skill audits finished papers ex post; this one audits designs while they can still change). Not a power analysis, not a preregistration template, not an ethics review. If the user needs the fuller preregistration audit (sampling plan, stopping rules, analysis plan, multiple-testing), note that this skill covers the treatment-construct mapping only and the rest belongs to a separate prereg workflow. If sampling issues surface during dimension 3, report them as composition findings; do not expand into a general sampling audit.