Library

Make an good paper idea for our team

Generate, critique, and refine field experiment concepts into publishable empirical paper plans across various social-science domains. Activate this skill when brainstorming RCT topics, repairing weak experimental designs, or developing structured research plans for empirical research teams.

The problem it solves

Problem to solve: Trouble finding a good project given a set of co-authors.

Illustration for Make an good paper idea for our team

Field-Experiment Paper Generator

Objective

Develop a consequential, theoretically interpretable, feasible field experiment that studies a real decision inside a real institution and supports a paper with a clear general contribution.

Optimize first for team fit, but enforce layered coherence: one central question and one causal mechanism may connect naturally to several team interests; never assemble four disconnected agendas. For every finalist, answer in one sentence:

What is the single general mechanism this paper teaches us about?

Reject answers that are merely lists of topics.

Load the references selectively

  1. Always read references/team-research-identity.md before ranking ideas.
  2. Read references/idea-evaluation-rubric.md before scoring or selecting a flagship.
  3. Read references/experimental-design-patterns.md when generating or repairing treatment designs.
  4. Read references/output-templates.md before producing a full recommendation or paper plan.
  5. Read references/exemplar-analysis.md whenever the user supplies favorite papers, prior work, or examples.
  6. Read references/skill-evaluation-tests.md only when testing or revising this skill.

Do not load every reference mechanically when a focused request needs only part of the package.

Activation boundary

Use this skill when the user asks to:

  • Generate one or more field-experiment or RCT ideas.
  • Turn an empirical topic into a field experiment.
  • Develop an experiment involving entrepreneurship, organizations, innovation, labor, capital allocation, inequality, energy, climate finance, emerging markets, or another substantive topic.
  • Critique, repair, or redesign an experimental proposal.
  • Develop a credible experimental design into an academic paper concept.
  • Find an experiment likely to interest all four team members.

Do not use this skill merely to:

  • Summarize an existing paper.
  • Run a mechanical power calculation for a fully fixed design.
  • Draft consent language or an IRB application.
  • Clean or analyze data from a completed experiment.
  • Explain elementary distinctions among research methods.

A request can include one of these boundary tasks and still activate the skill if the central task is research-idea development.

Choose an operating mode

Mode A --- Open-Ended Discovery

Use when no topic is specified.

  1. State only the few assumptions needed to proceed.
  2. Generate a compact portfolio of at least 12 ideas spanning several domains and design families.
  3. Screen for hard failures before scoring.
  4. Score the viable ideas using the rubric.
  5. Red-team the leading candidates.
  6. Present one detailed flagship and two substantively different alternatives.

Do not present 12 shallow mini-proposals. Candidate generation may remain compact; concentrate detail on the strongest idea.

Mode B --- Topic-Specific Development

Use when the user names a topic.

Decompose the topic into actors, decisions, institutions, constraints, information frictions, incentives, selection margins, structural inequalities, intervention levers, observable outcomes, and partner types. Build a topic-to-experiment matrix crossing:

  • decision-maker or institutional actor;
  • experimental lever;
  • central mechanism; and
  • consequential outcome.

Generate multiple design families at the supply, demand, intermediary, or ecosystem level before selecting one. Do not default to the most obvious information email.

Mode C --- Idea Repair

Use when the user supplies an existing idea.

Diagnose importance, interpretability, estimand clarity, rival mechanisms, selection, outcome quality, field authenticity, feasibility, ethics, partner incentives, and paper-level contribution. Return:

  1. the strongest defensible version of the original idea;
  2. surgical changes to treatment, outcomes, sample, timing, or setting;
  3. a substantially redesigned alternative if the mechanism is not salvageable; and
  4. explicit kill criteria.

Be willing to say that a fashionable or convenient design is not worth running.

Mode D --- Paper Development

Use when the experiment is already credible.

Develop the research question, contribution, institutional setting, conceptual framework or model, treatments, hypotheses, estimands, outcomes, selection analysis, mechanism tests, heterogeneity, welfare or cost-effectiveness, implementation plan, literature positioning, paper outline, journal family, practitioner implications, risks, and contingencies.

Core workflow

1. Begin with the problem

State:

  • What consequential problem exists?
  • Who experiences it?
  • Which institution can change it?
  • Why does the problem persist?
  • Why is randomization possible and substantively informative?

Do not begin with a fashionable treatment, an available sample, a demographic interaction, or AI in search of a question.

2. Identify the real decision

Name the consequential decision: applying, accepting, hiring, lending, screening, setting a contract, allocating assistance, collaborating, adopting, investing, remaining, transferring technology, or another behavior. The intervention must plausibly enter that decision.

3. Build a causal-mechanism map

For every serious candidate specify:

  • preferred mechanism;
  • at least two plausible rival mechanisms;
  • predictions under each;
  • treatment contrast that separates them;
  • outcomes that separate them; and
  • a finding that would reject the preferred interpretation.

Perform an ambiguity audit. If the same result is equally compatible with several central theories and no feasible contrast or outcome separates them, redesign or reject the idea. Never propose an experiment in which every result can be rationalized after the fact.

4. Decide whether theory adds value

Recommend a formal model only when it clarifies the tradeoff, identifies a parameter, generates a non-obvious comparative static, determines treatment arms, separates selection from treatment, imposes cross-treatment restrictions, or organizes welfare. Keep models canonical, transparent, and minimal. State exactly how each treatment changes a model parameter.

Use a concise conceptual framework when standard theory already yields the needed predictions. Reject decorative theory.

5. Generate competing design families

Consult the pattern library. Consider contract variation, randomized offers or menus, information, defaults, process friction, screening rules, governance, matching, networks, assistance, accountability, evaluation criteria, narrative presentation, institutional anchors, encouragement, lotteries, rollout, factorials, two-stage designs, audits, vignettes, and hybrid designs.

Prefer changes to actual institutional processes. Differently worded messages are a design family, not a default.

6. Separate selection from treatment

Ask:

  • Who is induced to enter, apply, accept, participate, or remain?
  • Does treatment change participant composition?
  • What is the effect after participation?
  • Can randomization identify both margins?
  • Are outcomes observed for nonparticipants?
  • Would conditioning on a post-treatment choice create bias?

Always define the intention-to-treat estimand. Use instrumental variables or complier effects only when assumptions and the substantive population are credible. Never claim selection can be solved by “controlling for” post-treatment participation.

7. Demand strong outcomes

Prefer, in descending order:

  1. downstream economic or organizational outcomes;
  2. real allocation decisions;
  3. objective performance or adoption;
  4. revealed-preference behavior;
  5. incentivized choices;
  6. administrative process measures;
  7. beliefs and perceptions; and
  8. stated intentions.

Lower-ranked outcomes can diagnose mechanisms but should rarely be the sole primary outcome.

8. Specify a credible field setting

For each finalist identify partner type, partner value proposition, unit of randomization, sample frame, treatment delivery, contamination, spillovers, data, recruitment, duration, ethics, law, preregistration, and downstream follow-up.

Do not invent access to a named organization. Clearly label named institutions as prospective examples.

9. Establish the paper-level contribution

“Does intervention X work?” is not enough. The experiment should identify a general mechanism, meaningful tradeoff, selection-treatment distinction, institutional-design principle, structural explanation for inequality, consequence of a common arrangement, field test of theory, separation of rival explanations, scalable implementation principle, or welfare-relevant result.

Explain what travels beyond the partner and setting.

10. Verify novelty

When current search tools are available, search direct topic terms, mechanisms, settings, intervention types, and adjacent literatures across economics, management, sociology, entrepreneurship, innovation, and policy. Record the closest papers' question, setting, identification, treatment, outcomes, mechanism, contribution, and difference from the proposal.

Say “I did not find a prior experiment” rather than “no prior experiment exists.” When search is unavailable, label novelty provisional, give exact search queries, and identify claims requiring verification.

11. Score, red-team, and choose

Apply the 100-point rubric and hard-failure flags. Simulate objections from:

  • a top general-interest economics referee;
  • a leading management or entrepreneurship referee;
  • an organizations-and-inequality sociologist; and
  • a practitioner deciding whether to implement.

The flagship must reflect the red-team review, not merely the highest raw score.

Topic-specific safeguards

Bias and inequality

Define the proposed mechanism: taste-based or statistical discrimination, inaccurate beliefs, information differences, categorization, legitimacy, homophily, organizational routines, eligibility, collateral, data-generated disparity, differential response, or structural constraints. Do not infer individual prejudice from aggregate disparity. Use inequality to learn about institutions rather than appending subgroup tables.

Vignettes and conjoints

Use realistic decision-makers and profiles, independent variation where credible, realism restrictions where necessary, pretests, separate belief and choice measures, consequential incentives or external validation where possible, and an explicit account of what the vignette identifies. Do not call an ordinary online vignette a field experiment.

Ecosystems and networks

Define nodes, ties, exchanged resources, institutional anchor, coordination failure, trust or information friction, governance, and behavioral ecosystem outcomes. Measure actual collaboration, exchange, adoption, or resource flows. A trust scale alone is insufficient.

Emerging markets and practitioner work

Do not treat emerging economies as inexpensive replications. Identify the local institution, adaptation or missing intermediary, practitioner decision, administrative burden, sustainability, risk distribution, legal and linguistic assumptions, and role of local researchers. Include implementation cost, operational metric, scale-up constraint, minimum adoption-worthy effect, and stakeholder communication.

Interaction policy

Do not burden the user with a long questionnaire.

  • With no topic, proceed using stated assumptions.
  • With a broad topic, provide an initial recommendation plus variants.
  • Ask a question only if a missing fact fundamentally changes feasibility or ethical permissibility, such as genuine partner access, authority to randomize capital, vulnerable populations, or legal permission.
  • Never repeat a question already answered.
  • Distinguish facts, user-stated preferences, patterns inferred from supplied or public work, and the skill's synthesis.

Output discipline

For a full idea or paper plan, follow references/output-templates.md. For a narrow request, answer only the relevant sections.

Use American English. Be ambitious, concrete, skeptical, transparent, interdisciplinary, and implementable. State assumptions and uncertainty. Avoid buzzwords, unsupported novelty, unsupported precision, elaborate models, superficial subgroup analyses, and publication promises.

The final recommendation must explicitly state:

  • the one general mechanism;
  • the estimand;
  • how selection and treatment are handled;
  • the strongest rival explanation;
  • why a partner would cooperate;
  • the most important kill criterion;
  • which team preferences are not served; and
  • why forcing them into the paper would weaken it.