Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when an American Economic Journal: Applied Economics (AEJ: Applied) manuscript's headline estimate must be shown to survive specification, sample, and inference choices before submission or in an R&R. Builds the robustness suite a sophisticated referee expects; it does not establish the primary identification (aeja-identification) or format the exhibits (aeja-tables-figures).
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 15% | 0% |
AEJ: Applied referees probe whether the headline number is stable, honestly inferred, and not the product of researcher degrees of freedom. Robustness here is not a wall of regressions — it is a targeted set of checks each tied to a specific threat to the design. Map every plausible objection to the one check that answers it, and report the checks so the reader sees the estimate barely moves.
| Threat to the result | The check that answers it | |----------------------|---------------------------| | Omitted confounders | Oster δ / coefficient-stability bounds; added controls in steps | | Specification search | a specification curve / multiverse; pre-registered primary spec | | Functional form | levels vs logs, alternative outcome definitions, nonparametric version | | Sample selection | drop influential units, alternative inclusion rules, balanced vs unbalanced panel | | Inference too narrow | clustered SEs at the right level, wild-cluster bootstrap (few clusters), randomization inference | | Design-specific fragility | DID: honest-DID bounds; RD: bandwidth/donut; IV: weak-IV-robust set | | Multiple outcomes/subgroups | Romano–Wolf / List–Shaikh–Wooldridge MHT adjustment |
Run the battery, don't just enumerate it. Full map: execution-with-mcp. AEJ: Applied is applied microeconomics — labor, health, education, and development field settings where a clean research design is the entry ticket.
romano_wolf (step-down FWER, accounts forcross-test correlation) or benjamini_hochberg — report the adjusted threshold.
oster_delta / sensemakr — the confounder strength that wouldoverturn the headline.
wild_cluster_bootstrap (few clusters), twoway_cluster / conley.audit_result(result_id) lists the missing checks and theexact suggest_function for each — no guessing the battery.
etable / did_summary_to_latex from the handle — no retyped numbers.Keep the decisive checks in the body and the exhaustive (now actually-run) battery in the appendix. See the executed chain in the JF execution walkthrough.
An IV estimate of the return to a training program is 0.11 (s.e. 0.04). The robustness suite: (i) effective F of 23 rules out weak instruments; (ii) the Anderson–Rubin 95% set is 0.04, 0.19], so inference is not weak-IV-fragile; (iii) Oster δ implies selection on unobservables would need to be 1.8× selection on observables to nullify it; (iv) wild-cluster bootstrap with 14 clusters keeps the CI away from zero; (v) dropping the largest region moves the estimate to 0.10. The point estimate barely moves — the AEJ: Applied target.
specification curve in which the point estimate barely moves.
bootstrap or randomization-inference p-value.
unobservables would have to be (relative to observables) to nullify the result.
【Primary spec】declared / pre-registered? [Y/N] — estimate: ___ (s.e. ___)
【Threat → check map】selection: ___ | spec-search: ___ | form: ___ | sample: ___ | inference: ___ | design: ___
【Inference】clustering level: ___; few-cluster/randomization: ___
【Design sensitivity】honest-DID / RD bandwidth / weak-IV set: ___
【Estimate stability】range across checks: [___, ___]; checks that move it: ___
【Next step】aeja-tables-figuresOther measured skills in the registry, with their headline benchmark lift.