Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when deciding which robustness checks belong in the few in-text words of an American Economic Review: Insights (AER: Insights) short-format manuscript versus the online Supplemental Appendix. Triages checks against the length cap; it does not design the identification (see aeri-identification).
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 2% | 0% |
A standard paper answers every threat in-text; AER: Insights cannot. With ≤7,000 words (minus 200 per exhibit, ≤5 exhibits), robustness must be triaged: keep in-text only the one or two checks that, if they failed, would kill the headline result, and move everything else to the Supplemental Appendix (which has no length cap but is read only by skeptics and the data editor). The skill is deciding the minimum credible set of in-text checks and signposting the rest in a single sentence.
For each candidate check, ask: "If a referee saw only the main text, would the absence of this check make them disbelieve the headline result?"
| Answer | Placement | |--------|-----------| | Yes — it defends the core identifying assumption | In-text (1–2 max), often a single robustness exhibit or a sentence | | No — it reassures but the result survives without seeing it | Supplemental Appendix, named in one in-text sentence | | It is really an extension / new question | Appendix or a different paper, not "robustness" |
Use one sentence to point to the appendix: "The result is robust to list] (Supplemental Appendix Section X; Tables X1–X6)." This buys credibility for the price of a sentence and keeps the word budget for the insight. Reproduce nothing in-text that the sentence already covers.
Run the battery, don't just enumerate it. Full map: execution-with-mcp. AER: Insights is a short format built around one decisive result, so the body/appendix split is even tighter — run the design cleanly the first time.
romano_wolf (step-down FWER, accounts forcross-test correlation) or benjamini_hochberg — report the adjusted threshold.
oster_delta / sensemakr — the confounder strength that wouldoverturn the headline.
wild_cluster_bootstrap (few clusters), twoway_cluster / conley.audit_result(result_id) lists the missing checks and theexact suggest_function for each — no guessing the battery.
etable / did_summary_to_latex from the handle — no retyped numbers.Keep the decisive checks in the body and the exhaustive (now actually-run) battery in the appendix. See the executed chain in the JF execution walkthrough.
【Core-assumption check (in-text)】<the 1–2 that must stay>
【Signpost sentence】"Robust to [list] (Suppl. Appendix §X, Tables …)."
【To the appendix】alt bandwidths / forms / placebos / heterogeneity / clustering
【Exhibit budget】in-text robustness uses __ of 5 exhibits
【Next step】aeri-tables-figuresOther measured skills in the registry, with their headline benchmark lift.