Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when building or revising exhibits for an AEJ: Economic Policy manuscript so they meet AEA house style and carry the policy message — no significance asterisks, a self-contained headline exhibit, and figures that show the policy effect with uncertainty. Designs exhibits; it does not run the estimation or write the surrounding prose.
.claude/skills/brycewang-stanford-aejpol-tables-figures/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 28% | 0% |
Every AEJ: Policy paper should have one exhibit a policymaker could screenshot: the policy effect in interpretable units with its welfare/cost-benefit reading where possible. Examples (illustrative formats):
esttab/booktabs), no vertical rules; align decimals; consistent digits.Generate exhibits from the fitted result, not by retyping numbers (the usual source of body-vs-appendix drift). Full map: execution-with-mcp.
etable (multi-model columns) or did_summary_to_latex straight from theresult_id — one variable definition, one set of numbers, body and appendix in sync.
plot_from_result / enhanced_event_study_plot / event_study_table —axis units and the SE/clustering note baked in.
states the magnitude in interpretable units.
See a full fitted-result → exhibit chain in the JF execution walkthrough.
| Design | Workhorse figure | What the notes must state | |---|---|---| | DID / event study | Event-study coefficients with CI bands; flat pre-period leads visible | estimator (CS/SA), comparison group, clustering level | | RDD | Binned means + fitted discontinuity + bandwidth | running variable, bandwidth, density-test result | | Bunching | Empirical vs. counterfactual density at the kink | counterfactual construction, excluded region | | RCT | Treatment-control means / dose-response with CIs | randomization unit, take-up, ITT vs. ToT | | Welfare | MVPF / cost-per-outcome across variants, with bands | which estimates feed the ledger, assumptions |
A tax-credit paper's main table reports a coefficient of 0.08 (s.e. 0.02) on the credit. Reworked for AEJ: Policy: the lead exhibit becomes a figure of employment around the credit's introduction with a CI band, the long-run effect annotated as "+3.1 pp employment (90% CI 1.9, 4.3])," and a companion row translating it into cost per additional job with its band — the number a policymaker takes away. No asterisks; SEs in parentheses throughout.
【Headline exhibit】figure/table + the policy magnitude it carries
【Significance reporting】SEs/CIs, no asterisks? [Y/N]
【Self-contained notes】sample/units/estimator/clustering/N/dep-mean present? [Y/N]
【Policy-units translation】coefficient → cost-per-X / MVPF / incidence shown? [Y/N]
【Next step】aejpol-writing-style| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,516 | 11,227 | -28% | 1 | 1 | 0% | 2,632 | 3,571 | +36% | 0 | 0 | — |
case-02 | fail→pass | 17,317 | 9,454 | -45% | 1 | 1 | 0% | 3,215 | 3,184 | -1% | 0 | 0 | — |
case-03 | fail→pass | 14,021 | 15,844 | +13% | 1 | 1 | 0% | 2,366 | 4,452 | +88% | 0 | 0 | — |
case-04 | pass→pass | 13,416 | 8,709 | -35% | 1 | 1 | 0% | 2,231 | 2,893 | +30% | 0 | 0 | — |
case-05 | fail→pass | 14,728 | 10,654 | -28% | 1 | 1 | 0% | 2,273 | 2,966 | +30% | 0 | 0 | — |
case-06 | fail→pass | 12,289 | 7,339 | -40% | 1 | 1 | 0% | 2,016 | 2,586 | +28% | 0 | 0 | — |
case-07 | pass→pass | 12,057 | 7,791 | -35% | 1 | 1 | 0% | 2,096 | 2,665 | +27% | 0 | 0 | — |
case-08 | fail→pass | 14,309 | 10,724 | -25% | 1 | 1 | 0% | 2,390 | 2,987 | +25% | 0 | 0 | — |
case-09 | fail→pass | 12,324 | 9,866 | -20% | 1 | 1 | 0% | 2,220 | 3,139 | +41% | 0 | 0 | — |
case-10 | fail→pass | 9,668 | 4,927 | -49% | 1 | 1 | 0% | 1,727 | 2,262 | +31% | 0 | 0 | — |
case-11 | pass→pass | 11,346 | 5,255 | -54% | 1 | 1 | 0% | 1,857 | 2,328 | +25% | 0 | 0 | — |
case-12 | pass→pass | 12,321 | 6,082 | -51% | 1 | 1 | 0% | 2,084 | 2,390 | +15% | 0 | 0 | — |
case-13 | pass→pass | 17,200 | 13,376 | -22% | 1 | 1 | 0% | 2,525 | 3,453 | +37% | 0 | 0 | — |
case-14 | fail→pass | 15,203 | 9,554 | -37% | 1 | 1 | 0% | 2,330 | 2,852 | +22% | 0 | 0 | — |
case-15 | pass→pass | 18,542 | 13,497 | -27% | 1 | 1 | 0% | 2,842 | 3,557 | +25% | 0 | 0 | — |
case-16 | pass→pass | 11,376 | 8,145 | -28% | 1 | 1 | 0% | 1,725 | 2,855 | +66% | 0 | 0 | — |
case-17 | fail→pass | 7,415 | 11,168 | +51% | 1 | 1 | 0% | 1,460 | 3,619 | +148% | 0 | 0 | — |
case-18 | pass→pass | 14,274 | 9,511 | -33% | 1 | 1 | 0% | 2,717 | 3,093 | +14% | 0 | 0 | — |
case-19 | pass→pass | 19,322 | 12,729 | -34% | 1 | 1 | 0% | 2,747 | 3,231 | +18% | 0 | 0 | — |
case-20 | pass→pass | 17,353 | 17,262 | -1% | 1 | 1 | 0% | 2,491 | 3,968 | +59% | 0 | 0 | — |
case-21 | pass→pass | 20,640 | 16,450 | -20% | 1 | 1 | 0% | 3,401 | 3,964 | +17% | 0 | 0 | — |
case-22 | pass→fail | 18,940 | 14,524 | -23% | 1 | 1 | 0% | 3,095 | 3,875 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.