Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the empirical identification strategy is the bottleneck for a 《中国农村经济》 manuscript — micro-household / village-level quasi-experimental designs (DID, IV, RDD, PSM). Stress-tests the design and the rural-specific endogeneity before drafting tables.
.claude/skills/brycewang-stanford-cre-identification/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 9% | 0% |
《中国农村经济》编委对农村微观研究的偏好排序(强 → 弱):
农村微观研究最常见的内生性来源,审稿人必查:
针对性策略至少给出一条:固定效应 + IV / 准实验冲击 / 匹配 + 安慰剂 / 双稳健(Heckman 仅作辅助,不能单独立住识别)。
把设计跑出来并审计,而不是只做描述。完整映射见 execution-with-mcp。《中国农村经济》是三农实证刊,政策评估与微观面板为主;突出识别与选择性偏误处理。
detect_design → recommend → 用 as_handle=true 拟合 → audit_result 列出尚欠的检查。callaway_santanna / sun_abraham + bacon_decomposition +honest_did_from_result);IV(effective_f_test + anderson_rubin_ci);RDD(rdrobust + mccrary_test)。
romano_wolf 做多结果族错误率控制。oster_delta / sensemakr。正文报告经济量级,完整 battery 进附录;每个数字都能复现。端到端真跑示例见 JF 执行 walkthrough。若 StatsPAI/Stata 未连接,改用 resources/code/ 并标注未验证数字。
【识别策略】DID / IV / RDD / PSM-DID / 结构估计 / 其他
【自选择处理】方式:[...]
【已完成检验】[平行趋势, 安慰剂, 弱工具, 匹配平衡, ...]
【缺失检验】[...]
【聚类层次】农户 / 村 / 县 / ...
【下一步】cre-mechanism| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | fail→pass | 23,270 | 21,333 | -8% | 1 | 1 | 0% | 2,384 | 3,618 | +52% | 0 | 0 | — |
case-18 | pass→pass | 25,014 | 17,021 | -32% | 1 | 1 | 0% | 2,770 | 3,517 | +27% | 0 | 0 | — |
case-01 | fail→pass | 32,878 | 26,152 | -20% | 1 | 1 | 0% | 3,926 | 4,764 | +21% | 0 | 0 | — |
case-02 | fail→pass | 28,686 | 26,402 | -8% | 1 | 1 | 0% | 3,661 | 4,716 | +29% | 0 | 0 | — |
case-03 | fail→pass | 33,924 | 21,692 | -36% | 1 | 1 | 0% | 3,957 | 4,960 | +25% | 0 | 0 | — |
case-04 | fail→pass | 35,386 | 27,769 | -22% | 1 | 1 | 0% | 4,267 | 4,665 | +9% | 0 | 0 | — |
case-05 | fail→pass | 25,102 | 42,760 | +70% | 1 | 1 | 0% | 3,154 | 4,948 | +57% | 0 | 0 | — |
case-06 | fail→fail | 29,264 | 28,159 | -4% | 1 | 1 | 0% | 3,171 | 5,055 | +59% | 0 | 0 | — |
case-07 | fail→pass | 34,937 | 20,125 | -42% | 1 | 1 | 0% | 2,829 | 4,145 | +47% | 0 | 0 | — |
case-08 | fail→fail | 26,273 | 27,595 | +5% | 1 | 1 | 0% | 2,522 | 4,394 | +74% | 0 | 0 | — |
case-09 | fail→pass | 21,804 | 27,994 | +28% | 1 | 1 | 0% | 2,627 | 4,711 | +79% | 0 | 0 | — |
case-10 | fail→pass | 32,805 | 26,397 | -20% | 1 | 1 | 0% | 3,470 | 4,746 | +37% | 0 | 0 | — |
case-11 | fail→pass | 30,021 | 35,099 | +17% | 1 | 1 | 0% | 3,741 | 5,743 | +54% | 0 | 0 | — |
case-12 | fail→pass | 22,535 | 16,561 | -27% | 1 | 1 | 0% | 2,346 | 3,417 | +46% | 0 | 0 | — |
case-14 | fail→pass | 25,484 | 18,906 | -26% | 1 | 1 | 0% | 3,111 | 4,465 | +44% | 0 | 0 | — |
case-15 | fail→pass | 27,350 | 33,180 | +21% | 1 | 1 | 0% | 4,123 | 5,214 | +26% | 0 | 0 | — |
case-16 | fail→pass | 25,426 | 19,548 | -23% | 1 | 1 | 0% | 2,772 | 4,503 | +62% | 0 | 0 | — |
case-17 | fail→pass | 30,693 | 23,640 | -23% | 1 | 1 | 0% | 3,646 | 4,046 | +11% | 0 | 0 | — |
case-19 | pass→fail | 55,079 | 47,517 | -14% | 1 | 1 | 0% | 7,603 | 9,856 | +30% | 0 | 0 | — |
case-20 | pass→fail | 32,529 | 40,054 | +23% | 1 | 1 | 0% | 4,357 | 7,038 | +62% | 0 | 0 | — |
case-21 | pass→pass | 13,963 | 15,535 | +11% | 1 | 1 | 0% | 1,921 | 3,055 | +59% | 0 | 0 | — |
case-22 | pass→pass | 24,342 | 24,698 | +1% | 1 | 1 | 0% | 2,613 | 5,010 | +92% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.