Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing or writing heterogeneity analysis for a 《中国农村经济》 manuscript. Enforces rural-relevant cut dimensions and theoretical-justification discipline.
.claude/skills/brycewang-stanford-cre-heterogeneity/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 19% | 0% |
按《中国农村经济》读者期待度排序:
至少切 3 个维度,每个维度至少 2 个子样本对比。避免只切"东中西"——这是审稿人最容易吐槽的"偷懒维度"。
本文进一步检验[维度,落到农户 / 区域 / 村庄]异质性。
理论上,[原因——为什么这类农户 / 地区效应不同]……
实证上,本文将样本按[维度]分为[组 1] 和 [组 2],分别估计主回归(结果见表 X 第 (1)(2) 列)。
结果显示,[组 1] 的处理效应为 X,而 [组 2] 为 Y,系数差异在 [显著水平] 上显著。
这一发现与本文的[机制]一致,说明该政策 / 行为对[何种农户]的作用更强。异质性 ≈ 机制的反向验证:
理想的实证文章应该:机制(M 强 → 效应强) 与 异质性(M 强的农户子样本 → 效应强) 互相印证。例如机制是"缓解信贷约束",则异质性上"原本信贷约束更紧的农户"效应应更强。
【异质性维度】X 个(是否含"东中西"以外的三农维度)
【系数差异检验】是 / 否
【与机制一致性】是 / 否
【最小子样本量】X
【下一步】cre-tables-figures先锁定农村问题、政策/制度场景、识别链条、机制证据和可执行含义,再判断稿件是否回应农村经济审稿人通常同时追问“三农”问题意识、政策场景、识别可信度和农村制度机制。
resources/official-source-map.md,列出仍可能改变建议的一个未核实事实。| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 21,099 | 29,836 | +41% | 1 | 1 | 0% | 3,165 | 5,248 | +66% | 0 | 0 | — |
case-02 | fail→pass | 32,558 | 32,802 | +1% | 1 | 1 | 0% | 4,151 | 5,719 | +38% | 0 | 0 | — |
case-03 | fail→pass | 29,600 | 28,377 | -4% | 1 | 1 | 0% | 4,472 | 5,162 | +15% | 0 | 0 | — |
case-04 | pass→fail | 27,887 | 35,407 | +27% | 1 | 1 | 0% | 4,046 | 5,962 | +47% | 0 | 0 | — |
case-05 | pass→fail | 24,584 | 27,416 | +12% | 1 | 1 | 0% | 3,259 | 5,403 | +66% | 0 | 0 | — |
case-06 | pass→fail | 18,694 | 21,292 | +14% | 1 | 1 | 0% | 2,489 | 3,679 | +48% | 0 | 0 | — |
case-07 | pass→pass | 24,619 | 25,976 | +6% | 1 | 1 | 0% | 3,095 | 4,535 | +47% | 0 | 0 | — |
case-08 | fail→pass | 44,550 | 25,535 | -43% | 1 | 1 | 0% | 2,922 | 4,638 | +59% | 0 | 0 | — |
case-09 | pass→pass | 22,276 | 21,118 | -5% | 1 | 1 | 0% | 2,737 | 3,781 | +38% | 0 | 0 | — |
case-10 | pass→pass | 28,155 | 22,453 | -20% | 1 | 1 | 0% | 3,130 | 4,764 | +52% | 0 | 0 | — |
case-11 | pass→pass | 20,727 | 22,236 | +7% | 1 | 1 | 0% | 2,349 | 3,756 | +60% | 0 | 0 | — |
case-12 | fail→fail | 19,124 | 28,487 | +49% | 1 | 1 | 0% | 2,938 | 4,243 | +44% | 0 | 0 | — |
case-13 | pass→pass | 27,113 | 30,793 | +14% | 1 | 1 | 0% | 3,039 | 4,644 | +53% | 0 | 0 | — |
case-14 | pass→pass | 27,540 | 27,309 | -1% | 1 | 1 | 0% | 2,916 | 4,394 | +51% | 0 | 0 | — |
case-15 | pass→pass | 25,876 | 32,730 | +26% | 1 | 1 | 0% | 3,348 | 5,053 | +51% | 0 | 0 | — |
case-16 | pass→pass | 25,610 | 25,655 | +0% | 1 | 1 | 0% | 3,318 | 5,162 | +56% | 0 | 0 | — |
case-17 | fail→pass | 19,036 | 12,597 | -34% | 1 | 1 | 0% | 2,544 | 3,021 | +19% | 0 | 0 | — |
case-18 | fail→pass | 18,410 | 5,450 | -70% | 1 | 1 | 0% | 2,617 | 2,065 | -21% | 0 | 0 | — |
case-19 | fail→pass | 23,054 | 28,133 | +22% | 1 | 1 | 0% | 3,344 | 4,726 | +41% | 0 | 0 | — |
case-20 | fail→pass | 13,834 | 2,480 | -82% | 1 | 1 | 0% | 2,082 | 1,579 | -24% | 0 | 0 | — |
case-21 | pass→pass | 20,553 | 27,330 | +33% | 1 | 1 | 0% | 2,787 | 4,942 | +77% | 0 | 0 | — |
case-22 | pass→pass | 19,141 | 26,115 | +36% | 1 | 1 | 0% | 2,774 | 4,819 | +74% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.