Install any skill in seconds. Free to start, no credit card required.
Get Started Free →A股模型风险/回测过拟合分析。当用户说"模型风险"、"model risk"、"过拟合"、"overfitting"、"回测失真"、"样本外失效"、"模型验证"时触发。基于 cn-stock-data 获取数据,评估量化模型的过拟合与模型风险。支持 formal/brief 两种输出风格。
.claude/skills/aifinlab-a-share-model-risk/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 15% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 13% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 17% | 0% |
| case-07 | ✓→✓ | = Same ✓ | 23% | 0% |
通过 cn-stock-data skill 获取数据:
# 模型风险评估报告
## 一、过拟合检测
| 指标 | 训练集 | 测试集 | 差距 |
|------|--------|--------|------|
## 二、回测陷阱
[各类偏差检查结果]
## 三、验证结果
[样本外/CV/排列检验]
## 四、风险管理建议## 模型风险速览
- 训练Sharpe 3.2 vs 测试 1.8,差距44%
- 过拟合风险:中等
- 未发现前视偏差
- 建议:简化模型,减少参数数量参考 references/model-risk-guide.md 获取详细方法论与 A股实证研究。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 36,325 | 35,608 | -2% | 1 | 1 | 0% | 5,418 | 5,606 | +3% | 0 | 0 | — |
case-02 | fail→fail | 21,720 | 12,824 | -41% | 1 | 1 | 0% | 2,795 | 2,158 | -23% | 0 | 0 | — |
case-03 | fail→fail | 35,333 | 43,153 | +22% | 1 | 1 | 0% | 4,755 | 6,371 | +34% | 0 | 0 | — |
case-04 | pass→pass | 18,426 | 20,253 | +10% | 1 | 1 | 0% | 2,796 | 3,217 | +15% | 0 | 0 | — |
case-05 | pass→pass | 19,835 | 20,679 | +4% | 1 | 1 | 0% | 2,814 | 3,186 | +13% | 0 | 0 | — |
case-06 | pass→pass | 22,784 | 34,404 | +51% | 1 | 1 | 0% | 3,310 | 3,871 | +17% | 0 | 0 | — |
case-07 | pass→pass | 24,502 | 23,482 | -4% | 1 | 1 | 0% | 3,045 | 3,759 | +23% | 0 | 0 | — |
case-08 | pass→pass | 21,189 | 43,847 | +107% | 1 | 1 | 0% | 3,285 | 4,141 | +26% | 0 | 0 | — |
case-09 | pass→pass | 20,628 | 23,289 | +13% | 1 | 1 | 0% | 2,600 | 3,196 | +23% | 0 | 0 | — |
case-10 | pass→pass | 12,709 | 14,849 | +17% | 1 | 1 | 0% | 1,989 | 2,850 | +43% | 0 | 0 | — |
case-11 | pass→pass | 7,346 | 10,892 | +48% | 1 | 1 | 0% | 944 | 2,176 | +131% | 0 | 0 | — |
case-12 | pass→pass | 22,537 | 23,249 | +3% | 1 | 1 | 0% | 3,246 | 4,075 | +26% | 0 | 0 | — |
case-13 | pass→pass | 26,657 | 30,209 | +13% | 1 | 1 | 0% | 3,249 | 4,215 | +30% | 0 | 0 | — |
case-14 | pass→pass | 25,530 | 24,601 | -4% | 1 | 1 | 0% | 3,049 | 3,769 | +24% | 0 | 0 | — |
case-15 | pass→pass | 22,119 | 27,056 | +22% | 1 | 1 | 0% | 3,052 | 3,937 | +29% | 0 | 0 | — |
case-16 | pass→pass | 15,958 | 18,142 | +14% | 1 | 1 | 0% | 2,730 | 3,129 | +15% | 0 | 0 | — |
case-17 | fail→pass | 25,117 | 10,661 | -58% | 1 | 1 | 0% | 3,094 | 2,289 | -26% | 0 | 0 | — |
case-18 | pass→pass | 16,115 | 20,322 | +26% | 1 | 1 | 0% | 2,267 | 3,170 | +40% | 0 | 0 | — |
case-19 | pass→pass | 17,446 | 19,556 | +12% | 1 | 1 | 0% | 2,626 | 3,539 | +35% | 0 | 0 | — |
case-20 | pass→pass | 34,372 | 29,374 | -15% | 1 | 1 | 0% | 5,364 | 5,618 | +5% | 0 | 0 | — |
case-21 | pass→pass | 40,284 | 29,720 | -26% | 1 | 1 | 0% | 5,309 | 4,997 | -6% | 0 | 0 | — |
case-22 | pass→pass | 12,489 | 10,516 | -16% | 1 | 1 | 0% | 2,261 | 2,241 | -1% | 0 | 0 | — |
case-23 | pass→pass | 8,699 | 9,176 | +5% | 1 | 1 | 0% | 1,334 | 1,965 | +47% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +4 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.