Install any skill in seconds. Free to start, no credit card required.
Get Started Free →このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します。
.claude/skills/affaan-m-benchmark/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -48% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 6% | 0% |
通过浏览器 MCP 测量真实浏览器指标:
1. 导航至每个目标 URL
2. 测量核心网页指标:
- LCP(最大内容绘制)— 目标 < 2.5 秒
- CLS(累积布局偏移)— 目标 < 0.1
- INP(与下一次绘制的交互)— 目标 < 200 毫秒
- FCP(首次内容绘制)— 目标 < 1.8 秒
- TTFB(首字节时间)— 目标 < 800 毫秒
3. 测量资源大小:
- 页面总重量(目标 < 1MB)
- JS 包大小(目标 < 200KB gzip 压缩后)
- CSS 大小
- 图片重量
- 第三方脚本重量
4. 统计网络请求数量
5. 检查阻塞渲染的资源对 API 端点进行基准测试:
1. 每个端点请求 100 次
2. 测量:p50、p95、p99 延迟
3. 追踪:响应大小、状态码
4. 负载测试:10 个并发请求
5. 与 SLA 目标进行对比测量开发反馈循环效率:
1. 冷构建时间
2. 热重载时间 (HMR)
3. 测试套件执行时间
4. TypeScript 检查时间
5. 代码检查时间
6. Docker 构建时间在变更前后运行以测量影响:
/benchmark baseline # 保存当前指标
# ... 进行更改 ...
/benchmark compare # 与基线进行比较输出结果:
| Metric | Before | After | Delta | Verdict |
|--------|--------|-------|-------|---------|
| LCP | 1.2s | 1.4s | +200ms | WARNING: WARN |
| Bundle | 180KB | 175KB | -5KB | ✓ BETTER |
| Build | 12s | 14s | +2s | WARNING: WARN |将基线数据以 JSON 格式存储在 .ecc/benchmarks/ 中。通过 Git 追踪,便于团队共享基线。
/benchmark compare/canary-watch 进行部署后监控/browser-qa 完成发布前完整检查清单| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 18,449 | 14,459 | -22% | 1 | 1 | 0% | 3,074 | 2,917 | -5% | 0 | 0 | — |
case-02 | fail→pass | 13,513 | 13,191 | -2% | 1 | 1 | 0% | 2,388 | 3,038 | +27% | 0 | 0 | — |
case-01 | fail→pass | 17,015 | 13,867 | -19% | 1 | 1 | 0% | 2,668 | 3,095 | +16% | 0 | 0 | — |
case-03 | fail→fail | 20,612 | 18,830 | -9% | 1 | 1 | 0% | 3,576 | 4,247 | +19% | 0 | 0 | — |
case-05 | pass→pass | 11,356 | 7,044 | -38% | 1 | 1 | 0% | 2,113 | 1,860 | -12% | 0 | 0 | — |
case-06 | fail→pass | 23,307 | 2,365 | -90% | 1 | 1 | 0% | 1,068 | 1,013 | -5% | 0 | 0 | — |
case-07 | fail→pass | 18,120 | 5,394 | -70% | 1 | 1 | 0% | 3,213 | 1,685 | -48% | 0 | 0 | — |
case-08 | pass→pass | 19,200 | 14,804 | -23% | 1 | 1 | 0% | 3,032 | 3,396 | +12% | 0 | 0 | — |
case-09 | fail→pass | 16,397 | 11,352 | -31% | 1 | 1 | 0% | 2,772 | 2,941 | +6% | 0 | 0 | — |
case-10 | pass→pass | 40,178 | 16,754 | -58% | 1 | 1 | 0% | 3,076 | 3,265 | +6% | 0 | 0 | — |
case-11 | pass→pass | 28,501 | 16,624 | -42% | 1 | 1 | 0% | 2,885 | 3,248 | +13% | 0 | 0 | — |
case-12 | fail→pass | 16,188 | 3,531 | -78% | 1 | 1 | 0% | 2,600 | 1,228 | -53% | 0 | 0 | — |
case-13 | fail→pass | 27,612 | 16,779 | -39% | 1 | 1 | 0% | 4,314 | 3,283 | -24% | 0 | 0 | — |
case-14 | fail→pass | 21,017 | 29,196 | +39% | 1 | 1 | 0% | 3,094 | 3,217 | +4% | 0 | 0 | — |
case-15 | pass→pass | 19,150 | 10,485 | -45% | 1 | 1 | 0% | 2,950 | 2,298 | -22% | 0 | 0 | — |
case-16 | fail→pass | 10,610 | 2,600 | -75% | 1 | 1 | 0% | 1,368 | 1,138 | -17% | 0 | 0 | — |
case-17 | fail→pass | 9,462 | 3,208 | -66% | 1 | 1 | 0% | 1,553 | 1,193 | -23% | 0 | 0 | — |
case-18 | fail→pass | 17,150 | 17,344 | +1% | 1 | 1 | 0% | 2,966 | 3,594 | +21% | 0 | 0 | — |
case-19 | pass→pass | 21,960 | 17,054 | -22% | 1 | 1 | 0% | 3,217 | 3,500 | +9% | 0 | 0 | — |
case-20 | pass→fail | 16,154 | 12,500 | -23% | 1 | 1 | 0% | 3,168 | 3,146 | -1% | 0 | 0 | — |
case-21 | pass→pass | 10,432 | 9,033 | -13% | 1 | 1 | 0% | 1,959 | 2,681 | +37% | 0 | 0 | — |
case-22 | pass→pass | 7,135 | 5,329 | -25% | 1 | 1 | 0% | 1,175 | 1,472 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.