Install any skill in seconds. Free to start, no credit card required.
Get Started Free →このスキルを使用して、デプロイメント、マージ、または依存関係アップグレード後にデプロイされたURLの回帰を監視します。
.claude/skills/affaan-m-canary-watch/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 121% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 20% | 0% |
监控已部署 URL 是否存在回归问题。循环运行,直至手动停止或监控窗口过期。
1. HTTP 状态 — 页面是否返回 200?
2. 控制台错误 — 是否出现之前没有的新错误?
3. 网络故障 — 是否存在失败的 API 调用、5xx 响应?
4. 性能 — LCP/CLS/INP 与基线相比是否有退化?
5. 内容 — 关键元素是否消失?(h1、导航、页脚、CTA)
6. API 健康 — 关键端点是否在 SLA 内响应?快速检查(默认):单次执行,报告结果
/canary-watch https://myapp.com持续监控:每 N 分钟检查一次,持续 M 小时
/canary-watch https://myapp.com --interval 5m --duration 2h差异模式:对比预发布环境与生产环境
/canary-watch --compare https://staging.myapp.com https://myapp.comyamlcritical: # immediate alert - HTTP status != 200 - Console error count > 5 (new errors only) - LCP > 4s - API endpoint returns 5xx warning: # flag in report - LCP increased > 500ms from baseline - CLS > 0.1 - New console warnings - Response time > 2x baseline info: # log only - Minor performance variance - New network requests (third-party scripts added?)
当超过关键阈值时:
~/.claude/canary-watch.logmarkdown## Canary 报告 — myapp.com — 2026-03-23 03:15 PST ### 状态:健康 ✓ | 检查项 | 结果 | 基线 | 偏差 | |-------|--------|----------|-------| | HTTP | 200 ✓ | 200 | — | | 控制台错误 | 0 ✓ | 0 | — | | LCP | 1.8s ✓ | 1.6s | +200ms | | CLS | 0.01 ✓ | 0.01 | — | | API /health | 145ms ✓ | 120ms | +25ms | ### 未检测到回归问题。部署状态良好。
配合使用:
/browser-qa 进行部署前验证git push 上添加 PostToolUse 钩子,部署后自动检查| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | pass→pass | 12,865 | 2,518 | -80% | 1 | 1 | 0% | 2,180 | 1,127 | -48% | 0 | 0 | — |
case-19 | fail→pass | 8,461 | 2,908 | -66% | 1 | 1 | 0% | 1,450 | 1,286 | -11% | 0 | 0 | — |
case-20 | pass→pass | 14,719 | 13,207 | -10% | 1 | 1 | 0% | 2,497 | 3,217 | +29% | 0 | 0 | — |
case-21 | pass→pass | 9,569 | 9,370 | -2% | 1 | 1 | 0% | 1,956 | 2,717 | +39% | 0 | 0 | — |
case-01 | fail→pass | 8,010 | 12,194 | +52% | 1 | 1 | 0% | 1,512 | 3,343 | +121% | 0 | 0 | — |
case-02 | fail→fail | 8,303 | 5,128 | -38% | 1 | 1 | 0% | 1,728 | 1,676 | -3% | 0 | 0 | — |
case-03 | fail→pass | 8,241 | 10,635 | +29% | 1 | 1 | 0% | 1,732 | 2,663 | +54% | 0 | 0 | — |
case-04 | pass→pass | 10,120 | 6,639 | -34% | 1 | 1 | 0% | 1,864 | 2,020 | +8% | 0 | 0 | — |
case-05 | fail→pass | 11,670 | 2,785 | -76% | 1 | 1 | 0% | 2,018 | 1,283 | -36% | 0 | 0 | — |
case-06 | pass→pass | 5,343 | 3,279 | -39% | 1 | 1 | 0% | 846 | 1,341 | +59% | 0 | 0 | — |
case-07 | pass→pass | 10,117 | 3,788 | -63% | 1 | 1 | 0% | 1,832 | 1,475 | -19% | 0 | 0 | — |
case-08 | pass→pass | 3,295 | 13,079 | +297% | 1 | 1 | 0% | 596 | 1,534 | +157% | 0 | 0 | — |
case-09 | fail→pass | 7,564 | 4,778 | -37% | 1 | 1 | 0% | 1,431 | 1,711 | +20% | 0 | 0 | — |
case-10 | fail→pass | 6,250 | 2,396 | -62% | 1 | 1 | 0% | 1,051 | 1,192 | +13% | 0 | 0 | — |
case-11 | fail→pass | 10,270 | 1,783 | -83% | 1 | 1 | 0% | 1,890 | 977 | -48% | 0 | 0 | — |
case-12 | fail→pass | 15,737 | 13,055 | -17% | 1 | 1 | 0% | 2,688 | 3,036 | +13% | 0 | 0 | — |
case-14 | fail→pass | 6,848 | 2,463 | -64% | 1 | 1 | 0% | 1,341 | 1,177 | -12% | 0 | 0 | — |
case-15 | pass→pass | 9,379 | 3,242 | -65% | 1 | 1 | 0% | 1,686 | 1,257 | -25% | 0 | 0 | — |
case-16 | pass→pass | 7,478 | 3,736 | -50% | 1 | 1 | 0% | 1,074 | 1,346 | +25% | 0 | 0 | — |
case-17 | pass→pass | 6,248 | 3,384 | -46% | 1 | 1 | 0% | 961 | 1,294 | +35% | 0 | 0 | — |
case-18 | fail→pass | 5,975 | 2,830 | -53% | 1 | 1 | 0% | 998 | 1,182 | +18% | 0 | 0 | — |
case-22 | fail→fail | 7,563 | 3,938 | -48% | 1 | 1 | 0% | 1,517 | 1,552 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.