Install any skill in seconds. Free to start, no credit card required.
Get Started Free →このスキルを使用して、機能をデプロイ後にブラウザ自動化を使用した自動ビジュアルテストとUI相互作用検証を自動化します。
.claude/skills/affaan-m-browser-qa/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-21 | ✗→✓ | ▲ Improved | -2% | 0% |
使用浏览器自动化 MCP(claude-in-chrome、Playwright 或 Puppeteer),像真实用户一样与线上页面交互。
1. 打开目标 URL
2. 检查控制台错误(过滤噪声:分析脚本、第三方库)
3. 验证网络请求中没有 4xx / 5xx
4. 在桌面和移动端视口截图首屏内容
5. 检查 Core Web Vitals:LCP < 2.5s,CLS < 0.1,INP < 200ms1. 点击所有导航链接,验证没有死链
2. 使用有效数据提交表单,验证成功态
3. 使用无效数据提交表单,验证错误态
4. 测试认证流程:登录 → 受保护页面 → 登出
5. 测试关键用户路径(结账、引导、搜索)1. 在 3 个断点(375px、768px、1440px)对关键页面截图
2. 与基线截图对比(如果已保存)
3. 标记 > 5px 的布局偏移、缺失元素、内容溢出
4. 如适用,检查暗色模式1. 在每个页面运行 axe-core 或等价工具
2. 标记 WCAG AA 违规(对比度、标签、焦点顺序)
3. 验证键盘导航可以端到端工作
4. 检查屏幕阅读器地标markdown## QA 报告 — [URL] — [timestamp] ### 冒烟测试 - 控制台错误:0 个严重错误,2 个警告(分析脚本噪声) - 网络:全部 200/304,无失败请求 - Core Web Vitals:LCP 1.2s,CLS 0.02,INP 89ms ### 交互 - [done] 导航链接:12/12 正常 - [issue] 联系表单:无效邮箱缺少错误态 - [done] 认证流程:登录 / 登出正常 ### 视觉 - [issue] Hero 区域在 375px 视口下溢出 - [done] 暗色模式:所有页面一致 ### 可访问性 - 2 个 AA 级违规:Hero 图片缺少 alt 文本,页脚链接对比度过低 ### 结论:修复后可发布(2 个问题,0 个阻塞项)
可与任意浏览器 MCP 配合:
mChild__claude-in-chrome__* 工具(推荐,直接使用你的真实 Chrome)mcp__browserbase__* 使用 Playwright可与 /canary-watch 搭配用于发布后的持续监控。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 22,001 | 13,148 | -40% | 1 | 1 | 0% | 3,930 | 2,991 | -24% | 0 | 0 | — |
case-02 | fail→fail | 19,808 | 14,605 | -26% | 1 | 1 | 0% | 3,392 | 2,999 | -12% | 0 | 0 | — |
case-03 | pass→pass | 20,059 | 19,558 | -2% | 1 | 1 | 0% | 4,141 | 5,421 | +31% | 0 | 0 | — |
case-04 | pass→pass | 9,387 | 10,755 | +15% | 1 | 1 | 0% | 1,986 | 2,875 | +45% | 0 | 0 | — |
case-05 | pass→pass | 20,363 | 18,049 | -11% | 1 | 1 | 0% | 3,709 | 4,387 | +18% | 0 | 0 | — |
case-06 | pass→pass | 18,707 | 14,756 | -21% | 1 | 1 | 0% | 3,084 | 3,393 | +10% | 0 | 0 | — |
case-07 | pass→pass | 18,410 | 11,930 | -35% | 1 | 1 | 0% | 2,916 | 2,742 | -6% | 0 | 0 | — |
case-08 | fail→pass | 17,214 | 18,078 | +5% | 1 | 1 | 0% | 2,523 | 2,776 | +10% | 0 | 0 | — |
case-09 | fail→pass | 15,616 | 8,054 | -48% | 1 | 1 | 0% | 2,206 | 2,124 | -4% | 0 | 0 | — |
case-14 | pass→pass | 17,180 | 17,426 | +1% | 1 | 1 | 0% | 2,787 | 3,517 | +26% | 0 | 0 | — |
case-10 | pass→pass | 18,491 | 18,174 | -2% | 1 | 1 | 0% | 2,618 | 3,221 | +23% | 0 | 0 | — |
case-11 | pass→pass | 15,845 | 16,268 | +3% | 1 | 1 | 0% | 2,539 | 3,325 | +31% | 0 | 0 | — |
case-12 | fail→pass | 18,344 | 15,087 | -18% | 1 | 1 | 0% | 2,684 | 3,135 | +17% | 0 | 0 | — |
case-13 | pass→pass | 14,758 | 13,837 | -6% | 1 | 1 | 0% | 2,067 | 2,625 | +27% | 0 | 0 | — |
case-15 | pass→pass | 20,088 | 17,926 | -11% | 1 | 1 | 0% | 2,977 | 3,729 | +25% | 0 | 0 | — |
case-16 | pass→pass | 11,256 | 3,662 | -67% | 1 | 1 | 0% | 1,773 | 1,277 | -28% | 0 | 0 | — |
case-17 | pass→pass | 13,541 | 3,742 | -72% | 1 | 1 | 0% | 2,224 | 1,361 | -39% | 0 | 0 | — |
case-18 | fail→fail | 12,742 | 10,600 | -17% | 1 | 1 | 0% | 2,345 | 2,694 | +15% | 0 | 0 | — |
case-19 | pass→pass | 20,220 | 16,621 | -18% | 1 | 1 | 0% | 3,032 | 3,171 | +5% | 0 | 0 | — |
case-20 | pass→pass | 16,368 | 10,244 | -37% | 1 | 1 | 0% | 2,523 | 2,292 | -9% | 0 | 0 | — |
case-21 | fail→pass | 21,989 | 15,431 | -30% | 1 | 1 | 0% | 3,295 | 3,215 | -2% | 0 | 0 | — |
case-22 | fail→pass | 13,232 | 5,241 | -60% | 1 | 1 | 0% | 1,980 | 1,539 | -22% | 0 | 0 | — |
case-23 | pass→pass | 17,065 | 12,712 | -26% | 1 | 1 | 0% | 2,620 | 2,737 | +4% | 0 | 0 | — |
case-24 | fail→pass | 15,952 | 8,790 | -45% | 1 | 1 | 0% | 2,335 | 2,107 | -10% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +29 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.