Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Test a playable browser game end to end with deterministic fixtures and real browser evidence. Use for gameplay QA, regression testing, controls, accessibility, responsive/mobile testing, save flows, console checks, performance smoke tests, and release verification.
.claude/skills/mengto-test-playable-web-games/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | -41% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -37% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 15% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 6% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 20% | 0% |
Combine deterministic tests with short real playthroughs. Do not treat a green build as gameplay proof.
Map the player journey: launch, new game, controls, core action, enemy encounter, reward, inventory/progression, save/continue, loss/retry, pause, settings, and completion. Add device variants for desktop keyboard/mouse, touch portrait, touch landscape, and reduced motion.
Create seedable fixtures, debug routes, or query parameters for combat phase, inventory loadout, boss phase, tutorial step, save migration, low health, and error states. Test important transitions directly rather than grinding through the campaign.
Use the repository-approved browser surface and honor local browser restrictions. Confirm screen-visible feedback after each meaningful action: control response, target state, damage/resource state, UI update, and navigation. Inspect console warnings/errors and sample performance only on a representative live scene.
Record reproduction steps, expected versus actual result, severity, device/viewport, and the shortest useful proof. Separate new regressions from pre-existing baseline issues. After testing, close idle browser pages, dev servers, and benchmarks while leaving active task resources alone.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,260 | 3,899 | -73% | 1 | 1 | 0% | 2,569 | 498 | -81% | 0 | 0 | — |
case-02 | fail→fail | 17,920 | 1,977 | -89% | 1 | 1 | 0% | 2,903 | 458 | -84% | 0 | 0 | — |
case-03 | fail→fail | 10,565 | 2,780 | -74% | 1 | 1 | 0% | 1,753 | 468 | -73% | 0 | 0 | — |
case-04 | pass→pass | 20,505 | 21,607 | +5% | 1 | 1 | 0% | 3,751 | 4,324 | +15% | 0 | 0 | — |
case-05 | pass→pass | 16,848 | 16,551 | -2% | 1 | 1 | 0% | 3,530 | 3,754 | +6% | 0 | 0 | — |
case-06 | pass→pass | 18,379 | 18,637 | +1% | 1 | 1 | 0% | 3,252 | 3,913 | +20% | 0 | 0 | — |
case-07 | pass→pass | 11,811 | 10,882 | -8% | 1 | 1 | 0% | 2,063 | 1,995 | -3% | 0 | 0 | — |
case-08 | pass→pass | 8,860 | 3,503 | -60% | 1 | 1 | 0% | 1,466 | 848 | -42% | 0 | 0 | — |
case-09 | pass→pass | 14,924 | 11,138 | -25% | 1 | 1 | 0% | 2,446 | 2,223 | -9% | 0 | 0 | — |
case-10 | pass→pass | 14,546 | 11,708 | -20% | 1 | 1 | 0% | 2,677 | 2,311 | -14% | 0 | 0 | — |
case-11 | pass→pass | 8,121 | 4,591 | -43% | 1 | 1 | 0% | 1,183 | 982 | -17% | 0 | 0 | — |
case-12 | fail→pass | 13,674 | 7,656 | -44% | 1 | 1 | 0% | 2,488 | 1,476 | -41% | 0 | 0 | — |
case-13 | fail→fail | 8,490 | 11,713 | +38% | 1 | 1 | 0% | 1,598 | 2,344 | +47% | 0 | 0 | — |
case-14 | pass→pass | 7,726 | 4,394 | -43% | 1 | 1 | 0% | 1,359 | 1,148 | -16% | 0 | 0 | — |
case-15 | fail→pass | 13,008 | 6,808 | -48% | 1 | 1 | 0% | 2,216 | 1,389 | -37% | 0 | 0 | — |
case-16 | pass→pass | 10,854 | 3,832 | -65% | 1 | 1 | 0% | 1,729 | 923 | -47% | 0 | 0 | — |
case-17 | pass→pass | 12,158 | 10,405 | -14% | 1 | 1 | 0% | 2,271 | 2,264 | -0% | 0 | 0 | — |
case-18 | pass→pass | 8,999 | 1,214 | -87% | 1 | 1 | 0% | 1,510 | 453 | -70% | 0 | 0 | — |
case-19 | pass→pass | 14,070 | 15,625 | +11% | 1 | 1 | 0% | 2,483 | 3,053 | +23% | 0 | 0 | — |
case-20 | pass→pass | 11,876 | 12,467 | +5% | 1 | 1 | 0% | 1,863 | 2,422 | +30% | 0 | 0 | — |
case-21 | pass→pass | 9,954 | 1,851 | -81% | 1 | 1 | 0% | 1,597 | 557 | -65% | 0 | 0 | — |
case-22 | pass→pass | 13,418 | 4,815 | -64% | 1 | 1 | 0% | 2,114 | 1,064 | -50% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.