Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when writing or structuring a game playtest report: follow the fixed nine-section schema with these exact field names and vocabularies.
.claude/skills/playtest-report/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 2 |
| Model | Lift | Δ tokens | Δ turns | Cases | Verified |
|---|---|---|---|---|---|
| gemini-3.5-flashbest | +67% | — | 0% | 24 | 87d ago |
| gemini-3.6-flash | +59% | +166% | 0% | 22 | 54d ago |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | — | — |
| case-14 | ✗→✓ | ▲ Improved | — | — |
| case-08 | ✗→✓ | ▲ Improved | — | — |
| case-03 | ✗→✓ | ▲ Improved | — | — |
| case-13 | ✗→✓ | ▲ Improved | — | — |
Enforces one fixed structure for every game playtest report: nine top-level ## sections in a fixed order, exact field labels, and a closed vocabulary for the graded fields. Apply whenever asked to produce a playtest report, a blank template, or to turn raw playtest notes into a structured writeup.
Every report has these nine ## sections and no others, in order. Never rename, merge, drop, reorder, or insert extra top-level sections.
## Session Info## Test Focus## First Impressions (First 5 minutes)## Gameplay Flow## Bugs Encountered## Feature-Specific Feedback## Quantitative Data (if available)## Overall Assessment## Top 3 Priorities from this sessionKeep the parentheticals verbatim: (First 5 minutes) on section 3 and (if available) on section 7.
A bullet list with EXACTLY these seven bold labels, in this order: Date, Build, Duration, Tester, Platform, Input Method, Session Type. The label for how the player controlled the game is always Input Method — never "Controller", never "Control Scheme", never "Input". The person label is Tester — never "Tester Name", never "Player".
One line stating what was being tested (a feature, a flow, or "general / first pass"). Comes second, directly after Session Info.
Bullets with these labels: Understood the goal?, Understood the controls?, Emotional response, Notes. Keep the (First 5 minutes) qualifier in the heading. The two "Understood…?" answers are one of Yes / No / Partially.
Exactly four ### subsections, in this order: What worked well, Pain points, Confusion points, Moments of delight. Each Pain point bullet ends with the suffix -- Severity: High/Medium/Low. Do not rename "Moments of delight" to "Highlights"/"Positives", and do not collapse the four into "Pros/Cons".
A Markdown table with EXACTLY these four columns, in this order: # | Description | Severity | Reproducible. The Reproducible cell is Yes or No (never "Always/Sometimes/Never", never a percentage). Do not add Status, Priority, Steps, or Owner columns.
One block per feature, each with bullets Understood purpose?, Found engaging?, Suggestions.
Bullets with labels Deaths, Time per area, Items used, Features discovered vs missed. Keep the section even when empty (write "n/a" per line); keep the (if available) qualifier.
Bullets: Would play again? (Yes / No / Maybe), Difficulty (Too Easy / Just Right / Too Hard), Pacing (Too Slow / Good / Too Fast), Session length preference (Shorter / Good / Longer). Each answer is one of its listed tokens, verbatim.
A numbered list of EXACTLY three items, most important first. Not two, not five, not "Top Priorities" — always exactly three, and always the final section.
Use these tokens exactly; no synonyms anywhere in the report:
When sorting findings for follow-up, use EXACTLY these four buckets, no others: Design changes needed (fun issues, player confusion, broken mechanics), Balance adjustments (numbers feel wrong, difficulty spiked or flat), Bug reports (reproducible implementation defects, crashes), Polish items (non-blocking friction/feel for later). Not "Must-fix / Nice-to-have", not "P0/P1/P2".
BEFORE (base default invents its own label):
## Session Info
- Date: 2026-06-26
- Controller: Gamepad
- Player Name: PriyaAFTER (exact labels + order):
## Session Info
- **Date**: 2026-06-26
- **Build**: v0.4.1
- **Duration**: 40 min
- **Tester**: Priya
- **Platform**: Steam Deck
- **Input Method**: Gamepad
- **Session Type**: First timeBEFORE (base merges into Pros/Cons, no severity):
## Gameplay Flow
### Pros
- Combat feels great
### Cons
- Double jump is unclearAFTER (four fixed subsections, severity suffix on pain points):
## Gameplay Flow
### What worked well
- Combat feels responsive and weighty
### Pain points
- Double-jump timing window is too tight -- Severity: Medium
### Confusion points
- Players didn't realize the crafting menu existed
### Moments of delight
- The boss-fight music drew an audible reactionBEFORE (base picks its own columns + reproducibility words):
## Bugs
| Bug | Steps | Priority | Status |
|-----|-------|----------|--------|
| Crash on pause | Pause in level 3 | P0 | Open |AFTER (exact four columns, Yes/No, High/Medium/Low):
## Bugs Encountered
| # | Description | Severity | Reproducible |
|---|-------------|----------|--------------|
| 1 | Game crashes when pausing in level 3 | High | Yes |BEFORE (free-form): - Difficulty: felt a bit on the easy side AFTER (closed token): - **Difficulty**: Too Easy
BEFORE: - Emotional response: kind of lost at first, then into it AFTER: - **Emotional response**: Confused
BEFORE (base front-loads bugs, drops qualifiers):
## Bugs
## Overview
## First ImpressionsAFTER: the nine sections in the R1 order, with Session Info first, Test Focus second, and Top 3 Priorities from this session last.
BEFORE (base lists five "key takeaways"):
## Key Takeaways
1. … 2. … 3. … 4. … 5. …AFTER:
## Top 3 Priorities from this session
1. Fix the level-3 pause crash
2. Clarify the double-jump tutorial
3. Add a difficulty bump to the early gameBEFORE: Must-fix: / Nice-to-have: AFTER: Design changes needed: / Balance adjustments: / Bug reports: / Polish items:
field labels; fill values with placeholders like [Date], [Yes/No/Partially], and leave the bug table header row with an empty body.
"n/a" or "not observed" rather than deleting the section.
pick "No" when not confirmed reproducible; do not invent "Unknown".
Low, "major"/"blocker"/"crash" → High, in-between → Medium.
Confusion points; the reproducible defect ALSO gets a row in Bugs Encountered. They are not mutually exclusive.
three bullets) once per feature under the single section heading.
Test Focus names the system; the schema is otherwise unchanged.
(First 5 minutes) and (if available) in the headings; never strip them.-- Severity: …; never leave a pain point unrated.# | Description | Severity | Reproducible; never add Status/Priority/Steps.(First 5 minutes) / (if available) qualifiers.-- Severity: High/Medium/Low.# | Description | Severity | Reproducible; Reproducible = Yes/No.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.5-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.5-flash | verified | 6/26/2026 | +67% |
Other measured skills in the registry, with their headline benchmark lift.