Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Plan a game night that actually works for the specific people coming — the right lineup for player count, weight tolerance, and time, sequenced from icebreaker to main event, with the fallback for when someone bails. Use when someone says 'planning a game night', 'what should six of us play', 'games for my family Christmas', 'my partner hates long games', or 'we always end up arguing over what to play'. Produces a sequenced lineup with reasoning, timings, and a plan B.
.claude/skills/mohitagw15856-game-night-planner/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 30% | 0% |
Game nights don't fail because the games are bad — they fail because someone brought a 3-hour economic engine to a table that wanted snacks and laughing, or because choosing took 40 minutes, or because seven people showed up for a 4-player game. This skill plans like a good host thinks: who is actually coming, what's their tolerance, sequence light-to-heavy, and always know what happens when the count changes at 7pm.
per-game: player count fit, realistic time including the teach, and why it suits this table
1–3 to-buy/borrow suggestions only if the shelf can't cover the night
when to call the last game
Ask for (if not already provided):
something to do"
weight ceiling for the main event; the most-enthusiastic one gets the nightcap. Say who each pick serves.
people arrive. Main event: the night's centrepiece, started while energy is high — never after 9pm for a heavy game. Nightcap: low-rules, high-laughs, quittable anytime.
90; say the real number per slot and total. If the lineup exceeds the stated window, cut a game, don't compress the estimates.
each lineup slot, name the swap when N±1 shows up (many great games break at exactly one count — flag those).
fit, note "check availability/price" rather than asserting either, and if the user's shelf covers the night, say so — the best recommendation is often already owned.
## The night at a glance
[Who's coming, the vibe, total window · one-line theme of the plan]
## Lineup
1. OPENER — [game] ([count] players, ~[real minutes] incl. teach)
Why for this table: … · Who it serves: …
2. MAIN EVENT — [game] …
3. NIGHTCAP (optional) — [game] …
Total honest time: [X]h[Y]m of your [window]
## When reality edits the guest list
- One fewer: … · One extra: … · Running an hour late: cut […], start at […]
## Logistics
[Who teaches what · setup while people arrive · the "last game" call time]beneficiary is a list, not a plan
wishful compression
"check availability"
decides whether there's a next game night
state prices as fact
actually happens
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 21,610 | 54,057 | +150% | 1 | 1 | 0% | 3,711 | 4,977 | +34% | 0 | 0 | — |
case-02 | fail→fail | 20,656 | 49,544 | +140% | 1 | 1 | 0% | 3,721 | 3,945 | +6% | 0 | 0 | — |
case-03 | fail→pass | 17,523 | 23,965 | +37% | 1 | 1 | 0% | 3,036 | 4,615 | +52% | 0 | 0 | — |
case-04 | pass→pass | 15,285 | 19,122 | +25% | 1 | 1 | 0% | 2,649 | 4,124 | +56% | 0 | 0 | — |
case-05 | fail→fail | 15,507 | 21,467 | +38% | 1 | 1 | 0% | 2,717 | 4,671 | +72% | 0 | 0 | — |
case-06 | fail→fail | 15,355 | 16,357 | +7% | 1 | 1 | 0% | 2,592 | 3,722 | +44% | 0 | 0 | — |
case-07 | pass→pass | 14,657 | 16,664 | +14% | 1 | 1 | 0% | 2,305 | 3,612 | +57% | 0 | 0 | — |
case-08 | fail→pass | 18,656 | 25,148 | +35% | 1 | 1 | 0% | 3,033 | 5,144 | +70% | 0 | 0 | — |
case-09 | pass→pass | 12,875 | 16,905 | +31% | 1 | 1 | 0% | 2,230 | 3,415 | +53% | 0 | 0 | — |
case-10 | pass→pass | 13,084 | 11,428 | -13% | 1 | 1 | 0% | 1,849 | 2,767 | +50% | 0 | 0 | — |
case-11 | fail→fail | 7,807 | 10,670 | +37% | 1 | 1 | 0% | 1,289 | 2,508 | +95% | 0 | 0 | — |
case-12 | pass→pass | 19,810 | 23,359 | +18% | 1 | 1 | 0% | 2,860 | 4,323 | +51% | 0 | 0 | — |
case-13 | pass→pass | 12,168 | 12,040 | -1% | 1 | 1 | 0% | 1,728 | 2,966 | +72% | 0 | 0 | — |
case-14 | fail→pass | 18,444 | 15,315 | -17% | 1 | 1 | 0% | 2,662 | 3,310 | +24% | 0 | 0 | — |
case-15 | pass→pass | 14,670 | 12,945 | -12% | 1 | 1 | 0% | 2,417 | 3,218 | +33% | 0 | 0 | — |
case-16 | fail→pass | 18,703 | 18,977 | +1% | 1 | 1 | 0% | 3,113 | 4,048 | +30% | 0 | 0 | — |
case-17 | pass→fail | 12,953 | 19,734 | +52% | 1 | 1 | 0% | 2,142 | 4,112 | +92% | 0 | 0 | — |
case-18 | fail→fail | 14,095 | 19,800 | +40% | 1 | 1 | 0% | 2,280 | 3,621 | +59% | 0 | 0 | — |
case-19 | fail→fail | 15,943 | 18,183 | +14% | 1 | 1 | 0% | 2,812 | 3,751 | +33% | 0 | 0 | — |
case-20 | pass→pass | 12,082 | 17,846 | +48% | 1 | 1 | 0% | 1,938 | 3,599 | +86% | 0 | 0 | — |
case-21 | pass→pass | 12,906 | 18,092 | +40% | 1 | 1 | 0% | 1,969 | 3,621 | +84% | 0 | 0 | — |
case-22 | pass→pass | 16,248 | 16,379 | +1% | 1 | 1 | 0% | 2,319 | 3,460 | +49% | 0 | 0 | — |
case-23 | pass→pass | 16,002 | 15,128 | -5% | 1 | 1 | 0% | 2,347 | 3,216 | +37% | 0 | 0 | — |
case-24 | pass→pass | 17,996 | 17,676 | -2% | 1 | 1 | 0% | 2,705 | 3,394 | +25% | 0 | 0 | — |
case-25 | pass→pass | 15,888 | 19,982 | +26% | 1 | 1 | 0% | 2,892 | 4,159 | +44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +16 percentage points is the difference between those two pass rates over the 25 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/15/2026 | +9% |
Other measured skills in the registry, with their headline benchmark lift.