Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build or modernize an employee handbook for a small business with no HR department -- interviews the owner for actual practices, drafts every core policy in plain English, flags the state-specific requirements that need local verification, and outputs a document employees might actually read.
.claude/skills/onewave-ai-employee-handbook-builder/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 64% | 0% |
Most small businesses run on unwritten rules until the first dispute, and then the absence of a handbook costs real money. Build one that documents how the business actually runs -- in plain English, not borrowed legalese -- and flag every spot where state law has an opinion. Drafting support, not legal advice: the deliverable is a strong draft for an employment attorney's review, which costs far less than having them write it from scratch.
[STATE CHECK: what to verify]. Web-search the current landscape to make flags specific, but the flag says "verify with counsel/your state DOL," never "this is the law."employee-handbook.md (hand to docx for the formatted document), the acknowledgment form, a one-page summary of the ten policies employees actually ask about, and the attorney-review punch list -- every [STATE CHECK] item collected with section references.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | fail→pass | 14,164 | 14,287 | +1% | 1 | 1 | 0% | 2,148 | 2,855 | +33% | 0 | 0 | — |
case-01 | fail→fail | 24,012 | 12,984 | -46% | 1 | 1 | 0% | 3,830 | 2,070 | -46% | 0 | 0 | — |
case-02 | fail→fail | 38,408 | 11,438 | -70% | 1 | 1 | 0% | 6,225 | 2,439 | -61% | 0 | 0 | — |
case-03 | fail→pass | 26,806 | 26,111 | -3% | 1 | 1 | 0% | 4,046 | 4,768 | +18% | 0 | 0 | — |
case-04 | fail→fail | 22,513 | 17,781 | -21% | 1 | 1 | 0% | 3,699 | 3,410 | -8% | 0 | 0 | — |
case-05 | fail→pass | 12,398 | 11,883 | -4% | 1 | 1 | 0% | 2,127 | 2,454 | +15% | 0 | 0 | — |
case-06 | fail→pass | 17,594 | 16,423 | -7% | 1 | 1 | 0% | 3,048 | 3,439 | +13% | 0 | 0 | — |
case-07 | pass→pass | 10,015 | 10,851 | +8% | 1 | 1 | 0% | 1,661 | 2,450 | +48% | 0 | 0 | — |
case-08 | pass→pass | 14,080 | 10,625 | -25% | 1 | 1 | 0% | 2,101 | 2,326 | +11% | 0 | 0 | — |
case-09 | pass→pass | 11,120 | 9,219 | -17% | 1 | 1 | 0% | 1,692 | 2,153 | +27% | 0 | 0 | — |
case-16 | fail→pass | 10,636 | 13,685 | +29% | 1 | 1 | 0% | 1,591 | 2,612 | +64% | 0 | 0 | — |
case-10 | fail→fail | 18,366 | 18,758 | +2% | 1 | 1 | 0% | 2,646 | 3,444 | +30% | 0 | 0 | — |
case-11 | fail→fail | 22,269 | 20,029 | -10% | 1 | 1 | 0% | 3,537 | 3,970 | +12% | 0 | 0 | — |
case-12 | fail→pass | 15,527 | 12,545 | -19% | 1 | 1 | 0% | 2,080 | 2,734 | +31% | 0 | 0 | — |
case-13 | fail→fail | 9,525 | 14,700 | +54% | 1 | 1 | 0% | 1,335 | 2,868 | +115% | 0 | 0 | — |
case-14 | fail→pass | 11,765 | 13,141 | +12% | 1 | 1 | 0% | 1,913 | 2,631 | +38% | 0 | 0 | — |
case-17 | fail→pass | 11,246 | 13,595 | +21% | 1 | 1 | 0% | 1,803 | 2,694 | +49% | 0 | 0 | — |
case-18 | pass→pass | 12,679 | 9,815 | -23% | 1 | 1 | 0% | 1,968 | 2,236 | +14% | 0 | 0 | — |
case-19 | fail→pass | 16,389 | 24,735 | +51% | 1 | 1 | 0% | 2,924 | 4,452 | +52% | 0 | 0 | — |
case-20 | fail→fail | 14,971 | 13,546 | -10% | 1 | 1 | 0% | 2,786 | 3,297 | +18% | 0 | 0 | — |
case-21 | fail→pass | 9,888 | 13,824 | +40% | 1 | 1 | 0% | 1,759 | 3,089 | +76% | 0 | 0 | — |
case-22 | pass→pass | 13,424 | 13,042 | -3% | 1 | 1 | 0% | 2,234 | 3,036 | +36% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.