Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generates the shared per-task fixture that every competing arm starts from:
.claude/skills/kwakseongjae-bench-fixture-gen-benchmark-internal-not-a-user-skill/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -39% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 172% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -44% | 0% |
Generates the shared per-task fixture that every competing arm starts from: a realistic dataset plus an image asset base. Fairness rule: fixtures are generated ONCE per task by this preprocessor and sealed (sha256) before any arm runs; arms never generate their own.
fixture-spec.json):task_iddataset: { global_key (window global), entities: { name, count,fields: {name, kind, enum?, range?}], relations? }], disclosure }
images: generate-task-assets item list ({file, prompt, style})node scripts/generate-task-dataset.mjs --spec <spec> --out <dir>the script VALIDATES structure (entity counts, field presence, enum membership, relation integrity), retries once on failure, then writes data.json + data.js (window global) + prints aggregate seeds.
node scripts/generate-task-assets.mjs --spec <images-spec> --out <assets-dir>every prepared cell's cell.json.
disclosure string carried inside the dataset itself.
fictional names.
sealed dataset at audit time, never hardcoded.
"open cases" ambiguity in wholesale-2026-08-18).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,902 | 43,431 | +191% | 1 | 1 | 0% | 2,741 | 6,726 | +145% | 0 | 0 | — |
case-02 | fail→pass | 36,480 | 31,892 | -13% | 1 | 1 | 0% | 7,684 | 7,585 | -1% | 0 | 0 | — |
case-03 | fail→pass | 26,405 | 18,484 | -30% | 1 | 1 | 0% | 5,080 | 3,092 | -39% | 0 | 0 | — |
case-04 | fail→pass | 28,695 | 13,085 | -54% | 1 | 1 | 0% | 1,042 | 2,834 | +172% | 0 | 0 | — |
case-05 | fail→pass | 11,589 | 5,271 | -55% | 1 | 1 | 0% | 1,671 | 1,317 | -21% | 0 | 0 | — |
case-06 | fail→pass | 19,939 | 12,367 | -38% | 1 | 1 | 0% | 1,819 | 1,024 | -44% | 0 | 0 | — |
case-07 | fail→pass | 11,471 | 8,880 | -23% | 1 | 1 | 0% | 1,789 | 976 | -45% | 0 | 0 | — |
case-08 | fail→pass | 9,191 | 5,576 | -39% | 1 | 1 | 0% | 1,332 | 1,057 | -21% | 0 | 0 | — |
case-09 | fail→pass | 16,230 | 5,606 | -65% | 1 | 1 | 0% | 2,018 | 1,219 | -40% | 0 | 0 | — |
case-10 | pass→pass | 14,958 | 11,603 | -22% | 1 | 1 | 0% | 2,103 | 1,075 | -49% | 0 | 0 | — |
case-11 | fail→pass | 15,044 | 10,770 | -28% | 1 | 1 | 0% | 2,221 | 2,019 | -9% | 0 | 0 | — |
case-12 | pass→pass | 18,134 | 42,301 | +133% | 1 | 1 | 0% | 2,471 | 2,673 | +8% | 0 | 0 | — |
case-13 | fail→pass | 16,813 | 8,139 | -52% | 1 | 1 | 0% | 2,374 | 1,589 | -33% | 0 | 0 | — |
case-14 | fail→pass | 12,913 | 5,680 | -56% | 1 | 1 | 0% | 1,866 | 1,143 | -39% | 0 | 0 | — |
case-15 | pass→pass | 15,085 | 18,154 | +20% | 1 | 1 | 0% | 2,277 | 2,474 | +9% | 0 | 0 | — |
case-16 | fail→pass | 12,113 | 11,998 | -1% | 1 | 1 | 0% | 1,740 | 1,128 | -35% | 0 | 0 | — |
case-17 | fail→pass | 15,431 | 2,713 | -82% | 1 | 1 | 0% | 2,296 | 797 | -65% | 0 | 0 | — |
case-18 | fail→pass | 24,070 | 7,178 | -70% | 1 | 1 | 0% | 2,215 | 1,168 | -47% | 0 | 0 | — |
case-19 | fail→pass | 21,384 | 2,714 | -87% | 1 | 1 | 0% | 3,156 | 835 | -74% | 0 | 0 | — |
case-20 | fail→pass | 18,930 | 8,180 | -57% | 1 | 1 | 0% | 2,324 | 1,660 | -29% | 0 | 0 | — |
case-21 | fail→pass | 19,654 | 12,438 | -37% | 1 | 1 | 0% | 1,959 | 1,250 | -36% | 0 | 0 | — |
case-22 | fail→pass | 13,458 | 2,816 | -79% | 1 | 1 | 0% | 1,662 | 854 | -49% | 0 | 0 | — |
case-23 | fail→fail | 15,549 | 17,216 | +11% | 1 | 1 | 0% | 2,643 | 3,385 | +28% | 0 | 0 | — |
case-24 | fail→fail | 37,817 | 51,593 | +36% | 1 | 1 | 0% | 3,861 | 4,824 | +25% | 0 | 0 | — |
case-25 | fail→fail | 14,162 | 7,405 | -48% | 1 | 1 | 0% | 2,625 | 1,768 | -33% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 24 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +72 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.