Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run an A/B test of the same task at multiple model tiers (Haiku, Sonnet, optionally Opus). Captures outputs, computes structural and semantic diffs, scores quality, writes a markdown comparison report. Use when the user wants to validate "is Haiku good enough for this task class?" or runs /tokenwise:ab "<task description>".
.claude/skills/bilal140202-tokenwiseab-ab-test-a-task-across-model-tiers/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -22% | 0% |
Run the same task at multiple tiers and compare outputs.
Expected form: <task description> [--tiers haiku,sonnet,opus]
-- flag (or the whole string if no flags)--tiers haiku,sonnet is the default (skip Opus by default — it's the baseline)--tiers haiku,sonnet,opus runs all threeIf $ARGUMENTS is empty, ask the user: > What task should I A/B test? Provide a task description (e.g., "rename getCwd to getCurrentWorkingDirectory across the codebase").
A/B test will run this task <N> times (once per tier). Estimated cost: $<rough estimate based on task size>. Proceed? [Y/n] Estimate by treating the task as ~10k input + ~1k output per tier and summing.
Task(description: <task>, subagent_type: "general-purpose", model: <tier>, prompt: <task>)model: param is silently overridden (Anthropic Issue #47488), warn user and abort with:> Cannot A/B test on this Claude Code build — subagent model routing is not honored. See /tokenwise:install probe results.
> Compare these N outputs for the task "<task>". Score each from 1-10 on (a) completeness, (b) correctness, (c) clarity. Return a JSON object: {"<tier>": {"completeness": N, "correctness": N, "clarity": N, "overall": N, "notes": "..."}}. Be honest — if two outputs are equivalent, give them the same score.
/tokenwise:report:./tokenwise-ab-<YYYYMMDD-HHMMSS>.md:markdown# A/B Test — <first 80 chars of task> Run: <ISO 8601 timestamp> Task: "<full task>" Tiers tested: <list> ## Tier comparison | Tier | Tokens (in/out) | Cost | Duration | Quality | |--------|------------------|----------|----------|---------| | Haiku | <in> / <out> | $<cost> | <ms> | <q>/10 | | Sonnet | ... | ... | ... | ... | ## Semantic scores (table with completeness/correctness/clarity per tier, plus notes from the judge) ## Recommendation (One paragraph: which tier is sufficient for this task class, what override rule to add to CLAUDE.md if any.) ## Output diffs (For each tier, include the full output verbatim, or a truncated version with full output in a fenced block.)
.tokenwise/log.ndjson with task_class: "ab-test" so it doesn't pollute regular routing stats. A/B test complete. Report: tokenwise-ab-<ts>.md Recommendation: <one-sentence recommendation>
Task (for spawning per-tier runs and the semantic judge), Read, Write, Bash.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,567 | 4,560 | -78% | 1 | 1 | 0% | 3,824 | 1,992 | -48% | 0 | 0 | — |
case-02 | fail→fail | 16,350 | 3,637 | -78% | 1 | 1 | 0% | 3,298 | 1,813 | -45% | 0 | 0 | — |
case-03 | fail→fail | 16,086 | 3,632 | -77% | 1 | 1 | 0% | 3,497 | 1,820 | -48% | 0 | 0 | — |
case-04 | fail→pass | 10,954 | 3,548 | -68% | 1 | 1 | 0% | 1,892 | 1,916 | +1% | 0 | 0 | — |
case-05 | pass→pass | 9,635 | 4,680 | -51% | 1 | 1 | 0% | 2,047 | 2,153 | +5% | 0 | 0 | — |
case-06 | pass→pass | 3,546 | 1,094 | -69% | 1 | 1 | 0% | 586 | 1,259 | +115% | 0 | 0 | — |
case-07 | fail→pass | 10,041 | 6,531 | -35% | 1 | 1 | 0% | 2,021 | 2,619 | +30% | 0 | 0 | — |
case-08 | fail→pass | 12,281 | 3,073 | -75% | 1 | 1 | 0% | 2,324 | 1,676 | -28% | 0 | 0 | — |
case-09 | fail→pass | 7,610 | 1,847 | -76% | 1 | 1 | 0% | 1,494 | 1,405 | -6% | 0 | 0 | — |
case-10 | fail→fail | 13,878 | 1,880 | -86% | 1 | 1 | 0% | 2,895 | 1,434 | -50% | 0 | 0 | — |
case-11 | fail→pass | 15,683 | 4,993 | -68% | 1 | 1 | 0% | 2,668 | 2,084 | -22% | 0 | 0 | — |
case-12 | pass→pass | 9,589 | 1,875 | -80% | 1 | 1 | 0% | 1,933 | 1,513 | -22% | 0 | 0 | — |
case-13 | pass→pass | 6,289 | 2,846 | -55% | 1 | 1 | 0% | 1,379 | 1,586 | +15% | 0 | 0 | — |
case-14 | fail→pass | 6,624 | 2,797 | -58% | 1 | 1 | 0% | 1,283 | 1,734 | +35% | 0 | 0 | — |
case-15 | pass→pass | 11,929 | 2,996 | -75% | 1 | 1 | 0% | 2,219 | 1,716 | -23% | 0 | 0 | — |
case-16 | fail→pass | 6,591 | 1,406 | -79% | 1 | 1 | 0% | 1,224 | 1,285 | +5% | 0 | 0 | — |
case-17 | fail→fail | 3,447 | 1,384 | -60% | 1 | 1 | 0% | 529 | 1,260 | +138% | 0 | 0 | — |
case-18 | fail→fail | 15,331 | 7,545 | -51% | 1 | 1 | 0% | 2,612 | 2,383 | -9% | 0 | 0 | — |
case-19 | fail→pass | 12,703 | 3,888 | -69% | 1 | 1 | 0% | 2,092 | 1,759 | -16% | 0 | 0 | — |
case-20 | pass→fail | 4,425 | 3,455 | -22% | 1 | 1 | 0% | 773 | 1,738 | +125% | 0 | 0 | — |
case-21 | pass→fail | 14,550 | 4,092 | -72% | 1 | 1 | 0% | 2,942 | 1,785 | -39% | 0 | 0 | — |
case-22 | fail→fail | 3,481 | 6,204 | +78% | 1 | 1 | 0% | 605 | 2,194 | +263% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.