Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design when the source names competing algorithms and they are about to become a related-work paragraph instead of arms. Covers running every named baseline at equal tuning effort, and sweeping the parameter you claim credit for.
.claude/skills/tangxiangru-math-equal-effort-baselines-and-knob-sweeps/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 43% | 0% |
An algorithmic claim is a comparison, so the baseline suite is a build item, not a related-work paragraph. Enumerate the named competitors from the references shipped with the problem, implement or install each as a real arm, and run them on identical instances at an identical budget, reported in the same table and the same figure. Give every arm the same tuning effort and the same optional machinery: if a baseline gets restarts, warm starts, preconditioning or a tuned step size, your method gets them too, and the reverse. An asymmetric enhancement is the commonest way an apparent speed-up turns out to be a configuration difference.
Attribute the gain. Ablate the novel component on and off, then sweep that component's own hyper-parameter across its range to locate the optimum and the point where gains saturate or reverse. A component reported only as present or absent leaves a reader unable to distinguish a mechanism from a lucky setting.
Report quality and cost as a pair for the same comparison -- objective value with iterations or time, solution quality with memory or node expansions, accuracy with throughput -- and give the break-even factor. Break results out per instance family and per problem regime the benchmark distinguishes, including families outside your method's design or training regime and families that are degenerate at the shipped settings. Each still gets its own reported row rather than being pooled away or dropped.
Four separate measured demands sit in criteria the bare agent never emitted: the full named baseline suite as separate arms at identical budget (4 tasks), a component ablation that additionally sweeps the component's own hyper-parameter to find its optimum and saturation (3 tasks), the quality/cost pair for the same comparison (3 tasks, one lost because sum-of-costs was computed but never paired with the collision-reduction claim), and per-instance-family rows including degenerate families (3 tasks). It also fixes the observed asymmetric-component failure, where an enhancement was implemented for the competitor but not for the proposed method.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 46,642 | 60,823 | +30% | 1 | 1 | 0% | 7,138 | 6,042 | -15% | 0 | 0 | — |
case-02 | fail→pass | 43,631 | 39,566 | -9% | 1 | 1 | 0% | 7,185 | 7,397 | +3% | 0 | 0 | — |
case-03 | fail→pass | 42,555 | 34,228 | -20% | 1 | 1 | 0% | 8,098 | 6,208 | -23% | 0 | 0 | — |
case-04 | fail→pass | 22,083 | 20,637 | -7% | 1 | 1 | 0% | 3,682 | 3,644 | -1% | 0 | 0 | — |
case-05 | fail→pass | 46,573 | 19,685 | -58% | 1 | 1 | 0% | 2,482 | 3,555 | +43% | 0 | 0 | — |
case-06 | pass→pass | 18,689 | 44,204 | +137% | 1 | 1 | 0% | 3,047 | 4,524 | +48% | 0 | 0 | — |
case-07 | fail→pass | 38,053 | 19,827 | -48% | 1 | 1 | 0% | 3,329 | 3,696 | +11% | 0 | 0 | — |
case-08 | fail→pass | 16,358 | 28,269 | +73% | 1 | 1 | 0% | 2,128 | 3,666 | +72% | 0 | 0 | — |
case-09 | fail→pass | 30,931 | 42,613 | +38% | 1 | 1 | 0% | 3,108 | 4,208 | +35% | 0 | 0 | — |
case-10 | fail→pass | 15,859 | 19,860 | +25% | 1 | 1 | 0% | 2,432 | 3,457 | +42% | 0 | 0 | — |
case-11 | fail→pass | 36,405 | 26,121 | -28% | 1 | 1 | 0% | 3,268 | 4,103 | +26% | 0 | 0 | — |
case-12 | fail→pass | 32,973 | 31,960 | -3% | 1 | 1 | 0% | 2,805 | 3,171 | +13% | 0 | 0 | — |
case-13 | fail→pass | 21,915 | 23,534 | +7% | 1 | 1 | 0% | 3,118 | 4,013 | +29% | 0 | 0 | — |
case-14 | fail→pass | 18,850 | 23,492 | +25% | 1 | 1 | 0% | 2,925 | 3,429 | +17% | 0 | 0 | — |
case-15 | fail→pass | 20,154 | 18,850 | -6% | 1 | 1 | 0% | 3,299 | 3,769 | +14% | 0 | 0 | — |
case-16 | fail→pass | 29,082 | 28,965 | -0% | 1 | 1 | 0% | 3,108 | 2,908 | -6% | 0 | 0 | — |
case-17 | fail→pass | 22,437 | 62,489 | +179% | 1 | 1 | 0% | 2,563 | 4,283 | +67% | 0 | 0 | — |
case-18 | fail→pass | 40,340 | 14,356 | -64% | 1 | 1 | 0% | 3,011 | 2,841 | -6% | 0 | 0 | — |
case-19 | fail→fail | 27,274 | 28,179 | +3% | 1 | 1 | 0% | 3,156 | 4,732 | +50% | 0 | 0 | — |
case-20 | pass→pass | 67,550 | 21,143 | -69% | 1 | 1 | 0% | 3,992 | 4,481 | +12% | 0 | 0 | — |
case-21 | pass→pass | 45,877 | 32,776 | -29% | 1 | 1 | 0% | 2,423 | 6,858 | +183% | 0 | 0 | — |
case-22 | pass→pass | 31,161 | 30,418 | -2% | 1 | 1 | 0% | 3,636 | 5,067 | +39% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +77 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.