Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when working with performance testing review multi agent review
.claude/skills/sickn33-performance-testing-review-multi-agent-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 99% | 0% |
| case-22 | ✓→✓ | = Same ✓ | 6% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 27% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 44% | 0% |
Compatibility alias of error-debugging-multi-agent-review; use that ID for new references when no existing contract requires this one. The full instructions and support files remain local so existing installations continue to work offline. This is one shared procedure, not an additional capability. Preserve the callable ID when an existing manifest or client configuration uses it. Modified in AAS on 2026-09-05; original metadata and license notices are retained.
Review a defined diff or subsystem from several relevant perspectives, such as correctness, authorization and performance. The skill describes a review procedure; it does not install an orchestration engine or prove compliance.
Pin repository/base/head, the changed paths, intended behavior and available tests. Use independent agents only if the user authorizes delegation and the host supports it. Otherwise conduct the perspectives sequentially. Do not spawn agents merely because this callable ID contains “multi-agent”. Read-only review is the default; fixes, external posts and deployment stay within the user's actual task authority.
bounded question, owned paths, time/effort limit and expected evidence format.
hypotheses separate from demonstrated failures; do not invent confidence scores.
repeating a claim are not evidence of correctness.
the original failing result; a rerun does not erase it.
For a tenant-scoped cache change, review key construction, authorization context and invalidation. Reproduce two tenants requesting the same prompt and different prompt versions within one tenant. Expected result: a finding only if an actual cross-scope hit or stale result can occur, with the smallest case showing it. A general style opinion must not be presented as a data-exposure defect.
No agent-routing code, compliance validator or quality-score calculator is bundled. Parallel review adds cost and can duplicate blind spots. This procedure cannot prove whole-repository safety from a diff, nor authorize production load tests or messages.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 28,937 | 24,154 | -17% | 1 | 1 | 0% | 2,389 | 2,532 | +6% | 0 | 0 | — |
case-01 | pass→pass | 30,539 | 24,405 | -20% | 1 | 1 | 0% | 5,367 | 6,830 | +27% | 0 | 0 | — |
case-02 | pass→pass | 23,682 | 21,913 | -7% | 1 | 1 | 0% | 4,018 | 5,767 | +44% | 0 | 0 | — |
case-03 | pass→pass | 15,925 | 12,647 | -21% | 1 | 1 | 0% | 2,547 | 3,676 | +44% | 0 | 0 | — |
case-04 | pass→pass | 5,560 | 4,735 | -15% | 1 | 1 | 0% | 1,135 | 2,530 | +123% | 0 | 0 | — |
case-05 | pass→pass | 5,091 | 6,139 | +21% | 1 | 1 | 0% | 1,077 | 2,634 | +145% | 0 | 0 | — |
case-06 | pass→pass | 15,612 | 13,607 | -13% | 1 | 1 | 0% | 2,650 | 3,892 | +47% | 0 | 0 | — |
case-07 | pass→pass | 22,355 | 19,053 | -15% | 1 | 1 | 0% | 4,046 | 5,474 | +35% | 0 | 0 | — |
case-08 | fail→fail | 16,361 | 18,196 | +11% | 1 | 1 | 0% | 2,965 | 4,941 | +67% | 0 | 0 | — |
case-09 | pass→pass | 16,205 | 16,697 | +3% | 1 | 1 | 0% | 2,933 | 4,539 | +55% | 0 | 0 | — |
case-10 | pass→pass | 15,008 | 15,737 | +5% | 1 | 1 | 0% | 2,318 | 4,205 | +81% | 0 | 0 | — |
case-11 | pass→pass | 16,404 | 17,267 | +5% | 1 | 1 | 0% | 2,867 | 4,954 | +73% | 0 | 0 | — |
case-12 | pass→pass | 16,485 | 19,593 | +19% | 1 | 1 | 0% | 2,781 | 4,895 | +76% | 0 | 0 | — |
case-13 | fail→pass | 14,264 | 17,166 | +20% | 1 | 1 | 0% | 2,541 | 4,710 | +85% | 0 | 0 | — |
case-14 | pass→pass | 15,192 | 16,671 | +10% | 1 | 1 | 0% | 2,456 | 4,221 | +72% | 0 | 0 | — |
case-15 | pass→pass | 9,723 | 7,816 | -20% | 1 | 1 | 0% | 1,687 | 2,987 | +77% | 0 | 0 | — |
case-16 | pass→pass | 7,741 | 4,177 | -46% | 1 | 1 | 0% | 1,288 | 2,342 | +82% | 0 | 0 | — |
case-17 | pass→pass | 10,363 | 4,563 | -56% | 1 | 1 | 0% | 1,940 | 2,302 | +19% | 0 | 0 | — |
case-18 | pass→pass | 13,642 | 4,126 | -70% | 1 | 1 | 0% | 2,233 | 2,118 | -5% | 0 | 0 | — |
case-19 | fail→pass | 12,315 | 16,408 | +33% | 1 | 1 | 0% | 2,101 | 4,180 | +99% | 0 | 0 | — |
case-20 | pass→pass | 5,039 | 2,245 | -55% | 1 | 1 | 0% | 879 | 1,938 | +120% | 0 | 0 | — |
case-21 | pass→pass | 8,886 | 10,772 | +21% | 1 | 1 | 0% | 1,523 | 3,476 | +128% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.