Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Deploy a bounded, profile-loaded adversarial review squad at an SDD point-cut (post-spec, post-plan, post-tasks, pre-merge, or an ad-hoc decision) so independent doctrine lenses converge on findings one reviewer would miss. Triggers: "deploy a squad", "adversarial squad", "post-tasks anti-laziness pass", "pre-spec investigation squad", "brownfield check", "second opinion on this design", "review squad", "run a multi-lens review". Does NOT handle: the implement-review loop (use spec-kitty-impleme
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -36% | 0% |
The operational HOW. The doctrinal WHEN/WHY is the procedure adversarial-squad-deployment (packs/built-in/procedures/adversarial-squad-deployment.procedure.yaml), which sits under the brownfield-onboarding paradigm. This skill changes no mission type or guard; it is a technique the orchestrator opts into.
A squad is worth its tokens at a high-leverage point-cut where one reviewer's blind spot is expensive:
/spec-kitty.specify → pre-spec investigation (scope, prior art, live repros)/spec-kitty.plan → post-planning brownfield check (foldable issues, split-brain, deprecations)/spec-kitty.tasks → post-tasks anti-laziness pass (fakeable DoDs, decomposition realism)Do NOT use it as a rubber stamp, and do NOT wire it as a mandatory gate.
not a vibe check.
architect-alphonso — structure / seams / topologydebugger-debbie — live-evidence, coverage, "would this catch the regression?"reviewer-renata — anti-laziness, contract-vs-implementation, fakeable assertionsrandy-reducer — duplication / dead code (⚠ duct-tape bias — read critically)paula-patterns — decomposition, boundaries, second-opinion adjudicationplanner-priti — scope, sequencing, tracker hygienepython-pedro — implementer feasibilitydoctrine-daphne — doctrine integrity / DRG wiringScale past 4 only for an explicit "audit / comprehensive" ask.
"FIRST run spec-kitty agent profile show <id> and spec-kitty charter context --action <action> --json; apply the resolved initialization, boundaries, directives, and tactics, then state which you applied." Loading the profile — not naming a persona — is the point. Only a read-only harness that cannot invoke the CLI may read packs/built-in/agent_profiles/<id>.agent.yaml; that degraded fallback can diverge because overlays, specializes_from lineage, and enhances/overrides semantics are not applied. Keep delegates read-only unless the task is an isolated implementation in its own worktree.
[SEVERITY] file:line — issue — recommendation, ending in a verdict, grounded in cited evidence, with honest concession of where its lens does not apply. A steelman that over-claims is weak; an adversary that concedes nothing is noise.
lighter tier for mechanical/tracker delegates.
consequential point, do NOT average — adjudicate from the source, or dispatch one focused second-opinion delegate. Be critical of any delegate with a known bias. If irreconcilable, escalate to the human with both positions.
findings doc, or memory). The value is convergent evidence that survived independent scrutiny — not a single opinion.
Invoke by name (adversarial-squad) with the point-cut + question, e.g. "adversarial-squad: post-tasks anti-laziness on WP01–WP08." This skill is the alias surface; the doctrine procedure is the canonical record of the technique.
Bounded (3–4) · profile-LOADED · structured output · model discipline · live-evidence-grounded · second-opinion on divergence · never a mission gate.
Other measured skills in the registry, with their headline benchmark lift.