Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Campaign: Multi-agent structured debate for adversarial validation. Core question: Can this artifact survive structured adversarial debate? Methods: Irving AI Safety via Debate, Du Society of Mind, Liang MAD, Toulmin Argumentation, D3 framework.
.claude/skills/yogsoth-ai-multiagent-debate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -50% | 0% |
Core question: Can this artifact survive structured adversarial debate?
| Artifact Type | Primary Strategy | Fallback Strategy | |---|---|---| | hypothesis, claim | critic-defender-judge | adversarial-escalation | | research-question | multi-perspective-panel | society-of-mind | | idea, approach | society-of-mind | courtroom-structured | | experiment-design | courtroom-structured | critic-defender-judge | | gap | multi-perspective-panel | adversarial-escalation |
| Parameter | S (Quick) | M (Standard) | L (Deep) | |---|---|---|---| | Debate rounds | 4 | 8 | 12 | | Participating agents | 3 | 5 | 8 | | Coverage dimensions | 3 | 5 | 7 | | External evidence searches | 2 | 5 | 10 |
Each subagent operates in isolated context. The debate-architect designs structure before execution. Transcripts are passed between rounds via structured markdown. Saturation detection terminates when novelty drops below threshold.
Produces DebateVerdict containing: survival assessment, key vulnerabilities, confidence score, debate transcript summary, and recommended mitigations.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | adversarial-escalation | Strategy: Progressive pressure escalation — starts with surface-level challenges and escalates to fundamental assumption attacks based on defender confidence decay. | | courtroom-structured | Strategy: Legal adversarial structure — prosecution presents case, defense responds, evidence is cross-examined, judge delivers verdict. Emphasizes evidence quality and procedural rigor. | | critic-defender-judge | Strategy: Classic triangular debate — Critic attacks, Defender responds, Judge adjudicates. Based on Irving AI Safety via Debate with Toulmin argumentation structure. | | multi-perspective-panel | Strategy: Multi-stakeholder review panel — diverse expert perspectives evaluate artifact simultaneously, then synthesize through structured deliberation. | | society-of-mind | Strategy: Multi-agent collaborative debate based on Du et al. Society of Mind. Agents share perspectives iteratively until convergence or divergence is detected. |
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | evidence-tournament | Tactic: Evidence gathering, cross-examination, and quality judgment. External evidence is collected, presented, challenged, and scored for relevance and reliability. | | stress-test-dialectical-escalation | Tactic: Progressive debate escalation based on confidence thresholds. Each round increases attack sophistication until defender collapses or proves resilient. | | stress-test-perspective-rotation | Tactic: Sequential perspective evaluation with divergence aggregation. Each agent evaluates from a distinct viewpoint, then disagreements are surfaced and resolved. |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. | | debate-transcript-analysis | Extracts key turning points, patterns, and insights from completed debate transcripts. Produces structured summary for verdict synthesis. | | stress-test-saturation-detection | Determines whether validation has reached saturation — no new weaknesses or failure modes being discovered. Used by all 5 campaigns as termination signal. | | verdict-synthesis | Synthesizes findings from a completed campaign into typed verdict reports. Produces DebateVerdict, RedTeamReport, FailureAnticipationReport, CounterfactualMap, or AdversarialStressReport depending on campaign. Also supports cross-campaign StressTestSummary. | | weakness-classification | Classifies discovered weaknesses into severity tiers (fatal/major/minor/cosmetic) with structured justification and exploitability assessment. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 42,184 | 50,185 | +19% | 1 | 1 | 0% | 5,650 | 8,339 | +48% | 0 | 0 | — |
case-02 | pass→pass | 44,231 | 41,252 | -7% | 1 | 1 | 0% | 6,745 | 4,944 | -27% | 0 | 0 | — |
case-03 | fail→fail | 46,619 | 12,676 | -73% | 1 | 1 | 0% | 6,967 | 1,694 | -76% | 0 | 0 | — |
case-04 | pass→pass | 18,665 | 18,631 | -0% | 1 | 1 | 0% | 2,083 | 3,382 | +62% | 0 | 0 | — |
case-05 | pass→pass | 34,366 | 7,969 | -77% | 1 | 1 | 0% | 1,965 | 1,655 | -16% | 0 | 0 | — |
case-06 | fail→pass | 43,384 | 3,796 | -91% | 1 | 1 | 0% | 1,272 | 1,847 | +45% | 0 | 0 | — |
case-07 | fail→pass | 23,249 | 10,174 | -56% | 1 | 1 | 0% | 2,947 | 3,090 | +5% | 0 | 0 | — |
case-08 | fail→pass | 43,808 | 8,331 | -81% | 1 | 1 | 0% | 1,834 | 1,782 | -3% | 0 | 0 | — |
case-09 | fail→pass | 18,908 | 7,702 | -59% | 1 | 1 | 0% | 1,776 | 1,629 | -8% | 0 | 0 | — |
case-10 | fail→pass | 20,694 | 7,843 | -62% | 1 | 1 | 0% | 3,149 | 1,581 | -50% | 0 | 0 | — |
case-11 | fail→pass | 11,710 | 7,824 | -33% | 1 | 1 | 0% | 1,817 | 1,646 | -9% | 0 | 0 | — |
case-12 | fail→fail | 37,669 | 29,120 | -23% | 1 | 1 | 0% | 3,916 | 5,838 | +49% | 0 | 0 | — |
case-13 | pass→pass | 17,866 | 16,799 | -6% | 1 | 1 | 0% | 2,924 | 3,697 | +26% | 0 | 0 | — |
case-14 | fail→pass | 29,833 | 12,464 | -58% | 1 | 1 | 0% | 1,136 | 2,773 | +144% | 0 | 0 | — |
case-15 | pass→pass | 12,942 | 3,734 | -71% | 1 | 1 | 0% | 2,248 | 1,801 | -20% | 0 | 0 | — |
case-16 | pass→pass | 29,022 | 2,714 | -91% | 1 | 1 | 0% | 1,622 | 1,624 | +0% | 0 | 0 | — |
case-17 | fail→pass | 17,494 | 2,911 | -83% | 1 | 1 | 0% | 2,894 | 1,636 | -43% | 0 | 0 | — |
case-18 | pass→pass | 13,498 | 3,594 | -73% | 1 | 1 | 0% | 2,106 | 1,787 | -15% | 0 | 0 | — |
case-19 | fail→pass | 18,551 | 24,156 | +30% | 1 | 1 | 0% | 3,226 | 5,313 | +65% | 0 | 0 | — |
case-20 | fail→pass | 13,891 | 3,462 | -75% | 1 | 1 | 0% | 2,121 | 1,657 | -22% | 0 | 0 | — |
case-21 | fail→pass | 13,716 | 2,424 | -82% | 1 | 1 | 0% | 2,106 | 1,616 | -23% | 0 | 0 | — |
case-22 | fail→pass | 17,148 | 4,000 | -77% | 1 | 1 | 0% | 2,553 | 1,814 | -29% | 0 | 0 | — |
case-23 | pass→pass | 8,539 | 10,341 | +21% | 1 | 1 | 0% | 1,799 | 3,221 | +79% | 0 | 0 | — |
case-24 | pass→pass | 13,610 | 22,512 | +65% | 1 | 1 | 0% | 3,189 | 5,996 | +88% | 0 | 0 | — |
case-25 | pass→fail | 31,210 | 11,951 | -62% | 1 | 1 | 0% | 2,019 | 2,909 | +44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 21 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +44 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.