Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Steel-Manning Campaign — adversarial verification of convergence decisions through resurrection advocacy, winner stress-testing, criteria interrogation, and multi-perspective attack using Devil's Advocacy, Pre-mortem, Red Teaming, Dialectical Inquiry methods.
.claude/skills/yogsoth-ai-steel-manning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 121% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 104% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 72% | 0% |
Adversarial verification of convergence decisions. This campaign subjects winners, criteria, and rejected alternatives to rigorous challenge — ensuring that final decisions survive the strongest possible counter-arguments rather than merely the weakest objections.
The campaign deploys four complementary attack vectors: resurrecting rejected candidates to test whether elimination was justified, stress-testing winners to find hidden weaknesses, interrogating the criteria framework itself, and simulating stakeholder objections for political feasibility.
| Signal | Strategy | |--------|----------| | resurrect rejected candidates / give losers a fair hearing | resurrection-advocacy | | stress-test the winner / find weaknesses in top pick | winner-stress-testing | | challenge the criteria themselves / meta-level questioning | criteria-interrogation | | simulate stakeholder objections / political feasibility | stakeholder-objection-simulation | | construct strongest counter-argument / dialectical challenge | counter-thesis-construction |
| Strategy | Method Lineage | |----------|---------------| | resurrection-advocacy | Devil's Advocacy, Dialectical Inquiry, Adversarial Collaboration (Kahneman) | | winner-stress-testing | Pre-mortem (Klein), Red Teaming, Failure Mode Analysis | | criteria-interrogation | Assumption-based Planning, Critical Systems Heuristics, Boundary Critique | | stakeholder-objection-simulation | Role-play, Stakeholder Analysis, Political Feasibility | | counter-thesis-construction | Dialectical Inquiry, Thesis-Antithesis-Synthesis, Adversarial Debate |
| Tactic | SOPs Used | |--------|-----------| | adversarial-debate-protocol | advocate-construction, critic-attack, judge-verdict | | assumption-excavation | assumption-extraction, assumption-challenge, conclusion-sensitivity | | multi-perspective-attack | perspective-assignment, perspective-attack, steel-manning-synthesis |
| SOP | Input | Output | Shareable | |-----|-------|--------|-----------| | advocate-construction | rejected_candidate, context | strongest_case_for_resurrection | validation | | critic-attack | winner, advocate_case | attack_arguments], severity_ratings | validation | | judge-verdict | advocate_case, critic_attacks | verdict, reasoning, conditions | validation | | assumption-extraction | decision, evidence | assumptions], confidence_levels | — | | assumption-challenge | assumption | challenge_argument, alternative, impact_if_wrong | — | | conclusion-sensitivity | assumptions], challenges] | sensitivity_map, critical_assumptions] | — | | perspective-assignment | decision, stakeholders | perspective_briefs] | — | | perspective-attack | decision, perspective_brief | attacks], constructive_alternatives] | — | | steel-manning-synthesis | all_attacks, all_verdicts | final_verdict, surviving_concerns, modifications | — |
| Metric | Minimum | |--------|---------| | Attack perspectives | >= 3 distinct angles | | Debate rounds | >= 2 per contested decision | | Assumptions challenged | >= 5 per winner | | Final verdict | Explicit ACCEPT / REJECT / REVISE with evidence |
mcp__wiki-vault__vault_search — retrieve prior decisions and contextmcp__wiki-vault__vault_query_graph — trace dependency chains for impact analysismcp__wiki-vault__vault_add_edge — record challenge relationshipsThe campaign maintains a Challenge Ledger tracking:
State is passed between strategies via the ledger. Each strategy updates it upon completion.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | counter-thesis-construction | Construct the strongest possible counter-argument to the convergence decision using Dialectical Inquiry and Thesis-Antithesis-Synthesis methods. | | criteria-interrogation | Challenge the evaluation criteria themselves using Assumption-based Planning, Critical Systems Heuristics, and Boundary Critique to ensure the framework is sound. | | resurrection-advocacy | Argue for rejected candidates using Devil's Advocacy, Dialectical Inquiry, and Adversarial Collaboration to ensure elimination was justified. | | stakeholder-objection-simulation | Simulate stakeholder objections through role-play and political feasibility analysis to test whether the decision survives real-world opposition. | | winner-stress-testing | Stress-test the winning candidate using Pre-mortem, Red Teaming, and Failure Mode Analysis to expose hidden weaknesses before commitment. |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. | | convergence-multi-stakeholder-simulation | Simulates diverse stakeholder perspectives and their strongest objections/support arguments. Shared across steel-manning and consensus campaigns. | | convergence-saturation-detection | Determines when to stop iterating — coverage threshold met or marginal returns diminishing. Shared across all campaigns. | | convergence-sensitivity-analysis | Tests conclusion robustness by perturbing parameters and observing rank changes. Shared across scoring, portfolio, and steel-manning campaigns. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 37,887 | 20,057 | -47% | 1 | 1 | 0% | 4,552 | 3,002 | -34% | 0 | 0 | — |
case-01 | fail→pass | 41,223 | 40,473 | -2% | 1 | 1 | 0% | 5,370 | 6,231 | +16% | 0 | 0 | — |
case-02 | fail→pass | 23,806 | 40,866 | +72% | 1 | 1 | 0% | 2,827 | 6,252 | +121% | 0 | 0 | — |
case-04 | fail→fail | 35,936 | 14,687 | -59% | 1 | 1 | 0% | 3,447 | 1,764 | -49% | 0 | 0 | — |
case-05 | fail→fail | 33,100 | 21,581 | -35% | 1 | 1 | 0% | 4,152 | 1,836 | -56% | 0 | 0 | — |
case-06 | fail→fail | 41,706 | 17,351 | -58% | 1 | 1 | 0% | 3,241 | 1,800 | -44% | 0 | 0 | — |
case-07 | fail→pass | 20,623 | 38,124 | +85% | 1 | 1 | 0% | 2,941 | 5,996 | +104% | 0 | 0 | — |
case-08 | pass→pass | 23,532 | 38,281 | +63% | 1 | 1 | 0% | 2,504 | 6,408 | +156% | 0 | 0 | — |
case-09 | fail→pass | 32,773 | 50,764 | +55% | 1 | 1 | 0% | 4,060 | 6,729 | +66% | 0 | 0 | — |
case-10 | fail→pass | 23,059 | 43,571 | +89% | 1 | 1 | 0% | 3,294 | 5,651 | +72% | 0 | 0 | — |
case-11 | fail→pass | 23,617 | 54,062 | +129% | 1 | 1 | 0% | 3,457 | 8,147 | +136% | 0 | 0 | — |
case-12 | fail→fail | 22,953 | 15,307 | -33% | 1 | 1 | 0% | 2,747 | 1,733 | -37% | 0 | 0 | — |
case-13 | fail→fail | 37,531 | 17,958 | -52% | 1 | 1 | 0% | 4,755 | 1,886 | -60% | 0 | 0 | — |
case-14 | fail→pass | 25,764 | 46,650 | +81% | 1 | 1 | 0% | 3,239 | 7,655 | +136% | 0 | 0 | — |
case-15 | fail→pass | 16,336 | 30,013 | +84% | 1 | 1 | 0% | 2,506 | 5,001 | +100% | 0 | 0 | — |
case-16 | fail→pass | 37,445 | 53,325 | +42% | 1 | 1 | 0% | 5,213 | 8,124 | +56% | 0 | 0 | — |
case-17 | fail→pass | 24,513 | 34,942 | +43% | 1 | 1 | 0% | 2,908 | 5,710 | +96% | 0 | 0 | — |
case-18 | fail→fail | 20,892 | 11,447 | -45% | 1 | 1 | 0% | 3,174 | 1,728 | -46% | 0 | 0 | — |
case-19 | pass→fail | 12,008 | 18,339 | +53% | 1 | 1 | 0% | 2,085 | 1,852 | -11% | 0 | 0 | — |
case-20 | pass→fail | 23,202 | 38,185 | +65% | 1 | 1 | 0% | 3,620 | 7,512 | +108% | 0 | 0 | — |
case-21 | pass→fail | 14,887 | 15,675 | +5% | 1 | 1 | 0% | 2,390 | 1,722 | -28% | 0 | 0 | — |
case-22 | fail→pass | 34,083 | 36,454 | +7% | 1 | 1 | 0% | 4,328 | 6,339 | +46% | 0 | 0 | — |
case-23 | fail→pass | 33,063 | 43,083 | +30% | 1 | 1 | 0% | 2,760 | 6,580 | +138% | 0 | 0 | — |
case-24 | fail→pass | 28,770 | 30,783 | +7% | 1 | 1 | 0% | 3,936 | 6,354 | +61% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 15 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +42 percentage points is the difference between those two pass rates over the 15 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.