Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Strategy: Systematic adversarial probing retuned for truth-seeking. Threat surface = the set of load-bearing claims. Output is NOT a resilience score and NOT a hardening list — it is, per claim, the specific observation/computation that would refute it, plus which attacks succeeded. Methods: UFMCS Key Assumptions Check (repurposed), CIA Devil's Advocacy, Platt strong inference.
.claude/skills/yogsoth-ai-red-team-truthseeking/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 264% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 252% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 77% | 0% |
A retuning of systematic red-teaming. Classic red-teaming enumerates a threat surface, fires attack vectors, and outputs a resilience score (0.0-1.0) plus a list of hardening actions. Two things make that wrong for research: (1) "resilience score" is a defense metric — it rewards un-attackability, the signature of an unfalsifiable claim; (2) "hardening" means patching the artifact to deflect future attacks — exactly the patchwork anti-pattern we reject. This variant keeps the systematic-probing machinery (it is genuinely good at enumeration and coverage) but changes what we enumerate and what we output.
| Element | Original (publication/defense) | This variant (truth-seeking) | |---|---|---| | Threat surface | Attackable weaknesses | The set of load-bearing CLAIMS (a claim, not a weakness, is the unit) | | Per-vector goal | Show the artifact can be attacked | Produce the concrete observation/computation that would refute THIS claim | | Primary output | Resilience score 0.0-1.0 | Refutation-condition per claim (falsifiable? what would break it?) | | Secondary output | Hardening / mitigation actions | NONE. Findings route to revise/demote/residue, never to patch-to-survive | | A claim no attack touches | High resilience (good) | UNFALSIFIABLE (RED — worst outcome) |
For each load-bearing claim, the red team does NOT ask "how can I make this look bad?" It asks Platt's strong-inference question: "What is the experiment/observation/computation whose result would force me to abandon this claim?" If a clean such condition exists, the claim is falsifiable and we record it (this is itself the most valuable product — it tells the next round / the sandbox exactly what to measure). If NO such condition can be constructed, the claim is UNFALSIFIABLE and flagged RED.
Enumerate every claim the artifact LEANS ON — not decorative restatements, the ones that, if false, collapse a downstream conclusion. Sort by load: how many downstream conclusions depend on each. Priority targets are the claims that carry the most weight and the claims stated most confidently relative to their evidence.
For each claim, surface the hidden assumptions it rides on. Classify each assumption: SUPPORTED (we have evidence) / ASSERTED (we just believe it) / CONVENIENT (it makes the story prettier — high suspicion, ties to elegance-trap). ASSERTED and CONVENIENT assumptions are the priority attack targets.
For each claim, attempt to construct its refutation-condition (the strong-inference test). Three results:
Where a refutation-condition is constructible by reasoning/compute NOW, actually attempt the refutation (counterexample search, derivation check, limiting-case evaluation). Record success/failure honestly. A successful refutation = BROKEN. A failed severe refutation = CORROBORATED.
One dedicated subagent argues the strongest case that the WHOLE artifact is a seductive product of our own framing — not to be balanced, but to make sure the prettiest claims got the hardest look.
resilience score. We do not summarize truth as a number between 0 and 1.hardening actions. If a claim is weak, we revise/demote/shelve it; we never armor it against scrutiny.RefutationSurfaceMap: per load-bearing claim → {load rank | hidden assumptions classified SUPPORTED/ASSERTED/CONVENIENT | refutation-condition (clean / oracle-only / none) | if attacked: refutation attempted + outcome bucket}. Plus the devil's-advocate brief on the artifact-as-framing-artifact risk.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | devils-advocacy | Construct the strongest possible counter-argument against a position, steelmanning the opposition before attacking. | | key-assumptions-check | Military ACT: systematically enumerate all assumptions, classify by type, and evaluate evidence strength supporting each. | | probe-execution | Execute a single attack probe against an artifact, record the result with evidence and severity classification. | | threat-surface-mapping | Enumerate all attackable surfaces of an artifact — logical, empirical, methodological, social, and practical dimensions. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-09 | fail→pass | 18,930 | 47,754 | +152% | 1 | 1 | 0% | 2,152 | 7,825 | +264% | 0 | 0 | — |
case-02 | fail→pass | 57,566 | 56,069 | -3% | 1 | 1 | 0% | 8,278 | 9,529 | +15% | 0 | 0 | — |
case-01 | fail→pass | 52,440 | 46,559 | -11% | 1 | 1 | 0% | 7,974 | 7,954 | -0% | 0 | 0 | — |
case-03 | pass→pass | 51,260 | 40,002 | -22% | 1 | 1 | 0% | 7,803 | 7,622 | -2% | 0 | 0 | — |
case-04 | fail→pass | 16,652 | 39,470 | +137% | 1 | 1 | 0% | 1,884 | 6,637 | +252% | 0 | 0 | — |
case-05 | pass→pass | 16,601 | 29,983 | +81% | 1 | 1 | 0% | 2,142 | 5,544 | +159% | 0 | 0 | — |
case-06 | pass→pass | 8,235 | 28,506 | +246% | 1 | 1 | 0% | 512 | 6,310 | +1132% | 0 | 0 | — |
case-07 | fail→pass | 25,420 | 31,382 | +23% | 1 | 1 | 0% | 2,847 | 5,041 | +77% | 0 | 0 | — |
case-08 | fail→pass | 13,477 | 47,569 | +253% | 1 | 1 | 0% | 1,187 | 7,836 | +560% | 0 | 0 | — |
case-10 | fail→pass | 35,099 | 42,698 | +22% | 1 | 1 | 0% | 6,058 | 7,404 | +22% | 0 | 0 | — |
case-11 | fail→pass | 29,066 | 26,342 | -9% | 1 | 1 | 0% | 3,741 | 4,617 | +23% | 0 | 0 | — |
case-12 | pass→pass | 12,633 | 40,180 | +218% | 1 | 1 | 0% | 1,192 | 6,873 | +477% | 0 | 0 | — |
case-13 | fail→fail | 23,470 | 25,797 | +10% | 1 | 1 | 0% | 2,698 | 4,995 | +85% | 0 | 0 | — |
case-14 | pass→pass | 28,792 | 25,635 | -11% | 1 | 1 | 0% | 3,542 | 4,629 | +31% | 0 | 0 | — |
case-15 | fail→pass | 15,772 | 12,928 | -18% | 1 | 1 | 0% | 2,448 | 2,621 | +7% | 0 | 0 | — |
case-16 | fail→pass | 31,838 | 31,673 | -1% | 1 | 1 | 0% | 4,103 | 5,717 | +39% | 0 | 0 | — |
case-17 | pass→fail | 23,588 | 37,210 | +58% | 1 | 1 | 0% | 2,864 | 5,872 | +105% | 0 | 0 | — |
case-18 | fail→pass | 30,544 | 28,557 | -7% | 1 | 1 | 0% | 4,092 | 5,017 | +23% | 0 | 0 | — |
case-19 | fail→pass | 21,589 | 28,351 | +31% | 1 | 1 | 0% | 2,345 | 4,754 | +103% | 0 | 0 | — |
case-20 | fail→pass | 25,539 | 37,437 | +47% | 1 | 1 | 0% | 2,869 | 6,522 | +127% | 0 | 0 | — |
case-21 | pass→pass | 22,193 | 21,659 | -2% | 1 | 1 | 0% | 2,531 | 3,727 | +47% | 0 | 0 | — |
case-22 | fail→pass | 14,807 | 28,244 | +91% | 1 | 1 | 0% | 2,281 | 4,865 | +113% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.