Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Campaign: Truth-seeking adversarial validation for scientific research artifacts (NOT publication defense). Core question: Where have we fooled ourselves, and is each load-bearing claim even falsifiable? Win-condition is INVERTED from survival/resilience to active refutation. Methods: Popper falsificationism, Lakatos Proofs and Refutations, Mayo severe testing, Platt strong inference.
.claude/skills/yogsoth-ai-falsification-first-stress-test/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 130% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 83% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 97% | 0% |
Core question: Where have we fooled ourselves — and is each load-bearing claim even falsifiable?
This campaign is a deliberate inversion of publication-oriented stress testing. In a publication frame, the artifact is a thing to be defended; success = it survives debate and gets hardened. That frame optimizes for persuasiveness and is actively dangerous for truth-seeking research: it rewards a claim for being un-attackable, which is exactly the failure mode of an unfalsifiable theory. Here the artifact is a suspect. We WANT to break it, because breaking it teaches us something true and cheap (compute/thought) before we spend expensive effort (sandbox, wet-lab) on a false premise. Confidence is not assumed and defended down; it is EARNED up, only by surviving honest assault.
Every load-bearing claim exits in exactly one bucket:
| Bucket | Meaning | Action | |---|---|---| | BROKEN | A concrete refutation (counterexample / disagreeing case / failed derivation) was found. | Revise the claim or demote it. Record what the refutation taught. | | CORROBORATED | The claim was stated falsifiably, attacked severely, and held. | Raise confidence. Record WHAT was forbidden-and-held (the more it forbids, the stronger). | | UNFALSIFIABLE | No statement of the claim could be found that some observation/computation could refute. | RED FLAG. Demote to conjecture or relabel (e.g. "isomorphism"→"analogy"). |
| Claim/artifact type | Primary strategy | Secondary | |---|---|---| | a beautiful unification / elegant backbone | elegance-trap-probe | adversarial-debate-truthseeking | | an isomorphism claim (A ≅ B ≅ …) | isomorphism-falsification | adversarial-stress-testing (sibling campaign) | | load-bearing propositions | counterfactual-probing (sibling campaign) | adversarial-stress-testing (sibling campaign) | | an "N paths independently converged" claim | independent-convergence-audit | red-team-truthseeking | | a validator / sandbox design | circular-validation-audit | red-team-truthseeking | | any sharp claim | red-team-truthseeking | adversarial-debate-truthseeking |
These are existing campaigns this campaign routes claims TO; they are referenced, not declared as dependencies (4-layer rule: a campaign lists only its own strategies and sops).
| Parameter | S (Quick) | M (Standard) | L (Deep) | |---|---|---|---| | Load-bearing claims attacked | 4 | 8 | 14 | | Refutation attempts per claim | 2 | 4 | 7 | | Severe tests designed per claim | 1 | 2 | 4 | | External evidence searches | 3 | 6 | 12 |
Each attack subagent runs in isolated context to prevent the defender's framing from contaminating the critic. The Falsification Ledger is the single accumulating artifact; every checkpoint appends, never overwrites. Saturation: stop attacking a claim when 2 consecutive rounds find no new refutation vector AND no new severe test.
Produces FalsificationLedger: a table of every load-bearing claim → {falsifiable? (Y/N) | what observation/computation would refute it | attacks attempted | outcome bucket | if BROKEN: the refutation + revision/demotion | if CORROBORATED: what was forbidden-and-held + severity of the test passed}. Plus a top-level honest-residue list: claims that are irreducibly unfalsifiable by compute and require an external oracle.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | adversarial-debate-truthseeking | Strategy: Dialectic engine retuned for truth-seeking, not survival. A defender steelmans a claim into its MOST falsifiable form, a critic attacks to refute it, a judge classifies the exchange into BROKEN/CORROBORATED/UNFALSIFIABLE — the judge does NOT pick a winner or score persuasiveness. Methods: Irving debate (repurposed), Toulmin argumentation, Mayo severe testing. | | circular-validation-audit | Strategy: Run BEFORE building any validator (sandbox/simulation/benchmark). Builds a non-circularity matrix of theory-claim × validator-assumption to detect when a validator would 'confirm' a theory only because it was built on the theory's own premises. A circular validator's PASS carries zero evidential weight. Methods: Cartwright nomological machines, Winsberg sanctioning-of-simulations, tautology detection. | | elegance-trap-probe | Strategy: Attack a beautiful unified result on the suspicion that its beauty is the bug. Distinguishes EARNED simplicity (forbids/predicts/subsumes) from DECORATIVE simplicity (re-describes/relabels/accommodates). Directly serves the Occam aesthetic by making it a falsifiable bar, not a vibe. Methods: Sober parsimony-as-evidence, MDL, Meehl risky prediction, accommodation-vs-prediction. | | independent-convergence-audit | Strategy: Attack the evidential weight of an 'independent convergence' claim. When N reasoning paths all reach the same conclusion, the confidence boost is real only if the paths were actually independent. Measures shared-prior / shared-blindspot contamination and corrects the over-counted confidence. Methods: Bayesian agreement-as-evidence, correlated-error analysis, jury theorem assumptions. | | isomorphism-falsification | Strategy: Attack an isomorphism claim by demanding an explicit structure-preserving map and trying to break it. Targets any multi-language claim of the form 'X ≅ Y ≅ … across N mathematical languages'. Forces the claim to either earn the word 'isomorphism' or be demoted to 'analogy'. Methods: category theory (functor/natural-iso criteria), model theory, Lakatos monster-barring. | | red-team-truthseeking | Strategy: Systematic adversarial probing retuned for truth-seeking. Threat surface = the set of load-bearing claims. Output is NOT a resilience score and NOT a hardening list — it is, per claim, the specific observation/computation that would refute it, plus which attacks succeeded. Methods: UFMCS Key Assumptions Check (repurposed), CIA Devil's Advocacy, Platt strong inference. |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. | | stress-test-saturation-detection | Determines whether validation has reached saturation — no new weaknesses or failure modes being discovered. Used by all 5 campaigns as termination signal. | | verdict-synthesis | Synthesizes findings from a completed campaign into typed verdict reports. Produces DebateVerdict, RedTeamReport, FailureAnticipationReport, CounterfactualMap, or AdversarialStressReport depending on campaign. Also supports cross-campaign StressTestSummary. | | weakness-classification | Classifies discovered weaknesses into severity tiers (fatal/major/minor/cosmetic) with structured justification and exploitability assessment. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 38,933 | 34,506 | -11% | 1 | 1 | 0% | 6,242 | 7,770 | +24% | 0 | 0 | — |
case-01 | fail→fail | 38,676 | 49,227 | +27% | 1 | 1 | 0% | 6,245 | 8,777 | +41% | 0 | 0 | — |
case-03 | fail→pass | 25,588 | 31,980 | +25% | 1 | 1 | 0% | 3,813 | 7,437 | +95% | 0 | 0 | — |
case-04 | fail→fail | 14,483 | 27,963 | +93% | 1 | 1 | 0% | 2,492 | 7,000 | +181% | 0 | 0 | — |
case-05 | fail→pass | 18,688 | 22,757 | +22% | 1 | 1 | 0% | 2,647 | 6,087 | +130% | 0 | 0 | — |
case-06 | fail→fail | 13,030 | 18,085 | +39% | 1 | 1 | 0% | 2,181 | 5,924 | +172% | 0 | 0 | — |
case-07 | pass→pass | 17,684 | 9,939 | -44% | 1 | 1 | 0% | 2,580 | 4,206 | +63% | 0 | 0 | — |
case-08 | fail→pass | 13,710 | 9,012 | -34% | 1 | 1 | 0% | 2,208 | 4,042 | +83% | 0 | 0 | — |
case-09 | pass→pass | 20,527 | 19,872 | -3% | 1 | 1 | 0% | 2,899 | 5,274 | +82% | 0 | 0 | — |
case-10 | fail→pass | 10,380 | 4,425 | -57% | 1 | 1 | 0% | 1,618 | 3,191 | +97% | 0 | 0 | — |
case-11 | fail→pass | 12,559 | 8,626 | -31% | 1 | 1 | 0% | 1,886 | 3,967 | +110% | 0 | 0 | — |
case-12 | pass→pass | 10,834 | 10,376 | -4% | 1 | 1 | 0% | 1,628 | 4,174 | +156% | 0 | 0 | — |
case-13 | fail→pass | 38,326 | 6,783 | -82% | 1 | 1 | 0% | 1,305 | 3,311 | +154% | 0 | 0 | — |
case-14 | fail→pass | 11,746 | 3,287 | -72% | 1 | 1 | 0% | 1,860 | 3,127 | +68% | 0 | 0 | — |
case-15 | fail→pass | 14,874 | 2,849 | -81% | 1 | 1 | 0% | 2,473 | 3,113 | +26% | 0 | 0 | — |
case-16 | fail→pass | 14,685 | 9,177 | -38% | 1 | 1 | 0% | 2,212 | 4,025 | +82% | 0 | 0 | — |
case-17 | fail→pass | 14,483 | 11,737 | -19% | 1 | 1 | 0% | 2,326 | 4,407 | +89% | 0 | 0 | — |
case-18 | pass→pass | 14,048 | 13,977 | -1% | 1 | 1 | 0% | 2,173 | 4,624 | +113% | 0 | 0 | — |
case-19 | fail→pass | 16,752 | 16,674 | -0% | 1 | 1 | 0% | 2,412 | 4,999 | +107% | 0 | 0 | — |
case-20 | pass→pass | 9,834 | 7,033 | -28% | 1 | 1 | 0% | 1,561 | 3,655 | +134% | 0 | 0 | — |
case-21 | fail→pass | 9,753 | 8,325 | -15% | 1 | 1 | 0% | 1,736 | 3,953 | +128% | 0 | 0 | — |
case-22 | pass→pass | 10,201 | 8,312 | -19% | 1 | 1 | 0% | 1,513 | 4,176 | +176% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.