Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Strategy: Attack the evidential weight of an 'independent convergence' claim. When N reasoning paths all reach the same conclusion, the confidence boost is real only if the paths were actually independent. Measures shared-prior / shared-blindspot contamination and corrects the over-counted confidence. Methods: Bayesian agreement-as-evidence, correlated-error analysis, jury theorem assumptions.
.claude/skills/yogsoth-ai-independent-convergence-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 100% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-10 | ✓→✓ | = Same ✓ | 197% | 0% |
A dedicated attacker for one specific piece of evidence artifacts often lean on heavily: N reasoning paths "independently" converged on the same conclusion, and that triple- (or N-fold) convergence is treated as strong corroboration. It IS strong — IF the paths were independent. But paths run by the same model, from the same context, under the same prompts, may carry the same blind spot and the same priors. Correlated agreement is weak evidence dressed as strong evidence. This strategy measures how independent the convergence really was and corrects the confidence we should draw from it.
The logic "many paths agree → likely true" is the Condorcet/jury intuition, and it has a hard precondition: the paths' errors must be (at least partly) INDEPENDENT. If three jurors all read the same biased newspaper, their unanimous verdict carries roughly the weight of one juror, not three. The "newspaper" candidates are: (1) the same base model with the same training priors, (2) the same conversation context seen by all subagents, (3) the same prompt framing imposed by the orchestrator, (4) the same source literature feeding all paths, (5) a shared aesthetic pull toward unification (the elegance-trap acting on all paths at once). Each shared source collapses N effective-independent-paths toward 1.
For the convergence claim, score each potential shared-source channel:
| Shared channel | Present? | How much it correlates the paths | Effect on effective N | |---|---|---|---| | same base model / priors | yes (structural) | high — same inductive biases | large collapse | | same conversation context | depends — were subagents isolated? | medium-high | medium collapse | | same prompt framing | depends — did orchestrator steer toward "find the unifier"? | medium | medium collapse | | same source literature | depends — did all paths cite the same sources? | medium | medium collapse | | shared elegance pull | likely | medium — all drawn to the pretty answer | medium collapse |
Effective independence N_eff = (naive N) discounted by the collapse factors. If N_eff ≈ 1, the convergence is essentially ONE observation and our confidence must drop accordingly.
From the convergence checkpoint, recover for each path: what context it was given, whether it ran in an isolated subagent, what literature or seeds it had, and crucially whether the prompt already named or hinted at the target answer. A path that was TOLD to find a particular kind of unifier did not independently discover one.
The killer question: is there a single upstream framing that MADE all paths converge? E.g. if the setup framed every input as "a symptom of a missing dynamical object," then all paths were pre-aimed at a dynamical unifier. That would make the convergence an artifact of the framing, not a property of the problem. Look specifically for the orchestrator's fingerprints.
Design (and if cheap, run) a genuinely independent additional path — different framing, blind to the conclusion the other paths reached, ideally a different reasoning mode (e.g. purely empirical / bottom-up from data rather than top-down from theory). If it ALSO lands on the same conclusion, that's real independent corroboration. If it lands somewhere else, the original convergence was framing-induced.
Even if the conclusion is right, do the N paths share the SAME potential error? If all paths would be wrong in the same way under some condition, their agreement provides no protection against that error. List the errors the convergence does NOT rule out.
State the corrected evidential weight: e.g. "N-fold convergence, discounted for shared model and framing, is worth approximately one-and-a-half independent observations, not N." Adjust any downstream confidence that cited the convergence.
This audit is NOT trying to prove the conclusion wrong — the conclusion may be entirely correct. It is trying to prevent us from being MORE confident than the evidence warrants. A good outcome is honest recalibration: "the convergence is partly framing-induced, so we should treat the conclusion as a strong hypothesis requiring a real test, not as already-established." That keeps us from skipping the real test because we felt too sure.
ConvergenceIndependenceReport: the independence ledger with N_eff estimate | identified common-cause framing (orchestrator fingerprints, if any) | result of the independent additional path (if run) or its design | correlated errors the convergence does not rule out | corrected confidence statement for the conclusion.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→pass | 20,012 | 25,761 | +29% | 1 | 1 | 0% | 2,306 | 4,623 | +100% | 0 | 0 | — |
case-10 | pass→pass | 22,043 | 28,152 | +28% | 1 | 1 | 0% | 1,866 | 5,534 | +197% | 0 | 0 | — |
case-01 | pass→pass | 67,815 | 31,119 | -54% | 1 | 1 | 0% | 2,719 | 3,779 | +39% | 0 | 0 | — |
case-02 | pass→pass | 23,769 | 49,026 | +106% | 1 | 1 | 0% | 2,286 | 3,404 | +49% | 0 | 0 | — |
case-03 | fail→pass | 19,937 | 31,889 | +60% | 1 | 1 | 0% | 2,102 | 4,132 | +97% | 0 | 0 | — |
case-04 | pass→pass | 45,802 | 39,929 | -13% | 1 | 1 | 0% | 2,419 | 4,759 | +97% | 0 | 0 | — |
case-06 | pass→pass | 18,789 | 24,472 | +30% | 1 | 1 | 0% | 2,265 | 3,911 | +73% | 0 | 0 | — |
case-07 | pass→pass | 14,573 | 26,418 | +81% | 1 | 1 | 0% | 1,929 | 3,880 | +101% | 0 | 0 | — |
case-08 | fail→pass | 20,394 | 19,767 | -3% | 1 | 1 | 0% | 2,124 | 4,113 | +94% | 0 | 0 | — |
case-09 | fail→pass | 34,008 | 23,581 | -31% | 1 | 1 | 0% | 2,049 | 4,237 | +107% | 0 | 0 | — |
case-11 | pass→pass | 23,100 | 40,581 | +76% | 1 | 1 | 0% | 4,002 | 8,265 | +107% | 0 | 0 | — |
case-12 | pass→pass | 16,818 | 24,355 | +45% | 1 | 1 | 0% | 2,077 | 5,400 | +160% | 0 | 0 | — |
case-13 | pass→pass | 18,151 | 18,231 | +0% | 1 | 1 | 0% | 3,088 | 3,684 | +19% | 0 | 0 | — |
case-14 | pass→pass | 22,671 | 25,549 | +13% | 1 | 1 | 0% | 2,642 | 4,399 | +67% | 0 | 0 | — |
case-15 | pass→pass | 23,700 | 14,972 | -37% | 1 | 1 | 0% | 2,325 | 3,569 | +54% | 0 | 0 | — |
case-16 | pass→pass | 19,697 | 20,291 | +3% | 1 | 1 | 0% | 2,385 | 4,153 | +74% | 0 | 0 | — |
case-17 | pass→pass | 12,708 | 23,298 | +83% | 1 | 1 | 0% | 1,607 | 4,286 | +167% | 0 | 0 | — |
case-18 | pass→pass | 36,370 | 32,109 | -12% | 1 | 1 | 0% | 4,801 | 6,262 | +30% | 0 | 0 | — |
case-19 | pass→pass | 16,818 | 12,992 | -23% | 1 | 1 | 0% | 2,347 | 3,057 | +30% | 0 | 0 | — |
case-20 | pass→pass | 10,951 | 14,553 | +33% | 1 | 1 | 0% | 1,964 | 3,388 | +73% | 0 | 0 | — |
case-21 | pass→pass | 17,120 | 17,948 | +5% | 1 | 1 | 0% | 2,280 | 3,626 | +59% | 0 | 0 | — |
case-22 | fail→fail | 14,261 | 15,594 | +9% | 1 | 1 | 0% | 2,071 | 3,439 | +66% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.