Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run an external frontier-model referee (Claude Code -> GPT-5.6 Sol Pro via the Oracle CLI) on a paper, proof, estimator, or replication package -- and adjudicate what comes back. Use when the user says "send this to oracle", "get an external review", "run a referee round", "deep-check this proof", or before a submission when an independent second opinion is worth more than another in-house pass. Never launches bare: brief first, evidence-forcing prompt, coverage manifest, then CONFIRMED/REFUTED/
.claude/skills/pedrohcgs-oracle-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 228% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 71% | 0% |
Never launch a bare Oracle run. This skill is the driver; the mechanics and the full contract live in external-oracle-process.md. Read it before the first run in a project — it carries setup, flags, artifact layout, the payload cliff, and the failure modes that have actually cost runs.
Compose with /credible-claims (the brief before, the claim record after) and /deep-audit (exhaustive in-house coverage first, so the oracle is confirmation, not discovery).
confirm N named fixes cleared?)
assumption concessions, reporting-language downgrades.
Maintain a statement inventory and a cross-round coverage ledger. Each round assigns what to audit and requires the referee to report what it actually verified, so union coverage reaches 100% instead of drifting toward whatever is easiest to read.
Mechanics, flags, and gotchas: the reference, §2–§3. Smoke-test first; check --files-report against the payload cliff; a run with no conversation URL never happened.
Every finding is a CANDIDATE. Hand the batch to /adjudicate-review: judge each against the actual text, compute the computable first, filter the HELD list, and assign CONFIRMED / REFUTED / DOWNGRADED.
> Oracle agreeing with your own reading is not independent confirmation — different models > correlate on the same wrong answer.
Batch every confirmed fix in one pass, re-verify, then run at most one confirmation round. Converged when a round returns no new CONFIRMED correctness defect — only held items and exposition taste. Close with a claim record: what was fixed (location + evidence), what was REFUTED and why, what is unresolved, and which decisions are the user's.
external-oracle-process.md — the mechanics, and the five credibility questions the findings must be sorted into/adjudicate-review — the triage half/credible-claims — brief before, claim record after/deep-audit — in-house coverage firstverification-ladder.md — rung 6; why the oracle comes last| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 30,436 | 6,834 | -78% | 1 | 1 | 0% | 5,464 | 1,234 | -77% | 0 | 0 | — |
case-02 | pass→fail | 52,657 | 10,700 | -80% | 1 | 1 | 0% | 9,120 | 1,111 | -88% | 0 | 0 | — |
case-03 | fail→fail | 7,184 | 11,851 | +65% | 1 | 1 | 0% | 283 | 1,056 | +273% | 0 | 0 | — |
case-04 | fail→pass | 16,781 | 33,085 | +97% | 1 | 1 | 0% | 2,148 | 7,037 | +228% | 0 | 0 | — |
case-05 | fail→fail | 50,571 | 11,146 | -78% | 1 | 1 | 0% | 8,222 | 957 | -88% | 0 | 0 | — |
case-06 | pass→pass | 17,785 | 22,356 | +26% | 1 | 1 | 0% | 2,546 | 3,978 | +56% | 0 | 0 | — |
case-07 | pass→pass | 15,427 | 9,745 | -37% | 1 | 1 | 0% | 2,296 | 1,994 | -13% | 0 | 0 | — |
case-08 | pass→fail | 25,551 | 10,179 | -60% | 1 | 1 | 0% | 2,441 | 2,132 | -13% | 0 | 0 | — |
case-09 | pass→pass | 12,932 | 12,118 | -6% | 1 | 1 | 0% | 1,756 | 1,693 | -4% | 0 | 0 | — |
case-10 | fail→pass | 18,059 | 12,456 | -31% | 1 | 1 | 0% | 2,586 | 2,850 | +10% | 0 | 0 | — |
case-11 | fail→fail | 9,156 | 6,859 | -25% | 1 | 1 | 0% | 1,112 | 1,699 | +53% | 0 | 0 | — |
case-12 | fail→pass | 11,057 | 5,420 | -51% | 1 | 1 | 0% | 1,415 | 1,515 | +7% | 0 | 0 | — |
case-13 | pass→pass | 15,195 | 23,459 | +54% | 1 | 1 | 0% | 2,179 | 2,401 | +10% | 0 | 0 | — |
case-14 | fail→fail | 12,963 | 8,038 | -38% | 1 | 1 | 0% | 1,839 | 1,955 | +6% | 0 | 0 | — |
case-15 | fail→pass | 8,747 | 2,866 | -67% | 1 | 1 | 0% | 1,243 | 1,075 | -14% | 0 | 0 | — |
case-16 | pass→pass | 10,197 | 4,991 | -51% | 1 | 1 | 0% | 1,258 | 1,467 | +17% | 0 | 0 | — |
case-17 | fail→pass | 46,239 | 7,358 | -84% | 1 | 1 | 0% | 1,133 | 1,941 | +71% | 0 | 0 | — |
case-18 | fail→pass | 21,708 | 7,114 | -67% | 1 | 1 | 0% | 895 | 1,799 | +101% | 0 | 0 | — |
case-19 | pass→pass | 11,856 | 5,062 | -57% | 1 | 1 | 0% | 1,584 | 1,312 | -17% | 0 | 0 | — |
case-20 | fail→pass | 19,229 | 6,410 | -67% | 1 | 1 | 0% | 2,202 | 1,667 | -24% | 0 | 0 | — |
case-21 | pass→pass | 10,936 | 8,315 | -24% | 1 | 1 | 0% | 1,444 | 1,769 | +23% | 0 | 0 | — |
case-22 | fail→pass | 16,780 | 6,248 | -63% | 1 | 1 | 0% | 1,436 | 1,536 | +7% | 0 | 0 | — |
case-23 | fail→pass | 150,150 | 11,584 | -92% | 1 | 1 | 0% | 2,186 | 2,162 | -1% | 0 | 0 | — |
case-24 | pass→pass | 11,513 | 8,789 | -24% | 1 | 1 | 0% | 1,671 | 1,825 | +9% | 0 | 0 | — |
case-25 | fail→pass | 14,301 | 25,129 | +76% | 1 | 1 | 0% | 1,987 | 1,864 | -6% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 19 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.