Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
.claude/skills/yogsoth-ai-ladder-quality-order/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 186% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -27% | 0% |
You rank ONE topic's 6 research-design samples (each a research_graph + research_result pair) by quality. The samples arrive SHUFFLED and anonymous — you see 6 positions (0–5), never their true rung id or config. You judge only on the D1–D5 standard:
Judge only on the D1–D5 standard above; never on academic-publication criteria of any kind. You never see any quality-check list.
You will be asked to compare two positions at a time. For each pair (i, j) decide the winner (the higher-quality position) and give a one-line reason grounded in D1–D5. Do not assign absolute scores — only pick a winner per pair. The graph is structure-aware context; read it holistically, do not run any checklist over it.
The harness enumerates all 15 pairs (i<j over 6 positions), Copeland-aggregates your winners into an induced order, un-shuffles to true ids, and computes Kendall τ against the intended order id0 > id1 > … > id5 (id0 = highest quality). You only emit {winner, reason} per pair.
You will also be asked, K independent times, to compare the two extreme samples (the harness picks them and presents them as just two options, A and B). Return {"winner": "A" | "B"} — exactly the label of the higher-quality one. Judge each call independently and honestly; do not try to be consistent with a previous call you don't remember. (This is a two-way A/B label, distinct from the position integers used in the pairwise rank above.)
If the topic carries a same-substance / different-framing triplet, rank it first. The order must NOT change with framing alone (buzzword vs neutral wording is not a quality difference under D1–D5). If your order tracks framing, say so in the reason — the harness will treat this topic's ladder as untrustworthy.
The harness assembles loss2.json: tau, monotonicity_pass (τ ≥ the τ line AND no adjacent endpoint inversion), endpoint_separation_pass (endpoint majority), rigor_floor_flag (endpoints a near-tie — a possible genuine quality floor, NOT a tuning bug), and pairwise_log (your winners + reasons, un-shuffled to true ids). Your only job: honest per-pair winners and D1–D5 reasons.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 10,108 | 11,678 | +16% | 1 | 1 | 0% | 896 | 1,944 | +117% | 0 | 0 | — |
case-02 | fail→fail | 14,233 | 13,900 | -2% | 1 | 1 | 0% | 1,855 | 2,355 | +27% | 0 | 0 | — |
case-03 | pass→fail | 22,350 | 15,816 | -29% | 1 | 1 | 0% | 2,673 | 2,718 | +2% | 0 | 0 | — |
case-04 | pass→pass | 29,152 | 28,714 | -2% | 1 | 1 | 0% | 5,015 | 5,004 | -0% | 0 | 0 | — |
case-05 | fail→pass | 12,342 | 7,283 | -41% | 1 | 1 | 0% | 1,186 | 1,040 | -12% | 0 | 0 | — |
case-06 | pass→pass | 9,023 | 7,684 | -15% | 1 | 1 | 0% | 815 | 1,138 | +40% | 0 | 0 | — |
case-07 | pass→pass | 12,769 | 7,432 | -42% | 1 | 1 | 0% | 1,237 | 1,097 | -11% | 0 | 0 | — |
case-08 | fail→fail | 8,545 | 10,173 | +19% | 1 | 1 | 0% | 566 | 1,327 | +134% | 0 | 0 | — |
case-09 | pass→pass | 11,938 | 8,705 | -27% | 1 | 1 | 0% | 1,365 | 1,264 | -7% | 0 | 0 | — |
case-10 | pass→pass | 9,588 | 7,326 | -24% | 1 | 1 | 0% | 778 | 1,077 | +38% | 0 | 0 | — |
case-11 | pass→pass | 10,537 | 8,083 | -23% | 1 | 1 | 0% | 1,007 | 1,215 | +21% | 0 | 0 | — |
case-12 | pass→pass | 14,915 | 8,422 | -44% | 1 | 1 | 0% | 1,934 | 1,308 | -32% | 0 | 0 | — |
case-13 | pass→pass | 7,971 | 7,585 | -5% | 1 | 1 | 0% | 510 | 968 | +90% | 0 | 0 | — |
case-14 | pass→pass | 16,960 | 7,686 | -55% | 1 | 1 | 0% | 1,808 | 1,014 | -44% | 0 | 0 | — |
case-15 | fail→pass | 18,235 | 12,202 | -33% | 1 | 1 | 0% | 1,891 | 1,897 | +0% | 0 | 0 | — |
case-16 | fail→pass | 9,676 | 13,095 | +35% | 1 | 1 | 0% | 790 | 2,258 | +186% | 0 | 0 | — |
case-17 | fail→pass | 17,551 | 9,307 | -47% | 1 | 1 | 0% | 2,013 | 1,472 | -27% | 0 | 0 | — |
case-18 | fail→fail | 9,523 | 8,168 | -14% | 1 | 1 | 0% | 976 | 1,239 | +27% | 0 | 0 | — |
case-19 | fail→pass | 17,620 | 10,483 | -41% | 1 | 1 | 0% | 2,257 | 1,708 | -24% | 0 | 0 | — |
case-20 | fail→pass | 15,720 | 13,095 | -17% | 1 | 1 | 0% | 1,890 | 2,273 | +20% | 0 | 0 | — |
case-21 | fail→pass | 13,883 | 9,720 | -30% | 1 | 1 | 0% | 1,589 | 1,562 | -2% | 0 | 0 | — |
case-22 | fail→pass | 9,494 | 7,544 | -21% | 1 | 1 | 0% | 849 | 1,135 | +34% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.