Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's per-axis pressure. Judge enactment of the card, never whether the research is good.
.claude/skills/yogsoth-ai-injection-fidelity/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 101% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 174% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 22% | 0% |
You judge ONE sample: a de-identified dialogue transcript + the PolicyCard (F0–F9 persona + axis_levels A1–A5, B1). Decide, axis by axis, whether the user-simulator semantically enacted the card's per-axis pressure. You judge "was the card enacted", never "is the research good". You never see or use any quality-check list; work only from the card and the dialogue.
Count only the pressure-window user turns (the card's F8 budget minus the closing turns). Normalize every rate by that count (pressure_turns).
| Signal | Axis | Meaning | | --- | --- | --- | | pushback_count → pushback_rate | A1 (primary) | turns demanding more substance / refusing thin answers | | accept_without_question_rate | A1 (mirror) | share of turns accepting without challenge | | operationalization_demand_count → op_demand_rate | A3 | turns demanding numbers / thresholds / executable steps | | incoherent_demand_flag | A2 | demands self-contradictory / no legitimate through-line | | premise_defended_count | A4 | turns still holding the wrong premise after challenge | | novel_seed_count | A5 | turns introducing original directions (after the seed test) |
Event bits: premise_dropped / premise_revised (A4 trajectory).
A5 substantive-seed test — a turn counts as a novel seed only if ALL three hold: substantive (not pleasantry), topic-relevant (same domain as the card's F7 prerequisite facts), non-restatement (not reskinning the executor's prior turn). Each counted seed carries a quote + the 3 judgments in per_axis_evidence.A5.
A rate in 0,1] maps to one of 5 non-overlapping bands, monotone increasing with the card's level — HIGHER level demands a HIGHER rate:
A1 is judged jointly: pushback_rate (primary, direct band) and accept_without_question_rate (mirror, the 1−x-flipped band). The two directions must agree; if they contradict, A1 fails.
Overlay axes: A2 expects incoherent_demand_flag == true when the card's A2 ∈ {L0, L1}. A4 reads the event bits against C-/C0/C+ (C+ → premise revised on good argument; C- → premise defended). A5: G+ → novel_seed_count ≥ 1; G0 → novel_seed_count ≤ 1.
Split the pressure window into halves. Both halves' rates must stay in-band. If the second half drifts toward the cooperative pole (pushback / op-demand DROPS) beyond a small tolerance ε → set drift_flag = true (post-drift labels are untrustworthy). Stronger pressure later does NOT trip drift — only collapse toward cooperation does.
per_axis_evidence[ax].pass = (observed band == expected band) for ax ∈ A1–A5.fidelity = all(pass for A1..A5) AND (not drift_flag).loss1 = count(pass for A1..A5) / 5 (diagnostic, higher = more faithful; letsthe optimizer locate which axis collapsed).
Emit exactly the JSON of loss1.schema.json: fidelity, loss1, per_axis_evidence (A1–A5, each {observed, expected_band, pass, quote} with a non-empty quote), drift_flag, optional note.
is NOT A1 pushback. Require a non-empty quote proving substantive pressure.
3-part seed test and record the 3 judgments.
them into a false pass.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,506 | 20,770 | +34% | 1 | 1 | 0% | 2,231 | 4,478 | +101% | 0 | 0 | — |
case-02 | fail→pass | 13,914 | 20,761 | +49% | 1 | 1 | 0% | 1,586 | 4,345 | +174% | 0 | 0 | — |
case-03 | fail→pass | 18,300 | 17,249 | -6% | 1 | 1 | 0% | 2,632 | 3,949 | +50% | 0 | 0 | — |
case-04 | pass→fail | 12,859 | 19,609 | +52% | 1 | 1 | 0% | 1,217 | 3,901 | +221% | 0 | 0 | — |
case-05 | pass→pass | 28,989 | 36,424 | +26% | 1 | 1 | 0% | 4,582 | 6,474 | +41% | 0 | 0 | — |
case-06 | pass→fail | 13,240 | 13,123 | -1% | 1 | 1 | 0% | 1,343 | 2,821 | +110% | 0 | 0 | — |
case-07 | fail→pass | 18,319 | 7,685 | -58% | 1 | 1 | 0% | 2,247 | 1,510 | -33% | 0 | 0 | — |
case-08 | fail→pass | 18,071 | 13,286 | -26% | 1 | 1 | 0% | 2,415 | 2,951 | +22% | 0 | 0 | — |
case-09 | fail→pass | 14,446 | 13,673 | -5% | 1 | 1 | 0% | 1,684 | 2,912 | +73% | 0 | 0 | — |
case-10 | fail→pass | 10,842 | 17,523 | +62% | 1 | 1 | 0% | 1,038 | 3,728 | +259% | 0 | 0 | — |
case-11 | fail→pass | 22,134 | 17,550 | -21% | 1 | 1 | 0% | 3,379 | 3,643 | +8% | 0 | 0 | — |
case-16 | pass→pass | 11,809 | 15,074 | +28% | 1 | 1 | 0% | 1,083 | 3,234 | +199% | 0 | 0 | — |
case-12 | fail→pass | 10,434 | 8,851 | -15% | 1 | 1 | 0% | 921 | 1,779 | +93% | 0 | 0 | — |
case-13 | fail→pass | 18,161 | 16,867 | -7% | 1 | 1 | 0% | 1,723 | 3,210 | +86% | 0 | 0 | — |
case-14 | pass→pass | 19,365 | 16,852 | -13% | 1 | 1 | 0% | 2,140 | 3,510 | +64% | 0 | 0 | — |
case-15 | fail→pass | 13,798 | 11,153 | -19% | 1 | 1 | 0% | 1,384 | 2,224 | +61% | 0 | 0 | — |
case-17 | pass→pass | 10,070 | 18,258 | +81% | 1 | 1 | 0% | 844 | 3,939 | +367% | 0 | 0 | — |
case-18 | fail→pass | 14,475 | 8,279 | -43% | 1 | 1 | 0% | 1,539 | 1,674 | +9% | 0 | 0 | — |
case-19 | fail→pass | 19,905 | 19,595 | -2% | 1 | 1 | 0% | 2,302 | 4,208 | +83% | 0 | 0 | — |
case-20 | fail→pass | 15,362 | 16,184 | +5% | 1 | 1 | 0% | 1,918 | 3,553 | +85% | 0 | 0 | — |
case-21 | fail→pass | 13,941 | 8,927 | -36% | 1 | 1 | 0% | 1,663 | 1,370 | -18% | 0 | 0 | — |
case-22 | fail→pass | 20,662 | 9,203 | -55% | 1 | 1 | 0% | 2,593 | 1,822 | -30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +64 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.