Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Adversarial Quarto-vs-Beamer parity QA. A critic agent compares the Quarto HTML render to the Beamer PDF benchmark for content/visual parity; a fixer agent applies fixes; loops until APPROVED (max 5 rounds). Use when user says "qa the quarto", "check parity", "does the html match the pdf?", "quarto matches beamer?", or after a translate-to-quarto run. Requires both the `.qmd` rendered and a `.pdf` benchmark.
.claude/skills/pedrohcgs-qa-quarto/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -21% | 0% |
Compare Quarto HTML slides against their Beamer PDF benchmark using an iterative critic/fixer loop.
Philosophy: The Beamer PDF is the gold standard. The Quarto translation must be at least as good in every dimension.
Phase 0: Pre-flight → Phase 1: Critic audit → Phase 2: Fixer → Phase 3: Re-audit → Loop until APPROVED (max 5 rounds)| Gate | Condition | |------|-----------| | Overflow | NO content cut off | | Plot Quality | Interactive charts >= static plots | | Content Parity | No missing slides/equations/text | | Visual Regression | Quarto >= Beamer in all dimensions | | Slide Centering | Content centered, no jumping | | Notation Fidelity | All math verbatim from Beamer |
Launch the quarto-critic agent to compare Beamer vs Quarto comprehensively. Report saved to quality_reports/[Lecture]_qa_critic_round1.md.
If not APPROVED, launch quarto-fixer agent to apply fixes (Critical → Major → Minor), re-render, and verify.
Re-launch critic to verify fixes. Loop back to Phase 2 if needed.
This is the loop-until-dry primitive from orchestrator-protocol.md: the critic returns FINDINGs (the hard-gate table is the CRITICAL roll-up, per orchestration-schemas.md); the loop converges when a round adds 0 new CRITICAL/MAJOR findings (deduped on id = sha1(file:line:locus)), not at a fixed round count.
summary-parity.md).Save to quality_reports/[Lecture]_qa_final.md with hard gate status, iteration summary, and remaining issues.
This skill's reviewers emit findings under the machine-checked contract in finding-schema.json. Reports are JSON arrays.
Smoke-test the harness before spending review effort — a run that fans out reviewers and then cannot write a valid report has wasted the whole pass:
bashecho '[]' | python3 scripts/validate-findings.py
Then, before presenting any summary:
bashpython3 scripts/validate-findings.py <report>.json # exit 0 required
What the contract forces, and why:
rule — the documented rule or standard violated. A finding citing no rule is anopinion, and opinions do not gate a commit.
failing_case — a concrete configuration under which the claim breaks, or the exactmissing hypothesis. "This could be clearer" does not validate.
id = sha1("<file>:<line>:<locus>") — deterministic, so dedup across rounds isexact and the two-strikes rule is checkable rather than eyeballed.
mechanical — true only for fixes that cannot change a result (typo, cross-reference,formatting, label). Never for an estimand, assumption, specification, inference procedure, sample definition, or reporting language: those return to the researcher.
Apply the per-lens evidence burdens and the "does NOT count" filters in orchestration-schemas.md §7 before verification, so known false alarms never reach the judge. The verifier pass is refute-biased: only verdict: "confirmed" findings ship; anything it cannot ground is dropped, not downgraded to a warning.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 55,122 | 66,267 | +20% | 1 | 1 | 0% | 4,640 | 1,500 | -68% | 0 | 0 | — |
case-02 | fail→fail | 5,295 | 33,915 | +541% | 1 | 1 | 0% | 265 | 1,272 | +380% | 0 | 0 | — |
case-03 | fail→fail | 6,797 | 4,056 | -40% | 1 | 1 | 0% | 229 | 1,263 | +452% | 0 | 0 | — |
case-04 | fail→pass | 7,819 | 1,713 | -78% | 1 | 1 | 0% | 1,188 | 1,349 | +14% | 0 | 0 | — |
case-05 | fail→pass | 13,210 | 4,685 | -65% | 1 | 1 | 0% | 2,228 | 1,824 | -18% | 0 | 0 | — |
case-06 | fail→pass | 8,531 | 33,712 | +295% | 1 | 1 | 0% | 1,175 | 1,706 | +45% | 0 | 0 | — |
case-07 | pass→pass | 9,047 | 4,630 | -49% | 1 | 1 | 0% | 1,401 | 1,813 | +29% | 0 | 0 | — |
case-08 | pass→pass | 13,795 | 5,634 | -59% | 1 | 1 | 0% | 2,095 | 1,954 | -7% | 0 | 0 | — |
case-09 | fail→pass | 9,103 | 2,573 | -72% | 1 | 1 | 0% | 1,485 | 1,504 | +1% | 0 | 0 | — |
case-10 | pass→pass | 5,359 | 2,829 | -47% | 1 | 1 | 0% | 729 | 1,583 | +117% | 0 | 0 | — |
case-11 | pass→pass | 7,287 | 2,820 | -61% | 1 | 1 | 0% | 1,076 | 1,489 | +38% | 0 | 0 | — |
case-12 | fail→pass | 10,921 | 2,350 | -78% | 1 | 1 | 0% | 1,862 | 1,476 | -21% | 0 | 0 | — |
case-13 | fail→fail | 7,018 | 2,293 | -67% | 1 | 1 | 0% | 1,181 | 1,418 | +20% | 0 | 0 | — |
case-14 | fail→fail | 12,719 | 2,967 | -77% | 1 | 1 | 0% | 2,138 | 1,513 | -29% | 0 | 0 | — |
case-15 | pass→pass | 6,720 | 3,753 | -44% | 1 | 1 | 0% | 961 | 1,732 | +80% | 0 | 0 | — |
case-16 | fail→pass | 11,238 | 3,844 | -66% | 1 | 1 | 0% | 1,820 | 1,768 | -3% | 0 | 0 | — |
case-17 | fail→pass | 15,841 | 2,843 | -82% | 1 | 1 | 0% | 2,475 | 1,494 | -40% | 0 | 0 | — |
case-18 | fail→pass | 3,533 | 2,744 | -22% | 1 | 1 | 0% | 595 | 1,513 | +154% | 0 | 0 | — |
case-19 | pass→fail | 10,382 | 2,671 | -74% | 1 | 1 | 0% | 1,182 | 1,467 | +24% | 0 | 0 | — |
case-20 | pass→pass | 16,653 | 2,712 | -84% | 1 | 1 | 0% | 1,200 | 1,500 | +25% | 0 | 0 | — |
case-21 | fail→fail | 10,720 | 1,901 | -82% | 1 | 1 | 0% | 1,800 | 1,290 | -28% | 0 | 0 | — |
case-22 | pass→pass | 19,042 | 8,399 | -56% | 1 | 1 | 0% | 2,108 | 2,451 | +16% | 0 | 0 | — |
case-23 | pass→pass | 9,429 | 7,578 | -20% | 1 | 1 | 0% | 1,811 | 2,429 | +34% | 0 | 0 | — |
case-24 | pass→pass | 10,616 | 6,251 | -41% | 1 | 1 | 0% | 2,088 | 2,223 | +6% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 21 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +29 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/13/2026 | +23% |
Other measured skills in the registry, with their headline benchmark lift.