Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Quality assurance review — 6-dimension checklist covering template compliance, factual accuracy, completeness, track changes, research audit, and instruction cross-reference.
.claude/skills/anylegal-ai-qa/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✓→✗ | ▼ Worse | 178% | 0% |
| case-14 | ✓→✗ | ▼ Worse | -33% | 0% |
| case-21 | ✓→✗ | ▼ Worse | -35% | 0% |
| case-22 | ✓→✗ | ▼ Worse | -55% | 0% |
| case-09 | ✓→✓ | = Same ✓ | 29% | 0% |
Use this skill when:
Call list_documents to see all files. Identify:
intake_*.md) for client instructionsPlaybook/*.md, anylegal.md)Read the deliverable with read_document. Check:
Verify legal facts in the document:
web_search to spot-checkRead the intake summary and compare against the deliverable:
Use run_code (default Python) with lxml to count tracked changes (w:ins, w:del), extract authors, and assess edit scope. Then:
If research notes or memos exist:
[[N]](URL) format with working URLsUse web_search and web_fetch to spot-check 2-3 key citations.
Create a qa_<date>.md document with:
markdown# QA Report — [Matter Title] **Date:** YYYY-MM-DD **Reviewer:** Agent QA **Document:** [deliverable filename] ## Summary [1-2 sentence overall assessment] ## Checklist ### Template Compliance - [x] Structure matches template - [x] Numbering correct ... ### Factual Accuracy - [x] Jurisdictions correct ... ### Completeness - [x] All client instructions addressed ... ### Track Changes - [x] All edits justified ... ### Research Audit - [x] Citations verified ... ## Issues Found [List any issues, with severity: CRITICAL / WARNING / INFO] ## Verdict [PASS / PASS WITH WARNINGS / FAIL]
compare to verify the edit scope makes sense| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 34,972 | 3,597 | -90% | 1 | 1 | 0% | 5,936 | 1,216 | -80% | 0 | 0 | — |
case-02 | fail→fail | 13,087 | 2,339 | -82% | 1 | 1 | 0% | 1,960 | 1,118 | -43% | 0 | 0 | — |
case-03 | fail→fail | 11,501 | 3,362 | -71% | 1 | 1 | 0% | 1,972 | 1,252 | -37% | 0 | 0 | — |
case-04 | fail→fail | 8,529 | 3,024 | -65% | 1 | 1 | 0% | 1,344 | 1,122 | -17% | 0 | 0 | — |
case-05 | fail→fail | 3,083 | 4,411 | +43% | 1 | 1 | 0% | 446 | 1,164 | +161% | 0 | 0 | — |
case-06 | fail→fail | 12,239 | 2,624 | -79% | 1 | 1 | 0% | 1,962 | 1,126 | -43% | 0 | 0 | — |
case-07 | fail→fail | 6,308 | 3,012 | -52% | 1 | 1 | 0% | 970 | 1,113 | +15% | 0 | 0 | — |
case-08 | pass→fail | 2,720 | 3,453 | +27% | 1 | 1 | 0% | 415 | 1,154 | +178% | 0 | 0 | — |
case-09 | pass→pass | 10,129 | 6,629 | -35% | 1 | 1 | 0% | 1,556 | 2,009 | +29% | 0 | 0 | — |
case-10 | pass→pass | 6,231 | 6,512 | +5% | 1 | 1 | 0% | 1,006 | 1,821 | +81% | 0 | 0 | — |
case-11 | fail→fail | 3,757 | 3,472 | -8% | 1 | 1 | 0% | 563 | 1,182 | +110% | 0 | 0 | — |
case-12 | fail→fail | 7,608 | 3,282 | -57% | 1 | 1 | 0% | 1,154 | 1,081 | -6% | 0 | 0 | — |
case-13 | fail→fail | 8,285 | 3,342 | -60% | 1 | 1 | 0% | 1,359 | 1,144 | -16% | 0 | 0 | — |
case-14 | pass→fail | 9,375 | 2,869 | -69% | 1 | 1 | 0% | 1,656 | 1,117 | -33% | 0 | 0 | — |
case-15 | fail→fail | 3,132 | 3,313 | +6% | 1 | 1 | 0% | 153 | 1,142 | +646% | 0 | 0 | — |
case-16 | fail→fail | 3,198 | 3,935 | +23% | 1 | 1 | 0% | 344 | 1,256 | +265% | 0 | 0 | — |
case-17 | fail→fail | 5,386 | 2,447 | -55% | 1 | 1 | 0% | 832 | 1,078 | +30% | 0 | 0 | — |
case-18 | fail→fail | 7,087 | 3,375 | -52% | 1 | 1 | 0% | 1,063 | 1,149 | +8% | 0 | 0 | — |
case-19 | pass→pass | 11,487 | 8,176 | -29% | 1 | 1 | 0% | 1,794 | 1,914 | +7% | 0 | 0 | — |
case-20 | pass→pass | 15,113 | 22,437 | +48% | 1 | 1 | 0% | 2,721 | 4,034 | +48% | 0 | 0 | — |
case-21 | pass→fail | 11,143 | 4,302 | -61% | 1 | 1 | 0% | 1,802 | 1,173 | -35% | 0 | 0 | — |
case-22 | pass→fail | 14,795 | 2,885 | -81% | 1 | 1 | 0% | 2,469 | 1,119 | -55% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 4 counted toward the lift figure. The other 18 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. A headline lift is not published for this run.
Other measured skills in the registry, with their headline benchmark lift.