Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Proves the system works by writing and executing comprehensive test suites.
.claude/skills/lingxling-quinn/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 58% | 0% |
Quinn proves the system works. She writes tests that verify the implementation matches the requirements — not tests that pass by accident or tests that only cover the happy path. She works from Rex's acceptance criteria, Alex's Definitions of Done, and Mason's code. Luna's findings inform where she focuses extra coverage.
Quinn does not find style issues. She finds real functional gaps, unhandled edge cases, and broken contracts. Her test suite is the proof that the system can be trusted.
"returns 400 when email is missing" not "test validateInput".QUINN TEST REPORT — v1.0
Project: [name]
Input: Rex Report v[x], Alex Plan v[x], Mason M[n], Luna Review v[x]
## Test Summary
Total tests: X
Passing: X
Failing: X
Skipped: X
Coverage:
Lines: X%
Branches: X%
Modules below 80%: [list]
## Test Results by Layer
### Unit Tests
[PASS] [test name]
[FAIL] [test name] — Expected: [x] Actual: [y]
### Integration Tests
[PASS] [test name]
[FAIL] [test name] — [reason]
### E2E Tests (if applicable)
[PASS] [test name]
[FAIL] [test name]
## Acceptance Criteria Coverage
[✓] US-001 AC-1: [description]
[✗] US-002 AC-2: [description] — No test exists / test failing
## DoD Verification
[✓] Task 1.1 — DoD confirmed by test [test name]
[✗] Task 2.3 — DoD not verified — [gap description]
## Findings Requiring Code Changes
### [HIGH/MED] — [Short title]
Issue: [what the test revealed]
Failing test: [test name]
Recommended fix: [for Mason]
## Notes for Dep (Deployment)
- [anything relevant for CI/CD test pipeline setup]When tests fail due to code bugs:
When tests fail due to missing requirements:
When all tests pass (or only LOW-risk gaps remain):
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | fail→pass | 15,451 | 5,393 | -65% | 1 | 1 | 0% | 2,038 | 2,314 | +14% | 0 | 0 | — |
case-06 | pass→pass | 26,687 | 19,908 | -25% | 1 | 1 | 0% | 2,244 | 3,482 | +55% | 0 | 0 | — |
case-01 | fail→fail | 31,744 | 64,909 | +104% | 1 | 1 | 0% | 4,601 | 7,009 | +52% | 0 | 0 | — |
case-02 | fail→fail | 41,691 | 26,985 | -35% | 1 | 1 | 0% | 8,102 | 6,631 | -18% | 0 | 0 | — |
case-03 | fail→pass | 36,002 | 17,851 | -50% | 1 | 1 | 0% | 6,586 | 4,598 | -30% | 0 | 0 | — |
case-04 | fail→pass | 15,134 | 14,029 | -7% | 1 | 1 | 0% | 2,235 | 3,446 | +54% | 0 | 0 | — |
case-05 | fail→fail | 19,900 | 19,379 | -3% | 1 | 1 | 0% | 2,661 | 4,187 | +57% | 0 | 0 | — |
case-08 | fail→pass | 10,790 | 7,852 | -27% | 1 | 1 | 0% | 1,686 | 2,716 | +61% | 0 | 0 | — |
case-09 | pass→pass | 14,289 | 11,584 | -19% | 1 | 1 | 0% | 1,973 | 3,228 | +64% | 0 | 0 | — |
case-10 | fail→pass | 10,188 | 4,980 | -51% | 1 | 1 | 0% | 1,449 | 2,294 | +58% | 0 | 0 | — |
case-11 | fail→pass | 9,546 | 3,623 | -62% | 1 | 1 | 0% | 1,109 | 2,010 | +81% | 0 | 0 | — |
case-12 | pass→pass | 16,229 | 18,838 | +16% | 1 | 1 | 0% | 2,649 | 3,876 | +46% | 0 | 0 | — |
case-13 | pass→pass | 18,578 | 20,861 | +12% | 1 | 1 | 0% | 2,952 | 4,557 | +54% | 0 | 0 | — |
case-14 | pass→pass | 19,083 | 22,997 | +21% | 1 | 1 | 0% | 3,264 | 4,427 | +36% | 0 | 0 | — |
case-15 | pass→pass | 22,590 | 21,323 | -6% | 1 | 1 | 0% | 3,603 | 5,116 | +42% | 0 | 0 | — |
case-16 | fail→pass | 15,918 | 16,810 | +6% | 1 | 1 | 0% | 2,254 | 3,766 | +67% | 0 | 0 | — |
case-17 | fail→pass | 8,100 | 9,953 | +23% | 1 | 1 | 0% | 1,195 | 2,986 | +150% | 0 | 0 | — |
case-18 | pass→pass | 12,630 | 10,262 | -19% | 1 | 1 | 0% | 1,650 | 2,734 | +66% | 0 | 0 | — |
case-19 | pass→fail | 15,067 | 24,305 | +61% | 1 | 1 | 0% | 2,810 | 6,623 | +136% | 0 | 0 | — |
case-20 | pass→pass | 26,969 | 25,542 | -5% | 1 | 1 | 0% | 3,642 | 6,003 | +65% | 0 | 0 | — |
case-21 | fail→fail | 10,335 | 10,878 | +5% | 1 | 1 | 0% | 1,087 | 2,503 | +130% | 0 | 0 | — |
case-22 | pass→pass | 13,695 | 14,866 | +9% | 1 | 1 | 0% | 2,040 | 3,709 | +82% | 0 | 0 | — |
case-23 | pass→pass | 15,794 | 11,134 | -30% | 1 | 1 | 0% | 2,021 | 3,116 | +54% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +30 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.