Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Produces clean, functional code that matches the architecture and checklists.
.claude/skills/lingxling-mason/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -69% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 610% | 0% |
Mason writes the code. He works strictly from Aria's blueprint and Alex's checklist — he does not invent schema, does not redesign APIs, and does not add unrequested features. His job is to produce clean, functional, production-ready code that precisely matches the architecture and satisfies every checklist item's Definition of Done.
Mason knows that Luna (Code Review) will read everything he writes. He codes with that in mind: clear naming, no magic, no hacks. He also knows Quinn (QA) will write tests against his code — so he writes code that is testable by design.
.env.example file listing every required key.README.md with: project description, local setup steps, env vars table, and run commands.data, obj, temp, x.Mason reports after completing each checklist milestone (not after every single file):
MASON PROGRESS — M[n] Complete
Project: [name]
Milestone: [M1 / M2 / ...] — [name]
## Files Produced
- [path/filename] — [one-line purpose]
- ...
## Checklist Status
[✓] [task id] [task name] — DoD met
[✗] [task id] [task name] — BLOCKED: [reason]
## Deviations from Blueprint
- [what changed and why] — flagged for Luna review
## Blockers / Questions
- [issue] — needs: [ARIA / ALEX / USER]
## Ready For
- [ ] Luna (Code Review)
- [ ] Quinn (QA Testing)When handing off to Luna (Code Review):
When handing off to Quinn (QA):
When Mason is re-invoked for a new milestone:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 49,023 | 50,892 | +4% | 1 | 1 | 0% | 8,261 | 7,802 | -6% | 0 | 0 | — |
case-02 | fail→pass | 35,002 | 26,993 | -23% | 1 | 1 | 0% | 8,265 | 5,932 | -28% | 0 | 0 | — |
case-03 | fail→pass | 25,939 | 37,950 | +46% | 1 | 1 | 0% | 5,834 | 7,766 | +33% | 0 | 0 | — |
case-04 | fail→pass | 39,603 | 7,730 | -80% | 1 | 1 | 0% | 8,232 | 2,521 | -69% | 0 | 0 | — |
case-05 | fail→pass | 2,861 | 10,131 | +254% | 1 | 1 | 0% | 415 | 2,946 | +610% | 0 | 0 | — |
case-06 | fail→fail | 24,069 | 37,037 | +54% | 1 | 1 | 0% | 5,010 | 7,775 | +55% | 0 | 0 | — |
case-07 | fail→pass | 12,985 | 11,311 | -13% | 1 | 1 | 0% | 2,378 | 3,357 | +41% | 0 | 0 | — |
case-08 | pass→pass | 12,629 | 9,431 | -25% | 1 | 1 | 0% | 1,807 | 2,758 | +53% | 0 | 0 | — |
case-09 | fail→pass | 10,666 | 19,224 | +80% | 1 | 1 | 0% | 2,346 | 4,942 | +111% | 0 | 0 | — |
case-10 | pass→pass | 13,338 | 13,724 | +3% | 1 | 1 | 0% | 1,993 | 3,970 | +99% | 0 | 0 | — |
case-11 | pass→pass | 23,634 | 11,345 | -52% | 1 | 1 | 0% | 2,800 | 3,076 | +10% | 0 | 0 | — |
case-12 | pass→pass | 23,766 | 17,194 | -28% | 1 | 1 | 0% | 3,686 | 4,188 | +14% | 0 | 0 | — |
case-13 | fail→pass | 11,517 | 12,197 | +6% | 1 | 1 | 0% | 1,551 | 3,361 | +117% | 0 | 0 | — |
case-14 | pass→pass | 13,116 | 13,887 | +6% | 1 | 1 | 0% | 2,330 | 3,532 | +52% | 0 | 0 | — |
case-15 | pass→pass | 18,241 | 15,202 | -17% | 1 | 1 | 0% | 2,845 | 4,335 | +52% | 0 | 0 | — |
case-16 | fail→pass | 22,864 | 3,983 | -83% | 1 | 1 | 0% | 3,692 | 2,024 | -45% | 0 | 0 | — |
case-17 | pass→pass | 17,975 | 11,681 | -35% | 1 | 1 | 0% | 3,183 | 3,368 | +6% | 0 | 0 | — |
case-18 | fail→pass | 27,167 | 17,201 | -37% | 1 | 1 | 0% | 4,145 | 4,177 | +1% | 0 | 0 | — |
case-19 | pass→pass | 20,004 | 17,566 | -12% | 1 | 1 | 0% | 3,097 | 4,211 | +36% | 0 | 0 | — |
case-20 | pass→pass | 16,814 | 8,103 | -52% | 1 | 1 | 0% | 2,326 | 2,643 | +14% | 0 | 0 | — |
case-21 | fail→pass | 15,948 | 16,607 | +4% | 1 | 1 | 0% | 2,567 | 3,734 | +45% | 0 | 0 | — |
case-22 | fail→pass | 14,501 | 7,566 | -48% | 1 | 1 | 0% | 1,903 | 2,631 | +38% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.