Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Validate, test, and score the quality of skills within the claude-skills ecosystem. Comprehensive meta-skill: structure validation, Python script testing (syntax + imports + runtime + output format), multi-dimensional quality scoring with letter grades and tier classification (BASIC/STANDARD/POWERFUL). Use when authoring a new skill, auditing existing skills for tier promotion, setting up pre-commit hooks for skill quality, or integrating skill QA into CI.
.claude/skills/alirezarezvani-skill-tester/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 15% | 0% |
Tier: POWERFUL · Category: Engineering Quality Assurance · Dependencies: None (Python stdlib only)
Meta-skill that validates, tests, and scores skills in this repository. Four tools, run from the repo root with full paths:
scripts/skill_validator.py — structure + documentation compliancescripts/script_tester.py — Python script syntax/imports/runtime/output testingscripts/quality_scorer.py — multi-dimensional scoring with letter gradescripts/security_scorer.py — security posture scoring (also available via quality_scorer.py --include-security)> Scope note: this skill's tier line-count minimums measure legacy skills. For authoring new skills, engineering/write-a-skill (SKILL.md under ~100 lines, Matt Pocock doctrine) is the binding standard — do not pad a new skill to satisfy a tier minimum here.
bash# 1. Validate structure (exit non-zero on failure — usable as a gate) python3 engineering/skills/skill-tester/scripts/skill_validator.py engineering/skills/self-eval --json # 2. Test the skill's Python scripts (30s default timeout per script) python3 engineering/skills/skill-tester/scripts/script_tester.py engineering/skills/self-eval --json # 3. Score quality (fail CI below threshold with --minimum-score) python3 engineering/skills/skill-tester/scripts/quality_scorer.py engineering/skills/self-eval --json --detailed --minimum-score 75
Consume the JSON: validator emits overall_score, compliance_level, per-check checks{}; scorer emits overall_score, letter_grade, tier_recommendation, dimensions, and an improvement_roadmap — work the roadmap top-down, then re-run until the target score is met.
For repo-wide auditing prefer scripts/audit_skills.py at the repo root (wraps the write-a-skill checklist runner across all skills).
--tier BASIC|STANDARD|POWERFUL)--timeout, default 30s)--help functionality verification; sample-data runs compared against expected_outputs/Four dimensions, 25% each: Documentation (depth, examples, references), Code Quality (complexity, error handling, output consistency), Completeness (required dirs, sample data, expected outputs), Usability (help text, example clarity). Outputs 0-100 + A-F grade + tier recommendation.
| Tier | SKILL.md | Scripts | CLI surface | |---|---|---|---| | BASIC | ≥ 100 lines | 1 (100-300 LOC) | basic argparse | | STANDARD | ≥ 200 lines | 1-2 (300-500 LOC) | subcommands, JSON + text output | | POWERFUL | ≥ 300 lines | 2-3 (500-800 LOC) | multiple modes, CI integration |
(Advisory for legacy skills; new skills follow write-a-skill — see scope note above.)
yaml# GitHub Actions: gate changed skills - name: "validate-changed-skills" run: | for skill in $changed_skills; do python3 engineering/skills/skill-tester/scripts/skill_validator.py "$skill" --json python3 engineering/skills/skill-tester/scripts/script_tester.py "$skill" python3 engineering/skills/skill-tester/scripts/quality_scorer.py "$skill" --minimum-score 75 done
Pre-commit hook: run the validator on the staged skill directory and block the commit on non-zero exit.
A skill "passes" when, in one run from repo root:
skill_validator.py <skill> --json exits 0,script_tester.py <skill> reports all scripts passing, andquality_scorer.py <skill> --minimum-score <target> exits 0.If any step fails, apply the top improvement_roadmap item and re-run all three — never report a partial pass.
--timeout or optimize the script under testReferences: references/ holds the structure specification, tier requirements matrix, and scoring rubric the tools implement.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 6,085 | 4,749 | -22% | 1 | 1 | 0% | 1,228 | 2,131 | +74% | 0 | 0 | — |
case-02 | fail→pass | 17,437 | 18,009 | +3% | 1 | 1 | 0% | 3,159 | 4,814 | +52% | 0 | 0 | — |
case-03 | fail→fail | 19,366 | 5,405 | -72% | 1 | 1 | 0% | 3,461 | 1,401 | -60% | 0 | 0 | — |
case-04 | fail→pass | 10,033 | 2,009 | -80% | 1 | 1 | 0% | 1,920 | 1,595 | -17% | 0 | 0 | — |
case-05 | pass→pass | 5,592 | 2,783 | -50% | 1 | 1 | 0% | 1,010 | 1,558 | +54% | 0 | 0 | — |
case-06 | fail→fail | 14,419 | 3,220 | -78% | 1 | 1 | 0% | 1,825 | 1,442 | -21% | 0 | 0 | — |
case-07 | fail→pass | 7,414 | 3,548 | -52% | 1 | 1 | 0% | 1,158 | 1,924 | +66% | 0 | 0 | — |
case-08 | pass→pass | 16,280 | 1,709 | -90% | 1 | 1 | 0% | 2,831 | 1,515 | -46% | 0 | 0 | — |
case-09 | pass→pass | 5,470 | 3,345 | -39% | 1 | 1 | 0% | 986 | 1,662 | +69% | 0 | 0 | — |
case-10 | fail→pass | 19,935 | 2,417 | -88% | 1 | 1 | 0% | 1,378 | 1,585 | +15% | 0 | 0 | — |
case-11 | pass→pass | 10,406 | 3,617 | -65% | 1 | 1 | 0% | 1,704 | 1,814 | +6% | 0 | 0 | — |
case-12 | fail→pass | 14,210 | 2,563 | -82% | 1 | 1 | 0% | 2,462 | 1,553 | -37% | 0 | 0 | — |
case-13 | fail→pass | 9,085 | 2,166 | -76% | 1 | 1 | 0% | 1,734 | 1,575 | -9% | 0 | 0 | — |
case-14 | fail→pass | 17,843 | 4,986 | -72% | 1 | 1 | 0% | 1,207 | 2,200 | +82% | 0 | 0 | — |
case-15 | pass→pass | 5,783 | 3,016 | -48% | 1 | 1 | 0% | 1,051 | 1,565 | +49% | 0 | 0 | — |
case-16 | pass→pass | 7,415 | 5,468 | -26% | 1 | 1 | 0% | 1,361 | 1,612 | +18% | 0 | 0 | — |
case-17 | pass→pass | 12,031 | 13,978 | +16% | 1 | 1 | 0% | 2,360 | 4,038 | +71% | 0 | 0 | — |
case-18 | fail→fail | 14,275 | 2,305 | -84% | 1 | 1 | 0% | 2,529 | 1,670 | -34% | 0 | 0 | — |
case-19 | fail→pass | 8,353 | 3,451 | -59% | 1 | 1 | 0% | 1,549 | 1,835 | +18% | 0 | 0 | — |
case-20 | fail→pass | 8,373 | 6,755 | -19% | 1 | 1 | 0% | 1,678 | 2,487 | +48% | 0 | 0 | — |
case-21 | pass→pass | 13,191 | 13,619 | +3% | 1 | 1 | 0% | 2,139 | 3,378 | +58% | 0 | 0 | — |
case-22 | pass→pass | 10,738 | 12,853 | +20% | 1 | 1 | 0% | 2,202 | 3,789 | +72% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.