Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audit an executed AI phase's evaluation coverage and produce an EVAL-REVIEW.md remediation plan.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 186% | 0% |
| case-08 | ✓→✗ | ▼ Worse | -82% | 0% |
| case-18 | ✓→✗ | ▼ Worse | -85% | 0% |
| case-20 | ✓→✗ | ▼ Worse | -85% | 0% |
| case-21 | ✓→✗ | ▼ Worse | -90% | 0% |
<objective> Conduct a retroactive evaluation coverage audit of a completed AI phase. Checks whether the evaluation strategy from AI-SPEC.md was implemented. Produces EVAL-REVIEW.md with score, verdict, gaps, and remediation plan. </objective>
<execution_context> @~/.claude/gsd-core/workflows/eval-review.md @~/.claude/gsd-core/references/ai-evals.md </execution_context>
<context> Phase: $ARGUMENTS — optional, defaults to last completed phase. </context>
<process> Execute end-to-end. Preserve all workflow gates. </process>
Other measured skills in the registry, with their headline benchmark lift.