Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when writing or rewriting goal statements for a team, project, department, or review period: make every goal SMART — one metric with its number and unit plus an explicit date or period stated inside the goal sentence itself — which cheaper models do not do by default. Do NOT use for OKR planning with objectives and scored key results, or for mission and vision statements.
.claude/skills/smart-goals/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 20 |
| gemini-3.1-pro-preview | 100% | 2 |
| Model | Lift | Δ tokens | Δ turns | Cases | Verified |
|---|---|---|---|---|---|
| gemini-3.6-flashbest | +40% | +69% | 0% | 25 | 54d ago |
| gemini-3.5-flash | pending re-run | — | |||
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-25 | ✗→✓ | ▲ Improved | — | — |
| case-19 | ✗→✓ | ▲ Improved | — | — |
| case-08 | ✗→✓ | ▲ Improved | — | — |
| case-12 | ✗→✓ | ▲ Improved | — | — |
| case-18 | ✗→✓ | ▲ Improved | — | — |
Enforces the SMART criteria on every goal statement produced: each goal is a single sentence that names one specific focus, carries a numeric target with its unit, and states an explicit date or period — Specific, Measurable, Achievable, Relevant, Time-bound. Applies to goal lists for teams, projects, and review periods; not to mission statements or objective/key-result planning structures.
abandonment in mobile checkout", never "get better at e-commerce".
currency amount, duration, or rating. Where the current value is known, express the move as baseline → target ("from 12% to 8%") rather than a bare delta.
plausible, not a fantasy multiple. Moving a metric 10–30% past its baseline is a typical ambitious band; an overnight 10× usually is not.
which priority it advances, cut it rather than keep it for symmetry.
30 September", "within Q2", "before the holiday season". "Soon", "ongoing", and "ASAP" do not qualify.
half-goals ("raise revenue and cut costs") hide which one moved.
"boost" are admissible only when a quantified target rides in the same sentence.
pushed into a separate KPI appendix, footnote, or "success criteria" block that the sentence merely gestures at.
Vague aspiration → complete goal.
BEFORE Improve weekend sales at the bakery.
AFTER Grow weekend pastry revenue from $2,100 to $2,600 per weekend by 30 November.The rewrite names the focus (weekend pastry revenue), the number and unit ($2,600 per weekend, from a $2,100 baseline), and the date (30 November) — all in one sentence.
Direction without a quantity → baseline and target.
BEFORE Get more residents recycling this year.
AFTER Raise curbside recycling participation from 41% to 55% of households by 31 December.Two aims welded together → two goals.
BEFORE Help club members get race-ready and grow the membership.
AFTER 1. Have 25 of the club's 60 members finish an official 10K race by 31 May.
2. Grow paid club membership from 60 to 80 members by 31 August.Vague verb dressed up with detail → quantified verb.
BEFORE Substantially enhance the quality of our member training plans over the coming months.
AFTER Cut the share of members reporting a training injury from 18% to under 10% by 1 July.Length and formality did not fix the BEFORE — only the number, unit, and date did.
score, a count of mentions, a retention figure. "Lift the staff survey engagement score from 6.8 to 7.5 by 15 December."
measure first. "Reach 500 newsletter subscribers by 31 March" works without a baseline.
quality threshold if bare completion could be gamed: "Launch the online store by 1 October with at least 40 products listed."
rather than one distant target no one can steer by.
or it fails the achievability test no matter the number.
table the sentence points at.
with no base stated.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-25 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-24 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-23 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +40 percentage points is the difference between those two pass rates over the 25 comparable cases.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.5-flash | verified | 7/10/2026 | +92% |
Other measured skills in the registry, with their headline benchmark lift.