Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when drafting quarterly or annual goals for a team or organization: write one qualitative Objective backed by 3-5 Key Results, each a metric moving from an explicit baseline to a target and scored 0.0-1.0 at period end with ~0.7 counting as success — never a list of tasks. Do NOT use for sprint planning, project task breakdowns, milestone schedules, or performance reviews.
.claude/skills/okr-format/SKILL.md| Model | Lift | Δ tokens | Δ turns | Cases | Verified |
|---|---|---|---|---|---|
| gemini-3.6-flashbest | +32% | +46% | 0% | 22 | 53d ago |
| gemini-3.5-flash | pending re-run | — | |||
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | — | — |
| case-15 | ✗→✓ | ▲ Improved | — | — |
| case-13 | ✗→✓ | ▲ Improved | — | — |
| case-02 | ✗→✓ | ▲ Improved | — | — |
| case-19 | ✗→✓ | ▲ Improved | — | — |
Enforces the Objective-and-Key-Results structure on every set of team or organization goals you draft for a quarter or year. Apply whenever the task is to write, restructure, or clean up a team's period goals; not for sprint task planning, project milestone schedules, or individual performance reviews.
statement of where the team is going this period — memorable, motivating, and short.
in the Objective line; every number lives in a Key Result. "Make every delivery something customers can count on" — not "Increase weekly deliveries 30%".
more than five means priorities have not been chosen.
metric name, its current value, and the value to reach — a from→to pair, e.g. "Rebuffering ratio from 1.8% to 0.6%". If the ask supplies current figures, use them as the baselines; never leave a Key Result as a bare direction like "improve latency".
"release" describe work, not results. The test: if the team did the activity and nothing changed in the metric, would you still call it success? Rewrite every task into the result the work is supposed to move:
Launch the loyalty program → Repeat-purchase rate from 18% to 26% Migrate search to the new engine → Search-to-purchase rate from 9% to 14% Hire three field technicians → Median install wait from 12 days to 5 days
fractional progress toward its target. Landing around 0.7 on stretch targets counts as success; a routine 1.0 means the targets were sandbagged. Include a one-line scoring note with the goal set.
BEFORE (task list — the default)
Goals for the quarter:
- Migrate the player to the new CDN
- Reduce rebuffering
- Ship the offline mode beta
- Fix the top crash
AFTER (conforming)
Objective: Make playback feel instant and uninterrupted, every time.
Key Results:
1. Rebuffering ratio from 1.8% to 0.6%
2. Median video start time from 2.4s to 1.2s
3. Playback-failure sessions from 0.9% to 0.3%
4. Crash-free viewing sessions from 99.2% to 99.8%
Scoring: each Key Result is graded 0.0–1.0 at quarter end; landing around
0.7 on these stretch targets counts as success.What changed: the CDN migration and offline beta are work items, not results — they only appear insofar as they move a listed metric; "reduce rebuffering" gained a baseline and a target; the Objective became one qualitative, number-free line.
is unknown, make instrumenting it the first Key Result with its own numeric target date coverage (e.g. "Metric X instrumented and reporting for 100% of traffic"), then target it next period.
the adoption or quality outcome the launch must produce (users, conversion, error rate).
do not pool eight Key Results under one vague umbrella Objective.
measured form so partial credit is gradable on the 0.0–1.0 scale.
compliance), note that it is expected at 1.0; the ~0.7 norm applies to the stretch ones.
uptime".
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.5-flash | verified | 7/10/2026 | +100% |
Other measured skills in the registry, with their headline benchmark lift.