Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transform a user's 'basic prompt' into an 'optimized prompt' by applying context engineering -> the right context beats the right prompt.
.claude/skills/strativd-a-prompt-optimizer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 83% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 22% | 0% |
| case-11 | ✓→✗ | ▼ Worse | -31% | 0% |
You are a Prompt Optimizer. Your job is to transform a user’s “basic prompt” into an “optimized prompt” by applying context engineering: the right context beats the right prompt.
Core principle (4W): ensure the final prompt includes:
A) OPTIMIZED PROMPT (copy/paste-ready) using this structure:
markdown ROLE OBJECTIVE CONTEXT PACKAGE (Audience, Voice/Tone, Length target, Must-use inputs, Constraints/Boundaries) WORKFLOW (Gap check → Plan → Draft → Review → Revise) OUTPUT FORMAT FIRST ACTION
B) OPTIONAL: A SHORT CHANGELOG (bullets) explaining what you added and why (no long essay).
Start with Step 0: GAP CHECK. However, if the user has not provided a general prompt then ask for the prompt first.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,084 | 5,727 | -19% | 1 | 1 | 0% | 1,301 | 1,592 | +22% | 0 | 0 | — |
case-02 | fail→pass | 12,502 | 16,177 | +29% | 1 | 1 | 0% | 2,220 | 3,314 | +49% | 0 | 0 | — |
case-03 | fail→fail | 8,196 | 6,123 | -25% | 1 | 1 | 0% | 1,516 | 1,576 | +4% | 0 | 0 | — |
case-04 | pass→pass | 7,178 | 3,642 | -49% | 1 | 1 | 0% | 1,252 | 1,163 | -7% | 0 | 0 | — |
case-05 | fail→pass | 6,170 | 9,684 | +57% | 1 | 1 | 0% | 1,047 | 2,276 | +117% | 0 | 0 | — |
case-06 | fail→fail | 8,497 | 5,203 | -39% | 1 | 1 | 0% | 1,781 | 1,645 | -8% | 0 | 0 | — |
case-07 | pass→fail | 7,953 | 31,950 | +302% | 1 | 1 | 0% | 1,376 | 1,684 | +22% | 0 | 0 | — |
case-08 | pass→pass | 6,746 | 10,819 | +60% | 1 | 1 | 0% | 1,167 | 2,541 | +118% | 0 | 0 | — |
case-09 | fail→fail | 10,795 | 5,237 | -51% | 1 | 1 | 0% | 1,752 | 1,436 | -18% | 0 | 0 | — |
case-10 | pass→pass | 9,667 | 5,216 | -46% | 1 | 1 | 0% | 1,723 | 1,477 | -14% | 0 | 0 | — |
case-11 | pass→fail | 12,621 | 5,840 | -54% | 1 | 1 | 0% | 2,188 | 1,514 | -31% | 0 | 0 | — |
case-12 | pass→fail | 10,746 | 6,424 | -40% | 1 | 1 | 0% | 1,946 | 1,670 | -14% | 0 | 0 | — |
case-13 | pass→pass | 10,309 | 14,928 | +45% | 1 | 1 | 0% | 1,769 | 3,193 | +80% | 0 | 0 | — |
case-14 | fail→pass | 9,585 | 15,380 | +60% | 1 | 1 | 0% | 1,739 | 3,188 | +83% | 0 | 0 | — |
case-15 | fail→fail | 6,736 | 5,379 | -20% | 1 | 1 | 0% | 1,131 | 1,456 | +29% | 0 | 0 | — |
case-16 | fail→fail | 9,852 | 5,539 | -44% | 1 | 1 | 0% | 1,750 | 1,509 | -14% | 0 | 0 | — |
case-17 | fail→fail | 10,542 | 6,632 | -37% | 1 | 1 | 0% | 1,607 | 1,700 | +6% | 0 | 0 | — |
case-18 | fail→fail | 11,575 | 4,833 | -58% | 1 | 1 | 0% | 2,121 | 1,387 | -35% | 0 | 0 | — |
case-19 | pass→fail | 10,823 | 5,925 | -45% | 1 | 1 | 0% | 1,843 | 1,578 | -14% | 0 | 0 | — |
case-20 | pass→fail | 13,218 | 4,659 | -65% | 1 | 1 | 0% | 2,215 | 1,262 | -43% | 0 | 0 | — |
case-21 | pass→fail | 8,004 | 8,924 | +11% | 1 | 1 | 0% | 1,596 | 2,395 | +50% | 0 | 0 | — |
case-22 | pass→fail | 17,465 | 15,379 | -12% | 1 | 1 | 0% | 1,891 | 2,584 | +37% | 0 | 0 | — |
case-23 | fail→fail | 15,584 | 5,325 | -66% | 1 | 1 | 0% | 3,055 | 1,394 | -54% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of -50 percentage points is the difference between those two pass rates over the 23 comparable cases. 11 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.