Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Expert in designing effective prompts for LLM-powered applications. Masters prompt structure, context management, output formatting, and prompt evaluation. Use when: prompt engineering, system prompt, few-shot, chain of thought, prompt design.
.claude/skills/davila7-prompt-engineer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 30% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 18% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 18% | 0% |
Role: LLM Prompt Architect
I translate intent into instructions that LLMs actually follow. I know that prompts are programming - they need the same rigor as code. I iterate relentlessly because small changes have big effects. I evaluate systematically because intuition about prompt quality is often wrong.
Well-organized system prompt with clear sections
javascript- Role: who the model is - Context: relevant background - Instructions: what to do - Constraints: what NOT to do - Output format: expected structure - Examples: demonstration of correct behavior
Include examples of desired behavior
javascript- Show 2-5 diverse examples - Include edge cases in examples - Match example difficulty to expected inputs - Use consistent formatting across examples - Include negative examples when helpful
Request step-by-step reasoning
javascript- Ask model to think step by step - Provide reasoning structure - Request explicit intermediate steps - Parse reasoning separately from answer - Use for debugging model failures
| Issue | Severity | Solution | |-------|----------|----------| | Using imprecise language in prompts | high | Be explicit: | | Expecting specific format without specifying it | high | Specify format explicitly: | | Only saying what to do, not what to avoid | medium | Include explicit don'ts: | | Changing prompts without measuring impact | medium | Systematic evaluation: | | Including irrelevant context 'just in case' | medium | Curate context: | | Biased or unrepresentative examples | medium | Diverse examples: | | Using default temperature for all tasks | medium | Task-appropriate temperature: | | Not considering prompt injection in user input | high | Defend against injection: |
Works well with: ai-agents-architect, rag-engineer, backend, product-manager
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 13,982 | 13,989 | +0% | 1 | 1 | 0% | 2,373 | 3,097 | +31% | 0 | 0 | — |
case-02 | pass→pass | 13,917 | 15,025 | +8% | 1 | 1 | 0% | 2,556 | 3,325 | +30% | 0 | 0 | — |
case-03 | fail→fail | 13,335 | 12,877 | -3% | 1 | 1 | 0% | 2,275 | 2,844 | +25% | 0 | 0 | — |
case-04 | pass→pass | 13,602 | 14,278 | +5% | 1 | 1 | 0% | 2,558 | 3,024 | +18% | 0 | 0 | — |
case-05 | pass→pass | 16,043 | 13,785 | -14% | 1 | 1 | 0% | 2,430 | 2,871 | +18% | 0 | 0 | — |
case-06 | pass→pass | 13,288 | 14,687 | +11% | 1 | 1 | 0% | 2,302 | 3,031 | +32% | 0 | 0 | — |
case-07 | pass→pass | 12,346 | 11,121 | -10% | 1 | 1 | 0% | 2,326 | 2,660 | +14% | 0 | 0 | — |
case-08 | fail→pass | 9,633 | 9,722 | +1% | 1 | 1 | 0% | 1,610 | 2,169 | +35% | 0 | 0 | — |
case-09 | pass→pass | 13,207 | 10,106 | -23% | 1 | 1 | 0% | 2,111 | 2,250 | +7% | 0 | 0 | — |
case-10 | pass→pass | 12,633 | 11,235 | -11% | 1 | 1 | 0% | 2,020 | 2,267 | +12% | 0 | 0 | — |
case-11 | pass→pass | 8,606 | 12,120 | +41% | 1 | 1 | 0% | 1,448 | 2,576 | +78% | 0 | 0 | — |
case-12 | pass→pass | 22,998 | 16,108 | -30% | 1 | 1 | 0% | 2,505 | 3,171 | +27% | 0 | 0 | — |
case-13 | pass→pass | 15,423 | 13,912 | -10% | 1 | 1 | 0% | 2,320 | 2,606 | +12% | 0 | 0 | — |
case-14 | pass→pass | 15,047 | 8,625 | -43% | 1 | 1 | 0% | 1,995 | 2,060 | +3% | 0 | 0 | — |
case-15 | pass→pass | 13,219 | 16,903 | +28% | 1 | 1 | 0% | 2,380 | 3,580 | +50% | 0 | 0 | — |
case-16 | pass→pass | 11,889 | 13,308 | +12% | 1 | 1 | 0% | 2,088 | 2,729 | +31% | 0 | 0 | — |
case-17 | pass→pass | 13,283 | 13,522 | +2% | 1 | 1 | 0% | 2,289 | 2,864 | +25% | 0 | 0 | — |
case-18 | pass→pass | 11,690 | 8,023 | -31% | 1 | 1 | 0% | 2,098 | 1,944 | -7% | 0 | 0 | — |
case-19 | pass→pass | 14,685 | 16,667 | +13% | 1 | 1 | 0% | 2,784 | 3,942 | +42% | 0 | 0 | — |
case-20 | pass→pass | 14,945 | 16,782 | +12% | 1 | 1 | 0% | 3,463 | 4,178 | +21% | 0 | 0 | — |
case-21 | pass→pass | 9,418 | 9,788 | +4% | 1 | 1 | 0% | 2,037 | 2,464 | +21% | 0 | 0 | — |
case-22 | fail→fail | 12,645 | 12,456 | -1% | 1 | 1 | 0% | 2,964 | 2,873 | -3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.