Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the user asks to optimize prompts, design prompt templates, evaluate LLM outputs with an eval set, measure RAG retrieval quality, validate agent/tool configurations, analyze token usage, or design structured-output contracts. Covers eval-driven prompt iteration, RAG metrics (relevance, faithfulness, coverage), agent workflow validation, and token/cost budgeting — all model-agnostic, with three stdlib Python tools.
.claude/skills/alirezarezvani-senior-prompt-engineer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 161% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 128% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 149% | 0% |
Eval-driven prompt engineering, RAG quality measurement, and agent workflow validation. Everything here is model-agnostic by design: techniques are framed by what they do, not by which model generation they were observed on, and the tools never hardcode model IDs or pricing — you supply your provider's current rates when you want dollar figures.
--analyze --output baseline.json), then compare every iteration against it.--price-per-mtok (never trust a cached price table — including any you remember).scripts/prompt_optimizer.pyStatic analysis: token estimate, clarity/structure scores (0–100), ambiguity + redundancy detection, few-shot example extraction.
bash# Full analysis (human-readable report) python3 scripts/prompt_optimizer.py prompt.txt --analyze # Save machine-readable baseline for later comparison python3 scripts/prompt_optimizer.py prompt.txt --analyze --json --output baseline.json # Token estimate; cost only if you supply your provider's current rate python3 scripts/prompt_optimizer.py prompt.txt --tokens --model claude --price-per-mtok 3.00 # Whitespace/redundancy-trimmed version python3 scripts/prompt_optimizer.py prompt.txt --optimize --output optimized.txt # Extract Input/Output few-shot pairs to JSON python3 scripts/prompt_optimizer.py prompt.txt --extract-examples --output examples.json # Compare a revision against the saved baseline python3 scripts/prompt_optimizer.py optimized.txt --analyze --compare baseline.json
--model accepts any string; only the tokenizer family is inferred (names containing "claude" → 3.5 chars/token, otherwise 4.0). Exit 0 on success, 1 on missing file.
scripts/rag_evaluator.pyMeasures retrieval and grounding quality from two JSON files (formats printed in --help).
bashpython3 scripts/rag_evaluator.py --contexts retrieved.json --questions eval_set.json python3 scripts/rag_evaluator.py --contexts ctx.json --questions q.json --k 10 --json python3 scripts/rag_evaluator.py --contexts ctx.json --questions q.json --output report.json --verbose python3 scripts/rag_evaluator.py --contexts ctx.json --questions q.json --compare baseline_report.json
Reports context relevance, precision@k, coverage, answer faithfulness, groundedness. Treat relevance < 0.80 as a retrieval problem (chunking/embedding/filtering), not a prompt problem — fix retrieval before rewriting the generation prompt.
scripts/agent_orchestrator.pyValidates agent configs (YAML/JSON): tool wiring, missing required config, loop risk, token estimates.
bashpython3 scripts/agent_orchestrator.py agent.yaml --validate python3 scripts/agent_orchestrator.py agent.yaml --visualize --format mermaid python3 scripts/agent_orchestrator.py agent.yaml --estimate-cost --runs 100 \ --input-price-per-mtok 3.00 --output-price-per-mtok 15.00
Without the two price flags, --estimate-cost reports token estimates only. The model: field in the config is informational — any model name is accepted.
python3 scripts/prompt_optimizer.py current_prompt.txt --analyze --json --output baseline.json| Symptom | Fix | |---------|-----| | Malformed/unparseable output | Native structured outputs / JSON schema if the API supports it; explicit schema-in-prompt otherwise | | Inconsistent answers across runs | Tighten instructions + add 2–3 contrastive examples (one near-miss showing what NOT to do) | | Misses edge cases | Enumerate the edge cases explicitly; add a "when uncertain, do X" rule | | Token bloat on repeated calls | Move stable prefix (system rules, examples) first so prompt caching applies; trim redundancy | | Wrong reasoning on hard cases | Ask for stepwise reasoning in a scratch field the consumer ignores, or use the provider's extended-thinking mode |
python3 scripts/prompt_optimizer.py revised.txt --analyze --compare baseline.jsoneval_results.json, then assert:bash python3 scripts/prompt_optimizer.py revised.txt --analyze --json --output revised.json \ && python3 -c " import json, sys r = json.load(open('revised.json')); b = json.load(open('baseline.json')) ok = r['clarity_score'] >= b['clarity_score'] and r['token_count'] <= b['token_count'] * 1.10 sys.exit(0 if ok else 1)" echo "gate exit=$?" # 0 = ship; 1 = regression, iterate again Pair this structural gate with your task-level eval: the revision must not lose any previously-passing eval case (no-regression rule).
python3 scripts/prompt_optimizer.py prompt_with_examples.txt --extract-examples --output examples.json and inspect that every extracted pair parses against your schema.python3 -c "import json,sys; [json.loads(l) for l in sys.stdin]" at minimum); 10/10 must parse, else return to step 2.questions.json (id, question, reference answer) and capture current retrievals to contexts.json.python3 scripts/rag_evaluator.py --contexts contexts.json --questions questions.json --output rag_baseline.jsonpython3 scripts/rag_evaluator.py --contexts new_contexts.json --questions questions.json --compare rag_baseline.json — every metric must be ≥ baseline; any regression blocks the change.python3 scripts/agent_orchestrator.py agent.yaml --validate — must exit with VALIDATION PASSED; fix every error and warning (missing tool config, unbounded iterations, loop risk).--estimate-cost --runs N with your current prices; if cost/run exceeds budget, cut tools or context before downgrading the model.| File | Contains | Load when user asks about | |------|----------|---------------------------| | references/prompt_engineering_patterns.md | 10 prompt patterns with input/output examples | "which pattern?", few-shot design, decomposition, meta-prompting | | references/llm_evaluation_frameworks.md | Eval metrics, scoring methods, A/B testing | "how to evaluate?", "measure quality", "compare prompts" | | references/agentic_system_design.md | Agent architectures (ReAct, Plan-Execute, Tool Use) | "build agent", "tool calling", "multi-agent" |
engineering-team/skills/senior-ml-engineer — model deployment and serving (this skill stops at the prompt/eval layer)engineering/rag-architect — RAG system architecture (this skill measures RAG quality; that one designs the pipeline)engineering/agent-designer — full agent system design (this skill validates configs; that one designs the architecture)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 11,450 | 4,793 | -58% | 1 | 1 | 0% | 2,457 | 2,562 | +4% | 0 | 0 | — |
case-02 | fail→pass | 9,682 | 10,990 | +14% | 1 | 1 | 0% | 1,728 | 4,505 | +161% | 0 | 0 | — |
case-03 | fail→fail | 18,534 | 5,351 | -71% | 1 | 1 | 0% | 3,585 | 2,662 | -26% | 0 | 0 | — |
case-04 | fail→fail | 15,574 | 16,040 | +3% | 1 | 1 | 0% | 3,194 | 5,487 | +72% | 0 | 0 | — |
case-05 | pass→pass | 12,623 | 13,031 | +3% | 1 | 1 | 0% | 2,323 | 4,685 | +102% | 0 | 0 | — |
case-06 | pass→pass | 16,189 | 12,754 | -21% | 1 | 1 | 0% | 3,189 | 4,732 | +48% | 0 | 0 | — |
case-07 | fail→pass | 11,162 | 5,790 | -48% | 1 | 1 | 0% | 1,790 | 3,449 | +93% | 0 | 0 | — |
case-08 | fail→pass | 8,533 | 5,964 | -30% | 1 | 1 | 0% | 1,439 | 3,277 | +128% | 0 | 0 | — |
case-09 | pass→pass | 9,786 | 7,032 | -28% | 1 | 1 | 0% | 1,643 | 3,521 | +114% | 0 | 0 | — |
case-10 | fail→pass | 12,955 | 8,770 | -32% | 1 | 1 | 0% | 2,214 | 3,828 | +73% | 0 | 0 | — |
case-11 | fail→pass | 6,886 | 5,088 | -26% | 1 | 1 | 0% | 1,348 | 3,351 | +149% | 0 | 0 | — |
case-12 | fail→pass | 9,231 | 2,392 | -74% | 1 | 1 | 0% | 1,763 | 2,797 | +59% | 0 | 0 | — |
case-13 | pass→pass | 8,709 | 2,248 | -74% | 1 | 1 | 0% | 1,519 | 2,704 | +78% | 0 | 0 | — |
case-14 | fail→pass | 11,786 | 3,475 | -71% | 1 | 1 | 0% | 2,325 | 2,977 | +28% | 0 | 0 | — |
case-15 | pass→pass | 10,457 | 5,839 | -44% | 1 | 1 | 0% | 1,740 | 3,347 | +92% | 0 | 0 | — |
case-16 | fail→pass | 3,759 | 1,759 | -53% | 1 | 1 | 0% | 614 | 2,692 | +338% | 0 | 0 | — |
case-17 | fail→fail | 13,417 | 10,213 | -24% | 1 | 1 | 0% | 2,067 | 4,042 | +96% | 0 | 0 | — |
case-18 | fail→pass | 11,545 | 7,844 | -32% | 1 | 1 | 0% | 1,842 | 3,547 | +93% | 0 | 0 | — |
case-19 | pass→pass | 14,341 | 9,342 | -35% | 1 | 1 | 0% | 2,393 | 4,023 | +68% | 0 | 0 | — |
case-20 | pass→pass | 14,530 | 10,867 | -25% | 1 | 1 | 0% | 2,603 | 4,228 | +62% | 0 | 0 | — |
case-21 | fail→pass | 12,143 | 7,846 | -35% | 1 | 1 | 0% | 1,982 | 3,746 | +89% | 0 | 0 | — |
case-22 | pass→pass | 12,402 | 6,041 | -51% | 1 | 1 | 0% | 1,900 | 3,365 | +77% | 0 | 0 | — |
case-23 | fail→pass | 10,633 | 5,086 | -52% | 1 | 1 | 0% | 1,794 | 3,194 | +78% | 0 | 0 | — |
case-24 | fail→pass | 9,556 | 2,020 | -79% | 1 | 1 | 0% | 1,620 | 2,637 | +63% | 0 | 0 | — |
case-25 | fail→pass | 7,779 | 2,879 | -63% | 1 | 1 | 0% | 1,399 | 2,841 | +103% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 23 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +52 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.