Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Model the cost and latency of an LLM feature before it ships and surprises the bill. Use when asked to estimate LLM API costs, set a latency/token budget, decide which model tier to use, or bring down the cost of an AI feature. Produces a cost & latency budget — token math per request, monthly cost projection, model tiering, caching/streaming levers, p95 latency targets, and a guardrail/alert plan.
.claude/skills/mohitagw15856-llm-cost-latency-budget/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 120% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 45% | 0% |
LLM features have a unit cost and a tail latency that demos hide and production exposes. This skill does the token math up front — what one request costs, what a million cost, where the p95 latency comes from — and lays out the levers (model tiering, caching, prompt trimming) so cost and speed are designed, not discovered.
Ask for these only if they aren't already provided:
1. Per-request token math — a table estimating tokens in/out per call, and the resulting cost at each candidate model's price.
| Component | Tokens | $ in | $ out | |---|---|---|---| | System prompt | | | | | Retrieved context | | | | | User input | | | | | Output | | | | | Per request | | $x | |
2. Monthly projection — per-request cost × volume, at current and target scale; the headline number leadership will ask for.
3. Model tiering — route easy requests to a cheaper/faster model and only escalate hard ones (cascade); show the blended cost. Often the single biggest saving.
4. Latency — where the p95 comes from (model TTFT + output length + retrieval + network), the target, and how streaming changes perceived latency even when total time is unchanged.
5. Cost levers — ranked by impact: prompt/context trimming, caching (prompt cache + response cache for repeats), shorter outputs (max_tokens), batching, tiering, and "do you need the model at all for this path."
6. Guardrails — per-user / per-day rate limits, a max-tokens cap, a spend alert threshold, and a kill switch — so a bug or abuse can't produce a surprise invoice.
LLM production cost/latency practice — token accounting, model cascades/tiering, prompt & response caching, and tail-latency budgeting.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 27,900 | 18,131 | -35% | 1 | 1 | 0% | 6,111 | 5,023 | -18% | 0 | 0 | — |
case-02 | pass→pass | 12,516 | 18,322 | +46% | 1 | 1 | 0% | 2,838 | 4,805 | +69% | 0 | 0 | — |
case-03 | pass→pass | 15,553 | 20,136 | +29% | 1 | 1 | 0% | 3,442 | 5,080 | +48% | 0 | 0 | — |
case-04 | fail→pass | 35,493 | 18,922 | -47% | 1 | 1 | 0% | 2,100 | 4,622 | +120% | 0 | 0 | — |
case-05 | fail→pass | 9,898 | 9,651 | -2% | 1 | 1 | 0% | 2,181 | 2,961 | +36% | 0 | 0 | — |
case-06 | pass→pass | 15,260 | 16,957 | +11% | 1 | 1 | 0% | 3,081 | 4,195 | +36% | 0 | 0 | — |
case-07 | fail→pass | 17,618 | 19,439 | +10% | 1 | 1 | 0% | 2,961 | 4,461 | +51% | 0 | 0 | — |
case-08 | pass→pass | 16,633 | 19,587 | +18% | 1 | 1 | 0% | 2,720 | 4,187 | +54% | 0 | 0 | — |
case-09 | pass→pass | 15,803 | 17,209 | +9% | 1 | 1 | 0% | 2,436 | 3,757 | +54% | 0 | 0 | — |
case-10 | pass→pass | 12,746 | 15,066 | +18% | 1 | 1 | 0% | 2,478 | 3,590 | +45% | 0 | 0 | — |
case-11 | pass→pass | 13,442 | 14,581 | +8% | 1 | 1 | 0% | 2,102 | 3,661 | +74% | 0 | 0 | — |
case-12 | pass→pass | 14,756 | 18,022 | +22% | 1 | 1 | 0% | 2,483 | 4,048 | +63% | 0 | 0 | — |
case-13 | pass→pass | 7,041 | 11,822 | +68% | 1 | 1 | 0% | 1,202 | 2,540 | +111% | 0 | 0 | — |
case-14 | pass→pass | 16,651 | 13,874 | -17% | 1 | 1 | 0% | 2,763 | 3,274 | +18% | 0 | 0 | — |
case-15 | pass→pass | 16,153 | 18,580 | +15% | 1 | 1 | 0% | 2,856 | 4,161 | +46% | 0 | 0 | — |
case-16 | pass→pass | 11,160 | 13,078 | +17% | 1 | 1 | 0% | 1,816 | 2,989 | +65% | 0 | 0 | — |
case-17 | pass→pass | 14,874 | 16,905 | +14% | 1 | 1 | 0% | 2,362 | 3,974 | +68% | 0 | 0 | — |
case-18 | fail→pass | 9,742 | 8,594 | -12% | 1 | 1 | 0% | 1,465 | 2,123 | +45% | 0 | 0 | — |
case-19 | pass→pass | 13,445 | 21,164 | +57% | 1 | 1 | 0% | 2,304 | 4,651 | +102% | 0 | 0 | — |
case-20 | pass→pass | 12,560 | 18,118 | +44% | 1 | 1 | 0% | 2,128 | 4,126 | +94% | 0 | 0 | — |
case-21 | pass→pass | 10,058 | 10,926 | +9% | 1 | 1 | 0% | 1,744 | 2,631 | +51% | 0 | 0 | — |
case-22 | fail→pass | 11,312 | 8,706 | -23% | 1 | 1 | 0% | 1,867 | 2,323 | +24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.