Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring. Use when analyzing Claude API billing, reducing costs, or implementing cost controls and budget alerts. Trigger with phrases like "anthropic cost", "claude billing", "reduce claude spend", "anthropic budget", "claude pricing optimize".
.claude/skills/jeremylongshore-anth-cost-tuning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 175% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 312% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 162% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 147% | 0% |
Optimize Claude API spend through model routing, prompt caching, the Message Batches API, and real-time cost tracking. The four biggest levers: model selection (4-19x), prompt caching (10x input), batches (2x), and max_tokens discipline.
| Model | Input | Output | Cache Read | Cache Write | |-------|-------|--------|------------|-------------| | Claude Haiku | $0.80 | $4.00 | $0.08 | $1.00 | | Claude Sonnet | $3.00 | $15.00 | $0.30 | $3.75 | | Claude Opus | $15.00 | $75.00 | $1.50 | $18.75 |
Message Batches: 50% off all model pricing for async processing.
pythondef estimate_cost( input_tokens: int, output_tokens: int, model: str = "claude-sonnet-4-20250514", cached_input: int = 0, use_batch: bool = False ) -> float: pricing = { "claude-haiku-4-20250514": {"input": 0.80, "output": 4.00, "cache_read": 0.08}, "claude-sonnet-4-20250514": {"input": 3.00, "output": 15.00, "cache_read": 0.30}, "claude-opus-4-20250514": {"input": 15.00, "output": 75.00, "cache_read": 1.50}, } rates = pricing[model] uncached_input = input_tokens - cached_input cost = ( uncached_input * rates["input"] + cached_input * rates["cache_read"] + output_tokens * rates["output"] ) / 1_000_000 if use_batch: cost *= 0.5 return cost # Example: 10K requests/day, 500 input + 200 output tokens each daily = estimate_cost(500, 200, "claude-sonnet-4-20250514") * 10_000 print(f"Daily: ${daily:.2f}") # ~$0.045 * 10K = $450/day print(f"Monthly: ${daily * 30:.2f}") # ~$13,500/month # Same with Haiku + batching daily_optimized = estimate_cost(500, 200, "claude-haiku-4-20250514", use_batch=True) * 10_000 print(f"Optimized: ${daily_optimized:.2f}/day") # ~$22/day (20x cheaper)
pythondef route_to_model(task: str, complexity: str) -> str: """Route tasks to cheapest adequate model.""" # Haiku: classification, extraction, yes/no, routing ($0.80/$4) if task in ("classify", "extract", "route", "validate"): return "claude-haiku-4-20250514" # Sonnet: general tasks, code, tool use ($3/$15) if complexity in ("low", "medium"): return "claude-sonnet-4-20250514" # Opus: only for complex reasoning, research ($15/$75) return "claude-opus-4-20250514"
python# Cache system prompts and reference documents (90% input savings) # Break-even: 2 requests with same cached content message = client.messages.create( model="claude-sonnet-4-20250514", max_tokens=256, system=[{ "type": "text", "text": large_reference_document, # 10K+ tokens "cache_control": {"type": "ephemeral"} }], messages=[{"role": "user", "content": user_question}] )
python# 50% cost reduction for anything that doesn't need immediate response # Ideal for: summarization pipelines, data extraction, content generation batch = client.messages.batches.create(requests=[...]) # Up to 100K requests
pythonimport anthropic from dataclasses import dataclass, field @dataclass class SpendTracker: budget_usd: float = 100.0 spent_usd: float = 0.0 requests: int = 0 def track(self, response): cost = estimate_cost( response.usage.input_tokens, response.usage.output_tokens, response.model, getattr(response.usage, "cache_read_input_tokens", 0) ) self.spent_usd += cost self.requests += 1 if self.spent_usd > self.budget_usd * 0.8: print(f"WARNING: 80% budget used (${self.spent_usd:.2f}/${self.budget_usd})") if self.spent_usd > self.budget_usd: raise RuntimeError(f"Budget exceeded: ${self.spent_usd:.2f}") tracker = SpendTracker(budget_usd=50.0)
max_tokens to realistic values (not maximum)max_tokens, concurrency, retries, and batch size. Enforce per-feature and per-workspace budgets before requests are sent; fail closed when a budget or scope check cannot be evaluated.Produce a cost-control receipt containing the pricing snapshot date, policy version, model/batch/cache decisions, token aggregates, projected and observed spend, quality and latency results, budget outcome, canary scope, approval, and rollback reference. Exclude prompt/response text, customer identifiers, API keys, and raw billing exports.
| Failure | Response | |---|---| | Unknown model price or usage field | Stop forecasting, refresh the official pricing/usage source, and mark the estimate provisional. | | Budget or quota exceeded | Reject or queue new work, alert the owner, and do not bypass the guard with another key or workspace. | | Quality regression after cheaper routing | Restore the prior route, quarantine affected output, and rerun the quality fixture before another canary. | | Cache or batch unsuitable for data/latency policy | Disable that optimization and use the approved synchronous, non-cached path. |
Evaluate 1,000 synthetic classification prompts in a sandbox with a fixed budget, compare pinned Sonnet against Haiku plus an approved batch policy, assert customer_content_logged=0, and emit budget=within_limit; quality=pass; canary=internal; rollback=route-v1. Do not use live customer prompts to tune pricing.
For architecture patterns, see anth-reference-architecture.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 8,918 | 6,625 | -26% | 1 | 1 | 0% | 1,885 | 3,005 | +59% | 0 | 0 | — |
case-01 | fail→fail | 14,474 | 10,217 | -29% | 1 | 1 | 0% | 3,598 | 4,199 | +17% | 0 | 0 | — |
case-02 | fail→fail | 15,101 | 13,787 | -9% | 1 | 1 | 0% | 2,968 | 4,474 | +51% | 0 | 0 | — |
case-03 | pass→pass | 20,977 | 19,208 | -8% | 1 | 1 | 0% | 4,995 | 6,570 | +32% | 0 | 0 | — |
case-04 | fail→pass | 8,598 | 2,536 | -71% | 1 | 1 | 0% | 1,851 | 2,149 | +16% | 0 | 0 | — |
case-05 | pass→pass | 6,314 | 2,969 | -53% | 1 | 1 | 0% | 1,245 | 2,231 | +79% | 0 | 0 | — |
case-06 | pass→pass | 6,777 | 2,938 | -57% | 1 | 1 | 0% | 1,401 | 2,216 | +58% | 0 | 0 | — |
case-07 | pass→pass | 2,401 | 2,853 | +19% | 1 | 1 | 0% | 436 | 2,158 | +395% | 0 | 0 | — |
case-12 | fail→pass | 5,303 | 5,182 | -2% | 1 | 1 | 0% | 985 | 2,713 | +175% | 0 | 0 | — |
case-08 | pass→pass | 8,837 | 7,596 | -14% | 1 | 1 | 0% | 1,899 | 3,379 | +78% | 0 | 0 | — |
case-09 | pass→pass | 9,424 | 7,117 | -24% | 1 | 1 | 0% | 1,841 | 3,086 | +68% | 0 | 0 | — |
case-10 | fail→pass | 3,301 | 3,034 | -8% | 1 | 1 | 0% | 543 | 2,235 | +312% | 0 | 0 | — |
case-11 | pass→pass | 8,505 | 2,853 | -66% | 1 | 1 | 0% | 1,557 | 2,237 | +44% | 0 | 0 | — |
case-13 | fail→pass | 7,053 | 8,372 | +19% | 1 | 1 | 0% | 1,275 | 3,339 | +162% | 0 | 0 | — |
case-14 | fail→pass | 6,061 | 4,380 | -28% | 1 | 1 | 0% | 1,014 | 2,508 | +147% | 0 | 0 | — |
case-15 | pass→pass | 8,462 | 5,021 | -41% | 1 | 1 | 0% | 1,349 | 2,643 | +96% | 0 | 0 | — |
case-16 | pass→pass | 3,942 | 5,682 | +44% | 1 | 1 | 0% | 635 | 2,654 | +318% | 0 | 0 | — |
case-17 | fail→fail | 11,104 | 7,788 | -30% | 1 | 1 | 0% | 2,284 | 3,474 | +52% | 0 | 0 | — |
case-18 | pass→pass | 7,584 | 3,423 | -55% | 1 | 1 | 0% | 1,442 | 2,338 | +62% | 0 | 0 | — |
case-19 | pass→pass | 4,685 | 4,800 | +2% | 1 | 1 | 0% | 962 | 2,650 | +175% | 0 | 0 | — |
case-20 | pass→pass | 5,120 | 3,503 | -32% | 1 | 1 | 0% | 1,070 | 2,297 | +115% | 0 | 0 | — |
case-21 | pass→pass | 11,788 | 11,835 | +0% | 1 | 1 | 0% | 2,363 | 4,231 | +79% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.