Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optimize Anthropic API costs — model selection, prompt caching, batches, Use when working with cost-tuning patterns. token reduction, and usage monitoring. Trigger with "anthropic pricing", "claude cost", "reduce anthropic spend", "anthropic billing", "claude cheaper".
.claude/skills/jeremylongshore-clade-cost-tuning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -30% | 0% |
Anthropic charges per token. Input tokens, output tokens, and cached tokens each have different prices. Here's how to minimize cost without losing quality.
| Model | Input | Output | Cached Input | Batch Input | Batch Output | |-------|-------|--------|-------------|-------------|--------------| | Claude Opus 4 | $15.00 | $75.00 | $1.50 | $7.50 | $37.50 | | Claude Sonnet 4 | $3.00 | $15.00 | $0.30 | $1.50 | $7.50 | | Claude Haiku 4.5 | $0.80 | $4.00 | $0.08 | $0.40 | $2.00 |
typescript// DON'T use Opus for everything // DO match model to task complexity: // Simple classification/extraction → Haiku (cheapest) const category = await classify(text, 'claude-haiku-4-5-20251001'); // General coding/writing → Sonnet (balanced) const code = await generate(spec, 'claude-sonnet-4-20250514'); // Complex multi-step reasoning → Opus (best quality) const analysis = await analyze(data, 'claude-opus-4-20250514');
typescript// Cache your system prompt — pays for itself after 2 calls const message = await client.messages.create({ model: 'claude-sonnet-4-20250514', max_tokens: 1024, system: [{ type: 'text', text: longSystemPrompt, // Must be 1024+ tokens cache_control: { type: 'ephemeral' }, // Cache for 5 minutes }], messages, }, { headers: { 'claude-beta': 'prompt-caching-2024-07-31' }, }); // First call: cache_creation_input_tokens charged at 1.25x // Subsequent calls: cache_read_input_tokens charged at 0.1x (90% savings!)
typescript// For non-urgent work — 50% cheaper, 24h processing SLA const batch = await client.messages.batches.create({ requests: prompts.map((p, i) => ({ custom_id: `job-${i}`, params: { model: 'claude-sonnet-4-20250514', max_tokens: 1024, messages: [{ role: 'user', content: p }], }, })), }); // Sonnet: $1.50/$7.50 per MTok instead of $3/$15
typescript// Trim conversation history — keep system + last N turns function trimMessages(messages: MessageParam[], maxTurns = 10) { if (messages.length <= maxTurns * 2) return messages; return messages.slice(-(maxTurns * 2)); } // Set tight max_tokens — don't pay for output you won't use const message = await client.messages.create({ model: 'claude-sonnet-4-20250514', max_tokens: 256, // Not 4096 if you only need a short answer messages, }); // Use concise system prompts system: 'Reply in 1-2 sentences.' // Not a 500-word personality description
typescript// Log every call's cost function logUsage(message: Anthropic.Message) { const { input_tokens, output_tokens } = message.usage; const cost = (input_tokens * 3 + output_tokens * 15) / 1_000_000; // Sonnet pricing console.log(`Tokens: ${input_tokens}in/${output_tokens}out | Cost: $${cost.toFixed(4)}`); }
Processing 10,000 documents (avg 500 tokens each, 200 token response):
| Strategy | Input Cost | Output Cost | Total | |----------|-----------|-------------|-------| | Opus, no optimization | $75.00 | $150.00 | $225.00 | | Sonnet, no optimization | $15.00 | $30.00 | $45.00 | | Sonnet + Batches | $7.50 | $15.00 | $22.50 | | Haiku + Batches | $2.00 | $4.00 | $6.00 | | Haiku + Batches + Caching | ~$1.00 | $4.00 | ~$5.00 |
| Error | Cause | Solution | |-------|-------|----------| | API Error | Check error type and status code | See clade-common-errors |
See Pricing table, five numbered strategy sections with code, and the Cost Comparison Example table showing savings from $225 to $5 for 10K documents.
See clade-performance-tuning for latency optimization.
clade-install-auth| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 18,359 | 12,037 | -34% | 1 | 1 | 0% | 4,217 | 4,305 | +2% | 0 | 0 | — |
case-02 | fail→pass | 17,776 | 10,803 | -39% | 1 | 1 | 0% | 3,523 | 3,755 | +7% | 0 | 0 | — |
case-03 | fail→pass | 10,092 | 5,164 | -49% | 1 | 1 | 0% | 1,830 | 2,527 | +38% | 0 | 0 | — |
case-04 | pass→pass | 5,065 | 3,875 | -23% | 1 | 1 | 0% | 931 | 2,285 | +145% | 0 | 0 | — |
case-05 | pass→pass | 4,898 | 4,687 | -4% | 1 | 1 | 0% | 882 | 2,389 | +171% | 0 | 0 | — |
case-06 | pass→pass | 8,982 | 1,756 | -80% | 1 | 1 | 0% | 1,503 | 1,769 | +18% | 0 | 0 | — |
case-07 | fail→pass | 5,962 | 2,020 | -66% | 1 | 1 | 0% | 1,053 | 1,847 | +75% | 0 | 0 | — |
case-08 | fail→pass | 21,853 | 7,545 | -65% | 1 | 1 | 0% | 4,317 | 3,027 | -30% | 0 | 0 | — |
case-09 | fail→pass | 6,733 | 2,243 | -67% | 1 | 1 | 0% | 709 | 1,849 | +161% | 0 | 0 | — |
case-10 | pass→pass | 4,770 | 1,654 | -65% | 1 | 1 | 0% | 841 | 1,772 | +111% | 0 | 0 | — |
case-11 | pass→pass | 10,848 | 6,783 | -37% | 1 | 1 | 0% | 2,250 | 2,987 | +33% | 0 | 0 | — |
case-12 | pass→pass | 12,042 | 8,218 | -32% | 1 | 1 | 0% | 2,192 | 3,054 | +39% | 0 | 0 | — |
case-13 | pass→pass | 10,316 | 6,662 | -35% | 1 | 1 | 0% | 1,621 | 2,623 | +62% | 0 | 0 | — |
case-14 | pass→pass | 5,375 | 2,526 | -53% | 1 | 1 | 0% | 899 | 1,882 | +109% | 0 | 0 | — |
case-15 | pass→pass | 4,225 | 3,559 | -16% | 1 | 1 | 0% | 760 | 2,175 | +186% | 0 | 0 | — |
case-16 | pass→pass | 3,902 | 2,560 | -34% | 1 | 1 | 0% | 538 | 1,792 | +233% | 0 | 0 | — |
case-17 | fail→pass | 5,234 | 1,492 | -71% | 1 | 1 | 0% | 895 | 1,716 | +92% | 0 | 0 | — |
case-18 | fail→pass | 5,578 | 1,728 | -69% | 1 | 1 | 0% | 929 | 1,760 | +89% | 0 | 0 | — |
case-19 | fail→pass | 6,115 | 2,791 | -54% | 1 | 1 | 0% | 1,054 | 2,083 | +98% | 0 | 0 | — |
case-20 | fail→fail | 17,921 | 13,655 | -24% | 1 | 1 | 0% | 3,473 | 4,162 | +20% | 0 | 0 | — |
case-21 | pass→pass | 10,021 | 8,573 | -14% | 1 | 1 | 0% | 1,953 | 3,081 | +58% | 0 | 0 | — |
case-22 | pass→pass | 15,267 | 10,382 | -32% | 1 | 1 | 0% | 3,351 | 3,508 | +5% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.