Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optimize Anthropic API latency — streaming, prompt caching, model selection, Use when working with performance-tuning patterns. connection reuse, and parallel requests. Trigger with "anthropic slow", "claude latency", "speed up anthropic", "anthropic performance", "claude response time".
.claude/skills/jeremylongshore-clade-performance-tuning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 23% | 0% |
Claude latency has two components: time to first token (TTFT) and tokens per second (TPS). Different strategies target each.
| Model | TTFT (p50) | TTFT (p95) | Output TPS | |-------|-----------|-----------|------------| | Claude Haiku 4.5 | 200ms | 600ms | ~150 | | Claude Sonnet 4 | 400ms | 1.2s | ~90 | | Claude Opus 4 | 800ms | 2.5s | ~40 |
typescript// Streaming delivers the first token ASAP — user sees response instantly // instead of waiting for the full response to generate const stream = client.messages.stream({ model: 'claude-sonnet-4-20250514', max_tokens: 1024, messages, }); // First token arrives in ~400ms (Sonnet) // Full response may take 5-10s, but user sees progress immediately for await (const event of stream) { if (event.type === 'content_block_delta') { yield event.delta.text; } }
typescript// Cached prompts skip re-processing — dramatically lower TTFT for large system prompts const message = await client.messages.create({ model: 'claude-sonnet-4-20250514', max_tokens: 1024, system: [{ type: 'text', text: largeSystemPrompt, // 10K+ tokens cache_control: { type: 'ephemeral' }, }], messages, }, { headers: { 'claude-beta': 'prompt-caching-2024-07-31' }, }); // TTFT drops from ~2s to ~500ms on cache hit with large prompts
typescript// Haiku is 2-4x faster than Sonnet with 80% quality for many tasks // Use for: classification, extraction, simple Q&A, routing decisions const route = await client.messages.create({ model: 'claude-haiku-4-5-20251001', // 200ms TTFT max_tokens: 10, system: 'Classify the intent. Reply with exactly one word: search, create, update, delete.', messages: [{ role: 'user', content: userInput }], }); // Then use Sonnet/Opus for the actual task
typescript// BAD — creates new connection pool per request app.get('/api/chat', async (req, res) => { const client = new Anthropic(); // DON'T // ... }); // GOOD — single client shared across requests const client = new Anthropic(); // Module-level singleton app.get('/api/chat', async (req, res) => { const message = await client.messages.create({ ... }); // ... });
typescript// When you need multiple independent Claude calls, fire them in parallel const [summary, sentiment, entities] = await Promise.all([ client.messages.create({ model: 'claude-haiku-4-5-20251001', max_tokens: 200, messages: [{ role: 'user', content: `Summarize: ${text}` }] }), client.messages.create({ model: 'claude-haiku-4-5-20251001', max_tokens: 20, messages: [{ role: 'user', content: `Sentiment (positive/negative/neutral): ${text}` }] }), client.messages.create({ model: 'claude-haiku-4-5-20251001', max_tokens: 200, messages: [{ role: 'user', content: `Extract named entities from: ${text}` }] }), ]);
typescript// Fewer output tokens = faster response system: 'Be extremely concise. Use bullet points, not paragraphs.', // Set tight max_tokens max_tokens: 256, // Don't use 4096 for short answers
| Issue | Cause | Fix | |-------|-------|-----| | TTFT > 3s | Large uncached prompt | Enable prompt caching | | Slow output | Using Opus for simple tasks | Downgrade to Haiku/Sonnet | | Timeouts | Long generation + default timeout | new Anthropic({ timeout: 120_000 }) | | 529 overloaded | API capacity | SDK auto-retries; add fallback model |
See Latency Benchmarks table and six numbered strategy sections above, each with complete TypeScript code examples.
See clade-deploy-integration for production deployment patterns.
clade-install-auth| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→pass | 9,930 | 9,145 | -8% | 1 | 1 | 0% | 1,709 | 2,228 | +30% | 0 | 0 | — |
case-01 | fail→pass | 11,608 | 8,012 | -31% | 1 | 1 | 0% | 2,223 | 2,969 | +34% | 0 | 0 | — |
case-02 | pass→fail | 12,720 | 9,134 | -28% | 1 | 1 | 0% | 2,526 | 3,066 | +21% | 0 | 0 | — |
case-04 | pass→pass | 6,422 | 4,263 | -34% | 1 | 1 | 0% | 1,084 | 2,131 | +97% | 0 | 0 | — |
case-05 | pass→pass | 10,896 | 7,773 | -29% | 1 | 1 | 0% | 2,089 | 2,892 | +38% | 0 | 0 | — |
case-06 | fail→fail | 12,478 | 6,619 | -47% | 1 | 1 | 0% | 1,917 | 2,524 | +32% | 0 | 0 | — |
case-07 | pass→pass | 8,366 | 6,943 | -17% | 1 | 1 | 0% | 1,649 | 2,671 | +62% | 0 | 0 | — |
case-08 | pass→pass | 10,780 | 7,473 | -31% | 1 | 1 | 0% | 1,988 | 2,795 | +41% | 0 | 0 | — |
case-09 | fail→pass | 6,298 | 1,948 | -69% | 1 | 1 | 0% | 876 | 1,724 | +97% | 0 | 0 | — |
case-10 | fail→pass | 7,019 | 1,851 | -74% | 1 | 1 | 0% | 1,166 | 1,678 | +44% | 0 | 0 | — |
case-11 | fail→pass | 8,289 | 2,123 | -74% | 1 | 1 | 0% | 1,413 | 1,738 | +23% | 0 | 0 | — |
case-12 | pass→pass | 12,184 | 8,776 | -28% | 1 | 1 | 0% | 2,042 | 2,848 | +39% | 0 | 0 | — |
case-13 | fail→pass | 12,653 | 9,135 | -28% | 1 | 1 | 0% | 1,987 | 2,966 | +49% | 0 | 0 | — |
case-14 | pass→pass | 13,040 | 8,317 | -36% | 1 | 1 | 0% | 2,008 | 2,764 | +38% | 0 | 0 | — |
case-15 | pass→pass | 6,811 | 2,669 | -61% | 1 | 1 | 0% | 1,223 | 1,907 | +56% | 0 | 0 | — |
case-16 | fail→pass | 14,649 | 9,183 | -37% | 1 | 1 | 0% | 2,346 | 3,194 | +36% | 0 | 0 | — |
case-17 | pass→pass | 3,383 | 5,073 | +50% | 1 | 1 | 0% | 491 | 2,235 | +355% | 0 | 0 | — |
case-18 | pass→pass | 2,791 | 2,092 | -25% | 1 | 1 | 0% | 377 | 1,749 | +364% | 0 | 0 | — |
case-19 | pass→pass | 7,357 | 5,021 | -32% | 1 | 1 | 0% | 1,451 | 2,502 | +72% | 0 | 0 | — |
case-20 | pass→pass | 9,861 | 7,595 | -23% | 1 | 1 | 0% | 1,862 | 2,843 | +53% | 0 | 0 | — |
case-21 | pass→pass | 8,336 | 7,954 | -5% | 1 | 1 | 0% | 1,514 | 2,811 | +86% | 0 | 0 | — |
case-22 | pass→pass | 5,651 | 5,079 | -10% | 1 | 1 | 0% | 1,046 | 2,421 | +131% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.