Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Implement load testing, auto-scaling, and capacity planning for Claude API. Use when running performance benchmarks, planning for traffic spikes, or configuring horizontal scaling for Claude-powered services. Trigger with phrases like "anthropic load test", "claude scaling", "anthropic capacity planning", "scale claude api".
.claude/skills/jeremylongshore-anth-load-scale/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 90% | 0% |
Capacity planning and load testing for Claude API integrations. Key constraint: your rate limits (RPM/ITPM/OTPM) are the ceiling, not your infrastructure.
python# Calculate required tier based on traffic def plan_capacity( requests_per_minute: int, avg_input_tokens: int, avg_output_tokens: int, model: str = "claude-sonnet-4-20250514" ) -> dict: itpm = requests_per_minute * avg_input_tokens otpm = requests_per_minute * avg_output_tokens # Estimate monthly cost pricing = { "claude-haiku-4-20250514": (0.80, 4.00), "claude-sonnet-4-20250514": (3.00, 15.00), "claude-opus-4-20250514": (15.00, 75.00), } rates = pricing[model] cost_per_request = (avg_input_tokens * rates[0] + avg_output_tokens * rates[1]) / 1_000_000 monthly_cost = cost_per_request * requests_per_minute * 60 * 24 * 30 return { "rpm_needed": requests_per_minute, "itpm_needed": itpm, "otpm_needed": otpm, "cost_per_request": f"${cost_per_request:.4f}", "monthly_estimate": f"${monthly_cost:,.0f}", "recommendation": "Contact Anthropic sales for Scale tier" if requests_per_minute > 500 else "Self-serve tiers sufficient", } print(plan_capacity(100, 500, 200))
pythonimport anthropic import asyncio import time from dataclasses import dataclass @dataclass class LoadTestResult: total_requests: int = 0 successful: int = 0 failed: int = 0 rate_limited: int = 0 avg_latency_ms: float = 0 p99_latency_ms: float = 0 total_input_tokens: int = 0 total_output_tokens: int = 0 async def load_test( concurrency: int = 10, total_requests: int = 100, model: str = "claude-haiku-4-20250514" ) -> LoadTestResult: client = anthropic.Anthropic() result = LoadTestResult() latencies = [] semaphore = asyncio.Semaphore(concurrency) async def single_request(): async with semaphore: start = time.monotonic() try: msg = client.messages.create( model=model, max_tokens=64, messages=[{"role": "user", "content": "Respond with exactly: OK"}] ) duration = (time.monotonic() - start) * 1000 latencies.append(duration) result.successful += 1 result.total_input_tokens += msg.usage.input_tokens result.total_output_tokens += msg.usage.output_tokens except anthropic.RateLimitError: result.rate_limited += 1 except Exception: result.failed += 1 result.total_requests += 1 tasks = [single_request() for _ in range(total_requests)] await asyncio.gather(*tasks) if latencies: latencies.sort() result.avg_latency_ms = sum(latencies) / len(latencies) result.p99_latency_ms = latencies[int(len(latencies) * 0.99)] return result # Run: asyncio.run(load_test(concurrency=10, total_requests=50))
| Strategy | When | Implementation | |----------|------|---------------| | Queue-based processing | > 50 RPM sustained | Redis/SQS queue + worker pool | | Model routing | Mixed workloads | Haiku for simple, Sonnet for complex | | Message Batches | Offline processing | 100K requests, 50% cheaper, no RPM impact | | Prompt caching | Repeated system prompts | 90% input token savings | | Request coalescing | Duplicate prompts | Cache identical request hashes |
python# Multiple application instances sharing the same API key # Rate limits are per-organization, NOT per-instance # Use a shared rate limiter (Redis) to coordinate import redis r = redis.Redis() def check_rate_limit(key: str = "claude:rpm", limit: int = 100, window: int = 60) -> bool: current = r.incr(key) if current == 1: r.expire(key, window) return current <= limit
| Issue | Cause | Fix | |-------|-------|-----| | 429 during load test | Exceeded tier limits | Reduce concurrency or upgrade tier | | Increasing latency under load | Output queue saturation | Reduce max_tokens | | Uneven request distribution | No load balancing | Use queue for fair distribution |
side_effects=0. Promote only after an owner approves the result; revert autoscaling/limiter changes on regression.Return a capacity receipt with workload class, model, concurrency steps, aggregate request/token counts, p50/p95/p99 latency, status/429 counts, queue depth, cost estimate, threshold decision, canary result, rollback reference, and cleanup status. Do not include payloads or secret material.
Run 50 requests using Respond with exactly: OK in the sandbox, cap concurrency at 10, and assert side_effects=0. A useful receipt is requests=50; successes=50; rate_limited=0; p99_ms=<redacted>; tokens=<aggregate>; canary=pass; cleanup=verified.
For reliability patterns, see anth-reliability-patterns.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 16,529 | 13,487 | -18% | 1 | 1 | 0% | 3,320 | 4,169 | +26% | 0 | 0 | — |
case-01 | fail→fail | 17,711 | 12,633 | -29% | 1 | 1 | 0% | 4,013 | 4,573 | +14% | 0 | 0 | — |
case-02 | fail→fail | 14,568 | 12,447 | -15% | 1 | 1 | 0% | 3,549 | 4,408 | +24% | 0 | 0 | — |
case-04 | pass→pass | 7,333 | 8,893 | +21% | 1 | 1 | 0% | 1,386 | 3,059 | +121% | 0 | 0 | — |
case-05 | pass→pass | 6,527 | 5,885 | -10% | 1 | 1 | 0% | 1,212 | 2,555 | +111% | 0 | 0 | — |
case-06 | fail→pass | 14,248 | 4,034 | -72% | 1 | 1 | 0% | 2,395 | 2,007 | -16% | 0 | 0 | — |
case-07 | fail→pass | 12,219 | 8,287 | -32% | 1 | 1 | 0% | 2,104 | 2,739 | +30% | 0 | 0 | — |
case-22 | pass→pass | 9,192 | 9,887 | +8% | 1 | 1 | 0% | 2,087 | 3,565 | +71% | 0 | 0 | — |
case-08 | pass→pass | 13,083 | 14,348 | +10% | 1 | 1 | 0% | 2,201 | 4,124 | +87% | 0 | 0 | — |
case-09 | fail→fail | 8,152 | 5,760 | -29% | 1 | 1 | 0% | 1,389 | 2,360 | +70% | 0 | 0 | — |
case-10 | pass→pass | 12,195 | 9,154 | -25% | 1 | 1 | 0% | 2,223 | 3,161 | +42% | 0 | 0 | — |
case-11 | pass→pass | 14,455 | 16,039 | +11% | 1 | 1 | 0% | 2,675 | 4,811 | +80% | 0 | 0 | — |
case-12 | pass→pass | 12,455 | 10,969 | -12% | 1 | 1 | 0% | 2,301 | 3,738 | +62% | 0 | 0 | — |
case-13 | pass→pass | 3,187 | 3,041 | -5% | 1 | 1 | 0% | 778 | 2,060 | +165% | 0 | 0 | — |
case-14 | fail→pass | 6,167 | 2,367 | -62% | 1 | 1 | 0% | 1,344 | 1,896 | +41% | 0 | 0 | — |
case-15 | fail→pass | 5,176 | 2,466 | -52% | 1 | 1 | 0% | 981 | 1,901 | +94% | 0 | 0 | — |
case-16 | fail→pass | 5,144 | 2,196 | -57% | 1 | 1 | 0% | 965 | 1,835 | +90% | 0 | 0 | — |
case-17 | fail→pass | 12,356 | 4,008 | -68% | 1 | 1 | 0% | 2,688 | 2,143 | -20% | 0 | 0 | — |
case-18 | fail→pass | 10,293 | 1,978 | -81% | 1 | 1 | 0% | 1,901 | 1,763 | -7% | 0 | 0 | — |
case-19 | fail→fail | 17,960 | 15,919 | -11% | 1 | 1 | 0% | 3,561 | 4,573 | +28% | 0 | 0 | — |
case-20 | pass→pass | 14,470 | 12,873 | -11% | 1 | 1 | 0% | 2,794 | 3,902 | +40% | 0 | 0 | — |
case-21 | pass→pass | 13,220 | 12,337 | -7% | 1 | 1 | 0% | 2,574 | 3,779 | +47% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.