Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Scale Claude usage for high-throughput applications — batches, queues, Use when working with load-scale patterns. concurrency control, and tier upgrades. Trigger with "anthropic scale", "claude high volume", "anthropic throughput", "scale claude api", "anthropic concurrent requests".
.claude/skills/jeremylongshore-clade-load-scale/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 83% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -30% | 0% |
Scale Claude usage for high-throughput applications. Covers four strategies: Message Batches (10K requests, 50% off, no rate limits), request queues with concurrency control via p-limit, tier upgrades (Tier 1-4 + Scale), and model selection for throughput (Haiku is 3-4x faster than Sonnet).
typescript// 10K requests per batch, 50% cheaper, no rate limits const batch = await client.messages.batches.create({ requests: items.map((item, i) => ({ custom_id: `${i}`, params: { model: 'claude-sonnet-4-20250514', max_tokens: 1024, messages: [{ role: 'user', content: item }] }, })), }); // Process up to 100 concurrent batches
typescriptimport pLimit from 'p-limit'; // Match your rate limit tier const limit = pLimit(10); // 10 concurrent requests const results = await Promise.all( inputs.map(input => limit(() => client.messages.create({ model: 'claude-sonnet-4-20250514', max_tokens: 1024, messages: [{ role: 'user', content: input }], })) ) );
Increase your spending to unlock higher tiers:
| Tier | RPM | Input TPM | How to Qualify | |------|-----|-----------|----------------| | 1 | 50 | 40K | Free | | 2 | 1,000 | 80K | $40+ total spend | | 3 | 2,000 | 160K | $200+ total spend | | 4 | 4,000 | 400K | $400+ total spend | | Scale | Custom | Custom | Contact sales |
typescript// Haiku processes 3-4x faster than Sonnet, 8x faster than Opus // Use the fastest model that meets quality requirements const model = taskComplexity === 'simple' ? 'claude-haiku-4-5-20251001' : 'claude-sonnet-4-20250514';
typescript// Track throughput metrics let requestCount = 0; let tokenCount = 0; setInterval(() => { console.log(`Throughput: ${requestCount} req/min, ${tokenCount} tokens/min`); requestCount = 0; tokenCount = 0; }, 60_000);
| Error | Cause | Solution | |-------|-------|----------| | API Error | Check error type and status code | See clade-common-errors |
See Message Batches example, p-limit concurrency control, Tier Upgrades table, and Monitoring at Scale metrics tracking above.
See clade-reliability-patterns for fault-tolerant high-scale patterns.
clade-rate-limits for understanding tier limits| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→pass | 17,652 | 12,467 | -29% | 1 | 1 | 0% | 3,943 | 3,525 | -11% | 0 | 0 | — |
case-01 | fail→pass | 19,389 | 14,780 | -24% | 1 | 1 | 0% | 3,632 | 4,128 | +14% | 0 | 0 | — |
case-02 | fail→pass | 19,496 | 16,953 | -13% | 1 | 1 | 0% | 3,154 | 4,212 | +34% | 0 | 0 | — |
case-04 | pass→pass | 14,216 | 7,124 | -50% | 1 | 1 | 0% | 2,899 | 2,425 | -16% | 0 | 0 | — |
case-05 | pass→pass | 10,705 | 5,141 | -52% | 1 | 1 | 0% | 1,632 | 1,840 | +13% | 0 | 0 | — |
case-06 | pass→pass | 4,874 | 2,791 | -43% | 1 | 1 | 0% | 825 | 1,405 | +70% | 0 | 0 | — |
case-07 | pass→pass | 8,244 | 3,899 | -53% | 1 | 1 | 0% | 1,618 | 1,563 | -3% | 0 | 0 | — |
case-08 | pass→pass | 3,827 | 2,079 | -46% | 1 | 1 | 0% | 634 | 1,298 | +105% | 0 | 0 | — |
case-09 | fail→pass | 4,023 | 2,232 | -45% | 1 | 1 | 0% | 721 | 1,322 | +83% | 0 | 0 | — |
case-10 | fail→pass | 18,812 | 3,027 | -84% | 1 | 1 | 0% | 2,222 | 1,554 | -30% | 0 | 0 | — |
case-11 | pass→pass | 8,049 | 2,029 | -75% | 1 | 1 | 0% | 1,415 | 1,203 | -15% | 0 | 0 | — |
case-12 | pass→pass | 7,284 | 2,155 | -70% | 1 | 1 | 0% | 1,389 | 1,320 | -5% | 0 | 0 | — |
case-22 | pass→pass | 7,885 | 2,754 | -65% | 1 | 1 | 0% | 1,357 | 1,303 | -4% | 0 | 0 | — |
case-13 | fail→pass | 11,847 | 5,189 | -56% | 1 | 1 | 0% | 1,857 | 1,920 | +3% | 0 | 0 | — |
case-14 | fail→pass | 12,531 | 3,349 | -73% | 1 | 1 | 0% | 2,093 | 1,481 | -29% | 0 | 0 | — |
case-15 | fail→pass | 9,876 | 2,375 | -76% | 1 | 1 | 0% | 1,598 | 1,204 | -25% | 0 | 0 | — |
case-16 | fail→fail | 4,204 | 4,133 | -2% | 1 | 1 | 0% | 793 | 1,686 | +113% | 0 | 0 | — |
case-17 | pass→pass | 16,846 | 7,536 | -55% | 1 | 1 | 0% | 3,286 | 2,350 | -28% | 0 | 0 | — |
case-18 | pass→pass | 11,689 | 8,931 | -24% | 1 | 1 | 0% | 2,467 | 2,603 | +6% | 0 | 0 | — |
case-19 | pass→pass | 13,042 | 11,806 | -9% | 1 | 1 | 0% | 2,352 | 3,092 | +31% | 0 | 0 | — |
case-20 | fail→pass | 21,047 | 13,873 | -34% | 1 | 1 | 0% | 2,616 | 3,317 | +27% | 0 | 0 | — |
case-21 | pass→pass | 7,203 | 2,541 | -65% | 1 | 1 | 0% | 1,071 | 1,339 | +25% | 0 | 0 | — |
case-23 | pass→pass | 8,118 | 6,580 | -19% | 1 | 1 | 0% | 1,393 | 2,202 | +58% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.