Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Handle CAST AI API rate limits with backoff and request queuing. Use when hitting 429 errors, optimizing API call patterns, or implementing rate-aware batch operations. Trigger with phrases like "cast ai rate limit", "cast ai 429", "cast ai throttle", "cast ai API limits".
.claude/skills/jeremylongshore-castai-rate-limits/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 84% | 0% |
Treat rate capacity as an endpoint-specific runtime signal, not a single invented requests-per-second number. Bound concurrency, retries, polling, and total work while preserving correct authentication and organization scope.
Use Read and Grep to identify direct API calls, Terraform operations, polling loops, scheduled jobs, retries, pagination, and per-cluster fan-out. Calculate worst-case attempts, not only successful requests.
Separate reads, idempotent updates, and non-idempotent mutations. Only retry operations whose semantics are proven safe. Keep auth failures, permission denials, malformed requests, and policy rejections outside the retry path.
Use Write or Edit to set per-endpoint concurrency, maximum attempts, maximum elapsed time, request timeout, queue capacity, and global job ceiling. Because CAST AI documents endpoint-dependent limits plus an upstream edge limiter, derive values from current specification and observed behavior.
On 429, honor a valid server-provided delay signal when present; otherwise use capped exponential backoff with jitter. Coordinate workers through one limiter, stop adding retries after the deadline, and avoid synchronized polling at fixed boundaries.
Cache stable reads within their acceptable staleness window, collapse duplicate requests, paginate deliberately, query only required organizations and clusters, and replace rapid status polling with a slower bounded schedule or documented notification path when suitable.
Add deterministic tests for 429 with and without delay metadata, repeated 5xx, timeout, cancellation, queue overflow, deadline exhaustion, non-idempotent mutation, and mixed-region credentials. Assert no infinite loop and no duplicate unsafe mutation.
Use Read and Grep for client and specification analysis. Use Write and Edit for limiter logic, tests, and operational documentation. This skill does not send live requests or guess undocumented quota values.
A nightly collector shares one limiter across cluster workers and stops at its batch deadline. A policy mutation returns 429 but is not replayed until its idempotency contract is established.
| Failure | Response | | ------------------------------- | ---------------------------------------------------------- | | No endpoint limit is documented | Start conservatively and calibrate from observed responses | | Delay metadata exceeds deadline | Fail with a resumable checkpoint | | Queue reaches its bound | Shed or defer low-priority work | | Mutation outcome is unknown | Reconcile state before any retry |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 13,272 | 12,098 | -9% | 1 | 1 | 0% | 2,877 | 3,819 | +33% | 0 | 0 | — |
case-02 | fail→pass | 19,563 | 7,302 | -63% | 1 | 1 | 0% | 4,092 | 2,876 | -30% | 0 | 0 | — |
case-03 | fail→pass | 19,957 | 14,351 | -28% | 1 | 1 | 0% | 3,555 | 4,221 | +19% | 0 | 0 | — |
case-04 | pass→pass | 13,670 | 14,856 | +9% | 1 | 1 | 0% | 2,621 | 4,231 | +61% | 0 | 0 | — |
case-05 | pass→pass | 9,846 | 10,374 | +5% | 1 | 1 | 0% | 1,621 | 3,165 | +95% | 0 | 0 | — |
case-06 | pass→pass | 14,834 | 14,535 | -2% | 1 | 1 | 0% | 2,885 | 4,165 | +44% | 0 | 0 | — |
case-07 | pass→pass | 10,921 | 5,439 | -50% | 1 | 1 | 0% | 1,747 | 2,244 | +28% | 0 | 0 | — |
case-08 | pass→pass | 4,220 | 2,680 | -36% | 1 | 1 | 0% | 692 | 1,697 | +145% | 0 | 0 | — |
case-09 | fail→pass | 8,033 | 3,073 | -62% | 1 | 1 | 0% | 1,512 | 1,740 | +15% | 0 | 0 | — |
case-10 | fail→pass | 14,510 | 2,847 | -80% | 1 | 1 | 0% | 1,519 | 1,686 | +11% | 0 | 0 | — |
case-11 | pass→pass | 8,199 | 1,906 | -77% | 1 | 1 | 0% | 1,253 | 1,573 | +26% | 0 | 0 | — |
case-12 | fail→pass | 4,631 | 1,636 | -65% | 1 | 1 | 0% | 818 | 1,505 | +84% | 0 | 0 | — |
case-13 | pass→pass | 1,964 | 2,457 | +25% | 1 | 1 | 0% | 290 | 1,607 | +454% | 0 | 0 | — |
case-14 | pass→pass | 4,176 | 3,797 | -9% | 1 | 1 | 0% | 620 | 1,542 | +149% | 0 | 0 | — |
case-15 | pass→pass | 4,572 | 2,869 | -37% | 1 | 1 | 0% | 733 | 1,780 | +143% | 0 | 0 | — |
case-16 | pass→pass | 4,505 | 3,676 | -18% | 1 | 1 | 0% | 609 | 1,906 | +213% | 0 | 0 | — |
case-17 | pass→pass | 5,959 | 2,148 | -64% | 1 | 1 | 0% | 901 | 1,681 | +87% | 0 | 0 | — |
case-18 | fail→pass | 6,780 | 3,631 | -46% | 1 | 1 | 0% | 1,181 | 1,903 | +61% | 0 | 0 | — |
case-19 | pass→pass | 7,567 | 2,982 | -61% | 1 | 1 | 0% | 1,256 | 1,628 | +30% | 0 | 0 | — |
case-20 | pass→pass | 19,266 | 4,218 | -78% | 1 | 1 | 0% | 1,699 | 1,894 | +11% | 0 | 0 | — |
case-21 | pass→pass | 6,433 | 2,851 | -56% | 1 | 1 | 0% | 1,047 | 1,715 | +64% | 0 | 0 | — |
case-22 | pass→pass | 7,314 | 3,641 | -50% | 1 | 1 | 0% | 1,100 | 1,884 | +71% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.