Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Diagnose and fix CAST AI agent, API, and autoscaler errors. Use when the CAST AI agent is offline, nodes are not scaling, or API calls return errors. Trigger with phrases like "cast ai error", "cast ai not working", "cast ai agent offline", "cast ai debug", "fix cast ai".
.claude/skills/jeremylongshore-castai-common-errors/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 62% | 0% |
Separate observation, connectivity, policy, capacity, and disruption failures before proposing a change. Preserve the failing state, use current component topology, and stop when the evidence requires cloud-provider or CAST AI support access.
castai-agent namespaceRecord expected versus actual behavior, timestamps, workload identity, pending-pod reason, and recent configuration changes. Use Read and Grep on runbooks and IaC to determine whether Cost Monitoring, Node Autoscaling, or Workload Autoscaling is actually enabled.
Use Bash(castctl:_) for version or non-mutating status commands supported by the installed client. Use Bash(helm:_) to inspect releases and values, then Bash(kubectl:\) to inspect workloads, readiness, events, and bounded logs in castai-agent. Do not restart components before collecting evidence.
| Plane | Evidence | Likely boundary | | ---------------- | ------------------------------------------------------ | ----------------------------------------------------------- | | Connection | Agent readiness, outbound failures, console disconnect | Identity, network, or cloud permissions | | Node scaling | Pending pods, policy bounds, node-template fit | Unsatisfied constraints or maximum CPU boundary | | Workload scaling | Missing recommendations, policy assignment, metrics | Metrics server, confidence, policy, or unsupported workload | | Disruption | Eviction denial, PDB events, deferred changes | PDB or selected apply mode | | Reporting | Missing cost or savings window | Ingestion, baseline, adoption, or pricing configuration |
Choose the smallest reversible check. Confirm regional endpoint alignment, effective scaling-policy assignment, metrics availability, supported workload type, node-template constraints, and cloud quota. Treat the deprecated cluster minimum CPU setting as migration debt, not a current control to add.
Map the evidence to the owning layer. Change repository-managed values only through their source of truth; do not mix console edits into Terraform or GitOps ownership. Escalate with a redacted bundle when the failure is inside the hosted control plane or an undocumented provider response.
Use Read and Grep for configuration and runbook evidence. Use Bash(kubectl:_), Bash(helm:_), and Bash(castctl:\) only for bounded inspection commands. Do not apply, upgrade, restart, connect, disconnect, or expose Secret objects during diagnosis.
Recommendations are absent because metrics-server is missing, so the remedy belongs to cluster observability. A node remains pending because every approved node template conflicts with its constraints; increasing a global limit without reviewing the workload is not the remedy.
| Failure | Response | | ------------------------------------- | ------------------------------------------------------------------- | | Kube context is ambiguous | Stop before any cluster command and resolve it | | Logs include credentials or inventory | Redact locally and do not attach raw output | | A PDB blocks Immediate mode | Preserve the PDB and evaluate Deferred mode with the workload owner | | Evidence points to cloud quota | Escalate to the cloud owner with the exact denied dimension |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 17,374 | 9,013 | -48% | 1 | 1 | 0% | 2,018 | 2,913 | +44% | 0 | 0 | — |
case-01 | fail→fail | 14,727 | 24,363 | +65% | 1 | 1 | 0% | 2,797 | 4,000 | +43% | 0 | 0 | — |
case-02 | fail→pass | 18,329 | 16,591 | -9% | 1 | 1 | 0% | 3,270 | 4,564 | +40% | 0 | 0 | — |
case-03 | fail→fail | 15,965 | 12,694 | -20% | 1 | 1 | 0% | 2,988 | 3,824 | +28% | 0 | 0 | — |
case-04 | pass→pass | 7,027 | 4,049 | -42% | 1 | 1 | 0% | 1,305 | 2,142 | +64% | 0 | 0 | — |
case-05 | pass→pass | 5,242 | 2,008 | -62% | 1 | 1 | 0% | 825 | 1,760 | +113% | 0 | 0 | — |
case-06 | fail→pass | 8,865 | 3,800 | -57% | 1 | 1 | 0% | 1,553 | 2,016 | +30% | 0 | 0 | — |
case-07 | pass→pass | 5,937 | 4,226 | -29% | 1 | 1 | 0% | 1,083 | 2,025 | +87% | 0 | 0 | — |
case-08 | fail→pass | 8,756 | 2,786 | -68% | 1 | 1 | 0% | 1,596 | 1,899 | +19% | 0 | 0 | — |
case-09 | fail→pass | 7,570 | 3,139 | -59% | 1 | 1 | 0% | 1,100 | 1,915 | +74% | 0 | 0 | — |
case-10 | pass→pass | 8,730 | 4,484 | -49% | 1 | 1 | 0% | 1,581 | 2,154 | +36% | 0 | 0 | — |
case-11 | fail→pass | 9,668 | 6,232 | -36% | 1 | 1 | 0% | 1,517 | 2,464 | +62% | 0 | 0 | — |
case-12 | fail→pass | 12,712 | 3,484 | -73% | 1 | 1 | 0% | 1,985 | 2,068 | +4% | 0 | 0 | — |
case-13 | fail→pass | 10,149 | 5,038 | -50% | 1 | 1 | 0% | 1,670 | 2,271 | +36% | 0 | 0 | — |
case-14 | fail→fail | 13,276 | 8,504 | -36% | 1 | 1 | 0% | 2,242 | 3,068 | +37% | 0 | 0 | — |
case-15 | fail→fail | 10,361 | 7,029 | -32% | 1 | 1 | 0% | 2,007 | 2,772 | +38% | 0 | 0 | — |
case-16 | pass→pass | 8,267 | 2,829 | -66% | 1 | 1 | 0% | 1,477 | 1,899 | +29% | 0 | 0 | — |
case-17 | fail→pass | 11,140 | 4,602 | -59% | 1 | 1 | 0% | 1,765 | 2,146 | +22% | 0 | 0 | — |
case-18 | pass→pass | 2,053 | 1,827 | -11% | 1 | 1 | 0% | 324 | 1,635 | +405% | 0 | 0 | — |
case-19 | pass→pass | 10,041 | 6,882 | -31% | 1 | 1 | 0% | 1,910 | 2,765 | +45% | 0 | 0 | — |
case-20 | pass→pass | 25,318 | 23,212 | -8% | 1 | 1 | 0% | 4,880 | 6,234 | +28% | 0 | 0 | — |
case-21 | pass→pass | 11,966 | 14,780 | +24% | 1 | 1 | 0% | 2,280 | 3,944 | +73% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.