Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Monitor Claude API calls — log tokens, latency, costs, errors, and Use when working with observability patterns. set up alerts for production Claude integrations. Trigger with "anthropic monitoring", "claude observability", "track claude usage", "anthropic logging".
.claude/skills/jeremylongshore-clade-observability/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -44% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -7% | 0% |
Every messages.create call should be instrumented. Track tokens, latency, cost, model, and errors.
typescriptimport Anthropic from '@claude-ai/sdk'; const client = new Anthropic(); async function trackedCreate(params: Anthropic.MessageCreateParams) { const start = performance.now(); try { const message = await client.messages.create(params); const durationMs = Math.round(performance.now() - start); const log = { timestamp: new Date().toISOString(), model: message.model, input_tokens: message.usage.input_tokens, output_tokens: message.usage.output_tokens, cache_read_tokens: message.usage.cache_read_input_tokens || 0, duration_ms: durationMs, stop_reason: message.stop_reason, estimated_cost: estimateCost(message.model, message.usage), }; console.log('anthropic_request', JSON.stringify(log)); return message; } catch (err) { const durationMs = Math.round(performance.now() - start); console.error('anthropic_error', JSON.stringify({ timestamp: new Date().toISOString(), model: params.model, error_type: err instanceof Anthropic.APIError ? err.error?.type : 'unknown', status: err instanceof Anthropic.APIError ? err.status : null, request_id: err instanceof Anthropic.APIError ? err.headers?.['request-id'] : null, duration_ms: durationMs, })); throw err; } } function estimateCost(model: string, usage: Anthropic.Usage): number { const rates: Record<string, [number, number]> = { 'claude-opus-4-20250514': [15, 75], 'claude-sonnet-4-20250514': [3, 15], 'claude-haiku-4-5-20251001': [0.80, 4], }; const [inputRate, outputRate] = rates[model] || [3, 15]; return (usage.input_tokens * inputRate + usage.output_tokens * outputRate) / 1_000_000; }
| Metric | Source | Alert Threshold | |--------|--------|----------------| | Error rate | error logs | > 5% over 5 minutes | | p95 latency | duration_ms | > 10s (Sonnet) | | Daily cost | estimated_cost sum | > 2x daily average | | 429 rate | error_type = rate_limit | > 10/minute | | 529 rate | error_type = overloaded | > 5/minute | | Token usage | input_tokens + output_tokens | > daily budget |
| Error | Cause | Solution | |-------|-------|----------| | API Error | Check error type and status code | See clade-common-errors |
See Logging Wrapper with trackedCreate(), estimateCost() function, Key Metrics table with alert thresholds, and Anthropic Console Monitoring section above.
See clade-incident-runbook for when things go wrong.
clade-install-authEach section contains production-ready code examples. Copy and adapt them to your use case.
Integrate the patterns that match your requirements. Test each change individually.
Run your test suite to confirm the integration works correctly.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 18,069 | 26,891 | +49% | 1 | 1 | 0% | 3,983 | 3,711 | -7% | 0 | 0 | — |
case-02 | fail→pass | 24,987 | 11,742 | -53% | 1 | 1 | 0% | 3,037 | 3,342 | +10% | 0 | 0 | — |
case-03 | fail→pass | 18,223 | 27,262 | +50% | 1 | 1 | 0% | 4,160 | 5,117 | +23% | 0 | 0 | — |
case-04 | pass→pass | 6,584 | 3,287 | -50% | 1 | 1 | 0% | 1,297 | 1,684 | +30% | 0 | 0 | — |
case-05 | pass→pass | 8,270 | 2,735 | -67% | 1 | 1 | 0% | 1,858 | 1,664 | -10% | 0 | 0 | — |
case-06 | fail→pass | 11,133 | 2,927 | -74% | 1 | 1 | 0% | 2,237 | 1,516 | -32% | 0 | 0 | — |
case-07 | pass→pass | 10,434 | 2,482 | -76% | 1 | 1 | 0% | 1,825 | 1,556 | -15% | 0 | 0 | — |
case-08 | fail→pass | 13,398 | 2,077 | -84% | 1 | 1 | 0% | 2,282 | 1,271 | -44% | 0 | 0 | — |
case-09 | pass→pass | 13,432 | 4,696 | -65% | 1 | 1 | 0% | 2,740 | 1,957 | -29% | 0 | 0 | — |
case-10 | pass→pass | 12,637 | 7,161 | -43% | 1 | 1 | 0% | 2,423 | 2,394 | -1% | 0 | 0 | — |
case-11 | pass→pass | 11,382 | 8,370 | -26% | 1 | 1 | 0% | 2,020 | 2,572 | +27% | 0 | 0 | — |
case-12 | pass→pass | 11,365 | 2,474 | -78% | 1 | 1 | 0% | 1,902 | 1,442 | -24% | 0 | 0 | — |
case-13 | pass→pass | 11,046 | 4,216 | -62% | 1 | 1 | 0% | 2,083 | 1,804 | -13% | 0 | 0 | — |
case-14 | fail→pass | 11,535 | 3,934 | -66% | 1 | 1 | 0% | 1,963 | 1,825 | -7% | 0 | 0 | — |
case-15 | fail→pass | 21,650 | 1,984 | -91% | 1 | 1 | 0% | 1,834 | 1,309 | -29% | 0 | 0 | — |
case-16 | fail→pass | 24,053 | 2,650 | -89% | 1 | 1 | 0% | 1,995 | 1,432 | -28% | 0 | 0 | — |
case-17 | fail→pass | 9,959 | 1,834 | -82% | 1 | 1 | 0% | 1,424 | 1,310 | -8% | 0 | 0 | — |
case-18 | fail→pass | 14,516 | 3,176 | -78% | 1 | 1 | 0% | 2,430 | 1,399 | -42% | 0 | 0 | — |
case-19 | fail→pass | 6,687 | 2,020 | -70% | 1 | 1 | 0% | 975 | 1,364 | +40% | 0 | 0 | — |
case-20 | pass→pass | 8,683 | 5,505 | -37% | 1 | 1 | 0% | 1,805 | 2,082 | +15% | 0 | 0 | — |
case-21 | pass→pass | 14,810 | 11,159 | -25% | 1 | 1 | 0% | 3,014 | 3,626 | +20% | 0 | 0 | — |
case-22 | pass→pass | 7,208 | 4,611 | -36% | 1 | 1 | 0% | 1,320 | 1,942 | +47% | 0 | 0 | — |
case-23 | pass→pass | 10,607 | 8,078 | -24% | 1 | 1 | 0% | 1,899 | 2,409 | +27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +43 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.