Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Set up comprehensive observability for Deepgram integrations. Use when implementing monitoring, setting up dashboards, or configuring alerting for Deepgram integration health. Trigger: "deepgram monitoring", "deepgram metrics", "deepgram observability", "monitor deepgram", "deepgram alerts", "deepgram dashboard".
.claude/skills/jeremylongshore-deepgram-observability/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-20 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 125% | 0% |
Full observability stack for Deepgram: Prometheus metrics (request counts, latency histograms, audio processed, cost tracking), OpenTelemetry distributed tracing, structured JSON logging with Pino, Grafana dashboard JSON, and AlertManager rules.
| Pillar | Tool | What It Tracks | |--------|------|----------------| | Metrics | Prometheus | Request rate, latency, error rate, audio minutes, estimated cost | | Traces | OpenTelemetry | End-to-end request flow, Deepgram API span timing | | Logs | Pino (JSON) | Request details, errors, audit trail | | Alerts | AlertManager | Error rate >5%, P95 latency >10s, rate limit hits |
typescriptimport { Counter, Histogram, Gauge, Registry, collectDefaultMetrics } from 'prom-client'; const registry = new Registry(); collectDefaultMetrics({ register: registry }); // Request metrics const requestsTotal = new Counter({ name: 'deepgram_requests_total', help: 'Total Deepgram API requests', labelNames: ['method', 'model', 'status'] as const, registers: [registry], }); const latencyHistogram = new Histogram({ name: 'deepgram_request_duration_seconds', help: 'Deepgram API request duration', labelNames: ['method', 'model'] as const, buckets: [0.1, 0.5, 1, 2, 5, 10, 30, 60], registers: [registry], }); // Usage metrics const audioProcessedSeconds = new Counter({ name: 'deepgram_audio_processed_seconds_total', help: 'Total audio seconds processed', labelNames: ['model'] as const, registers: [registry], }); const estimatedCostDollars = new Counter({ name: 'deepgram_estimated_cost_dollars_total', help: 'Estimated cost in USD', labelNames: ['model', 'method'] as const, registers: [registry], }); // Operational metrics const activeConnections = new Gauge({ name: 'deepgram_active_websocket_connections', help: 'Currently active WebSocket connections', registers: [registry], }); const rateLimitHits = new Counter({ name: 'deepgram_rate_limit_hits_total', help: 'Number of 429 rate limit responses', registers: [registry], }); export { registry, requestsTotal, latencyHistogram, audioProcessedSeconds, estimatedCostDollars, activeConnections, rateLimitHits };
typescriptimport { createClient, DeepgramClient } from '@deepgram/sdk'; class InstrumentedDeepgram { private client: DeepgramClient; private costPerMinute: Record<string, number> = { 'nova-3': 0.0043, 'nova-2': 0.0043, 'base': 0.0048, 'whisper-large': 0.0048, }; constructor(apiKey: string) { this.client = createClient(apiKey); } async transcribeUrl(url: string, options: Record<string, any> = {}) { const model = options.model ?? 'nova-3'; const timer = latencyHistogram.startTimer({ method: 'prerecorded', model }); try { const { result, error } = await this.client.listen.prerecorded.transcribeUrl( { url }, { model, smart_format: true, ...options } ); const status = error ? 'error' : 'success'; timer(); requestsTotal.inc({ method: 'prerecorded', model, status }); if (error) { if ((error as any).status === 429) rateLimitHits.inc(); throw error; } // Track usage const duration = result.metadata.duration; audioProcessedSeconds.inc({ model }, duration); estimatedCostDollars.inc( { model, method: 'prerecorded' }, (duration / 60) * (this.costPerMinute[model] ?? 0.0043) ); return result; } catch (err) { timer(); requestsTotal.inc({ method: 'prerecorded', model, status: 'error' }); throw err; } } // Live transcription with connection tracking connectLive(options: Record<string, any>) { const model = options.model ?? 'nova-3'; activeConnections.inc(); const connection = this.client.listen.live(options); const originalFinish = connection.finish.bind(connection); connection.finish = () => { activeConnections.dec(); return originalFinish(); }; return connection; } }
typescriptimport { NodeSDK } from '@opentelemetry/sdk-node'; import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http'; import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node'; import { Resource } from '@opentelemetry/resources'; import { SEMRESATTRS_SERVICE_NAME } from '@opentelemetry/semantic-conventions'; import { trace } from '@opentelemetry/api'; const sdk = new NodeSDK({ resource: new Resource({ [SEMRESATTRS_SERVICE_NAME]: 'deepgram-service', 'deployment.environment': process.env.NODE_ENV ?? 'development', }), traceExporter: new OTLPTraceExporter({ url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT ?? 'http://localhost:4318/v1/traces', }), instrumentations: [ getNodeAutoInstrumentations({ '@opentelemetry/instrumentation-http': { ignoreIncomingPaths: ['/health', '/metrics'], }, }), ], }); sdk.start(); // Add custom spans for Deepgram operations const tracer = trace.getTracer('deepgram'); async function tracedTranscribe(url: string, model: string) { return tracer.startActiveSpan('deepgram.transcribe', async (span) => { span.setAttribute('deepgram.model', model); span.setAttribute('deepgram.audio_url', url.substring(0, 100)); try { const instrumented = new InstrumentedDeepgram(process.env.DEEPGRAM_API_KEY!); const result = await instrumented.transcribeUrl(url, { model }); span.setAttribute('deepgram.duration_seconds', result.metadata.duration); span.setAttribute('deepgram.request_id', result.metadata.request_id); span.setAttribute('deepgram.confidence', result.results.channels[0].alternatives[0].confidence); return result; } catch (err: any) { span.recordException(err); span.setStatus({ code: 2, message: err.message }); throw err; } finally { span.end(); } }); }
typescriptimport pino from 'pino'; const logger = pino({ level: process.env.LOG_LEVEL ?? 'info', formatters: { level: (label) => ({ level: label }), }, timestamp: pino.stdTimeFunctions.isoTime, base: { service: 'deepgram-integration', env: process.env.NODE_ENV, }, }); // Child loggers per component const transcriptionLog = logger.child({ component: 'transcription' }); const metricsLog = logger.child({ component: 'metrics' }); // Usage: transcriptionLog.info({ action: 'transcribe', model: 'nova-3', audioUrl: url.substring(0, 100), requestId: result.metadata.request_id, duration: result.metadata.duration, confidence: result.results.channels[0].alternatives[0].confidence, }, 'Transcription completed'); transcriptionLog.error({ action: 'transcribe', model: 'nova-3', error: err.message, statusCode: err.status, }, 'Transcription failed');
json{ "title": "Deepgram Observability", "panels": [ { "title": "Request Rate", "type": "timeseries", "targets": [{ "expr": "rate(deepgram_requests_total[5m])" }] }, { "title": "P95 Latency", "type": "gauge", "targets": [{ "expr": "histogram_quantile(0.95, rate(deepgram_request_duration_seconds_bucket[5m]))" }] }, { "title": "Error Rate %", "type": "stat", "targets": [{ "expr": "rate(deepgram_requests_total{status='error'}[5m]) / rate(deepgram_requests_total[5m]) * 100" }] }, { "title": "Audio Processed (min/hr)", "type": "timeseries", "targets": [{ "expr": "rate(deepgram_audio_processed_seconds_total[1h]) / 60" }] }, { "title": "Estimated Daily Cost", "type": "stat", "targets": [{ "expr": "increase(deepgram_estimated_cost_dollars_total[24h])" }] }, { "title": "Active WebSocket Connections", "type": "gauge", "targets": [{ "expr": "deepgram_active_websocket_connections" }] } ] }
yamlgroups: - name: deepgram-alerts rules: - alert: DeepgramHighErrorRate expr: > rate(deepgram_requests_total{status="error"}[5m]) / rate(deepgram_requests_total[5m]) > 0.05 for: 5m labels: { severity: critical } annotations: summary: "Deepgram error rate > 5% for 5 minutes" - alert: DeepgramHighLatency expr: > histogram_quantile(0.95, rate(deepgram_request_duration_seconds_bucket[5m]) ) > 10 for: 5m labels: { severity: warning } annotations: summary: "Deepgram P95 latency > 10 seconds" - alert: DeepgramRateLimited expr: rate(deepgram_rate_limit_hits_total[1h]) > 10 for: 10m labels: { severity: warning } annotations: summary: "Deepgram rate limit hits > 10/hour" - alert: DeepgramCostSpike expr: > increase(deepgram_estimated_cost_dollars_total[24h]) > 2 * increase(deepgram_estimated_cost_dollars_total[24h] offset 1d) for: 30m labels: { severity: warning } annotations: summary: "Deepgram daily cost > 2x yesterday" - alert: DeepgramZeroRequests expr: rate(deepgram_requests_total[15m]) == 0 for: 15m labels: { severity: warning } annotations: summary: "No Deepgram requests for 15 minutes"
typescriptimport express from 'express'; const app = express(); app.get('/metrics', async (req, res) => { res.set('Content-Type', registry.contentType); res.send(await registry.metrics()); });
| Issue | Cause | Solution | |-------|-------|----------| | Metrics not appearing | Registry not exported | Check /metrics endpoint | | High cardinality | Too many label values | Limit labels to known set | | Alert storms | Thresholds too sensitive | Add for: duration, tune values | | Missing traces | OTEL exporter not configured | Set OTEL_EXPORTER_OTLP_ENDPOINT |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | fail→pass | 13,881 | 6,794 | -51% | 1 | 1 | 0% | 2,566 | 4,343 | +69% | 0 | 0 | — |
case-01 | fail→pass | 29,708 | 87,520 | +195% | 1 | 1 | 0% | 6,240 | 9,384 | +50% | 0 | 0 | — |
case-02 | fail→pass | 29,386 | 28,186 | -4% | 1 | 1 | 0% | 6,228 | 9,243 | +48% | 0 | 0 | — |
case-03 | fail→fail | 24,853 | 19,494 | -22% | 1 | 1 | 0% | 5,373 | 7,582 | +41% | 0 | 0 | — |
case-04 | fail→fail | 13,085 | 16,384 | +25% | 1 | 1 | 0% | 2,340 | 5,967 | +155% | 0 | 0 | — |
case-05 | pass→pass | 10,293 | 5,900 | -43% | 1 | 1 | 0% | 1,755 | 4,314 | +146% | 0 | 0 | — |
case-06 | pass→pass | 18,337 | 20,173 | +10% | 1 | 1 | 0% | 3,515 | 7,239 | +106% | 0 | 0 | — |
case-07 | fail→fail | 12,997 | 9,501 | -27% | 1 | 1 | 0% | 2,296 | 4,788 | +109% | 0 | 0 | — |
case-08 | fail→pass | 13,659 | 9,048 | -34% | 1 | 1 | 0% | 2,324 | 4,703 | +102% | 0 | 0 | — |
case-09 | fail→pass | 11,033 | 6,334 | -43% | 1 | 1 | 0% | 1,920 | 4,326 | +125% | 0 | 0 | — |
case-10 | fail→fail | 12,734 | 9,375 | -26% | 1 | 1 | 0% | 2,235 | 4,926 | +120% | 0 | 0 | — |
case-11 | fail→fail | 11,949 | 11,725 | -2% | 1 | 1 | 0% | 2,238 | 5,433 | +143% | 0 | 0 | — |
case-12 | fail→pass | 15,017 | 44,840 | +199% | 1 | 1 | 0% | 2,698 | 5,083 | +88% | 0 | 0 | — |
case-13 | fail→pass | 15,546 | 15,424 | -1% | 1 | 1 | 0% | 2,793 | 6,307 | +126% | 0 | 0 | — |
case-19 | fail→fail | 8,580 | 7,025 | -18% | 1 | 1 | 0% | 1,510 | 4,434 | +194% | 0 | 0 | — |
case-14 | pass→pass | 7,021 | 7,566 | +8% | 1 | 1 | 0% | 1,166 | 4,303 | +269% | 0 | 0 | — |
case-15 | fail→pass | 54,906 | 6,874 | -87% | 1 | 1 | 0% | 1,792 | 4,342 | +142% | 0 | 0 | — |
case-16 | fail→pass | 9,241 | 4,476 | -52% | 1 | 1 | 0% | 1,637 | 3,909 | +139% | 0 | 0 | — |
case-17 | pass→pass | 9,774 | 6,452 | -34% | 1 | 1 | 0% | 1,737 | 4,042 | +133% | 0 | 0 | — |
case-18 | fail→pass | 6,975 | 2,725 | -61% | 1 | 1 | 0% | 1,273 | 3,621 | +184% | 0 | 0 | — |
case-21 | pass→pass | 9,199 | 3,649 | -60% | 1 | 1 | 0% | 1,569 | 3,616 | +130% | 0 | 0 | — |
case-22 | fail→pass | 43,919 | 3,081 | -93% | 1 | 1 | 0% | 1,350 | 3,610 | +167% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.