Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute Deepgram production deployment checklist. Use when preparing for production launch, auditing production readiness, or verifying deployment configurations. Trigger: "deepgram production", "deploy deepgram", "deepgram prod checklist", "deepgram go-live", "production ready deepgram".
.claude/skills/jeremylongshore-deepgram-prod-checklist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 91% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 234% | 0% |
Comprehensive go-live checklist for Deepgram integrations. Covers singleton client, health checks, Prometheus metrics, alert rules, error handling, and a phased go-live timeline.
| Category | Item | Status | |----------|------|--------| | Auth | Production API key with scoped permissions | ] | | Auth | Key stored in secret manager (not env file) | ] | | Auth | Key rotation schedule (90-day) configured | ] | | Auth | Fallback key provisioned and tested | ] | | Resilience | Retry with exponential backoff on 429/5xx | ] | | Resilience | Circuit breaker for cascade failure prevention | ] | | Resilience | Request timeout set (30s pre-recorded, 10s TTS) | ] | | Resilience | Graceful degradation when API unavailable | ] | | Performance | Singleton client (not creating per-request) | ] | | Performance | Concurrency limited (50-80% of plan limit) | ] | | Performance | Audio preprocessed (16kHz mono for best results) | ] | | Performance | Large files use callback URL (async) | ] | | Monitoring | Health check endpoint testing Deepgram API | ] | | Monitoring | Prometheus metrics: latency, error rate, usage | ] | | Monitoring | Alerts: error rate >5%, latency >10s, circuit open | ] | | Security | PII redaction enabled if handling sensitive audio | ] | | Security | Audio URLs validated (HTTPS, no private IPs) | ] | | Security | Audit logging on all operations | ] |
typescriptimport { createClient, DeepgramClient } from '@deepgram/sdk'; class ProductionDeepgram { private static client: DeepgramClient | null = null; static getClient(): DeepgramClient { if (!this.client) { const key = process.env.DEEPGRAM_API_KEY; if (!key) throw new Error('DEEPGRAM_API_KEY required for production'); this.client = createClient(key); } return this.client; } // Force re-init (for key rotation) static reset() { this.client = null; } }
typescriptimport express from 'express'; import { createClient } from '@deepgram/sdk'; const app = express(); const deepgram = createClient(process.env.DEEPGRAM_API_KEY!); app.get('/health', async (req, res) => { const start = Date.now(); try { // Test API connectivity by listing projects const { error } = await deepgram.manage.getProjects(); const latency = Date.now() - start; if (error) { return res.status(503).json({ status: 'unhealthy', deepgram: 'error', error: error.message, latency_ms: latency, }); } res.json({ status: 'healthy', deepgram: 'connected', latency_ms: latency, timestamp: new Date().toISOString(), }); } catch (err: any) { res.status(503).json({ status: 'unhealthy', deepgram: 'unreachable', error: err.message, latency_ms: Date.now() - start, }); } });
typescriptimport { Counter, Histogram, Gauge, Registry } from 'prom-client'; const registry = new Registry(); const transcriptionRequests = new Counter({ name: 'deepgram_requests_total', help: 'Total Deepgram API requests', labelNames: ['method', 'model', 'status'], registers: [registry], }); const transcriptionLatency = new Histogram({ name: 'deepgram_latency_seconds', help: 'Deepgram API request latency', labelNames: ['method', 'model'], buckets: [0.5, 1, 2, 5, 10, 30], registers: [registry], }); const audioProcessed = new Counter({ name: 'deepgram_audio_seconds_total', help: 'Total audio seconds processed', labelNames: ['model'], registers: [registry], }); const activeConnections = new Gauge({ name: 'deepgram_active_connections', help: 'Active WebSocket connections', registers: [registry], }); // Instrumented transcription async function instrumentedTranscribe(url: string, model = 'nova-3') { const timer = transcriptionLatency.startTimer({ method: 'prerecorded', model }); try { const { result, error } = await deepgram.listen.prerecorded.transcribeUrl( { url }, { model, smart_format: true } ); timer(); transcriptionRequests.inc({ method: 'prerecorded', model, status: error ? 'error' : 'ok' }); if (result?.metadata?.duration) { audioProcessed.inc({ model }, result.metadata.duration); } if (error) throw error; return result; } catch (err) { timer(); transcriptionRequests.inc({ method: 'prerecorded', model, status: 'error' }); throw err; } } // Expose metrics endpoint app.get('/metrics', async (req, res) => { res.set('Content-Type', registry.contentType); res.send(await registry.metrics()); });
yamlgroups: - name: deepgram rules: - alert: DeepgramHighErrorRate expr: rate(deepgram_requests_total{status="error"}[5m]) / rate(deepgram_requests_total[5m]) > 0.05 for: 5m labels: severity: critical annotations: summary: "Deepgram error rate > 5%" - alert: DeepgramHighLatency expr: histogram_quantile(0.95, rate(deepgram_latency_seconds_bucket[5m])) > 10 for: 5m labels: severity: warning annotations: summary: "Deepgram P95 latency > 10s" - alert: DeepgramHealthCheckFailed expr: up{job="deepgram-service"} == 0 for: 2m labels: severity: critical annotations: summary: "Deepgram health check failed for 2+ minutes"
typescriptasync function safeTranscribe(url: string, options: Record<string, any> = {}) { const timeout = options.timeout ?? 30000; const controller = new AbortController(); const timeoutId = setTimeout(() => controller.abort(), timeout); try { const result = await Promise.race([ instrumentedTranscribe(url, options.model ?? 'nova-3'), new Promise((_, reject) => setTimeout(() => reject(new Error('Transcription timeout')), timeout) ), ]); clearTimeout(timeoutId); return result; } catch (err: any) { clearTimeout(timeoutId); // Log structured error console.error(JSON.stringify({ level: 'error', service: 'deepgram', message: err.message, url: url.substring(0, 100), timestamp: new Date().toISOString(), })); throw err; } }
| Phase | When | Actions | |-------|------|---------| | D-7 | 1 week before | Load test at 2x expected volume, security review | | D-3 | 3 days before | Smoke test with production key, verify all alerts fire | | D-1 | Day before | Confirm on-call rotation, validate dashboards | | D-0 | Launch | Shadow mode (10% traffic), monitoring open | | D+1 | Day after | Review error rate, latency, verify no anomalies | | D+7 | 1 week after | Full traffic, tune alert thresholds based on baselines |
| Issue | Cause | Solution | |-------|-------|----------| | Health check 503 | API key expired | Rotate key, check secret manager | | Metrics not scraped | Wrong port/path | Verify Prometheus target config | | Alert storms | Thresholds too tight | Add for: duration, tune values | | Timeout on large files | Sync mode too slow | Switch to callback URL pattern |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 17,370 | 11,638 | -33% | 1 | 1 | 0% | 3,697 | 4,816 | +30% | 0 | 0 | — |
case-02 | fail→fail | 29,337 | 21,191 | -28% | 1 | 1 | 0% | 5,862 | 6,440 | +10% | 0 | 0 | — |
case-03 | fail→pass | 35,525 | 24,135 | -32% | 1 | 1 | 0% | 5,900 | 6,714 | +14% | 0 | 0 | — |
case-04 | pass→pass | 23,590 | 18,911 | -20% | 1 | 1 | 0% | 4,763 | 6,081 | +28% | 0 | 0 | — |
case-05 | pass→pass | 15,123 | 16,169 | +7% | 1 | 1 | 0% | 3,058 | 5,259 | +72% | 0 | 0 | — |
case-06 | pass→pass | 14,026 | 16,116 | +15% | 1 | 1 | 0% | 3,356 | 5,631 | +68% | 0 | 0 | — |
case-07 | pass→pass | 9,242 | 9,360 | +1% | 1 | 1 | 0% | 1,775 | 4,066 | +129% | 0 | 0 | — |
case-08 | fail→pass | 12,332 | 11,516 | -7% | 1 | 1 | 0% | 2,363 | 4,515 | +91% | 0 | 0 | — |
case-09 | fail→fail | 10,458 | 8,622 | -18% | 1 | 1 | 0% | 1,895 | 3,965 | +109% | 0 | 0 | — |
case-10 | fail→fail | 10,572 | 10,547 | -0% | 1 | 1 | 0% | 1,898 | 3,892 | +105% | 0 | 0 | — |
case-11 | fail→pass | 8,351 | 6,007 | -28% | 1 | 1 | 0% | 1,382 | 3,249 | +135% | 0 | 0 | — |
case-12 | fail→pass | 11,536 | 12,340 | +7% | 1 | 1 | 0% | 2,212 | 4,795 | +117% | 0 | 0 | — |
case-13 | pass→pass | 8,143 | 8,318 | +2% | 1 | 1 | 0% | 1,304 | 3,752 | +188% | 0 | 0 | — |
case-14 | fail→pass | 6,067 | 35,680 | +488% | 1 | 1 | 0% | 928 | 3,096 | +234% | 0 | 0 | — |
case-15 | fail→fail | 6,211 | 4,154 | -33% | 1 | 1 | 0% | 877 | 2,914 | +232% | 0 | 0 | — |
case-16 | fail→fail | 9,658 | 10,964 | +14% | 1 | 1 | 0% | 1,800 | 4,437 | +147% | 0 | 0 | — |
case-17 | fail→pass | 6,241 | 10,191 | +63% | 1 | 1 | 0% | 1,218 | 4,002 | +229% | 0 | 0 | — |
case-18 | pass→pass | 75,602 | 12,992 | -83% | 1 | 1 | 0% | 2,200 | 4,651 | +111% | 0 | 0 | — |
case-19 | fail→fail | 16,882 | 19,024 | +13% | 1 | 1 | 0% | 2,569 | 5,753 | +124% | 0 | 0 | — |
case-20 | fail→pass | 24,180 | 19,678 | -19% | 1 | 1 | 0% | 3,729 | 5,492 | +47% | 0 | 0 | — |
case-21 | fail→pass | 18,411 | 18,399 | -0% | 1 | 1 | 0% | 2,814 | 5,110 | +82% | 0 | 0 | — |
case-22 | fail→pass | 13,849 | 12,047 | -13% | 1 | 1 | 0% | 2,137 | 4,093 | +92% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.