Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optimize Deepgram costs and usage for budget-conscious deployments. Use when reducing transcription costs, implementing usage controls, or optimizing pricing tier utilization. Trigger: "deepgram cost", "reduce deepgram spending", "deepgram pricing", "deepgram budget", "optimize deepgram usage", "deepgram billing".
.claude/skills/jeremylongshore-deepgram-cost-tuning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 124% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 153% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 198% | 0% |
Optimize Deepgram API costs through smart model selection, audio preprocessing to reduce billable minutes, usage monitoring via the Deepgram API, budget guardrails, and feature-aware cost estimation. Deepgram bills per audio minute processed.
| Product | Model | Price/Minute | Notes | |---------|-------|-------------|-------| | STT (Batch) | Nova-3 | $0.0043 | Best accuracy | | STT (Batch) | Nova-2 | $0.0043 | Proven stable | | STT (Streaming) | Nova-3 | $0.0059 | Real-time | | STT (Streaming) | Nova-2 | $0.0059 | Real-time | | STT (Batch) | Base | $0.0048 | Fastest | | STT (Batch) | Whisper | $0.0048 | Multilingual | | TTS | Aura-2 | Pay-per-character | See TTS pricing | | Intelligence | Summarize/Topics/Sentiment | Included with STT | No extra cost |
Add-on costs:
typescriptimport { createClient } from '@deepgram/sdk'; interface BudgetConfig { monthlyLimitUsd: number; warningThreshold: number; // 0.0-1.0 (e.g., 0.8 = warn at 80%) costPerMinute: number; // Base STT cost } class BudgetAwareTranscriber { private client: ReturnType<typeof createClient>; private config: BudgetConfig; private monthlySpendUsd = 0; private monthlyMinutes = 0; constructor(apiKey: string, config: BudgetConfig) { this.client = createClient(apiKey); this.config = config; } async transcribe(source: any, options: any) { // Estimate cost before transcription const estimatedCost = this.estimateCost(options); const projected = this.monthlySpendUsd + estimatedCost; if (projected > this.config.monthlyLimitUsd) { throw new Error( `Budget exceeded: $${this.monthlySpendUsd.toFixed(2)} spent, ` + `$${this.config.monthlyLimitUsd} limit` ); } if (projected > this.config.monthlyLimitUsd * this.config.warningThreshold) { console.warn( `Budget warning: ${((projected / this.config.monthlyLimitUsd) * 100).toFixed(0)}% ` + `of $${this.config.monthlyLimitUsd} limit` ); } const { result, error } = await this.client.listen.prerecorded.transcribeUrl( source, options ); if (error) throw error; // Track actual usage const duration = result.metadata.duration / 60; // Convert to minutes const actualCost = this.calculateCost(duration, options); this.monthlyMinutes += duration; this.monthlySpendUsd += actualCost; return result; } private estimateCost(options: any): number { // Conservative estimate — assume 5 minutes per file return this.calculateCost(5, options); } private calculateCost(minutes: number, options: any): number { let cost = minutes * this.config.costPerMinute; if (options.diarize) cost += minutes * 0.0044; // Diarization add-on return cost; } getUsageSummary() { return { minutesUsed: this.monthlyMinutes.toFixed(1), spentUsd: this.monthlySpendUsd.toFixed(4), remainingUsd: (this.config.monthlyLimitUsd - this.monthlySpendUsd).toFixed(4), utilizationPercent: ((this.monthlySpendUsd / this.config.monthlyLimitUsd) * 100).toFixed(1), }; } } // Usage: const transcriber = new BudgetAwareTranscriber(process.env.DEEPGRAM_API_KEY!, { monthlyLimitUsd: 100, warningThreshold: 0.8, costPerMinute: 0.0043, });
bash# Remove silence — can save 10-40% of billable minutes ffmpeg -i input.wav \ -af "silenceremove=stop_periods=-1:stop_duration=0.5:stop_threshold=-30dB" \ -ar 16000 -ac 1 -acodec pcm_s16le \ trimmed.wav # Speed up audio (1.25x) — saves 20% of billable minutes # Deepgram handles slightly sped-up audio well ffmpeg -i input.wav \ -filter:a "atempo=1.25" \ -ar 16000 -ac 1 -acodec pcm_s16le \ faster.wav
typescriptimport { execSync } from 'child_process'; function measureSavings(inputPath: string) { // Get original duration const origDuration = parseFloat( execSync(`ffprobe -v quiet -show_entries format=duration -of csv=p=0 "${inputPath}"`) .toString().trim() ); // Remove silence execSync(`ffmpeg -y -i "${inputPath}" \ -af "silenceremove=stop_periods=-1:stop_duration=0.5:stop_threshold=-30dB" \ -ar 16000 -ac 1 -acodec pcm_s16le /tmp/trimmed.wav 2>/dev/null`); const trimmedDuration = parseFloat( execSync(`ffprobe -v quiet -show_entries format=duration -of csv=p=0 /tmp/trimmed.wav`) .toString().trim() ); const savings = ((1 - trimmedDuration / origDuration) * 100).toFixed(1); const costSaved = ((origDuration - trimmedDuration) / 60 * 0.0043).toFixed(4); console.log(`Original: ${origDuration.toFixed(1)}s`); console.log(`Trimmed: ${trimmedDuration.toFixed(1)}s`); console.log(`Savings: ${savings}% (${costSaved}/file at $0.0043/min)`); }
typescriptimport { createClient } from '@deepgram/sdk'; async function getUsageDashboard(projectId: string) { const client = createClient(process.env.DEEPGRAM_API_KEY!); // Get usage for current month const now = new Date(); const monthStart = new Date(now.getFullYear(), now.getMonth(), 1); const { result } = await client.manage.getUsage(projectId, { start: monthStart.toISOString(), end: now.toISOString(), }); // Aggregate by model const byModel: Record<string, { minutes: number; cost: number }> = {}; for (const entry of (result as any).results ?? []) { const model = entry.model ?? 'unknown'; if (!byModel[model]) byModel[model] = { minutes: 0, cost: 0 }; byModel[model].minutes += (entry.hours ?? 0) * 60 + (entry.minutes ?? 0); } console.log('=== Monthly Usage ==='); for (const [model, data] of Object.entries(byModel)) { const cost = data.minutes * 0.0043; console.log(`${model}: ${data.minutes.toFixed(1)} min ($${cost.toFixed(2)})`); } // Monthly projection const dayOfMonth = now.getDate(); const daysInMonth = new Date(now.getFullYear(), now.getMonth() + 1, 0).getDate(); const totalMinutes = Object.values(byModel).reduce((s, d) => s + d.minutes, 0); const projectedMinutes = (totalMinutes / dayOfMonth) * daysInMonth; const projectedCost = projectedMinutes * 0.0043; console.log(`\nProjected monthly: ${projectedMinutes.toFixed(0)} min ($${projectedCost.toFixed(2)})`); }
typescriptfunction recommendModel(params: { qualityNeeded: 'high' | 'medium' | 'low'; isRealtime: boolean; languages: string[]; budgetPerMinute?: number; }): { model: string; pricePerMin: number; reason: string } { const { qualityNeeded, isRealtime, languages, budgetPerMinute } = params; // Multilingual -> Whisper if (languages.length > 1 || !['en', 'es', 'fr', 'de'].includes(languages[0])) { return { model: 'whisper-large', pricePerMin: 0.0048, reason: 'Multilingual support' }; } // Budget constraint if (budgetPerMinute !== undefined && budgetPerMinute < 0.005) { return { model: 'nova-2', pricePerMin: 0.0043, reason: 'Best price per quality' }; } // Real-time -> Nova-3 (streaming price $0.0059/min) if (isRealtime) { return { model: 'nova-3', pricePerMin: 0.0059, reason: 'Best real-time accuracy' }; } // Quality based switch (qualityNeeded) { case 'high': return { model: 'nova-3', pricePerMin: 0.0043, reason: 'Highest accuracy' }; case 'medium': return { model: 'nova-2', pricePerMin: 0.0043, reason: 'Good accuracy, proven' }; case 'low': return { model: 'base', pricePerMin: 0.0048, reason: 'Fastest processing' }; } }
typescript// Feature cost breakdown per minute of audio const featureCosts: Record<string, { cost: number; description: string }> = { // Free features (included with STT) smart_format: { cost: 0, description: 'Punctuation + paragraphs + numerals' }, punctuate: { cost: 0, description: 'Punctuation only' }, paragraphs: { cost: 0, description: 'Paragraph formatting' }, summarize: { cost: 0, description: 'AI summary (included with STT)' }, detect_topics: { cost: 0, description: 'Topic detection (included)' }, sentiment: { cost: 0, description: 'Sentiment analysis (included)' }, intents: { cost: 0, description: 'Intent recognition (included)' }, redact: { cost: 0, description: 'PII redaction (included)' }, // Paid add-ons diarize: { cost: 0.0044, description: 'Speaker identification (+$0.0044/min)' }, multichannel: { cost: 0.0043, description: 'Per-channel billing (1x STT cost per channel)' }, }; function estimateJobCost(params: { durationMinutes: number; model: string; features: string[]; channels?: number; }): number { const baseCost = params.durationMinutes * 0.0043; let addOnCost = 0; for (const feature of params.features) { addOnCost += (featureCosts[feature]?.cost ?? 0) * params.durationMinutes; } // Multichannel: billed per channel const channelMultiplier = params.channels ?? 1; return (baseCost + addOnCost) * channelMultiplier; } // Example: 60 min meeting with diarization // estimateJobCost({ durationMinutes: 60, model: 'nova-3', features: ['diarize'] }) // = (60 * 0.0043) + (60 * 0.0044) = $0.258 + $0.264 = $0.522
| Strategy | Savings | Effort | |----------|---------|--------| | Remove silence from audio | 10-40% | Low (ffmpeg one-liner) | | Disable diarization when not needed | ~50% | Low (remove option) | | Use callback for long files | Indirect (no timeouts) | Low | | Cache repeated transcriptions | 20-60% | Medium (Redis) | | Speed up audio 1.25x | 20% | Low (ffmpeg) | | Use Nova-2 instead of Nova-3 | 0% (same price) | None | | Batch pre-recorded vs streaming | 37% ($0.0043 vs $0.0059) | Medium |
| Issue | Cause | Solution | |-------|-------|----------| | Budget exceeded | No controls | Enable budget check before transcription | | Unexpected charges | Diarization always on | Make diarization opt-in | | Usage API empty | Wrong project ID | Get ID from getProjects() | | Cost spike | Batch job without limits | Set concurrency limits + budget cap |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,891 | 14,228 | -10% | 1 | 1 | 0% | 3,203 | 6,549 | +104% | 0 | 0 | — |
case-02 | fail→pass | 20,617 | 14,532 | -30% | 1 | 1 | 0% | 4,737 | 6,750 | +42% | 0 | 0 | — |
case-03 | fail→pass | 18,916 | 15,538 | -18% | 1 | 1 | 0% | 4,077 | 6,991 | +71% | 0 | 0 | — |
case-09 | fail→pass | 13,946 | 11,766 | -16% | 1 | 1 | 0% | 2,621 | 5,869 | +124% | 0 | 0 | — |
case-23 | pass→pass | 17,144 | 15,720 | -8% | 1 | 1 | 0% | 3,848 | 7,013 | +82% | 0 | 0 | — |
case-24 | pass→pass | 15,817 | 13,361 | -16% | 1 | 1 | 0% | 2,799 | 6,004 | +115% | 0 | 0 | — |
case-04 | fail→pass | 9,020 | 5,709 | -37% | 1 | 1 | 0% | 1,802 | 4,563 | +153% | 0 | 0 | — |
case-05 | fail→pass | 8,131 | 6,572 | -19% | 1 | 1 | 0% | 1,590 | 4,742 | +198% | 0 | 0 | — |
case-06 | pass→pass | 9,491 | 7,131 | -25% | 1 | 1 | 0% | 1,933 | 4,843 | +151% | 0 | 0 | — |
case-07 | pass→pass | 7,172 | 3,592 | -50% | 1 | 1 | 0% | 1,436 | 3,982 | +177% | 0 | 0 | — |
case-08 | fail→fail | 10,651 | 5,800 | -46% | 1 | 1 | 0% | 2,105 | 4,501 | +114% | 0 | 0 | — |
case-10 | pass→pass | 6,905 | 4,511 | -35% | 1 | 1 | 0% | 1,338 | 4,361 | +226% | 0 | 0 | — |
case-11 | fail→pass | 11,507 | 2,274 | -80% | 1 | 1 | 0% | 2,221 | 3,862 | +74% | 0 | 0 | — |
case-12 | fail→fail | 18,725 | 10,058 | -46% | 1 | 1 | 0% | 3,762 | 5,697 | +51% | 0 | 0 | — |
case-13 | fail→pass | 8,238 | 6,811 | -17% | 1 | 1 | 0% | 1,599 | 4,717 | +195% | 0 | 0 | — |
case-14 | fail→pass | 8,439 | 2,587 | -69% | 1 | 1 | 0% | 1,628 | 3,843 | +136% | 0 | 0 | — |
case-15 | pass→pass | 8,256 | 2,808 | -66% | 1 | 1 | 0% | 1,413 | 3,931 | +178% | 0 | 0 | — |
case-16 | fail→pass | 10,614 | 12,272 | +16% | 1 | 1 | 0% | 2,190 | 6,285 | +187% | 0 | 0 | — |
case-17 | pass→pass | 8,731 | 6,677 | -24% | 1 | 1 | 0% | 1,693 | 4,727 | +179% | 0 | 0 | — |
case-18 | fail→pass | 4,447 | 4,238 | -5% | 1 | 1 | 0% | 752 | 4,201 | +459% | 0 | 0 | — |
case-19 | fail→pass | 6,167 | 1,560 | -75% | 1 | 1 | 0% | 1,049 | 3,665 | +249% | 0 | 0 | — |
case-20 | fail→pass | 16,102 | 8,054 | -50% | 1 | 1 | 0% | 3,655 | 5,295 | +45% | 0 | 0 | — |
case-21 | fail→pass | 9,147 | 4,758 | -48% | 1 | 1 | 0% | 1,553 | 4,283 | +176% | 0 | 0 | — |
case-22 | pass→pass | 20,605 | 16,844 | -18% | 1 | 1 | 0% | 4,084 | 6,968 | +71% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +54 percentage points is the difference between those two pass rates over the 24 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.