Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or adopting SpeechAnalyzer, SpeechTranscriber, DictationTranscriber, AssetInventory, and async result streams on iOS 26+.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 127% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 260% | 0% |
Transcribe live and pre-recorded audio to text using Apple's Speech framework. Covers SpeechAnalyzer / SpeechTranscriber (iOS 26+) and SFSpeechRecognizer (iOS 10+) fallback guidance.
Scope boundary: Use this skill for speech-to-text recognition, speech authorization, microphone capture plumbing, and result handling. Hand off text analysis, language identification after transcription, sentiment, embeddings, and translation to natural-language; hand off audio playback UI to avkit; hand off summarization or generation over transcripts to apple-on-device-ai.
Use SpeechAnalyzer for modern iOS 26+ speech analysis, especially long-form recordings, live transcription, time-indexed transcripts, and fully on-device flows. Keep SFSpeechRecognizer for iOS 10+ deployment targets, server-backed locale coverage, or existing callback/delegate implementations.
Read SpeechAnalyzer patterns when implementing an iOS 26+ transcription pipeline, model asset handling, volatile results, or file/buffer examples.
SpeechTranscriber for the newer general-purpose on-device model.DictationTranscriber when SpeechTranscriber is unavailable for thecurrent device or locale and dictation-compatible support is acceptable.
SpeechDetector only in conjunction with a transcriber when voiceactivity detection is worth the accuracy/power tradeoff.
SpeechTranscriber.isAvailableSpeechTranscriber.supportedLocale(equivalentTo:)SpeechTranscriber.installedLocales / supportedLocales when showinglanguage choices.
.transcription for basic accurate transcription..progressiveTranscription for live UI updates..timeIndexedProgressiveTranscription when playback highlighting needsaudioTimeRange.
AssetInventory.assetInstallationRequest.SpeechAnalyzer.bestAvailableAudioFormat(compatibleWith:) before yielding AnalyzerInput.
AsyncSequence in a separate task.finalizeAndFinish(through:),finalizeAndFinishThroughEndOfInput(), or cancelAndFinishNow().
Do not use an offlineTranscription preset; Apple does not document one. Finishing an AsyncStream input sequence does not finish the analyzer session.
swiftimport Speech // Default locale (user's current language) let recognizer = SFSpeechRecognizer() // Specific locale let recognizer = SFSpeechRecognizer(locale: Locale(identifier: "en-US")) // Check if recognition is available for this locale guard let recognizer, recognizer.isAvailable else { print("Speech recognition not available") return }
swiftfinal class SpeechManager: NSObject, SFSpeechRecognizerDelegate { private let recognizer = SFSpeechRecognizer()! override init() { super.init() recognizer.delegate = self } func speechRecognizer( _ speechRecognizer: SFSpeechRecognizer, availabilityDidChange available: Bool ) { // Update UI — disable record button when unavailable } }
Request both speech recognition and microphone permissions before starting live transcription. Add these keys to Info.plist:
NSSpeechRecognitionUsageDescriptionNSMicrophoneUsageDescriptionswiftimport Speech import AVFoundation func requestPermissions() async -> Bool { let speechStatus = await withCheckedContinuation { continuation in SFSpeechRecognizer.requestAuthorization { status in continuation.resume(returning: status) } } guard speechStatus == .authorized else { return false } let micStatus: Bool if #available(iOS 17, *) { micStatus = await AVAudioApplication.requestRecordPermission() } else { micStatus = await withCheckedContinuation { continuation in AVAudioSession.sharedInstance().requestRecordPermission { granted in continuation.resume(returning: granted) } } } return micStatus }
The standard pattern: AVAudioEngine captures microphone audio → buffers are appended to SFSpeechAudioBufferRecognitionRequest → results stream in.
swiftimport Speech import AVFoundation final class LiveTranscriber { private let recognizer = SFSpeechRecognizer(locale: Locale(identifier: "en-US"))! private let audioEngine = AVAudioEngine() private var recognitionRequest: SFSpeechAudioBufferRecognitionRequest? private var recognitionTask: SFSpeechRecognitionTask? func startTranscribing() throws { // Cancel any in-progress task recognitionTask?.cancel() recognitionTask = nil // Configure audio session let audioSession = AVAudioSession.sharedInstance() try audioSession.setCategory(.record, mode: .measurement, options: .duckOthers) try audioSession.setActive(true, options: .notifyOthersOnDeactivation) // Create request let request = SFSpeechAudioBufferRecognitionRequest() request.shouldReportPartialResults = true self.recognitionRequest = request // Start recognition task recognitionTask = recognizer.recognitionTask(with: request) { result, error in if let result { let text = result.bestTranscription.formattedString print("Transcription: \(text)") if result.isFinal { self.stopTranscribing() } } if let error { print("Recognition error: \(error)") self.stopTranscribing() } } // Install audio tap let inputNode = audioEngine.inputNode let recordingFormat = inputNode.outputFormat(forBus: 0) inputNode.installTap(onBus: 0, bufferSize: 1024, format: recordingFormat) { buffer, _ in request.append(buffer) } audioEngine.prepare() try audioEngine.start() } func stopTranscribing() { audioEngine.stop() audioEngine.inputNode.removeTap(onBus: 0) recognitionRequest?.endAudio() recognitionRequest = nil recognitionTask?.cancel() recognitionTask = nil } }
Use SFSpeechURLRecognitionRequest for audio files on disk:
swiftfunc transcribeFile(at url: URL) async throws -> String { guard let recognizer = SFSpeechRecognizer(), recognizer.isAvailable else { throw SpeechError.unavailable } let request = SFSpeechURLRecognitionRequest(url: url) request.shouldReportPartialResults = false return try await withCheckedThrowingContinuation { continuation in var didResume = false recognizer.recognitionTask(with: request) { result, error in guard !didResume else { return } if let error { didResume = true continuation.resume(throwing: error) } else if let result, result.isFinal { didResume = true continuation.resume( returning: result.bestTranscription.formattedString ) } } } }
SFSpeechRecognizer can use on-device recognition for supported locales on iOS 13+. If supportsOnDeviceRecognition is false, the recognizer requires a network connection. requiresOnDeviceRecognition only has effect when the recognizer supports it.
swiftlet recognizer = SFSpeechRecognizer(locale: Locale(identifier: "en-US"))! // Check if on-device is supported for this locale if recognizer.supportsOnDeviceRecognition { let request = SFSpeechAudioBufferRecognitionRequest() request.requiresOnDeviceRecognition = true // Force on-device }
SFSpeechRecognizer requests may still be a poor fit for long-form capture. Apple documents a roughly one-minute task limit for speech recognition and other service limits. For long recordings on iOS 26+, prefer SpeechAnalyzer; otherwise chunk or restart recognition before the limit and preserve transcript state across tasks.
With shouldReportPartialResults, replace the displayed partial transcript until SFSpeechRecognitionResult.isFinal commits it. This is separate from SpeechTranscriber.Result.isFinal, whose volatile attributed range must be replaced by the final result for that range. The live example and references/speechanalyzer-patterns.md contain the canonical loops.
swiftrecognizer.recognitionTask(with: request) { result, error in guard let result else { return } // Best transcription let best = result.bestTranscription // All alternatives (sorted by confidence, descending) for transcription in result.transcriptions { for segment in transcription.segments { print("\(segment.substring): \(segment.confidence)") } } }
swiftlet request = SFSpeechAudioBufferRecognitionRequest() request.addsPunctuation = true
Improve recognition of domain-specific terms:
swiftlet request = SFSpeechAudioBufferRecognitionRequest() request.contextualStrings = ["SwiftUI", "Xcode", "CloudKit"]
| Mistake | Fix | |---|---| | Live audio requests only speech authorization | Require both speech and microphone permission. | | Recognizer availability is checked once | Observe delegate availability changes and model loss of service. | | Final/error/rollover paths perform different cleanup | Use one stopTranscribing() owner for engine stop, tap removal, endAudio, and task cancellation. | | On-device mode is forced for every locale | Check supportsOnDeviceRecognition and provide a fallback. | | One SFSpeechRecognizer task is used for long-form capture | Prefer SpeechAnalyzer on iOS 26+ or roll bounded segments while preserving committed transcript. | | Finishing analyzer input is treated as session completion | Explicitly finalize through the last sample or cancel-and-finish. | | Volatile SpeechAnalyzer results are appended | Replace the volatile range until a final result commits it. | | A second recognition task starts before the first ends | Cancel/finish the active task and complete cleanup before replacement. |
Load references/speechanalyzer-patterns.md for complete analyzer finalization and volatile-result code.
NSSpeechRecognitionUsageDescription is in Info.plistNSMicrophoneUsageDescription is in Info.plist (if using live audio)SFSpeechRecognizerDelegate is set to handle availabilityDidChangerecognitionRequest.endAudio() is called when done recordingrecognitionTask is canceled before starting a new onesupportsOnDeviceRecognition is checked before requiring on-device modeisFinal) resultsSFSpeechRecognizer one-minute/service limits are accounted forAssetInventory assets are installed before using SpeechAnalyzerSpeechTranscriber.isAvailable and locale support are checkedOther measured skills in the registry, with their headline benchmark lift.