Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Fully local, offline text-to-speech using Kyutai's Pocket TTS model. Generate high-quality audio from text without any API calls or internet connection. Features 8 built-in voices, voice cloning support, and runs entirely on CPU.
.claude/skills/sundial-org-pocket-tts/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -29% | 0% |
Fully local, offline text-to-speech using Kyutai's Pocket TTS model. Generate high-quality audio from text without any API calls or internet connection. Features 8 built-in voices, voice cloning support, and runs entirely on CPU.
bash# 1. Accept the model license on Hugging Face # https://huggingface.co/kyutai/pocket-tts # 2. Install the package pip install pocket-tts # Or use uv for automatic dependency management uvx pocket-tts generate "Hello world"
bash# Basic usage pocket-tts "Hello, I am your AI assistant" # With specific voice pocket-tts "Hello" --voice alba --output hello.wav # With custom voice file (voice cloning) pocket-tts "Hello" --voice-file myvoice.wav --output output.wav # Adjust speed pocket-tts "Hello" --speed 1.2 # Start local server pocket-tts --serve # List available voices pocket-tts --list-voices
pythonfrom pocket_tts import TTSModel import scipy.io.wavfile # Load model tts_model = TTSModel.load_model() # Get voice state voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) # Generate audio audio = tts_model.generate_audio(voice_state, "Hello world!") # Save to WAV scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) # Check sample rate print(f"Sample rate: {tts_model.sample_rate} Hz")
| Voice | Description | |-------|-------------| | alba | Casual female voice | | marius | Male voice | | javert | Clear male voice | | jean | Natural male voice | | fantine | Female voice | | cosette | Female voice | | eponine | Female voice | | azelma | Female voice |
Or use --voice-file /path/to/wav.wav for custom voice cloning.
| Option | Description | Default | |--------|-------------|---------| | text | Text to convert | Required | | -o, --output | Output WAV file | output.wav | | -v, --voice | Voice preset | alba | | -s, --speed | Speech speed (0.5-2.0) | 1.0 | | --voice-file | Custom WAV for cloning | None | | --serve | Start HTTP server | False | | --list-voices | List all voices | False |
Other measured skills in the registry, with their headline benchmark lift.