Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Automate ElevenLabs text-to-speech workflows -- generate speech from text, browse and inspect voices, check subscription limits, list models, stream audio, and retrieve history via the Composio MCP integration.
.claude/skills/composiohq-elevenlabs-automation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 165% | 0% |
Automate your ElevenLabs text-to-speech workflows -- convert text to natural speech, browse the voice library, inspect voice details, check subscription credits, select TTS models, stream audio for low-latency delivery, and retrieve previously generated audio from history.
Toolkit docs: composio.dev/toolkits/elevenlabs
https://rube.app/mcpUse ELEVENLABS_TEXT_TO_SPEECH to convert text into a downloadable audio file.
Tool: ELEVENLABS_TEXT_TO_SPEECH
Inputs:
- voice_id: string (required) -- obtain from ELEVENLABS_GET_VOICES
- text: string (required) -- max ~10,000 chars (most models), 30,000 (Flash/Turbo v2), 40,000 (v2.5)
- model_id: string (default "eleven_monolingual_v1") -- e.g., "eleven_multilingual_v2"
- output_format: string (default "mp3_44100_128") -- see formats below
- optimize_streaming_latency: integer (0-4; NOT supported with eleven_v3)
- seed: integer (optional, for reproducibility -- not guaranteed)
- pronunciation_dictionary_locators: array (optional, up to 3 dictionaries)Output formats:
mp3_22050_32, mp3_44100_32, mp3_44100_64, mp3_44100_96, mp3_44100_128, mp3_44100_192 (Creator+)pcm_16000, pcm_22050, pcm_24000, pcm_44100 (Pro+)ulaw_8000 (for Twilio)Important: Output is a file object with a presigned download link at data.file.s3url (expires in ~1 hour). Download promptly.
Use ELEVENLABS_GET_VOICES to list all voices with their attributes and settings.
Tool: ELEVENLABS_GET_VOICES
Inputs: (none)Returns an array at data.voices[] with voice_id, name, labels (gender, accent, use_case), and settings.
Use ELEVENLABS_GET_VOICE to get detailed metadata for a candidate voice before synthesis.
Tool: ELEVENLABS_GET_VOICE
Inputs:
- voice_id: string (required) -- e.g., "21m00Tcm4TlvDq8ikWAM"
- with_settings: boolean (default false) -- include detailed voice settingsUse ELEVENLABS_GET_USER_SUBSCRIPTION_INFO to verify plan limits and remaining credits before bulk generation.
Tool: ELEVENLABS_GET_USER_SUBSCRIPTION_INFO
Inputs: (none)Use ELEVENLABS_GET_MODELS to discover compatible models and filter by can_do_text_to_speech: true.
Tool: ELEVENLABS_GET_MODELS
Inputs: (none)Use ELEVENLABS_TEXT_TO_SPEECH_STREAM for low-latency streamed delivery, and ELEVENLABS_GET_AUDIO_FROM_HISTORY_ITEM to re-download previously generated audio.
Tool: ELEVENLABS_TEXT_TO_SPEECH_STREAM
- Same core inputs as TEXT_TO_SPEECH but returns a stream for low-latency playback
Tool: ELEVENLABS_GET_AUDIO_FROM_HISTORY_ITEM
- history_item_id: string (required) -- ID from a previous generation| Pitfall | Detail | |---------|--------| | Text length limits | Most models cap at ~10,000-20,000 chars per request. Oversized input returns HTTP 400. Split long text into chunks (~5000 chars) and generate per chunk. | | Output is a presigned URL | ELEVENLABS_TEXT_TO_SPEECH returns data.file.s3url with a ~1 hour expiry (X-Amz-Expires=3600). Download the audio file promptly. | | Quota and credit errors | HTTP 401 with quota_exceeded or HTTP 402 payment_required means insufficient credits or tier restrictions. Check with ELEVENLABS_GET_USER_SUBSCRIPTION_INFO before bulk jobs. | | Voice permissions | HTTP 401 with missing_permissions means the API key lacks voices_read scope. Verify key permissions. | | Model compatibility | Not all models support TTS. Use ELEVENLABS_GET_MODELS and filter by can_do_text_to_speech: true. The optimize_streaming_latency parameter is NOT supported with eleven_v3. | | Large voice list truncation | ELEVENLABS_GET_VOICES may return a large list. Select from the full data.voices[] payload -- previews may appear truncated. |
| Tool Slug | Description | |-----------|-------------| | ELEVENLABS_TEXT_TO_SPEECH | Convert text to speech, returns downloadable audio file | | ELEVENLABS_GET_VOICES | List all available voices with attributes | | ELEVENLABS_GET_VOICE | Get detailed info for a specific voice | | ELEVENLABS_GET_USER_SUBSCRIPTION_INFO | Check subscription plan and remaining credits | | ELEVENLABS_GET_MODELS | List available TTS models and capabilities | | ELEVENLABS_TEXT_TO_SPEECH_STREAM | Stream audio for low-latency delivery | | ELEVENLABS_GET_AUDIO_FROM_HISTORY_ITEM | Re-download audio from generation history |
Powered by Composio
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,996 | 2,439 | -69% | 1 | 1 | 0% | 1,850 | 1,804 | -2% | 0 | 0 | — |
case-02 | fail→fail | 2,509 | 1,967 | -22% | 1 | 1 | 0% | 465 | 1,877 | +304% | 0 | 0 | — |
case-03 | fail→pass | 8,639 | 5,556 | -36% | 1 | 1 | 0% | 1,757 | 2,791 | +59% | 0 | 0 | — |
case-04 | pass→pass | 7,188 | 2,249 | -69% | 1 | 1 | 0% | 1,252 | 1,988 | +59% | 0 | 0 | — |
case-05 | pass→pass | 5,468 | 2,338 | -57% | 1 | 1 | 0% | 1,011 | 2,015 | +99% | 0 | 0 | — |
case-06 | pass→pass | 8,070 | 2,688 | -67% | 1 | 1 | 0% | 1,688 | 2,081 | +23% | 0 | 0 | — |
case-07 | pass→pass | 6,632 | 3,357 | -49% | 1 | 1 | 0% | 1,305 | 2,217 | +70% | 0 | 0 | — |
case-08 | fail→pass | 6,125 | 2,143 | -65% | 1 | 1 | 0% | 1,345 | 2,006 | +49% | 0 | 0 | — |
case-09 | fail→pass | 11,262 | 3,176 | -72% | 1 | 1 | 0% | 2,253 | 2,169 | -4% | 0 | 0 | — |
case-10 | fail→pass | 10,809 | 3,750 | -65% | 1 | 1 | 0% | 1,965 | 2,351 | +20% | 0 | 0 | — |
case-11 | fail→fail | 9,822 | 4,871 | -50% | 1 | 1 | 0% | 1,724 | 1,948 | +13% | 0 | 0 | — |
case-12 | pass→pass | 3,942 | 2,663 | -32% | 1 | 1 | 0% | 590 | 1,945 | +230% | 0 | 0 | — |
case-13 | fail→pass | 8,661 | 1,270 | -85% | 1 | 1 | 0% | 661 | 1,754 | +165% | 0 | 0 | — |
case-14 | fail→pass | 8,722 | 2,527 | -71% | 1 | 1 | 0% | 1,701 | 1,760 | +3% | 0 | 0 | — |
case-15 | fail→fail | 8,197 | 2,142 | -74% | 1 | 1 | 0% | 1,496 | 1,921 | +28% | 0 | 0 | — |
case-16 | fail→fail | 10,332 | 2,749 | -73% | 1 | 1 | 0% | 1,881 | 2,034 | +8% | 0 | 0 | — |
case-17 | fail→fail | 9,777 | 4,365 | -55% | 1 | 1 | 0% | 1,775 | 2,377 | +34% | 0 | 0 | — |
case-18 | pass→pass | 5,446 | 8,118 | +49% | 1 | 1 | 0% | 1,062 | 2,428 | +129% | 0 | 0 | — |
case-19 | pass→pass | 9,312 | 2,724 | -71% | 1 | 1 | 0% | 1,624 | 2,046 | +26% | 0 | 0 | — |
case-20 | fail→fail | 3,610 | 6,334 | +75% | 1 | 1 | 0% | 663 | 2,782 | +320% | 0 | 0 | — |
case-21 | fail→pass | 4,490 | 4,544 | +1% | 1 | 1 | 0% | 748 | 2,448 | +227% | 0 | 0 | — |
case-22 | fail→pass | 5,752 | 6,584 | +14% | 1 | 1 | 0% | 993 | 2,798 | +182% | 0 | 0 | — |
case-23 | fail→fail | 9,324 | 7,300 | -22% | 1 | 1 | 0% | 1,576 | 2,842 | +80% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +35 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.