🎤Micdrop

ElevenLabs

@micdrop/elevenlabs streams ElevenLabs text to speech for @micdrop/server. It sends the agent’s answer to ElevenLabs over a WebSocket as the agent writes it, so the audio comes back before the answer is finished. The package works with Eleven v4 Turbo, Eleven v4, Eleven v3, Flash v2.5, Turbo v2.5 and Multilingual v2, and picks the WebSocket each model needs.

ElevenLabs and every other voice engine in Micdrop implement the same interface, so moving to an ElevenLabs alternative only means replacing new ElevenLabsTTS(...) with the other engine’s class.

Installation

Terminal window
npm install @micdrop/elevenlabs

ElevenLabs TTS (Text-to-Speech)

Usage with MicdropServer

import { ElevenLabsTTS } from '@micdrop/elevenlabs'
import { MicdropServer } from '@micdrop/server'
const tts = new ElevenLabsTTS({
apiKey: process.env.ELEVENLABS_API_KEY || '',
voiceId: '21m00Tcm4TlvDq8ikWAM', // ElevenLabs voice ID
modelId: 'eleven_v4_turbo', // Optional: model to use
language: 'en', // Optional: language code
voiceSettings: {
stability: 0.5,
},
})
// Use with MicdropServer
new MicdropServer(socket, {
tts,
// ... other options
})

Usage without MicdropServer

import { ElevenLabsTTS } from '@micdrop/elevenlabs'
import { Readable } from 'stream'
const tts = new ElevenLabsTTS({
apiKey: process.env.ELEVENLABS_API_KEY || '',
voiceId: '21m00Tcm4TlvDq8ikWAM',
})
// Audio is raw PCM, 16 bits, 16 kHz, mono
tts.on('Audio', (chunk) => console.log('Audio:', chunk.length, 'bytes'))
tts.on('Failed', (texts) => console.error('Failed:', texts))
tts.speak(Readable.from(['Hello! ', 'What can I do for you?']))

Events

EventPayloadDescription
AudioBufferA chunk of audio, PCM 16 bits, 16 kHz, mono, ready to be played.
Failedstring[]The text left unspoken when ElevenLabsTTS gives up after its retries.

See the TTS interface for the full contract.

Choosing a model

ElevenLabs streams speech over two WebSockets, each with its own protocol. Eleven v3 and v4 only stream over the Text to Dialogue WebSocket, while Flash v2.5, Turbo v2.5 and Multilingual v2 use the Text to Speech one. ElevenLabsTTS opens the right one for the modelId you pass, so moving from Flash to v4 Turbo only means changing the model ID.

ModelWebSocketLanguagesAudio tagsBest for
eleven_v4_turboText to Dialogue90+PerformedExpressive voice agents, with a median inference latency of about 100 ms
eleven_v4Text to Dialogue90+PerformedNarration and characters, where the quality of the voice comes first
eleven_v3_conversationalText to Dialogue70+PerformedExpressive real-time speech, the generation before v4 Turbo
eleven_v3Text to Dialogue70+PerformedExpressive speech, the generation before v4
eleven_flash_v2_5Text to Speech32Read outThe lowest latency, about 75 ms
eleven_multilingual_v2Text to Speech29Read outLong narration, where the voice stays the most stable
eleven_turbo_v2_5Text to Speech32Read outExisting servers, since ElevenLabs deprecated it in favor of Flash v2.5

The latencies are the ones ElevenLabs gives, network excluded. A Micdrop test measured the first audio in a streamed voice agent, network included: about 90 ms on v4 Turbo and about 850 ms on v3. eleven_turbo_v2_5 stays the default of the package, so an existing server keeps its voice until you pick another model.

Eleven v4 and audio tags

Eleven v3 and v4 perform audio tags, short directions in square brackets. A tag like [whispers], [laughs] or [excited] goes right before the words it changes. Eleven v4 also plays sound effects like [door slams] and follows free directions like [strong French accent]. ElevenLabs lists the tags it tested with Eleven v4.

The LLM writes the tags, so tell it in the system prompt which ones it may use:

import { OpenaiAgent } from '@micdrop/openai'
const agent = new OpenaiAgent({
apiKey: process.env.OPENAI_API_KEY || '',
systemPrompt: `You are a warm and playful assistant. Your answers are spoken out loud.
Your voice performs audio tags: short directions in square brackets, placed right before the words they change. Use one or two per answer, where they fit:
- How you speak: [whispers], [excited], [curious], [sarcastic]
- How you react: [laughs], [sighs], [clears throat]
Use an ellipsis for a pause. The words you say always stay outside the brackets.`,
})

The tags reach the client with the rest of the answer, in the conversation and in the partial messages. Run the storyteller demo to hear what the tags change. Its page shows each tag apart from the words, and lets you switch the voice to Flash v2.5, where the same storyteller speaks without tags.

Only Eleven v3 and v4 perform the tags. Flash v2.5, Turbo v2.5 and Multilingual v2 read them out loud, so leave the tag instructions out of the system prompt when the voice runs on one of those models. When one of those models only serves as a backup voice, keep the tag instructions and strip the audio tags from the text it receives.

How Eleven v3 and v4 stream

The Text to Dialogue WebSocket holds text back until it has about 40 characters and 8 words. ElevenLabsTTS flushes it at the end of every sentence, so a short first sentence such as “Sure!” starts playing without waiting for the next one. The package counts an ellipsis as part of the sentence, since Eleven v4 performs it as a pause.

The Text to Dialogue WebSocket has no message to stop a synthesis halfway. When the user interrupts the assistant, ElevenLabsTTS closes the connection and opens a new one while the user is still speaking, so the interrupted answer stops where the user cut it. ElevenLabs closes a connection after 20 seconds without a message, so the package sends a keep-alive every 15 seconds.

ElevenLabs counts each connection to the Text to Dialogue WebSocket as one dialogue session for as long as it stays open, speaking or silent, so each call holds one session from start to end. ElevenLabs allows 14 sessions at once on its free plan and 105 on Scale, and publishes the session limit of each plan. Past the limit, it rejects new connections with too_many_concurrent_requests. ElevenLabs counts only the time spent generating audio on the Text to Speech WebSocket, so the same plan usually allows more calls at once with Flash.

Options

OptionTypeDefaultDescription
apiKeystringRequiredYour ElevenLabs API key
voiceIdstringRequiredElevenLabs voice ID
modelIdElevenLabsModelId'eleven_turbo_v2_5'Model to use for speech synthesis, one of the supported models
languagestringOptionalLanguage code (e.g., ‘en’, ‘fr’)
outputFormatTextToSpeechStreamRequestOutputFormat'pcm_16000'Audio output format
voiceSettingsVoiceSettingsOptionalVoice customization settings. Eleven v3 and v4 only read stability
connectionTimeoutnumber5000Timeout in milliseconds for WebSocket connection
retryDelaynumber1000Delay in milliseconds between reconnection attempts
maxRetrynumber3Maximum number of reconnection attempts before failing

Voice Settings

The voiceSettings option takes the VoiceSettings type of the ElevenLabs SDK, with its keys in camel case:

const tts = new ElevenLabsTTS({
apiKey: 'your-api-key',
voiceId: 'your-voice-id',
modelId: 'eleven_flash_v2_5',
voiceSettings: {
stability: 0.5, // 0 to 1, lower is more expressive, higher is steadier
similarityBoost: 0.75, // 0 to 1, how closely to match the original voice
style: 0.5, // 0 to 1, exaggeration of the style
useSpeakerBoost: true, // Boost the similarity to the speaker
speed: 1, // 1 is the normal pace
},
})

Eleven v3 and v4 only read stability. A lower value gives the voice a broader emotional range, so it responds more to the audio tags.

Supported Languages

The languages depend on the model: more than 90 for Eleven v4 and v4 Turbo, more than 70 for Eleven v3, 32 for Flash v2.5 and Turbo v2.5, and 29 for Multilingual v2. ElevenLabs keeps the list of languages of each model. Every model speaks these languages:

CodeLanguageCodeLanguageCodeLanguage
enEnglishesSpanishfrFrench
deGermanitItalianptPortuguese
plPolishtrTurkishruRussian
nlDutchcsCzecharArabic
zhChinesejaJapanesesvSwedish
koKoreanhiHindifiFinnish

Getting Started

  1. Sign up for an ElevenLabs account and get your API key
  2. Choose a voice from the ElevenLabs voice library or create a custom voice
  3. Install the package and configure with your credentials
import { ElevenLabsTTS } from '@micdrop/elevenlabs'
const tts = new ElevenLabsTTS({
apiKey: 'your-elevenlabs-api-key',
voiceId: 'your-voice-id', // Get this from ElevenLabs dashboard
modelId: 'eleven_v4_turbo', // Choose based on your needs
language: 'en',
})
// Use with MicdropServer
new MicdropServer(socket, {
tts,
// ... other options
})

Finding Voice IDs

You can find voice IDs in the ElevenLabs dashboard, by browsing the voices of your account, or through the public voice library API.