🎤Micdrop

Cartesia

Cartesia implementation for @micdrop/server.

This package provides high-quality real-time text-to-speech implementation using Cartesia’s streaming API. Before you pick a voice engine, you can compare Cartesia and ElevenLabs for a voice agent on latency, voices, price and concurrency.

Installation

Terminal window
npm install @micdrop/cartesia

Cartesia TTS (Text-to-Speech)

Usage with MicdropServer

import { CartesiaTTS } from '@micdrop/cartesia'
import { MicdropServer } from '@micdrop/server'
const tts = new CartesiaTTS({
apiKey: process.env.CARTESIA_API_KEY || '',
modelId: 'sonic-3.6', // Cartesia model ID
voiceId: 'a0e99841-438c-4a64-b679-ae501e7d6091', // Voice ID
language: 'en', // Optional: specify language
generationConfig: { speed: 1.1 }, // Optional: speed, volume and emotion
})
// Use with MicdropServer
new MicdropServer(socket, {
tts,
// ... other options
})

Usage without MicdropServer

import { CartesiaTTS } from '@micdrop/cartesia'
import { Readable } from 'stream'
const tts = new CartesiaTTS({
apiKey: process.env.CARTESIA_API_KEY || '',
modelId: 'sonic-3.6',
voiceId: 'a0e99841-438c-4a64-b679-ae501e7d6091',
})
// Audio is raw PCM, 16 bits, 16 kHz, mono
tts.on('Audio', (chunk) => console.log('Audio:', chunk.length, 'bytes'))
tts.on('Failed', (texts) => console.error('Failed:', texts))
tts.speak(Readable.from(['Hello! ', 'What can I do for you?']))

Events

EventPayloadDescription
AudioBufferA chunk of audio, PCM 16 bits, 16 kHz, mono, ready to be played.
Failedstring[]Synthesis gave up after its retries, with the text that stayed unspoken.

See the TTS interface for the full contract.

Options

OptionTypeDefaultDescription
apiKeystringRequiredYour Cartesia API key
modelIdstringRequiredCartesia model ID to use, such as sonic-3.6
voiceIdstringRequiredVoice ID for speech synthesis
languageCartesiaLanguageOptionalLanguage code for speech
generationConfigCartesiaGenerationConfigOptionalSpeed, volume and emotion of the voice, see below
speed'fast' | 'normal' | 'slow'OptionalDeprecated, Sonic 3 models read generationConfig.speed
connectionTimeoutnumber5000Timeout in milliseconds for WebSocket connection
retryDelaynumber1000Delay in milliseconds between reconnection attempts
maxRetrynumber3Maximum number of reconnection attempts before failing

Models

Use sonic-3.6, the current Sonic model, or pin its snapshot sonic-3.6-2026-08-27 in production. Cartesia stops serving sonic-turbo and sonic-2 after October 20, 2026, so an agent configured with one of them has to switch its modelId before that date. The Cartesia model list gives every snapshot and its sunset date.

Speed, volume and emotion

generationConfig applies to every sentence the agent speaks. Cartesia treats its values as guidance and keeps the pacing natural.

FieldTypeRangeDescription
speednumber0.6 to 1.5Speech speed, 1 by default
volumenumber0.5 to 2.0Volume, 1 by default
emotionCartesiaEmotionEmotion the voice leans toward, such as calm or content, English only
const tts = new CartesiaTTS({
apiKey: process.env.CARTESIA_API_KEY || '',
modelId: 'sonic-3.6',
voiceId: 'a0e99841-438c-4a64-b679-ae501e7d6091',
generationConfig: { speed: 0.9, volume: 1.2, emotion: 'calm' },
})

To change the tone sentence by sentence, the LLM can write Cartesia’s emotion and speed tags into its answer, such as <emotion value="calm"/>.

Supported Languages

Sonic 3.6 speaks 44 languages:

CodeLanguageCodeLanguageCodeLanguage
enEnglishfrFrenchdeGerman
esSpanishptPortuguesezhChinese
jaJapanesehiHindiitItalian
koKoreannlDutchplPolish
ruRussiansvSwedishtrTurkish
tlTagalogbgBulgarianroRomanian
arArabiccsCzechelGreek
fiFinnishhrCroatianmsMalay
skSlovakdaDanishtaTamil
ukUkrainianhuHungariannoNorwegian
viVietnamesebnBengalithThai
heHebrewkaGeorgianidIndonesian
teTeluguguGujaratiknKannada
mlMalayalammrMarathipaPunjabi
orOdiaurUrdu

Getting Started

  1. Sign up for a Cartesia account and get your API key
  2. Choose a model ID and voice ID from the Cartesia dashboard
  3. Install the package and configure with your credentials
import { CartesiaTTS } from '@micdrop/cartesia'
const tts = new CartesiaTTS({
apiKey: 'your-cartesia-api-key',
modelId: 'sonic-3.6',
voiceId: 'your-preferred-voice-id',
language: 'en',
})
// Use with MicdropServer
new MicdropServer(socket, {
tts,
// ... other options
})