Cartesia
Cartesia implementation for @micdrop/server.
This package provides high-quality real-time text-to-speech implementation using Cartesia’s streaming API. Before you pick a voice engine, you can compare Cartesia and ElevenLabs for a voice agent on latency, voices, price and concurrency.
Installation
npm install @micdrop/cartesiaCartesia TTS (Text-to-Speech)
Usage with MicdropServer
import { CartesiaTTS } from '@micdrop/cartesia'import { MicdropServer } from '@micdrop/server'
const tts = new CartesiaTTS({ apiKey: process.env.CARTESIA_API_KEY || '', modelId: 'sonic-3.6', // Cartesia model ID voiceId: 'a0e99841-438c-4a64-b679-ae501e7d6091', // Voice ID language: 'en', // Optional: specify language generationConfig: { speed: 1.1 }, // Optional: speed, volume and emotion})
// Use with MicdropServernew MicdropServer(socket, { tts, // ... other options})Usage without MicdropServer
import { CartesiaTTS } from '@micdrop/cartesia'import { Readable } from 'stream'
const tts = new CartesiaTTS({ apiKey: process.env.CARTESIA_API_KEY || '', modelId: 'sonic-3.6', voiceId: 'a0e99841-438c-4a64-b679-ae501e7d6091',})
// Audio is raw PCM, 16 bits, 16 kHz, monotts.on('Audio', (chunk) => console.log('Audio:', chunk.length, 'bytes'))tts.on('Failed', (texts) => console.error('Failed:', texts))
tts.speak(Readable.from(['Hello! ', 'What can I do for you?']))Events
| Event | Payload | Description |
|---|---|---|
Audio | Buffer | A chunk of audio, PCM 16 bits, 16 kHz, mono, ready to be played. |
Failed | string[] | Synthesis gave up after its retries, with the text that stayed unspoken. |
See the TTS interface for the full contract.
Options
| Option | Type | Default | Description |
|---|---|---|---|
apiKey | string | Required | Your Cartesia API key |
modelId | string | Required | Cartesia model ID to use, such as sonic-3.6 |
voiceId | string | Required | Voice ID for speech synthesis |
language | CartesiaLanguage | Optional | Language code for speech |
generationConfig | CartesiaGenerationConfig | Optional | Speed, volume and emotion of the voice, see below |
speed | 'fast' | 'normal' | 'slow' | Optional | Deprecated, Sonic 3 models read generationConfig.speed |
connectionTimeout | number | 5000 | Timeout in milliseconds for WebSocket connection |
retryDelay | number | 1000 | Delay in milliseconds between reconnection attempts |
maxRetry | number | 3 | Maximum number of reconnection attempts before failing |
Models
Use sonic-3.6, the current Sonic model, or pin its snapshot sonic-3.6-2026-08-27 in production. Cartesia stops serving sonic-turbo and sonic-2 after October 20, 2026, so an agent configured with one of them has to switch its modelId before that date. The Cartesia model list gives every snapshot and its sunset date.
Speed, volume and emotion
generationConfig applies to every sentence the agent speaks. Cartesia treats its values as guidance and keeps the pacing natural.
| Field | Type | Range | Description |
|---|---|---|---|
speed | number | 0.6 to 1.5 | Speech speed, 1 by default |
volume | number | 0.5 to 2.0 | Volume, 1 by default |
emotion | CartesiaEmotion | Emotion the voice leans toward, such as calm or content, English only |
const tts = new CartesiaTTS({ apiKey: process.env.CARTESIA_API_KEY || '', modelId: 'sonic-3.6', voiceId: 'a0e99841-438c-4a64-b679-ae501e7d6091', generationConfig: { speed: 0.9, volume: 1.2, emotion: 'calm' },})To change the tone sentence by sentence, the LLM can write Cartesia’s emotion and speed tags into its answer, such as <emotion value="calm"/>.
Supported Languages
Sonic 3.6 speaks 44 languages:
| Code | Language | Code | Language | Code | Language |
|---|---|---|---|---|---|
en | English | fr | French | de | German |
es | Spanish | pt | Portuguese | zh | Chinese |
ja | Japanese | hi | Hindi | it | Italian |
ko | Korean | nl | Dutch | pl | Polish |
ru | Russian | sv | Swedish | tr | Turkish |
tl | Tagalog | bg | Bulgarian | ro | Romanian |
ar | Arabic | cs | Czech | el | Greek |
fi | Finnish | hr | Croatian | ms | Malay |
sk | Slovak | da | Danish | ta | Tamil |
uk | Ukrainian | hu | Hungarian | no | Norwegian |
vi | Vietnamese | bn | Bengali | th | Thai |
he | Hebrew | ka | Georgian | id | Indonesian |
te | Telugu | gu | Gujarati | kn | Kannada |
ml | Malayalam | mr | Marathi | pa | Punjabi |
or | Odia | ur | Urdu |
Getting Started
- Sign up for a Cartesia account and get your API key
- Choose a model ID and voice ID from the Cartesia dashboard
- Install the package and configure with your credentials
import { CartesiaTTS } from '@micdrop/cartesia'
const tts = new CartesiaTTS({ apiKey: 'your-cartesia-api-key', modelId: 'sonic-3.6', voiceId: 'your-preferred-voice-id', language: 'en',})
// Use with MicdropServernew MicdropServer(socket, { tts, // ... other options})