Piper
Local text to speech for @micdrop/server, driving the Piper binary as a subprocess.
Piper voices are small VITS models, a few dozen megabytes each, covering around forty languages including French. They generate much faster than real time even on a modest CPU, which makes Piper the option to reach for outside English. In English itself, Kokoro sounds better, and English is all it speaks in Micdrop.
Installation
Install the package:
npm install @micdrop/piperInstall the binary:
pip install piper-ttsThen download a voice. It comes as a pair of files that have to sit next to each other:
mkdir -p voices && cd voicesBASE=https://huggingface.co/rhasspy/piper-voices/resolve/main/fr/fr_FR/siwis/mediumcurl -LO $BASE/fr_FR-siwis-medium.onnxcurl -LO $BASE/fr_FR-siwis-medium.onnx.jsonPiper ships a download_voices module, but calling it means finding the exact
interpreter pip installed it under, which is rarely the python3 in your
path. Downloading the two files is shorter and gives the same result.
Usage with MicdropServer
import { PiperTTS } from '@micdrop/piper'import { MicdropServer } from '@micdrop/server'
const tts = new PiperTTS({ modelPath: './voices/fr_FR-siwis-medium.onnx',})
// Use with MicdropServernew MicdropServer(socket, { tts, // ... other options})Usage without MicdropServer
import { PiperTTS } from '@micdrop/piper'import { Readable } from 'stream'
const tts = new PiperTTS({ modelPath: './voices/fr_FR-siwis-medium.onnx',})
// Audio is raw PCM, 16 bits, 16 kHz, monotts.on('Audio', (chunk) => console.log('Audio:', chunk.length, 'bytes'))tts.on('Failed', (texts) => console.error('Failed:', texts))
tts.speak(Readable.from(['Hello! ', 'What can I do for you?']))Events
| Event | Payload | Description |
|---|---|---|
Audio | Buffer | A chunk of audio, PCM 16 bits, 16 kHz, mono, ready to be played. |
Failed | string[] | The Piper process exited with an error. The payload stays empty. |
See the TTS interface for the full contract.
Options
| Option | Type | Default | Description |
|---|---|---|---|
modelPath | string | Required | Path to the .onnx voice file |
configPath | string | <modelPath>.json | Voice configuration |
binaryPath | string | 'piper' in PATH | Piper executable |
speaker | number | Optional | Speaker index, for voices with several speakers |
lengthScale | number | Optional | Duration multiplier, above 1 to slow the voice down |
noiseScale | number | Optional | Variability of the generated speech |
noiseWidth | number | Optional | Variability of the phoneme durations |
sentenceSilence | number | Optional | Silence added after each sentence, in seconds |
Voices
Voices are named after their language, their speaker and their quality, such as
fr_FR-siwis-medium or en_GB-alba-medium. The
voice samples page lets you listen
to them before downloading.
Each voice ships as <name>.onnx and <name>.onnx.json. The JSON file gives
the sample rate the model generates at, and Micdrop reads that file at startup
so the audio reaches the client at the rate it expects.
One process for the whole call
Loading a voice takes about a second, and synthesizing a sentence adds almost
nothing on top of it. PiperTTS starts the process as soon as you construct it
and keeps it alive across sentences, so the voice loads once per call.
An interruption kills the process, which cuts the audio at once, and starts its replacement right away. The new process loads its voice while the user is still speaking and is ready by the time the next answer begins.
The whole answer is usually ready before the user has finished hearing its first sentence.
Documentation
Read the guide on running Micdrop with local models to see how Piper fits with a local LLM and local speech to text. The comparison of local TTS engines in Node measures Piper against Kokoro and Pocket TTS on latency, memory and languages.
License
The Piper engine is released under GPL-3.0 since it embedded espeak-ng. The
archived rhasspy/piper repository was MIT. Micdrop drives the binary as a
separate process, and redistributing that binary yourself brings the GPL terms
with it. Each voice carries its own license, some of them restrictive, so read
the model card of the voice you pick.