🎤Micdrop

Local Models

Micdrop splits a call into an agent, a transcription and a voice, and each one has a local counterpart. Run them all locally and the whole conversation stays on your machine, with no API key and no network call.

This page sets up a complete local stack. Five other pages cover the rest:

  • Local LLM helps you choose the LLM and lists the mistakes local models make.
  • Local STT helps you choose the Whisper checkpoint for your language.
  • Local TTS compares the four local engines.
  • Latency and memory gives the delays and memory of a local call, and explains what runs on the CPU and on the GPU.
  • Explorations lists the models that were tested and not kept, and why.

All the figures in this section were measured on one machine, described on the latency and memory page, and the recommendations are made for that machine.

What runs locally

PartPackageRuns
Agent@micdrop/ai-sdkAny local server speaking the OpenAI protocol
Speech to text@micdrop/whisperWhisper, in your Node process
Text to speech@micdrop/kokoroKokoro, in your Node process, English only
Text to speech@micdrop/piperPiper, as a subprocess, around forty languages
Text to speech@micdrop/pocket-ttsPocket TTS, in your Node process, English only, cloning a voice
Text to speech@micdrop/qwen-ttsQwen3-TTS, on an mlx-audio server, ten languages

An npm install is enough for Whisper and Kokoro. They download their weights on first use and run inside Node. The LLM needs a server such as Ollama, Qwen3-TTS needs an mlx-audio server, Piper needs its binary, and Pocket TTS needs its archive extracted somewhere the server can read.

Setting it up

Install Ollama and pull a model:

Terminal window
brew install ollama
brew services start ollama
ollama pull qwen3:4b-instruct

ollama pull sends a request to the Ollama daemon, so the daemon has to be up first. brew services start puts it in the background and keeps it across reboots. ollama serve also works, but it runs in the foreground until you stop it.

The local LLM page explains why the tag matters, and which small models answer without reasoning first.

Install the Micdrop packages:

Terminal window
npm install @micdrop/ai-sdk @micdrop/whisper @micdrop/kokoro @ai-sdk/openai

Then assemble the call:

import { createOpenAI } from '@ai-sdk/openai'
import { AiSdkAgent } from '@micdrop/ai-sdk'
import { KokoroTTS } from '@micdrop/kokoro'
import { MicdropServer } from '@micdrop/server'
import { WhisperSTT } from '@micdrop/whisper'
// Ollama serves the OpenAI protocol on /v1
const ollama = createOpenAI({
baseURL: 'http://localhost:11434/v1',
apiKey: 'ollama', // Unused, the SDK refuses to start without one
})
new MicdropServer(socket, {
// .chat() rather than the provider itself: the default of the OpenAI
// provider is the Responses API, which a local server does not serve
agent: new AiSdkAgent({
model: ollama.chat('qwen3:4b-instruct'),
systemPrompt: 'You are a helpful voice assistant.',
}),
stt: new WhisperSTT({
model: 'base',
language: 'en',
}),
tts: new KokoroTTS({
voice: 'britishFemale',
}),
})

LM Studio and the llama.cpp server answer on the same routes, so point baseURL at their port to run them instead of Ollama.

First run

Each part downloads its weights before it can answer, from 45 MB for the smallest transcription checkpoint to several gigabytes for an LLM. Some download on first use, others during the install, and each integration page says which. Run one call before a demo, so that the downloads are already done.

Licenses

The recommended stack only uses models with permissive licenses, so you can use it commercially. This was a choice made for these integrations, and not every local model allows it. Each integration page gives the license of its model. Check the license of any model you add, since licenses can change.

Most restrictions concern voice cloning. A permissive license covers the weights, not the voice: cloning someone’s voice requires their consent. Several of the best cloning models are for non-commercial use only.

Trying it

The demo in examples/advanced exposes every provider in the selects at the top of the page, local ones included. A provider whose key, binary or server is missing is greyed out, instead of failing during the call. You can use the demo to compare a local stack with a hosted one on the same conversation.