🎤Micdrop

Let your users talk to your app

Micdrop streams your users’ speech to your Node server and plays the answer from the AI models you choose.

npm install @micdrop/web @micdrop/react
import { Micdrop } from '@micdrop/web'
import { useMicdropState } from '@micdrop/react'
const start = () => Micdrop.start({ url: 'wss://your-server.com' })
function Call() {
const { isStarted, conversation } = useMicdropState()
return (
<>
<button onClick={isStarted ? Micdrop.stop : start}>
{isStarted ? 'Hang up' : 'Start the call'}
</button>
{conversation.map((item, i) =>
'content' in item && (
<p key={i}>{item.role}: {item.content}</p>
)
)}
</>
)
}

Hear Micdrop hold a conversation

Every demo is open source in the repository, ready to run with your own API keys.

Micdrop handles what happens between two sentences

A voice agent feels slow or rude when it cuts you off mid-thought, waits too long to answer, or talks over you.

  1. The browser hears you

    Voice activity detection runs in the browser, which streams the audio to your server while you are still talking.

  2. The agent waits through a pause

    Turn detection sees that the sentence is unfinished, so the agent waits for you to say the day.

  3. The agent runs your code

    Before it answers, the agent calls a tool you wrote, which runs in your Node server next to your database.

  4. The voice starts early

    Text to speech reads the first words of the answer while the model is still writing the rest.

  5. You can interrupt the assistant

    Playback stops as soon as you talk. Echo cancellation keeps the assistant’s voice out of your microphone.

Your app code stays the same whichever models answer

import { MicdropServer } from '@micdrop/server'
import { GladiaSTT } from '@micdrop/gladia'
import { OpenaiAgent } from '@micdrop/openai'
import { ElevenLabsTTS } from '@micdrop/elevenlabs'
new MicdropServer(socket, {
stt: new GladiaSTT({ apiKey: GLADIA_KEY }),
agent: new OpenaiAgent({ apiKey: OPENAI_KEY, systemPrompt }),
tts: new ElevenLabsTTS({ apiKey: ELEVENLABS_KEY, voiceId }),
})

Each stage streams into the next. To swap a provider, you change its line. Read the docs →

import { MicdropServer } from '@micdrop/server'
import { GeminiLive } from '@micdrop/gemini'
new MicdropServer(socket, {
realtime: new GeminiLive({
apiKey: GEMINI_KEY,
systemPrompt,
}),
})

One model hears the audio and answers in audio. Micdrop still handles the turns, the interruptions and the order of the transcript. Read the docs →

import { MicdropServer } from '@micdrop/server'
import { GladiaSTT } from '@micdrop/gladia'
import { MistralAgent } from '@micdrop/mistral'
import { GradiumTTS } from '@micdrop/gradium'
new MicdropServer(socket, {
stt: new GladiaSTT({ apiKey: GLADIA_KEY }),
agent: new MistralAgent({ apiKey: MISTRAL_KEY, systemPrompt }),
tts: new GradiumTTS({ apiKey: GRADIUM_KEY, voiceId, region: 'eu' }),
})

Three European providers keep the audio and the conversation inside the EU. They handle French and 90 other languages. Read the docs →

import { MicdropServer } from '@micdrop/server'
import { WhisperSTT } from '@micdrop/whisper'
import { AiSdkAgent } from '@micdrop/ai-sdk'
import { KokoroTTS } from '@micdrop/kokoro'
import { createOpenAI } from '@ai-sdk/openai'
// Ollama serves the OpenAI protocol on /v1
const ollama = createOpenAI({ baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' })
new MicdropServer(socket, {
stt: new WhisperSTT({ model: 'base' }),
agent: new AiSdkAgent({ model: ollama.chat('qwen3:4b-instruct'), systemPrompt }),
tts: new KokoroTTS({ voice: 'britishFemale' }),
})

Whisper and Kokoro run inside your Node process while Ollama runs the agent. The call works offline and needs no API key. The audio stays on your machine. Read the docs →

Micdrop is a library that runs inside your server

MicdropHosted platformsVapi, Retell AIAgent frameworksPipecat, LiveKit Agents
Where the call runsIn the Node server you already deployOn the vendor’s serversIn an agent service next to your app, often behind a WebRTC server
LanguagesTypeScript, from the browser and the phone to the serverAny, through a REST API and webhooksPython first. LiveKit Agents also has a Node SDK.
Your agent logicA class in your codebase, next to your database and your authA public endpoint that the platform callsCode in the agent process, apart from your app
CostFree under the MIT license. You pay each AI provider directly.A fee per minute on top of the providers, $0.05 on VapiFree and open source. You pay for the servers that carry the audio, or for a cloud plan.
Phone numbers and dashboardYou add them yourself. Micdrop covers calls in web and mobile apps.Included, with compliance certifications and an SLATelephony through SIP or Twilio

Read the detailed comparisons withVapiRetell AIPipecatLiveKit Agents

Raconte runs its voice interviews on Micdrop

With Raconte, a team sends a link, an AI interviews each person by voice, and the team gets back transcripts, sentiment and insights.

Each interview runs on Micdrop in the browser, with Gladia, Mistral, Gradium, OpenAI and ElevenLabs on the server. Raconte and Micdrop have the same author.

Visit raconte.ai

Raconte showing the transcript of a voice interview, with the recording, the AI summary and the sentiment over time

Micdrop also handles the rest of the call

Just call Micdrop.start() and your app holds a conversation.

Start building

Micdrop is free and open source under the MIT license. Bring your own API keys, or run the models on your own machine.