October 3, 2026
Smart Turn and End of Turn Detection for Voice Agents
A voice agent ends your turn on silence, on the sound of your sentence or on its words. Compare the three, then run Smart Turn in TypeScript.
Building real-time voice AI on the web: architecture, providers, latency and data sovereignty.
October 3, 2026
A voice agent ends your turn on silence, on the sound of your sentence or on its words. Compare the three, then run Smart Turn in TypeScript.
October 2, 2026
We turn the ElevenLabs API prices of October 2026 into a cost per minute of agent speech and count how many calls you can run at once on each plan.
October 1, 2026
We compare Cartesia Sonic 3.6 with ElevenLabs Flash and v4 Turbo on latency, voices, price and concurrency, then run both in one TypeScript voice agent.
September 30, 2026
Only Eleven v3 and v4 perform ElevenLabs audio tags. In a voice agent, prompt the LLM to write them, show them as labels and strip them for a backup voice.
September 29, 2026
Eleven v4 Turbo streams only over the Text to Dialogue WebSocket. With a flush per sentence, it started speaking in 138 ms in our launch-day test.
September 28, 2026
OpenAI, Gemini, Voxtral and Gladia compared for live transcription in a voice agent: price per minute, languages, streaming accuracy and time to transcript.
September 27, 2026
Keywords, an LLM or a dedicated model can classify each turn of a voice call. Compare their latency and route the answer in TypeScript before the LLM speaks.
September 26, 2026
Jev, the System One model of TypeSafe AI, answers typed questions about a text with calibrated probabilities in 70 to 500 ms. It writes no text.
September 25, 2026
Classify each turn of a voice call with Jev in a few hundred ms, then pick who answers before the LLM starts: a scripted line, a human, plain code or the LLM.
September 24, 2026
Eleven voice AI SDKs for web and React Native apps, ranked on TypeScript coverage, React Native support and provider choice. Micdrop, the publisher, is second.
September 22, 2026
The Realtime SDK needs a WebRTC transport, an audio session and an ephemeral key on the phone. Hold the session on your server and the app only streams PCM.
September 18, 2026
GPT-Live-1 listens while it speaks, GPT-Realtime-2.1 reasons and calls tools in one session, and Gemini 3.8 Live costs about a quarter of GPT-Realtime-2.1.
September 15, 2026
Whisper transcribes in Node with no API key but pads every sentence to thirty seconds. Vosk, Moonshine and Parakeet read audio at its real length.
September 13, 2026
This ranking compares ten AI voice agent platforms on price per minute, telephony, provider choice and EU hosting. Micdrop publishes it and sits fourth.
September 12, 2026
Cartesia, Gradium, OpenAI, Gemini and four local engines share one TypeScript interface. Switching engine takes one line. FallbackTTS keeps a backup ready.
September 9, 2026
The Web Speech API transcribes in Chrome and Safari and sends your audio to Google or Apple. Firefox reached it only in 2025, off by default.
September 2, 2026
Three engines speak from a Node server with no API key. Pocket TTS speaks first, Piper covers 43 languages, Kokoro installs in one command.
August 30, 2026
The system recogniser, a model on the device, or streaming to a server. Three ways to transcribe voice in a React Native app, and what each one costs.
August 24, 2026
Micdrop runs the voice loop in your Node server, so you pay your AI providers directly. Retell AI bills it per minute and extra concurrent calls monthly.
August 19, 2026
The Realtime API gives you speech to speech in one connection. A pipeline gives you provider choice, voices and cost control. Here is how to pick between them.
August 16, 2026
Vapi runs your voice agent on its servers for $0.05 a minute. Micdrop runs the same loop inside the Node server you already ship, with no platform fee.
August 13, 2026
Micdrop runs a voice call in TypeScript over a WebSocket to your own Node server. LiveKit Agents works in Node too, as a worker that joins a WebRTC room.
August 13, 2026
Thirteen voice agent frameworks and hosted platforms compared on language, transport, licence and cost. Micdrop publishes this ranking and sits third in it.
August 13, 2026
A browser voice agent needs voice activity detection to tell when the user speaks. Compare volume thresholds, WebRTC VAD and Silero, then tune your choice.
March 3, 2026
Micdrop runs the whole voice AI pipeline in TypeScript, browser and Node.js side, with provider fallback and semantic turn detection built in.
March 3, 2026
Intégrez Gradium, le TTS français issu de Kyutai, dans une application web avec Micdrop. Vos données passent par des serveurs européens, avec un fallback.