Voice Activity Detection (VAD)
Micdrop uses a VAD (Voice Activity Detection) to detect speech and silence and send chunks of audio to the server only when speech is detected.
Each VAD setting trades latency against false positives. Silero VAD, WebRTC VAD and a volume threshold misfire on background noise at very different rates, so the detector you pick limits how far you can tune for speed.
Whichever VAD decides, Micdrop keeps the last moments of audio in reserve, so the syllable spoken before it reacted is sent along with the rest of the sentence.
When the user can interrupt the assistant, everything they say over it is recorded, and echo cancellation keeps the assistant voice out of it. When interruptions are disabled, what the microphone hears while the assistant speaks, and for 300 ms after, is never sent to the server. Micdrop disables interruptions by itself when the browser reports that echo cancellation is off.
Supported VAD Types
Micdrop supports the following VADs by name:
'volume': Volume-based VAD (default)'silero': AI-based VAD using Silero
You can also pass instances of these VADs, or combine them in an array. See below for details.
Note: Only
'volume'and'silero'are supported as string names. Custom VADs must be passed as instances.
Quick Start
Configure VAD when starting a call:
import { Micdrop } from '@micdrop/web'
// Use volume-based detection (default)await Micdrop.start({ url: 'ws://localhost:8081', vad: 'volume',})
// Use AI-based detection for better accuracyawait Micdrop.start({ url: 'ws://localhost:8081', vad: 'silero',})
// Combine multiple VADs for best resultsawait Micdrop.start({ url: 'ws://localhost:8081', vad: ['volume', 'silero'],})Or when starting the microphone (before starting the call):
Micdrop.startMic({ vad: 'volume' })Volume VAD: Speech detection based on volume
By default, MicdropClient uses VolumeVAD for speech detection. You can use it explicitly when starting Micdrop:
Micdrop.start({ vad: 'volume' })or when starting the microphone (before starting the call):
Micdrop.startMic({ vad: 'volume' })It is inspired by hark and triggers speech detection events based on volume changes.
The level VolumeVAD follows is the loudest frequency rather than the overall energy: a voice concentrates its power in a few bands while a fan spreads it over all of them, which separates the two by about 25 dB instead of the 7 dB a plain average would give.
You can also pass an instance of VolumeVAD to MicdropClient:
const vad = new VolumeVAD({ history: 5, // Number of frames to consider for volume calculation threshold: -55, // Threshold in decibels for speech detection})Micdrop.start({ vad })- Default options:
{ history: 5, threshold: -55 } - Persistence: Options are saved to
localStoragein a browser, and whereversetMicdropStorage()points on a phone.
When to use Volume VAD:
- ✅ Low latency requirements
- ✅ Quiet environments
- ✅ Clear speech patterns
- ❌ Noisy environments
- ❌ Soft-spoken users
Silero VAD: Human speech detection with AI
To use SileroVAD for speech detection:
Micdrop.start({ vad: 'silero' })It runs the Silero model on ONNX Runtime Web in a browser and on the native runtime on a phone. The model gives each window of sound a probability of speech. Micdrop decides from those probabilities when speech starts and stops, with the same algorithm as ricky0123/vad, the open source project that runs Silero in the browser. The model is about two megabytes. Micdrop fetches it once, on the first call, from the @ricky0123/vad-web npm package on jsDelivr.
SileroVAD is more accurate than VolumeVAD and works better with low voice, since it hears the difference between a voice and a noise rather than measuring how loud the room is.
The ONNX runtime weighs more than the rest of Micdrop put together, so it only enters your bundle if you ask for it:
import '@micdrop/web/silero'To serve the model from your own domain rather than from jsDelivr:
import { setSileroOptions } from '@micdrop/web/silero'
setSileroOptions({ model: '/assets/silero_vad_v5.onnx' })On a phone the model runs on the native runtime, which asks for one more package and a patch. See Voice activity detection on React Native.
You can also pass an instance of SileroVAD to MicdropClient:
const vad = new SileroVAD({ positiveSpeechThreshold: 0.18, // Threshold for positive speech detection negativeSpeechThreshold: 0.11, // Threshold for negative speech detection minSpeechFrames: 4, // Windows of speech before a turn is confirmed redemptionFrames: 8, // Windows of silence before a turn is closed})Micdrop.start({ vad })- Default options:
{ positiveSpeechThreshold: 0.18, negativeSpeechThreshold: 0.11, minSpeechFrames: 4, redemptionFrames: 8 } - Persistence: Options are saved to
localStoragein a browser, and whereversetMicdropStorage()points on a phone.
redemptionFrames is eight windows, about a quarter of a second, so a turn closes quickly and leans on turn detection to keep the floor when a speaker only paused. Running Silero without a turn detector, raise it to twenty, or expect the agent to answer during hesitations. See Reducing Latency for what each window costs and what it buys.
When to use Silero VAD:
- ✅ Noisy environments
- ✅ Soft-spoken users
- ✅ Multiple speakers
- ✅ Background music/TV
- ❌ Extremely low latency needs (adds ~50ms processing)
Multiple VAD: Combine multiple VADs
Combining multiple VADs is useful to get more accurate speech detection:
- Volume to ignore low voice
- Silero to detect human speech
A turn opens once every VAD has heard the speech and at least one of them is sure of it. A short word like “No” is over before the volume VAD confirms it, and still gets through as long as it was loud enough to be noticed.
You can combine multiple VADs by passing an array of VAD names:
Micdrop.start({ vad: ['volume', 'silero'] })Or with instances:
const vad = [new VolumeVAD(), new SileroVAD()]Micdrop.start({ vad })Or mix names and instances:
await Micdrop.start({ vad: ['volume', new SileroVAD({ positiveSpeechThreshold: 0.15 })],})How it works:
StartSpeakingis emitted when any VAD detects possible speech.ConfirmSpeakingis emitted only when all VADs confirm speech.StopSpeakingis emitted when all VADs detect silence.CancelSpeakingis emitted if all VADs agree speech was a false positive.
This approach reduces false positives while maintaining quick response times.
VAD Events
VADs emit the following events:
StartSpeaking: Possible speech detected (not yet confirmed)ConfirmSpeaking: Speech confirmedCancelSpeaking: Speech start was a false positive (noise, etc.)StopSpeaking: Speech endedChangeStatus: Status changed (Silence,MaybeSpeaking,Speaking)
Monitor VAD activity in your application:
Micdrop.vad.on('StartSpeaking', () => { console.log('🎤 Possible speech detected...') showListeningIndicator()})
Micdrop.vad.on('ConfirmSpeaking', () => { console.log('✅ Speech confirmed - recording') highlightMicrophoneButton()})
Micdrop.vad.on('StopSpeaking', () => { console.log('🔇 Speech ended') resetMicrophoneButton()})
Micdrop.vad.on('CancelSpeaking', () => { console.log('❌ False positive - not speech') hideListeningIndicator()})
Micdrop.vad.on('ChangeStatus', (status) => { console.log('VAD status:', status) // 'Silence', 'MaybeSpeaking', 'Speaking'})Custom VAD
The VAD class is exported, so another method can be plugged in. It listens to the microphone and reports four moments: speech may have started, it is confirmed, it was only noise, it has ended.
import { MicSource, VAD } from '@micdrop/web'
class MyVAD extends VAD { private mic?: MicSource
get isStarted() { return !!this.mic }
get isPaused() { return false }
async start(mic: MicSource) { this.mic = mic mic.on('Frames', this.onFrames) }
async stop() { this.mic?.off('Frames', this.onFrames) this.mic = undefined }
async pause() {} async resume() {}
private onFrames = (frames: Float32Array, sampleRate: number) => { // Decide, then report // this.emit('StartSpeaking') // this.emit('ConfirmSpeaking') // this.emit('CancelSpeaking') // this.emit('StopSpeaking') }}
Micdrop.start({ vad: new MyVAD() })Frames carries mono samples between -1 and 1, Volume carries the level in dBFS. The class lives in @micdrop/client, so the same detector runs in a browser and on a phone, where it is imported from @micdrop/react-native.
See VolumeVAD as a complete example.
VAD Delay
Every VAD declares in delay the longest time it may take to notice that someone started speaking, in milliseconds. Micdrop keeps that much audio in reserve and sends it with the turn, so the first word arrives whole. Silero asks for about 320 ms, the volume VAD for 500 ms, and a combination of VADs for the largest of its members. A custom VAD that needs several samples before it reacts should raise it (the default is 100 ms).
Tuning VAD Performance
Volume VAD Tuning
Adjust sensitivity based on environment:
// Quiet environment - more sensitiveconst quietVad = new VolumeVAD({ threshold: -65, // Lower threshold for quiet voices history: 3, // Faster response})
// Noisy environment - less sensitiveconst noisyVad = new VolumeVAD({ threshold: -45, // Higher threshold to ignore noise history: 8, // More frames for stability})Silero VAD Tuning
Fine-tune AI detection:
// More sensitive - catches quiet speechconst sensitiveVad = new SileroVAD({ positiveSpeechThreshold: 0.15, // Lower threshold minSpeechFrames: 6, // Faster confirmation})
// More conservative - reduces false positivesconst conservativeVad = new SileroVAD({ positiveSpeechThreshold: 0.22, // Higher threshold minSpeechFrames: 8, // More confirmation needed redemptionFrames: 20, // Longer silence confirmation})Dynamic VAD Configuration
You can update VAD settings in real-time without restarting:
// Update Volume VAD settingsconst volumeVad = Micdrop.vad as VolumeVADvolumeVad.setOptions({ threshold: -45 })
// Update Silero VAD settingsconst sileroVad = Micdrop.vad as SileroVADsileroVad.setOptions({ positiveSpeechThreshold: 0.15 })
// Reset to default optionsvolumeVad.resetOptions()sileroVad.resetOptions()Persistent Settings
Both VolumeVAD and SileroVAD settings are automatically saved to localStorage and restored when loading with their names ('volume' or 'silero') and not instances.
React VAD Settings UI
See a complete React component for VAD configuration based on the demo client: VADSettings