🎤Micdrop

Voice Activity Detection (VAD)

Micdrop uses a VAD (Voice Activity Detection) to detect speech and silence and send chunks of audio to the server only when speech is detected.

Each VAD setting trades latency against false positives. Silero VAD, WebRTC VAD and a volume threshold misfire on background noise at very different rates, so the detector you pick limits how far you can tune for speed.

Whichever VAD decides, Micdrop keeps the last moments of audio in reserve, so the syllable spoken before it reacted is sent along with the rest of the sentence.

When the user can interrupt the assistant, everything they say over it is recorded, and echo cancellation keeps the assistant voice out of it. When interruptions are disabled, what the microphone hears while the assistant speaks, and for 300 ms after, is never sent to the server. Micdrop disables interruptions by itself when the browser reports that echo cancellation is off.

Supported VAD Types

Micdrop supports the following VADs by name:

  • 'volume': Volume-based VAD (default)
  • 'silero': AI-based VAD using Silero

You can also pass instances of these VADs, or combine them in an array. See below for details.

Note: Only 'volume' and 'silero' are supported as string names. Custom VADs must be passed as instances.

Quick Start

Configure VAD when starting a call:

import { Micdrop } from '@micdrop/web'
// Use volume-based detection (default)
await Micdrop.start({
url: 'ws://localhost:8081',
vad: 'volume',
})
// Use AI-based detection for better accuracy
await Micdrop.start({
url: 'ws://localhost:8081',
vad: 'silero',
})
// Combine multiple VADs for best results
await Micdrop.start({
url: 'ws://localhost:8081',
vad: ['volume', 'silero'],
})

Or when starting the microphone (before starting the call):

Micdrop.startMic({ vad: 'volume' })

Volume VAD: Speech detection based on volume

By default, MicdropClient uses VolumeVAD for speech detection. You can use it explicitly when starting Micdrop:

Micdrop.start({ vad: 'volume' })

or when starting the microphone (before starting the call):

Micdrop.startMic({ vad: 'volume' })

It is inspired by hark and triggers speech detection events based on volume changes.

The level VolumeVAD follows is the loudest frequency rather than the overall energy: a voice concentrates its power in a few bands while a fan spreads it over all of them, which separates the two by about 25 dB instead of the 7 dB a plain average would give.

You can also pass an instance of VolumeVAD to MicdropClient:

const vad = new VolumeVAD({
history: 5, // Number of frames to consider for volume calculation
threshold: -55, // Threshold in decibels for speech detection
})
Micdrop.start({ vad })
  • Default options: { history: 5, threshold: -55 }
  • Persistence: Options are saved to localStorage in a browser, and wherever setMicdropStorage() points on a phone.

When to use Volume VAD:

  • ✅ Low latency requirements
  • ✅ Quiet environments
  • ✅ Clear speech patterns
  • ❌ Noisy environments
  • ❌ Soft-spoken users

Silero VAD: Human speech detection with AI

To use SileroVAD for speech detection:

Micdrop.start({ vad: 'silero' })

It runs the Silero model on ONNX Runtime Web in a browser and on the native runtime on a phone. The model gives each window of sound a probability of speech. Micdrop decides from those probabilities when speech starts and stops, with the same algorithm as ricky0123/vad, the open source project that runs Silero in the browser. The model is about two megabytes. Micdrop fetches it once, on the first call, from the @ricky0123/vad-web npm package on jsDelivr.

SileroVAD is more accurate than VolumeVAD and works better with low voice, since it hears the difference between a voice and a noise rather than measuring how loud the room is.

The ONNX runtime weighs more than the rest of Micdrop put together, so it only enters your bundle if you ask for it:

import '@micdrop/web/silero'

To serve the model from your own domain rather than from jsDelivr:

import { setSileroOptions } from '@micdrop/web/silero'
setSileroOptions({ model: '/assets/silero_vad_v5.onnx' })

On a phone the model runs on the native runtime, which asks for one more package and a patch. See Voice activity detection on React Native.

You can also pass an instance of SileroVAD to MicdropClient:

const vad = new SileroVAD({
positiveSpeechThreshold: 0.18, // Threshold for positive speech detection
negativeSpeechThreshold: 0.11, // Threshold for negative speech detection
minSpeechFrames: 4, // Windows of speech before a turn is confirmed
redemptionFrames: 8, // Windows of silence before a turn is closed
})
Micdrop.start({ vad })
  • Default options: { positiveSpeechThreshold: 0.18, negativeSpeechThreshold: 0.11, minSpeechFrames: 4, redemptionFrames: 8 }
  • Persistence: Options are saved to localStorage in a browser, and wherever setMicdropStorage() points on a phone.

redemptionFrames is eight windows, about a quarter of a second, so a turn closes quickly and leans on turn detection to keep the floor when a speaker only paused. Running Silero without a turn detector, raise it to twenty, or expect the agent to answer during hesitations. See Reducing Latency for what each window costs and what it buys.

When to use Silero VAD:

  • ✅ Noisy environments
  • ✅ Soft-spoken users
  • ✅ Multiple speakers
  • ✅ Background music/TV
  • ❌ Extremely low latency needs (adds ~50ms processing)

Multiple VAD: Combine multiple VADs

Combining multiple VADs is useful to get more accurate speech detection:

  • Volume to ignore low voice
  • Silero to detect human speech

A turn opens once every VAD has heard the speech and at least one of them is sure of it. A short word like “No” is over before the volume VAD confirms it, and still gets through as long as it was loud enough to be noticed.

You can combine multiple VADs by passing an array of VAD names:

Micdrop.start({ vad: ['volume', 'silero'] })

Or with instances:

const vad = [new VolumeVAD(), new SileroVAD()]
Micdrop.start({ vad })

Or mix names and instances:

await Micdrop.start({
vad: ['volume', new SileroVAD({ positiveSpeechThreshold: 0.15 })],
})

How it works:

  • StartSpeaking is emitted when any VAD detects possible speech.
  • ConfirmSpeaking is emitted only when all VADs confirm speech.
  • StopSpeaking is emitted when all VADs detect silence.
  • CancelSpeaking is emitted if all VADs agree speech was a false positive.

This approach reduces false positives while maintaining quick response times.

VAD Events

VADs emit the following events:

  • StartSpeaking: Possible speech detected (not yet confirmed)
  • ConfirmSpeaking: Speech confirmed
  • CancelSpeaking: Speech start was a false positive (noise, etc.)
  • StopSpeaking: Speech ended
  • ChangeStatus: Status changed (Silence, MaybeSpeaking, Speaking)

Monitor VAD activity in your application:

Micdrop.vad.on('StartSpeaking', () => {
console.log('🎤 Possible speech detected...')
showListeningIndicator()
})
Micdrop.vad.on('ConfirmSpeaking', () => {
console.log('✅ Speech confirmed - recording')
highlightMicrophoneButton()
})
Micdrop.vad.on('StopSpeaking', () => {
console.log('🔇 Speech ended')
resetMicrophoneButton()
})
Micdrop.vad.on('CancelSpeaking', () => {
console.log('❌ False positive - not speech')
hideListeningIndicator()
})
Micdrop.vad.on('ChangeStatus', (status) => {
console.log('VAD status:', status) // 'Silence', 'MaybeSpeaking', 'Speaking'
})

Custom VAD

The VAD class is exported, so another method can be plugged in. It listens to the microphone and reports four moments: speech may have started, it is confirmed, it was only noise, it has ended.

import { MicSource, VAD } from '@micdrop/web'
class MyVAD extends VAD {
private mic?: MicSource
get isStarted() {
return !!this.mic
}
get isPaused() {
return false
}
async start(mic: MicSource) {
this.mic = mic
mic.on('Frames', this.onFrames)
}
async stop() {
this.mic?.off('Frames', this.onFrames)
this.mic = undefined
}
async pause() {}
async resume() {}
private onFrames = (frames: Float32Array, sampleRate: number) => {
// Decide, then report
// this.emit('StartSpeaking')
// this.emit('ConfirmSpeaking')
// this.emit('CancelSpeaking')
// this.emit('StopSpeaking')
}
}
Micdrop.start({ vad: new MyVAD() })

Frames carries mono samples between -1 and 1, Volume carries the level in dBFS. The class lives in @micdrop/client, so the same detector runs in a browser and on a phone, where it is imported from @micdrop/react-native.

See VolumeVAD as a complete example.

VAD Delay

Every VAD declares in delay the longest time it may take to notice that someone started speaking, in milliseconds. Micdrop keeps that much audio in reserve and sends it with the turn, so the first word arrives whole. Silero asks for about 320 ms, the volume VAD for 500 ms, and a combination of VADs for the largest of its members. A custom VAD that needs several samples before it reacts should raise it (the default is 100 ms).

Tuning VAD Performance

Volume VAD Tuning

Adjust sensitivity based on environment:

// Quiet environment - more sensitive
const quietVad = new VolumeVAD({
threshold: -65, // Lower threshold for quiet voices
history: 3, // Faster response
})
// Noisy environment - less sensitive
const noisyVad = new VolumeVAD({
threshold: -45, // Higher threshold to ignore noise
history: 8, // More frames for stability
})

Silero VAD Tuning

Fine-tune AI detection:

// More sensitive - catches quiet speech
const sensitiveVad = new SileroVAD({
positiveSpeechThreshold: 0.15, // Lower threshold
minSpeechFrames: 6, // Faster confirmation
})
// More conservative - reduces false positives
const conservativeVad = new SileroVAD({
positiveSpeechThreshold: 0.22, // Higher threshold
minSpeechFrames: 8, // More confirmation needed
redemptionFrames: 20, // Longer silence confirmation
})

Dynamic VAD Configuration

You can update VAD settings in real-time without restarting:

// Update Volume VAD settings
const volumeVad = Micdrop.vad as VolumeVAD
volumeVad.setOptions({ threshold: -45 })
// Update Silero VAD settings
const sileroVad = Micdrop.vad as SileroVAD
sileroVad.setOptions({ positiveSpeechThreshold: 0.15 })
// Reset to default options
volumeVad.resetOptions()
sileroVad.resetOptions()

Persistent Settings

Both VolumeVAD and SileroVAD settings are automatically saved to localStorage and restored when loading with their names ('volume' or 'silero') and not instances.

React VAD Settings UI

See a complete React component for VAD configuration based on the demo client: VADSettings