Micdrop Protocol
WebSocket Protocol
Micdrop uses a simple custom protocol over WebSocket for real-time communication between the client and server.
sequenceDiagram
participant W as Client
participant B as Server
Note left of W: Audio setup
W -->> B: Create call
B -->> W: Message (First assistant message)
Note right of B: Start TTS : Generate voice for first message
B -->> W: Audio chunk (First assistant message)
Note left of W: Play assistant speech
B -->> W: Audio chunk (First assistant message)
Note right of B: Stop TTS - Sent audio for first message
loop
Note left of W: Wait until user speaks
W -->> B: StartSpeaking
Note right of B: Start STT : Transcribe user speech
W -->> B: Audio chunk (User speech)
W -->> B: Audio chunk (User speech)
Note left of W: Silence
W -->> B: StopSpeaking - User stops speaking
Note right of B: Stop STT : User speech transcribed
B -->> W: Message (Transcribed user message)
Note right of B: Start Agent : Generate answer
Note right of B: Start TTS : Generate voice for answer
B ->> W: Audio chunk (Assistant answer)
B ->> W: Audio chunk (Assistant answer)
B -->> W: Message (Assistant answer)
Note right of B: Stop Agent : Finished answering
B ->> W: Audio chunk (Assistant answer)
B ->> W: Audio chunk (Assistant answer)
Note right of B: Stop TTS : Sent audio for answer
Note left of W: Play assistant speech
end
Why WebSocket?
While WebRTC is a powerful protocol for real-time communication, Micdrop uses a simple custom protocol over WebSocket for several reasons:
-
๐ฏ Focused on our use case: WebRTC is designed for peer-to-peer communication, with features we donโt need. Our client-server architecture is simpler.
-
๐ Efficient audio transmission: By using Voice Activity Detection (VAD) on the client side, we only send audio when the user is actually speaking. This reduces bandwidth usage and processing load compared to continuous streaming.
-
๐ก Simple implementation: WebSocket provides a straightforward, reliable way to send both text and binary data. The protocol is easy to implement and debug on both client and server.
-
๐ Bidirectional communication: WebSocket allows for real-time bidirectional messaging, which is perfect for our text and audio exchange needs.
-
๐ ๏ธ Custom protocol control: Our simple protocol lets us optimize exactly how and when audio/text messages are sent, without the overhead of WebRTCโs full feature set.
This approach gives us the real-time capabilities we need while keeping the implementation lean and efficient.
Calls without a voice
A call configured without a text to speech carries no audio chunk back. The
answer arrives as a Message, followed by SkipAnswer, which is what tells the
client to stop waiting and listen again. A call configured without an agent
sends SkipAnswer right after the transcribed user message.
See Dictation and text-only calls.
Client Commands
The client can send the following commands to the server:
MicdropClientCommands.StartSpeaking- The user starts speakingMicdropClientCommands.StopSpeaking- The user stops speakingMicdropClientCommands.Mute- The user mutes the microphone
Server Commands
The server can send the following commands to the client:
MicdropServerCommands.Message- A settled message, from the user or the assistant.MicdropServerCommands.PartialAssistantMessage- The answer being written, before it is finished. Sent whenpartialMessagesis on.MicdropServerCommands.CancelLastUserMessage- Cancel the last user message.MicdropServerCommands.SkipAnswer- Notify that the last generated answer was ignored, itโs listening again.MicdropServerCommands.EndCall- End the call.MicdropServerCommands.ToolCall- A tool the agent called, with its result.MicdropServerCommands.Classification- What the classifier made of the user turn, as JSON. Sent when the server has a classifier andclassifierOptions.sendToClientis on.