🎤Micdrop

OpenAI

OpenAI implementation for @micdrop/server.

This package provides AI agent, speech-to-text and text-to-speech implementations using OpenAI’s API.

Each of the three is a separate argument, so any of them can move to another provider. OpenaiRealtime runs all three in one connection to the Realtime API, which changes what you can swap and what you pay.

Installation

Terminal window
npm install @micdrop/openai

OpenAI Agent

Usage with MicdropServer

import { OpenaiAgent } from '@micdrop/openai'
import { MicdropServer } from '@micdrop/server'
const agent = new OpenaiAgent({
apiKey: process.env.OPENAI_API_KEY || '',
model: 'gpt-4o', // Default model
systemPrompt: 'You are a helpful assistant',
// Custom OpenAI Responses API settings (optional)
settings: {
temperature: 0.7,
max_output_tokens: 150,
},
})
// Use with MicdropServer
new MicdropServer(socket, {
agent,
// ... other options
})

Usage without MicdropServer

import { OpenaiAgent } from '@micdrop/openai'
const agent = new OpenaiAgent({
apiKey: process.env.OPENAI_API_KEY || '',
systemPrompt: 'You are a helpful assistant',
})
agent.on('Message', (message) => console.log('Message:', message))
agent.addUserMessage('Hello, what can you do?')
// The answer is a text stream, written as the model generates it
agent.answer().on('data', (chunk) => process.stdout.write(chunk))

Events

EventPayloadDescription
MessageMicdropConversationItemA message, a tool call or a tool result was added to the conversation.
ToolCallMicdropToolCallA tool declared with emitOutput ran, with its parameters and output.
CancelLastUserMessagenoneThe last user message was dropped because it carried no intent.
SkipAnswernoneThe agent stays silent and waits for the user to finish their sentence.
EndCallnoneThe agent decided that the call is over.
FailednoneThe agent gave up generating an answer after its retries.

See the Agent interface for the full contract.

Options

OptionTypeDefaultDescription
apiKeystringRequired*Your OpenAI API key (required if openai not provided)
openaiOpenAIOptionalOpenAI instance (alternative to apiKey)
modelstring'gpt-4o'OpenAI model to use
systemPromptstringRequiredSystem prompt for the agent
retryDelaynumber1000Delay in milliseconds between retry attempts
maxRetrynumber3Maximum number of retries on API failures
maxStepsnumber5Maximum number of steps (for tool calls)
autoEndCallboolean | stringfalseAuto-detect when user wants to end call
autoSemanticTurnboolean | stringfalseHandle incomplete user sentences
autoIgnoreUserNoiseboolean | stringfalseFilter meaningless user sounds
extractExtractJsonOptions | ExtractTagOptionsundefinedExtract structured data from responses
onBeforeAnswerfunctionundefinedHook called before answer generation, returns true to skip it or a text to answer instead
settingsobject{}Additional OpenAI Responses API parameters

The OpenAI Agent supports adding and removing custom tools to extend its capabilities. For detailed information about tool management, see the Tools documentation.

Advanced Features

The OpenAI Agent supports advanced features for improved conversation handling:

Langfuse Integration

You can integrate Langfuse for observability by using the openai option with a Langfuse-wrapped OpenAI client:

import { OpenaiAgent } from '@micdrop/openai'
import { Langfuse, observeOpenAI } from 'langfuse'
import OpenAI from 'openai'
// Initialize Langfuse
const langfuse = new Langfuse({
secretKey: process.env.LANGFUSE_SECRET_KEY,
publicKey: process.env.LANGFUSE_PUBLIC_KEY,
baseUrl: process.env.LANGFUSE_BASE_URL, // Optional, defaults to https://cloud.langfuse.com
})
// Get system prompt from Langfuse
const systemPrompt = await langfuse.getPrompt('voice-assistant-system-prompt')
// Create OpenAI client and wrap with Langfuse observability
const openai = observeOpenAI(
new OpenAI({ apiKey: process.env.OPENAI_API_KEY }),
{
sessionId: 'session-123',
userId: 'user-456',
}
)
// Create agent with Langfuse-wrapped OpenAI client
const agent = new OpenaiAgent({
openai,
model: 'gpt-4o',
systemPrompt: systemPrompt.prompt,
})

This integration will automatically track all OpenAI API calls, token usage, and conversation flows in your Langfuse dashboard with session and user context.

OpenAI STT (Speech-to-Text)

Real-time speech-to-text implementation using OpenAI’s WebSocket-based real-time transcription API.

Leave out the agent and the voice, and OpenaiSTT alone gives you speech to text with nothing answering back.

OpenAI deprecated gpt-4o-transcribe (the default), gpt-4o-mini-transcribe and whisper-1, which shut down on February 26, 2027. Their replacements, gpt-live-transcribe and gpt-transcribe, work with the same options: language is sent as the list of expected languages these models take. gpt-live-transcribe is recommended for new integrations.

Usage with MicdropServer

import { OpenaiSTT } from '@micdrop/openai'
import { MicdropServer } from '@micdrop/server'
const stt = new OpenaiSTT({
apiKey: process.env.OPENAI_API_KEY || '',
model: 'gpt-live-transcribe', // Default 'gpt-4o-transcribe', shut down on February 26, 2027
language: 'en', // Optional: specify language for better accuracy
prompt: 'Transcribe the incoming audio in real time.', // Optional: custom prompt
transcriptionTimeout: 4000, // Optional: timeout in ms for transcription
})
// Use with MicdropServer
new MicdropServer(socket, {
stt,
// ... other options
})

Usage without MicdropServer

import { OpenaiSTT } from '@micdrop/openai'
import { createReadStream } from 'fs'
const stt = new OpenaiSTT({
apiKey: process.env.OPENAI_API_KEY || '',
language: 'en',
})
stt.on('Transcript', (text) => console.log('Transcript:', text))
stt.on('Failed', (chunks) => console.error('Failed:', chunks.length, 'chunks'))
// Audio is raw PCM, 16 bits, 16 kHz, mono
stt.transcribe(createReadStream('speech.pcm'))

Events

EventPayloadDescription
TranscriptstringTranscription of one utterance. The text is empty when nothing was recognized.
FailedBuffer[]Transcription gave up after its retries, with the audio chunks left pending.

See the STT interface for the full contract.

Options

OptionTypeDefaultDescription
apiKeystringRequiredYour OpenAI API key
modelstring'gpt-4o-transcribe'Real-time transcription model: gpt-live-transcribe and gpt-transcribe, or until February 26, 2027 gpt-4o-transcribe and gpt-4o-mini-transcribe
languagestring'en'Language code for transcription
promptstring'Transcribe the incoming audio in real time.'Custom prompt to guide transcription behavior
connectionTimeoutnumber5000Timeout in milliseconds for WebSocket connection
transcriptionTimeoutnumber4000Timeout in milliseconds to wait for transcription result
retryDelaynumber1000Delay in milliseconds between reconnection attempts
maxRetrynumber3Maximum number of reconnection attempts before failing

OpenAI TTS (Text-to-Speech)

Two ways to synthesize, picked by model:

  • Speech endpoint (default): gpt-4o-mini-tts, its dated snapshots, tts-1 and tts-1-hd. OpenAI deprecated all of them, and the endpoint shuts down on January 6, 2027. Until then, they work as before.
  • Realtime API: gpt-realtime-2.1-mini, the replacement OpenAI names, recommended for new integrations. Each sentence is sent as an out-of-band response with instructions to read it word for word. The model is a voice model rather than a reader, so it can reword a sentence now and then (a contraction such as “I’m” for “I am”), and OpenaiTTS logs the transcript of any sentence it reads differently.

Either way, the incoming text is buffered into sentences and each sentence is synthesized as soon as it is complete, so playback can start without waiting for the whole answer.

Usage with MicdropServer

import { OpenaiTTS } from '@micdrop/openai'
import { MicdropServer } from '@micdrop/server'
const tts = new OpenaiTTS({
apiKey: process.env.OPENAI_API_KEY || '',
// Recommended: the speech endpoint models (default 'gpt-4o-mini-tts')
// shut down on January 6, 2027
model: 'gpt-realtime-2.1-mini',
voice: 'marin', // Default with gpt-realtime-*, 'alloy' otherwise
// Delivery control, not for tts-1 / tts-1-hd (optional)
instructions: 'Speak in a calm and friendly tone',
})
// Use with MicdropServer
new MicdropServer(socket, {
tts,
// ... other options
})

Usage without MicdropServer

import { OpenaiTTS } from '@micdrop/openai'
import { Readable } from 'stream'
const tts = new OpenaiTTS({
apiKey: process.env.OPENAI_API_KEY || '',
voice: 'alloy',
})
// Audio is raw PCM, 16 bits, 16 kHz, mono
tts.on('Audio', (chunk) => console.log('Audio:', chunk.length, 'bytes'))
tts.on('Failed', (texts) => console.error('Failed:', texts))
tts.speak(Readable.from(['Hello! ', 'What can I do for you?']))

Events

EventPayloadDescription
AudioBufferA chunk of audio, PCM 16 bits, 16 kHz, mono, ready to be played.
Failedstring[]Synthesis gave up after its retries, with the text that stayed unspoken.

See the TTS interface for the full contract.

Options

OptionTypeDefaultDescription
apiKeystringRequired*Your OpenAI API key (required if openai not provided)
openaiOpenAIOptionalOpenAI instance (alternative to apiKey)
modelstring'gpt-4o-mini-tts'gpt-realtime-2.1-mini (recommended), or until January 6, 2027 the speech endpoint models gpt-4o-mini-tts, tts-1, tts-1-hd
voicestring'marin' / 'alloy'marin with gpt-realtime-*, alloy otherwise. fable, nova and onyx only exist on the speech endpoint
instructionsstringundefinedDelivery control (accent, emotion, tone). Works with gpt-realtime-* and gpt-4o-mini-tts
speednumberundefinedSpeech speed, from 0.25 to 1.5 with gpt-realtime-* (clamped), from 0.25 to 4.0 with tts-1/tts-1-hd
reasoningEffortstring'minimal'Reasoning before speaking, gpt-realtime-* only: minimal, low, medium, high, xhigh
connectionTimeoutnumber5000Timeout in milliseconds to open the Realtime connection, gpt-realtime-* only
ℹ️ Note

Voices: the Realtime API offers alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin and cedar. fable, nova and onyx only exist on the speech endpoint.

Language: neither the speech endpoint nor the Realtime API has a language parameter, the voice follows the language of the input text. To influence the accent, use instructions (e.g. 'Speak with a Parisian accent') with gpt-realtime-* or gpt-4o-mini-tts.

OpenAI Realtime

A realtime model running on the OpenAI Realtime API. It hears the user and answers with its own voice, in place of a speech to text, an agent and a text to speech.

GPT-Live-1, OpenAI’s full-duplex voice model, needs its own endpoint and events, which OpenaiRealtime does not implement. GPT-Live-1 differs from GPT-Realtime-2.1 and Gemini 3.8 Live in price, turn-taking and tools.

Usage with MicdropServer

import { OpenaiRealtime } from '@micdrop/openai'
import { MicdropServer } from '@micdrop/server'
const realtime = new OpenaiRealtime({
apiKey: process.env.OPENAI_API_KEY || '',
model: 'gpt-realtime-2.1', // Default model
voice: 'marin', // Default voice
systemPrompt: 'You are a helpful assistant',
language: 'en', // Optional, helps the transcription of the user
// Advanced features (optional)
autoEndCall: true, // Automatically end call when user requests
})
new MicdropServer(socket, {
realtime,
generateFirstMessage: true,
})

The Micdrop client detects when the user speaks, so the turn detection of the Realtime API is turned off. The audio of each turn is committed when the client closes it, and the answer is asked for right after. Tools are added with addTool(), as with OpenaiAgent.

When the user interrupts, the answer is cancelled and the conversation of the model is cut where the user most likely stopped hearing it.

Events

OpenaiRealtime emits the events of an agent, and two of its own.

EventPayloadDescription
AudioBufferA chunk of the voice of the model, PCM 16 bits, 16 kHz, mono.
PartialMessagestringThe transcript of the answer so far.
MessageMicdropConversationItemA message, a tool call or a tool result was added to the conversation.
ToolCallMicdropToolCallA tool declared with emitOutput ran, with its parameters and output.
CancelLastUserMessagenoneThe last user message was dropped because it carried no intent.
SkipAnswernoneThe model gave no spoken answer.
EndCallnoneThe model decided that the call is over.
FailednoneThe connection could not be restored after its retries.

Options

OptionTypeDefaultDescription
apiKeystringRequiredYour OpenAI API key
systemPromptstringRequiredInstructions of the model
modelstring'gpt-realtime-2.1'Realtime model to use (gpt-realtime-2.1, gpt-realtime-mini, …)
voicestring'marin'Voice of the model
languagestringundefinedLanguage of the user, as an ISO 639-1 code, for the transcription
transcriptionModelstring'gpt-4o-transcribe'Model transcribing the user: gpt-live-transcribe or gpt-transcribe recommended, gpt-4o-transcribe shuts down on February 26, 2027
autoEndCallboolean | stringfalseEnds the call when the user asks for it
autoSemanticTurnboolean | stringfalseLets the model wait silently when the user did not finish their sentence
autoIgnoreUserNoiseboolean | stringfalseDrops the turn when it carries no meaning
connectionTimeoutnumber5000Time to open the connection, in milliseconds
retryDelaynumber1000Delay before reconnecting, in milliseconds
maxRetrynumber3Reconnection attempts before giving up

A Realtime session lasts sixty minutes at most. When the connection closes, OpenaiRealtime reconnects and sends the conversation written so far to the new session.