OpenAI
OpenAI implementation for @micdrop/server.
This package provides AI agent, speech-to-text and text-to-speech implementations using OpenAI’s API.
Each of the three is a separate argument, so any of them can move to another provider. OpenaiRealtime runs all three in one connection to the Realtime API, which changes what you can swap and what you pay.
Installation
npm install @micdrop/openaiOpenAI Agent
Usage with MicdropServer
import { OpenaiAgent } from '@micdrop/openai'import { MicdropServer } from '@micdrop/server'
const agent = new OpenaiAgent({ apiKey: process.env.OPENAI_API_KEY || '', model: 'gpt-4o', // Default model systemPrompt: 'You are a helpful assistant',
// Custom OpenAI Responses API settings (optional) settings: { temperature: 0.7, max_output_tokens: 150, },})
// Use with MicdropServernew MicdropServer(socket, { agent, // ... other options})Usage without MicdropServer
import { OpenaiAgent } from '@micdrop/openai'
const agent = new OpenaiAgent({ apiKey: process.env.OPENAI_API_KEY || '', systemPrompt: 'You are a helpful assistant',})
agent.on('Message', (message) => console.log('Message:', message))
agent.addUserMessage('Hello, what can you do?')
// The answer is a text stream, written as the model generates itagent.answer().on('data', (chunk) => process.stdout.write(chunk))Events
| Event | Payload | Description |
|---|---|---|
Message | MicdropConversationItem | A message, a tool call or a tool result was added to the conversation. |
ToolCall | MicdropToolCall | A tool declared with emitOutput ran, with its parameters and output. |
CancelLastUserMessage | none | The last user message was dropped because it carried no intent. |
SkipAnswer | none | The agent stays silent and waits for the user to finish their sentence. |
EndCall | none | The agent decided that the call is over. |
Failed | none | The agent gave up generating an answer after its retries. |
See the Agent interface for the full contract.
Options
| Option | Type | Default | Description |
|---|---|---|---|
apiKey | string | Required* | Your OpenAI API key (required if openai not provided) |
openai | OpenAI | Optional | OpenAI instance (alternative to apiKey) |
model | string | 'gpt-4o' | OpenAI model to use |
systemPrompt | string | Required | System prompt for the agent |
retryDelay | number | 1000 | Delay in milliseconds between retry attempts |
maxRetry | number | 3 | Maximum number of retries on API failures |
maxSteps | number | 5 | Maximum number of steps (for tool calls) |
autoEndCall | boolean | string | false | Auto-detect when user wants to end call |
autoSemanticTurn | boolean | string | false | Handle incomplete user sentences |
autoIgnoreUserNoise | boolean | string | false | Filter meaningless user sounds |
extract | ExtractJsonOptions | ExtractTagOptions | undefined | Extract structured data from responses |
onBeforeAnswer | function | undefined | Hook called before answer generation, returns true to skip it or a text to answer instead |
settings | object | {} | Additional OpenAI Responses API parameters |
The OpenAI Agent supports adding and removing custom tools to extend its capabilities. For detailed information about tool management, see the Tools documentation.
Advanced Features
The OpenAI Agent supports advanced features for improved conversation handling:
- Auto End Call: Automatically detect when users want to end the conversation
- Semantic Turn Detection: Handle incomplete sentences for natural flow
- User Noise Filtering: Filter out meaningless sounds and filler words
- Extract Value from Answer: Extract structured data from responses
- Tools: Add custom tools to the agent
Langfuse Integration
You can integrate Langfuse for observability by using the openai option with a Langfuse-wrapped OpenAI client:
import { OpenaiAgent } from '@micdrop/openai'import { Langfuse, observeOpenAI } from 'langfuse'import OpenAI from 'openai'
// Initialize Langfuseconst langfuse = new Langfuse({ secretKey: process.env.LANGFUSE_SECRET_KEY, publicKey: process.env.LANGFUSE_PUBLIC_KEY, baseUrl: process.env.LANGFUSE_BASE_URL, // Optional, defaults to https://cloud.langfuse.com})
// Get system prompt from Langfuseconst systemPrompt = await langfuse.getPrompt('voice-assistant-system-prompt')
// Create OpenAI client and wrap with Langfuse observabilityconst openai = observeOpenAI( new OpenAI({ apiKey: process.env.OPENAI_API_KEY }), { sessionId: 'session-123', userId: 'user-456', })
// Create agent with Langfuse-wrapped OpenAI clientconst agent = new OpenaiAgent({ openai, model: 'gpt-4o', systemPrompt: systemPrompt.prompt,})This integration will automatically track all OpenAI API calls, token usage, and conversation flows in your Langfuse dashboard with session and user context.
OpenAI STT (Speech-to-Text)
Real-time speech-to-text implementation using OpenAI’s WebSocket-based real-time transcription API.
Leave out the agent and the voice, and OpenaiSTT alone gives you speech to text with nothing answering back.
OpenAI deprecated gpt-4o-transcribe (the default), gpt-4o-mini-transcribe and whisper-1, which shut down on February 26, 2027. Their replacements, gpt-live-transcribe and gpt-transcribe, work with the same options: language is sent as the list of expected languages these models take. gpt-live-transcribe is recommended for new integrations.
Usage with MicdropServer
import { OpenaiSTT } from '@micdrop/openai'import { MicdropServer } from '@micdrop/server'
const stt = new OpenaiSTT({ apiKey: process.env.OPENAI_API_KEY || '', model: 'gpt-live-transcribe', // Default 'gpt-4o-transcribe', shut down on February 26, 2027 language: 'en', // Optional: specify language for better accuracy prompt: 'Transcribe the incoming audio in real time.', // Optional: custom prompt transcriptionTimeout: 4000, // Optional: timeout in ms for transcription})
// Use with MicdropServernew MicdropServer(socket, { stt, // ... other options})Usage without MicdropServer
import { OpenaiSTT } from '@micdrop/openai'import { createReadStream } from 'fs'
const stt = new OpenaiSTT({ apiKey: process.env.OPENAI_API_KEY || '', language: 'en',})
stt.on('Transcript', (text) => console.log('Transcript:', text))stt.on('Failed', (chunks) => console.error('Failed:', chunks.length, 'chunks'))
// Audio is raw PCM, 16 bits, 16 kHz, monostt.transcribe(createReadStream('speech.pcm'))Events
| Event | Payload | Description |
|---|---|---|
Transcript | string | Transcription of one utterance. The text is empty when nothing was recognized. |
Failed | Buffer[] | Transcription gave up after its retries, with the audio chunks left pending. |
See the STT interface for the full contract.
Options
| Option | Type | Default | Description |
|---|---|---|---|
apiKey | string | Required | Your OpenAI API key |
model | string | 'gpt-4o-transcribe' | Real-time transcription model: gpt-live-transcribe and gpt-transcribe, or until February 26, 2027 gpt-4o-transcribe and gpt-4o-mini-transcribe |
language | string | 'en' | Language code for transcription |
prompt | string | 'Transcribe the incoming audio in real time.' | Custom prompt to guide transcription behavior |
connectionTimeout | number | 5000 | Timeout in milliseconds for WebSocket connection |
transcriptionTimeout | number | 4000 | Timeout in milliseconds to wait for transcription result |
retryDelay | number | 1000 | Delay in milliseconds between reconnection attempts |
maxRetry | number | 3 | Maximum number of reconnection attempts before failing |
OpenAI TTS (Text-to-Speech)
Two ways to synthesize, picked by model:
- Speech endpoint (default):
gpt-4o-mini-tts, its dated snapshots,tts-1andtts-1-hd. OpenAI deprecated all of them, and the endpoint shuts down on January 6, 2027. Until then, they work as before. - Realtime API:
gpt-realtime-2.1-mini, the replacement OpenAI names, recommended for new integrations. Each sentence is sent as an out-of-band response with instructions to read it word for word. The model is a voice model rather than a reader, so it can reword a sentence now and then (a contraction such as “I’m” for “I am”), andOpenaiTTSlogs the transcript of any sentence it reads differently.
Either way, the incoming text is buffered into sentences and each sentence is synthesized as soon as it is complete, so playback can start without waiting for the whole answer.
Usage with MicdropServer
import { OpenaiTTS } from '@micdrop/openai'import { MicdropServer } from '@micdrop/server'
const tts = new OpenaiTTS({ apiKey: process.env.OPENAI_API_KEY || '', // Recommended: the speech endpoint models (default 'gpt-4o-mini-tts') // shut down on January 6, 2027 model: 'gpt-realtime-2.1-mini', voice: 'marin', // Default with gpt-realtime-*, 'alloy' otherwise
// Delivery control, not for tts-1 / tts-1-hd (optional) instructions: 'Speak in a calm and friendly tone',})
// Use with MicdropServernew MicdropServer(socket, { tts, // ... other options})Usage without MicdropServer
import { OpenaiTTS } from '@micdrop/openai'import { Readable } from 'stream'
const tts = new OpenaiTTS({ apiKey: process.env.OPENAI_API_KEY || '', voice: 'alloy',})
// Audio is raw PCM, 16 bits, 16 kHz, monotts.on('Audio', (chunk) => console.log('Audio:', chunk.length, 'bytes'))tts.on('Failed', (texts) => console.error('Failed:', texts))
tts.speak(Readable.from(['Hello! ', 'What can I do for you?']))Events
| Event | Payload | Description |
|---|---|---|
Audio | Buffer | A chunk of audio, PCM 16 bits, 16 kHz, mono, ready to be played. |
Failed | string[] | Synthesis gave up after its retries, with the text that stayed unspoken. |
See the TTS interface for the full contract.
Options
| Option | Type | Default | Description |
|---|---|---|---|
apiKey | string | Required* | Your OpenAI API key (required if openai not provided) |
openai | OpenAI | Optional | OpenAI instance (alternative to apiKey) |
model | string | 'gpt-4o-mini-tts' | gpt-realtime-2.1-mini (recommended), or until January 6, 2027 the speech endpoint models gpt-4o-mini-tts, tts-1, tts-1-hd |
voice | string | 'marin' / 'alloy' | marin with gpt-realtime-*, alloy otherwise. fable, nova and onyx only exist on the speech endpoint |
instructions | string | undefined | Delivery control (accent, emotion, tone). Works with gpt-realtime-* and gpt-4o-mini-tts |
speed | number | undefined | Speech speed, from 0.25 to 1.5 with gpt-realtime-* (clamped), from 0.25 to 4.0 with tts-1/tts-1-hd |
reasoningEffort | string | 'minimal' | Reasoning before speaking, gpt-realtime-* only: minimal, low, medium, high, xhigh |
connectionTimeout | number | 5000 | Timeout in milliseconds to open the Realtime connection, gpt-realtime-* only |
Voices: the Realtime API offers alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin and cedar. fable, nova and onyx only exist on the speech endpoint.
Language: neither the speech endpoint nor the Realtime API has a language parameter, the voice follows the language of the input text. To influence the accent, use instructions (e.g. 'Speak with a Parisian accent') with gpt-realtime-* or gpt-4o-mini-tts.
OpenAI Realtime
A realtime model running on the OpenAI Realtime API. It hears the user and answers with its own voice, in place of a speech to text, an agent and a text to speech.
GPT-Live-1, OpenAI’s full-duplex voice model, needs its own endpoint and events, which OpenaiRealtime does not implement. GPT-Live-1 differs from GPT-Realtime-2.1 and Gemini 3.8 Live in price, turn-taking and tools.
Usage with MicdropServer
import { OpenaiRealtime } from '@micdrop/openai'import { MicdropServer } from '@micdrop/server'
const realtime = new OpenaiRealtime({ apiKey: process.env.OPENAI_API_KEY || '', model: 'gpt-realtime-2.1', // Default model voice: 'marin', // Default voice systemPrompt: 'You are a helpful assistant', language: 'en', // Optional, helps the transcription of the user
// Advanced features (optional) autoEndCall: true, // Automatically end call when user requests})
new MicdropServer(socket, { realtime, generateFirstMessage: true,})The Micdrop client detects when the user speaks, so the turn detection of the Realtime API is turned off. The audio of each turn is committed when the client closes it, and the answer is asked for right after. Tools are added with addTool(), as with OpenaiAgent.
When the user interrupts, the answer is cancelled and the conversation of the model is cut where the user most likely stopped hearing it.
Events
OpenaiRealtime emits the events of an agent, and two of its own.
| Event | Payload | Description |
|---|---|---|
Audio | Buffer | A chunk of the voice of the model, PCM 16 bits, 16 kHz, mono. |
PartialMessage | string | The transcript of the answer so far. |
Message | MicdropConversationItem | A message, a tool call or a tool result was added to the conversation. |
ToolCall | MicdropToolCall | A tool declared with emitOutput ran, with its parameters and output. |
CancelLastUserMessage | none | The last user message was dropped because it carried no intent. |
SkipAnswer | none | The model gave no spoken answer. |
EndCall | none | The model decided that the call is over. |
Failed | none | The connection could not be restored after its retries. |
Options
| Option | Type | Default | Description |
|---|---|---|---|
apiKey | string | Required | Your OpenAI API key |
systemPrompt | string | Required | Instructions of the model |
model | string | 'gpt-realtime-2.1' | Realtime model to use (gpt-realtime-2.1, gpt-realtime-mini, …) |
voice | string | 'marin' | Voice of the model |
language | string | undefined | Language of the user, as an ISO 639-1 code, for the transcription |
transcriptionModel | string | 'gpt-4o-transcribe' | Model transcribing the user: gpt-live-transcribe or gpt-transcribe recommended, gpt-4o-transcribe shuts down on February 26, 2027 |
autoEndCall | boolean | string | false | Ends the call when the user asks for it |
autoSemanticTurn | boolean | string | false | Lets the model wait silently when the user did not finish their sentence |
autoIgnoreUserNoise | boolean | string | false | Drops the turn when it carries no meaning |
connectionTimeout | number | 5000 | Time to open the connection, in milliseconds |
retryDelay | number | 1000 | Delay before reconnecting, in milliseconds |
maxRetry | number | 3 | Reconnection attempts before giving up |
A Realtime session lasts sixty minutes at most. When the connection closes, OpenaiRealtime reconnects and sends the conversation written so far to the new session.