🎤Micdrop

Mistral

Mistral AI implementation for @micdrop/server.

This package provides an AI agent implementation using Mistral AI’s API and a real-time speech-to-text implementation using Mistral’s Voxtral realtime transcription.

Installation

Terminal window
npm install @micdrop/mistral

Mistral Agent

Usage with MicdropServer

import { MistralAgent } from '@micdrop/mistral'
import { MicdropServer } from '@micdrop/server'
const agent = new MistralAgent({
apiKey: process.env.MISTRAL_API_KEY || '',
model: 'mistral-large-latest',
systemPrompt: 'You are a helpful assistant',
})
// Use with MicdropServer
new MicdropServer(socket, {
agent,
// ... other options
})

Usage without MicdropServer

import { MistralAgent } from '@micdrop/mistral'
const agent = new MistralAgent({
apiKey: process.env.MISTRAL_API_KEY || '',
systemPrompt: 'You are a helpful assistant',
})
agent.on('Message', (message) => console.log('Message:', message))
agent.addUserMessage('Hello, what can you do?')
// The answer is a text stream, written as the model generates it
agent.answer().on('data', (chunk) => process.stdout.write(chunk))

Events

EventPayloadDescription
MessageMicdropConversationItemA message, a tool call or a tool result was added to the conversation.
ToolCallMicdropToolCallA tool declared with emitOutput ran, with its parameters and output.
CancelLastUserMessagenoneThe last user message was dropped because it carried no intent.
SkipAnswernoneThe agent stays silent and waits for the user to finish their sentence.
EndCallnoneThe agent decided that the call is over.
FailednoneThe agent gave up generating an answer after its retries.

See the Agent interface for the full contract.

Options

OptionTypeDefaultDescription
apiKeystringRequiredYour Mistral AI API key
modelstring'mistral-large-latest'Mistral AI model to use
systemPromptstringRequiredSystem prompt for the agent
maxRetrynumber3Maximum number of retries on API failures
retryDelaynumber500Delay in milliseconds between retries
maxStepsnumber5Maximum number of steps (for tool calls)
autoEndCallboolean | stringfalseAuto-detect when user wants to end call
autoSemanticTurnboolean | stringfalseHandle incomplete user sentences
autoIgnoreUserNoiseboolean | stringfalseFilter meaningless user sounds
extractExtractJsonOptions | ExtractTagOptionsundefinedExtract structured data from responses
onBeforeAnswerfunctionundefinedHook called before answer generation, returns true to skip it or a text to answer instead
settingsobject{}Additional Mistral AI API parameters

Available Models

Mistral AI offers several models you can choose from:

  • ministral-8b-latest - Fast and efficient 8B parameter model (default)
  • mistral-large-latest - Most capable model for complex tasks
  • mistral-small-latest - Balanced performance and cost
  • codestral-latest - Specialized for code generation

Settings Object

The settings parameter accepts any additional options from the Mistral AI Chat Completions API:

const agent = new MistralAgent({
apiKey: process.env.MISTRAL_API_KEY || '',
systemPrompt: 'You are a helpful assistant',
settings: {
temperature: 0.7, // Controls randomness (0-1)
max_tokens: 1000, // Maximum tokens in response
top_p: 0.9, // Nucleus sampling parameter
random_seed: 42, // For reproducible outputs
safe_prompt: false, // Enable/disable safety filtering
},
})

Advanced Features

The Mistral Agent supports advanced features for improved conversation handling:

Mistral STT (Speech-to-Text)

Real-time transcription using Mistral’s Voxtral realtime models.

Voxtral also runs on its own. With stt as the server’s only option, you get a dictation tool that transcribes each sentence and stays quiet.

Usage with MicdropServer

import { MistralSTT } from '@micdrop/mistral'
import { MicdropServer } from '@micdrop/server'
const stt = new MistralSTT({
apiKey: process.env.MISTRAL_API_KEY || '',
})
// Use with MicdropServer
new MicdropServer(socket, {
stt,
// ... other options
})

Usage without MicdropServer

import { MistralSTT } from '@micdrop/mistral'
import { createReadStream } from 'fs'
const stt = new MistralSTT({
apiKey: process.env.MISTRAL_API_KEY || '',
})
stt.on('Transcript', (text) => console.log('Transcript:', text))
stt.on('Failed', (chunks) => console.error('Failed:', chunks.length, 'chunks'))
// Audio is raw PCM, 16 bits, 16 kHz, mono
stt.transcribe(createReadStream('speech.pcm'))

Events

EventPayloadDescription
TranscriptstringTranscription of one utterance. The text is empty when nothing was recognized.
FailedBuffer[]Transcription gave up after its retries, with the audio chunks left pending.

See the STT interface for the full contract.

Options

OptionTypeDefaultDescription
apiKeystringRequiredYour Mistral AI API key
modelstring'voxtral-mini-transcribe-realtime-2602'Realtime transcription model to use
encodingMistralAudioEncoding'pcm_s16le'Audio encoding of the incoming stream
targetStreamingDelayMsnumberOptionalTarget streaming delay in milliseconds
connectionTimeoutnumber5000Timeout in milliseconds for WebSocket connection
transcriptionTimeoutnumber4000Timeout in milliseconds to wait for transcription
retryDelaynumber1000Delay in milliseconds between reconnection attempts
maxRetrynumber3Maximum number of reconnection attempts before failing

The Micdrop client streams 16kHz PCM16 mono audio, which matches Voxtral’s default input format, so no resampling is required.