Realtime voice

Open a WebSocket to /realtime?model=<id> (wss://) and speak OpenAI's Realtime protocol: session.update, input_audio_buffer.append, response.create… Authenticate with the Authorization header, or from a browser with the subprotocol openai-insecure-api-key.<key> (only with a key restricted to your domain). Each response is billed from the usage the model reports; input transcription supports gpt-4o-transcribe, gpt-4o-mini-transcribe and whisper-1. A session closes after 60 minutes, after 10 idle minutes, or when your balance or the key's spend limit runs out (an error event, then close code 1008).

import WebSocket from 'ws'

const ws = new WebSocket('wss://api.ultragpt.pro/v1/realtime?model=gpt-realtime', {
  headers: { Authorization: `Bearer ${process.env.ULTRAGPT_API_KEY}` },
})

ws.on('open', () => {
  ws.send(JSON.stringify({
    type: 'session.update',
    session: { instructions: 'Be brief.', voice: 'alloy' },
  }))
  ws.send(JSON.stringify({ type: 'response.create', response: { modalities: ['text'] } }))
})
ws.on('message', (data) => {
  const event = JSON.parse(data.toString())
  if (event.type === 'response.done') console.log(event.response.usage)
})