Skip to content
import WebSocket from 'ws'

const ws = new WebSocket(
  'wss://api.wxiai.com/xai/v1/realtime?model=grok-voice-latest',
  { headers: { Authorization: 'Bearer YOUR_API_KEY' } },
)

ws.on('open', () => {
  ws.send(JSON.stringify({
    type: 'session.update',
    session: {
      voice: 'eve',
      instructions: '你是一个耐心的中文客服助手。',
      turn_detection: { type: 'server_vad' },
      tools: [{ type: 'web_search' }],
    },
  }))
})

ws.on('message', (raw) => {
  const event = JSON.parse(raw.toString())
  if (event.type === 'response.output_audio_transcript.delta') {
    process.stdout.write(event.delta)
  }
  if (event.type === 'response.output_audio.delta') {
    // event.delta 是 base64 音频,解码后播放
  }
})
# 浏览器的 WebSocket API 不能自定义请求头,所以先用你的 Key
# 换一个短期临时凭证出来,再拿它建连。这个端点本身不计费。
curl -X POST https://api.wxiai.com/xai/v1/realtime/client_secrets \
  -H "Authorization: Bearer $WXIAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "expires_after": { "seconds": 300 } }'
// 子协议字段是浏览器唯一能自定义的握手信息
const ws = new WebSocket('wss://api.wxiai.com/xai/v1/realtime', [
  `xai-client-secret.${ephemeralToken}`,
])
// 必须一字不差的台词不用过模型,直接走 TTS
ws.send(JSON.stringify({
  type: 'conversation.item.create',
  item: {
    type: 'force_message',
    role: 'assistant',
    interruptible: false,
    content: [{ type: 'output_text', text: '本次通话将被录音。' }],
  },
}))
{ "type": "session.created" }
{ "type": "conversation.created" }
{ "type": "session.updated" }
{ "type": "input_audio_buffer.speech_started" }
{ "type": "conversation.item.input_audio_transcription.completed", "transcript": "你好" }
{ "type": "response.created" }
{ "type": "response.output_audio_transcript.delta", "delta": "你好" }
{ "type": "response.output_audio.delta", "delta": "<base64 音频>" }
{ "type": "response.output_audio.done" }
{ "type": "response.done" }
原生透传层

实时语音

WebSocket 长连接,双向实时语音。事件原样转发给 Grok,不做转换,报错是官方原文。

WSS/wss:/api.wxiai.com/xai/v1/realtime

端点 ​

WSS  /xai/v1/realtime?model=grok-voice-latest
POST /xai/v1/realtime/client_secrets     # 换临时凭证(不计费)

这是 原生透传路径。事件原样转发给 Grok,不做转换——Grok 新加的会话参数和工具类型当天就能用。

同一能力在 OpenAI 兼容层的写法见 实时语音(OpenAI 兼容)。

什么情况该用它 ​

你要做的用哪个
语音对话助手(像打电话一样)用这个
一次性的「录音 → 转文字」音生文 STT,更简单
普通的文本流式输出对话补全 就够,不用上 WebSocket

建连 ​

javascript
import WebSocket from 'ws'

const ws = new WebSocket(
  'wss://api.wxiai.com/xai/v1/realtime?model=grok-voice-latest',
  { headers: { Authorization: 'Bearer YOUR_API_KEY' } },
)

ws.on('open', () => {
  ws.send(JSON.stringify({
    type: 'session.update',
    session: {
      voice: 'eve',
      instructions: '你是一个耐心的中文客服助手。',
      turn_detection: { type: 'server_vad' },
      tools: [{ type: 'web_search' }],
    },
  }))
})

model 是 URL 查询参数,不是 session.update 里的字段。

模型说明
grok-voice-latest别名,指向当前旗舰,推荐
grok-voice-think-fast-2.0旗舰语音模型
grok-voice-think-fast-1.0上一代语音模型

浏览器怎么连 ​

浏览器原生的 WebSocket 构造函数不支持自定义请求头,没法带 Authorization。xAI 的做法是两步走:

  1. 你的后端用 API Key 调 POST /xai/v1/realtime/client_secrets,换一个短期临时凭证
  2. 前端拿这个凭证,用子协议字段建连
bash
curl -X POST https://api.wxiai.com/xai/v1/realtime/client_secrets \
  -H "Authorization: Bearer $WXIAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "expires_after": { "seconds": 300 } }'
javascript
// 浏览器端
const ws = new WebSocket('wss://api.wxiai.com/xai/v1/realtime', [
  `xai-client-secret.${ephemeralToken}`,
])

网关同时支持两种子协议前缀:

前缀用途
xai-client-secret.<临时凭证>xAI 官方风格,配合 client_secrets 使用
openai-insecure-api-key.<密钥>OpenAI 风格

不要把长期 API Key 写进前端

那样任何访问者都能从源码里拿走它。client_secrets 这个端点存在的唯一目的就是避免这件事,而且它本身不计费。

会话配置 ​

建连后第一件事就是发 session.update:

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    voice: 'eve',
    instructions: '你是一个耐心的中文客服助手。',
    turn_detection: {
      type: 'server_vad',
      threshold: 0.85,
      silence_duration_ms: 600,
      prefix_padding_ms: 333,
    },
    audio: {
      input:  { format: { type: 'audio/pcm', rate: 24000 } },
      output: { format: { type: 'audio/pcm', rate: 24000 } },
    },
    tools: [{ type: 'web_search' }, { type: 'x_search' }],
  },
}))

全部可配字段与5 类工具(file_search / web_search / x_search / mcp / function)见实时语音总览。音色 ID 与 TTS 共用同一套,见音频处理总览。

事件流 ​

你发过去:session.update、input_audio_buffer.append、input_audio_buffer.commit(仅手动模式)、input_audio_buffer.clear、conversation.item.create / .delete / .truncate、response.create、response.cancel。

服务端发回来:session.created / session.updated、conversation.created、input_audio_buffer.*、conversation.item.*、response.created、response.output_audio.delta / .done、response.output_audio_transcript.delta / .done、response.output_text.delta、response.done、error。

完整事件表、字段说明与工具调用流程见实时语音总览。

结束事件是 response.done

不是 OpenAI 的 response.completed。另外文本分片会以 response.text.delta 和 response.output_text.delta 两个名字出现,客户端两个都要处理。

固定播报(force_message) ​

合规播报、IVR 提示音、固定开场白这类必须一字不差的台词,不需要经过模型:

javascript
ws.send(JSON.stringify({
  type: 'conversation.item.create',
  item: {
    type: 'force_message',
    role: 'assistant',
    interruptible: false,
    content: [{ type: 'output_text', text: '本次通话将被录音。' }],
  },
}))
// 不要再发 response.create —— force_message 本身就是一轮回应

这一层的注意点 ​

  • model 放错地方:它是 URL 查询参数,不是 session.update 字段。
  • 在 server VAD 模式下自己发 commit:只有 turn_detection 为 null 时才需要手动提交。
  • 等 response.completed:Grok 用的是 response.done。
  • format 和 transport 搞混:format 选编码,transport 选这些字节怎么在 WebSocket 上走。
  • 断线不重连:120 分钟上限 + 网络抖动,断开是常态不是异常。

相关页 ​

基于 Apache-2.0 许可发布