Skip to content
curl -X POST https://api.wxiai.com/v1/audio/speech \
  -H "Authorization: Bearer $WXIAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-tts",
    "input": "你好,这是一段语音合成测试。",
    "voice": "eve",
    "response_format": "mp3",
    "language": "zh"
  }' \
  --output speech.mp3
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.wxiai.com/v1",
)

with client.audio.speech.with_streaming_response.create(
    model="grok-tts",
    voice="eve",
    input="你好,这是一段语音合成测试。",
) as resp:
    resp.stream_to_file("speech.mp3")
curl -X POST https://api.wxiai.com/v1/audio/speech \
  -H "Authorization: Bearer $WXIAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-tts",
    "input": "Hello world.",
    "voice": "eve",
    "language": "en",
    "with_timestamps": true
  }' \
  --output response.json
# 成功时返回的是音频二进制流(Content-Type 如 audio/mpeg),不是 JSON。
# 用 curl 的 --output 存文件,或让 SDK 直接写文件。
#
# 传了 with_timestamps: true 时才改成返回 JSON:
# { "audio": "<base64>", "content_type": "audio/mpeg",
#   "duration": 0.92, "audio_timestamps": {...} }
{
  "error": {
    "message": "language is required",
    "type": "wxi_api_error",
    "param": "language",
    "code": "invalid_request_error"
  }
}
OpenAI 兼容层

文生音 TTS

把文本合成为语音,直接返回音频流。同步接口,一次请求拿结果。

POST/v1/audio/speech

端点 ​

http
POST /v1/audio/speech

这是 OpenAI 兼容路径。同一能力在原生透传层的写法见 文生音 TTS(原生透传)。

请求 ​

bash
curl -X POST https://api.wxiai.com/v1/audio/speech \
  -H "Authorization: Bearer $WXIAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-tts",
    "input": "你好,这是一段语音合成测试。",
    "voice": "eve",
    "response_format": "mp3",
    "language": "zh"
  }' \
  --output speech.mp3

用 OpenAI SDK 可以直接流式写文件:

python
with client.audio.speech.with_streaming_response.create(
    model="grok-tts", voice="eve", input="你好,这是一段语音合成测试。",
) as resp:
    resp.stream_to_file("speech.mp3")

请求参数 ​

三个核心字段用 OpenAI 的叫法,网关会转成 Grok 字段:

OpenAI 字段转成 Grok 字段
inputtext
voicevoice_id
response_formatoutput_format.codec

其余字段原样透传,所以 Grok 自己的参数在这里照样能用:

参数必填说明
input是要合成的文本,最长 15,000 字符。支持语音标签
voice否音色 ID,不传默认 eve。大小写不敏感
language是BCP-47 语言码(en、zh、pt-BR)或 auto。不传会报错
response_format否mp3(默认)/ wav / pcm / mulaw / alaw
output_format否Grok 原生的完整格式对象(codec + sample_rate + bit_rate),优先于 response_format
speed否语速倍率,0.7–1.5,默认 1.0
with_timestamps否为 true 时返回 JSON 信封,含 base64 音频 + 逐字符时间戳
text_normalization否为 true 时先把数字、缩写、符号转成口语形式再合成
replace否发音替换表,见 TTS 总览
model否不传默认 grok-tts

返回 ​

成功时返回音频二进制流,Content-Type 取决于编码:

  • curl 用 --output speech.mp3 存文件
  • Python 用 resp.stream_to_file("speech.mp3")
  • 别用 response.json(),会解析失败

例外:传了 with_timestamps: true 时改成返回 application/json:

json
{
  "audio": "<base64 音频>",
  "content_type": "audio/mpeg",
  "duration": 0.92,
  "audio_timestamps": {
    "graph_chars": ["H", "e", "l", "l", "o"],
    "graph_times": [
      [0.00, 0.06], [0.06, 0.12], [0.12, 0.18], [0.18, 0.24], [0.24, 0.34]
    ]
  }
}

graph_chars[i] 与 graph_times[i] 按下标一一对应,每个元素是 [开始秒, 结束秒]。做字幕、卡拉OK 高亮、口型同步都用它。

这一层的注意点 ​

  • language 是必填的:不传直接报参数错误。不想指定就传 auto。
  • 把返回当 JSON 解析会失败:默认返回音频流,直接存文件。
  • model 传 grok-tts 即可:TTS 不挑模型名,网关用它做计费和路由。
  • 长文本会被拒:单次上限 15,000 字符,超了改用流式接口(见总览)。

相关页 ​

基于 Apache-2.0 许可发布