跳转至

文本转语音字段参考

Field reference

Provider apiKey fields, including personas.<id>.providers.<provider>.apiKey, can be raw strings or SecretRefs in global, per-agent, and Discord voice TTS config. During cold Gateway startup, an unavailable TTS SecretRef marks the built-in TTS capability configured-unavailable instead of stopping the Gateway. tts.speak then returns UNAVAILABLE with reason SECRET_SURFACE_UNAVAILABLE, and no provider request is sent. Status and doctor list the degraded TTS owner and its config paths. The explicit refs remain in the runtime snapshot, so environment or profile credentials cannot silently select a different account. Reloads and config-write preflight apply the owner-aware degradation policy: an unchanged eligible TTS owner may keep its last-known-good credentials as stale, while a new or changed failure becomes cold without blocking healthy owners. Structurally invalid refs and resolved values still fail startup or reject the update.

Top-level tts.*
auto "off" | "always" | "inbound" | "tagged" (path)
Auto-TTS mode. inbound only sends audio after an inbound voice message; tagged only sends audio when the reply includes [[tts:...]] directives or a [[tts:text]] block.
enabled boolean (path) deprecated
Legacy toggle. openclaw doctor --fix migrates this to auto.
mode "final" | "all" (path) default: final
"all" includes tool/block replies in addition to final replies.
provider string (path)
Speech provider id. When unset, OpenClaw uses the first configured provider in registry auto-select order. Legacy provider: "edge" is rewritten to "microsoft" by openclaw doctor --fix.
persona string (path)

Active persona id from personas. Normalized to lowercase.

Stable spoken identity. Fields: label, description, provider, fallbackPolicy, providers.<provider>. See Personas.

summaryModel string (path)
Cheap model for auto-summary; defaults to agents.defaults.model.primary. Accepts provider/model or a configured model alias.
modelOverrides object (path)

Allow the model to emit TTS directives. enabled defaults to true; allowProvider defaults to false.

Provider-owned settings keyed by speech provider id. Legacy direct blocks (tts.openai, .elevenlabs, .microsoft, .edge) are rewritten by openclaw doctor --fix; commit only tts.providers.<id>.

maxTextLength number (path) default: 4096
Hard cap for TTS input characters. /tts audio, tts.convert, and tts.speak fail if exceeded.
timeoutMs number (path) default: 30000
Request timeout in milliseconds. A per-call timeoutMs (agent tool, gateway) wins when set; otherwise an explicitly configured tts.timeoutMs wins over any plugin-authored provider default.
Azure Speech
apiKey string (path)
Env: AZURE_SPEECH_KEY, AZURE_SPEECH_API_KEY, or SPEECH_KEY.
region string (path)
Azure Speech region (e.g. eastus). Env: AZURE_SPEECH_REGION or SPEECH_REGION.
endpoint string (path)
Optional Azure Speech endpoint override (alias baseUrl).
speakerVoice string (path)
Azure voice ShortName. Default en-US-JennyNeural. Legacy alias: voice.
lang string (path)
SSML language code. Default en-US.
outputFormat string (path)
Azure X-Microsoft-OutputFormat for standard audio. Default audio-24khz-48kbitrate-mono-mp3.
voiceNoteOutputFormat string (path)
Azure X-Microsoft-OutputFormat for voice-note output. Default ogg-24khz-16bit-mono-opus.
ElevenLabs
apiKey string (path)
Falls back to ELEVENLABS_API_KEY or XI_API_KEY.
modelId string (path)
Model id. Default eleven_multilingual_v2; set modelId: "eleven_v3" for v3. The config key model is ignored. Legacy ids eleven_turbo_v2_5/eleven_turbo_v2 are normalized to the matching flash model.
speakerVoiceId string (path)
ElevenLabs voice id. Default pMsXgVXv3BLzUgSXRplE. Legacy alias: voiceId.
voiceSettings object (path)
stability, similarityBoost, style (each 0..1, defaults 0.5/0.75/0), useSpeakerBoost (true|false, default true), speed (0.5..2.0, default 1.0).
applyTextNormalization "auto" | "on" | "off" (path)
Text normalization mode.
languageCode string (path)
2-letter ISO 639-1 (e.g. en, de).
seed number (path)
Integer 0..4294967295 for best-effort determinism.
baseUrl string (path)
Override ElevenLabs API base URL.
Google Gemini
apiKey string (path)
Falls back to GEMINI_API_KEY / GOOGLE_API_KEY. If omitted, TTS can reuse models.providers.google.apiKey before env fallback.
model string (path)
Gemini TTS model. Default gemini-3.1-flash-tts-preview. Set gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts to opt in to Gemini 3.8, which OpenClaw sends through the Interactions API. gemini-2.5-flash-preview-tts and gemini-2.5-pro-preview-tts also work.
speakerVoice string (path)
Gemini prebuilt voice name. Default Kore. Legacy aliases: voiceName, voice.
audioProfile string (path)
Natural-language delivery style. Gemini 3.8 sends it as speech_metadata.style. Gemini 3.1 and 2.5 preview models prepend it to the spoken text.
speakerName string (path)
Optional speaker label. Gemini 3.8 sends it as the structured speech_metadata.speaker label alongside the single configured voice. Older preview models prepend Speaker name: before the spoken text.
speakers array (path)
Exactly two { speaker, voice, style? } entries. On Gemini 3.8, lines that start with one of the two configured names and a colon (with or without a following space) become conversational turns and the name is not spoken; any other line, including other Word: text prose, stays inside the current turn. Transcripts without configured labels stay single-voice.
promptTemplate "audio-profile-v1" (path)
On Gemini 3.1 and 2.5 preview models, wrap active persona fields in a deterministic prompt. On Gemini 3.8, personaPrompt is sent as speech_metadata.style and is not read aloud; the persona label is not sent.
personaPrompt string (path)
Google-specific persona direction. Gemini 3.8 sends it as style metadata. Older preview models append it to the audio-profile template's Director's Notes.
baseUrl string (path)
Only https://generativelanguage.googleapis.com is accepted.
Gradium
apiKey string (path)
Env: GRADIUM_API_KEY.
baseUrl string (path)
HTTPS Gradium API URL on api.gradium.ai. Default https://api.gradium.ai.
speakerVoiceId string (path)
Default Emma (YTpq7expH9539ERJ). Legacy alias: voiceId.
Inworld
<a id="inworld-primary" />
apiKey string (path)
Env: INWORLD_API_KEY.
baseUrl string (path)
Default https://api.inworld.ai.
modelId string (path)
Default inworld-tts-1.5-max. Also: inworld-tts-1.5-mini, inworld-tts-1-max, inworld-tts-1.
speakerVoiceId string (path)
Default Sarah. Legacy alias: voiceId.
temperature number (path)
Sampling temperature 0..2 (exclusive of 0).
Local CLI (tts-local-cli)
command string (path)
Local executable or command string for CLI TTS.
args string[] (path)
Command arguments. Supports {{Text}}, {{OutputPath}}, {{OutputDir}}, {{OutputBase}} placeholders.
outputFormat "mp3" | "opus" | "wav" (path)
Expected CLI output format. Default mp3 for audio attachments.
timeoutMs number (path)
Command timeout in milliseconds. Overrides the resolved TTS request timeout when set. When omitted, follows the request timeout; the plugin default is 120000.
cwd string (path)

Optional command working directory.

Optional environment overrides for the command.

Command stdout and generated or converted audio are limited to 50 MiB. Diagnostic stderr is limited to 1 MiB. OpenClaw terminates the command and fails synthesis when either limit is exceeded.

Microsoft (no API key)
enabled boolean (path)
Allow Microsoft speech usage.
speakerVoice string (path)
Microsoft neural voice name (e.g. en-US-MichelleNeural). Legacy alias: voice. If the default English voice is in effect and reply text is CJK-dominant, OpenClaw auto-switches to zh-CN-XiaoxiaoNeural.
lang string (path)
Language code (e.g. en-US).
outputFormat string (path)
Microsoft output format. Default audio-24khz-48kbitrate-mono-mp3. Not all formats are supported by the bundled Edge-backed transport.
rate / pitch / volume string (path)
Percent strings (e.g. +10%, -5%).
saveSubtitles boolean (path)
Write JSON subtitles alongside the audio file.
proxy string (path)
Proxy URL for Microsoft speech requests.
timeoutMs number (path)
Request timeout override (ms).
edge.* object (path) deprecated
Legacy alias. Run openclaw doctor --fix to rewrite persisted config to providers.microsoft.
MiniMax
apiKey string (path)
Falls back to MINIMAX_API_KEY. Token Plan auth via MINIMAX_OAUTH_TOKEN, MINIMAX_CODE_PLAN_KEY, or MINIMAX_CODING_API_KEY.
baseUrl string (path)
Default https://api.minimax.io. Env: MINIMAX_API_HOST.
model string (path)
Default speech-2.8-hd. Env: MINIMAX_TTS_MODEL.
speakerVoiceId string (path)
Default English_expressive_narrator. Env: MINIMAX_TTS_VOICE_ID. Legacy alias: voiceId.
speed number (path)
0.5..2.0. Default 1.0.
vol number (path)
(0, 10]. Default 1.0.
pitch number (path)
Integer -12..12. Default 0. Fractional values are truncated before the request.
OpenAI
apiKey string (path)
Falls back to OPENAI_API_KEY.
model string (path)
OpenAI TTS model id. Default gpt-4o-mini-tts.
speakerVoice string (path)
Voice name (e.g. alloy, cedar). Default coral. Legacy alias: voice.
instructions string (path)
Explicit OpenAI instructions field. When set, persona prompt fields are not auto-mapped.
responseFormat "mp3" | "opus" | "wav" (path)

Explicit response format. When omitted, OpenClaw selects Opus for voice-note targets and MP3 otherwise. Use wav for compatible local endpoints that do not encode compressed audio.

Extra JSON fields merged into /audio/speech request bodies after generated OpenAI TTS fields. Use this for OpenAI-compatible endpoints such as Kokoro that require provider-specific keys like lang; unsafe prototype keys are ignored.

baseUrl string (path)
Override the OpenAI TTS endpoint. Resolution order: config → OPENAI_TTS_BASE_URL → https://api.openai.com/v1. Non-default values are treated as OpenAI-compatible TTS endpoints, so custom model and voice names are accepted, and speed loses its 0.25..4.0 range check.
OpenRouter
apiKey string (path)
Env: OPENROUTER_API_KEY. Can reuse models.providers.openrouter.apiKey.
baseUrl string (path)
Default https://openrouter.ai/api/v1. Legacy https://openrouter.ai/v1 is normalized.
model string (path)
Default hexgrad/kokoro-82m. Alias: modelId.
speakerVoice string (path)
Default af_alloy. Legacy aliases: voice, voiceId.
responseFormat "mp3" | "pcm" (path)
Default mp3.
speed number (path)
Provider-native speed override.
Volcengine (BytePlus Seed Speech)
apiKey string (path)
Env: VOLCENGINE_TTS_API_KEY or BYTEPLUS_SEED_SPEECH_API_KEY.
resourceId string (path)
Default seed-tts-1.0. Env: VOLCENGINE_TTS_RESOURCE_ID. Use seed-tts-2.0 when your project has TTS 2.0 entitlement.
appKey string (path)
App key header. Default aGjiRDfUWi. Env: VOLCENGINE_TTS_APP_KEY.
baseUrl string (path)
Override the Seed Speech TTS HTTP endpoint. Env: VOLCENGINE_TTS_BASE_URL.
speakerVoice string (path)
Voice type. Default en_female_anna_mars_bigtts. Env: VOLCENGINE_TTS_VOICE. Legacy alias: voice.
speedRatio number (path)
Provider-native speed ratio, 0.2..3.
emotion string (path)
Provider-native emotion tag.
appId / token / cluster string (path) deprecated
Legacy Volcengine Speech Console fields. Env: VOLCENGINE_TTS_APPID, VOLCENGINE_TTS_TOKEN, VOLCENGINE_TTS_CLUSTER (default volcano_tts).
xAI
apiKey string (path)
Env: XAI_API_KEY.
baseUrl string (path)
Default https://api.x.ai/v1. Env: XAI_BASE_URL.
speakerVoiceId string (path)
Default eve. With auth, openclaw infer tts voices --provider xai fetches the current built-in catalog; without auth it lists offline fallbacks ara, eve, leo, rex, and sal. Account custom voice IDs are forwarded even when absent from the built-in list. Legacy alias: voiceId.
language string (path)
BCP-47 language code or auto. Default en.
responseFormat "mp3" | "wav" | "pcm" | "mulaw" | "alaw" (path)
Default mp3.
speed number (path)
Provider-native speed override, 0.7..1.5.
Xiaomi MiMo
apiKey string (path)
Env: XIAOMI_API_KEY.
baseUrl string (path)
Default https://api.xiaomimimo.com/v1. Env: XIAOMI_BASE_URL.
model string (path)
Default mimo-v2.5-tts. Env: XIAOMI_TTS_MODEL. Also supports mimo-v2.5-tts-voicedesign.
speakerVoice string (path)
Default mimo_default for preset-voice models. Env: XIAOMI_TTS_VOICE. Legacy alias: voice. Not sent for mimo-v2.5-tts-voicedesign.
format "mp3" | "wav" (path)
Default mp3. Env: XIAOMI_TTS_FORMAT.
style string (path)
Optional natural-language style instruction sent as the user message; not spoken. For mimo-v2.5-tts-voicedesign, this is the voice-design prompt; OpenClaw supplies a default when omitted.

本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw