文本转语音字段参考
Field reference¶
Provider apiKey fields, including personas.<id>.providers.<provider>.apiKey,
can be raw strings or SecretRefs in global, per-agent, and Discord voice TTS config.
During cold Gateway startup, an unavailable TTS SecretRef marks the built-in TTS capability
configured-unavailable instead of stopping the Gateway. tts.speak then returns
UNAVAILABLE with reason SECRET_SURFACE_UNAVAILABLE, and no provider request is
sent. Status and doctor list the degraded TTS owner and its config paths. The
explicit refs remain in the runtime snapshot, so environment or profile
credentials cannot silently select a different account. Reloads and config-write
preflight apply the owner-aware degradation policy: an unchanged eligible TTS
owner may keep its last-known-good credentials as stale, while a new or changed
failure becomes cold without blocking healthy owners. Structurally invalid refs
and resolved values still fail startup or reject the update.
Top-level tts.*
auto"off" | "always" | "inbound" | "tagged" (path)- Auto-TTS mode.
inboundonly sends audio after an inbound voice message;taggedonly sends audio when the reply includes[[tts:...]]directives or a[[tts:text]]block. enabledboolean (path) deprecated- Legacy toggle.
openclaw doctor --fixmigrates this toauto. mode"final" | "all" (path) default:final"all"includes tool/block replies in addition to final replies.providerstring (path)- Speech provider id. When unset, OpenClaw uses the first configured provider in registry auto-select order. Legacy
provider: "edge"is rewritten to"microsoft"byopenclaw doctor --fix. personastring (path)-
Active persona id from
personas. Normalized to lowercase.Stable spoken identity. Fields: label,description,provider,fallbackPolicy,providers.<provider>. See Personas. summaryModelstring (path)- Cheap model for auto-summary; defaults to
agents.defaults.model.primary. Acceptsprovider/modelor a configured model alias. modelOverridesobject (path)-
Allow the model to emit TTS directives.
enableddefaults totrue;allowProviderdefaults tofalse.Provider-owned settings keyed by speech provider id. Legacy direct blocks ( tts.openai,.elevenlabs,.microsoft,.edge) are rewritten byopenclaw doctor --fix; commit onlytts.providers.<id>. maxTextLengthnumber (path) default:4096- Hard cap for TTS input characters.
/tts audio,tts.convert, andtts.speakfail if exceeded. timeoutMsnumber (path) default:30000- Request timeout in milliseconds. A per-call
timeoutMs(agent tool, gateway) wins when set; otherwise an explicitly configuredtts.timeoutMswins over any plugin-authored provider default.
Azure Speech
apiKeystring (path)- Env:
AZURE_SPEECH_KEY,AZURE_SPEECH_API_KEY, orSPEECH_KEY. regionstring (path)- Azure Speech region (e.g.
eastus). Env:AZURE_SPEECH_REGIONorSPEECH_REGION. endpointstring (path)- Optional Azure Speech endpoint override (alias
baseUrl). speakerVoicestring (path)- Azure voice ShortName. Default
en-US-JennyNeural. Legacy alias:voice. langstring (path)- SSML language code. Default
en-US. outputFormatstring (path)- Azure
X-Microsoft-OutputFormatfor standard audio. Defaultaudio-24khz-48kbitrate-mono-mp3. voiceNoteOutputFormatstring (path)- Azure
X-Microsoft-OutputFormatfor voice-note output. Defaultogg-24khz-16bit-mono-opus.
ElevenLabs
apiKeystring (path)- Falls back to
ELEVENLABS_API_KEYorXI_API_KEY. modelIdstring (path)- Model id. Default
eleven_multilingual_v2; setmodelId: "eleven_v3"for v3. The config keymodelis ignored. Legacy idseleven_turbo_v2_5/eleven_turbo_v2are normalized to the matchingflashmodel. speakerVoiceIdstring (path)- ElevenLabs voice id. Default
pMsXgVXv3BLzUgSXRplE. Legacy alias:voiceId. voiceSettingsobject (path)stability,similarityBoost,style(each0..1, defaults0.5/0.75/0),useSpeakerBoost(true|false, defaulttrue),speed(0.5..2.0, default1.0).applyTextNormalization"auto" | "on" | "off" (path)- Text normalization mode.
languageCodestring (path)- 2-letter ISO 639-1 (e.g.
en,de). seednumber (path)- Integer
0..4294967295for best-effort determinism. baseUrlstring (path)- Override ElevenLabs API base URL.
Google Gemini
apiKeystring (path)- Falls back to
GEMINI_API_KEY/GOOGLE_API_KEY. If omitted, TTS can reusemodels.providers.google.apiKeybefore env fallback. modelstring (path)- Gemini TTS model. Default
gemini-3.1-flash-tts-preview. Setgemini-3.8-flash-ttsorgemini-3.8-flash-lite-ttsto opt in to Gemini 3.8, which OpenClaw sends through the Interactions API.gemini-2.5-flash-preview-ttsandgemini-2.5-pro-preview-ttsalso work. speakerVoicestring (path)- Gemini prebuilt voice name. Default
Kore. Legacy aliases:voiceName,voice. audioProfilestring (path)- Natural-language delivery style. Gemini 3.8 sends it as
speech_metadata.style. Gemini 3.1 and 2.5 preview models prepend it to the spoken text. speakerNamestring (path)- Optional speaker label. Gemini 3.8 sends it as the structured
speech_metadata.speakerlabel alongside the single configured voice. Older preview models prependSpeaker name:before the spoken text. speakersarray (path)- Exactly two
{ speaker, voice, style? }entries. On Gemini 3.8, lines that start with one of the two configured names and a colon (with or without a following space) become conversational turns and the name is not spoken; any other line, including otherWord: textprose, stays inside the current turn. Transcripts without configured labels stay single-voice. promptTemplate"audio-profile-v1" (path)- On Gemini 3.1 and 2.5 preview models, wrap active persona fields in a deterministic prompt. On Gemini 3.8,
personaPromptis sent asspeech_metadata.styleand is not read aloud; the persona label is not sent. personaPromptstring (path)- Google-specific persona direction. Gemini 3.8 sends it as style metadata. Older preview models append it to the audio-profile template's Director's Notes.
baseUrlstring (path)- Only
https://generativelanguage.googleapis.comis accepted.
Gradium
apiKeystring (path)- Env:
GRADIUM_API_KEY. baseUrlstring (path)- HTTPS Gradium API URL on
api.gradium.ai. Defaulthttps://api.gradium.ai. speakerVoiceIdstring (path)- Default Emma (
YTpq7expH9539ERJ). Legacy alias:voiceId.
Inworld
<a id="inworld-primary" />
apiKeystring (path)- Env:
INWORLD_API_KEY. baseUrlstring (path)- Default
https://api.inworld.ai. modelIdstring (path)- Default
inworld-tts-1.5-max. Also:inworld-tts-1.5-mini,inworld-tts-1-max,inworld-tts-1. speakerVoiceIdstring (path)- Default
Sarah. Legacy alias:voiceId. temperaturenumber (path)- Sampling temperature
0..2(exclusive of 0).
Local CLI (tts-local-cli)
commandstring (path)- Local executable or command string for CLI TTS.
argsstring[] (path)- Command arguments. Supports
{{Text}},{{OutputPath}},{{OutputDir}},{{OutputBase}}placeholders. outputFormat"mp3" | "opus" | "wav" (path)- Expected CLI output format. Default
mp3for audio attachments. timeoutMsnumber (path)- Command timeout in milliseconds. Overrides the resolved TTS request timeout when set. When omitted, follows the request timeout; the plugin default is
120000. cwdstring (path)-
Optional command working directory.
Optional environment overrides for the command. Command stdout and generated or converted audio are limited to 50 MiB. Diagnostic stderr is limited to 1 MiB. OpenClaw terminates the command and fails synthesis when either limit is exceeded.
Microsoft (no API key)
enabledboolean (path)- Allow Microsoft speech usage.
speakerVoicestring (path)- Microsoft neural voice name (e.g.
en-US-MichelleNeural). Legacy alias:voice. If the default English voice is in effect and reply text is CJK-dominant, OpenClaw auto-switches tozh-CN-XiaoxiaoNeural. langstring (path)- Language code (e.g.
en-US). outputFormatstring (path)- Microsoft output format. Default
audio-24khz-48kbitrate-mono-mp3. Not all formats are supported by the bundled Edge-backed transport. rate / pitch / volumestring (path)- Percent strings (e.g.
+10%,-5%). saveSubtitlesboolean (path)- Write JSON subtitles alongside the audio file.
proxystring (path)- Proxy URL for Microsoft speech requests.
timeoutMsnumber (path)- Request timeout override (ms).
edge.*object (path) deprecated- Legacy alias. Run
openclaw doctor --fixto rewrite persisted config toproviders.microsoft.
MiniMax
apiKeystring (path)- Falls back to
MINIMAX_API_KEY. Token Plan auth viaMINIMAX_OAUTH_TOKEN,MINIMAX_CODE_PLAN_KEY, orMINIMAX_CODING_API_KEY. baseUrlstring (path)- Default
https://api.minimax.io. Env:MINIMAX_API_HOST. modelstring (path)- Default
speech-2.8-hd. Env:MINIMAX_TTS_MODEL. speakerVoiceIdstring (path)- Default
English_expressive_narrator. Env:MINIMAX_TTS_VOICE_ID. Legacy alias:voiceId. speednumber (path)0.5..2.0. Default1.0.volnumber (path)(0, 10]. Default1.0.pitchnumber (path)- Integer
-12..12. Default0. Fractional values are truncated before the request.
OpenAI
apiKeystring (path)- Falls back to
OPENAI_API_KEY. modelstring (path)- OpenAI TTS model id. Default
gpt-4o-mini-tts. speakerVoicestring (path)- Voice name (e.g.
alloy,cedar). Defaultcoral. Legacy alias:voice. instructionsstring (path)- Explicit OpenAI
instructionsfield. When set, persona prompt fields are not auto-mapped. responseFormat"mp3" | "opus" | "wav" (path)-
Explicit response format. When omitted, OpenClaw selects Opus for voice-note targets and MP3 otherwise. Use
wavfor compatible local endpoints that do not encode compressed audio.Extra JSON fields merged into /audio/speechrequest bodies after generated OpenAI TTS fields. Use this for OpenAI-compatible endpoints such as Kokoro that require provider-specific keys likelang; unsafe prototype keys are ignored. baseUrlstring (path)- Override the OpenAI TTS endpoint. Resolution order: config →
OPENAI_TTS_BASE_URL→https://api.openai.com/v1. Non-default values are treated as OpenAI-compatible TTS endpoints, so custom model and voice names are accepted, andspeedloses its0.25..4.0range check.
OpenRouter
apiKeystring (path)- Env:
OPENROUTER_API_KEY. Can reusemodels.providers.openrouter.apiKey. baseUrlstring (path)- Default
https://openrouter.ai/api/v1. Legacyhttps://openrouter.ai/v1is normalized. modelstring (path)- Default
hexgrad/kokoro-82m. Alias:modelId. speakerVoicestring (path)- Default
af_alloy. Legacy aliases:voice,voiceId. responseFormat"mp3" | "pcm" (path)- Default
mp3. speednumber (path)- Provider-native speed override.
Volcengine (BytePlus Seed Speech)
apiKeystring (path)- Env:
VOLCENGINE_TTS_API_KEYorBYTEPLUS_SEED_SPEECH_API_KEY. resourceIdstring (path)- Default
seed-tts-1.0. Env:VOLCENGINE_TTS_RESOURCE_ID. Useseed-tts-2.0when your project has TTS 2.0 entitlement. appKeystring (path)- App key header. Default
aGjiRDfUWi. Env:VOLCENGINE_TTS_APP_KEY. baseUrlstring (path)- Override the Seed Speech TTS HTTP endpoint. Env:
VOLCENGINE_TTS_BASE_URL. speakerVoicestring (path)- Voice type. Default
en_female_anna_mars_bigtts. Env:VOLCENGINE_TTS_VOICE. Legacy alias:voice. speedRationumber (path)- Provider-native speed ratio,
0.2..3. emotionstring (path)- Provider-native emotion tag.
appId / token / clusterstring (path) deprecated- Legacy Volcengine Speech Console fields. Env:
VOLCENGINE_TTS_APPID,VOLCENGINE_TTS_TOKEN,VOLCENGINE_TTS_CLUSTER(defaultvolcano_tts).
xAI
apiKeystring (path)- Env:
XAI_API_KEY. baseUrlstring (path)- Default
https://api.x.ai/v1. Env:XAI_BASE_URL. speakerVoiceIdstring (path)- Default
eve. With auth,openclaw infer tts voices --provider xaifetches the current built-in catalog; without auth it lists offline fallbacksara,eve,leo,rex, andsal. Account custom voice IDs are forwarded even when absent from the built-in list. Legacy alias:voiceId. languagestring (path)- BCP-47 language code or
auto. Defaulten. responseFormat"mp3" | "wav" | "pcm" | "mulaw" | "alaw" (path)- Default
mp3. speednumber (path)- Provider-native speed override,
0.7..1.5.
Xiaomi MiMo
apiKeystring (path)- Env:
XIAOMI_API_KEY. baseUrlstring (path)- Default
https://api.xiaomimimo.com/v1. Env:XIAOMI_BASE_URL. modelstring (path)- Default
mimo-v2.5-tts. Env:XIAOMI_TTS_MODEL. Also supportsmimo-v2.5-tts-voicedesign. speakerVoicestring (path)- Default
mimo_defaultfor preset-voice models. Env:XIAOMI_TTS_VOICE. Legacy alias:voice. Not sent formimo-v2.5-tts-voicedesign. format"mp3" | "wav" (path)- Default
mp3. Env:XIAOMI_TTS_FORMAT. stylestring (path)- Optional natural-language style instruction sent as the user message; not spoken. For
mimo-v2.5-tts-voicedesign, this is the voice-design prompt; OpenClaw supplies a default when omitted.
本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw