核心运行时辅助工具
api.runtime 向插件代码暴露的内容,以及相应提供商注册的归属规则。本文是插件架构内部机制指南的一部分。
运行时辅助工具¶
插件可以通过 api.runtime 访问选定的核心辅助工具。对于 TTS:
const clip = await api.runtime.tts.textToSpeech({
text: "Hello from OpenClaw",
cfg: api.config,
});
const result = await api.runtime.tts.textToSpeechTelephony({
text: "Hello from OpenClaw",
cfg: api.config,
});
const voices = await api.runtime.tts.listVoices({
provider: "elevenlabs",
cfg: api.config,
});
注意:
textToSpeech返回用于文件/语音笔记场景的标准核心 TTS 输出负载。- 使用核心
tts配置和提供商选择。 - 返回 PCM 音频缓冲区和采样率。插件必须针对提供商进行重采样/编码。
listVoices按提供商可选。用于厂商自有的语音选择器或设置流程。- 核心会将已解析的请求截止时间传递给提供商的
listVoices钩子;提供商特定的超时设置可能会覆盖它。 - 语音列表可包含更丰富的元数据,例如区域设置、性别和个性标签,以支持感知提供商的选取器。
- 电话场景需要实现
synthesizeTelephony的提供商;核心会跳过未实现该方法的提供商,并报告unsupported_for_telephony。实现该方法的内置提供商包括:azure-speech、elevenlabs、fish-audio-speech、google、gradium、inworld、openai、tts-local-cli和xai。microsoft未实现。
插件也可以通过 api.registerSpeechProvider(...) 注册语音提供商。
api.registerSpeechProvider({
id: "acme-speech",
label: "Acme Speech",
isConfigured: ({ config }) => Boolean(config.messages?.tts),
synthesize: async (req) => {
return {
audioBuffer: Buffer.from([]),
outputFormat: "mp3",
fileExtension: ".mp3",
voiceCompatible: false,
};
},
});
注意:
- 将 TTS 策略、回退和回复下发保留在核心中。
- 使用语音提供商来处理厂商自有的合成行为。
- 旧的 Microsoft
edge输入会被规范化为microsoft提供商 ID。 - 首选的归属模型以公司为导向:随着 OpenClaw 增加这些能力契约,一个厂商插件可以同时拥有文本、语音、图像以及未来的媒体提供商。
对于图像/音频/视频理解,插件应注册一个类型化的媒体理解提供商,而不是使用通用的键/值集合:
api.registerMediaUnderstandingProvider({
id: "google",
capabilities: ["image", "audio", "video"],
describeImage: async (req) => ({ text: "..." }),
transcribeAudio: async (req) => ({ text: "..." }),
describeVideo: async (req) => ({ text: "..." }),
});
注意:
- 将编排、回退、配置和渠道接线保留在核心中。
- 将厂商行为保留在提供商插件中。
- 增量扩展应保持类型化:新增可选方法、新的可选结果字段、新的可选能力。
- 视频生成已遵循相同的模式:
- 核心拥有能力契约和运行时辅助工具
- 厂商插件注册
api.registerVideoGenerationProvider(...) - 功能/渠道插件消费
api.runtime.videoGeneration.*
对于媒体理解运行时辅助工具,插件可以调用:
const image = await api.runtime.mediaUnderstanding.describeImageFile({
filePath: "/tmp/inbound-photo.jpg",
cfg: api.config,
agentDir: "/tmp/agent",
});
const video = await api.runtime.mediaUnderstanding.describeVideoFile({
filePath: "/tmp/inbound-video.mp4",
cfg: api.config,
});
// receiptImageBuffer is your own image bytes, not an SDK-provided value.
const extraction = await api.runtime.mediaUnderstanding.extractStructuredWithModel({
provider: "codex",
model: "gpt-6-astra",
input: [
{
type: "image",
buffer: receiptImageBuffer,
fileName: "receipt.png",
mime: "image/png",
},
{ type: "text", text: "Use the printed fields as the source of truth." },
],
instructions: "Return entities and searchable tags.",
schemaName: "example.evidence",
jsonSchema: {
type: "object",
properties: {
entities: { type: "array", items: { type: "string" } },
tags: { type: "array", items: { type: "string" } },
},
},
cfg: api.config,
});
对于音频转写,插件可以使用媒体理解运行时,也可以使用旧的 STT 别名:
const { text } = await api.runtime.mediaUnderstanding.transcribeAudioFile({
filePath: "/tmp/inbound-audio.ogg",
cfg: api.config,
// Optional when MIME cannot be inferred reliably:
mime: "audio/ogg",
});
注意:
api.runtime.mediaUnderstanding.*是用于图像/音频/视频理解的首选共享接口。extractStructuredWithModel(...)是面向插件的接缝,用于有界、提供商自有的图像优先提取。至少包含一个图像输入;文本输入作为补充上下文。产品插件拥有自己的路由和模式,而 OpenClaw 拥有提供商/运行时边界。- 使用核心媒体理解音频配置(
tools.media.audio)和提供商回退顺序。 - 当未产生任何转写输出时(例如跳过/不支持的输入),返回
{ text: undefined }。
插件还可以通过 api.runtime.subagent 启动后台子代理运行:
const result = await api.runtime.subagent.run({
sessionKey: "agent:main:subagent:search-helper",
message: "Expand this query into focused follow-up searches.",
toolsAlsoAllow: ["my_plugin_progress"],
provider: "openai",
model: "gpt-4.1-mini",
deliver: false,
});
注意:
provider和model是每次运行的可选覆盖项,不是持久的会话变更。toolsAlsoAllow接受调用插件注册的精确且唯一拥有的工具名称。核心名称和歧义名称会被拒绝。它是对正常配置文件的补充,但操作员的允许列表和拒绝列表仍然具有最终决定权。- OpenClaw 仅对受信任的调用者兑现这些覆盖字段。
- 对于插件自有的回退运行,操作员必须通过
plugins.entries.<id>.subagent.allowModelOverride: true主动选择启用。 - 使用
plugins.entries.<id>.subagent.allowedModels将受信任插件限制为特定的规范provider/model目标,或使用"*"显式允许任何目标。 - 不受信任的插件子代理运行仍然有效,但覆盖请求会被拒绝,而不是静默回退。
- 插件创建的子代理会话会标记创建该会话的插件 ID。回退的
api.runtime.subagent.deleteSession(...)只能删除这些拥有的会话;任意会话删除仍需要管理员范围的 Gateway 请求。
对于 Web 搜索,插件可以使用共享的运行时助手(runtime helper),而无需直接深入 agent 工具的内部接线:
const providers = api.runtime.webSearch.listProviders({
config: api.config,
});
const result = await api.runtime.webSearch.search({
config: api.config,
args: {
query: "OpenClaw plugin runtime helpers",
count: 5,
},
});
插件还可以通过 api.registerWebSearchProvider(...) 注册 Web 搜索提供商。
说明:
- 将提供商的选取、凭据解析和共享请求语义保留在核心(core)中。
- 使用 Web 搜索提供商来处理厂商特定的搜索传输。
api.runtime.webSearch.*是功能/渠道插件在需要搜索行为时的首选共享接口,无需依赖 agent 工具包装器。
api.runtime.imageGeneration¶
const result = await api.runtime.imageGeneration.generate({
config: api.config,
args: { prompt: "A friendly lobster mascot", size: "1024x1024" },
});
const providers = api.runtime.imageGeneration.listProviders({
config: api.config,
});
generate(...):使用已配置的图像生成提供商链路生成图像。listProviders(...):列出可用的图像生成提供商及其能力。
本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw