跳转至

核心运行时辅助工具

api.runtime 向插件代码暴露的内容,以及相应提供商注册的归属规则。本文是插件架构内部机制指南的一部分。

运行时辅助工具

插件可以通过 api.runtime 访问选定的核心辅助工具。对于 TTS:

const clip = await api.runtime.tts.textToSpeech({
  text: "Hello from OpenClaw",
  cfg: api.config,
});

const result = await api.runtime.tts.textToSpeechTelephony({
  text: "Hello from OpenClaw",
  cfg: api.config,
});

const voices = await api.runtime.tts.listVoices({
  provider: "elevenlabs",
  cfg: api.config,
});

注意:

  • textToSpeech 返回用于文件/语音笔记场景的标准核心 TTS 输出负载。
  • 使用核心 tts 配置和提供商选择。
  • 返回 PCM 音频缓冲区和采样率。插件必须针对提供商进行重采样/编码。
  • listVoices 按提供商可选。用于厂商自有的语音选择器或设置流程。
  • 核心会将已解析的请求截止时间传递给提供商的 listVoices 钩子;提供商特定的超时设置可能会覆盖它。
  • 语音列表可包含更丰富的元数据,例如区域设置、性别和个性标签,以支持感知提供商的选取器。
  • 电话场景需要实现 synthesizeTelephony 的提供商;核心会跳过未实现该方法的提供商,并报告 unsupported_for_telephony。实现该方法的内置提供商包括:azure-speech、elevenlabs、fish-audio-speech、google、gradium、inworld、openai、tts-local-cli 和 xai。microsoft 未实现。

插件也可以通过 api.registerSpeechProvider(...) 注册语音提供商。

api.registerSpeechProvider({
  id: "acme-speech",
  label: "Acme Speech",
  isConfigured: ({ config }) => Boolean(config.messages?.tts),
  synthesize: async (req) => {
    return {
      audioBuffer: Buffer.from([]),
      outputFormat: "mp3",
      fileExtension: ".mp3",
      voiceCompatible: false,
    };
  },
});

注意:

  • 将 TTS 策略、回退和回复下发保留在核心中。
  • 使用语音提供商来处理厂商自有的合成行为。
  • 旧的 Microsoft edge 输入会被规范化为 microsoft 提供商 ID。
  • 首选的归属模型以公司为导向:随着 OpenClaw 增加这些能力契约,一个厂商插件可以同时拥有文本、语音、图像以及未来的媒体提供商。

对于图像/音频/视频理解,插件应注册一个类型化的媒体理解提供商,而不是使用通用的键/值集合:

api.registerMediaUnderstandingProvider({
  id: "google",
  capabilities: ["image", "audio", "video"],
  describeImage: async (req) => ({ text: "..." }),
  transcribeAudio: async (req) => ({ text: "..." }),
  describeVideo: async (req) => ({ text: "..." }),
});

注意:

  • 将编排、回退、配置和渠道接线保留在核心中。
  • 将厂商行为保留在提供商插件中。
  • 增量扩展应保持类型化:新增可选方法、新的可选结果字段、新的可选能力。
  • 视频生成已遵循相同的模式:
  • 核心拥有能力契约和运行时辅助工具
  • 厂商插件注册 api.registerVideoGenerationProvider(...)
  • 功能/渠道插件消费 api.runtime.videoGeneration.*

对于媒体理解运行时辅助工具,插件可以调用:

const image = await api.runtime.mediaUnderstanding.describeImageFile({
  filePath: "/tmp/inbound-photo.jpg",
  cfg: api.config,
  agentDir: "/tmp/agent",
});

const video = await api.runtime.mediaUnderstanding.describeVideoFile({
  filePath: "/tmp/inbound-video.mp4",
  cfg: api.config,
});

// receiptImageBuffer is your own image bytes, not an SDK-provided value.
const extraction = await api.runtime.mediaUnderstanding.extractStructuredWithModel({
  provider: "codex",
  model: "gpt-6-astra",
  input: [
    {
      type: "image",
      buffer: receiptImageBuffer,
      fileName: "receipt.png",
      mime: "image/png",
    },
    { type: "text", text: "Use the printed fields as the source of truth." },
  ],
  instructions: "Return entities and searchable tags.",
  schemaName: "example.evidence",
  jsonSchema: {
    type: "object",
    properties: {
      entities: { type: "array", items: { type: "string" } },
      tags: { type: "array", items: { type: "string" } },
    },
  },
  cfg: api.config,
});

对于音频转写,插件可以使用媒体理解运行时,也可以使用旧的 STT 别名:

const { text } = await api.runtime.mediaUnderstanding.transcribeAudioFile({
  filePath: "/tmp/inbound-audio.ogg",
  cfg: api.config,
  // Optional when MIME cannot be inferred reliably:
  mime: "audio/ogg",
});

注意:

  • api.runtime.mediaUnderstanding.* 是用于图像/音频/视频理解的首选共享接口。
  • extractStructuredWithModel(...) 是面向插件的接缝,用于有界、提供商自有的图像优先提取。至少包含一个图像输入;文本输入作为补充上下文。产品插件拥有自己的路由和模式,而 OpenClaw 拥有提供商/运行时边界。
  • 使用核心媒体理解音频配置(tools.media.audio)和提供商回退顺序。
  • 当未产生任何转写输出时(例如跳过/不支持的输入),返回 { text: undefined }。

插件还可以通过 api.runtime.subagent 启动后台子代理运行:

const result = await api.runtime.subagent.run({
  sessionKey: "agent:main:subagent:search-helper",
  message: "Expand this query into focused follow-up searches.",
  toolsAlsoAllow: ["my_plugin_progress"],
  provider: "openai",
  model: "gpt-4.1-mini",
  deliver: false,
});

注意:

  • provider 和 model 是每次运行的可选覆盖项,不是持久的会话变更。
  • toolsAlsoAllow 接受调用插件注册的精确且唯一拥有的工具名称。核心名称和歧义名称会被拒绝。它是对正常配置文件的补充,但操作员的允许列表和拒绝列表仍然具有最终决定权。
  • OpenClaw 仅对受信任的调用者兑现这些覆盖字段。
  • 对于插件自有的回退运行,操作员必须通过 plugins.entries.<id>.subagent.allowModelOverride: true 主动选择启用。
  • 使用 plugins.entries.<id>.subagent.allowedModels 将受信任插件限制为特定的规范 provider/model 目标,或使用 "*" 显式允许任何目标。
  • 不受信任的插件子代理运行仍然有效,但覆盖请求会被拒绝,而不是静默回退。
  • 插件创建的子代理会话会标记创建该会话的插件 ID。回退的 api.runtime.subagent.deleteSession(...) 只能删除这些拥有的会话;任意会话删除仍需要管理员范围的 Gateway 请求。

对于 Web 搜索,插件可以使用共享的运行时助手(runtime helper),而无需直接深入 agent 工具的内部接线:

const providers = api.runtime.webSearch.listProviders({
  config: api.config,
});

const result = await api.runtime.webSearch.search({
  config: api.config,
  args: {
    query: "OpenClaw plugin runtime helpers",
    count: 5,
  },
});

插件还可以通过 api.registerWebSearchProvider(...) 注册 Web 搜索提供商。

说明:

  • 将提供商的选取、凭据解析和共享请求语义保留在核心(core)中。
  • 使用 Web 搜索提供商来处理厂商特定的搜索传输。
  • api.runtime.webSearch.* 是功能/渠道插件在需要搜索行为时的首选共享接口,无需依赖 agent 工具包装器。

api.runtime.imageGeneration

const result = await api.runtime.imageGeneration.generate({
  config: api.config,
  args: { prompt: "A friendly lobster mascot", size: "1024x1024" },
});

const providers = api.runtime.imageGeneration.listProviders({
  config: api.config,
});
  • generate(...):使用已配置的图像生成提供商链路生成图像。
  • listProviders(...):列出可用的图像生成提供商及其能力。

本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw