跳转至

媒体辅助

语音、媒体理解、生成、网络搜索以及底层媒体工具。是插件运行时辅助参考的一部分。

FFmpeg 命令发现

官方插件源码使用来自私有 openclaw/plugin-sdk/media-ffmpeg 运行时的 resolveFfmpegBin。它会解析与媒体运行时相同的受信任系统路径,并在 FFmpeg 不可用时抛出安装提示。这个窄入口点可避免在源音频 worker 启动时加载媒体生成和 agent 运行时。它只有 JavaScript 主机导出;其声明被排除在包之外。第三方插件保留来自 openclaw/plugin-sdk/media-runtime 的现有 resolveFfmpegBin 导出。

官方插件发布构建器会为这个私有导入输出 media-runtime,包括 worker 条目,因此已发布的插件在早于 media-ffmpeg 的受支持主机上仍可正常工作。两条路径都使用主机的 FFmpeg 解析器。

实时语音播放

官方音频 worker 使用来自私有 openclaw/plugin-sdk/realtime-voice-playback 运行时的 createRealtimeVoiceOutputActivityTracker 和 isRealtimeVoiceAudioAudible,以避免从源码加载语音会话和 agent 运行时。发布构建器会为这些辅助函数输出现有的 openclaw/plugin-sdk/realtime-voice 主机导入,包括 worker 条目,以兼容受支持的主机。源门面具有仅 JavaScript 的主机导出;其声明被排除在包之外。

媒体与生成命名空间

api.runtime.tts

文本转语音合成。

// Standard TTS
const clip = await api.runtime.tts.textToSpeech({
  text: "Hello from OpenClaw",
  cfg: api.config,
});

// Telephony-optimized TTS
const telephonyClip = await api.runtime.tts.textToSpeechTelephony({
  text: "Hello from OpenClaw",
  cfg: api.config,
});

// List available voices
const voices = await api.runtime.tts.listVoices({
  provider: "elevenlabs",
  cfg: api.config,
});

使用核心 tts 配置和提供商选择。返回 PCM 音频缓冲区 + 采样率。也可使用 textToSpeechStream 进行流式合成。

api.runtime.mediaUnderstanding

图像、音频和视频分析。

// Describe an image
const image = await api.runtime.mediaUnderstanding.describeImageFile({
  filePath: "/tmp/inbound-photo.jpg",
  cfg: api.config,
  agentDir: "/tmp/agent",
});

// Prepare a capture limit before installing audio receive listeners.
const budget = await api.runtime.mediaUnderstanding.resolveAudioInputBudget({
  cfg: api.config,
});
// budget.enabled is false when audio understanding is disabled. Otherwise,
// budget.maxBytes includes the container header and covers the largest fallback.

// Transcribe audio
const { text } = await api.runtime.mediaUnderstanding.transcribeAudioFile({
  filePath: "/tmp/inbound-audio.ogg",
  cfg: api.config,
  mime: "audio/ogg", // optional, for when MIME cannot be inferred
});

// Describe a video
const video = await api.runtime.mediaUnderstanding.describeVideoFile({
  filePath: "/tmp/inbound-video.mp4",
  cfg: api.config,
});

// Generic file analysis
const result = await api.runtime.mediaUnderstanding.runFile({
  filePath: "/tmp/inbound-file.pdf",
  cfg: api.config,
});

// Structured image extraction through a specific provider/model.
// Include at least one image; text inputs are supplemental context.
// receiptImageBuffer is your own image bytes, not an SDK-provided value.
const evidence = await api.runtime.mediaUnderstanding.extractStructuredWithModel({
  provider: "codex",
  model: "gpt-6-astra",
  input: [
    {
      type: "image",
      buffer: receiptImageBuffer,
      fileName: "receipt.png",
      mime: "image/png",
    },
    { type: "text", text: "Prefer the printed total over handwritten notes." },
  ],
  instructions: "Extract vendor, total, and searchable tags.",
  schemaName: "receipt.evidence",
  jsonSchema: {
    type: "object",
    properties: {
      vendor: { type: "string" },
      total: { type: "number" },
      tags: { type: "array", items: { type: "string" } },
    },
    required: ["vendor", "total"],
  },
  cfg: api.config,
});

当没有产生输出时(例如跳过输入),返回 { text: undefined }。

describeImageFileWithModel(...) 通过特定提供商/模型描述已知图像,绕过 describeImageFile(...) 使用的默认活动模型解析。

api.runtime.imageGeneration

图像生成。

const result = await api.runtime.imageGeneration.generate({
  prompt: "A robot painting a sunset",
  cfg: api.config,
});

const providers = api.runtime.imageGeneration.listProviders({ cfg: api.config });
api.runtime.videoGeneration

视频生成,其结构与图像生成一致。

const result = await api.runtime.videoGeneration.generate({
  prompt: "A drone shot flying over a coastline at sunrise",
  cfg: api.config,
});

const providers = api.runtime.videoGeneration.listProviders({ cfg: api.config });
api.runtime.musicGeneration

音乐生成,其结构与图像生成一致。

const result = await api.runtime.musicGeneration.generate({
  prompt: "An upbeat lo-fi track for a coding session",
  cfg: api.config,
});

const providers = api.runtime.musicGeneration.listProviders({ cfg: api.config });
api.runtime.webSearch

网络搜索。

const providers = api.runtime.webSearch.listProviders({ config: api.config });

const result = await api.runtime.webSearch.search({
  config: api.config,
  args: { query: "OpenClaw plugin SDK", count: 5 },
});

搜索调用方可以随 signal 提供一个同步的 assertCurrent 回调, 以便在 provider 准备期间保留其权限。受保护的 HTTP 请求 会在传输准备之后、每次请求或重定向之前检查它。 使用其他传输的已注册搜索 provider 必须在每次外部副作用之前、 在已等待的准备完成之后调用执行上下文的 assertCurrent。 该回调属于调用方,并在搜索结束时过期;provider 不得替换它, 也不得将其缺失视为许可。

api.runtime.media

低级媒体工具。

const webMedia = await api.runtime.media.loadWebMedia(url);
const mime = await api.runtime.media.detectMime({ buffer });
const kind = api.runtime.media.mediaKindFromMime("image/jpeg"); // "image"
const isVoice = api.runtime.media.isVoiceCompatibleAudio({ fileName: filePath });
const metadata = await api.runtime.media.getImageMetadata(buffer);
const resized = await api.runtime.media.resizeToJpeg({ buffer, maxSide: 800, quality: 85 });

QR 辅助函数由 openclaw/plugin-sdk/media-runtime 导出:

import { resolvePreferredOpenClawTmpDir } from "openclaw/plugin-sdk/temp-path";

const qr = await import("openclaw/plugin-sdk/media-runtime");
const terminalQr = await qr.renderQrTerminal("https://openclaw.ai");
const pngQr = await qr.renderQrPngBase64("https://openclaw.ai", {
  scale: 6, // 1-12
  marginModules: 4, // 0-16
});
const pngQrDataUrl = await qr.renderQrPngDataUrl("https://openclaw.ai");
const tmpRoot = resolvePreferredOpenClawTmpDir();
const pngQrFile = await qr.writeQrPngTempFile("https://openclaw.ai", {
  tmpRoot,
  dirPrefix: "my-plugin-qr-",
  fileName: "qr.png",
});

本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw