跳转至

推理 CLI

openclaw infer 是用于提供方支持的推理的规范无头接口。它暴露能力族(model、image、audio、tts、video、web、embedding),而不是原始网关 RPC 名称或代理工具 ID。openclaw capability ... 是同一命令树的别名。

相比一次性提供方封装,优先使用它的原因:

  • 复用 OpenClaw 中已配置的提供方和模型。
  • 为脚本和代理驱动的自动化提供稳定的 --json 封装(参见 JSON 输出)。
  • 对于大多数子命令,无需网关即可运行正常的本地路径。
  • 对于端到端提供方检查,它会在提供方请求发出之前,验证随附的 CLI、配置加载、默认代理解析、捆绑插件激活以及共享能力运行时。

命令树

 openclaw infer
  list
  inspect

  model
    run
    list
    inspect
    providers
    auth login
    auth logout
    auth status

  image
    generate
    edit
    describe
    describe-many
    providers

  audio
    transcribe
    providers

  tts
    convert
    voices
    providers
    personas
    status
    enable
    disable
    set-provider
    set-persona

  video
    generate
    describe
    providers

  web
    search
    fetch
    providers

  embedding
    create
    providers

infer list / infer inspect --name <capability> 以数据形式显示此树(能力 ID、传输方式、描述)。

父命令和子命令帮助在不加载提供方执行运行时的情况下,暴露完整的推理命令树。这些命令定义还提供推理 shell 补全元数据。

常见任务

任务 命令 说明
运行文本/模型 Prompt openclaw infer model run --prompt "..." --json 默认本地
对图像运行模型 Prompt openclaw infer model run --prompt "Describe this" --file ./image.png --model provider/model 对多张图像重复 --file
生成图像 openclaw infer image generate --prompt "..." --json 从现有文件开始时使用 image edit
描述图像文件或 URL openclaw infer image describe --file ./image.png --prompt "..." --json --model 必须是支持图像的 <provider/model>
转录音频 openclaw infer audio transcribe --file ./memo.m4a --json --model 必须是 <provider/model>
合成语音 openclaw infer tts convert --text "..." --output ./speech.mp3 --json tts status 仅通过网关运行
生成视频 openclaw infer video generate --prompt "..." --json 支持提供方提示,例如 --resolution
描述视频文件 openclaw infer video describe --file ./clip.mp4 --json --model 必须是 <provider/model>
搜索网络 openclaw infer web search --query "..." --json
获取网页 openclaw infer web fetch --url https://example.com --json
创建嵌入 openclaw infer embedding create --text "..." --json

行为

  • 当输出要传递给另一个命令或脚本时,使用 --json;否则使用文本输出。
  • 使用 --provider 或 --model provider/model 固定特定后端。
  • 对 --limit、--count、--duration 和 --timeout-ms 显式设置为空值或仅空白值会报错。省略该标志可保留其默认行为。
  • image edit 和 image describe-many 至少需要一个 --file;embedding create 至少需要一个 --text。对多个输入重复该标志。省略它属于用法错误,而不是空的成功结果,并且不会发送推理请求。
  • 使用 model run --thinking <level> 进行一次性思考/推理覆盖:off、minimal、low、medium、high、adaptive、xhigh 或 max。
  • 对于 image describe、audio transcribe 和 video describe,--model 必须使用 <provider/model> 形式。
  • 对于 image describe,--file 接受本地路径和 HTTP(S) URL;远程 URL 会经过正常的媒体获取 SSRF 策略。
  • 无状态执行命令(model run、image *、audio *、video *、web *、embedding *)默认本地。Gateway 管理的状态命令(tts status)默认 Gateway。
  • 本地路径从不要求网关正在运行。
  • 对于提供方清单命令,如果其 configured 状态可能来自已保存的代理身份验证,则接受 --agent <id>。如果没有它,它们会使用 agents.defaults.systemAgent.agentId 或唯一已配置的代理;没有系统所有者的显式多代理集群必须传递 --agent。提供方目录仍保持聚合;--agent 限定已保存身份验证和按代理选择的事实。Gateway 拥有的 TTS 提供方状态仍保持 Gateway 全局,因此 tts providers --gateway 不接受 --agent。
  • 解析代理拥有的模型或身份验证状态的命令(model run、image generate、image edit、image describe、image describe-many、audio transcribe、video generate、video describe、embedding create 和 model auth login/logout/status)也接受 --agent <id>。它们首先解析显式 ID,然后是 agents.defaults.systemAgent.agentId,然后是唯一已配置的代理。
  • 生成的图像和视频 --output 文件会先暂存到目标旁边,并在完整缓冲区写入后才替换目标;写入失败时,现有目标保持不变。
  • 本地 model run 是精简的一次性提供方补全:它会解析已配置的代理模型和身份验证,但不会启动聊天代理回合、加载工具或打开捆绑的 MCP 服务器。
  • model run --file 将图像文件(自动检测 MIME 类型)附加到 Prompt;对多张图像重复 --file。非图像文件会被拒绝——请改用 infer audio transcribe 或 infer video describe。
  • model run --gateway 会验证 Gateway 路由、已保存身份验证、提供方选择和嵌入式运行时,但仍保持原始模型探测:没有先前会话转录、bootstrap/AGENTS 上下文、工具或捆绑的 MCP 服务器。
  • model run --gateway --model <provider/model> 需要可信操作员的网关凭据,因为它要求 Gateway 运行一次性的提供方/模型覆盖。

模型

文本推理以及模型/提供商检查。

model list、model inspect 和 model providers 读取所选代理的模型目录。它们保留已配置的模型详情,并排除该目录之外的模型。使用 infer model --agent <id> list --json 或 infer model --agent <id> inspect --model <provider/model> --json 选择代理。提供商计数使用相同的清单。

openclaw infer model run --prompt "Reply with exactly: smoke-ok" --json
openclaw infer model run --prompt "Summarize this changelog entry" --model openai/gpt-5.4 --json
openclaw infer model run --prompt "Describe this image in one sentence" --file ./photo.jpg --model google/gemini-2.5-flash --json
openclaw infer model run --prompt "Use more reasoning here" --thinking high --json
openclaw infer model providers --agent <id> --json
openclaw infer model inspect --model gpt-6-astra --json

使用完整的 <provider/model> 引用配合 --local,可在不启动 Gateway 或加载代理工具界面的情况下,对单个提供商进行冒烟测试:

openclaw infer model run --local --model anthropic/claude-sonnet-4-6 --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model cerebras/zai-glm-4.7 --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model google/gemini-2.5-flash --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model groq/llama-3.1-8b-instant --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model llmman/qwen3.8 --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model llmman/gemma4:e4b --prompt "Describe this image." --file ./photo.jpg --json
openclaw infer model run --local --model mistral/mistral-medium-3-5 --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model mistral/mistral-small-latest --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model openai/gpt-5.6-luna --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model ollama/qwen2.5vl:7b --prompt "Describe this image." --file ./photo.jpg --json

说明:

  • 本地 model run 是用于提供商/模型/身份验证健康状况的最小范围 CLI 冒烟测试:对于非 ChatGPT-Codex 提供商,它仅发送所提供的提示。
  • 本地 model run --model <provider/model> 可以在该提供商写入配置之前,解析精确的捆绑静态目录条目(与 openclaw models list --all 显示的条目相同)。仍然需要提供商身份验证;缺少凭据会作为身份验证错误失败,而不是 Unknown model。
  • 对于 Mistral Medium 3.5 推理探测,请保持 temperature 未设置/默认。Mistral 会拒绝 reasoning_effort="high" 与 temperature: 0 的组合;请使用默认 temperature 或非零值,例如 0.7。
  • OpenAI ChatGPT/Codex OAuth(openai-chatgpt-responses API)本地探测会添加一条最小系统指令,以便传输层填充其必需的 instructions 字段——不包含完整的代理上下文、工具、内存或会话转录。
  • model run --file 会将图像内容直接附加到单条用户消息。当 MIME 类型被检测为 image/* 时,常见格式(PNG、JPEG、WebP)可用;不支持或无法识别的文件会在调用提供商之前失败。如果你希望使用 OpenClaw 的图像模型路由和回退,而不是直接探测多模态模型,请改用 infer image describe。
  • 所选模型必须支持图像输入;仅文本模型可能在提供商层拒绝请求。
  • model run --prompt 必须包含非空白文本;空提示会在任何提供商或 Gateway 调用之前被拒绝。
  • 当提供商未返回文本输出时,本地 model run 会以非零状态退出,因此无法访问的提供商和空补全不会表现为成功的探测。
  • 使用 model run --gateway 可在保持模型输入原始状态的同时,测试 Gateway 路由或代理运行时配置。如需完整的代理上下文、工具、内存和会话转录,请使用 openclaw agent 或聊天界面。
  • --thinking adaptive 映射到补全运行时级别 medium;对于支持原生 max 努力的 OpenAI 模型,--thinking max 映射到 max,否则映射到 xhigh。
  • model auth login、model auth logout 和 model auth status 管理已保存的提供商身份验证状态。

图像

生成、编辑和描述。

openclaw infer image generate --prompt "friendly lobster illustration" --json
openclaw infer image generate --prompt "cinematic product photo of headphones" --json
openclaw infer image generate --model openai/gpt-image-1.5 --output-format png --background transparent --prompt "simple red circle sticker on a transparent background" --json
openclaw infer image generate --model openai/gpt-image-2 --quality low --openai-moderation low --prompt "low-cost draft poster" --json
openclaw infer image generate --prompt "slow image backend" --timeout-ms 180000 --json
openclaw infer image edit --file ./logo.png --model openai/gpt-image-1.5 --output-format png --background transparent --prompt "keep the logo, remove the background" --json
openclaw infer image edit --file ./poster.png --prompt "make this a vertical story ad" --size 2160x3840 --aspect-ratio 9:16 --resolution 4K --json
openclaw infer image describe --file ./photo.jpg --json
openclaw infer image describe --file https://example.com/photo.png --json
openclaw infer image describe --file ./receipt.jpg --prompt "Extract the merchant, date, and total" --json
openclaw infer image describe-many --file ./before.png --file ./after.png --prompt "Compare the screenshots and list visible UI changes" --json
openclaw infer image describe --file ./ui-screenshot.png --model openai/gpt-5.4-mini --json
openclaw infer image describe --file ./photo.jpg --model ollama/qwen2.5vl:7b --prompt "Describe the image in one sentence" --timeout-ms 300000 --json

说明:

  • 当从现有输入文件开始时,请使用 image edit;对于支持这些选项的提供商/模型,--size、--aspect-ratio 或 --resolution 会添加几何提示。
  • 将 --output-format png --background transparent 与 --model openai/gpt-image-1.5 一起使用,可获得透明背景的 OpenAI PNG 输出;--openai-background 是同一提示的 OpenAI 专用别名。未声明支持背景的提供商会将其报告为被忽略的覆盖项(参见 JSON 信封 中的 ignoredOverrides)。
  • --quality low|medium|high|auto 适用于支持图像质量提示的提供商,包括 OpenAI。OpenAI 还接受 --openai-moderation low|auto。
  • image providers --json 列出哪些捆绑图像提供商可发现、已配置、已选中,以及每个提供商暴露哪些生成/编辑能力。
  • image generate --model <provider/model> --json 是用于图像生成变更的最小范围实时冒烟测试:
  openclaw infer image providers --json
  openclaw infer image generate \
    --model google/gemini-3.1-flash-image \
    --prompt "Minimal flat test image: one blue square on a white background, no text." \
    --output ./openclaw-infer-image-smoke.png \
    --json
  ```

  响应会报告 `ok`、`provider`、`model`、`attempts` 以及已写入的输出路径。当设置 `--output` 时,最终扩展名可能会遵循提供方返回的 MIME 类型。

- 对于 `image describe` 和 `image describe-many`,使用 `--prompt` 提供任务特定指令(OCR、比较、UI 检查、简洁字幕)。
- 对于较慢的本地视觉模型或 Ollama 冷启动,使用 `--timeout-ms`。
- 对于 `image describe`,显式指定的 `--model`(必须是支持图像的 `<provider/model>`)会先运行,如果该调用失败,则尝试已配置的 `agents.defaults.imageModel.fallbacks`。输入准备错误(文件缺失、不支持的 URL)会在任何回退尝试之前失败,并且模型必须在模型目录或提供方配置中支持图像。
- 对于本地 Ollama 视觉模型,请先拉取模型,并将 `OLLAMA_API_KEY` 设置为任意占位值,例如 `ollama-local`。参见 [Ollama](../providers/ollama.md#vision-and-image-description)。
- 对于 llmman 视觉模型(例如 `llmman/gemma4:e4b`),请在模型条目上为 `llmman` 提供方配置 `input: ["text", "image"]`,并将 `LLMMAN_API_KEY` 设置为占位值,例如 `llmman-local`。参见 [llmman](../providers/llmman.md#vision-and-image-description)。

## 音频 {#audio}

文件转录(不是实时会话管理)。

```bash
openclaw infer audio transcribe --file ./memo.m4a --json
openclaw infer audio transcribe --agent <id> --file ./memo.m4a --json
openclaw infer audio transcribe --file ./team-sync.m4a --language en --prompt "Focus on names and action items" --json
openclaw infer audio transcribe --file ./memo.m4a --model openai/whisper-1 --json

--model 必须是 <provider/model>。

对于基于 CLI 的转录,结果中的 provider 标识工具族,model 报告执行的命令。自动检测的工具会报告其解析后的可执行文件路径;显式 CLI 条目会保留其编写的命令值。此字段不会标识工具内部加载的语音模型。

TTS

语音合成以及 TTS 提供方/角色状态。

openclaw infer tts convert --text "hello from openclaw" --output ./hello.mp3 --json
openclaw infer tts convert --text "Your build is complete" --output ./build-complete.mp3 --json
openclaw infer tts convert --provider xiaomi --text "Provider-only selection" --output ./xiaomi.mp3 --json
openclaw infer tts providers --json
openclaw infer tts personas --json
openclaw infer tts status --json

说明:

  • tts status 仅支持 --gateway(它反映由网关管理的 TTS 状态)。
  • 本地和 loopback-Gateway 的 tts convert --output 会在目标文件旁暂存副本,并仅在成功后替换它;复制失败时,现有文件保持不变。
  • Remote-Gateway 的 tts convert --output 会在请求语音合成之前被拒绝。
  • 当选择提供方但不覆盖其模型时,使用 tts convert --provider <id>。
  • 使用 tts providers、tts voices、tts personas、tts set-provider 和 tts set-persona 来检查和配置 TTS 行为。

视频

生成和描述。

openclaw infer video generate --prompt "cinematic sunset over the ocean" --json
openclaw infer video generate --prompt "slow drone shot over a forest lake" --resolution 768P --duration 6 --json
openclaw infer video describe --file ./clip.mp4 --json
openclaw infer video describe --agent <id> --file ./clip.mp4 --json
openclaw infer video describe --file ./clip.mp4 --model openai/gpt-5.4-mini --json

说明:

  • video generate 接受 --size、--aspect-ratio、--resolution、--duration、--audio、--watermark 和 --timeout-ms,并转发到视频生成运行时。
  • 提供方托管的视频下载会拒绝空响应、文本响应和 JSON 响应,而不是将不可用的文件报告为成功输出。
  • 使用 --output 时,基于 URL 的视频会流式传输到同级临时文件,并仅在完整且非空的下载成功后替换目标文件;流式传输失败时,现有目标文件保持不变。
  • 对于 video describe,--model 必须是 <provider/model>。

Web

搜索和抓取。

openclaw infer web search --query "OpenClaw docs" --json
openclaw infer web search --query "OpenClaw infer web providers" --json
openclaw infer web fetch --url https://docs.openclaw.ai/cli/infer --json
openclaw infer web providers --agent <id> --json

web providers 列出搜索和抓取可用的、已配置的以及已选择的提供方。

嵌入

向量创建和嵌入提供方检查。

openclaw infer embedding create --text "friendly lobster" --json
openclaw infer embedding create --text "customer support ticket: delayed shipment" --model openai/text-embedding-3-large --json
openclaw infer embedding providers --agent <id> --json

JSON 输出

Infer 命令会在共享信封下规范化 JSON 输出:

{
  "ok": true,
  "capability": "image.generate",
  "transport": "local",
  "provider": "openai",
  "model": "gpt-image-2",
  "attempts": [],
  "outputs": []
}

稳定的顶层字段:

  • ok
  • capability
  • transport
  • provider
  • model
  • attempts
  • inputs(适用时,随请求发送的图像附件)
  • outputs
  • ignoredOverrides(适用时,提供方不支持的提示键)
  • error

对于生成的媒体命令,outputs 包含由 OpenClaw 写入的文件。在自动化中,请使用该数组中的 path、mimeType、size 以及任何媒体特定尺寸,而不是解析人类可读的 stdout。

常见陷阱

# Bad
openclaw infer media image generate --prompt "friendly lobster"

# Good
openclaw infer image generate --prompt "friendly lobster"
# Bad
openclaw infer audio transcribe --file ./memo.m4a --model whisper-1 --json

# Good
openclaw infer audio transcribe --file ./memo.m4a --model openai/whisper-1 --json

将 infer 转化为技能

将此内容复制并粘贴给智能体:

Read https://docs.openclaw.ai/cli/infer, then create a skill that routes my common workflows to `openclaw infer`.
Focus on model runs, image generation, video generation, audio transcription, TTS, web search, and embeddings.

一个好的基于 infer 的技能会将常见用户意图映射到正确的子命令,为每个工作流包含若干标准示例,优先使用 openclaw infer ... 而非更低层的替代命令,并且不会在技能正文中重新记录整个 infer 功能面。

本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw