跳转至

Prometheus

OpenClaw 可以通过官方 diagnostics-prometheus 插件暴露诊断指标。它会监听受信任的诊断事件,以及内部标记的、由 dispatcher 拥有的诊断事件(队列、内存和会话恢复信号),并在以下位置提供 Prometheus 文本端点:

GET /api/diagnostics/prometheus

内容类型为 text/plain; version=0.0.4; charset=utf-8,即标准的 Prometheus 暴露格式。

Warning

该路由使用 Gateway 身份验证(operator 作用域,受信任 operator 接口面),并要求调用方的有效作用域包含 operator.read(由 operator.write 或 operator.admin 隐含)。不要将其暴露为公开的未认证 /metrics 端点。请通过你用于其他 operator API 的同一身份验证路径对其进行抓取。

有关跟踪、日志、OTLP 推送以及 OpenTelemetry GenAI 语义属性,请参阅 OpenTelemetry 导出。

快速开始

1. 安装插件

openclaw plugins install clawhub:@openclaw/diagnostics-prometheus

2. 启用插件

{
  plugins: {
    allow: ["diagnostics-prometheus"],
    entries: {
      "diagnostics-prometheus": { enabled: true },
    },
  },
  diagnostics: {
    enabled: true,
  },
}
openclaw plugins enable diagnostics-prometheus

3. 重启 Gateway

HTTP 路由在插件启动时注册,因此启用后需要重新加载。

openclaw gateway restart

4. 抓取受保护的路由

发送你的 operator 客户端使用的相同 Gateway 身份验证:

curl -H "Authorization: Bearer $OPENCLAW_GATEWAY_TOKEN" \
  http://127.0.0.1:18789/api/diagnostics/prometheus

5. 接入 Prometheus

# prometheus.yml
scrape_configs:
  - job_name: openclaw
    scrape_interval: 30s
    metrics_path: /api/diagnostics/prometheus
    authorization:
      credentials_file: /etc/prometheus/openclaw-gateway-token
    static_configs:
      - targets: ["openclaw-gateway:18789"]

Note

diagnostics.enabled 默认为 true;仅在严格受限的环境中将其设置为 false。如果导出器启动时该值为 false,插件仍会注册 HTTP 路由,但不会记录任何诊断事件或运行时身份,因此响应为空。

导出的指标

指标 类型 标签
openclaw_gateway_build_info gauge process_instance_id,可选 build_id
openclaw_gc_duration_seconds histogram 无
openclaw_gateway_rpc_requests_total counter method
openclaw_gateway_rpc_first_response_seconds histogram method
openclaw_gateway_rpc_handler_seconds histogram method
openclaw_gateway_rpc_admission_seconds histogram method
openclaw_gateway_rpc_queue_wait_seconds histogram method
openclaw_gateway_rpc_stage_seconds histogram method、phase
openclaw_gateway_rpc_stage_thread_cpu_seconds histogram method、phase
openclaw_gateway_rpc_outcomes_total counter phase、outcome
openclaw_run_completed_total counter channel、model、outcome、provider、trigger
openclaw_run_duration_seconds histogram channel、model、outcome、provider、trigger
openclaw_model_call_total counter api、error_category、model、observation_unit、outcome、provider、transport
openclaw_model_call_duration_seconds histogram api、error_category、model、observation_unit、outcome、provider、transport
openclaw_model_failover_total counter from_model、from_provider、lane、reason、suspended、to_model、to_provider
openclaw_model_tokens_total counter agent、channel、model、provider、token_type
openclaw_gen_ai_client_token_usage histogram model、provider、token_type
openclaw_model_cost_usd_total counter agent、channel、model、provider
openclaw_model_usage_duration_seconds histogram agent、channel、model、provider
openclaw_skill_used_total counter activation、agent、skill、source
openclaw_tool_execution_total counter error_category、outcome、params_kind、tool、tool_owner、tool_source
指标 类型 标签
openclaw_tool_execution_duration_seconds histogram error_category, outcome, params_kind, tool, tool_owner, tool_source
openclaw_tool_execution_blocked_total counter denied_reason, params_kind, tool, tool_owner, tool_source
openclaw_harness_run_total counter channel, error_category, harness, model, outcome, phase, plugin, provider
openclaw_harness_run_duration_seconds histogram channel, error_category, harness, model, outcome, phase, plugin, provider
openclaw_webhook_received_total counter channel, webhook
openclaw_webhook_error_total counter channel, webhook
openclaw_webhook_duration_seconds histogram channel, webhook
openclaw_message_received_total counter channel, source
openclaw_message_dispatch_started_total counter channel, source
openclaw_message_dispatch_completed_total counter channel, outcome, reason, source
openclaw_message_dispatch_duration_seconds histogram channel, outcome, reason, source
openclaw_message_processed_total counter channel, outcome, reason
openclaw_message_processed_duration_seconds histogram channel, outcome, reason
openclaw_message_delivery_started_total counter channel, delivery_kind
openclaw_message_delivery_total counter channel, delivery_kind, error_category, outcome
openclaw_message_delivery_duration_seconds histogram channel, delivery_kind, error_category, outcome
openclaw_talk_event_total counter brain, event_type, mode, provider, transport
openclaw_talk_event_duration_seconds histogram brain, event_type, mode, provider, transport
openclaw_talk_audio_bytes histogram brain, event_type, mode, provider, transport
openclaw_queue_lane_size gauge lane
openclaw_queue_lane_wait_seconds histogram lane
openclaw_session_state_total counter reason, state
openclaw_session_queue_depth gauge state
openclaw_session_turn_created_total counter agent, channel, trigger
openclaw_session_stuck_total counter reason, state
openclaw_session_stuck_age_seconds histogram reason, state
openclaw_session_recovery_total counter action, active_work_kind, state, status
openclaw_session_recovery_age_seconds histogram action, active_work_kind, state, status
openclaw_gateway_event_loop_delay_max_seconds histogram 无
openclaw_gateway_event_loop_observed_seconds_total counter 无
openclaw_liveness_warning_total counter reason
openclaw_liveness_sessions gauge state
openclaw_liveness_event_loop_delay_p99_seconds histogram reason
openclaw_liveness_event_loop_delay_max_seconds histogram reason
openclaw_liveness_event_loop_utilization_ratio histogram reason
openclaw_liveness_cpu_core_ratio histogram reason
指标 类型 标签
openclaw_payload_large_total 计数器 action, channel, plugin, reason, surface
openclaw_payload_large_bytes 直方图 action, channel, plugin, reason, surface
openclaw_memory_bytes 仪表 kind
openclaw_worker_count 仪表 无
openclaw_worker_heap_sampled_count 仪表 无
openclaw_worker_heap_used_bytes 仪表 script
openclaw_worker_started_total 计数器 script
openclaw_worker_retired_total 计数器 script, reason
openclaw_child_process_spawn_total 计数器 family
openclaw_memory_rss_bytes 直方图 无
openclaw_memory_pressure_total 计数器 level, reason
openclaw_telemetry_exporter_total 计数器 exporter, reason, signal, status
openclaw_prometheus_series_dropped_total 计数器 无
openclaw_diagnostic_async_queue_dropped_total 计数器 drop_class
openclaw_diagnostic_async_queue_length 仪表 无

对于模型调用指标,observation_unit="request" 衡量一次可观察的提供商请求。observation_unit="turn" 衡量一个合成的 Claude Code 或 Codex CLI 智能体回合,其中可能包含多个隐藏的提供商请求。比较延迟时,请将这些序列分开。

网关 RPC 指标涵盖有效的已认证 WebSocket 请求,包括后续拒绝。first_response 衡量从接收到发送方接受的第一帧为止;不可用或被抑制的发送没有持续时间样本。handler 衡量实际处理程序调用直至返回或抛出,admission 衡量从接收到该调用为止。queue_wait 仅衡量操作者请求启动队列等待,与命令/会话通道指标分开。它们衡量的是经过时间,而不是 CPU 时间。早期确认以及处理程序返回后的响应,与已完成的智能体工作不同。参见 网关 RPC 计时语义。

接收在已连接客户端的请求帧通过验证后开始。这些计时不包括 CLI 启动、本地诊断、连接/认证设置以及请求分发前的事件循环延迟。直方图记录已完成的观测:未完成的处理程序尚无处理程序持续时间样本。调查超时时,请比较请求数、已完成的计时和事件循环观测;仅低处理程序延迟不足以证明客户端路径响应迅速。

RPC 方法标签包含规范的核心方法名称、插件方法的 other,或 unknown。结果总计按阶段和结果聚合,没有方法维度。每个具有全部四种计时的方法在共享的 2,048 个样本上限中占用五个聚合样本。持续时间直方图占用一个样本,但会扩展为 19 个抓取序列(桶、总和和计数)。当上限填满时,现有样本会继续更新;未见的 RPC 或其他操作样本会被拒绝,并递增 openclaw_prometheus_series_dropped_total。监控该计数器:覆盖每个核心方法都可能填满上限,因此在解释总计或延迟百分位数时,零值很重要。异步诊断队列饱和也可能丢弃观测,由 openclaw_diagnostic_async_queue_dropped_total 报告。

目录列表阶段

阶段直方图目前仅覆盖 sessions.catalog.list。它们使用六个固定的阶段标签:

阶段 观测
projection_initial 请求必须等待时的初始共享投影就绪状态
planning 领导者提供商规划快照的同步创建和冻结
provider 领导者枚举,包括提供商准入、等待的源工作及其完成
coalesced 跟随者等待同一调用者和请求的现有枚举
projection_final 枚举后所需的共享投影就绪状态
delivery 同步最终投影、可见性过滤和响应回调
阶段 观测

每个已进入的阶段贡献一个已完成的耗时观测,包括该阶段抛出异常时。未访问的阶段不存在。仅元数据请求不会进入这些阶段。提供者和投影的持续时间包含等待的工作和调度;它们不是 CPU 时间,也不是对共享工作的独占所有权。

线程 CPU 直方图仅记录 planning 和 delivery。其区间从不跨越 await,并在遥测发出前结束。它们包含同线程原生工作和垃圾回收,但排除 worker CPU、后台物化和进度发布。它们是选定的 CPU 区间,不是完整的请求 CPU 总计。如果 CPU 计数器读取失败,则该 CPU 观测被省略,而耗时计时和请求结果仍可用。计算均值时,比较每个指标自身的计数。

这些可信的、无负载的观测在诊断和感兴趣的可信消费者处于活动状态时,使用现有的有界诊断队列和导出器端点。它们没有慢日志阈值,也不添加请求、会话、提供者或主机标签。六个耗时样本群体和两个 CPU 样本群体在现有 2,048 个样本上限下最多消耗八个聚合样本,或 152 个暴露的直方图序列。启动阶段发出及其整个进程的 CPU 语义保持不变。

禁用 traces 的 OpenTelemetry 导出器不请求阶段事件。已配置的 Prometheus 导出器将这些观测记录为指标,无需 OpenTelemetry traces。

运行时身份

openclaw_gateway_build_info 的值为 1,用于标识提供抓取的服务进程。其 process_instance_id 是由 system.info 返回的同一进程拥有的 UUID;进程重启时会变化,包括 PID 被复用时。build_id 与 hello.server.buildId 报告的已加载构建匹配,并在该来源不可用时省略。更新磁盘上的文件不会更改运行中进程的身份。

启用诊断时,导出器会在服务启动时捕获这些事实,早于记录事件。 该指标在现有上限下使用一个聚合样本。不具备可选运行时身份能力的旧主机会省略它。UUID 仅限于此 info 指标;它不会添加到 RPC 或其他指标标签中。

使用同一抓取中的 info 样本,为新测量进行归属,并在进程变化时拆分计数器区间。它不是健康信号、请求 ID 或导出器纪元:在同一进程中重启导出器会重置其计数器,同时保留进程身份。它无法重新标记旧样本,也无法建立完整的诊断丢失覆盖。

事件循环观测窗口

openclaw_liveness_cpu_core_ratio 以核心当量衡量整个进程的 CPU 使用率,包括 worker 和原生线程,并且可能超过 1。应结合主线程延迟和利用率来解释它;参见 CPU 压力和事件循环延迟。

事件循环直方图记录每个已完成的 Gateway 健康监控窗口的最大延迟。计数器累加这些窗口所代表的秒数。两者均为累计值:后来的健康窗口不会抹去较早的高延迟观测。就绪、状态和抓取请求会消费已完成的观测,而不会推进或重置采样窗口。

监控器每 20 毫秒采样一次已流逝的事件循环区间,并在至少一秒后完成一个窗口,或在延迟警告时更早完成。 它在普通窗口重置期间保留待处理区间,因此在逾期样本之前读取健康状态不会抹去该延迟。直方图计数是窗口计数,不是停顿 计数。直方图分位数描述窗口最大值,而不是采样的事件循环 延迟分布或其总体 p99。这些指标没有请求标签或 trace 归属,也不会标识阻塞的 JavaScript 函数。

采集使用上述插件启用机制。当感兴趣的指标导出器运行时启动;它不会回填更早的窗口。有意的 监控重置会丢弃未完成的窗口。诊断队列丢弃、 导出器的序列上限和进程重启也可能丢失观测。评估 覆盖范围时,关注现有的丢弃计数器和所代表时长计数器。就绪决策和持续性存活警告阈值保持不变。

内存与进程周转

openclaw_memory_bytes 暴露 rss、heap_total、heap_used、external、 array_buffers、worker_heap_total 和 worker_heap_used。RSS 覆盖 整个进程。无前缀的堆和原生缓冲区值覆盖主 isolate;array_buffers 包含在 external 中,因此不要将它们相加。 这些值未计入所有原生分配或分配器 arena。

Worker 总计汇总资源注册表启动后创建的存活 Worker 的已完成原生堆样本。在 Node 上,这包括直接插件 Worker; 嵌套 Worker、原生库线程池(例如 Discord DAVE 的 Rayon 池) 和 V8 内部线程位于父注册表之外。 30 秒诊断心跳启动非阻塞刷新,每个 Worker 最多保留一个未完成请求。样本在 60 秒后过期, 并在 Worker 退出时移除。比较 openclaw_worker_heap_sampled_count 与 openclaw_worker_count:启动、不可用的 API 和停滞的 Worker 可能 产生部分总计。不会创建堆快照或额外采样计时器。 openclaw_worker_heap_used_bytes{script="..."} 对每个 Worker 脚本的新鲜堆样本求和。标签使用固定的允许列表,涵盖运行时入口点、池化和捆绑插件 Worker 的基本名称,并在源码和打包运行中规范化为 .js;未知、eval 和未包装的第三方 Worker 使用 other。从不记录完整路径和 eval 源码。当脚本没有存活的新鲜样本时,其序列 消失。内存压力日志包含相同的字节计数和 Worker 覆盖计数,外加 workerHeaps:以 {script, heapUsed, heapTotal}(字节)形式列出五个最大的 单个新鲜 Worker 堆,使用相同的有界脚本名称。

openclaw_worker_started_total{script="..."} 和 openclaw_worker_retired_total{script="...",reason="..."} 通过同一次心跳暴露累计 启动次数和已确认的原生退出次数。退役原因包括 idle_timeout、memory_pressure、closed、rotation、cancelled、 failure,或者当所有者未指定原因时的 exit。退役请求 在 Worker 退出之前不计入;终止失败和重试不会 重复计数。计数器在 exporter 重启后保留,并随进程重置。 按脚本统计每分钟启动次数时,使用 60 * rate(openclaw_worker_started_total[5m])。这些计数共享上述注册表覆盖范围限制;它们不 测量退出时释放的常驻内存。

openclaw gateway call diagnostics.lanes --json 还会报告 workerCount、 workerPoolCount 和 workerPools。每个池条目包含进程本地的 poolId、一个允许列表中的 script,以及其存活的 workerCount。响应会列出 最大的 100 个存活池;workerPoolCount 包含所有存活池。计数 包括待退役项,直到原生退出为止;当某个池没有存活 Worker 时,这些计数会消失。直接 Worker 会计入 workerCount,但没有池条目。 这些是 JavaScript Worker 计数,而不是操作系统线程普查。

openclaw_child_process_spawn_total{family="..."} 统计通过 OpenClaw 共享的 spawn 和 exec 所有者成功启动的次数,包括代理启动。 必须启用诊断。现有心跳会在至少一分钟之后发布累计 计数,调试日志会使用实际经过的时间间隔报告计数和速率。失败的启动、绕过这些所有者的直接调用,以及由子进程启动的后代进程均被排除。family 是 固定的可执行文件名允许列表;无法识别的命令会变成 other。 参数和路径永远不会被记录。统计每分钟启动次数时,使用 60 * rate(openclaw_child_process_spawn_total[5m]);该窗口可容纳 按分钟批量发布的机制。这两种核算路径都不会改变压力 阈值或用户工具执行。

垃圾回收持续时间

openclaw_gc_duration_seconds 记录 Node.js 为承载 JavaScript isolate 报告的垃圾回收(GC)经过时长。每个观测值是一条 GC 条目,而不是 CPU 时间、已分配字节数,或保证的 stop-the-world 暂停。 将其桶计数与事件循环窗口最大值进行比较,以调查 GC 是否是导致卡顿的可能因素;采集间隔匹配并不能证明因果关系。

采集使用现有的诊断启用机制,并在诊断心跳观察到感兴趣的消费者(例如 metrics exporter)时开始。在心跳启动之后添加的消费者可能需要等待到下一个 30 秒 tick,如果事件循环停滞,则可能等待更长时间。观察者激活之前的条目不会被回填。 需求会在条目交付时检查,因此在下一次心跳之前短暂的消费者缺失仍可能导致延迟观测。失去最后一个 消费者会抑制新的导出;观察者会在下一次心跳时断开。 禁用诊断或停止心跳会立即断开它。

在第一次观测之前,该直方图不存在,因此缺失并不能证明 GC 为零。队列丢弃、序列上限、观测间隙和进程重启都会限制 覆盖范围。禁用/重新启用诊断会保留 exporter 的现有 计数器;重启 exporter 会像往常一样重置它们。不会采集额外的计时器、GC 触发、trace 归因或应用负载。

标签策略

有界、低基数标签

Prometheus 标签保持有界且低基数。exporter 不会输出原始诊断标识符,例如 runId、sessionKey、sessionId、callId、toolCallId、消息 ID、聊天 ID 或 provider 请求 ID。

标签值会被脱敏,并且必须符合 OpenClaw 的低基数字符策略。不符合策略的值会根据指标被替换为 unknown、other 或 none。看起来像作用域 agent 会话键的标签也会被替换为 unknown。

序列上限与溢出核算

exporter 将内存中保留的时间序列上限设为 2048 条,涵盖计数器、gauge 和直方图的总和。超出该上限的新序列会被丢弃,并且每次都会使 openclaw_prometheus_series_dropped_total 递增 1。

将此计数器视为上游某个属性正在泄漏高基数值的硬性信号。exporter 永远不会自动提高上限;如果它持续上升,请修复源头,而不是禁用上限。

Prometheus 输出中永不出现的内容
  • prompt 文本、响应文本、工具输入、工具输出、system prompt
  • Talk 转录、音频负载、call id、room id、handoff token、turn id 和原始 session id
  • 原始 provider 请求 ID(仅在适用时,span 上会有有界哈希——绝不在 metrics 上)
  • session key 和 session ID
  • 主机名、文件路径、secret 值

PromQL 配方

# Gateway RPC requests per second by method
sum by (method) (rate(openclaw_gateway_rpc_requests_total[5m]))

# 95th percentile first-response latency by method
histogram_quantile(
  0.95,
  sum by (le, method) (rate(openclaw_gateway_rpc_first_response_seconds_bucket[5m]))
)

# 95th percentile operator request start-queue wait by method
histogram_quantile(
  0.95,
  sum by (le, method) (rate(openclaw_gateway_rpc_queue_wait_seconds_bucket[5m]))
)

# Tokens per second, split by provider
sum by (provider) (rate(openclaw_model_tokens_total[1m]))

# Spend (USD) over the last hour, by model
sum by (model) (increase(openclaw_model_cost_usd_total[1h]))

# 95th percentile model run duration
histogram_quantile(
  0.95,
  sum by (le, provider, model)
    (rate(openclaw_run_duration_seconds_bucket[5m]))
)

# Queue wait time SLO (95p under 2s)
histogram_quantile(
  0.95,
  sum by (le, lane) (rate(openclaw_queue_lane_wait_seconds_bucket[5m]))
) < 2

# Skill usage, split by bounded source
sum by (skill, source) (increase(openclaw_skill_used_total[24h]))

# Dropped Prometheus series (cardinality alarm)
increase(openclaw_prometheus_series_dropped_total[15m]) > 0

# Completed windows whose maximum delay exceeded one second
increase(openclaw_gateway_event_loop_delay_max_seconds_count[5m])
  - increase(openclaw_gateway_event_loop_delay_max_seconds_bucket{le="1"}[5m])

# Seconds represented by exported event-loop windows
increase(openclaw_gateway_event_loop_observed_seconds_total[5m])

# Observed GC entries whose elapsed duration exceeded one second
increase(openclaw_gc_duration_seconds_count[5m])
  - increase(openclaw_gc_duration_seconds_bucket{le="1"}[5m])

Tip

对于跨提供商仪表盘,优先使用 gen_ai_client_token_usage:它遵循 OpenTelemetry GenAI 语义约定,并且与非 OpenClaw GenAI 服务的指标保持一致。

在 Prometheus 与 OpenTelemetry 导出之间选择

OpenClaw 独立支持这两种导出表面。你可以启用其中一种、两者都启用,或都不启用。

  • 拉取 模型:Prometheus 抓取 /api/diagnostics/prometheus。
  • 无需外部收集器。
  • 通过常规 Gateway 认证进行身份验证。
  • 该表面仅提供指标(不包含追踪或日志)。
  • 最适合已标准化使用 Prometheus + Grafana 的技术栈。
  • 推送 模型:OpenClaw 通过 OTLP/HTTP 向收集器或兼容 OTLP 的后端发送数据。
  • 该表面包含指标、追踪和日志。
  • 当你需要两者时,可通过 OpenTelemetry Collector(prometheus 或 prometheusremotewrite 导出器)桥接到 Prometheus。
  • 完整目录参见 OpenTelemetry 导出。

故障排查

空响应体
  • 检查配置中 diagnostics.enabled 是否未设置为 false(默认值为 true)。
  • 使用 openclaw plugins list --enabled 确认插件已启用并加载。
  • 生成一些流量;计数器和直方图只有在至少发生一次事件后才会输出行。
401 / 未授权

该端点需要 Gateway 操作员范围(auth: "gateway" 且 gatewayRuntimeScopeSurface: "trusted-operator")。请使用 Prometheus 访问其他 Gateway 操作员路由时使用的同一令牌或密码。没有公开的未认证模式。

403 missing scope: operator.read

调用方已完成身份验证,但其有效的操作员范围不包含 operator.read。当 trusted-proxy 等携带身份信息的认证模式将抓取器映射到某个 命名角色,且该角色的范围上限排除了读取权限时,就会出现这种情况。请为抓取器角色授予 operator.read(或 operator.write / operator.admin,它们隐含该权限)。

openclaw_prometheus_series_dropped_total 持续增长

某个新属性正在超过 2048 个序列的上限。检查最近的指标,查找意外高基数的标签,并在源头修复。导出器会故意丢弃新序列,而不是静默重写标签。

Prometheus 在重启后显示过期序列

该插件仅在内存中保留状态。Gateway 重启后,计数器会重置为零,仪表值将从下一次报告值重新开始。请使用 PromQL rate() 和 increase() 来平滑处理重置。

本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw