Prometheus
OpenClaw 可以通过官方
diagnostics-prometheus 插件暴露诊断指标。它会监听受信任的诊断事件,以及内部标记的、由 dispatcher 拥有的诊断事件(队列、内存和会话恢复信号),并在以下位置提供 Prometheus 文本端点:
内容类型为 text/plain; version=0.0.4; charset=utf-8,即标准的 Prometheus 暴露格式。
Warning
该路由使用 Gateway 身份验证(operator 作用域,受信任 operator 接口面),并要求调用方的有效作用域包含 operator.read(由 operator.write 或 operator.admin 隐含)。不要将其暴露为公开的未认证 /metrics 端点。请通过你用于其他 operator API 的同一身份验证路径对其进行抓取。
有关跟踪、日志、OTLP 推送以及 OpenTelemetry GenAI 语义属性,请参阅 OpenTelemetry 导出。
快速开始¶
1. 安装插件
2. 启用插件
3. 重启 Gateway
HTTP 路由在插件启动时注册,因此启用后需要重新加载。
4. 抓取受保护的路由
发送你的 operator 客户端使用的相同 Gateway 身份验证:
curl -H "Authorization: Bearer $OPENCLAW_GATEWAY_TOKEN" \
http://127.0.0.1:18789/api/diagnostics/prometheus
5. 接入 Prometheus
# prometheus.yml
scrape_configs:
- job_name: openclaw
scrape_interval: 30s
metrics_path: /api/diagnostics/prometheus
authorization:
credentials_file: /etc/prometheus/openclaw-gateway-token
static_configs:
- targets: ["openclaw-gateway:18789"]
Note
diagnostics.enabled 默认为 true;仅在严格受限的环境中将其设置为 false。如果导出器启动时该值为 false,插件仍会注册 HTTP 路由,但不会记录任何诊断事件或运行时身份,因此响应为空。
导出的指标¶
| 指标 | 类型 | 标签 |
|---|---|---|
openclaw_gateway_build_info |
gauge | process_instance_id,可选 build_id |
openclaw_gc_duration_seconds |
histogram | 无 |
openclaw_gateway_rpc_requests_total |
counter | method |
openclaw_gateway_rpc_first_response_seconds |
histogram | method |
openclaw_gateway_rpc_handler_seconds |
histogram | method |
openclaw_gateway_rpc_admission_seconds |
histogram | method |
openclaw_gateway_rpc_queue_wait_seconds |
histogram | method |
openclaw_gateway_rpc_stage_seconds |
histogram | method、phase |
openclaw_gateway_rpc_stage_thread_cpu_seconds |
histogram | method、phase |
openclaw_gateway_rpc_outcomes_total |
counter | phase、outcome |
openclaw_run_completed_total |
counter | channel、model、outcome、provider、trigger |
openclaw_run_duration_seconds |
histogram | channel、model、outcome、provider、trigger |
openclaw_model_call_total |
counter | api、error_category、model、observation_unit、outcome、provider、transport |
openclaw_model_call_duration_seconds |
histogram | api、error_category、model、observation_unit、outcome、provider、transport |
openclaw_model_failover_total |
counter | from_model、from_provider、lane、reason、suspended、to_model、to_provider |
openclaw_model_tokens_total |
counter | agent、channel、model、provider、token_type |
openclaw_gen_ai_client_token_usage |
histogram | model、provider、token_type |
openclaw_model_cost_usd_total |
counter | agent、channel、model、provider |
openclaw_model_usage_duration_seconds |
histogram | agent、channel、model、provider |
openclaw_skill_used_total |
counter | activation、agent、skill、source |
openclaw_tool_execution_total |
counter | error_category、outcome、params_kind、tool、tool_owner、tool_source |
| 指标 | 类型 | 标签 |
|---|---|---|
openclaw_tool_execution_duration_seconds |
histogram | error_category, outcome, params_kind, tool, tool_owner, tool_source |
openclaw_tool_execution_blocked_total |
counter | denied_reason, params_kind, tool, tool_owner, tool_source |
openclaw_harness_run_total |
counter | channel, error_category, harness, model, outcome, phase, plugin, provider |
openclaw_harness_run_duration_seconds |
histogram | channel, error_category, harness, model, outcome, phase, plugin, provider |
openclaw_webhook_received_total |
counter | channel, webhook |
openclaw_webhook_error_total |
counter | channel, webhook |
openclaw_webhook_duration_seconds |
histogram | channel, webhook |
openclaw_message_received_total |
counter | channel, source |
openclaw_message_dispatch_started_total |
counter | channel, source |
openclaw_message_dispatch_completed_total |
counter | channel, outcome, reason, source |
openclaw_message_dispatch_duration_seconds |
histogram | channel, outcome, reason, source |
openclaw_message_processed_total |
counter | channel, outcome, reason |
openclaw_message_processed_duration_seconds |
histogram | channel, outcome, reason |
openclaw_message_delivery_started_total |
counter | channel, delivery_kind |
openclaw_message_delivery_total |
counter | channel, delivery_kind, error_category, outcome |
openclaw_message_delivery_duration_seconds |
histogram | channel, delivery_kind, error_category, outcome |
openclaw_talk_event_total |
counter | brain, event_type, mode, provider, transport |
openclaw_talk_event_duration_seconds |
histogram | brain, event_type, mode, provider, transport |
openclaw_talk_audio_bytes |
histogram | brain, event_type, mode, provider, transport |
openclaw_queue_lane_size |
gauge | lane |
openclaw_queue_lane_wait_seconds |
histogram | lane |
openclaw_session_state_total |
counter | reason, state |
openclaw_session_queue_depth |
gauge | state |
openclaw_session_turn_created_total |
counter | agent, channel, trigger |
openclaw_session_stuck_total |
counter | reason, state |
openclaw_session_stuck_age_seconds |
histogram | reason, state |
openclaw_session_recovery_total |
counter | action, active_work_kind, state, status |
openclaw_session_recovery_age_seconds |
histogram | action, active_work_kind, state, status |
openclaw_gateway_event_loop_delay_max_seconds |
histogram | 无 |
openclaw_gateway_event_loop_observed_seconds_total |
counter | 无 |
openclaw_liveness_warning_total |
counter | reason |
openclaw_liveness_sessions |
gauge | state |
openclaw_liveness_event_loop_delay_p99_seconds |
histogram | reason |
openclaw_liveness_event_loop_delay_max_seconds |
histogram | reason |
openclaw_liveness_event_loop_utilization_ratio |
histogram | reason |
openclaw_liveness_cpu_core_ratio |
histogram | reason |
| 指标 | 类型 | 标签 |
|---|---|---|
openclaw_payload_large_total |
计数器 | action, channel, plugin, reason, surface |
openclaw_payload_large_bytes |
直方图 | action, channel, plugin, reason, surface |
openclaw_memory_bytes |
仪表 | kind |
openclaw_worker_count |
仪表 | 无 |
openclaw_worker_heap_sampled_count |
仪表 | 无 |
openclaw_worker_heap_used_bytes |
仪表 | script |
openclaw_worker_started_total |
计数器 | script |
openclaw_worker_retired_total |
计数器 | script, reason |
openclaw_child_process_spawn_total |
计数器 | family |
openclaw_memory_rss_bytes |
直方图 | 无 |
openclaw_memory_pressure_total |
计数器 | level, reason |
openclaw_telemetry_exporter_total |
计数器 | exporter, reason, signal, status |
openclaw_prometheus_series_dropped_total |
计数器 | 无 |
openclaw_diagnostic_async_queue_dropped_total |
计数器 | drop_class |
openclaw_diagnostic_async_queue_length |
仪表 | 无 |
对于模型调用指标,observation_unit="request" 衡量一次可观察的提供商请求。observation_unit="turn" 衡量一个合成的 Claude Code 或 Codex CLI 智能体回合,其中可能包含多个隐藏的提供商请求。比较延迟时,请将这些序列分开。
网关 RPC 指标涵盖有效的已认证 WebSocket 请求,包括后续拒绝。first_response 衡量从接收到发送方接受的第一帧为止;不可用或被抑制的发送没有持续时间样本。handler 衡量实际处理程序调用直至返回或抛出,admission 衡量从接收到该调用为止。queue_wait 仅衡量操作者请求启动队列等待,与命令/会话通道指标分开。它们衡量的是经过时间,而不是 CPU 时间。早期确认以及处理程序返回后的响应,与已完成的智能体工作不同。参见 网关 RPC 计时语义。
接收在已连接客户端的请求帧通过验证后开始。这些计时不包括 CLI 启动、本地诊断、连接/认证设置以及请求分发前的事件循环延迟。直方图记录已完成的观测:未完成的处理程序尚无处理程序持续时间样本。调查超时时,请比较请求数、已完成的计时和事件循环观测;仅低处理程序延迟不足以证明客户端路径响应迅速。
RPC 方法标签包含规范的核心方法名称、插件方法的 other,或 unknown。结果总计按阶段和结果聚合,没有方法维度。每个具有全部四种计时的方法在共享的 2,048 个样本上限中占用五个聚合样本。持续时间直方图占用一个样本,但会扩展为 19 个抓取序列(桶、总和和计数)。当上限填满时,现有样本会继续更新;未见的 RPC 或其他操作样本会被拒绝,并递增 openclaw_prometheus_series_dropped_total。监控该计数器:覆盖每个核心方法都可能填满上限,因此在解释总计或延迟百分位数时,零值很重要。异步诊断队列饱和也可能丢弃观测,由 openclaw_diagnostic_async_queue_dropped_total 报告。
目录列表阶段¶
阶段直方图目前仅覆盖 sessions.catalog.list。它们使用六个固定的阶段标签:
| 阶段 | 观测 |
|---|---|
projection_initial |
请求必须等待时的初始共享投影就绪状态 |
planning |
领导者提供商规划快照的同步创建和冻结 |
provider |
领导者枚举,包括提供商准入、等待的源工作及其完成 |
coalesced |
跟随者等待同一调用者和请求的现有枚举 |
projection_final |
枚举后所需的共享投影就绪状态 |
delivery |
同步最终投影、可见性过滤和响应回调 |
| 阶段 | 观测 |
|---|---|
每个已进入的阶段贡献一个已完成的耗时观测,包括该阶段抛出异常时。未访问的阶段不存在。仅元数据请求不会进入这些阶段。提供者和投影的持续时间包含等待的工作和调度;它们不是 CPU 时间,也不是对共享工作的独占所有权。
线程 CPU 直方图仅记录 planning 和 delivery。其区间从不跨越 await,并在遥测发出前结束。它们包含同线程原生工作和垃圾回收,但排除 worker CPU、后台物化和进度发布。它们是选定的 CPU 区间,不是完整的请求 CPU 总计。如果 CPU 计数器读取失败,则该 CPU 观测被省略,而耗时计时和请求结果仍可用。计算均值时,比较每个指标自身的计数。
这些可信的、无负载的观测在诊断和感兴趣的可信消费者处于活动状态时,使用现有的有界诊断队列和导出器端点。它们没有慢日志阈值,也不添加请求、会话、提供者或主机标签。六个耗时样本群体和两个 CPU 样本群体在现有 2,048 个样本上限下最多消耗八个聚合样本,或 152 个暴露的直方图序列。启动阶段发出及其整个进程的 CPU 语义保持不变。
禁用 traces 的 OpenTelemetry 导出器不请求阶段事件。已配置的 Prometheus 导出器将这些观测记录为指标,无需 OpenTelemetry traces。
运行时身份¶
openclaw_gateway_build_info 的值为 1,用于标识提供抓取的服务进程。其 process_instance_id 是由 system.info 返回的同一进程拥有的 UUID;进程重启时会变化,包括 PID 被复用时。build_id 与 hello.server.buildId 报告的已加载构建匹配,并在该来源不可用时省略。更新磁盘上的文件不会更改运行中进程的身份。
启用诊断时,导出器会在服务启动时捕获这些事实,早于记录事件。 该指标在现有上限下使用一个聚合样本。不具备可选运行时身份能力的旧主机会省略它。UUID 仅限于此 info 指标;它不会添加到 RPC 或其他指标标签中。
使用同一抓取中的 info 样本,为新测量进行归属,并在进程变化时拆分计数器区间。它不是健康信号、请求 ID 或导出器纪元:在同一进程中重启导出器会重置其计数器,同时保留进程身份。它无法重新标记旧样本,也无法建立完整的诊断丢失覆盖。
事件循环观测窗口¶
openclaw_liveness_cpu_core_ratio 以核心当量衡量整个进程的 CPU 使用率,包括 worker 和原生线程,并且可能超过 1。应结合主线程延迟和利用率来解释它;参见
CPU 压力和事件循环延迟。
事件循环直方图记录每个已完成的 Gateway 健康监控窗口的最大延迟。计数器累加这些窗口所代表的秒数。两者均为累计值:后来的健康窗口不会抹去较早的高延迟观测。就绪、状态和抓取请求会消费已完成的观测,而不会推进或重置采样窗口。
监控器每 20 毫秒采样一次已流逝的事件循环区间,并在至少一秒后完成一个窗口,或在延迟警告时更早完成。 它在普通窗口重置期间保留待处理区间,因此在逾期样本之前读取健康状态不会抹去该延迟。直方图计数是窗口计数,不是停顿 计数。直方图分位数描述窗口最大值,而不是采样的事件循环 延迟分布或其总体 p99。这些指标没有请求标签或 trace 归属,也不会标识阻塞的 JavaScript 函数。
采集使用上述插件启用机制。当感兴趣的指标导出器运行时启动;它不会回填更早的窗口。有意的 监控重置会丢弃未完成的窗口。诊断队列丢弃、 导出器的序列上限和进程重启也可能丢失观测。评估 覆盖范围时,关注现有的丢弃计数器和所代表时长计数器。就绪决策和持续性存活警告阈值保持不变。
内存与进程周转¶
openclaw_memory_bytes 暴露 rss、heap_total、heap_used、external、
array_buffers、worker_heap_total 和 worker_heap_used。RSS 覆盖
整个进程。无前缀的堆和原生缓冲区值覆盖主
isolate;array_buffers 包含在 external 中,因此不要将它们相加。
这些值未计入所有原生分配或分配器 arena。
Worker 总计汇总资源注册表启动后创建的存活 Worker 的已完成原生堆样本。在 Node 上,这包括直接插件 Worker;
嵌套 Worker、原生库线程池(例如 Discord DAVE 的 Rayon 池)
和 V8 内部线程位于父注册表之外。
30 秒诊断心跳启动非阻塞刷新,每个 Worker 最多保留一个未完成请求。样本在 60 秒后过期,
并在 Worker 退出时移除。比较 openclaw_worker_heap_sampled_count
与 openclaw_worker_count:启动、不可用的 API 和停滞的 Worker 可能
产生部分总计。不会创建堆快照或额外采样计时器。
openclaw_worker_heap_used_bytes{script="..."} 对每个
Worker 脚本的新鲜堆样本求和。标签使用固定的允许列表,涵盖运行时入口点、池化和捆绑插件 Worker 的基本名称,并在源码和打包运行中规范化为
.js;未知、eval 和未包装的第三方 Worker
使用 other。从不记录完整路径和 eval 源码。当脚本没有存活的新鲜样本时,其序列
消失。内存压力日志包含相同的字节计数和 Worker 覆盖计数,外加 workerHeaps:以 {script, heapUsed, heapTotal}(字节)形式列出五个最大的
单个新鲜 Worker 堆,使用相同的有界脚本名称。
openclaw_worker_started_total{script="..."} 和
openclaw_worker_retired_total{script="...",reason="..."} 通过同一次心跳暴露累计
启动次数和已确认的原生退出次数。退役原因包括 idle_timeout、memory_pressure、closed、rotation、cancelled、
failure,或者当所有者未指定原因时的 exit。退役请求
在 Worker 退出之前不计入;终止失败和重试不会
重复计数。计数器在 exporter 重启后保留,并随进程重置。
按脚本统计每分钟启动次数时,使用 60 * rate(openclaw_worker_started_total[5m])。这些计数共享上述注册表覆盖范围限制;它们不
测量退出时释放的常驻内存。
openclaw gateway call diagnostics.lanes --json 还会报告 workerCount、
workerPoolCount 和 workerPools。每个池条目包含进程本地的
poolId、一个允许列表中的 script,以及其存活的 workerCount。响应会列出
最大的 100 个存活池;workerPoolCount 包含所有存活池。计数
包括待退役项,直到原生退出为止;当某个池没有存活
Worker 时,这些计数会消失。直接 Worker 会计入 workerCount,但没有池条目。
这些是 JavaScript Worker 计数,而不是操作系统线程普查。
openclaw_child_process_spawn_total{family="..."} 统计通过 OpenClaw 共享的 spawn 和 exec 所有者成功启动的次数,包括代理启动。
必须启用诊断。现有心跳会在至少一分钟之后发布累计
计数,调试日志会使用实际经过的时间间隔报告计数和速率。失败的启动、绕过这些所有者的直接调用,以及由子进程启动的后代进程均被排除。family 是
固定的可执行文件名允许列表;无法识别的命令会变成 other。
参数和路径永远不会被记录。统计每分钟启动次数时,使用
60 * rate(openclaw_child_process_spawn_total[5m]);该窗口可容纳
按分钟批量发布的机制。这两种核算路径都不会改变压力
阈值或用户工具执行。
垃圾回收持续时间¶
openclaw_gc_duration_seconds 记录 Node.js 为承载 JavaScript isolate 报告的垃圾回收(GC)经过时长。每个观测值是一条
GC 条目,而不是 CPU 时间、已分配字节数,或保证的 stop-the-world 暂停。
将其桶计数与事件循环窗口最大值进行比较,以调查 GC 是否是导致卡顿的可能因素;采集间隔匹配并不能证明因果关系。
采集使用现有的诊断启用机制,并在诊断心跳观察到感兴趣的消费者(例如 metrics exporter)时开始。在心跳启动之后添加的消费者可能需要等待到下一个 30 秒 tick,如果事件循环停滞,则可能等待更长时间。观察者激活之前的条目不会被回填。 需求会在条目交付时检查,因此在下一次心跳之前短暂的消费者缺失仍可能导致延迟观测。失去最后一个 消费者会抑制新的导出;观察者会在下一次心跳时断开。 禁用诊断或停止心跳会立即断开它。
在第一次观测之前,该直方图不存在,因此缺失并不能证明 GC 为零。队列丢弃、序列上限、观测间隙和进程重启都会限制 覆盖范围。禁用/重新启用诊断会保留 exporter 的现有 计数器;重启 exporter 会像往常一样重置它们。不会采集额外的计时器、GC 触发、trace 归因或应用负载。
标签策略¶
有界、低基数标签
Prometheus 标签保持有界且低基数。exporter 不会输出原始诊断标识符,例如 runId、sessionKey、sessionId、callId、toolCallId、消息 ID、聊天 ID 或 provider 请求 ID。
标签值会被脱敏,并且必须符合 OpenClaw 的低基数字符策略。不符合策略的值会根据指标被替换为 unknown、other 或 none。看起来像作用域 agent 会话键的标签也会被替换为 unknown。
序列上限与溢出核算
exporter 将内存中保留的时间序列上限设为 2048 条,涵盖计数器、gauge 和直方图的总和。超出该上限的新序列会被丢弃,并且每次都会使 openclaw_prometheus_series_dropped_total 递增 1。
将此计数器视为上游某个属性正在泄漏高基数值的硬性信号。exporter 永远不会自动提高上限;如果它持续上升,请修复源头,而不是禁用上限。
Prometheus 输出中永不出现的内容
- prompt 文本、响应文本、工具输入、工具输出、system prompt
- Talk 转录、音频负载、call id、room id、handoff token、turn id 和原始 session id
- 原始 provider 请求 ID(仅在适用时,span 上会有有界哈希——绝不在 metrics 上)
- session key 和 session ID
- 主机名、文件路径、secret 值
PromQL 配方¶
# Gateway RPC requests per second by method
sum by (method) (rate(openclaw_gateway_rpc_requests_total[5m]))
# 95th percentile first-response latency by method
histogram_quantile(
0.95,
sum by (le, method) (rate(openclaw_gateway_rpc_first_response_seconds_bucket[5m]))
)
# 95th percentile operator request start-queue wait by method
histogram_quantile(
0.95,
sum by (le, method) (rate(openclaw_gateway_rpc_queue_wait_seconds_bucket[5m]))
)
# Tokens per second, split by provider
sum by (provider) (rate(openclaw_model_tokens_total[1m]))
# Spend (USD) over the last hour, by model
sum by (model) (increase(openclaw_model_cost_usd_total[1h]))
# 95th percentile model run duration
histogram_quantile(
0.95,
sum by (le, provider, model)
(rate(openclaw_run_duration_seconds_bucket[5m]))
)
# Queue wait time SLO (95p under 2s)
histogram_quantile(
0.95,
sum by (le, lane) (rate(openclaw_queue_lane_wait_seconds_bucket[5m]))
) < 2
# Skill usage, split by bounded source
sum by (skill, source) (increase(openclaw_skill_used_total[24h]))
# Dropped Prometheus series (cardinality alarm)
increase(openclaw_prometheus_series_dropped_total[15m]) > 0
# Completed windows whose maximum delay exceeded one second
increase(openclaw_gateway_event_loop_delay_max_seconds_count[5m])
- increase(openclaw_gateway_event_loop_delay_max_seconds_bucket{le="1"}[5m])
# Seconds represented by exported event-loop windows
increase(openclaw_gateway_event_loop_observed_seconds_total[5m])
# Observed GC entries whose elapsed duration exceeded one second
increase(openclaw_gc_duration_seconds_count[5m])
- increase(openclaw_gc_duration_seconds_bucket{le="1"}[5m])
Tip
对于跨提供商仪表盘,优先使用 gen_ai_client_token_usage:它遵循 OpenTelemetry GenAI 语义约定,并且与非 OpenClaw GenAI 服务的指标保持一致。
在 Prometheus 与 OpenTelemetry 导出之间选择¶
OpenClaw 独立支持这两种导出表面。你可以启用其中一种、两者都启用,或都不启用。
- 拉取 模型:Prometheus 抓取
/api/diagnostics/prometheus。 - 无需外部收集器。
- 通过常规 Gateway 认证进行身份验证。
- 该表面仅提供指标(不包含追踪或日志)。
- 最适合已标准化使用 Prometheus + Grafana 的技术栈。
- 推送 模型:OpenClaw 通过 OTLP/HTTP 向收集器或兼容 OTLP 的后端发送数据。
- 该表面包含指标、追踪和日志。
- 当你需要两者时,可通过 OpenTelemetry Collector(
prometheus或prometheusremotewrite导出器)桥接到 Prometheus。 - 完整目录参见 OpenTelemetry 导出。
故障排查¶
空响应体
- 检查配置中
diagnostics.enabled是否未设置为false(默认值为true)。 - 使用
openclaw plugins list --enabled确认插件已启用并加载。 - 生成一些流量;计数器和直方图只有在至少发生一次事件后才会输出行。
401 / 未授权
该端点需要 Gateway 操作员范围(auth: "gateway" 且 gatewayRuntimeScopeSurface: "trusted-operator")。请使用 Prometheus 访问其他 Gateway 操作员路由时使用的同一令牌或密码。没有公开的未认证模式。
403 missing scope: operator.read
调用方已完成身份验证,但其有效的操作员范围不包含 operator.read。当 trusted-proxy 等携带身份信息的认证模式将抓取器映射到某个 命名角色,且该角色的范围上限排除了读取权限时,就会出现这种情况。请为抓取器角色授予 operator.read(或 operator.write / operator.admin,它们隐含该权限)。
openclaw_prometheus_series_dropped_total 持续增长
某个新属性正在超过 2048 个序列的上限。检查最近的指标,查找意外高基数的标签,并在源头修复。导出器会故意丢弃新序列,而不是静默重写标签。
Prometheus 在重启后显示过期序列
该插件仅在内存中保留状态。Gateway 重启后,计数器会重置为零,仪表值将从下一次报告值重新开始。请使用 PromQL rate() 和 increase() 来平滑处理重置。
相关¶
- 诊断导出 — 用于支持包的本地诊断 zip
- 健康检查和就绪检查 —
/healthz和/readyz探针 - 日志记录 — 基于文件的日志记录
- OpenTelemetry 导出 — 用于追踪、指标和日志的 OTLP 推送
本页原文 Markdown:在 AtomGit 查看·内容源自开源项目 cl/openclaw