跳转到内容

边缘观测数据传输模型(现行目标版)

本文是「Agent ↔ Server 观测数据怎么传」的最新权威说明。
读完应能回答:传什么、从哪采、多久采一次、Server 怎么存、产品指标从哪查。
协议字段与 DDL 细节另见 观测上报协议与表结构;问题背景见 边缘可观测与业务流量统计


0. 先记住三层

回答的问题唯一数据来源产品例子
L1 业务交付提供了多少数据?多少请求?access.log 明细已提供数据、请求数、UV、状态码、Top 域名
L2 边缘健康OpenResty 活着吗?现在多少连接?本机 /openflare/observability节点健康、当前连接
L3 宿主机资源CPU/内存/磁盘/网卡怎样?操作系统读数容量趋势、宿主机网卡

三层互不对账。
「已提供数据」≠「当前连接」≠「宿主机网卡出站」。


1. 总览:谁采集、谁上报、谁聚合

text
┌─────────────────────────────────────────────────────────────┐
│ 边缘节点                                                      │
│                                                               │
│  访客请求 ──► OpenResty                                        │
│                 │                                              │
│                 ├─ access.log(每请求一行)  ←── L1 采集点      │
│                 │                                              │
│                 └─ 连接状态(进程内维护)                         │
│                        │                                       │
│                        ▼                                       │
│              GET /openflare/observability  ←── L2 读快照       │
│              (不扫日志、不重算业务量)                            │
│                                                               │
│  操作系统 /proc 等  ────────────────────── L3 读快照            │
│                                                               │
│              ┌────────── Agent ──────────┐                     │
│              │ 默认每 3s 组一包 NodePayload  │                     │
│              │  · tail access.log 增量    │                     │
│              │  · GET 本机 observability  │                     │
│              │  · 读 host_metrics         │                     │
│              └────────────┬──────────────┘                     │
└─────────────────────────────│──────────────────────────────────┘
                              │ HTTP 心跳 或 WebSocket status

┌─────────────────────────────────────────────────────────────┐
│ Server(控制面)                                               │
│  · 明细 → ClickHouse of_node_access_logs                      │
│  · 健康 → 节点最新态 + of_node_edge_health                      │
│  · 主机 → of_node_metric_snapshots                            │
│  · 业务趋势 / Zone 统计 = 只对 access_logs 做 sum/count/uniq   │
└─────────────────────────────────────────────────────────────┘
角色做什么不做什么
OpenResty写 access.log;维护连接数不向控制面直接上报
Agent采集事实并上报不算 UV/TopN/24h 已提供数据
Server入库 + 聚合解释不信任边缘业务预汇总

2. 采集频率(默认)

动作默认频率配置
Agent → Server 上报每 3 秒 一次完整 payloadheartbeat_interval / 控制面 agent_heartbeat_interval(毫秒,默认 3000
组包时 tail access.log随上报(两次上报之间的新行)同上
组包时 GET /openflare/observability随上报(读当前连接快照)同上
组包时读主机指标随上报同上
OpenResty 写 access.log每个请求结束时 1 行与心跳无关
连接数在进程内更新连接变化时(内核维护)与心跳无关
离线补传窗口默认保留约 60 分钟observability_replay_minutes
节点离线判定60 秒 无成功心跳node_offline_threshold(默认 60000 毫秒)

说明:

  • Agent 没有单独的「采样时钟」;采样点 = 上报点(默认 3s)。
  • access.log 是「请求级连续写入」;Agent 只是周期性 搬运增量行
  • /openflare/observability 不是「被调用才开始统计业务」;对连接而言是 读 Nginx 已有瞬时值

传输通道:

  • HTTP 心跳:按间隔 POST 整包。
  • WebSocket:连通后按同一间隔发 status 消息(内容同构);此时不再走 HTTP 心跳双发。

3. Agent → Server 数据包(NodePayload v2)

3.1 结构骨架

json
{
  "schema_version": 2,
  "node_id": "n_01hxyz",
  "name": "edge-shanghai-1",
  "ip": "203.0.113.10",
  "version": "3.4.0",
  "ext_version": "",
  "current_version": "20260718-abc",
  "last_error": "",
  "profile": { },
  "host_metrics": { },
  "edge_health": { },
  "access_logs": [ ],
  "buffered": [ ],
  "health_events": [ ],
  "waf_ip_group_checksums": { }
}
字段含义
身份/版本/last_error控制节点是谁、跑什么版本
profile低频概况主机名、核数等(变化才报)
access_logsL1访问明细增量
edge_healthL2OpenResty 健康 + 当前连接
host_metricsL3CPU/内存/磁盘/网卡读数
buffered补传离线期间攒的事实批次
health_events事件如 openresty_unhealthy
waf_ip_group_checksums同步非观测湖

协议已删除(无兼容层,旧 Agent 必须升级):

  • traffic_report
  • openresty_observation(含 rx/tx)
  • snapshot / buffered_observability
  • 业务含义的 openresty 吞吐字段

4. L1 业务:access_logs

4.1 采集从哪里来

步骤位置说明
1OpenResty log_format openflare_json每请求写一行 JSON 到 access_log_path
2Agent 按文件 offset tail 增量两次心跳之间的新行
3解析后放入 access_logs[]可截断过长 path;不做 sum/count

日志格式(OpenResty 变量):

text
ts            ← $time_iso8601
host          ← $host
path          ← $request_uri
remote_addr   ← $remote_addr
status        ← $status
request_time  ← $request_time
bytes_sent    ← $body_bytes_sent     【已提供数据 = 响应体字节】
request_length← $request_length      【接收数据】
user_agent    ← $http_user_agent
cache_status  ← $upstream_cache_status  【缓存状态;UI 可推导命中/回源/未缓存】

观测端口请求 不写 业务 access.log(独立 server access_log off)。

4.2 上报示例

json
"access_logs": [
  {
    "logged_at_unix": 1721289601,
    "remote_addr": "198.51.100.20",
    "host": "www.example.com",
    "path": "/api/v1/ping",
    "status_code": 200,
    "bytes_sent": 1024,
    "request_length": 128,
    "request_time_ms": 15,
    "user_agent": "curl/8.0",
    "cache_status": "MISS"
  },
  {
    "logged_at_unix": 1721289602,
    "remote_addr": "198.51.100.21",
    "host": "www.example.com",
    "path": "/index.html",
    "status_code": 200,
    "bytes_sent": 8192,
    "request_length": 300,
    "request_time_ms": 8,
    "user_agent": "Mozilla/5.0",
    "cache_status": "HIT"
  }
]
字段解释
bytes_sent已提供数据(单请求);全局/Zone 合计 = Server sum
request_length接收数据(单请求)
logged_at_unix请求完成时间(业务时间轴)
host用于 Zone 域名过滤
cache_status$upstream_cache_status 原样;详情/列表可推导三态(命中/回源/未缓存);不上报 upstream 地址
regionServer 入库时 GeoIP 写入

4.3 Server 如何用(产品指标)

产品指标算法(仅 L1)
已提供数据sum(bytes_sent)
接收数据sum(request_length)
请求数count()
UVuniqExact(remote_addr)
状态码分布group by status_code
Top 域名group by host
Zone 页同上 + host IN (该 Zone 域名)
看板业务区同上,全局或 Top 过滤

落库表:of_node_access_logs(可选 Server 侧 of_access_log_hourly 加速,Agent 不写)。

4.4 上报频率

text
请求发生 ──立即──► 写 access.log
Agent 每 3s ──搬运──► 这 3s 内新行(可能 0 行,也可能很多行)
Server ──立即/批量──► CH

业务量正确性 不依赖 3s 对齐;3s 只影响「明细到达控制面的延迟」和单包条数。


5. L2 健康:edge_health 与 /openflare/observability

5.1 本机监测口

数据采集接口:

http
GET http://127.0.0.1:{openresty_observability_port}/openflare/observability

默认端口:18081openresty_observability_port)。

职责: 回答「OpenResty 此刻怎样」,回答业务已提供多少数据。

返回示例

json
{
  "ok": true,
  "captured_at_unix": 1721289600,
  "connections": {
    "active": 42,
    "reading": 0,
    "writing": 1,
    "waiting": 41
  }
}
字段是否瞬时从哪来说明
ok当次探测能返回 200 即 true探活
captured_at_unix采样时刻ngx.time()与上报对齐
connections.active瞬时Nginx 连接状态(原 stub_status Active)当前活跃连接
reading / writing / waiting瞬时同上细分可选但建议带

不返回(已删除):

旧字段原因
request_count / error_count / UV / status_codes / top_domains业务窗汇总,改由 access log
openresty_rx_bytes / openresty_tx_bytes与已提供/接收数据重复且易错
source_countries从未实现;国家走 Server GeoIP
server.accepts/handled/requests进程累计 counter,易与业务请求混淆;主路径不收录

/openflare/stub_status 保留;/openflare/observability 内部读取该口组装连接数 JSON,Agent 健康检查也直接探测该口。

5.2 采集机制(读快照)

text
Nginx 在连接建立/释放时维护 Active connections 等

Agent GET /openflare/observability

只读取「当前值」拼 JSON 返回
  • 不扫 access.log、不算 60 秒业务均值。
  • 返回 瞬时 gauge 快照

5.3 上报示例(装进 NodePayload)

json
"edge_health": {
  "captured_at_unix": 1721289600,
  "status": "healthy",
  "message": "",
  "connections": 42
}
字段来源
status / messageAgent 健康探测(配置校验/进程等,可与观测口 ok 配合);须与顶层 openresty_status / openresty_message 对齐
connections观测口 connections.active

落库拆分(权威源):

内容写入
最新 status + messagePG 节点表(UI / 列表 / 告警)
时序 status + connectionsCH of_node_edge_health无 message

6. L3 主机:host_metrics

6.1 采集从哪里来

Agent 读本机(如 /proc、磁盘统计等),每次组包时读一次

字段语义说明
cpu_usage_percent瞬时当前 CPU%
memory_* / storage_*瞬时用量/总量占用率在 Server 或展示层算
disk_read_bytes / disk_write_bytes累计 counter内核累计 IO
network_rx_bytes / network_tx_bytes累计 counter宿主机网卡,不是已提供数据

6.2 上报示例

json
"host_metrics": {
  "captured_at_unix": 1721289600,
  "cpu_usage_percent": 12.5,
  "memory_used_bytes": 4294967296,
  "memory_total_bytes": 16106127360,
  "storage_used_bytes": 50000000000,
  "storage_total_bytes": 107374182400,
  "disk_read_bytes": 9000000000,
  "disk_write_bytes": 12000000000,
  "network_rx_bytes": 500000000000,
  "network_tx_bytes": 800000000000
}

6.3 Server 如何处理累计字段

text
存原值时间序列
展示「这段时间网卡出站」时:
  delta = 本次 - 上次
  若 delta < 0 → 视为重启/计数器归零,本段增量记 0,从新基线继续
  若 delta >= 0 → 记入该时段增量
  • Agent 上报原值,不在边缘算 24h 总量。
  • 禁止 对累计原值做 sum 当业务量。
  • 文案必须是 「宿主机网卡」,禁止叫「已提供数据 / OpenResty 出站」。

落库:of_node_metric_snapshots(可选 capacity hourly MV)。


7. 一次完整上报示例

json
{
  "schema_version": 2,
  "node_id": "n_01hxyz",
  "name": "edge-shanghai-1",
  "ip": "203.0.113.10",
  "version": "3.4.0",
  "ext_version": "",
  "current_version": "20260718-abc",
  "last_error": "",
  "host_metrics": {
    "captured_at_unix": 1721289600,
    "cpu_usage_percent": 12.5,
    "memory_used_bytes": 4294967296,
    "memory_total_bytes": 16106127360,
    "storage_used_bytes": 50000000000,
    "storage_total_bytes": 107374182400,
    "disk_read_bytes": 9000000000,
    "disk_write_bytes": 12000000000,
    "network_rx_bytes": 500000000000,
    "network_tx_bytes": 800000000000
  },
  "edge_health": {
    "captured_at_unix": 1721289600,
    "status": "healthy",
    "message": "",
    "connections": 42
  },
  "access_logs": [
    {
      "logged_at_unix": 1721289595,
      "remote_addr": "198.51.100.20",
      "host": "www.example.com",
      "path": "/",
      "status_code": 200,
      "bytes_sent": 4096,
      "request_length": 200,
      "request_time_ms": 12
    }
  ],
  "buffered": [],
  "health_events": [],
  "waf_ip_group_checksums": {
    "1": "d41d8cd98f00b204e9800998ecf8427e"
  }
}

Server 落库示意:

payload 块写入
access_logs[0]CH 一行,bytes_sent=4096region 由 GeoIP 填
edge_health节点 openresty_status=healthy,connections=42
host_metricsCH metric 一行累计/瞬时字段

产品查询示意(24h):

  • 已提供数据 = 该节点(或全局)日志 sum(bytes_sent)
  • 当前连接 = 最新 edge_health.connections
  • 宿主机网卡出站 = metric 上 network_tx 非负差分之和

三者数字 不必相等


8. 离线补传 buffered

Agent 上报失败时,把 同一类事实 按窗口缓存在本地(默认约 60 分钟),恢复后塞进 buffered[]

json
"buffered": [
  {
    "captured_at_unix": 1721289500,
    "host_metrics": { },
    "edge_health": { },
    "access_logs": [ ]
  }
]
  • 只装事实,不装旧 TrafficReport。
  • Server 处理逻辑与主字段相同。

9. 端到端时序(默认 3s)

text
t=0.0s   访客请求完成 → 写 access.log 一行;连接数可能变化
t=0.1s   又一请求 → 又一行 log

t=3s     Agent 心跳:
           · 读走 2 行 access_logs
           · GET observability → connections=42
           · 读 host_metrics
           · 发给 Server
t=3s+    Server 入库;看板/Zone 查询时聚合日志
t=6s     下一轮…

10. 旧模型对照

旧做法新模型
Lua dict 60s 窗 request_count + Agent 10s 拉 + Server sum删除;请求数 = 日志 count
openresty_tx 当「出站」删除;已提供数据 = sum(bytes_sent)
两个口 observability + stub_status数据采集统一走 observability;stub_status 保留为探活与内部读取口
TrafficReport 预聚合删除;协议与 API 均无此路径
业务与网卡混称「流量」分文案、分 API、分表
健康 status/messagePG 最新态权威;CH 仅 status+连接时序

11. 配置与实现索引

位置/键
心跳间隔Agent heartbeat_interval;控制面 agent_heartbeat_interval(默认 3000ms)
离线阈值控制面 node_offline_threshold(默认 60000ms)
观测端口openresty_observability_port(默认 18081)
access.log 路径access_log_path
补传分钟数observability_replay_minutes(默认 60)
协议类型pkg/protocol/agent.go(落地时按 v2 演进)
表结构 DDLobservability-data-model.md

12. 修订记录

日期说明
2026-07-18初稿:作为「最新传输模型」单页说明——三层、频率、示例 JSON、采集来源、与旧模型对照
2026-07-18默认上报间隔 3s;离线阈值 60s;补传窗口 60 分钟
2026-07-18M5:edge_health 表、access_log_hourly、废弃 request_reports/obs_openresty 吞吐表
2026-07-18无兼容层:删除「兼容期可忽略」表述;健康 message 仅 PG、CH 无 message

基于 Apache License 2.0 发布