文档目录

8.4 配套代码:SLI 采集与错误预算

对应小节:8.4 生产观测:SLI 与错误预算 三部分:① SLO 定义(可执行);② 错误预算 PromQL;③ 看板与告警的数据源。

一、可执行的 SLO 定义

# slo/slos.yaml —— 机器可读的 SLO 定义
# 用于生成:Prometheus 记录规则、Grafana 看板、错误预算告警
slos:
  # ═══ 可用性 ═══
  - name: orders-availability
    title: "订单查询可用性"
    description: "订单查询接口的非 5xx 比例"
    sli:
      type: ratio
      good: 'sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))'
      total: 'sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))'
    objective: 0.999
    window: 30d
    error_budget: 0.001

  # ═══ 延迟 ═══
  - name: orders-latency
    title: "订单查询延迟"
    description: "订单查询 P99 低于 200ms 的比例"
    sli:
      type: threshold
      metric: 'http_request_duration_seconds_bucket{uri="/orders/{id}"}'
      threshold: 0.2
      # good = 落在 200ms 桶内的请求数
      good: 'sum(rate(http_request_duration_seconds_bucket{uri="/orders/{id}", le="0.2"}[5m]))'
      total: 'sum(rate(http_request_duration_seconds_count{uri="/orders/{id}"}[5m]))'
    objective: 0.999
    window: 30d
    error_budget: 0.001
    # ⚠️ 负载前提:低流量时 P99 不稳定
    precondition: 'sum(rate(http_requests_total{uri="/orders/{id}"}[5m])) > 100'

  # ═══ 业务 SLI(比技术 SLI 更接近真实价值)═══
  - name: order-creation-success
    title: "下单成功率"
    description: "下单请求中业务成功的比例"
    sli:
      type: ratio
      good: 'sum(rate(business_orders_success_total[5m]))'
      total: 'sum(rate(business_orders_total[5m]))'
    objective: 0.995
    window: 30d
    error_budget: 0.005

注意三项的设计差异:

SLO 类型 特点
可用性 ratio 最常见
延迟 threshold 用「落在阈值内的比例」而不是百分位数值(更稳定、更易计算预算)
业务 ratio 最接近真实价值

为什么延迟 SLO 用 threshold 而不是百分位:

❌ 用百分位:histogram_quantile(0.99, ...) < 0.2
   → 这是一个布尔值(达标/不达标),无法计算"消耗了多少预算"

✅ 用 threshold:P99 < 200ms 的【请求比例】> 99.9%
   → 可以直接算出"有多少请求违约" → 能计算错误预算

二、错误预算的 PromQL

-- ═══════════════════════════════════════════════════════════
-- ① 可用性 SLI(当前值)
-- ═══════════════════════════════════════════════════════════
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))
  / sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))

-- ═══════════════════════════════════════════════════════════
-- ② 错误预算剩余(30 天窗口,可用性)
-- ═══════════════════════════════════════════════════════════
1 - (
  (1 - (
    sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[30d]))
      / sum(rate(http_requests_total{uri="/orders/{id}"}[30d]))
  )) / (1 - 0.999)
)

-- ═══════════════════════════════════════════════════════════
-- ③ 错误预算剩余(延迟 SLO)
-- ═══════════════════════════════════════════════════════════
1 - (
  (1 - (
    sum(rate(http_request_duration_seconds_bucket{uri="/orders/{id}", le="0.2"}[30d]))
      / sum(rate(http_request_duration_seconds_count{uri="/orders/{id}"}[30d]))
  )) / (1 - 0.999)
)

-- ═══════════════════════════════════════════════════════════
-- ④ 燃烧率(多个窗口,用于告警)
-- ═══════════════════════════════════════════════════════════
-- 1 小时窗口
(
  1 - (
    sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[1h]))
      / sum(rate(http_requests_total{uri="/orders/{id}"}[1h]))
  )
) / (1 - 0.999)

-- 5 分钟窗口(短窗口,保证灵敏)
(
  1 - (
    sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))
      / sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))
  )
) / (1 - 0.999)

三、Prometheus 记录规则

# observability/recording-rules.yml
groups:
  - name: slo-recording
    interval: 30s
    rules:
      # ═══ 可用性 SLI ═══
      - record: slo:orders_availability:ratio_5m
        expr: |
          sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))
            / sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))

      - record: slo:orders_availability:ratio_1h
        expr: |
          sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[1h]))
            / sum(rate(http_requests_total{uri="/orders/{id}"}[1h]))

      - record: slo:orders_availability:ratio_30d
        expr: |
          sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[30d]))
            / sum(rate(http_requests_total{uri="/orders/{id}"}[30d]))

      # ═══ 延迟 SLI(threshold 型)═══
      - record: slo:orders_latency:ratio_5m
        expr: |
          sum(rate(http_request_duration_seconds_bucket{uri="/orders/{id}", le="0.2"}[5m]))
            / sum(rate(http_request_duration_seconds_count{uri="/orders/{id}"}[5m]))

      # ═══ 错误预算剩余 ═══
      - record: slo:orders_availability:budget_remaining
        expr: |
          1 - ((1 - slo:orders_availability:ratio_30d) / (1 - 0.999))

      - record: slo:orders_latency:budget_remaining
        expr: |
          1 - ((1 - slo:orders_latency:ratio_30d) / (1 - 0.999))

      # ═══ 燃烧率(多窗口)═══
      - record: slo:orders_availability:burn_rate_5m
        expr: (1 - slo:orders_availability:ratio_5m) / (1 - 0.999)

      - record: slo:orders_availability:burn_rate_1h
        expr: (1 - slo:orders_availability:ratio_1h) / (1 - 0.999)

四、Grafana 看板(JSON 骨架)

{
  "title": "SLO 与错误预算",
  "panels": [
    {
      "title": "错误预算剩余(30 天窗口)",
      "type": "gauge",
      "targets": [
        { "expr": "slo:orders_availability:budget_remaining", "legendFormat": "可用性" },
        { "expr": "slo:orders_latency:budget_remaining", "legendFormat": "延迟" }
      ],
      "fieldConfig": {
        "defaults": {
          "unit": "percentunit",
          "min": 0, "max": 1,
          "thresholds": {
            "steps": [
              { "color": "red", "value": 0 },
              { "color": "yellow", "value": 0.2 },
              { "color": "green", "value": 0.5 }
            ]
          }
        }
      }
    },
    {
      "title": "燃烧率(1h / 6h / 1d)",
      "type": "timeseries",
      "targets": [
        { "expr": "slo:orders_availability:burn_rate_5m", "legendFormat": "5m" },
        { "expr": "slo:orders_availability:burn_rate_1h", "legendFormat": "1h" }
      ]
    },
    {
      "title": "按接口分解的错误预算消耗",
      "type": "table",
      "targets": [
        {
          "expr": "topk(5, (1 - sum by (uri) (rate(http_requests_total{status!~\"5..\"}[30d])) / sum by (uri) (rate(http_requests_total[30d]))) / (1 - 0.999))",
          "format": "table"
        }
      ]
    },
    {
      "title": "饱和度(领先指标)",
      "type": "timeseries",
      "targets": [
        { "expr": "db_pool_pending", "legendFormat": "连接池 pending" },
        { "expr": "executor_queue_depth", "legendFormat": "队列深度" },
        { "expr": "rate(container_cpu_cfs_throttled_periods_total[5m]) / rate(container_cpu_cfs_periods_total[5m])", "legendFormat": "CPU 节流比例" }
      ]
    }
  ]
}

注意第四个面板:饱和度与错误预算放在同一个看板上——因为:

✅ 正确做法:
   错误预算开始消耗 → 看相邻的饱和度面板 → 立刻知道是哪个资源的问题

❌ 错误做法:
   错误预算和饱和度在两个不同的看板 → 排查时要来回切换

五、SLO 定义校验脚本

# tools/validate-slos.py <slos.yaml>
"""
校验 SLO 定义是否完整、可用。

检查项:
  ① 每项 SLO 都有 objective / window / SLI
  ② 延迟类 SLO 有负载前提(precondition)
  ③ objective 不是 100%(不可能达成的目标会导致预算为 0)
  ④ 延迟用 threshold 类型而不是百分位(否则无法算预算)
"""
import sys

try:
    import yaml
except ImportError:
    print("需要 pyyaml:pip install pyyaml")
    sys.exit(1)


def validate(path):
    data = yaml.safe_load(open(path, encoding="utf-8"))
    slos = data.get("slos", [])

    problems = []
    warnings = []

    for slo in slos:
        name = slo.get("name", "?")
        print(f"检查 {name}...")

        # ① 必需字段
        for field in ("objective", "window", "sli"):
            if field not in slo:
                problems.append(f"{name}: 缺少 {field}")

        # ② objective 合理性
        obj = slo.get("objective", 0)
        if obj >= 1.0:
            problems.append(f"{name}: objective = {obj}(100% 不可能达成,会导致预算为 0)")
        elif obj < 0.9:
            warnings.append(f"{name}: objective = {obj} 偏低(低于 90% 说明目标太松)")

        # ③ 延迟类 SLO 必须有负载前提
        sli = slo.get("sli", {})
        if sli.get("type") == "threshold" and "precondition" not in slo:
            warnings.append(f"{name}: 延迟类 SLO 缺少 precondition(低流量时 P99 不稳定)")

        # ④ 延迟 SLI 应该用 threshold 而不是百分位
        if "histogram_quantile" in str(sli.get("good", "")):
            problems.append(f"{name}: 延迟 SLI 用了百分位 —— 无法计算错误预算,应改用 threshold 类型")

        # ⑤ 窗口合理性
        window = slo.get("window", "")
        if window and not any(window.endswith(u) for u in ("d", "w")):
            warnings.append(f"{name}: window = {window}(建议用 30d 或更长的窗口)")

    print()
    print("═" * 70)
    if problems:
        print(f"❌ {len(problems)} 个问题:")
        for p in problems:
            print(f"   - {p}")
    if warnings:
        print(f"⚠️  {len(warnings)} 个警告:")
        for w in warnings:
            print(f"   - {w}")
    if not problems and not warnings:
        print("✅ SLO 定义完整")
    print("═" * 70)

    return 1 if problems else 0


if __name__ == "__main__":
    sys.exit(validate(sys.argv[1] if len(sys.argv) > 1 else "slo/slos.yaml"))

六、动手改造

改动 观察什么
用 validate-slos.py 校验你的 SLO 定义 看它检出多少问题(尤其"缺少负载前提")
把某条 SLO 的 objective 改成 1.0 校验会报错——100% 的目标会让预算为 0
把延迟 SLI 改成用 histogram_quantile 校验会报"无法计算错误预算"
在 Grafana 里把饱和度和错误预算放同一个看板 排查时更快定位
故意让服务返回 5xx 一部分请求 观察错误预算与燃烧率的变化

七、这段代码的局限

  • slo/slos.yaml 是自定义格式:需要自己写生成器把它转成 Prometheus 规则(或手工维护规则)。
  • 30 天窗口的 PromQL 很重:rate(...[30d]) 需要扫描大量数据,生产建议用预聚合(recording rules 定时计算,看板查记录规则)。
  • 错误预算的"百分比"含义依赖请求量:低流量服务(每天 100 个请求)的 0.1% 预算是 0.1 个请求——小流量服务的 SLO 要按「绝对数量」设,而不是百分比。
  • 业务 SLI 需要额外埋点:business_orders_success_total 这类指标需要业务代码配合上报。
  • Grafana 看板 JSON 需要按你的数据源调整:datasource、unit、thresholds 都可能需要改。