8.4 配套代码:SLI 采集与错误预算
对应小节:8.4 生产观测:SLI 与错误预算 三部分:① SLO 定义(可执行);② 错误预算 PromQL;③ 看板与告警的数据源。
一、可执行的 SLO 定义
# slo/slos.yaml —— 机器可读的 SLO 定义
# 用于生成:Prometheus 记录规则、Grafana 看板、错误预算告警
slos:
# ═══ 可用性 ═══
- name: orders-availability
title: "订单查询可用性"
description: "订单查询接口的非 5xx 比例"
sli:
type: ratio
good: 'sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))'
total: 'sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))'
objective: 0.999
window: 30d
error_budget: 0.001
# ═══ 延迟 ═══
- name: orders-latency
title: "订单查询延迟"
description: "订单查询 P99 低于 200ms 的比例"
sli:
type: threshold
metric: 'http_request_duration_seconds_bucket{uri="/orders/{id}"}'
threshold: 0.2
# good = 落在 200ms 桶内的请求数
good: 'sum(rate(http_request_duration_seconds_bucket{uri="/orders/{id}", le="0.2"}[5m]))'
total: 'sum(rate(http_request_duration_seconds_count{uri="/orders/{id}"}[5m]))'
objective: 0.999
window: 30d
error_budget: 0.001
# ⚠️ 负载前提:低流量时 P99 不稳定
precondition: 'sum(rate(http_requests_total{uri="/orders/{id}"}[5m])) > 100'
# ═══ 业务 SLI(比技术 SLI 更接近真实价值)═══
- name: order-creation-success
title: "下单成功率"
description: "下单请求中业务成功的比例"
sli:
type: ratio
good: 'sum(rate(business_orders_success_total[5m]))'
total: 'sum(rate(business_orders_total[5m]))'
objective: 0.995
window: 30d
error_budget: 0.005
注意三项的设计差异:
| SLO | 类型 | 特点 |
|---|---|---|
| 可用性 | ratio | 最常见 |
| 延迟 | threshold | 用「落在阈值内的比例」而不是百分位数值(更稳定、更易计算预算) |
| 业务 | ratio | 最接近真实价值 |
为什么延迟 SLO 用 threshold 而不是百分位:
❌ 用百分位:histogram_quantile(0.99, ...) < 0.2
→ 这是一个布尔值(达标/不达标),无法计算"消耗了多少预算"
✅ 用 threshold:P99 < 200ms 的【请求比例】> 99.9%
→ 可以直接算出"有多少请求违约" → 能计算错误预算
二、错误预算的 PromQL
-- ═══════════════════════════════════════════════════════════
-- ① 可用性 SLI(当前值)
-- ═══════════════════════════════════════════════════════════
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))
/ sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))
-- ═══════════════════════════════════════════════════════════
-- ② 错误预算剩余(30 天窗口,可用性)
-- ═══════════════════════════════════════════════════════════
1 - (
(1 - (
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[30d]))
/ sum(rate(http_requests_total{uri="/orders/{id}"}[30d]))
)) / (1 - 0.999)
)
-- ═══════════════════════════════════════════════════════════
-- ③ 错误预算剩余(延迟 SLO)
-- ═══════════════════════════════════════════════════════════
1 - (
(1 - (
sum(rate(http_request_duration_seconds_bucket{uri="/orders/{id}", le="0.2"}[30d]))
/ sum(rate(http_request_duration_seconds_count{uri="/orders/{id}"}[30d]))
)) / (1 - 0.999)
)
-- ═══════════════════════════════════════════════════════════
-- ④ 燃烧率(多个窗口,用于告警)
-- ═══════════════════════════════════════════════════════════
-- 1 小时窗口
(
1 - (
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[1h]))
/ sum(rate(http_requests_total{uri="/orders/{id}"}[1h]))
)
) / (1 - 0.999)
-- 5 分钟窗口(短窗口,保证灵敏)
(
1 - (
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))
/ sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))
)
) / (1 - 0.999)
三、Prometheus 记录规则
# observability/recording-rules.yml
groups:
- name: slo-recording
interval: 30s
rules:
# ═══ 可用性 SLI ═══
- record: slo:orders_availability:ratio_5m
expr: |
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[5m]))
/ sum(rate(http_requests_total{uri="/orders/{id}"}[5m]))
- record: slo:orders_availability:ratio_1h
expr: |
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[1h]))
/ sum(rate(http_requests_total{uri="/orders/{id}"}[1h]))
- record: slo:orders_availability:ratio_30d
expr: |
sum(rate(http_requests_total{uri="/orders/{id}", status!~"5.."}[30d]))
/ sum(rate(http_requests_total{uri="/orders/{id}"}[30d]))
# ═══ 延迟 SLI(threshold 型)═══
- record: slo:orders_latency:ratio_5m
expr: |
sum(rate(http_request_duration_seconds_bucket{uri="/orders/{id}", le="0.2"}[5m]))
/ sum(rate(http_request_duration_seconds_count{uri="/orders/{id}"}[5m]))
# ═══ 错误预算剩余 ═══
- record: slo:orders_availability:budget_remaining
expr: |
1 - ((1 - slo:orders_availability:ratio_30d) / (1 - 0.999))
- record: slo:orders_latency:budget_remaining
expr: |
1 - ((1 - slo:orders_latency:ratio_30d) / (1 - 0.999))
# ═══ 燃烧率(多窗口)═══
- record: slo:orders_availability:burn_rate_5m
expr: (1 - slo:orders_availability:ratio_5m) / (1 - 0.999)
- record: slo:orders_availability:burn_rate_1h
expr: (1 - slo:orders_availability:ratio_1h) / (1 - 0.999)
四、Grafana 看板(JSON 骨架)
{
"title": "SLO 与错误预算",
"panels": [
{
"title": "错误预算剩余(30 天窗口)",
"type": "gauge",
"targets": [
{ "expr": "slo:orders_availability:budget_remaining", "legendFormat": "可用性" },
{ "expr": "slo:orders_latency:budget_remaining", "legendFormat": "延迟" }
],
"fieldConfig": {
"defaults": {
"unit": "percentunit",
"min": 0, "max": 1,
"thresholds": {
"steps": [
{ "color": "red", "value": 0 },
{ "color": "yellow", "value": 0.2 },
{ "color": "green", "value": 0.5 }
]
}
}
}
},
{
"title": "燃烧率(1h / 6h / 1d)",
"type": "timeseries",
"targets": [
{ "expr": "slo:orders_availability:burn_rate_5m", "legendFormat": "5m" },
{ "expr": "slo:orders_availability:burn_rate_1h", "legendFormat": "1h" }
]
},
{
"title": "按接口分解的错误预算消耗",
"type": "table",
"targets": [
{
"expr": "topk(5, (1 - sum by (uri) (rate(http_requests_total{status!~\"5..\"}[30d])) / sum by (uri) (rate(http_requests_total[30d]))) / (1 - 0.999))",
"format": "table"
}
]
},
{
"title": "饱和度(领先指标)",
"type": "timeseries",
"targets": [
{ "expr": "db_pool_pending", "legendFormat": "连接池 pending" },
{ "expr": "executor_queue_depth", "legendFormat": "队列深度" },
{ "expr": "rate(container_cpu_cfs_throttled_periods_total[5m]) / rate(container_cpu_cfs_periods_total[5m])", "legendFormat": "CPU 节流比例" }
]
}
]
}
注意第四个面板:饱和度与错误预算放在同一个看板上——因为:
✅ 正确做法:
错误预算开始消耗 → 看相邻的饱和度面板 → 立刻知道是哪个资源的问题
❌ 错误做法:
错误预算和饱和度在两个不同的看板 → 排查时要来回切换
五、SLO 定义校验脚本
# tools/validate-slos.py <slos.yaml>
"""
校验 SLO 定义是否完整、可用。
检查项:
① 每项 SLO 都有 objective / window / SLI
② 延迟类 SLO 有负载前提(precondition)
③ objective 不是 100%(不可能达成的目标会导致预算为 0)
④ 延迟用 threshold 类型而不是百分位(否则无法算预算)
"""
import sys
try:
import yaml
except ImportError:
print("需要 pyyaml:pip install pyyaml")
sys.exit(1)
def validate(path):
data = yaml.safe_load(open(path, encoding="utf-8"))
slos = data.get("slos", [])
problems = []
warnings = []
for slo in slos:
name = slo.get("name", "?")
print(f"检查 {name}...")
# ① 必需字段
for field in ("objective", "window", "sli"):
if field not in slo:
problems.append(f"{name}: 缺少 {field}")
# ② objective 合理性
obj = slo.get("objective", 0)
if obj >= 1.0:
problems.append(f"{name}: objective = {obj}(100% 不可能达成,会导致预算为 0)")
elif obj < 0.9:
warnings.append(f"{name}: objective = {obj} 偏低(低于 90% 说明目标太松)")
# ③ 延迟类 SLO 必须有负载前提
sli = slo.get("sli", {})
if sli.get("type") == "threshold" and "precondition" not in slo:
warnings.append(f"{name}: 延迟类 SLO 缺少 precondition(低流量时 P99 不稳定)")
# ④ 延迟 SLI 应该用 threshold 而不是百分位
if "histogram_quantile" in str(sli.get("good", "")):
problems.append(f"{name}: 延迟 SLI 用了百分位 —— 无法计算错误预算,应改用 threshold 类型")
# ⑤ 窗口合理性
window = slo.get("window", "")
if window and not any(window.endswith(u) for u in ("d", "w")):
warnings.append(f"{name}: window = {window}(建议用 30d 或更长的窗口)")
print()
print("═" * 70)
if problems:
print(f"❌ {len(problems)} 个问题:")
for p in problems:
print(f" - {p}")
if warnings:
print(f"⚠️ {len(warnings)} 个警告:")
for w in warnings:
print(f" - {w}")
if not problems and not warnings:
print("✅ SLO 定义完整")
print("═" * 70)
return 1 if problems else 0
if __name__ == "__main__":
sys.exit(validate(sys.argv[1] if len(sys.argv) > 1 else "slo/slos.yaml"))
六、动手改造
| 改动 | 观察什么 |
|---|---|
用 validate-slos.py 校验你的 SLO 定义 |
看它检出多少问题(尤其"缺少负载前提") |
| 把某条 SLO 的 objective 改成 1.0 | 校验会报错——100% 的目标会让预算为 0 |
把延迟 SLI 改成用 histogram_quantile |
校验会报"无法计算错误预算" |
| 在 Grafana 里把饱和度和错误预算放同一个看板 | 排查时更快定位 |
| 故意让服务返回 5xx 一部分请求 | 观察错误预算与燃烧率的变化 |
七、这段代码的局限
slo/slos.yaml是自定义格式:需要自己写生成器把它转成 Prometheus 规则(或手工维护规则)。- 30 天窗口的 PromQL 很重:
rate(...[30d])需要扫描大量数据,生产建议用预聚合(recording rules 定时计算,看板查记录规则)。 - 错误预算的"百分比"含义依赖请求量:低流量服务(每天 100 个请求)的 0.1% 预算是 0.1 个请求——小流量服务的 SLO 要按「绝对数量」设,而不是百分比。
- 业务 SLI 需要额外埋点:
business_orders_success_total这类指标需要业务代码配合上报。 - Grafana 看板 JSON 需要按你的数据源调整:
datasource、unit、thresholds都可能需要改。