9.2 配套代码:SLO 表与延迟预算校验
对应小节:9.2 步骤一:定义目标
一、SLO 定义文件
# perf/slo.yaml —— 单一事实源:文档、门禁、告警都从这里读
service: shortlink
version: "1.0"
owner: "@backend-team"
updated: "2024-06-01"
slis:
# ── 接口级 ──
- name: redirect_latency
description: "跳转接口的端到端延迟"
metric: "http_server_requests_seconds"
labels: { uri: "/{code}", method: GET }
unit: ms
- name: create_latency
description: "创建接口的端到端延迟"
metric: "http_server_requests_seconds"
labels: { uri: "/links", method: POST }
unit: ms
- name: redirect_availability
description: "跳转接口的成功率"
metric: "app_requests_failed / app.requests.total"
labels: { uri: "/{code}" }
unit: ratio
- name: dependency_latency
description: "数据库查询延迟(依赖级)"
metric: "app_db_query_seconds"
unit: ms
# ── 关键:错误预算的分子要排除"正常业务失败" ──
exclusions:
- "短码冲突(业务正常,重试成功)"
- "用户输入的非法 URL(归 4xx,不计失败)"
slos:
- sli: redirect_latency
objective: "P99 < 100ms"
window: "28d"
error_budget: "1% 的请求可超 100ms"
- sli: create_latency
objective: "P99 < 200ms" # 写路径预算更宽(有 INSERT)
window: "28d"
error_budget: "1%"
- sli: redirect_availability
objective: ">= 99.9%"
window: "28d"
error_budget: "0.1% = 43.2min/月"
- sli: dependency_latency
objective: "P99 < 60ms" # 依赖预算 = 端到端预算的一部分
window: "28d"
budget:
latency_budget:
total_ms: 100
allocations:
- { component: "网关/负载均衡", ms: 5, type: serial }
- { component: "Ktor 路由与序列化", ms: 5, type: serial }
- { component: "业务逻辑", ms: 10, type: serial }
- { component: "数据库查询", ms: 60, type: critical, parallel_with: ["缓存查询"] }
- { component: "缓存查询", ms: 2, type: parallel }
- { component: "余量", ms: 18, type: reserve }
notes: |
数据库与缓存是【互斥】的(缓存未命中才查库),
所以最大路径 = 串行部分 + max(数据库, 缓存) = 80 + 60 = 140ms?❌
不对 —— 串行部分只有 网关+路由+业务 = 20ms,
加上 max(60, 2) = 60ms,总计 80ms,留 20ms 余量 ✅
二、延迟预算校验器
#!/usr/bin/env python3
"""tools/validate-budget.py —— 校验延迟预算是否自洽
只做三件事:
1. 串行项相加 + 并行项取最大 <= 总预算
2. 每项都有明确的 type(serial/parallel/critical/reserve)
3. 余量 >= 20%(否则没有容错空间)
"""
import sys
import yaml
def validate(path: str) -> int:
with open(path) as f:
slo = yaml.safe_load(f)
budget = slo["budget"]["latency_budget"]
total = budget["total_ms"]
allocs = budget["allocations"]
errors = []
# ── ① 每项必须声明 type ──
for a in allocs:
if "type" not in a:
errors.append(f"「{a['component']}」缺少 type 字段")
serial = [a for a in allocs if a["type"] in ("serial", "critical")]
parallel = [a for a in allocs if a["type"] == "parallel"]
reserve = [a for a in allocs if a["type"] == "reserve"]
serial_sum = sum(a["ms"] for a in serial)
parallel_max = max((a["ms"] for a in parallel), default=0)
# ── ② 并行项必须真的并行(不能有依赖关系)──
# 简化:只检查是否显式声明了 parallel_with
for a in parallel:
if "parallel_with" not in a:
errors.append(
f"「{a['component']}」标为 parallel 但未声明 parallel_with —— "
f"无法证明它真的并行"
)
worst_path = serial_sum + parallel_max
reserve_ms = sum(a["ms"] for a in reserve)
reserve_pct = reserve_ms / total * 100
# ── ③ 最坏路径 + 余量 <= 总预算 ──
if worst_path + reserve_ms > total:
errors.append(
f"预算超支:最坏路径 {worst_path}ms + 余量 {reserve_ms}ms "
f"= {worst_path + reserve_ms}ms > 总预算 {total}ms"
)
# ── ④ 余量比例 ──
if reserve_pct < 20:
errors.append(
f"余量只有 {reserve_pct:.0f}%(< 20%)—— "
f"没有容错空间,任何一项超支都会击穿 SLO"
)
# ── 输出 ──
print(f"总预算:{total}ms")
print(f"串行路径:{serial_sum}ms " + " + ".join(f"{a['ms']}" for a in serial))
print(f"并行取最大:{parallel_max}ms " + ", ".join(a["component"] for a in parallel))
print(f"最坏路径:{worst_path}ms")
print(f"余量:{reserve_ms}ms ({reserve_pct:.0f}%)")
print()
if errors:
print("❌ 预算校验失败:")
for e in errors:
print(f" • {e}")
return 1
print("✅ 预算自洽")
print()
print(f"解读:最坏路径 {worst_path}ms,留 {total - worst_path}ms 余量")
print(f" → 只要依赖不超过各自预算,端到端不会击穿 {total}ms")
return 0
if __name__ == "__main__":
sys.exit(validate(sys.argv[1] if len(sys.argv) > 1 else "perf/slo.yaml"))
三、把 SLO 变成压测阈值
#!/usr/bin/env bash
# tools/gen-k6-thresholds.sh —— 从 slo.yaml 生成 k6 的 thresholds 片段
# 为什么这么做:阈值和 SLO 必须是【同一个数字】,否则会漂移
set -euo pipefail
SLO="${1:-perf/slo.yaml}"
OUT="perf/thresholds.js"
python3 - "$SLO" > "$OUT" <<'PY'
import sys, yaml
with open(sys.argv[1]) as f:
slo = yaml.safe_load(f)
print("// 本文件由 tools/gen-k6-thresholds.sh 自动生成 —— 请勿手改")
print("// 改 SLO 请改 perf/slo.yaml,然后重跑生成脚本")
print("export const thresholds = {")
for s in slo["slos"]:
sli = s["sli"]
obj = s["objective"] # 例如 "P99 < 100ms"
if "P99 <" in obj:
ms = obj.split("<")[1].strip().rstrip("ms")
# 门禁阈值 = SLO 阈值 × 1.0(不做额外收紧)
# 为什么不像常规那样收紧 1.5 倍:这里是在【验证 SLO】,
# 不是在【卡回归】——回归门禁见第 9.9 节
print(f" '{sli}': ['p(99)<{ms}'],")
print("};")
PY
echo "✅ 已生成 $OUT"
echo
cat "$OUT"
生成的片段:
// perf/thresholds.js
export const thresholds = {
redirect_latency: ['p(99)<100'],
create_latency: ['p(99)<200'],
redirect_availability: ['rate>0.999'],
dependency_latency: ['p(99)<60'],
};
四、目标检查清单(自动核对)
#!/usr/bin/env python3
"""tools/check-goals.py —— 检查步骤一的四项交付物是否齐全
步骤一的产出物必须是【可被机器检查】的,否则它只是一段愿望。
"""
import sys
import yaml
REQUIRED = {
"slis": "至少一个 SLI 定义",
"slos": "至少一个 SLO(含 objective 与 window)",
"budget": "延迟预算",
"exclusions": "错误预算的排除项",
}
def main(path: str) -> int:
with open(path) as f:
slo = yaml.safe_load(f)
ok = True
print("═══ 步骤一交付物检查 ═══\n")
for key, desc in REQUIRED.items():
if key in slo and slo[key]:
print(f"✅ {key:12s} {desc}")
else:
print(f"❌ {key:12s} {desc} ← 缺失")
ok = False
# ── SLO 必须写清 window(否则"28 天"和"1 小时"差别巨大)──
print()
for s in slo.get("slos", []):
missing = [k for k in ("sli", "objective", "window") if k not in s]
if missing:
print(f"❌ SLO「{s.get('sli', '?')}」缺少:{', '.join(missing)}")
ok = False
else:
print(f"✅ SLO「{s['sli']}」{s['objective']} @ {s['window']}")
print()
print("═══════════════════════════════")
print("✅ 步骤一完成" if ok else "❌ 步骤一未完成 —— 不要往下走")
if not ok:
print()
print("为什么不能跳过:")
print(" 没有 SLO → 后面的「优化成功」无法定义")
print(" 没有预算 → 优化时不知道该改哪一段")
return 0 if ok else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1] if len(sys.argv) > 1 else "perf/slo.yaml"))
五、动手改造
| 改动 | 观察什么 |
|---|---|
| 把数据库预算从 60ms 改成 100ms | 校验器报"预算超支"——理解预算是硬约束 |
把"缓存查询"的 parallel_with 删掉 |
校验器报错——并行必须可证明,不能只是声称 |
| 把余量从 18ms 改成 5ms | 校验器报"余量 < 20%"——理解余量的作用 |
| 把 SLO 的 P99 从 100ms 改成 90ms | 生成的门禁阈值随之变——单一事实源生效 |
给 create_latency 也加 dependency_latency |
会发现它写路径的依赖预算没定义 → 暴露出遗漏 |
六、这段代码的局限
validate-budget.py不能验证"预算是否真实":它只检查自洽性。真实值要靠测量(步骤五的延迟分解)。exclusions只是文档:没有代码强制 k6 按它计算错误率——k6 需要写 tag 逻辑(见code/01-metrics-and-slo/04-writing-slo.md)。- 并行项检查很弱:只检查有没有声明
parallel_with,不检查实际调用链——要真正验证得看 trace。 - 没有多窗口 SLO(比如同时定义 1h 和 28d)——生产常见,这里简化。
- 门禁阈值直接用 SLO:实际上门禁应该比 SLO 更严(在 SLO 被击穿前就拦住)——第 9.9 节会处理。