6.1 配套代码:分析工作流与假设台账
对应小节:6.1 分析流程五步法 两个工具:① 现象量化表模板;② 假设台账管理(让每个假设都有验证状态)。
一、现象量化表模板
<!-- docs/experiments/E07-incident-xxx/README.md 的「现象量化」章节 -->
## 现象量化
### 基本信息
| 维度 | 观测值 |
| --- | --- |
| 发现时间 | YYYY-MM-DD HH:MM |
| 报告人 | |
| 影响 | 用户投诉 / 监控告警 / 例行检查发现 |
### 谁慢
| 维度 | 观测值 |
| --- | --- |
| 接口 | |
| 实例范围 | 全部 / 部分(哪几台) |
| 请求范围 | 全部 / 特定参数 / 特定用户 |
| 依赖范围 | 是否与某个下游相关 |
### 多慢
| 指标 | 基线 | 当前 | 变化 |
| --- | --- | --- | --- |
| P50 | | | |
| P95 | | | |
| P99 | | | |
| 错误率 | | | |
| 超时率 | | | |
| QPS | | | |
### 从何时开始
| 问题 | 答案 |
| --- | --- |
| 突变还是渐变? | |
| 与某次发布吻合? | (版本号 + 时间) |
| 与数据量增长吻合? | |
| 与外部事件吻合? | (大促、其他服务上线) |
### 资源与饱和度(用于排除)
| 指标 | 基线 | 当前 | 是否异常 |
| --- | --- | --- | --- |
| CPU(user) | | | |
| CPU(sys) | | | |
| 内存 / 堆 | | | |
| GC 停顿 P99 | | | |
| 线程数 | | | |
| 连接池 pending | | | |
| 队列深度 | | | |
| 容器节流比例 | | | |
### 初步推论(由上面的数据排除/指向什么)
- QPS 未变 + CPU 未变 → 排除「过载」
- 只在特定参数上慢 → 指向数据访问路径
- 与发布吻合 → 高度怀疑该变更
为什么这张表这么重要:它把「接口慢了」这个模糊现象变成了一组可以排除假设的事实。填完它,假设空间通常能缩小到 2–3 个。
二、假设台账管理
# tools/hypotheses.py — 假设台账管理
"""
用法:
python3 tools/hypotheses.py init <EXP_ID>
python3 tools/hypotheses.py add <EXP_ID> "假设内容" "验证命令" "否证条件"
python3 tools/hypotheses.py update <EXP_ID> <H编号> <状态> [证据]
python3 tools/hypotheses.py list <EXP_ID>
python3 tools/hypotheses.py report <EXP_ID>
"""
import json
import pathlib
import sys
from datetime import datetime
STATUSES = ["待验证", "已验证成立", "已排除", "暂缓"]
def ledger_path(exp_id):
d = pathlib.Path("docs/experiments") / exp_id
d.mkdir(parents=True, exist_ok=True)
return d / "hypotheses.json"
def load(exp_id):
p = ledger_path(exp_id)
if p.exists():
return json.loads(p.read_text(encoding="utf-8"))
return {"exp_id": exp_id, "hypotheses": []}
def save(exp_id, data):
ledger_path(exp_id).write_text(
json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")
def cmd_init(exp_id):
save(exp_id, {"exp_id": exp_id, "hypotheses": []})
print(f"✅ 已初始化假设台账 → {ledger_path(exp_id)}")
def cmd_add(exp_id, text, command, falsify):
data = load(exp_id)
hid = f"H{len(data['hypotheses']) + 1}"
data["hypotheses"].append({
"id": hid,
"hypothesis": text,
"verify_command": command,
"falsify_condition": falsify,
"status": "待验证",
"evidence": "",
"updated_at": datetime.now().isoformat(timespec="seconds"),
})
save(exp_id, data)
print(f"✅ 已添加 {hid}: {text}")
def cmd_update(exp_id, hid, status, evidence=""):
if status not in STATUSES:
print(f"❌ 状态必须是:{'/'.join(STATUSES)}")
return
data = load(exp_id)
for h in data["hypotheses"]:
if h["id"] == hid:
h["status"] = status
h["evidence"] = evidence
h["updated_at"] = datetime.now().isoformat(timespec="seconds")
save(exp_id, data)
print(f"✅ {hid} → {status}")
return
print(f"❌ 找不到 {hid}")
def cmd_list(exp_id):
data = load(exp_id)
if not data["hypotheses"]:
print("(台账为空)")
return
print(f"{'ID':<5}{'状态':<12}{'假设'}")
print("-" * 78)
for h in data["hypotheses"]:
print(f"{h['id']:<5}{h['status']:<12}{h['hypothesis'][:60]}")
print()
for h in data["hypotheses"]:
if h["status"] == "待验证":
print(f"⏳ {h['id']} 待验证")
print(f" 验证命令: {h['verify_command']}")
print(f" 否证条件: {h['falsify_condition']}")
print()
def cmd_report(exp_id):
data = load(exp_id)
hs = data["hypotheses"]
print("═" * 78)
print(f"假设台账报告:{exp_id}")
print("═" * 78)
print()
by_status = {}
for h in hs:
by_status.setdefault(h["status"], []).append(h)
for status in STATUSES:
items = by_status.get(status, [])
if items:
print(f"【{status}】{len(items)} 条")
for h in items:
print(f" {h['id']} {h['hypothesis']}")
if h["evidence"]:
print(f" 证据: {h['evidence']}")
print()
# 关键提示
if not by_status.get("已排除"):
print("⚠️ 还没有任何「已排除」的假设 —— 结论的可信度会打折扣")
print(" (第 6.8 节:被排除的假设与找到的根因同样重要)")
if by_status.get("待验证"):
print(f"⚠️ 还有 {len(by_status['待验证'])} 条待验证 —— 结论可能不完整")
print()
print("结论模板:")
print(" 根因:(已验证成立的假设)")
print(" 被排除:(已排除的假设 + 依据)")
print("═" * 78)
if __name__ == "__main__":
if len(sys.argv) < 3:
print(__doc__)
sys.exit(1)
cmd, exp = sys.argv[1], sys.argv[2]
try:
if cmd == "init":
cmd_init(exp)
elif cmd == "add":
cmd_add(exp, sys.argv[3], sys.argv[4], sys.argv[5])
elif cmd == "update":
cmd_update(exp, sys.argv[3], sys.argv[4], sys.argv[5] if len(sys.argv) > 5 else "")
elif cmd == "list":
cmd_list(exp)
elif cmd == "report":
cmd_report(exp)
else:
print(__doc__)
except IndexError:
print(__doc__)
使用示例:
# 初始化
python3 tools/hypotheses.py init E07-incident-orders
# 添加假设(每条都必须有验证命令与否证条件)
python3 tools/hypotheses.py add E07-incident-orders \
"GC 停顿造成 P99 尖刺" \
"对比 gc.log 的 Pause 时间戳与 P99 尖刺时间戳" \
"若两者不对齐则排除 GC"
python3 tools/hypotheses.py add E07-incident-orders \
"连接池排队导致延迟" \
"curl -s localhost:8080/metrics | grep db_pool_pending" \
"若 pending 始终为 0 则排除"
python3 tools/hypotheses.py add E07-incident-orders \
"下游 user-service 变慢" \
"jfr print --events jdk.SocketRead recording.jfr | grep user-service" \
"若无 user-service 相关阻塞则排除"
# 验证后更新状态
python3 tools/hypotheses.py update E07-incident-orders H1 "已排除" \
"GC 停顿时间戳与尖刺不对齐(差值 > 2 秒),且停顿幅度仅 15ms"
python3 tools/hypotheses.py update E07-incident-orders H2 "已验证成立" \
"pending 从 0 涨到 18,且与 P99 上升时间吻合"
# 查看报告
python3 tools/hypotheses.py report E07-incident-orders
预期输出:
══════════════════════════════════════════════════════════════════════════════
假设台账报告:E07-incident-orders
══════════════════════════════════════════════════════════════════════════════
【已验证成立】1 条
H2 连接池排队导致延迟
证据: pending 从 0 涨到 18,且与 P99 上升时间吻合
【已排除】1 条
H1 GC 停顿造成 P99 尖刺
证据: GC 停顿时间戳与尖刺不对齐(差值 > 2 秒),且停顿幅度仅 15ms
【待验证】1 条
H3 下游 user-service 变慢
⚠️ 还有 1 条待验证 —— 结论可能不完整
结论模板:
根因:(已验证成立的假设)
被排除:(已排除的假设 + 依据)
══════════════════════════════════════════════════════════════════════════════
三、把台账接进分析流程
#!/usr/bin/env bash
# tools/analyze-incident.sh <EXP_ID>
#
# 引导式分析流程:每步都要求你产出可验证的结果。
set -uo pipefail
EXP_ID="${1:?usage: analyze-incident.sh <EXP_ID>}"
DIR="docs/experiments/${EXP_ID}"
mkdir -p "$DIR/results"
echo "═══════════════════════════════════════════════════════════════"
echo "分析流程:$EXP_ID"
echo "═══════════════════════════════════════════════════════════════"
echo
echo "【第 1 步:现象量化】"
echo " 打开 $DIR/README.md,填入「现象量化」章节(见模板)"
echo " 必须填写:谁慢、多慢、何时开始、影响范围、资源指标"
read -r -p " 完成后按回车继续..." _
echo
echo "【第 2 步:延迟分解】"
echo " 运行:tools/decompose-latency.sh $EXP_ID"
read -r -p " 完成后按回车继续..." _
echo
echo "【第 3 步:收集假设】"
echo " 根据前两步的数据,列出 3–5 个可能原因"
echo " 每条都要写清「验证命令」和「否证条件」:"
echo " python3 tools/hypotheses.py add $EXP_ID \"假设\" \"验证命令\" \"否证条件\""
read -r -p " 完成后按回车继续..." _
echo
echo "【第 4 步:逐条验证】"
python3 tools/hypotheses.py list "$EXP_ID"
echo " 对每条假设执行验证,并更新状态:"
echo " python3 tools/hypotheses.py update $EXP_ID H1 已排除 \"依据\""
read -r -p " 完成后按回车继续..." _
echo
echo "【第 5 步:写结论】"
python3 tools/hypotheses.py report "$EXP_ID"
echo " 用 tools/conclusion-template.md 的结构写结论"
echo " 必须包含:证据链、贡献占比、被排除的假设、未覆盖的场景"
echo
echo "═══════════════════════════════════════════════════════════════"
echo "✅ 流程完成"
echo "═══════════════════════════════════════════════════════════════"
四、动手改造
| 改动 | 观察什么 |
|---|---|
| 用台账管理一次真实的排查 | 体会「假设状态化」如何避免遗漏 |
| 故意不加「否证条件」 | 你会发现验证变成"看起来像就算成立"——这正是不加否证条件的后果 |
在 report 里加上贡献占比字段 |
让结论更完整 |
| 把台账接进团队的问题跟踪系统 | 从流程上保证每个假设都有结论 |
五、这段代码的局限
- 台账管理是流程工具,不是技术工具:它不能替你验证假设,只能避免遗漏。
- JSON 格式不适合人工编辑:建议只用命令行操作,或者用 Markdown 表格(更好读但不好解析)。
analyze-incident.sh是引导式的:交互步骤多,适合学习阶段;熟练后可以直接跳到具体命令。- 假设的质量取决于人:「验证命令」和「否证条件」写得好不好,取决于你对系统的理解——工具无法提升这一点。