文档目录

6.1 配套代码:分析工作流与假设台账

对应小节:6.1 分析流程五步法 两个工具:① 现象量化表模板;② 假设台账管理(让每个假设都有验证状态)。

一、现象量化表模板

<!-- docs/experiments/E07-incident-xxx/README.md 的「现象量化」章节 -->

## 现象量化

### 基本信息
| 维度 | 观测值 |
| --- | --- |
| 发现时间 | YYYY-MM-DD HH:MM |
| 报告人 | |
| 影响 | 用户投诉 / 监控告警 / 例行检查发现 |

### 谁慢
| 维度 | 观测值 |
| --- | --- |
| 接口 | |
| 实例范围 | 全部 / 部分(哪几台) |
| 请求范围 | 全部 / 特定参数 / 特定用户 |
| 依赖范围 | 是否与某个下游相关 |

### 多慢
| 指标 | 基线 | 当前 | 变化 |
| --- | --- | --- | --- |
| P50 | | | |
| P95 | | | |
| P99 | | | |
| 错误率 | | | |
| 超时率 | | | |
| QPS | | | |

### 从何时开始
| 问题 | 答案 |
| --- | --- |
| 突变还是渐变? | |
| 与某次发布吻合? | (版本号 + 时间) |
| 与数据量增长吻合? | |
| 与外部事件吻合? | (大促、其他服务上线) |

### 资源与饱和度(用于排除)
| 指标 | 基线 | 当前 | 是否异常 |
| --- | --- | --- | --- |
| CPU(user) | | | |
| CPU(sys) | | | |
| 内存 / 堆 | | | |
| GC 停顿 P99 | | | |
| 线程数 | | | |
| 连接池 pending | | | |
| 队列深度 | | | |
| 容器节流比例 | | | |

### 初步推论(由上面的数据排除/指向什么)
- QPS 未变 + CPU 未变 → 排除「过载」
- 只在特定参数上慢 → 指向数据访问路径
- 与发布吻合 → 高度怀疑该变更

为什么这张表这么重要:它把「接口慢了」这个模糊现象变成了一组可以排除假设的事实。填完它,假设空间通常能缩小到 2–3 个。

二、假设台账管理

# tools/hypotheses.py — 假设台账管理
"""
用法:
    python3 tools/hypotheses.py init <EXP_ID>
    python3 tools/hypotheses.py add <EXP_ID> "假设内容" "验证命令" "否证条件"
    python3 tools/hypotheses.py update <EXP_ID> <H编号> <状态> [证据]
    python3 tools/hypotheses.py list <EXP_ID>
    python3 tools/hypotheses.py report <EXP_ID>
"""
import json
import pathlib
import sys
from datetime import datetime

STATUSES = ["待验证", "已验证成立", "已排除", "暂缓"]


def ledger_path(exp_id):
    d = pathlib.Path("docs/experiments") / exp_id
    d.mkdir(parents=True, exist_ok=True)
    return d / "hypotheses.json"


def load(exp_id):
    p = ledger_path(exp_id)
    if p.exists():
        return json.loads(p.read_text(encoding="utf-8"))
    return {"exp_id": exp_id, "hypotheses": []}


def save(exp_id, data):
    ledger_path(exp_id).write_text(
        json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")


def cmd_init(exp_id):
    save(exp_id, {"exp_id": exp_id, "hypotheses": []})
    print(f"✅ 已初始化假设台账 → {ledger_path(exp_id)}")


def cmd_add(exp_id, text, command, falsify):
    data = load(exp_id)
    hid = f"H{len(data['hypotheses']) + 1}"
    data["hypotheses"].append({
        "id": hid,
        "hypothesis": text,
        "verify_command": command,
        "falsify_condition": falsify,
        "status": "待验证",
        "evidence": "",
        "updated_at": datetime.now().isoformat(timespec="seconds"),
    })
    save(exp_id, data)
    print(f"✅ 已添加 {hid}: {text}")


def cmd_update(exp_id, hid, status, evidence=""):
    if status not in STATUSES:
        print(f"❌ 状态必须是:{'/'.join(STATUSES)}")
        return
    data = load(exp_id)
    for h in data["hypotheses"]:
        if h["id"] == hid:
            h["status"] = status
            h["evidence"] = evidence
            h["updated_at"] = datetime.now().isoformat(timespec="seconds")
            save(exp_id, data)
            print(f"✅ {hid} → {status}")
            return
    print(f"❌ 找不到 {hid}")


def cmd_list(exp_id):
    data = load(exp_id)
    if not data["hypotheses"]:
        print("(台账为空)")
        return
    print(f"{'ID':<5}{'状态':<12}{'假设'}")
    print("-" * 78)
    for h in data["hypotheses"]:
        print(f"{h['id']:<5}{h['status']:<12}{h['hypothesis'][:60]}")
    print()
    for h in data["hypotheses"]:
        if h["status"] == "待验证":
            print(f"⏳ {h['id']} 待验证")
            print(f"   验证命令: {h['verify_command']}")
            print(f"   否证条件: {h['falsify_condition']}")
            print()


def cmd_report(exp_id):
    data = load(exp_id)
    hs = data["hypotheses"]
    print("═" * 78)
    print(f"假设台账报告:{exp_id}")
    print("═" * 78)
    print()

    by_status = {}
    for h in hs:
        by_status.setdefault(h["status"], []).append(h)

    for status in STATUSES:
        items = by_status.get(status, [])
        if items:
            print(f"【{status}】{len(items)} 条")
            for h in items:
                print(f"  {h['id']}  {h['hypothesis']}")
                if h["evidence"]:
                    print(f"       证据: {h['evidence']}")
            print()

    # 关键提示
    if not by_status.get("已排除"):
        print("⚠️  还没有任何「已排除」的假设 —— 结论的可信度会打折扣")
        print("   (第 6.8 节:被排除的假设与找到的根因同样重要)")
    if by_status.get("待验证"):
        print(f"⚠️  还有 {len(by_status['待验证'])} 条待验证 —— 结论可能不完整")
    print()
    print("结论模板:")
    print("  根因:(已验证成立的假设)")
    print("  被排除:(已排除的假设 + 依据)")
    print("═" * 78)


if __name__ == "__main__":
    if len(sys.argv) < 3:
        print(__doc__)
        sys.exit(1)
    cmd, exp = sys.argv[1], sys.argv[2]
    try:
        if cmd == "init":
            cmd_init(exp)
        elif cmd == "add":
            cmd_add(exp, sys.argv[3], sys.argv[4], sys.argv[5])
        elif cmd == "update":
            cmd_update(exp, sys.argv[3], sys.argv[4], sys.argv[5] if len(sys.argv) > 5 else "")
        elif cmd == "list":
            cmd_list(exp)
        elif cmd == "report":
            cmd_report(exp)
        else:
            print(__doc__)
    except IndexError:
        print(__doc__)

使用示例:

# 初始化
python3 tools/hypotheses.py init E07-incident-orders

# 添加假设(每条都必须有验证命令与否证条件)
python3 tools/hypotheses.py add E07-incident-orders \
  "GC 停顿造成 P99 尖刺" \
  "对比 gc.log 的 Pause 时间戳与 P99 尖刺时间戳" \
  "若两者不对齐则排除 GC"

python3 tools/hypotheses.py add E07-incident-orders \
  "连接池排队导致延迟" \
  "curl -s localhost:8080/metrics | grep db_pool_pending" \
  "若 pending 始终为 0 则排除"

python3 tools/hypotheses.py add E07-incident-orders \
  "下游 user-service 变慢" \
  "jfr print --events jdk.SocketRead recording.jfr | grep user-service" \
  "若无 user-service 相关阻塞则排除"

# 验证后更新状态
python3 tools/hypotheses.py update E07-incident-orders H1 "已排除" \
  "GC 停顿时间戳与尖刺不对齐(差值 > 2 秒),且停顿幅度仅 15ms"
python3 tools/hypotheses.py update E07-incident-orders H2 "已验证成立" \
  "pending 从 0 涨到 18,且与 P99 上升时间吻合"

# 查看报告
python3 tools/hypotheses.py report E07-incident-orders

预期输出:

══════════════════════════════════════════════════════════════════════════════
假设台账报告:E07-incident-orders
══════════════════════════════════════════════════════════════════════════════

【已验证成立】1 条
  H2  连接池排队导致延迟
       证据: pending 从 0 涨到 18,且与 P99 上升时间吻合

【已排除】1 条
  H1  GC 停顿造成 P99 尖刺
       证据: GC 停顿时间戳与尖刺不对齐(差值 > 2 秒),且停顿幅度仅 15ms

【待验证】1 条
  H3  下游 user-service 变慢

⚠️  还有 1 条待验证 —— 结论可能不完整

结论模板:
  根因:(已验证成立的假设)
  被排除:(已排除的假设 + 依据)
══════════════════════════════════════════════════════════════════════════════

三、把台账接进分析流程

#!/usr/bin/env bash
# tools/analyze-incident.sh <EXP_ID>
#
# 引导式分析流程:每步都要求你产出可验证的结果。
set -uo pipefail

EXP_ID="${1:?usage: analyze-incident.sh <EXP_ID>}"
DIR="docs/experiments/${EXP_ID}"
mkdir -p "$DIR/results"

echo "═══════════════════════════════════════════════════════════════"
echo "分析流程:$EXP_ID"
echo "═══════════════════════════════════════════════════════════════"
echo

echo "【第 1 步:现象量化】"
echo "  打开 $DIR/README.md,填入「现象量化」章节(见模板)"
echo "  必须填写:谁慢、多慢、何时开始、影响范围、资源指标"
read -r -p "  完成后按回车继续..." _

echo
echo "【第 2 步:延迟分解】"
echo "  运行:tools/decompose-latency.sh $EXP_ID"
read -r -p "  完成后按回车继续..." _

echo
echo "【第 3 步:收集假设】"
echo "  根据前两步的数据,列出 3–5 个可能原因"
echo "  每条都要写清「验证命令」和「否证条件」:"
echo "    python3 tools/hypotheses.py add $EXP_ID \"假设\" \"验证命令\" \"否证条件\""
read -r -p "  完成后按回车继续..." _

echo
echo "【第 4 步:逐条验证】"
python3 tools/hypotheses.py list "$EXP_ID"
echo "  对每条假设执行验证,并更新状态:"
echo "    python3 tools/hypotheses.py update $EXP_ID H1 已排除 \"依据\""
read -r -p "  完成后按回车继续..." _

echo
echo "【第 5 步:写结论】"
python3 tools/hypotheses.py report "$EXP_ID"
echo "  用 tools/conclusion-template.md 的结构写结论"
echo "  必须包含:证据链、贡献占比、被排除的假设、未覆盖的场景"

echo
echo "═══════════════════════════════════════════════════════════════"
echo "✅ 流程完成"
echo "═══════════════════════════════════════════════════════════════"

四、动手改造

改动 观察什么
用台账管理一次真实的排查 体会「假设状态化」如何避免遗漏
故意不加「否证条件」 你会发现验证变成"看起来像就算成立"——这正是不加否证条件的后果
在 report 里加上贡献占比字段 让结论更完整
把台账接进团队的问题跟踪系统 从流程上保证每个假设都有结论

五、这段代码的局限

  • 台账管理是流程工具,不是技术工具:它不能替你验证假设,只能避免遗漏。
  • JSON 格式不适合人工编辑:建议只用命令行操作,或者用 Markdown 表格(更好读但不好解析)。
  • analyze-incident.sh 是引导式的:交互步骤多,适合学习阶段;熟练后可以直接跳到具体命令。
  • 假设的质量取决于人:「验证命令」和「否证条件」写得好不好,取决于你对系统的理解——工具无法提升这一点。