文档目录

4.9 配套代码:Lab 4 观测回路的编排与验收

对应小节:4.9 Lab 4 把本章的工具串成一条回路,并自动化验证。

一、观测栈(最小的 Compose 配置)

# docker-compose.observability.yml
services:
  prometheus:
    image: prom/prometheus:v2.54.1
    ports: ["9090:9090"]
    volumes:
      - ./observability/prometheus.yml:/etc/prometheus/prometheus.yml:ro
      - ./observability/alerts.yml:/etc/prometheus/alerts.yml:ro
      - prometheus-data:/prometheus
    command:
      - --config.file=/etc/prometheus/prometheus.yml
      - --storage.tsdb.retention.time=15d
      - --web.enable-lifecycle

  grafana:
    image: grafana/grafana:11.2.0
    ports: ["3000:3000"]
    environment:
      GF_SECURITY_ADMIN_PASSWORD: admin
      GF_USERS_ALLOW_SIGN_UP: "false"
    volumes:
      - grafana-data:/var/lib/grafana
    depends_on: [prometheus]

volumes:
  prometheus-data:
  grafana-data:
# observability/prometheus.yml
global:
  scrape_interval: 5s          # 压测期间用 5s;生产可放宽到 15~30s
  evaluation_interval: 5s

rule_files:
  - /etc/prometheus/alerts.yml

scrape_configs:
  - job_name: kotlin-app
    metrics_path: /metrics
    static_configs:
      # macOS/Windows 用 host.docker.internal;Linux 用宿主机 IP
      - targets: ["host.docker.internal:8080"]
# observability/alerts.yml —— 本章涉及的四个关键告警
groups:
  - name: saturation
    rules:
      - alert: ConnectionPoolSaturated
        expr: db_pool_pending > 0
        for: 5m
        labels: { severity: ticket }
        annotations:
          summary: "连接池排队(pending > 0 持续 5m)"
          runbook: "① 查慢查询(pg_stat_statements)② 查长事务 ③ 最后才考虑调池(第 4.7 节)"

      - alert: ContainerCpuThrottled
        expr: |
          rate(container_cpu_cfs_throttled_periods_total[5m])
          / rate(container_cpu_cfs_periods_total[5m]) > 0.25
        for: 10m
        labels: { severity: ticket }
        annotations:
          summary: "容器被 CPU 配额节流 > 25%(第 4.8 节)"
          runbook: "① 确认 limit ② 看 CPU 需求是否合理 ③ 优化代码或提 limit"

      - alert: GcPauseHigh
        expr: |
          histogram_quantile(0.99,
            sum by (le) (rate(jvm_gc_pause_seconds_bucket[5m]))) > 0.2
        for: 10m
        labels: { severity: ticket }
        annotations:
          summary: "GC 停顿 P99 > 200ms"
          runbook: "① 与 P99 尖刺对齐验证 ② 采 alloc 火焰图 ③ 检查堆与 GC 选型"

      - alert: MetricCardinalityExplosion
        expr: sum(count by (__name__) ({__name__=~".+"})) > 500000
        for: 10m
        labels: { severity: ticket }
        annotations:
          summary: "Prometheus 序列数超过 50 万,可能存在基数泄漏(第 4.6 节)"

二、一键跑完 Lab 4

#!/usr/bin/env bash
# tools/run-lab4.sh <EXP_ID>
set -euo pipefail

EXP_ID="${1:?usage: run-lab4.sh <EXP_ID>}"
DIR="docs/experiments/${EXP_ID}"
mkdir -p "$DIR/results" "$DIR/scripts"

echo "═══ 步骤 0:起观测栈 ═══"
docker compose -f docker-compose.observability.yml up -d
sleep 5
curl -sf http://localhost:9090/-/ready > /dev/null && echo "✅ Prometheus 就绪" || echo "⚠️  Prometheus 未就绪"
echo

echo "═══ 步骤 1:环境元数据 ═══"
tools/collect-env.sh "$EXP_ID"
tools/check-throttling.sh 5 | tee "$DIR/results/throttling-before.txt"
echo

echo "═══ 步骤 2:验证四层指标 ═══"
M=$(curl -s http://127.0.0.1:8080/metrics)
check() { echo "$M" | grep -qE "$1" && echo "  ✅ $2" || echo "  ❌ $2(缺这组指标)"; }
check "^app_requests_total"                "① 业务层"
check "^app_request_duration.*bucket"      "② 延迟层(直方图)"
check "^jvm_memory_used_bytes"             "③ 资源层"
check "^db_pool_pending"                   "④ 饱和度层"
echo "  直方图桶数量: $(echo "$M" | grep -c 'app_request_duration_seconds_bucket')"
echo

echo "═══ 步骤 3:启动服务并预热 ═══"
scripts/restart-app.sh "$DIR/results/gc.log"
sleep 5
BASE_URL=http://127.0.0.1:8080 k6 run --quiet --vus 20 --duration 60s loadtest/profile-constant.js > /dev/null
echo "✅ 预热完成"
echo

echo "═══ 步骤 4:压测 + 采四类火焰图(并行)═══"
# 后台启动压测
BASE_URL=http://127.0.0.1:8080 RATE=400 DURATION=8m \
k6 run --out json="$DIR/results/k6-raw.json" \
       --summary-export="$DIR/results/k6-summary.json" \
       loadtest/profile-constant.js > "$DIR/results/k6-stdout.txt" 2>&1 &
K6_PID=$!

sleep 60   # 等压测进入稳态

PID=$(jcmd | grep app.jar | awk '{print $1}')
echo "采集中(PID=$PID)..."
tools/profile-all.sh "$EXP_ID" 30 | tee "$DIR/results/profile-log.txt"

echo
echo "═══ 步骤 5:JFR 录制 ═══"
jcmd "$PID" JFR.start name=lab4 settings=profile duration=60s filename="$DIR/results/lab4.jfr"
echo "等待 60 秒..."
sleep 65
jfr summary "$DIR/results/lab4.jfr" > "$DIR/results/jfr-summary.txt" 2>&1
echo "✅ JFR 摘要:" && head -20 "$DIR/results/jfr-summary.txt"
echo

echo "═══ 步骤 6:等压测结束 ═══"
wait $K6_PID || true
echo

echo "═══ 步骤 7:数据可用性检查 ═══"
tools/check-load-generator.sh "$DIR/results/k6-summary.json" 400
echo

echo "═══ 步骤 8:节流复查 ═══"
tools/check-throttling.sh 5 | tee "$DIR/results/throttling-after.txt"
echo

echo "✅ Lab 4 完成 → $DIR"
echo
echo "接下来(手工):"
echo "  1. 打开 Grafana http://localhost:3000(admin/admin)看面板"
echo "  2. 打开 $DIR/results/{cpu,wall,alloc,lock}.html 填对比表"
echo "  3. 用 JMC 打开 $DIR/results/lab4.jfr 找一条火焰图没给出的线索"
echo "  4. 填写 $DIR/README.md"

三、Grafana 面板的最小配置

如果不想手工建图表,可以用一个精简的 dashboard JSON(关键的四类面板):

{
  "title": "Lab 4 · 观测回路",
  "panels": [
    {
      "title": "① 到达率 vs 服务端 QPS",
      "targets": [
        { "expr": "sum(rate(http_server_requests_seconds_count[1m]))", "legendFormat": "服务端 QPS" }
      ]
    },
    {
      "title": "② 延迟分布(P50/P95/P99/P999)",
      "targets": [
        { "expr": "histogram_quantile(0.50, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P50" },
        { "expr": "histogram_quantile(0.95, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P95" },
        { "expr": "histogram_quantile(0.99, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P99" },
        { "expr": "histogram_quantile(0.999, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P999" }
      ]
    },
    {
      "title": "③ 饱和度(连接池 + 队列)",
      "targets": [
        { "expr": "db_pool_pending", "legendFormat": "连接池 pending" },
        { "expr": "db_pool_active", "legendFormat": "连接池 active" },
        { "expr": "executor_queue_depth", "legendFormat": "队列深度" }
      ]
    },
    {
      "title": "④ 资源(CPU + GC 停顿)",
      "targets": [
        { "expr": "process_cpu_usage", "legendFormat": "进程 CPU" },
        { "expr": "histogram_quantile(0.99, sum by (le) (rate(jvm_gc_pause_seconds_bucket[5m])))", "legendFormat": "GC 停顿 P99" }
      ]
    }
  ]
}

四、Lab 4 的验收脚本

#!/usr/bin/env bash
# tools/verify-lab4.sh <EXP_ID>
#
# 自动检查 Lab 4 的产出是否齐全。
set -uo pipefail

EXP_ID="${1:?usage: verify-lab4.sh <EXP_ID>}"
DIR="docs/experiments/${EXP_ID}"
PASS=0; FAIL=0

ok()   { echo "  ✅ $1"; PASS=$((PASS+1)); }
bad()  { echo "  ❌ $1"; FAIL=$((FAIL+1)); }

echo "═══ Lab 4 验收 ═══"
echo

echo "① 环境元数据"
[ -f "$DIR/results/env.txt" ] && ok "env.txt 存在" || bad "缺 env.txt"
echo

echo "② 四层指标(从压测前的检查记录里找)"
[ -f "$DIR/results/profile-log.txt" ] && ok "profiling 日志存在" || bad "缺 profile-log.txt"
echo

echo "③ k6 数据"
[ -f "$DIR/results/k6-summary.json" ] && ok "k6-summary.json 存在" || bad "缺 k6-summary.json"
if [ -f "$DIR/results/k6-summary.json" ]; then
  python3 - "$DIR/results/k6-summary.json" <<'PY'
import json, sys
m = json.load(open(sys.argv[1]))["metrics"]
tw = m["http_reqs"]["rate"]
print(f"     到达率 {tw:.1f} req/s, 错误率 {m['http_req_failed']['rate']:.2%}")
print(f"     P50 {m['http_req_duration']['med']:.1f} ms, P99 {m['http_req_duration']['p(99)']:.1f} ms")
PY
fi
echo

echo "④ 四类火焰图"
for E in cpu wall alloc lock; do
  [ -f "$DIR/results/$E.html" ] && ok "$E.html 存在 ($(du -h "$DIR/results/$E.html" | cut -f1))" \
                                || bad "缺 $E.html"
done
echo

echo "⑤ JFR"
[ -f "$DIR/results/lab4.jfr" ] && ok "lab4.jfr 存在 ($(du -h "$DIR/results/lab4.jfr" | cut -f1))" \
                              || bad "缺 lab4.jfr"
echo

echo "⑥ 容器节流记录"
[ -f "$DIR/results/throttling-after.txt" ] && ok "节流记录存在" || bad "缺节流记录"
echo

echo "⑦ 实验档案"
[ -f "$DIR/README.md" ] && ok "README.md 存在" || bad "缺 README.md(用 _TEMPLATE 创建)"
echo

echo "═══════════════════════"
echo "通过 $PASS 项,失败 $FAIL 项"
[ "$FAIL" -eq 0 ] && echo "✅ Lab 4 产出齐全" || echo "❌ 还有 $FAIL 项待补"

五、实验档案模板

<!-- docs/experiments/E04-observability-loop/README.md -->
---
id: E04
title: 最小观测回路搭建与验证
date: 2025-xx-xx
chapter: 4
kind: probe
status: done
hypothesis: "只采 CPU 火焰图会看不到阻塞类瓶颈;wall 火焰图能看到;容器节流若存在会影响 P99"
variable: "无(工具链验证)"
control: "故意加的 50ms 阻塞接口 /slow"
commit: <hash>
jvm_args: "-Xms1g -Xmx1g -XX:StartFlightRecording=name=rolling,settings=profile,maxsize=256m,maxage=1h,dumponexit=true"
container: "cpu.max=<值>, memory.max=<值>"
tags: [toolchain, profiling, jfr, lab4]
---

## 1. 假设与预期
- CPU 火焰图:看不到 /slow 的 Thread.sleep
- wall 火焰图:能看到大量栈停在 Thread.sleep
- alloc 火焰图:能看到 JSON 序列化的分配
- 容器节流:<预期>

## 2. 环境元数据
(粘 results/env.txt + throttling-before.txt)

## 3. 原始结果
- k6: results/k6-summary.json
- 火焰图: results/{cpu,wall,alloc,lock}.html
- JFR: results/lab4.jfr
- 节流: results/throttling-{before,after}.txt

## 4. 四类火焰图对比表

| | cpu | wall | alloc | lock |
| --- | --- | --- | --- | --- |
| 最宽的栈 | | | | |
| 能看到 Thread.sleep 吗 | | | | |
| 能看到 JSON 序列化吗 | | | | |
| 能看到锁等待吗 | | | | |

## 5. 观察与解释

| 我以为 | 实际是 | 结论 |
| --- | --- | --- |
| | | |

### 从 JFR 得到的、火焰图没给出的线索
(例如:某个 SocketRead 阻塞 480ms,远端是 PostgreSQL:5432)

## 6. 三个必答问题
1. 只采 CPU 火焰图会得出什么错误结论?
2. 怎么确认容器 CPU 节流?
3. `pending > 0` 意味着什么,接下来该做什么?

## 7. 结论

## 8. 被推翻的假设

## 9. 遗留问题与下一步

六、动手改造

改动 观察什么
把 scrape_interval 从 5s 改成 30s Grafana 曲线变粗糙,细粒度尖刺消失——理解采样频率的代价
去掉 /slow 接口的阻塞调用,重新采 wall 图 对比两张 wall 图(采一次 before、一次 after)
用 docker run --cpus=0.5 跑服务再执行 Lab 节流比例飙升,P99 抖动——把第 4.8 节的理论变成实测
在 verify-lab4.sh 里加上「火焰图文件大小 > 100KB」的检查 防止采到空图(太小说明没采到栈)
用 -XX:ActiveProcessorCount=1 配合 --cpus=1 观察 JVM 线程池大小与协程并行度的变化

七、这段代码的局限

  • Compose 方案适合本地验证,不适配生产(生产用 K8s + Prometheus Operator + 持久化存储)。
  • host.docker.internal 在 Linux 上默认不可用,需要用宿主机 IP 或 --add-host。
  • profile-all.sh 会串行采四类图(每类 30 秒),总共 2 分钟;如果压测时间短,可能采不完——这时应该缩短每类的时长(-d 15)或减少类型。
  • 本 Lab 的验证是「工具是否可用」,不是「性能结论」——真正的性能分析从第 5、6 章开始。