4.9 配套代码:Lab 4 观测回路的编排与验收
对应小节:4.9 Lab 4 把本章的工具串成一条回路,并自动化验证。
一、观测栈(最小的 Compose 配置)
# docker-compose.observability.yml
services:
prometheus:
image: prom/prometheus:v2.54.1
ports: ["9090:9090"]
volumes:
- ./observability/prometheus.yml:/etc/prometheus/prometheus.yml:ro
- ./observability/alerts.yml:/etc/prometheus/alerts.yml:ro
- prometheus-data:/prometheus
command:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.retention.time=15d
- --web.enable-lifecycle
grafana:
image: grafana/grafana:11.2.0
ports: ["3000:3000"]
environment:
GF_SECURITY_ADMIN_PASSWORD: admin
GF_USERS_ALLOW_SIGN_UP: "false"
volumes:
- grafana-data:/var/lib/grafana
depends_on: [prometheus]
volumes:
prometheus-data:
grafana-data:
# observability/prometheus.yml
global:
scrape_interval: 5s # 压测期间用 5s;生产可放宽到 15~30s
evaluation_interval: 5s
rule_files:
- /etc/prometheus/alerts.yml
scrape_configs:
- job_name: kotlin-app
metrics_path: /metrics
static_configs:
# macOS/Windows 用 host.docker.internal;Linux 用宿主机 IP
- targets: ["host.docker.internal:8080"]
# observability/alerts.yml —— 本章涉及的四个关键告警
groups:
- name: saturation
rules:
- alert: ConnectionPoolSaturated
expr: db_pool_pending > 0
for: 5m
labels: { severity: ticket }
annotations:
summary: "连接池排队(pending > 0 持续 5m)"
runbook: "① 查慢查询(pg_stat_statements)② 查长事务 ③ 最后才考虑调池(第 4.7 节)"
- alert: ContainerCpuThrottled
expr: |
rate(container_cpu_cfs_throttled_periods_total[5m])
/ rate(container_cpu_cfs_periods_total[5m]) > 0.25
for: 10m
labels: { severity: ticket }
annotations:
summary: "容器被 CPU 配额节流 > 25%(第 4.8 节)"
runbook: "① 确认 limit ② 看 CPU 需求是否合理 ③ 优化代码或提 limit"
- alert: GcPauseHigh
expr: |
histogram_quantile(0.99,
sum by (le) (rate(jvm_gc_pause_seconds_bucket[5m]))) > 0.2
for: 10m
labels: { severity: ticket }
annotations:
summary: "GC 停顿 P99 > 200ms"
runbook: "① 与 P99 尖刺对齐验证 ② 采 alloc 火焰图 ③ 检查堆与 GC 选型"
- alert: MetricCardinalityExplosion
expr: sum(count by (__name__) ({__name__=~".+"})) > 500000
for: 10m
labels: { severity: ticket }
annotations:
summary: "Prometheus 序列数超过 50 万,可能存在基数泄漏(第 4.6 节)"
二、一键跑完 Lab 4
#!/usr/bin/env bash
# tools/run-lab4.sh <EXP_ID>
set -euo pipefail
EXP_ID="${1:?usage: run-lab4.sh <EXP_ID>}"
DIR="docs/experiments/${EXP_ID}"
mkdir -p "$DIR/results" "$DIR/scripts"
echo "═══ 步骤 0:起观测栈 ═══"
docker compose -f docker-compose.observability.yml up -d
sleep 5
curl -sf http://localhost:9090/-/ready > /dev/null && echo "✅ Prometheus 就绪" || echo "⚠️ Prometheus 未就绪"
echo
echo "═══ 步骤 1:环境元数据 ═══"
tools/collect-env.sh "$EXP_ID"
tools/check-throttling.sh 5 | tee "$DIR/results/throttling-before.txt"
echo
echo "═══ 步骤 2:验证四层指标 ═══"
M=$(curl -s http://127.0.0.1:8080/metrics)
check() { echo "$M" | grep -qE "$1" && echo " ✅ $2" || echo " ❌ $2(缺这组指标)"; }
check "^app_requests_total" "① 业务层"
check "^app_request_duration.*bucket" "② 延迟层(直方图)"
check "^jvm_memory_used_bytes" "③ 资源层"
check "^db_pool_pending" "④ 饱和度层"
echo " 直方图桶数量: $(echo "$M" | grep -c 'app_request_duration_seconds_bucket')"
echo
echo "═══ 步骤 3:启动服务并预热 ═══"
scripts/restart-app.sh "$DIR/results/gc.log"
sleep 5
BASE_URL=http://127.0.0.1:8080 k6 run --quiet --vus 20 --duration 60s loadtest/profile-constant.js > /dev/null
echo "✅ 预热完成"
echo
echo "═══ 步骤 4:压测 + 采四类火焰图(并行)═══"
# 后台启动压测
BASE_URL=http://127.0.0.1:8080 RATE=400 DURATION=8m \
k6 run --out json="$DIR/results/k6-raw.json" \
--summary-export="$DIR/results/k6-summary.json" \
loadtest/profile-constant.js > "$DIR/results/k6-stdout.txt" 2>&1 &
K6_PID=$!
sleep 60 # 等压测进入稳态
PID=$(jcmd | grep app.jar | awk '{print $1}')
echo "采集中(PID=$PID)..."
tools/profile-all.sh "$EXP_ID" 30 | tee "$DIR/results/profile-log.txt"
echo
echo "═══ 步骤 5:JFR 录制 ═══"
jcmd "$PID" JFR.start name=lab4 settings=profile duration=60s filename="$DIR/results/lab4.jfr"
echo "等待 60 秒..."
sleep 65
jfr summary "$DIR/results/lab4.jfr" > "$DIR/results/jfr-summary.txt" 2>&1
echo "✅ JFR 摘要:" && head -20 "$DIR/results/jfr-summary.txt"
echo
echo "═══ 步骤 6:等压测结束 ═══"
wait $K6_PID || true
echo
echo "═══ 步骤 7:数据可用性检查 ═══"
tools/check-load-generator.sh "$DIR/results/k6-summary.json" 400
echo
echo "═══ 步骤 8:节流复查 ═══"
tools/check-throttling.sh 5 | tee "$DIR/results/throttling-after.txt"
echo
echo "✅ Lab 4 完成 → $DIR"
echo
echo "接下来(手工):"
echo " 1. 打开 Grafana http://localhost:3000(admin/admin)看面板"
echo " 2. 打开 $DIR/results/{cpu,wall,alloc,lock}.html 填对比表"
echo " 3. 用 JMC 打开 $DIR/results/lab4.jfr 找一条火焰图没给出的线索"
echo " 4. 填写 $DIR/README.md"
三、Grafana 面板的最小配置
如果不想手工建图表,可以用一个精简的 dashboard JSON(关键的四类面板):
{
"title": "Lab 4 · 观测回路",
"panels": [
{
"title": "① 到达率 vs 服务端 QPS",
"targets": [
{ "expr": "sum(rate(http_server_requests_seconds_count[1m]))", "legendFormat": "服务端 QPS" }
]
},
{
"title": "② 延迟分布(P50/P95/P99/P999)",
"targets": [
{ "expr": "histogram_quantile(0.50, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P50" },
{ "expr": "histogram_quantile(0.95, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P95" },
{ "expr": "histogram_quantile(0.99, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P99" },
{ "expr": "histogram_quantile(0.999, sum by (le) (rate(app_request_duration_seconds_bucket[1m])))", "legendFormat": "P999" }
]
},
{
"title": "③ 饱和度(连接池 + 队列)",
"targets": [
{ "expr": "db_pool_pending", "legendFormat": "连接池 pending" },
{ "expr": "db_pool_active", "legendFormat": "连接池 active" },
{ "expr": "executor_queue_depth", "legendFormat": "队列深度" }
]
},
{
"title": "④ 资源(CPU + GC 停顿)",
"targets": [
{ "expr": "process_cpu_usage", "legendFormat": "进程 CPU" },
{ "expr": "histogram_quantile(0.99, sum by (le) (rate(jvm_gc_pause_seconds_bucket[5m])))", "legendFormat": "GC 停顿 P99" }
]
}
]
}
四、Lab 4 的验收脚本
#!/usr/bin/env bash
# tools/verify-lab4.sh <EXP_ID>
#
# 自动检查 Lab 4 的产出是否齐全。
set -uo pipefail
EXP_ID="${1:?usage: verify-lab4.sh <EXP_ID>}"
DIR="docs/experiments/${EXP_ID}"
PASS=0; FAIL=0
ok() { echo " ✅ $1"; PASS=$((PASS+1)); }
bad() { echo " ❌ $1"; FAIL=$((FAIL+1)); }
echo "═══ Lab 4 验收 ═══"
echo
echo "① 环境元数据"
[ -f "$DIR/results/env.txt" ] && ok "env.txt 存在" || bad "缺 env.txt"
echo
echo "② 四层指标(从压测前的检查记录里找)"
[ -f "$DIR/results/profile-log.txt" ] && ok "profiling 日志存在" || bad "缺 profile-log.txt"
echo
echo "③ k6 数据"
[ -f "$DIR/results/k6-summary.json" ] && ok "k6-summary.json 存在" || bad "缺 k6-summary.json"
if [ -f "$DIR/results/k6-summary.json" ]; then
python3 - "$DIR/results/k6-summary.json" <<'PY'
import json, sys
m = json.load(open(sys.argv[1]))["metrics"]
tw = m["http_reqs"]["rate"]
print(f" 到达率 {tw:.1f} req/s, 错误率 {m['http_req_failed']['rate']:.2%}")
print(f" P50 {m['http_req_duration']['med']:.1f} ms, P99 {m['http_req_duration']['p(99)']:.1f} ms")
PY
fi
echo
echo "④ 四类火焰图"
for E in cpu wall alloc lock; do
[ -f "$DIR/results/$E.html" ] && ok "$E.html 存在 ($(du -h "$DIR/results/$E.html" | cut -f1))" \
|| bad "缺 $E.html"
done
echo
echo "⑤ JFR"
[ -f "$DIR/results/lab4.jfr" ] && ok "lab4.jfr 存在 ($(du -h "$DIR/results/lab4.jfr" | cut -f1))" \
|| bad "缺 lab4.jfr"
echo
echo "⑥ 容器节流记录"
[ -f "$DIR/results/throttling-after.txt" ] && ok "节流记录存在" || bad "缺节流记录"
echo
echo "⑦ 实验档案"
[ -f "$DIR/README.md" ] && ok "README.md 存在" || bad "缺 README.md(用 _TEMPLATE 创建)"
echo
echo "═══════════════════════"
echo "通过 $PASS 项,失败 $FAIL 项"
[ "$FAIL" -eq 0 ] && echo "✅ Lab 4 产出齐全" || echo "❌ 还有 $FAIL 项待补"
五、实验档案模板
<!-- docs/experiments/E04-observability-loop/README.md -->
---
id: E04
title: 最小观测回路搭建与验证
date: 2025-xx-xx
chapter: 4
kind: probe
status: done
hypothesis: "只采 CPU 火焰图会看不到阻塞类瓶颈;wall 火焰图能看到;容器节流若存在会影响 P99"
variable: "无(工具链验证)"
control: "故意加的 50ms 阻塞接口 /slow"
commit: <hash>
jvm_args: "-Xms1g -Xmx1g -XX:StartFlightRecording=name=rolling,settings=profile,maxsize=256m,maxage=1h,dumponexit=true"
container: "cpu.max=<值>, memory.max=<值>"
tags: [toolchain, profiling, jfr, lab4]
---
## 1. 假设与预期
- CPU 火焰图:看不到 /slow 的 Thread.sleep
- wall 火焰图:能看到大量栈停在 Thread.sleep
- alloc 火焰图:能看到 JSON 序列化的分配
- 容器节流:<预期>
## 2. 环境元数据
(粘 results/env.txt + throttling-before.txt)
## 3. 原始结果
- k6: results/k6-summary.json
- 火焰图: results/{cpu,wall,alloc,lock}.html
- JFR: results/lab4.jfr
- 节流: results/throttling-{before,after}.txt
## 4. 四类火焰图对比表
| | cpu | wall | alloc | lock |
| --- | --- | --- | --- | --- |
| 最宽的栈 | | | | |
| 能看到 Thread.sleep 吗 | | | | |
| 能看到 JSON 序列化吗 | | | | |
| 能看到锁等待吗 | | | | |
## 5. 观察与解释
| 我以为 | 实际是 | 结论 |
| --- | --- | --- |
| | | |
### 从 JFR 得到的、火焰图没给出的线索
(例如:某个 SocketRead 阻塞 480ms,远端是 PostgreSQL:5432)
## 6. 三个必答问题
1. 只采 CPU 火焰图会得出什么错误结论?
2. 怎么确认容器 CPU 节流?
3. `pending > 0` 意味着什么,接下来该做什么?
## 7. 结论
## 8. 被推翻的假设
## 9. 遗留问题与下一步
六、动手改造
| 改动 | 观察什么 |
|---|---|
| 把 scrape_interval 从 5s 改成 30s | Grafana 曲线变粗糙,细粒度尖刺消失——理解采样频率的代价 |
去掉 /slow 接口的阻塞调用,重新采 wall 图 |
对比两张 wall 图(采一次 before、一次 after) |
用 docker run --cpus=0.5 跑服务再执行 Lab |
节流比例飙升,P99 抖动——把第 4.8 节的理论变成实测 |
在 verify-lab4.sh 里加上「火焰图文件大小 > 100KB」的检查 |
防止采到空图(太小说明没采到栈) |
用 -XX:ActiveProcessorCount=1 配合 --cpus=1 |
观察 JVM 线程池大小与协程并行度的变化 |
七、这段代码的局限
- Compose 方案适合本地验证,不适配生产(生产用 K8s + Prometheus Operator + 持久化存储)。
host.docker.internal在 Linux 上默认不可用,需要用宿主机 IP 或--add-host。profile-all.sh会串行采四类图(每类 30 秒),总共 2 分钟;如果压测时间短,可能采不完——这时应该缩短每类的时长(-d 15)或减少类型。- 本 Lab 的验证是「工具是否可用」,不是「性能结论」——真正的性能分析从第 5、6 章开始。