记录源 Linear→Plane (2026-09-03 起, Plane MCP) + plane.chans.xyz 服务行/upstream 段; inventory + hosts/synapse.chans.xyz.md 补 Plane 部署事实 (Helm plane-ce-1.8.0 / app v1.4.1, ns plane, IngressRoute/自有证书 issuer/PVC 5+5Gi local-path/无备份层); 新增 runbooks/plane-health.md (只读健康检查) 与 docs/plane-hardening/ 草稿 (values.hardened.yaml、secrets.yaml.example 占位、backup/ CronJob), 均为未应用设计稿; .gitignore 增加 .tmp-* agent 临时文件。
3.4 KiB
Runbook: Plane Health Check
Purpose
Read-only health check of the self-hosted Plane project-management instance
(plane.chans.xyz) running on the synapse K3s cluster.
Scope
- Applicable: synapse.chans.xyz, namespace
plane. - Read-only: does not change pods, ingress, certificates, secrets, or configuration.
- Not applicable: Plane upgrade, values changes, or data recovery — those need a
reviewed change (see
ansible-operations.md/release.md).
Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-09-02
- Related systems: synapse.chans.xyz (Helm
plane-app, chartplane-ce-1.8.0, appv1.4.1)
Pass criteria
All of the following must hold; any conflict means STOP and record evidence.
sudo helm list -Ashowsplane-appin nsplane, STATUSdeployed.- All 7 Deployments + 4 StatefulSets in ns
planeare1/1 Runningwith 0 recent restarts. - Init Jobs
api-migrate-*/minio-bucket-*areCompleted. - Certificate
plane-app-ssl-certisREADY=True(CNplane.chans.xyz). https://plane.chans.xyz/returns HTTP 200 with a valid Let's Encrypt cert.- Root disk usage below the 80% warning threshold.
Procedure
1. Release and workloads
ssh -4 windy@synapse.chans.xyz 'sudo helm list -A'
ssh -4 windy@synapse.chans.xyz 'sudo k3s kubectl -n plane get deploy,sts,pods -o wide'
Expected: plane-app deployed; all workloads 1/1 Running, RESTARTS low;
no CrashLoopBackOff / Evicted. Otherwise STOP and record evidence.
2. TLS certificate and ingress
ssh -4 windy@synapse.chans.xyz 'sudo k3s kubectl -n plane get certificate,issuer,ingressroute'
Expected: plane-app-ssl-cert READY=True; plane-app-ingress present with
routes for /, /api, /spaces, /god-mode, /live, /uploads.
If READY=False with a pending HTTP-01 challenge, the usual cause is the node
DNS chain (coreDNS → systemd-resolved → uplink) failing to resolve
plane.chans.xyz — check resolvectl query plane.chans.xyz vs
dig +short plane.chans.xyz @8.8.8.8. If the record exists publicly but the
node fails, sudo resolvectl flush-caches and wait for the cert-manager retry;
do not mutate the issuer.
3. Endpoint verification
curl -4 -s -o /dev/null -w '%{http_code}\n' https://plane.chans.xyz/
echo | openssl s_client -connect plane.chans.xyz:443 -servername plane.chans.xyz 2>/dev/null | openssl x509 -noout -subject -issuer -dates
Expected: HTTP 200, cert CN=plane.chans.xyz issued by Let's Encrypt with a
future notAfter. https://plane.chans.xyz/api/ may 404 — the API serves
under /api/... paths only, so a bare 404 there is not a failure.
4. Storage and disk
ssh -4 windy@synapse.chans.xyz 'sudo k3s kubectl -n plane get pvc'
ssh -4 windy@synapse.chans.xyz 'df -hP /'
Expected: all 4 PVCs Bound (minio 5Gi, pgdb 5Gi, rabbitmq 100Mi, redis 100Mi
on local-path); root disk < 80%.
Safety
- Read-only: never mutate pods, ingress, certificates, or configuration during this check.
- Plane data (ns
planePostgres + MinIO PVCs) has no backup tier; treat the instance as at-risk until a backup design exists. - If live state conflicts with the expected values above,
STOPand record evidence; do not "fix in passing".
References
- hosts/synapse.chans.xyz.md — plane stack facts
- matrix-health.md — sibling service on the same cluster