From e957bc2bb1f451bf8943670f6f3c8785424cfd7e Mon Sep 17 00:00:00 2001 From: windyboy Date: Wed, 2 Sep 2026 18:01:43 +0800 Subject: [PATCH] docs: add host-disk-cleanup runbook + hk2 disk facts; record source -> Plane vps (VPS-81) --- AGENTS.md | 14 +- hosts/hk2.chans.xyz.md | 18 +++ runbooks/README.md | 1 + runbooks/host-disk-cleanup.md | 276 ++++++++++++++++++++++++++++++++++ scripts/validate-repo.sh | 2 +- 5 files changed, 306 insertions(+), 5 deletions(-) create mode 100644 runbooks/host-disk-cleanup.md diff --git a/AGENTS.md b/AGENTS.md index a84f947..6ceb5a5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -30,9 +30,12 @@ This repo is the **agent ops handbook + fact source** for maintaining personal V > read-only unless a change is approved). 7. Record each material VPS operation, incident, configuration change, or - verification outcome in the corresponding **Linear `vps` project**. Include - scope, action, verification, and remaining follow-up; never put passwords, - tokens, private keys, recovery keys, or private room IDs in Linear. + verification outcome in the corresponding **Plane `vps` project** (via the + Plane MCP, following the `plane-workflow` skill). This supersedes the former + Linear `vps` mirror for new ops; legacy mirrored items stay in Linear. + Include scope, action, verification, and remaining follow-up; never put + passwords, tokens, private keys, recovery keys, or private room IDs in + either tracker. ### Runbook execution rules @@ -93,6 +96,7 @@ not maintain a second copy of the machine table here. | Matrix health | [runbooks/matrix-health.md](runbooks/matrix-health.md) | | RustDesk health (hk2) | [runbooks/rustdesk-health.md](runbooks/rustdesk-health.md) | | AdGuard Home health | [runbooks/adguard-home-health.md](runbooks/adguard-home-health.md) | +| Host disk cleanup | [runbooks/host-disk-cleanup.md](runbooks/host-disk-cleanup.md) | | Home Assistant maintenance | [runbooks/home-assistant-maintenance.md](runbooks/home-assistant-maintenance.md) + [scripts/ha-maintenance.sh](runbooks/scripts/ha-maintenance.sh) | | matrix_e2ee update (hass.windy.lan) | [runbooks/matrix-e2ee-update.md](runbooks/matrix-e2ee-update.md) | | Matrix upstream refs | [docs/matrix-upstream.md](docs/matrix-upstream.md) | @@ -116,7 +120,9 @@ Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report. ### Issue tracker -Issues are tracked in Linear and created/updated via the Linear MCP (`vps` project). See `docs/agents/issue-tracker.md`. +Issues are tracked in **Plane** (`vps` project) and created/updated via the +Plane MCP, following the `plane-workflow` skill (supersedes the Linear `vps` +mirror for new ops). See `docs/agents/issue-tracker.md`. ### Triage labels diff --git a/hosts/hk2.chans.xyz.md b/hosts/hk2.chans.xyz.md index ab02179..98c341d 100644 --- a/hosts/hk2.chans.xyz.md +++ b/hosts/hk2.chans.xyz.md @@ -185,6 +185,23 @@ dig @202.91.35.141 SOA wsvc.info +short On-server docs: `/opt/pdns/README.md`, `CHANGELOG.md`. +## Disk / logging (VPS-81, 2026-09-02) + +Root disk cleanup performed (runbook: [host-disk-cleanup](../runbooks/host-disk-cleanup.md)): + +- Root `/` (20G vda1): 76% used → **38% used** (15G → 7.1G; free 4.7G → 12G). +- **AGH log flood root cause fixed**: `/opt/adguard/conf/AdGuardHome.yaml` + `log.verbose: true → false` (backup `AdGuardHome.yaml.bak-20260902-vps81`). + Verbose debug was streaming to stderr → container `json.log` (~120MB/day); + `log.file: ""` makes AGH's own rotation keys inert. Restart only (no recreate). +- Journald capped: `/etc/systemd/journald.conf.d/00-vps81.conf` + `SystemMaxUse=200M`; journal vacuumed to ~96M. +- Docker: engine **29.7.2**; 14 unused images removed (kept `pdns-auth-50:5.0.5` + rollback pin); 12 orphan anonymous volumes + build cache pruned. In-use + volumes intact (`pdns_dbdata`, `b594d738…` PG data, `e855d078…` backup). +- Follow-up: re-check AGH `json.log` growth **2026-09-09** (one-week checkpoint); + global docker log rotation only if still needed. + ## Verified Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded. @@ -194,3 +211,4 @@ Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded. - `only-notify=` + `also-notify=202.91.35.141`; MASTER `domains.master` cleared - https://pdns.wsvc.info → **302**; https://pgweb.wsvc.info → **401** - Hardening backlog: API/DB credential rotation + TSIG rotate (see upstream doc) +- 2026-09-02 (VPS-81): post-cleanup verified — 10 containers Up (adguardhome healthy), DNS SOA/NS + web endpoints OK; see Disk/logging section above. diff --git a/runbooks/README.md b/runbooks/README.md index 614b9ec..0c11c10 100644 --- a/runbooks/README.md +++ b/runbooks/README.md @@ -20,6 +20,7 @@ runbooks start from [`_template.md`](_template.md). | pgdb DB restore (pg_restore) | [pgdb-restore.md](pgdb-restore.md) | change (procedure) | | pgdb image/compose update | [pgdb-update.md](pgdb-update.md) | change (gated) | | AdGuard Home health check | [adguard-home-health.md](adguard-home-health.md) | read-only | +| Host disk cleanup (logs/apt/docker) | [host-disk-cleanup.md](host-disk-cleanup.md) | change (gated) | | Matter packet capture | [matter-packet-capture.md](matter-packet-capture.md) | read-only | | Home Assistant maintenance | [home-assistant-maintenance.md](home-assistant-maintenance.md) | change (gated) | | matrix_e2ee integration update | [matrix-e2ee-update.md](matrix-e2ee-update.md) | change (gated) | diff --git a/runbooks/host-disk-cleanup.md b/runbooks/host-disk-cleanup.md new file mode 100644 index 0000000..b053b81 --- /dev/null +++ b/runbooks/host-disk-cleanup.md @@ -0,0 +1,276 @@ +# Runbook: Host disk cleanup (unbounded container logs / apt cache / docker artifacts) + +## Purpose + +Reclaim space on a root filesystem that is filling up (≥70% used) on a Docker +Compose host, by fixing unbounded container log growth at the source, clearing +apt/journal caches, and removing unused Docker images/volumes. Success: root +usage drops to a safe band (≤55% used, or per acceptance in the tracking issue) +and log growth stays bounded afterwards. + +## Scope + +- 适用环境: production single-root-fs hosts running Docker Compose stacks + (first application: `hk2.chans.xyz`; reusable for `mx2.windy.me` / `us2.wsvc.info` + which run the same unbounded-`json.log` pattern). +- 适用对象: root filesystem usage; container stdout/stderr log files + (`/var/lib/docker/containers/*/*-json.log`); `/var/cache/apt`; systemd journal; + unused Docker images / anonymous volumes / build cache. +- 不适用情形: hosts without systemd-journald or without Docker; LAN/HAOS hosts + (use their own runbooks); cases needing disk *growth* (provider resize) rather + than cleanup; anything touching service data volumes or `/opt/*` configs + (STOP and use the service-specific runbook instead). + +## Ownership + +- Owner: windy (operator) + agent executing per approval +- Last reviewed: 2026-09-02 +- Related systems: hk2.chans.xyz (PowerDNS auth / AdGuard Home / Traefik / RustDesk compose stacks) + +## Preconditions + +- SSH access to the target host with **passwordless sudo** (`sudo -n true` must succeed). +- A recorded `df -h` baseline and `docker system df` baseline. +- **Explicit user approval** for every service touch listed in Approval gates + (recorded in the tracking issue, e.g. Plane `vps` VPS-81). +- No open incident on the target host. +- Container log growth root cause identified in Diagnose before mutating. + +## Inputs + +| Input | Source | Required | Validation | +|---|---:|---| +| Target host | inventory/hosts.md | yes | SSH login + `uname -r` | +| df/docker baseline | live read-only probe | yes | recorded before first mutation | +| Approved service touches | user confirmation in tracking issue | yes | issue comment states approval | +| Image keep-list (rollback pins) | operator decision in issue | yes | review `docker image ls` before rmi | +| Backup of any config edited | local copy with timestamp | yes | exists before edit | + +## Safety + +### Non-negotiable rules + +- Prefer read-only diagnosis before mutation (never mutate on an unmeasured disk). +- Never use `rm` on a live container log — use `truncate -s 0` (keeps the fd valid). +- Never run `docker image prune -a` when a keep-list is intended — no keep-list + exists; delete explicitly with `docker rmi`. +- Never run `docker volume prune -a` — plain `docker volume prune` (no `-a`) + removes only unused anonymous volumes; named/in-use volumes stay. +- After every mutation, verify the expected state (`df -h`, container status). +- Destructive actions require explicit approval (Approval gates). + +### Stop conditions + +- Live state conflicts with this runbook's preconditions or expectations (e.g. + root usage differs wildly from baseline, or a container is unhealthy). +- Missing approval, missing backup, or missing rollback ability. +- A verification step fails with no documented next step. +- Any step would touch a volume/container/mount that is not on the approved list. + +### Approval gates + +| Action | Risk | Explicit approval | Approval record | +|---|---:|---|---| +| `docker restart ` | low (sec-level blip of that service only) | yes | tracking issue (VPS-81 T1) | +| `systemctl restart systemd-journald` | low (sec-level, no state loss) | yes | tracking issue (VPS-81 T2) | +| `apt-get clean` | low (re-downloadable) | no | — | +| `journalctl --vacuum-*` / journald drop-in | low | no (restart above is gated) | — | +| `docker rmi` of unused images | medium (rollback pin removed unless kept) | yes (keep-list) | tracking issue (VPS-81 T3) | +| `docker volume prune` | medium (data in anonymous volumes lost) | yes | tracking issue (VPS-81 T4) | +| `docker builder prune` | low | no | — | + +## Procedure + +### Step 1 — Diagnose + +**Action** + +Read-only: `df -h`, `df -i`, `sudo du -x -h --max-depth=1 /`, `docker system df`, +and locate oversized container logs: +`sudo ls -la /var/lib/docker/containers/*/*-json.log`. Map a big log to its +container (`docker inspect -f '{{.Name}} {{.LogPath}}' `), then inspect what +it logs (`sudo tail -c 400000 `; count `[debug]` lines) and find the +config flag driving it (e.g. AGH `log.verbose` in its YAML; note `log.file: ""` +means the app's own rotation keys are inert and output goes to the container log). + +**Expected** + +A full accounting of root usage and identification of: (a) any unbounded +container log and its root-cause flag; (b) reclaimable apt cache; (c) journal +size and journald limits; (d) unused images (0 dangling expected) and unused +anonymous volumes. + +**Decision** + +- If root is ≥70% used or any container log is unbounded → Step 2. +- If root is healthy and logs are bounded → STOP (no change needed; record evidence). +- If state conflicts with expectations (e.g. missing sudo, unexpected mount) → STOP. + +### Step 2 — Fix noisy container logging at the source, then truncate + +**Action** + +1. Back up the app config: `sudo cp .bak-YYYYMMDD-`. +2. Disable the debug/verbose flag (e.g. `log.verbose: true → false` in the AGH YAML). +3. Apply config with a container restart: `docker restart ` (config-level + change; **no recreate** needed and daemon.json rotation would not apply anyway). +4. Truncate the accumulated logs: `sudo truncate -s 0 ` for the chatty + container(s) (and any other oversized ones, e.g. traefik). +5. Record `df -h` before/after. + +**Expected** + +`docker logs ` no longer shows the `[debug]` flood; the `*-json.log` +stops growing; several GB reclaimed. + +**Verification** + +- `sudo tail -c 200000 ` after ≥1 minute → no new debug lines. +- `df -h` improvement recorded. +- Container still `Up (healthy)`. + +**Rollback** + +- Trigger: log volume unchanged, service degraded, or debug output is actually needed. +- Action: restore the config backup and `docker restart `. +- Verify: original verbose behaviour back; container healthy. + +### Step 3 — Clear apt cache and cap journald + +**Action** + +1. `sudo apt-get clean` (clears only `/var/cache/apt/archives`; `/var/lib/apt/lists` + is not cleared by it and regenerates on `apt update` — optional/low value, skip). +2. `sudo journalctl --vacuum-size=100M`. +3. Write drop-in `/etc/systemd/journald.conf.d/00-disk-.conf`: + `[Journal]` + `SystemMaxUse=200M`. +4. `sudo systemctl restart systemd-journald` (approved service touch). +5. Record `df -h` before/after. + +**Expected** + +Archives cleared (~1.4G on hk2), journal ≤100M, future journal capped at 200M. + +**Verification** + +- `du -sh /var/cache/apt/archives` → ~0. +- `journalctl --disk-usage` → ≤100M. +- `systemctl show systemd-journald -p ...` or restart log confirms new limit; + `journalctl -b` still readable. + +**Rollback** + +- Trigger: journald fails to start or logs lost unexpectedly. +- Action: remove the drop-in, `sudo systemctl restart systemd-journald`. +- Verify: journald active, prior journal entries still listed. + +### Step 4 — Remove unused Docker images (explicit keep-list) + +**Action** + +1. Enumerate unused images: `docker image ls` cross-checked against the images of + running containers (`docker ps --format '{{.Image}}'`). Re-enumerate at + execution time — the list drifts. +2. Present the exact removal list to the operator; keep the agreed rollback pin(s) + (e.g. `powerdns/pdns-auth-50:5.0.5`) and delete the rest explicitly: + `docker rmi ...` (per image). +3. Record `df -h` before/after. + +**Expected** + +Only in-use images + kept pins remain; ~1–2.5G reclaimed (reclaim is an upper +bound — layers shared with kept images are not freed; measure with `df`, do not +promise the estimate). + +**Verification** + +- `docker image ls` shows only the expected set. +- `docker system df` images reclaimable ≈ 0 for the removed set. +- All containers still `Up`. + +**Rollback** + +- Trigger: an image that was actually needed was removed. +- Action: re-pull it from the registry (`docker pull `); if a kept pin + must change, update the compose pin and `up -d`. +- Verify: image present; affected service healthy. + +### Step 5 — Remove unused anonymous volumes and build cache + +**Action** + +1. Enumerate volumes: `docker volume ls`, and confirm which are referenced by + containers (`docker inspect` Mounts). Expected targets: anonymous volumes with + no container reference. +2. `docker volume prune` (**no `-a`**) — engine ≥ v23 removes only unused + anonymous volumes; in-use volumes (e.g. PG data) are protected by container + references in every version. +3. `docker builder prune -f`. +4. Record `df -h` before/after. + +**Expected** + +Unused anonymous volumes (~1.2G on hk2) and build cache gone; in-use volumes intact. + +**Verification** + +- `docker volume ls` shows only in-use volumes. +- Services that own volumes (e.g. postgres) report healthy and data present. +- `df -h` improvement recorded. + +**Rollback** + +- Trigger: data loss suspected in a removed volume. +- Action: restore from backup if the volume ever contained data; verify against + the pre-prune enumeration (targets must be anonymous + unreferenced before prune). +- Note: this is why target enumeration is recorded before pruning. + +## Troubleshooting + +### Troubleshooting A — Log still grows after disabling verbose + +- Evidence: `sudo tail -c 200000 ` still shows new lines; app config re-checked. +- Allowed actions: check for a second verbose source (container entrypoint flags, + other apps in the same log); check `docker inspect --format '{{.HostConfig.LogConfig}}'`. +- Next step: back to Step 2 or STOP if a container-level log-opts change (recreate) + would be needed — that is a separate approval. + +### Troubleshooting B — `docker rmi` fails (image in use) + +- Evidence: `image is being used by stopped container ...`. +- Allowed actions: identify the stopped container (`docker ps -a`); confirm it is + not needed; remove it only with explicit approval. +- Next step: re-run rmi for the remaining images; never force-delete blindly. + +### Troubleshooting C — `docker volume prune` would remove more than expected + +- Evidence: prune dry-run/listing includes a named or referenced volume. +- Allowed actions: abort; do not add `-a`; re-check references. +- Next step: STOP and report to the operator with the enumeration. + +## Final Verification + +The flow is successful only when all of the following hold: + +- `df -h` root usage is in the agreed band (VPS-81: 76% → ≤55% used; measure, do not assume). +- `docker system df` shows reclaimable ≈ 0 for images/volumes targeted. +- All containers `Up` (health checks pass); public services verified + (`dig @ SOA ` for DNS hosts; service URLs reachable). +- Tracking issue updated with before/after `df`, actions, and the one-week + observation checkpoint for log growth. + +## Failure Handling + +If the flow cannot complete: + +1. Stop further mutation. +2. Collect command output, timestamps, and the exact step that failed. +3. Record completed steps, actual results, unmet expectations, and whether a + rollback ran. +4. Hand over per the tracking issue with evidence; do not guess further. + +## References + +- Plane `vps` issue VPS-81 "hk2: 释放根盘空间" (+ subtasks VPS-82…88) — plan, review findings, approvals. +- [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) — host facts. +- [RUNBOOKS.md](../RUNBOOKS.md) — runbook spec; [runbooks/README.md](README.md) — index. diff --git a/scripts/validate-repo.sh b/scripts/validate-repo.sh index ec4bff3..13ca199 100755 --- a/scripts/validate-repo.sh +++ b/scripts/validate-repo.sh @@ -138,7 +138,7 @@ fi # --------------------------------------------------------------------------- say "== 4. Runbook spec compliance (RUNBOOKS.md) ==" -PROCEDURE_RUNBOOKS="fix-ci.md issue-to-merge.md network-change.md network-recovery.md pgdb-restore.md release.md rollback.md" +PROCEDURE_RUNBOOKS="fix-ci.md host-disk-cleanup.md issue-to-merge.md network-change.md network-recovery.md pgdb-restore.md release.md rollback.md" GATED_REF_RUNBOOKS="mailcow-update.md ansible-operations.md home-assistant-maintenance.md matrix-e2ee-update.md pgdb-update.md vaultwarden-sqlite-to-postgres.md" RUNBOOK_INDEX="runbooks/README.md" for f in runbooks/*.md; do