- runbooks/pgdb-health.md: Last reviewed 2026-08-29 → 2026-09-22;Related systems 加入 scribe.csg_history 永久归档(job 1010/1011 维护) - hosts/pgdb.md: Runbooks 行注明 9 项(第 9 项 CSG 归档新鲜度)
14 KiB
Runbook: pgdb health (TimescaleDB + pgweb + pg-backup)
Purpose
Read-only health check of the pgdb TimescaleDB compose stack (PG18 + pgweb GUI + nightly custom-format backups). Confirms the stack is serving Home Assistant (hass/scribe) and that backups are current and restorable.
Scope
- Applicable: pgdb (
192.168.55.15),/opt/databasecompose stack. - Read-only: never mutates containers, databases, backups, or secrets.
- Not applicable: restoring data (use pgdb-restore), upgrading images (use pgdb-update), HA-side changes (see
hosts/hass.windy.lan.md).
Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-09-22(新增 check 9:CSG 归档新鲜度)
- Related systems: pgdb (
/opt/database, TimescaleDB 18.6 / TS 2.29.2), HA192.168.55.11(hass/scribe clients),scribe.csg_history永久归档(南方电网逐日/逐月,由 job 1010/1011 维护)
Access
SSH to pgdb (key auth works from the WSL agent shell as of 2026-08-29):
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15
From the agent sandbox, always use -F /dev/null (system ssh config is unreadable there) and prefer IPv4. The compose project lives at /opt/database — prefix every docker compose call with cd /opt/database. Never print .env values or bookmarks (they contain DB passwords); compare or use them only inside commands that output non-secret signals (status codes, counts, names).
Safety
Non-negotiable rules
- Read-only diagnosis only; never "fix while checking".
- Never print passwords or secrets — redact/consume them inside commands.
- If live state conflicts with an expected value,
STOPand record evidence; do not invent parameters or bypass a failed check. - Restore/update work belongs to the change runbooks, not this one.
Stop conditions
- Any container
Exited,Restarting, or nothealthywhere expected. - Expected database/table/hypertable missing or a key query errors.
- Write-activity sample does not increase (scribe
states_rawstatic). - Latest daily backup older than today, not custom format, or
pg_restore -lfails. - Disk usage near full on
/srv/pgdataor/. - New
ERROR/FATALlines in the timescaledb log or backup failures in the pg-backup log. - CSG 归档新鲜度断言失败(见 check 9):
csg_history最新 day 行落后超过 2 天、period上出现缺口、或任一 CSG job 非Success/total_failures > 0。
Pass criteria
docker compose ps -a:timescaledb+pg-backupUp (healthy),pgwebUp; ports bound to192.168.55.15:5432and:8081.- PG 18.x; databases
hass,scribe,postgrespresent; HA (192.168.55.11) connected ashassto bothhassandscribe. hass.statesandscribe.states_rawrow counts grow between two samples (scribe writes continuously).- TimescaleDB extension 2.29.x; scribe hypertables
states_raw+events(1-dim,time); compression configured (segmentby/orderby rows intimescaledb_information.compression_settings);entitiestable exists. - pgweb: no credentials → HTTP 401; with credentials → HTTP 200;
/api/bookmarks→["hass","scribe"]. daily/*-latest.dumpsymlinks point to today's dumps;file -LreportsPostgreSQL custom database dump./srv/pgdata(/dev/sdb1, 32G) and/not near full; fstab mounts/srv/pgdatabyUUID=c9e12e79-1f66-404c-ab7f-b8809be81d86withdefaults,noatime.- timescaledb log: no new
ERROR/FATAL; pg-backup log: recent successful backup. scribe.csg_history:最新 day 行period >= current_date - 2、period上无缺口;csg_daily_snapshot(1010)与csg_backfill_missing_days(1011)均Success、total_failures = 0; 各月|sum(day cost) − month cost| ≤ 0.05(逐日舍入的累积,非缺陷)。
Checks
1. Containers
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose ps -a --format "table {{.Name}}\t{{.Status}}\t{{.Ports}}"'
Expected
timescaledbUp (healthy),pg-backupUp (healthy),pgwebUp.- Ports:
192.168.55.15:5432->5432/tcp(timescaledb),192.168.55.15:8081->8081/tcp(pgweb).
Stop if any container is Exited/Restarting/unhealthy, or a port binding changed.
2. PG core and HA clients
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select version();" | head -1'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select datname from pg_database where datistemplate=false order by 1;"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select datname, usename, client_addr from pg_stat_activity where client_addr is not null group by 1,2,3 order by 1;"'
Expected
PostgreSQL 18.x(verified: 18.6).- Databases:
hass,postgres,scribe. - HA sessions:
hass|hass|192.168.55.11andscribe|hass|192.168.55.11(the HAOS recorder/scribe clients from192.168.55.11).
Stop if a database is missing, the version is not 18.x, or HA has no live sessions (scribe connectivity is part of the HA pipeline).
3. Write activity
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database
A=$(docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select count(*) from states_raw;")
sleep 30
B=$(docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select count(*) from states_raw;")
echo "states_raw $A -> $B"'
Also sample hass.states once (recorder table, bulk-writes on HA restart): docker compose exec -T timescaledb psql -U postgres -d hass -Atc "select count(*) from states;".
Expected — states_raw increases between samples (verified: 2665 → 2693 in 30 s). states count is sane (thousands).
Stop if states_raw is static across samples while HA is up — writes have stalled.
4. TimescaleDB (extension, hypertables, compression)
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select extversion from pg_extension where extname='"'"'timescaledb'"'"';"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select hypertable_name, num_dimensions from timescaledb_information.hypertables order by 1;"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select hypertable_name, attname, segmentby_column_index, orderby_column_index from timescaledb_information.compression_settings order by 1,3,4;"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select to_regclass('"'"'public.entities'"'"');"'
Expected
- Extension version
2.29.x(verified: 2.29.2). - Scribe hypertables:
states_rawandevents, both1dimension. - Compression configured for
states_raw(segmentbymetadata_ididx 1, orderbytimeidx 1) andevents(segmentbyevent_type, orderbytime). Note: TimescaleDB 2.29.x has nocompression_enabledcolumn in this view — row presence is the enabled signal. entitiesresolves (scribe registry table).
Stop if the extension version differs from the pinned 2.29.x line, a hypertable is missing, compression rows vanish, or entities is absent (scribe schema broke — see Known issues in hosts/pgdb.md).
5. pgweb GUI
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database
curl -s -o /dev/null -w "no-auth:%{http_code}\n" --max-time 8 http://192.168.55.15:8081/
U=$(grep -E "^PGWEB_AUTH_USER=" .env | cut -d= -f2-); P=$(grep -E "^PGWEB_AUTH_PASS=" .env | cut -d= -f2-)
curl -s -o /dev/null -w "auth:%{http_code}\n" --max-time 8 -u "$U:$P" http://192.168.55.15:8081/
echo -n "bookmarks:"; curl -s --max-time 8 -u "$U:$P" http://192.168.55.15:8081/api/bookmarks; echo'
Use 192.168.55.15:8081 (pgweb binds the VM IP only — loopback is not bound). Credentials are read from .env on the host and never printed.
Expected — no-auth:401, auth:200, bookmarks:["hass","scribe"].
Stop if pgweb is unreachable, unauthenticated access is not 401, or bookmarks diverge from ["hass","scribe"].
6. Backups
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'ls -l --time-style=long-iso /opt/database/backups/daily/'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'file -L /opt/database/backups/daily/hass-latest.dump'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T pg-backup pg_restore -l /backups/daily/hass-latest.dump | head -4'
Expected
daily/*-latest.dumpsymlinks point to today's*-YYYYMMDD.dump(nightly 02:00 localAsia/Shanghai; a fresh container also firesBACKUP_ON_START).file -LreportsPostgreSQL custom database dump(pg_restore format; verified v1.16-0).pg_restore -lfrom the pg-backup container lists the archive TOC without error (timescaledb does not mount/backups).
Stop if the latest dump is not from today, is not custom format, or pg_restore -l fails. Stray non-daily/ dumps at the /opt/database/backups/ root are pre-compose leftovers — ignore for health, flag for cleanup.
7. Disk and mount
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'df -h /srv/pgdata /'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'grep -E "srv/pgdata" /etc/fstab'
Expected
/srv/pgdata=/dev/sdb132G (verified: 88M used / 30G avail) and/with comfortable headroom.- fstab:
UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86 /srv/pgdata ext4 defaults,noatime 0 2.
Stop if either filesystem is near full (define threshold before acting) or the fstab entry is missing/changed.
8. Logs
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose logs --since 24h timescaledb 2>&1 | grep -E "ERROR|FATAL" | tail -10'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose logs --since 24h pg-backup 2>&1 | tail -5'
Expected
- timescaledb: no new
ERROR/FATAL. Known benign history: an oldrelation "hass.states" does not existfrom a wrong-schema probe andcompression_enabledcolumn errors from an outdated query — neither recurs with the commands above. - pg-backup: recent successful run (
SQL backup created successfullyfor each database, no restore/cleanup errors).
Stop if repeated ERROR/FATAL appear or a backup run failed.
9. CSG 长期归档新鲜度(scribe.csg_history)
为什么必须单独看这张表:它承载南方电网的永久逐日/逐月归档,但写入它的 job 曾经 294 次运行 0 成功、静默造成 23 天缺口而无人发现。更隐蔽的是: 上游停更时 job 会反复 upsert 同一行、
last_run_status仍是Success、 日期上也不会出现"缺口" —— 只有本节的"最新日期"断言能发现它。
# a) 最新 day 行日期(期望 >= 今天 - 2;CSG 数据本身滞后一天)
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select max(period) as last_day, current_date - 2 as min_expected from csg_history where kind='\''day'\'';"'
# b) period 上是否有缺口(期望 0 行)
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select period from generate_series('\''2026-07-01'\''::date, (select max(period) from csg_history where kind='\''day'\''), '\''1 day'\'') period except select period from csg_history where kind='\''day'\'';"'
# c) 两个 CSG job 的状态(期望均为 Success / 0 failures)
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select j.job_id, j.proc_name, s.last_run_status, s.total_runs, s.total_failures from timescaledb_information.jobs j join timescaledb_information.job_stats s on s.job_id = j.job_id where j.proc_name like '\''csg%'\'' order by j.job_id;"'
# d) 日费用与月费用的口径一致性(期望 |diff| <= 0.05)
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select m.period, m.usage_kwh, m.cost, round(d.s,2) as day_sum, round(d.s - m.cost,2) as diff from csg_history m left join (select date_trunc('\''month'\'',period)::date as mon, sum(cost) as s from csg_history where kind='\''day'\'' group by 1) d on d.mon = m.period where m.kind='\''month'\'' and m.period >= '\''2026-07-01'\'' order by m.period;"'
Expected
- a)
last_day >= current_date - 2。 - b) 0 行。
- c)
csg_daily_snapshot(job 1010,14:30 UTC = 22:30 CST)与csg_backfill_missing_days(job 1011,15:10 UTC = 23:10 CST)均Success、total_failures = 0。 - d) 每个月
|diff| <= 0.05元。这是逐日四舍五入到 2 位的累积,不是缺陷; 月行(由csg_ladder_cost单次舍入)才是权威值。
Stop if
last_day落后超过 2 天 → 上游停更或 job 未运行。先看 c) 的last_run_started_at与ha core logs里的china_southern_power_grid_stat错误,再决定是否需要 重新加载集成 config entry(见 home-assistant-maintenance)。- b) 出现缺口 → 等
csg_backfill_missing_days下一轮(或手动CALL run_job(1011);) 后复查;仍缺则查timescaledb_information.job_errors。 - c) 任一 job 非
Success或total_failures > 0→ 查job_errors的err_message; 历史上 job 1008 的失败原因是函数签名与 TimescaleDB 自定义 job 约定不符 (必须是(job_id integer, config jsonb)),修复记录见compose/pgdb/csg-snapshot-v3.sql/-v5.sql。 - d) 某个
|diff| > 0.05→ 说明有新写入路径绕过了 v5 的边际差分口径,停止并复核。