Author SHA1 Message Date
windyboy ff1a92110c docs(us2): Soft Serve → Gitea 迁移事实与参考镜像 (Plane VPS-94)
- hosts/us2: Gitea 1.27.3-rootless 部署实况 (repo.windy.me SSH:2222/Web), 16 仓迁移核对, 备份/回滚; soft-serve 停用保留作回滚
- compose/gitea: 参考镜像 (rootless compose + 备份 sidecar + 一次性迁移脚本留档)
- AGENTS/inventory/compose README: 服务表与索引同步
2026-09-18 17:24:35 +08:00
windyboy cc3fb99c14 docs(us2): 根盘清理 71%→23% 事实记录 (Plane VPS-93) 2026-09-18 11:46:17 +08:00
windyboy 6879d79cc6 docs(hass,pgdb): scribe 4.4.0 升级核对 + stats_io_interval 300 + sensor_minute 体积诊断
- hass: Scribe 3.8.0 → 4.4.0(HACS jonathan-gtd/scribe,2026-09-13 随 Core
  2026.9.1 / HAOS 18.2 升级)。记录 4.0 两个 breaking change 在本机均无需动作
  (3.x 结构 DB + PK 在、TimescaleDB 2.29.2 已装)、配置优先级
  YAML > options > entry data > 默认值、以及「YAML 改动必须重启 Core」。
- hass: 新增 stats_io_interval: 300(备份 scribe.yaml.bak-20260913-191558)。
  变更前 24h scribe 自写 11 019/87 461 行(12.6%);实测发布间隔 300s。
  retention_states/retention_events 可用但刻意不设;flush_interval 仍被 entry
  data 钉在 5s(采用新默认 30s 需显式写 YAML)。
- pgdb: sensor_minute 2.26 GB 是压缩窗口内的正常暂存(compress_after 按 chunk
  结束时间判断,09-17 才合格),稳态 4–5 GB 平台期,暂不处理;复查点 09-17
  之后。顺带修正 hass/scribe 库体积事实。

校验:scripts/validate-repo.sh PASS (0 warnings)。
2026-09-13 19:34:04 +08:00
windyboy c445c5f512 Merge remote-tracking branch 'origin/main' into main
AGENTS.md 冲突(两侧都改了 Linear→Plane 记录源):取本地更完整的表述
(self-hosted plane.chans.xyz + mcp__plane__* + 2026-09-03 停用日期 +
issue-tracker.md 已过时),并吸收远端的 `plane-workflow` skill 指引。
2026-09-13 19:13:06 +08:00
windyboy 3c83246f24 docs(pgdb,gfw): pgdb VM 重启根因 (W1N-263) + gfw Quad9 上游移除记录 2026-09-13 19:12:32 +08:00
windyboy c0c975584a docs(plane): 自托管 Plane 落地事实入仓库 + plane-health runbook + hardening 草稿
记录源 Linear→Plane (2026-09-03 起, Plane MCP) + plane.chans.xyz 服务行/upstream 段;
inventory + hosts/synapse.chans.xyz.md 补 Plane 部署事实 (Helm plane-ce-1.8.0 / app v1.4.1,
ns plane, IngressRoute/自有证书 issuer/PVC 5+5Gi local-path/无备份层);
新增 runbooks/plane-health.md (只读健康检查) 与 docs/plane-hardening/ 草稿
(values.hardened.yaml、secrets.yaml.example 占位、backup/ CronJob), 均为未应用设计稿;
.gitignore 增加 .tmp-* agent 临时文件。
2026-09-13 19:12:32 +08:00
windyboy 0034cec925 Merge remote-tracking branch 'origin/main' into HEAD 2026-09-13 19:11:13 +08:00
windyboy c9dcde1274 docs(hass): Quick 面板时间范围扩容 — 功率/环境/人体感应加 3d-30d 档,用电量加 3mo (VPS-92)
卡片 JS 只接受 <n>m|<n>h|<n>d 与命名档(today/week/month/3mo/6mo/year/custom);
记录 scribe sensor_minute 数据下界 2026-08-29 对 >15d 档的影响。
2026-09-13 19:09:54 +08:00
windyboy 037c4ccaa5 docs(hass): 马桶换气电源 Matter 插座接入 + 功率积分补电量 (VPS-92)
Matter Smart Plug (SIXWGH model 3596) 的 cluster 0x0091 声明 IMPE+CUME+PERE 但
CumulativeEnergyImported 恒为 null -> HA 能量实体永久 unknown。功率计量正常
(实测 24.7 W),故用 Integration (Riemann sum) 辅助元素补电量实体
sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_energy:
能源仪表盘 grid 源 [8] 与 Quick「用电量(按插座)」图换到该实体,Quick 另加
span-2「开关」区块。记录 config-flow 经 supervisor 代理走 REST 的 Agent 方法、
备份路径与回滚方式;energy/validate 已全绿。
2026-09-13 18:39:22 +08:00
windyboy 17bb171578 docs(hass): CSG 面板重构入维护 runbook — avg price 传感器 + 季节提醒 automation + WS 改法 (VPS-90) 2026-09-04 21:55:40 +08:00
windyboy cecf7e6331 docs(hass): CSG 电力监控面板重构 — avg price 传感器 + 去重/改名 + 季节切换提醒 (VPS-90) 2026-09-04 21:44:43 +08:00
windyboy e957bc2bb1 docs: add host-disk-cleanup runbook + hk2 disk facts; record source -> Plane vps (VPS-81) 2026-09-02 18:02:22 +08:00
windyboy de52cb8b57 docs(hass): 地图卡 CARTO 水印修复 — custom:map-card v1.16.0 + keyed tiles (W1N-261) 2026-08-30 13:39:42 +08:00
windyboy b15e19bce9 docs(pgdb): compose 开机竞态故障修复 + 自愈 unit (W1N-260)
- 根因:开机时 docker 恢复容器绑定 192.168.55.15:5432/8081 失败(EADDRNOTAVAIL,IP 尚未可绑)→ timescaledb/pgweb 启动失败且不重试,停摆 3h15m;pg-backup 开机备份失败 → unhealthy
- 处置:docker compose up -d --force-recreate(三容器回 database_default、端口发布、备份恢复、pgweb 恢复);用户重启 HA Core 后写入管道恢复
- 防复发:新增开机自愈 systemd oneshot pgdb-compose.service(enabled),源码 compose/pgdb/pgdb-compose.service
2026-08-30 13:19:05 +08:00
windyboy bee54a6858 docs(hass): CSG 长期归档 csg_history + recorder 365d 补录 (W1N-243) 2026-08-30 13:19:05 +08:00
windyboy d6747028b4 docs(soft-serve): 镜像固定 v0.12.2 + 备份 sidecar + 非 root 运行 (W1N-244..248)
- 镜像 pinned charmcli/soft-serve:v0.12.2(GHCR 为 dev/nightly 源,无 v0.12.x tag)
- soft-serve-backup sidecar:每日 02:00 sqlite .backup + repos-config 打包,03:00 prune 保留 14 份
- 非 root 运行(user 1000:1000),data chown;ssh.public_url 修复
- compose 源码参考:compose/soft-serve/(服务器文件为准)
2026-08-30 13:19:05 +08:00
windyboy f174aa1219 docs(hass): CSG 复核遗留修复 — 模板 days[-1] 补排序 + 本月日均/预测进度上屏 (W1N-242) 2026-08-29 21:08:13 +08:00
windyboy 5b5f6042e6 docs(hass): CSG 模板/面板修复记录 — availability 硬化 (W1N-239)、off-by-one + 新传感器 (W1N-241)、gauge 阶梯对齐 + 年度统计 (W1N-240)
- hosts/hass.windy.lan.md: csg_sensors 三次变更记录 + 季节性 gauge 切换已知事项 (11-01/5-01)
- runbooks/home-assistant-maintenance.md: float(0) fake-zero follow-up 标记已解决
2026-08-29 21:02:33 +08:00
windyboy 50136b2ffd docs(hass): W1N-238 — scribe config split to scribe.yaml + templates/ merge include
- Scribe 3.8.0 block moved verbatim from configuration.yaml to
  /homeassistant/scribe.yaml (scribe: !include scribe.yaml); YAML stays
  authoritative, import semantics unchanged.
- template: switched to !include_dir_merge_list templates; new
  quick_sensors.yaml scaffold (top-level list, quick_ prefix, unique_id
  required; pure sums stay min_max per W1N-233).
- Verified post-restart 2026-08-29 20:19 CST: core check ok, scribe
  connection on, states_written 18581→19426, template entities = 12,
  no scribe/template log errors. Backup configuration.yaml.bak-20260829-201724-w1n238.
2026-08-29 20:25:53 +08:00
windyboy 6707cebc88 chore: gitignore .agent-work/ agent scratch dir
Untracked agent scratch made validate-repo.sh link scan fail; same
category as the already-ignored .agents/ and .claude/ dirs.
2026-08-29 20:25:53 +08:00
windyboy 908ff5412a docs(hass): Quick dashboard round-2 state — badges, total-power helper, span-2 power pair, fill/colors (W1N-231)
Sync Quick dashboard section with live config: 2x2 mushroom light grid,
kong_diao AC entity fix, heading badges (env temps / AC / PC / total power),
min_max sum helper sensor.dang_qian_zong_gong_lu (fail-closed), selective
tozeroy fill on base-load chart, power charts paired as span-2 sections with
card titles removed, motion per-entity colors.
2026-08-29 19:54:35 +08:00
windyboy e58283210a chore(vaultwarden,healthcheck): upgrade 1.37.2 (Bitwarden 2026.8+); fix vps-health checks
- vaultwarden/server:1.37.1 -> 1.37.2 (required for Bitwarden clients 2026.8.0+)
- compose probe: flag only active services (config --services) so debug-profile
  pgweb 'Exited' no longer false-positives
- runner: build aggregate args line-by-line (robust vs Jinja trim_blocks)
- SMTP AUTH probe moved host-side (vaultwarden image has no python3); never
  prints the SMTP password
- us2 facts: probe refresh 2026-08-29, image/version, vps-health install
2026-08-29 19:54:32 +08:00
windyboy 8c73d1f894 docs(hass,pgdb): timescale-plotly-card chart stack — reader+card install, sensor_minute pipeline, Quick dashboard
- reader timescale_database_reader v1.1.0 (bb8776a) + card timescale-plotly-card
  2.2.0 (217961d), manual installs; config entry, Lovelace resource id recorded
- pgdb scribe: sensor_minute_aggregate cagg + sensor_minute hypertable + jobs
  1005/1006/1007; states_raw 3-month retention/compression statements deliberately
  skipped (permanent archive per host doc)
- sensor_minute_refresh local patch ELSE 0 → ELSE NULL (unavailable-minute zeros
  poison diff-mode energy charts) + one-time cleanup (505 head rows, 26 impossible
  zeros); re-apply after re-running upstream 02 SQL
- Quick dashboard: 5 chart cards via WS lovelace/config/save; documented section
  column_span (absent → span 1) vs card grid_options sizing rules
2026-08-29 15:54:40 +08:00
windyboy bc0a86245d docs: Scribe retention v4.x status (user declined RCs, 2026-08-29); HA Core 2026.8.3 verified 2026-08-29 14:56:05 +08:00
windyboy 13032fd0bb Merge origin/main (production Makefile) into pgdb runbooks delivery 2026-08-29 14:52:10 +08:00
windyboy d2063e7496 Merge pgdb ops runbooks — health/restore/update (W1N-228) 2026-08-29 14:51:31 +08:00
windyboy 2d95f87897 docs(runbooks): pgdb ops runbooks — health / restore / update + facts refresh (W1N-228)
- runbooks/pgdb-health.md: read-only health check (8 diagnostics) — containers,
  PG core + HA clients, write activity, TimescaleDB hypertables/compression,
  pgweb auth/bookmarks, daily custom-format backups, disk/fstab, logs
- runbooks/pgdb-restore.md: procedure-type restore (pg_restore -Fc, temp-DB swap,
  approval gates, rollback) — precondition command verified live
- runbooks/pgdb-update.md: gated command reference (pull -> config -q -> up -> verify;
  rollback = /opt/database/run + old volumes)
- index + validate-repo.sh classification updated; hosts/pgdb.md refreshed
  (SSH key auth works, scribe events hypertable, runbook cross-refs)
2026-08-29 14:51:20 +08:00
windyboy b61513c93e chore(pgdb): land W1N-227 compose-化 leftovers (compose source, host facts, inventory, scribe notes) 2026-08-29 14:51:20 +08:00
windyboy ab808088b8 Add production Makefile for routine VPS ops
Wrap validate-repo.sh and routine Ansible playbooks with safe-by-default
targets: read-only health/audit flows, CONFIRM=1 gates for mutating work,
and LIMIT/TARGETS guards. Document entry point in AGENTS.md.
2026-08-26 11:27:02 +08:00
windyboy c7dc4fd25c chore: remove duplicate hook setup
Keep pre-commit configuration as the single validation path and tolerate deleted tracked Markdown during link validation.
2026-08-23 17:58:19 +08:00
windyboy 7bc7d3f99b docs: align runbooks and validation structure 2026-08-23 17:42:31 +08:00
windyboy fea9a6560f chore: gitignore .opencode/ and .zcode/ local tool caches 2026-08-23 17:10:07 +08:00
windyboy a95b626636 docs: Matter bulbs failure mode C — both bulbs announce mDNS but refuse TCP 5540 (08-23 read-only verification); record 08-22 add/loop saga, working bulb MAC change, 3-fabric map, PD rotation to 238:4812; refresh stale DHCP-reservation note on gw (W1N-207) 2026-08-23 11:06:35 +08:00
windyboy 27fe9c078e docs: archive historical planning documents and fix references
Move self-described historical/upstream docs to docs/archive/:
- agent-runbook-guide.md
- lan-core-switch-upgrade-plan.md
- lan-rb5009-upgrade.md
- se5420-review-claim-verification-2026-08.md

Update archive/README.md manifest and fix relative links in active docs
and archived docs. Update AGENTS.md docs/ layout description.
2026-08-22 19:34:05 +08:00
windyboy aaa4ee312e docs: verify gw switch0 as limited capture point — SE5420 single-uplink (eth1 up, eth2/3 down), LAN55 wired hosts behind SE5420; record EdgeOS 3 CLI/access quirks (W1N-207) 2026-08-22 10:44:20 +08:00
windyboy 6b298491a8 docs: Matter packet-capture runbook v2 — hass end0 commissioner capture point, corrected AP-point scope (no wired↔wired unicast), UAP-AC-Lite model fix, SE5420 live status; fix hass interface end1→end0 (W1N-207) 2026-08-22 10:14:29 +08:00
windyboy 32631e2996 docs: add Matter pairing troubleshooting handbook; record bulb state + DHCP reservation mismatch (W1N-207) 2026-08-22 08:18:23 +08:00
windyboy 1426b4ecfe docs: matrix_e2ee v0.3.12/v0.3.9 notes; gw/ubnt IPv6 re-verification; agent sandbox SSH quirk (2026-08-20) 2026-08-21 08:57:53 +08:00
windyboy 1f5e58bf17 docs: record Matter/IPv6 findings — stale matter-server mDNS address, SSID cleanup, ER-X ULA infeasibility (W1N-207) 2026-08-21 08:56:26 +08:00
windyboy e5819eeba3 docs: correct hass hardware to x88 Pro physical box; sync CSG v1.3.2
- hass.windy.lan is a physical x88 Pro box (HAOS bare-metal, machine: green,
  CPE x88pro20, virtualization empty) — not PVE VM 180 (verified live 2026-08-18)
- Record CSG v1.3.2 (934f58c, W1N-118) deploy in maintenance runbook verify
  section and hosts live-tree section (backups now include w1n118)
2026-08-18 13:32:33 +08:00
windyboy 079332e082 docs: record matrix_e2ee v0.3.0 deploy; fix and rename matrix-e2ee update runbook
- Deploy v0.3.0 (main 216cc99, W1N-180 bot-initiated device verification
  wizard) on hass.windy.lan; backup matrix_e2ee.bak-20260818-v0.2.10
- Fix runbook: tag-only prerequisite (v0.3.0 was untagged), working-tree
  HEAD check, rsync exit-23 note, actual setup log line, post-deploy record
  step, ssh_config.d -F /dev/null gotcha
- Rename runbook matrix-e2e-update.md -> matrix-e2ee-update.md and update
  AGENTS.md/hosts references (domain is matrix_e2ee, double-e)
- Unify matrix_e2ee naming and update version history in
  docs/home-assistant-matrix.md
2026-08-18 13:31:15 +08:00
windyboy 7cedba7f51 docs: rename integration name from matrix_e2ee to matrix_e2e
- Rename runbook: matrix-e2ee-update.md -> matrix-e2e-update.md
- Update all references in AGENTS.md, hass.windy.lan.md,
  home-assistant-matrix.md to use the short name matrix_e2e
- The code domain stays matrix_e2ee (E2EE) in source; all
  doc prose and command references now use matrix_e2e
2026-08-18 13:31:15 +08:00
windyboyandCursor 343c5db415 feat: add gated Compose deploy and make inventory the host source of truth
Keep sanitized Compose sources in-repo with a confirmation-gated Ansible
playbook, add repo-wide validation, tighten runbook ownership/STOP/review
metadata, and archive stale research docs.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-17 17:36:39 +08:00
windyboy 885d977531 docs(runbooks): light-enhance home-assistant-maintenance and index it
Rebase onto origin/main surfaced runbooks/home-assistant-maintenance.md
(W1N-69) which predated the runbook reorg. Add Purpose/Scope/Safety
headers and add it to the README routing index (now 17/17 consistent).
2026-08-17 16:02:33 +08:00
windyboy b0c01b2551 docs(runbooks): add runbook spec, template, index and 6 first-batch runbooks; light-enhance existing 10
- RUNBOOKS.md: repo-level spec (six-field model, naming, safety, maturity path)
- runbooks/_template.md + README.md: standard template and 16-entry routing index
- new: issue-to-merge, fix-ci, release, rollback, network-change, network-recovery
- light-enhance 10 existing runbooks with Purpose/Scope/Safety headers
- AGENTS.md: point step 3 at index/spec, add runbook execution rules
- docs/agent-runbook-guide.md: archive of Manus AI guide
2026-08-17 15:59:46 +08:00
windyboy 047ac03346 chore(skills): remove vendored encrypted-dns-skill (installed globally via skills CLI) 2026-08-15 12:35:55 +08:00
windyboy 1ec9246156 docs: treat ha backups list as ignored arg, not a subcommand
ha backups --help has no list; extra positional args still print
the default backup list with exit 0. Keep the ha host update fix.
2026-08-14 22:44:02 +08:00
windyboy 1936b8f5fe docs: correct ha backups list CLI note in HA runbook
ha backups list exists on this host; only ha host update is missing.
2026-08-14 22:42:44 +08:00
windyboy eda6536ddb docs: record CSG v1.3.1 zip install and correct ha-maintenance restart failure
ha-maintenance.sh --restart-core --yes exited 1 in <1s without restarting
Core. Empty output is ssh failure hidden by 2>/dev/null + pipefail, not a
MOTD-strip after a successful restart. Direct `ha core restart` is the
working path.
2026-08-14 22:31:31 +08:00
windyboy 1dc880362d docs: add HA Matrix integration notes and record .local rewrite removals 2026-08-14 18:17:26 +08:00
windyboy 70aea6cd72 feat(ednsdiag): add DoQ/DoH3/DNSCrypt transports, proxy support, probe & compare 2026-08-14 18:17:26 +08:00
windyboy 8303d78caf Record W1N-105 CSG network step, P1, and fork master→main rename on hass.windy.lan. 2026-08-14 18:15:28 +08:00
windyboy ebfe7b8488 docs(gw): document EdgeOS PPPoE redial procedure 2026-08-14 16:26:23 +08:00
windyboy 88eaefda33 Record W1N-104 CSG auto dual-stack deploy on hass.windy.lan.
end1 IPv6 is on, wlan0 stays off, and the live custom component is de01914
with ip_family=auto after the IPv4 blackhole.
2026-08-14 16:11:03 +08:00
windyboy fcb76d3d5a Record W1N-102 CSG deploy and IPv4 blackhole on hass.windy.lan.
The host note now states the aiohttp IPv4 client is live, Core loaded it,
and home PPPoE IPv4 to 95598.csg.cn is currently blackholed while IPv6
works on gw. HA still has IPv6 disabled (W1N-85).
2026-08-14 15:47:08 +08:00
windyboy 6ae835037b hass.windy.lan: resolve health snapshot issues (W1N-70..76) and document findings
- Add tianqi weather recorder patch notes (W1N-75: _unrecorded_attributes)
- Document Bluetooth hci0 RTL8821CS instability (W1N-74) and eMMC lifetime
  10% (W1N-76) as known issues
- Rewrite runbook known-issues section: all snapshot items resolved; link
  hosts doc for the two remaining known issues
2026-08-13 19:48:18 +08:00
windyboy b5617fd3a9 docs(ha): add Home Assistant maintenance runbook + ha CLI script (W1N-69)
- runbooks/home-assistant-maintenance.md: access pattern (sudo -n -i ha),
  command reference verified on host, recovery ops, families not scripted,
  docs-vs-CLI discrepancies
- runbooks/scripts/ha-maintenance.sh: read-only health/logs + --yes-gated
  update/restart/rebuild/rollback/reboot/backup/restore/app modes
- AGENTS.md: register runbook in table; hosts/hass: access pattern + link
2026-08-13 17:54:33 +08:00
windyboy f5842568b9 docs: add Home Assistant API access notes 2026-08-13 17:20:49 +08:00
windyboy 6f8a4918f0 feat(ednsdiag): support custom DoH endpoint via --url
Allow overriding the provider preset with an explicit HTTPS DoH URL,
including validation that custom endpoints apply only to DoH queries.
2026-08-13 15:34:54 +08:00
windyboy 5b0f7950e6 docs: add hass.windy.lan Home Assistant host documentation
Document HAOS on PVE VM 180 (LAN55) with verified SSH access,
network details, and cross-links from lan-overview, inventory,
and AGENTS quick map.
2026-08-13 14:10:39 +08:00
windyboy c0cf82d4af tools(skills): add encrypted-dns-skill (ednsdiag CLI + agent skill) 2026-08-13 12:52:20 +08:00
windyboy 95ec2350af docs(dns): record mosdns foreign DoH multi-upstream redundancy (W1N-62) 2026-08-13 10:56:40 +08:00
windyboy 3de4beb028 docs: add low-volume mono laser MFP buying guide (2026-08)
Decision tree for occasional B&W laser MFP purchases: Brother L1638W/L1848W
as default, cloud-subscription models as opt-in only, and one-veto checks
for AirPrint, Ethernet, and duplex/ADF needs.
2026-08-13 10:54:29 +08:00
windyboy d54ec71aea docs(dns): record gfw foreign branch DoH change (W1N-62)
Document encrypted DoH upstream for mosdns foreign queries and note that
DoH traffic goes direct to hk2, not via OpenClash proxy.
2026-08-13 10:50:16 +08:00
windyboy 2ffd9f9f9c docs(se5420): align review claim verification with current guide (W1N-63)
Add historical snapshot header (baseline 35577d0), rewrite outdated
"current guide" assertions for post-ffb37a9 revisions, and add a
12-row status table mapping review claims to current §4.3/§11 sections.
2026-08-13 10:50:16 +08:00
windyboy f255785b72 docs(se5420): add review claim verification record (2026-08-10)
Documents which deployment-guide review claims are confirmed by specs,
field read-only checks on gfw, and remaining pre-change evidence needs.
2026-08-13 10:12:06 +08:00
windyboy 2fd354c2a9 docs(dns): record mosdns fallback hardening for AGH outage (W1N-56) 2026-08-12 22:24:54 +08:00
windyboy 82203038f0 docs(dns): record mosdns sequence misconfig found+fixed (W1N-56) 2026-08-12 22:19:20 +08:00
windyboy 096e1ce8b6 docs(dns): correct idle-mosdns premise in alternatives research (W1N-56) 2026-08-12 22:12:32 +08:00
windyboy 8550053287 docs(dns): record Phase 0 verification evidence + final decision alignment (W1N-56) 2026-08-12 22:11:53 +08:00
windyboy e501b93d65 docs(dns): correct gfw/.1 facts (W1N-56)
- hosts/gfw.windy.lan.md: 3 NICs (eth2/VLAN10 ubunt_upg live), mosdns is
  now OpenClash's nameserver (not idle), rewrite VLAN10 Wi-Fi section to
  live-verified state
- docs/lan-dns-architecture.md: mosdns on gfw no longer 闲置; note the
  recommended AGH+.36 companion architecture is still pending review
2026-08-12 22:11:53 +08:00
windyboy 62b8fbb8b7 docs(dns): add LAN DNS architecture research + recommendation (W1N-56) 2026-08-12 22:11:53 +08:00
windyboy efa6cf0899 docs(dns): record agh_ui_access LAN55 allow for Home Assistant (2026-08-12) 2026-08-12 22:11:53 +08:00
windyboy 086740b16e docs: record pdns PDA removal, us4 firewalld ops, LAN DNS alternatives
- runbooks/pdns-health.md: note the legacy powerdns-admin (PDA) orphan was
  removed 2026-08-12 (W1N-59).
- runbooks/ansible-operations.md: document the us4 firewalld reconciliation
  playbook scope (audited public zone only, fail-closed, no reload).
- docs/agents/domain.md: single-context repo layout for domain docs.
- docs/lan-dns-alternatives.md: notes on LAN DNS alternatives.
- .gitignore: exclude local agent-harness config (.agents/ .claude/ .omp/
  .mcp.json WATCHDOG.yml skills-lock.json) from the repo.
2026-08-12 21:16:31 +08:00
windyboy 1f6d028ab5 feat(us4): firewall audit + safe reconciliation playbook, host doc
Add playbooks/us4-firewalld.yml, a narrow reconciliation of the audited us4
public zone: fails closed on drift or unknown allowances, never reloads or
restarts firewalld, and does not manage Docker rules. Requires explicit
apply + provider-console confirmations, backs up the firewalld config and
ruleset, schedules an automatic 15-minute rollback via at, and verifies SSH,
HTTPS routes, containers, Fail2ban jails, and the WireGuard health check before
cancelling rollback. Pins ansible.posix 2.2.2 in requirements.yml.

Also expand hosts/us4.wsvc.info.md with deployment config and a live audit
snapshot (2026-08-12).
2026-08-12 21:16:31 +08:00
windyboy 035587e3bf feat(rustdesk): onboard self-hosted RustDesk server on hk2
Deploy hbbs + hbbr via a safe-by-default Ansible role (playbooks/rustdesk.yml):
with rustdesk_confirm=false it only reports whether compose.yml matches live
state and refuses to recreate the stack; with rustdesk_confirm=true it deploys
and recreates. The relay assert rejects the known-bad hk2.wsvc.info hostname.

Add health profiles rustdesk (hbbs/hbbr health, relay DNS) and hk2aux (co-located
traefik/adguard/remark42 on hk2), plus the rustdesk-health runbook and AGENTS.md
entry. Server image pinned rustdesk/rustdesk-server:1.1.14.
2026-08-12 21:16:30 +08:00
windyboy e7296e664a feat(ansible): move hosts to healthcheck_profiles list; decouple audit/restic
Inventory now declares the plural healthcheck_profiles list per host (hk2 runs
pdns, rustdesk, hk2aux) instead of a single healthcheck_profile, and carries
the rustdesk server vars (rustdesk_compose_dir/relay/image) plus the rustdesk
host group.

The restic role relied on the removed singular healthcheck_profile var; it now
uses its own restic_backup_profile (set per host to vaultwarden on us2 and pdns
on hk2), so the health-check rename no longer breaks it. audit.yml's summary
labels the host's profile list instead of the singular var.
2026-08-12 21:16:30 +08:00
windyboy d0d5e5a704 refactor(healthcheck): support multiple profiles per host with aggregate result
Replace the single healthcheck_profile with a healthcheck_profiles list so a
host can run several checks (e.g. hk2: pdns, rustdesk, hk2aux). Profiles emit
per-check JSON to latest-<check>.json; the dispatcher clears stale per-check
files, runs every profile, and merges them into latest.json.

Contract: a single complete check keeps the historical verbatim latest.json
shape; several checks produce a worst-status aggregate retaining every check's
detail. A profile that crashes before reporting is aggregated as unknown so
latest.json can never go stale while the dispatcher fails. The dispatcher exits
with the worst (max) profile exit code.
2026-08-12 21:16:30 +08:00
windyboy baca89be83 docs(hk2): document Traefik dashboard auth and password rotation 2026-08-12 21:16:30 +08:00
110 changed files with 8762 additions and 305 deletions
+13
View File
@@ -15,8 +15,21 @@ id_*
.ansible/
facts/
# Local agent-harness / tooling config (not repo content).
.agents/
.claude/
.omp/
.opencode/
.zcode/
.mcp.json
WATCHDOG.yml
skills-lock.json
# Editor and operating-system files.
.DS_Store
.vscode/
.idea/
*~
# Agent working scratch (not repo content).
.agent-work/
.tmp-*
+8
View File
@@ -0,0 +1,8 @@
repos:
- repo: local
hooks:
- id: validate-repo
name: validate repository
entry: scripts/validate-repo.sh
language: system
pass_filenames: false
+79 -23
View File
@@ -2,36 +2,58 @@
This repo is the **agent ops handbook + fact source** for maintaining personal VPS hosts. Prefer verifying live state over assuming docs are complete.
Also readable as `agent.md` (symlink → this file).
## How to work
1. Read [`inventory/hosts.md`](inventory/hosts.md) for the machine list.
2. Open the matching [`hosts/<name>.md`](hosts/) for SSH, roles, paths, and quirks.
3. For common tasks, follow a runbook under [`runbooks/`](runbooks/).
3. For common tasks, follow a runbook under [`runbooks/`](runbooks/). Pick the
most specific applicable one from [`runbooks/README.md`](runbooks/README.md);
the spec is [`RUNBOOKS.md`](RUNBOOKS.md) and new runbooks start from
[`runbooks/_template.md`](runbooks/_template.md).
4. Prefer read-only checks first; change only after confirming current state.
5. For routine checks and approved service reconciliation, run the matching
Ansible playbook from `ansible/`; see [routine Ansible operations](runbooks/ansible-operations.md).
6. Default SSH access (`ssh -4 windy@<host>`) is for focused diagnostics,
imperative upstream procedures, and incident work. Prefer **IPv4** from this
WSL client (AAAA often exists but IPv6 route does not).
> **Agent sandbox SSH quirk (verified 2026-08-20):** the agent shell runs in
> a sandboxed user namespace — system files such as
> `/etc/ssh/ssh_config.d/20-systemd-ssh-proxy.conf` appear owned by `nobody`,
> so plain `ssh` aborts with `Bad owner or permissions on ...`. Always use
> `ssh -F /dev/null` from the agent shell and pass options explicitly
> (`~/.ssh/config` is skipped; e.g. `ssh -F /dev/null -p 2222
> -i ~/.ssh/id_ed25519 windy@repo.windy.me`). `sudo` never works in the
> sandbox (`NoNewPrivs`, no capabilities, `/` read-only). The host itself is
> healthy — to inspect or act on the real host from the sandbox use
> `/mnt/c/WINDOWS/system32/wsl.exe -u root -- <cmd>` (real root: keep
> read-only unless a change is approved).
7. Record each material VPS operation, incident, configuration change, or
verification outcome in the corresponding **Linear `vps` project**. Include
scope, action, verification, and remaining follow-up; never put passwords,
tokens, private keys, recovery keys, or private room IDs in Linear.
verification outcome in the corresponding **Plane `vps` project**
(self-hosted `plane.chans.xyz`, Plane MCP `mcp__plane__*`, following the
`plane-workflow` skill). **Linear is retired as a record source (2026-09-03)
— do not create Linear issues;** existing W1N-* entries are read-only
history. Include scope, action, verification, and remaining follow-up; never
put passwords, tokens, private keys, recovery keys, or private room IDs in
Plane or Linear.
## Active hosts (quick map)
### Runbook execution rules
| Host | Role | SSH | Facts |
|------|------|-----|--------|
| **mx2.windy.me** | mailcow (`/opt/mail`, project `cow`) | `ssh -4 windy@mx2.windy.me` | [hosts/mx2.windy.me.md](hosts/mx2.windy.me.md) |
| **us2.wsvc.info** | Vaultwarden + Traefik (+ Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | [hosts/us2.wsvc.info.md](hosts/us2.wsvc.info.md) |
| **hk2.chans.xyz** | PowerDNS auth ns1 (`/opt/pdns`) | `ssh -4 windy@hk2.chans.xyz` | [hosts/hk2.chans.xyz.md](hosts/hk2.chans.xyz.md) |
| **synapse.chans.xyz** | Matrix ESS (Synapse + MAS + Element) on K3s | `ssh -4 windy@synapse.chans.xyz` | [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy | `ssh -4 windy@192.168.66.36` | [hosts/dns.windy.lan.md](hosts/dns.windy.lan.md) |
| **gfw.windy.lan** | OpenWrt LAN gateway / OpenClash | `ssh -4 root@192.168.66.1` | [hosts/gfw.windy.lan.md](hosts/gfw.windy.lan.md) |
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | [hosts/gw.md](hosts/gw.md) |
| **ubnt** | UniFi Network Controller | `ssh -4 windy@192.168.66.46` | [hosts/ubnt.md](hosts/ubnt.md) |
Before operational work: inspect `runbooks/`, select the most specific
applicable runbook, follow its steps in order, do not skip verification steps,
and respect its STOP and approval conditions. If no runbook applies, diagnose
only — do not mutate production state. When live state conflicts with a
runbook's assumptions, `STOP` and report; never invent missing parameters or
bypass failed checks. The spec is [`RUNBOOKS.md`](RUNBOOKS.md).
## Active hosts
The canonical machine list (roles, SSH endpoints, Ansible coverage, status) is
[`inventory/hosts.md`](inventory/hosts.md) — the single human-readable source
of truth. Per-host facts live in [`hosts/`](hosts/). The Ansible execution
inventory is [`ansible/inventory/hosts.yml`](ansible/inventory/hosts.yml). Do
not maintain a second copy of the machine table here.
### Public services
@@ -41,7 +63,7 @@ Also readable as `agent.md` (symlink → this file).
| SMTP `mx2.windy.me:587` (STARTTLS) or `:465` | mx2 | client submission; full email + mailbox password — [runbook](runbooks/mailcow-smtp-client.md) |
| IMAP `mx2.windy.me:993` | mx2 | same mailbox credentials |
| https://auth.wsvc.info | us2 (`/opt/vaultwarden`) | Vaultwarden (Postgres, **operational**) — client Server URL |
| `repo.windy.me:2222` | us2 (`/opt/soft-serve`) | Soft Serve (stub details) |
| `repo.windy.me` (git SSH `:2222` / web HTTPS) | us2 (`/opt/gitea`) | Gitea — 1.27.3-rootless pinned, backup sidecar; details in [hosts/us2.wsvc.info.md](hosts/us2.wsvc.info.md) |
| DNS `ns1.wsvc.info:53` | hk2 (`/opt/pdns`, Auth **5.0.6**) | PowerDNS auth — zones `windy.me`, `wsvc.info`, `chans.xyz` |
| https://pdns.wsvc.info | hk2 (`poweradmin`) | Poweradmin UI |
| https://pgweb.wsvc.info | hk2 (`pgweb`) | PowerDNS Postgres browser |
@@ -49,6 +71,7 @@ Also readable as `agent.md` (symlink → this file).
| https://synapse.chans.xyz | synapse | Synapse Client-Server + Federation API |
| https://account.chans.xyz | synapse | Matrix Authentication Service (local passwords) |
| https://admin.chans.xyz | synapse | Element Admin console (MAS admin auth) |
| https://plane.chans.xyz | synapse (`plane`, Helm `plane-ce` 1.8.0 / v1.4.1) | Plane project management (self-hosted, K3s) |
### Upstream docs
@@ -58,6 +81,10 @@ Also readable as `agent.md` (symlink → this file).
**Matrix (ESS on synapse):** Matrix homeserver running on `synapse.chans.xyz` via the official ESS (Element Server Suite) Helm chart with Synapse + MAS + Element Web + Admin. DNS zone `chans.xyz` managed by hk2 PowerDNS. Before changing config, read [docs/matrix-upstream.md](docs/matrix-upstream.md) and [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md). K3s cluster on this node has hostPort 80/443 for Traefik (no ServiceLB). Health: [matrix-health](runbooks/matrix-health.md).
**Plane (on synapse):** Self-hosted Plane project management at `plane.chans.xyz`, Helm release `plane-app` (chart `plane-ce-1.8.0`, app `v1.4.1`) in ns `plane` on the same K3s node as Matrix. Config from `/home/windy/plane-k3s/values.yaml`; workload/cert/ingress details in [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md). Its Postgres/MinIO PVCs are **not** backed up.
**RustDesk:** Self-hosted RustDesk server on `hk2.chans.xyz` (`/opt/rustdesk`, containers `hbbs`/`hbbr`, image pinned `1.1.14`). The `hbbs -r` relay hostname must resolve to the host's public IP `154.36.174.161` — use `hk2.chans.xyz` (never `hk2.wsvc.info`, which has no DNS record). Health: [rustdesk-health](runbooks/rustdesk-health.md).
## Runbooks & scripts
| Task | Path |
@@ -71,12 +98,26 @@ Also readable as `agent.md` (symlink → this file).
| PowerDNS health (hk2) | [runbooks/pdns-health.md](runbooks/pdns-health.md) |
| PowerDNS upstream refs | [docs/pdns-upstream.md](docs/pdns-upstream.md) |
| Matrix health | [runbooks/matrix-health.md](runbooks/matrix-health.md) |
| Plane health | [runbooks/plane-health.md](runbooks/plane-health.md) |
| RustDesk health (hk2) | [runbooks/rustdesk-health.md](runbooks/rustdesk-health.md) |
| AdGuard Home health | [runbooks/adguard-home-health.md](runbooks/adguard-home-health.md) |
| Host disk cleanup | [runbooks/host-disk-cleanup.md](runbooks/host-disk-cleanup.md) |
| Home Assistant maintenance | [runbooks/home-assistant-maintenance.md](runbooks/home-assistant-maintenance.md) + [scripts/ha-maintenance.sh](runbooks/scripts/ha-maintenance.sh) |
| matrix_e2ee update (hass.windy.lan) | [runbooks/matrix-e2ee-update.md](runbooks/matrix-e2ee-update.md) |
| Matrix upstream refs | [docs/matrix-upstream.md](docs/matrix-upstream.md) |
| Hermes Agent Matrix channel | [docs/hermes-matrix.md](docs/hermes-matrix.md) |
| UniFi local-service proxy bypass | [docs/unifi-openclash-localhost.md](docs/unifi-openclash-localhost.md) |
| UniFi SSO login setting (Ansible) | `cd ansible && ansible-playbook playbooks/unifi-sso.yml --limit unifi` |
| Routine Ansible operations | [runbooks/ansible-operations.md](runbooks/ansible-operations.md) |
| Routine make commands | `make help` (wraps `ansible-operations.md` read-only + gated flows) |
| Issue → mergeable change | [runbooks/issue-to-merge.md](runbooks/issue-to-merge.md) |
| Fix failing health/playbook run | [runbooks/fix-ci.md](runbooks/fix-ci.md) |
| Release a reviewed change | [runbooks/release.md](runbooks/release.md) |
| Roll back a change | [runbooks/rollback.md](runbooks/rollback.md) |
| Controlled network change | [runbooks/network-change.md](runbooks/network-change.md) |
| Network outage recovery | [runbooks/network-recovery.md](runbooks/network-recovery.md) |
Full index: [runbooks/README.md](runbooks/README.md). Spec: [RUNBOOKS.md](RUNBOOKS.md).
Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.yml --limit mailcow`. The local stub resolver is flaky; DNS probes use `1.1.1.1` / `8.8.8.8`.
@@ -84,12 +125,22 @@ Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.
### Issue tracker
Issues are tracked in Linear and created/updated via the Linear MCP (`vps` project). See `docs/agents/issue-tracker.md`.
Issues are tracked in **Plane** — self-hosted at `plane.chans.xyz`, project
`vps` — and created/updated via the Plane MCP (`mcp__plane__*`), following the
`plane-workflow` skill. **Linear is retired as a record source (2026-09-03); do
not create Linear issues.** Existing W1N-* entries are read-only history.
`docs/agents/issue-tracker.md` documents the retired Linear workflow and is
stale; treat this section as authoritative.
### Triage labels
Default triage labels: needs-triage, needs-info, ready-for-agent, ready-for-human, wontfix. See `docs/agents/triage-labels.md`.
### Domain docs
Domain-documentation conventions, including lazily created `CONTEXT.md` and
`docs/adr/` entries when needed, are described in [`docs/agents/domain.md`](docs/agents/domain.md).
## Safety
- Never commit secrets: passwords, API keys, private keys, `.env`, `mailcow.conf` DB passwords, Vaultwarden `ADMIN_TOKEN` / `.smtp-credentials`.
@@ -121,9 +172,14 @@ Bills, rough notes, and personal clutter stay in the Obsidian vault. This repo h
## Layout
```
AGENTS.md / agent.md # this entry (agent.md → AGENTS.md)
inventory/hosts.md # machine index
AGENTS.md # this entry
RUNBOOKS.md # runbook spec (six-field model, naming, review rules)
inventory/hosts.md # machine index (human-readable source of truth)
ansible/ # playbooks, roles, sanitized control-plane inventory
compose/ # repo-owned non-secret Compose sources (+ .env.example)
hosts/ # per-host facts
runbooks/ # step-by-step ops
docs/ # upstream doc indexes / design notes
runbooks/ # step-by-step ops (README.md = index, _template.md = template)
docs/ # upstream refs / design notes / research records (active + archive/)
scripts/validate-repo.sh # repo-wide validation (run before merging)
Makefile # routine validate / health / gated ansible wrappers
```
+183
View File
@@ -0,0 +1,183 @@
# VPS ops hub — routine validate / health / gated Ansible wrappers.
# See runbooks/ansible-operations.md for playbook semantics.
SHELL := /usr/bin/env bash
.SHELLFLAGS := -eu -o pipefail -c
.DEFAULT_GOAL := help
REPO_ROOT := $(CURDIR)
ANSIBLE_DIR := $(REPO_ROOT)/ansible
export ANSIBLE_LOCAL_TEMP := $(REPO_ROOT)/.ansible/tmp
export ANSIBLE_HOME := $(REPO_ROOT)/.ansible
LIMIT ?=
EXTRA ?=
VERBOSE ?= 0
CONFIRM ?= 0
TARGETS ?=
TRAEFIK ?= 0
LIMIT_FLAG := $(if $(LIMIT),--limit $(LIMIT),)
VERBOSE_FLAG := $(if $(filter 1,$(VERBOSE)),-v,$(if $(filter 2,$(VERBOSE)),-vvv,))
.PHONY: help validate check deps galaxy syntax ansible-prep \
ping inventory audit health health-mailcow health-matrix \
maint-preview baseline compose-check \
install-healthchecks install-matrix-healthchecks compose-deploy reconcile
help:
@printf '%s\n' \
'VPS ops hub — make targets (run from repo root)' \
'' \
'Variables: LIMIT=<group|host> CONFIRM=1 TARGETS=<svc[,svc]> TRAEFIK=1 VERBOSE=0|1|2 EXTRA=...' \
'' \
'Local / repo:' \
' validate, check scripts/validate-repo.sh (pre-merge gate)' \
' deps, galaxy ansible-galaxy collection install' \
' syntax ansible-playbook --syntax-check all playbooks' \
'' \
'Read-only remote (ansible):' \
' ping ansible managed -m ping' \
' inventory ansible-inventory --graph' \
' audit playbooks/audit.yml' \
' health [LIMIT=…] playbooks/health-report.yml' \
' health-mailcow health --limit mailcow' \
' health-matrix health --limit matrix' \
' maint-preview playbooks/maintenance-preview.yml' \
' baseline playbooks/baseline.yml' \
' compose-check compose-deploy --check --diff (requires LIMIT=)' \
'' \
'Mutating (require CONFIRM=1; host-scoped targets require LIMIT=):' \
' install-healthchecks playbooks/healthchecks.yml' \
' install-matrix-healthchecks playbooks/matrix-healthchecks.yml' \
' compose-deploy playbooks/compose-deploy.yml' \
' reconcile playbooks/compose-reconcile.yml (requires TARGETS=)' \
'' \
'Examples:' \
' make validate' \
' make health LIMIT=mailcow' \
' make compose-check LIMIT=vaultwarden' \
' make compose-deploy LIMIT=vaultwarden CONFIRM=1' \
' make reconcile LIMIT=powerdns TARGETS=auth CONFIRM=1' \
' make reconcile LIMIT=vaultwarden TARGETS=vaultwarden TRAEFIK=1 CONFIRM=1' \
'' \
'Advanced (not wrapped — use ansible-playbook directly):' \
' us4-firewalld, unifi-sso, k3s-server, matrix-stack, wireguard-harden,' \
' restic, rustdesk, email-alerts, mailcow update runbook'
validate check:
@bash "$(REPO_ROOT)/scripts/validate-repo.sh"
deps galaxy: ansible-prep
@command -v ansible-galaxy >/dev/null 2>&1 || { echo "ansible-galaxy not found; install Ansible first." >&2; exit 1; }
@cd "$(ANSIBLE_DIR)" && ansible-galaxy collection install -r requirements.yml
syntax: ansible-prep
@if ! command -v ansible-playbook >/dev/null 2>&1; then \
echo "ansible-playbook not found; syntax check skipped." >&2; \
exit 0; \
fi
@fail=0; \
for p in "$(ANSIBLE_DIR)"/playbooks/*.yml; do \
if ! (cd "$(ANSIBLE_DIR)" && ansible-playbook --syntax-check "playbooks/$$(basename "$$p")" >/dev/null 2>&1); then \
echo "syntax-check failed: $$p" >&2; \
fail=1; \
fi; \
done; \
exit $$fail
ansible-prep:
@mkdir -p "$(ANSIBLE_HOME)/tmp" "$(ANSIBLE_HOME)/ssh-control"
define require_ansible
@command -v ansible-playbook >/dev/null 2>&1 || { echo "ansible-playbook not found; install Ansible first." >&2; exit 1; }
endef
define require_limit
@if [ -z "$(LIMIT)" ]; then \
echo "LIMIT is required (e.g. LIMIT=mailcow, LIMIT=vaultwarden, LIMIT=powerdns)." >&2; \
exit 1; \
fi
endef
define require_confirm
@if [ "$(CONFIRM)" != "1" ]; then \
echo "Mutating operation blocked. Re-run with CONFIRM=1" >&2; \
exit 1; \
fi
endef
define require_targets
@if [ -z "$(TARGETS)" ]; then \
echo "TARGETS is required (comma-separated service names, e.g. TARGETS=auth or TARGETS=vaultwarden)." >&2; \
exit 1; \
fi
endef
ping: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible managed -m ping $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
inventory: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-inventory --graph $(EXTRA)
audit: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/audit.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
health: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
health-mailcow: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml --limit mailcow $(VERBOSE_FLAG) $(EXTRA)
health-matrix: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml --limit matrix $(VERBOSE_FLAG) $(EXTRA)
maint-preview: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/maintenance-preview.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
baseline: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/baseline.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
compose-check: ansible-prep
$(require_ansible)
$(require_limit)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-deploy.yml --check --diff --limit $(LIMIT) $(VERBOSE_FLAG) $(EXTRA)
install-healthchecks: ansible-prep
$(require_ansible)
$(require_confirm)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/healthchecks.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
install-matrix-healthchecks: ansible-prep
$(require_ansible)
$(require_confirm)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/matrix-healthchecks.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
compose-deploy: ansible-prep
$(require_ansible)
$(require_limit)
$(require_confirm)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-deploy.yml --limit $(LIMIT) \
-e '{"compose_deploy_confirm": true}' $(VERBOSE_FLAG) $(EXTRA)
reconcile: ansible-prep
$(require_ansible)
$(require_limit)
$(require_targets)
$(require_confirm)
@json=$$(python3 -c 'import json,sys; t=[x.strip() for x in sys.argv[1].split(",") if x.strip()]; \
(not t) and sys.exit("TARGETS must contain at least one non-empty service name"); \
d={"service_reconcile_confirm": True, "service_reconcile_targets": t}; \
(sys.argv[2]=="1") and d.update({"service_reconcile_restart_traefik": True}); \
print(json.dumps(d))' "$(TARGETS)" "$(TRAEFIK)"); \
cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-reconcile.yml --limit $(LIMIT) \
-e "$$json" $(VERBOSE_FLAG) $(EXTRA)
+89
View File
@@ -0,0 +1,89 @@
# RUNBOOKS — 仓库级规范
本文件统一所有 Runbook 的字段、命名、评审与变更规则。上游参考:[docs/archive/agent-runbook-guide.md](docs/archive/agent-runbook-guide.md)。
## 目录结构
```text
runbooks/
├── README.md # 意图 → 文件 路由索引(本目录的入口)
├── _template.md # 新建 runbook 的标准模板(复制后填写)
├── <intent>.md # 每份 runbook 只描述一种可识别的操作意图
└── ...
```
## 最小字段模型
每份 runbook 必须显式包含以下控制信息,否则盲目执行或错误恢复的风险会升高:
| 字段 | 作用 | 写作要求 |
|---|---|---|
| **Action** | 定义当前要执行的动作 | 可观察、可执行的动词;避免“检查一下”“适当调整” |
| **Expected** | 描述正常状态或预期输出 | 具体信号、阈值、状态码、测试结果或页面表现 |
| **Decision** | 定义分支与下一跳 | “条件 → 下一步”;无法判断时指向 `STOP` |
| **Verification** | 确认变更真正生效 | 每个有副作用的步骤后执行,不可跳过 |
| **Stop condition** | 规定何时不得继续 | 列出信息缺失、状态冲突、权限不足、验证失败等 |
| **Rollback** | 如何恢复到变更前状态 | 触发条件、前提、撤销步骤、回滚后验证 |
> 只读类 runbook 不产生副作用,可省略 Rollback;但必须保留 Stop condition(状态与预期冲突即 `STOP` 并记录证据)。
**只读类变体(read-only variant**:只读 runbookhealth 类、参考类)不强制
六字段模型,但必须包含以下最小结构,否则不视为达标:
- `## Purpose`12 行)+ `## Scope`(适用/不适用)
- `## Safety` 或等效章节,其中**必须**含显式 Stop condition(状态与预期冲突即
`STOP` 并记录证据;不得在执行中自行"顺手修复")
- 只读健康类另含可观察的 `## Pass criteria`(或等效的 Expected 信号)
- 每份 runbook 顶部/元信息区必须标注 `Last reviewed: <YYYY-MM-DD>`
## 命名与拆分规则
- 文件名采用小写连字符,反映**操作意图**而非目标主机,例如 `mailcow-health.md``release.md`
- 一份文件只描述一种意图。流程出现明显分叉时拆分为独立文件,不堆叠“万能流程”。
- 只读诊断与变更操作应分离:health 类 runbook 保持只读,变更走 `ansible-operations.md``release.md``rollback.md` 或对应 gated playbook。
## 章节约定
- 每份 runbook 顶部含 `## Purpose`12 行)与 `## Scope`(适用/不适用情形)。
- 变更型 runbook 必须记录明确的审批门:门控命令式使用 `## Approval gates` 表;
流程式在步骤中记录审批动作、证据位置和未批准时的 `STOP`。破坏性/不可逆操作必须获得明确批准。
- 语言约定:**runbook 正文统一使用英文**(由 agent 逐字执行,降低二义性);
元规范文件(AGENTS.md / RUNBOOKS.md / 模板注释)可保留中文。
- 变更型 runbook 的两种形态:
- **流程式(Procedure 型)**:使用六字段模型,适用多分支/多步骤变更
(现有:`fix-ci.md``issue-to-merge.md``network-change.md`
`network-recovery.md``release.md``rollback.md`)。
- **门控命令式(gated command reference**:已稳定、低歧义、可验证的
操作以命令集 + 门控呈现(现有:`mailcow-update.md`
`ansible-operations.md``home-assistant-maintenance.md`
`matrix-e2ee-update.md``vaultwarden-sqlite-to-postgres.md`),必须含 Approval gates 或确认变量
要求 + 显式 STOP,不替代流程式形态。新写的变更 runbook 默认用流程式。
- 统一在 `## Safety` 或正文中复用以下通用安全规则(更严格要求优先)。
```markdown
## Safety Rules
- Never delete an existing configuration as the first recovery action.
- Prefer read-only diagnosis before mutation.
- After every mutation, verify the expected state.
- If actual state conflicts with this runbook, STOP.
- Do not invent missing parameters.
- Do not bypass failed tests.
- Destructive actions require explicit approval.
```
## 评审与变更规则
- 新建/修改 runbook 与代码同仓评审,随系统演进更新。
- 每份 runbook 标注 `Last reviewed`;流程执行过程中发现的偏差记入对应的 Linear `vps` 项目 issue。
- 破坏性流程(迁移、删除、DNS 变更、网络变更)保持人工审批,不自动下沉。
## 成熟路径
1. **人工处理** → 现场处置与复盘,记录证据。
2. **Markdown runbook** → 固化步骤与证据要求,Agent 可辅助诊断。
3. **Agent + runbook** → 严格按流程执行,受 Stop/Approval 约束。
4. **Script / Ansible / Skill** → 把已稳定、低歧义、可验证的操作程序化(本仓库的执行层是 Ansible playbook)。
5. **人工审批 + 自动执行** → 审批门控下的自动变更(如 gated playbook + 确认变量)。
原则:先证据后变更,先小范围后扩大,先验证后结束,不确定则停止。
+5
View File
@@ -11,3 +11,8 @@ host_key_checking = True
become = True
become_method = sudo
become_ask_pass = False
[ssh_connection]
# Keep SSH control sockets inside the repo (gitignored .ansible/) so playbook
# runs work in sandboxed/CI environments without touching ~/.ansible.
ssh_args = -C -o ControlMaster=auto -o ControlPersist=60s -o ControlPath=.ansible/ssh-control/%h-%p-%r
+25 -5
View File
@@ -15,18 +15,23 @@ all:
mx2:
ansible_host: mx2.windy.me
ansible_host_ipv4: 194.163.160.244
display_name: mx2.windy.me
service_role: mailcow
compose_project_dir: /opt/mail
healthcheck_profile: mailcow
healthcheck_profiles: [mailcow]
service_reconcile_services:
all:
compose_args: [--force-recreate]
us2:
ansible_host: us2.wsvc.info
ansible_host_ipv4: 193.9.44.165
display_name: us2.wsvc.info
service_role: vaultwarden
compose_project_dir: /opt/vaultwarden
healthcheck_profile: vaultwarden
compose_repo_project: vaultwarden
compose_remote_file: docker-compose.yml
healthcheck_profiles: [vaultwarden]
restic_backup_profile: vaultwarden
service_reconcile_services:
vaultwarden:
compose_args: [--force-recreate]
@@ -34,9 +39,13 @@ all:
hk2:
ansible_host: hk2.chans.xyz
ansible_host_ipv4: 154.36.174.161
display_name: hk2.chans.xyz
service_role: powerdns
compose_project_dir: /opt/pdns
healthcheck_profile: pdns
compose_repo_project: pdns
compose_remote_file: compose.yml
healthcheck_profiles: [pdns, rustdesk, hk2aux]
restic_backup_profile: pdns
service_reconcile_services:
auth:
compose_args: [--force-recreate]
@@ -45,12 +54,17 @@ all:
backup:
compose_args: [--no-deps, --force-recreate]
service_reconcile_traefik_restart_targets: [poweradmin]
# RustDesk server (same host, separate compose project)
rustdesk_compose_dir: /opt/rustdesk
rustdesk_relay: hk2.chans.xyz:21117
rustdesk_image: rustdesk/rustdesk-server:1.1.14
us4:
ansible_host: us4.wsvc.info
ansible_host_ipv4: 185.201.226.122
display_name: us4.wsvc.info
service_role: wireguard
compose_project_dir: /opt/wireguard
healthcheck_profile: wireguard
healthcheck_profiles: [wireguard]
wireguard_image: >-
lscr.io/linuxserver/wireguard@sha256:ac43e1226878d2611315172d6ea357a95cb326ee73124b91108118efc8666889
service_reconcile_services:
@@ -59,9 +73,10 @@ all:
dns_windy_lan:
ansible_host: 192.168.66.36
ansible_host_ipv4: 192.168.66.36
display_name: dns.windy.lan
service_role: adguardhome
compose_project_dir: /opt/adguardhome
healthcheck_profile: adguardhome
healthcheck_profiles: [adguardhome]
service_reconcile_services:
adguardhome:
compose_args: [--no-deps, --force-recreate]
@@ -74,6 +89,9 @@ all:
powerdns:
hosts:
hk2:
rustdesk:
hosts:
hk2:
wireguard:
hosts:
us4:
@@ -85,6 +103,7 @@ all:
ubnt:
ansible_host: 192.168.66.46
ansible_host_ipv4: 192.168.66.46
display_name: ubnt
vars:
service_role: unifi
compose_project_dir: /home/windy/unifi-9
@@ -105,6 +124,7 @@ all:
matrix_vps:
ansible_host: 169.58.86.13
ansible_host_ipv4: 169.58.86.13
display_name: synapse.chans.xyz
service_role: matrix_k3s
matrix_server_name: chans.xyz
matrix_synapse_host: synapse.chans.xyz
+1 -1
View File
@@ -64,7 +64,7 @@
ansible.builtin.debug:
msg:
host: "{{ inventory_hostname }}"
profile: "{{ healthcheck_profile }}"
profiles: "{{ healthcheck_profiles | default([]) | join(', ') }}"
os: "{{ ansible_distribution }} {{ ansible_distribution_version }}"
kernel: "{{ ansible_kernel }}"
compose_rc: "{{ audit_compose_ps.rc }}"
+13
View File
@@ -0,0 +1,13 @@
---
# Deploy repo-owned Compose declarations (compose/<project>/compose.yml) to
# inventory hosts. Non-secret source; server-local .env provides the values.
# Gated: apply requires compose_deploy_confirm=true; --check is a read-only
# diff + validation. See runbooks/ansible-operations.md.
- name: Deploy repo-owned Compose declarations
hosts: docker_hosts
become: true
gather_facts: false
serial: 1
roles:
- role: compose_deploy
tags: [compose, deploy, mutating]
+20
View File
@@ -0,0 +1,20 @@
---
# Deploy/reconcile the self-hosted RustDesk server (hbbs + hbbr) on hk2.
#
# Safe by default: run with --check for a read-only report, or supply
# rustdesk_confirm=true to deploy the compose file and recreate the stack.
#
# # Read-only report
# ansible-playbook playbooks/rustdesk.yml --limit rustdesk --check
#
# # Apply (deploy compose + recreate hbbs/hbbr)
# ansible-playbook playbooks/rustdesk.yml --limit rustdesk \
# -e '{"rustdesk_confirm": true}'
- name: Deploy and reconcile RustDesk server
hosts: rustdesk
become: true
gather_facts: false
serial: 1
roles:
- role: rustdesk
tags: [rustdesk, mutating]
+456
View File
@@ -0,0 +1,456 @@
---
# Narrow reconciliation for the audited us4 public zone. This playbook never
# reloads or restarts firewalld and deliberately does not manage Docker rules.
- name: Safely remove audited stale firewalld allowances from us4
hosts: wireguard
become: true
gather_facts: false
serial: 1
any_errors_fatal: true
vars:
us4_firewalld_confirm: false
us4_console_confirm: false
us4_firewalld_zone: public
us4_firewalld_keep_services:
- dhcpv6-client
- http
- https
- smtp
- ssh
us4_firewalld_stale_services:
- imap
- imaps
- smtp-submission
- smtps
us4_firewalld_stale_ports:
- 24/tcp
- 6443/tcp
- 8443/tcp
us4_firewalld_expected_containers:
- nghttpx-proxy
- semaphoreui-postgres-1
- semaphoreui-semaphore-1
- squid-backend
- traefik
- trlm-server-trilium-1
- wireguard
us4_firewalld_backup_root: /var/backups/us4-firewall
tasks:
- name: Require the audited host and explicit apply confirmations
ansible.builtin.assert:
that:
- inventory_hostname == 'us4'
- ansible_host == 'us4.wsvc.info'
- ansible_host_ipv4 == '185.201.226.122'
- ansible_check_mode or (us4_firewalld_confirm | bool)
- ansible_check_mode or (us4_console_confirm | bool)
fail_msg: >-
Apply is allowed only for audited host us4 after the provider console
has been tested. Set both us4_firewalld_confirm=true and
us4_console_confirm=true. Check mode does not require confirmation.
- name: Verify the remote host identity
ansible.builtin.command:
argv: [hostname, -f]
check_mode: false
changed_when: false
register: us4_firewalld_hostname
- name: Reject an unexpected remote host
ansible.builtin.assert:
that:
- us4_firewalld_hostname.stdout == 'us4.wsvc.info'
- name: Verify required services are active
ansible.builtin.command:
argv: [systemctl, is-active, --quiet, "{{ item }}"]
check_mode: false
changed_when: false
loop:
- atd
- firewalld
- name: Verify firewalld Python bindings used by ansible.posix
ansible.builtin.command:
argv: [python3, -c, "import dbus, firewall, firewall.client"]
check_mode: false
changed_when: false
- name: Verify the default firewalld zone
ansible.builtin.command:
argv: [firewall-cmd, --get-default-zone]
check_mode: false
changed_when: false
register: us4_firewalld_default_zone
- name: Read runtime public-zone services
ansible.builtin.command:
argv: [firewall-cmd, --zone=public, --list-services]
check_mode: false
changed_when: false
register: us4_firewalld_runtime_services
- name: Read permanent public-zone services
ansible.builtin.command:
argv: [firewall-cmd, --permanent, --zone=public, --list-services]
check_mode: false
changed_when: false
register: us4_firewalld_permanent_services
- name: Read runtime public-zone ports
ansible.builtin.command:
argv: [firewall-cmd, --zone=public, --list-ports]
check_mode: false
changed_when: false
register: us4_firewalld_runtime_ports
- name: Read permanent public-zone ports
ansible.builtin.command:
argv: [firewall-cmd, --permanent, --zone=public, --list-ports]
check_mode: false
changed_when: false
register: us4_firewalld_permanent_ports
- name: Normalize the audited public-zone state
ansible.builtin.set_fact:
us4_firewalld_pre_services: "{{ us4_firewalld_runtime_services.stdout.split() | sort }}"
us4_firewalld_pre_permanent_services: "{{ us4_firewalld_permanent_services.stdout.split() | sort }}"
us4_firewalld_pre_ports: "{{ us4_firewalld_runtime_ports.stdout.split() | sort }}"
us4_firewalld_pre_permanent_ports: "{{ us4_firewalld_permanent_ports.stdout.split() | sort }}"
- name: Fail closed on public-zone drift or unknown allowances
ansible.builtin.assert:
that:
- us4_firewalld_default_zone.stdout == us4_firewalld_zone
- us4_firewalld_pre_services == us4_firewalld_pre_permanent_services
- us4_firewalld_pre_ports == us4_firewalld_pre_permanent_ports
- us4_firewalld_keep_services | difference(us4_firewalld_pre_services) | length == 0
- us4_firewalld_pre_services | difference(us4_firewalld_keep_services + us4_firewalld_stale_services) | length == 0
- us4_firewalld_pre_ports | difference(us4_firewalld_stale_ports) | length == 0
fail_msg: >-
The public zone differs from the audited baseline. Stop and review it;
this playbook will not infer whether an unknown allowance is required.
- name: Select only audited stale entries that currently exist
ansible.builtin.set_fact:
us4_firewalld_cleanup_services: >-
{{ us4_firewalld_stale_services | intersect(us4_firewalld_pre_services) | sort }}
us4_firewalld_cleanup_ports: >-
{{ us4_firewalld_stale_ports | intersect(us4_firewalld_pre_ports) | sort }}
- name: Report the proposed reconciliation
ansible.builtin.debug:
msg:
keep_services: "{{ us4_firewalld_keep_services }}"
remove_services: "{{ us4_firewalld_cleanup_services }}"
remove_ports: "{{ us4_firewalld_cleanup_ports }}"
reload_or_restart: false
- name: Create rollback material when cleanup is required
when:
- not ansible_check_mode
- us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0
block:
- name: Create the protected firewall backup root
ansible.builtin.file:
path: "{{ us4_firewalld_backup_root }}"
state: directory
owner: root
group: root
mode: "0700"
- name: Create a backup timestamp
ansible.builtin.command:
argv: [date, +%Y%m%dT%H%M%S%z]
changed_when: false
register: us4_firewalld_backup_timestamp
- name: Set the protected backup directory
ansible.builtin.set_fact:
us4_firewalld_backup_dir: >-
{{ us4_firewalld_backup_root }}/{{ us4_firewalld_backup_timestamp.stdout }}
us4_firewalld_rollback_command: >-
{{ us4_firewalld_backup_root }}/{{ us4_firewalld_backup_timestamp.stdout }}/rollback-phase1.sh
- name: Create the protected backup directory
ansible.builtin.file:
path: "{{ us4_firewalld_backup_dir }}"
state: directory
owner: root
group: root
mode: "0700"
- name: Back up the complete firewalld configuration
ansible.builtin.command:
argv:
- tar
- --create
- --gzip
- "--file={{ us4_firewalld_backup_dir }}/firewalld.tgz"
- --directory=/etc
- firewalld
changed_when: true
- name: Capture the pre-change runtime ruleset
ansible.builtin.shell:
cmd: >-
umask 077 && nft list ruleset >
{{ us4_firewalld_backup_dir | quote }}/nft-ruleset.txt
executable: /bin/bash
changed_when: true
- name: Install the exact pre-change rollback script
ansible.builtin.copy:
dest: "{{ us4_firewalld_rollback_command }}"
owner: root
group: root
mode: "0700"
content: |
#!/bin/sh
set -eu
exec >>/var/log/us4-firewalld-phase1-rollback.log 2>&1
printf '%s rollback start\n' "$(date -Is)"
add_service() {
service=$1
/usr/bin/firewall-cmd --permanent --zone=public \
--query-service="$service" >/dev/null 2>&1 ||
/usr/bin/firewall-cmd --permanent --zone=public \
--add-service="$service"
/usr/bin/firewall-cmd --zone=public \
--query-service="$service" >/dev/null 2>&1 ||
/usr/bin/firewall-cmd --zone=public --add-service="$service"
}
add_port() {
port=$1
/usr/bin/firewall-cmd --permanent --zone=public \
--query-port="$port" >/dev/null 2>&1 ||
/usr/bin/firewall-cmd --permanent --zone=public \
--add-port="$port"
/usr/bin/firewall-cmd --zone=public \
--query-port="$port" >/dev/null 2>&1 ||
/usr/bin/firewall-cmd --zone=public --add-port="$port"
}
{% for service in us4_firewalld_cleanup_services %}
add_service {{ service }}
{% endfor %}
{% for port in us4_firewalld_cleanup_ports %}
add_port {{ port }}
{% endfor %}
/usr/bin/firewall-cmd --check-config
printf '%s rollback complete\n' "$(date -Is)"
- name: Schedule the 15-minute automatic rollback
ansible.builtin.shell:
cmd: |
set -euo pipefail
output=$(printf '%s\n' {{ us4_firewalld_rollback_command | quote }} | at now + 15 minutes 2>&1)
job_id=$(printf '%s\n' "$output" | sed -n 's/^job \([0-9][0-9]*\).*/\1/p')
test -n "$job_id"
printf '%s\n' "$job_id"
executable: /bin/bash
changed_when: true
register: us4_firewalld_rollback_job
- name: Record the automatic rollback job
ansible.builtin.set_fact:
us4_firewalld_rollback_job_id: "{{ us4_firewalld_rollback_job.stdout }}"
us4_firewalld_rollback_cancelled: false
- name: Persist the rollback job ID beside the backup
ansible.builtin.copy:
dest: "{{ us4_firewalld_backup_dir }}/phase1-at-job-id"
owner: root
group: root
mode: "0600"
content: "{{ us4_firewalld_rollback_job_id }}\n"
- name: Reconcile and verify the audited public zone
block:
- name: Remove audited stale firewalld services
ansible.posix.firewalld:
zone: "{{ us4_firewalld_zone }}"
service: "{{ item }}"
state: disabled
permanent: true
immediate: true
loop: "{{ us4_firewalld_stale_services }}"
- name: Remove audited stale firewalld ports
ansible.posix.firewalld:
zone: "{{ us4_firewalld_zone }}"
port: "{{ item }}"
state: disabled
permanent: true
immediate: true
loop: "{{ us4_firewalld_stale_ports }}"
- name: Verify the permanent firewalld configuration
ansible.builtin.command:
argv: [firewall-cmd, --check-config]
when: not ansible_check_mode
changed_when: false
- name: Read reconciled runtime services
ansible.builtin.command:
argv: [firewall-cmd, --zone=public, --list-services]
changed_when: false
when: not ansible_check_mode
register: us4_firewalld_after_runtime_services
- name: Read reconciled permanent services
ansible.builtin.command:
argv: [firewall-cmd, --permanent, --zone=public, --list-services]
changed_when: false
when: not ansible_check_mode
register: us4_firewalld_after_permanent_services
- name: Read reconciled runtime ports
ansible.builtin.command:
argv: [firewall-cmd, --zone=public, --list-ports]
changed_when: false
when: not ansible_check_mode
register: us4_firewalld_after_runtime_ports
- name: Read reconciled permanent ports
ansible.builtin.command:
argv: [firewall-cmd, --permanent, --zone=public, --list-ports]
changed_when: false
when: not ansible_check_mode
register: us4_firewalld_after_permanent_ports
- name: Require the exact audited post-change public zone
ansible.builtin.assert:
that:
- us4_firewalld_after_runtime_services.stdout.split() | sort == us4_firewalld_keep_services | sort
- us4_firewalld_after_permanent_services.stdout.split() | sort == us4_firewalld_keep_services | sort
- us4_firewalld_after_runtime_ports.stdout.split() | length == 0
- us4_firewalld_after_permanent_ports.stdout.split() | length == 0
when: not ansible_check_mode
- name: Verify a fresh independent SSH and sudo path
ansible.builtin.command:
argv:
- ssh
- -4
- -o
- BatchMode=yes
- -o
- ConnectTimeout=10
- -o
- ControlMaster=no
- -o
- ControlPath=none
- windy@us4.wsvc.info
- sudo -n true
delegate_to: localhost
become: false
changed_when: false
when: not ansible_check_mode
vars:
ansible_become: false
- name: Verify public HTTPS routes
ansible.builtin.uri:
url: "{{ item.url }}"
follow_redirects: all
status_code: "{{ item.status }}"
validate_certs: true
use_proxy: false
loop:
- {url: https://update.wsvc.info/, status: 200}
- {url: https://us4-gate.wsvc.info/, status: 401}
- {url: https://trlm.wsvc.info/, status: 200}
delegate_to: localhost
become: false
when: not ansible_check_mode
vars:
ansible_become: false
- name: Verify the secondary MX TCP listener externally
ansible.builtin.wait_for:
host: "{{ ansible_host_ipv4 }}"
port: 25
state: started
connect_timeout: 5
timeout: 10
delegate_to: localhost
become: false
when: not ansible_check_mode
vars:
ansible_become: false
- name: Verify all expected containers are running
ansible.builtin.command:
argv: [docker, ps, --format, "{{ '{{.Names}}' }}"]
changed_when: false
when: not ansible_check_mode
register: us4_firewalld_running_containers
- name: Reject missing application containers
ansible.builtin.assert:
that:
- us4_firewalld_expected_containers | difference(us4_firewalld_running_containers.stdout_lines) | length == 0
when: not ansible_check_mode
- name: Verify Fail2ban remains active
ansible.builtin.command:
argv: [fail2ban-client, status]
changed_when: false
when: not ansible_check_mode
register: us4_firewalld_fail2ban
- name: Require all audited Fail2ban jails
ansible.builtin.assert:
that:
- item in us4_firewalld_fail2ban.stdout
loop:
- postfix-postscreen
- postfix-sasl
- recidive
- sshd
when: not ansible_check_mode
- name: Verify the deployed WireGuard health check
ansible.builtin.command:
argv: [/usr/local/lib/vps-health/run]
changed_when: false
when: not ansible_check_mode
register: us4_firewalld_wireguard_health
- name: Cancel automatic rollback only after all checks pass
ansible.builtin.command:
argv: [at, -r, "{{ us4_firewalld_rollback_job_id }}"]
changed_when: true
when:
- not ansible_check_mode
- us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0
- name: Mark the automatic rollback as cancelled
ansible.builtin.set_fact:
us4_firewalld_rollback_cancelled: true
when:
- not ansible_check_mode
- us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0
rescue:
- name: Preserve the automatic rollback and stop
ansible.builtin.fail:
msg: >-
A reconciliation or verification task failed. No reload was
attempted. If cleanup was required, its automatic rollback remains
scheduled; do not remove it manually.
always:
- name: Report backup and rollback disposition
ansible.builtin.debug:
msg:
backup: >-
{{ us4_firewalld_backup_dir |
default('not-created-in-check-mode' if ansible_check_mode else 'not-required') }}
automatic_rollback: >-
{{ 'not-created-in-check-mode' if ansible_check_mode else
('cancelled-after-success' if (us4_firewalld_rollback_cancelled | default(false)) else
'scheduled-or-executed') if
(us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0)
else 'not-required' }}
+5
View File
@@ -0,0 +1,5 @@
---
collections:
# us4-firewalld.yml was source-reviewed and exercised with this version.
- name: ansible.posix
version: 2.2.2
@@ -0,0 +1,8 @@
---
# Allowlist of compose/ projects this playbook may deploy. A host may only
# reference a project listed here (see tasks: "Require a repo compose project").
compose_repo_projects:
- vaultwarden
- pdns
- adguardhome
- unifi
@@ -0,0 +1,84 @@
---
# Deploy the repo-owned, sanitized Compose declaration to the host.
#
# Safety model:
# - Only hosts with an inventory `compose_repo_project` (allowlisted) are valid.
# - The repo file is staged to `<file>.dsh-new` and validated with
# `docker compose config --quiet` against the server-local .env BEFORE it
# replaces anything. A failed validation never touches the live file.
# - The current file is kept as `*.bak-<timestamp>` before promotion.
# - Apply mode requires `compose_deploy_confirm=true`; `--check` gives a
# read-only diff + validation without writes.
# - The playbook never writes, reads, or transfers the server .env.
- name: Require an allowlisted repo compose project for this host
ansible.builtin.assert:
that:
- compose_repo_project is defined
- compose_repo_project in compose_repo_projects
fail_msg: >-
No allowlisted compose_repo_project for {{ inventory_hostname }}.
Supported: {{ compose_repo_projects | join(', ') }}.
- name: Require explicit confirmation for apply mode
ansible.builtin.assert:
that:
- ansible_check_mode or (compose_deploy_confirm | bool)
fail_msg: >-
This playbook replaces the server compose file and may recreate
containers. Run with --check for a read-only diff, or supply
compose_deploy_confirm=true to apply.
- name: Stage the repo compose file next to the live one
ansible.builtin.copy:
src: "{{ playbook_dir }}/../../compose/{{ compose_repo_project }}/compose.yml"
dest: "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
mode: "0644"
diff: true
register: compose_stage
- name: Validate staged compose against the server .env (read-only)
ansible.builtin.command:
argv:
- docker
- compose
- -f
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
- --project-directory
- "{{ compose_project_dir }}"
- config
- --quiet
register: compose_validate
changed_when: false
failed_when: compose_validate.rc != 0
- name: Show staged-vs-live difference
ansible.builtin.debug:
msg: "{{ compose_stage.diff | default('(no change)') }}"
when: ansible_check_mode
- name: Back up the current compose file (apply mode)
ansible.builtin.shell:
cmd: >-
cp -a '{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}'
'{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.bak-$(date +%Y%m%d-%H%M%S)'
when: not ansible_check_mode
- name: Promote the validated compose file (apply mode)
ansible.builtin.command:
argv:
- mv
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}"
when: not ansible_check_mode
- name: Apply the compose declaration (apply mode)
ansible.builtin.command:
argv:
- docker
- compose
- --project-directory
- "{{ compose_project_dir }}"
- up
- -d
when: not ansible_check_mode
+5 -1
View File
@@ -7,9 +7,13 @@ healthcheck_timer_on_calendar: '*-*-* 06:15:00'
healthcheck_timer_randomized_delay_sec: 15m
healthcheck_backup_max_age_hours: 30
healthcheck_tls_warn_days: 21
healthcheck_profiles:
# Map of profile name -> installed script filename. A host selects which
# profiles it runs via the `healthcheck_profiles` list (inventory).
healthcheck_profile_scripts:
mailcow: mailcow.sh
vaultwarden: vaultwarden.sh
pdns: pdns.sh
wireguard: wireguard.sh
adguardhome: adguardhome.sh
rustdesk: rustdesk.sh
hk2aux: hk2aux.sh
+8 -6
View File
@@ -1,9 +1,10 @@
---
- name: Validate known health-check profile
- name: Validate known health-check profiles
ansible.builtin.assert:
that:
- healthcheck_profile in healthcheck_profiles
fail_msg: "Unsupported healthcheck_profile: {{ healthcheck_profile }}"
- item in healthcheck_profile_scripts
fail_msg: "Unsupported healthcheck_profile: {{ item }}"
loop: "{{ healthcheck_profiles }}"
- name: Install health-check directories
ansible.builtin.file:
@@ -28,13 +29,14 @@
group: root
mode: "0755"
- name: Install service health-check script
- name: Install service health-check scripts
ansible.builtin.template:
src: "{{ healthcheck_profiles[healthcheck_profile] }}.j2"
dest: "{{ healthcheck_install_root }}/{{ healthcheck_profiles[healthcheck_profile] }}"
src: "{{ healthcheck_profile_scripts[item] }}.j2"
dest: "{{ healthcheck_install_root }}/{{ healthcheck_profile_scripts[item] }}"
owner: root
group: root
mode: "0755"
loop: "{{ healthcheck_profiles }}"
- name: Install health-check dispatcher
ansible.builtin.template:
@@ -7,7 +7,7 @@ set -uo pipefail
RESULT_DIR='{{ healthcheck_state_dir }}'
LOG_DIR='{{ healthcheck_log_dir }}'
HOST_NAME="$(hostname -f 2>/dev/null || hostname)"
CHECK_NAME='{{ healthcheck_profile }}'
CHECK_NAME="$(basename "$0" .sh)"
STATUS=ok
EXIT_CODE=0
DETAILS=()
@@ -31,9 +31,27 @@ compose_ps() {
}
check_compose() {
local output
output="$(compose_ps)" || { record critical 'compose_ps_failed'; return; }
if grep -qiE 'Exited|Restarting|[[:space:]]Dead[[:space:]]' <<<"$output"; then
local output services bad
# Only flag containers of *active* services (config --services excludes
# debug/profile-gated services such as vaultwarden's pgweb, which is
# intentionally stopped unless started with --profile debug).
services="$(docker compose --project-directory '{{ compose_project_dir }}' config --services 2>/dev/null)" || { record critical 'compose_ps_failed'; return; }
output="$(docker compose --project-directory '{{ compose_project_dir }}' ps --all --format json 2>&1)" || { record critical 'compose_ps_failed'; return; }
bad="$(printf '%s\n' "$output" | python3 -c '
import json, sys
services = set(sys.argv[1].split())
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
c = json.loads(line)
except Exception:
continue
if c.get("Service") in services and c.get("State") in ("exited", "restarting", "dead"):
print(c.get("Service"))
' "$services")"
if [[ -n "$bad" ]]; then
record critical 'compose_unhealthy_container'
else
record ok 'compose_ok'
@@ -77,8 +95,11 @@ check_tls_days() {
}
emit_result() {
local tmp detail_json
tmp="$(mktemp "${RESULT_DIR}/latest.json.XXXXXX")"
# Per-check JSON at latest-<check>.json. The dispatcher merges these into
# latest.json so multiple profiles on one host do not overwrite each other.
local tmp path detail_json
path="${RESULT_DIR}/latest-${CHECK_NAME}.json"
tmp="$(mktemp "${RESULT_DIR}/.latest-${CHECK_NAME}.XXXXXX")"
detail_json="$(printf '%s\n' "${DETAILS[@]:-unknown:no_details}" | python3 -c 'import json,sys; print(json.dumps([line.rstrip() for line in sys.stdin if line.strip()]))')"
python3 - "$tmp" "$HOST_NAME" "$CHECK_NAME" "$STATUS" "$EXIT_CODE" "$detail_json" <<'PY'
import json, sys
@@ -90,6 +111,75 @@ with open(path, 'w', encoding='utf-8') as f:
f.write('\n')
PY
chmod 0640 "$tmp"
mv "$tmp" "${RESULT_DIR}/latest.json"
mv "$tmp" "$path"
cat "$path"
}
aggregate_result() {
# Merge the just-run per-check files into latest.json. With a single complete
# check this is a verbatim copy, preserving the historical one-object shape.
# With several checks it emits one object whose status is the worst of all
# checks; each check's own status/details are retained under `checks`. An
# expected check with no fresh result file (profile crashed before writing)
# is aggregated as `unknown`, so latest.json can never go stale while the
# dispatcher reports a failure.
case "$#" in
0) return 0 ;;
1) if [[ -f "${RESULT_DIR}/latest-$1.json" ]]; then
cp -f "${RESULT_DIR}/latest-$1.json" "${RESULT_DIR}/latest.json"
else
python3 - "$RESULT_DIR" "$HOST_NAME" "$@" <<'PY'
import json, os, sys
rdir, host = sys.argv[1], sys.argv[2]
checks = sys.argv[3:]
levels = {'ok': 0, 'warning': 1, 'unknown': 2, 'critical': 3}
worst, worst_code = 'ok', 0
items = []
for c in checks:
p = os.path.join(rdir, 'latest-%s.json' % c)
if os.path.exists(p):
d = json.load(open(p))
st, code = d['status'], d['exit_code']
items.append({'check': d['check'], 'status': st,
'exit_code': code, 'details': d['details']})
else:
st, code = 'unknown', 3
items.append({'check': c, 'status': st, 'exit_code': code,
'details': ['unknown:check_did_not_complete']})
if levels[st] > levels[worst]:
worst, worst_code = st, code
out = {'schema': 1, 'host': host, 'check': 'aggregate', 'status': worst,
'exit_code': worst_code, 'checks': items}
open(os.path.join(rdir, 'latest.json'), 'w').write(
json.dumps(out, sort_keys=True, separators=(',', ':')) + '\n')
PY
fi ;;
*) python3 - "$RESULT_DIR" "$HOST_NAME" "$@" <<'PY'
import json, os, sys
rdir, host = sys.argv[1], sys.argv[2]
checks = sys.argv[3:]
levels = {'ok': 0, 'warning': 1, 'unknown': 2, 'critical': 3}
worst, worst_code = 'ok', 0
items = []
for c in checks:
p = os.path.join(rdir, 'latest-%s.json' % c)
if os.path.exists(p):
d = json.load(open(p))
st, code = d['status'], d['exit_code']
items.append({'check': d['check'], 'status': st,
'exit_code': code, 'details': d['details']})
else:
st, code = 'unknown', 3
items.append({'check': c, 'status': st, 'exit_code': code,
'details': ['unknown:check_did_not_complete']})
if levels[st] > levels[worst]:
worst, worst_code = st, code
out = {'schema': 1, 'host': host, 'check': 'aggregate', 'status': worst,
'exit_code': worst_code, 'checks': items}
open(os.path.join(rdir, 'latest.json'), 'w').write(
json.dumps(out, sort_keys=True, separators=(',', ':')) + '\n')
PY
esac
chmod 0640 "${RESULT_DIR}/latest.json"
cat "${RESULT_DIR}/latest.json"
}
@@ -1,4 +1,24 @@
#!/usr/bin/env bash
set -o pipefail
'{{ healthcheck_install_root }}/{{ healthcheck_profiles[healthcheck_profile] }}' 2>&1 | tee -a '{{ healthcheck_log_dir }}/healthcheck.log'
exit "${PIPESTATUS[0]}"
source '{{ healthcheck_install_root }}/health-common.sh'
# Run every enabled health-check profile, exit with the worst (max) code, and
# merge the per-check results into /var/lib/vps-health/latest.json.
rc=0
# Drop per-check results from any prior run so a profile that crashes before
# reporting cannot leak a stale healthy result into the aggregate.
{% for profile in healthcheck_profiles %}
rm -f '{{ healthcheck_state_dir }}/latest-{{ healthcheck_profile_scripts[profile] | replace('.sh', '') }}.json'
{% endfor %}
{% for profile in healthcheck_profiles %}
'{{ healthcheck_install_root }}/{{ healthcheck_profile_scripts[profile] }}' 2>&1 | tee -a '{{ healthcheck_log_dir }}/healthcheck.log'
this_rc="${PIPESTATUS[0]}"
[ "$this_rc" -gt "$rc" ] && rc="$this_rc"
{% endfor %}
# Collect profile check names line-by-line (robust against Jinja trim_blocks
# whitespace control, which would otherwise merge this into one line).
aggregate_args=""
{% for profile in healthcheck_profiles %}
aggregate_args="$aggregate_args {{ healthcheck_profile_scripts[profile] | replace('.sh', '') }}"
{% endfor %}
aggregate_result $aggregate_args
exit "$rc"
@@ -0,0 +1,56 @@
#!/usr/bin/env bash
set -uo pipefail
source '{{ healthcheck_install_root }}/health-common.sh'
require_command docker
require_command ss
# Auxiliary services co-located on hk2.chans.xyz (separate compose projects
# under /opt, fronted by Traefik). Verified live 2026-08-12.
# traefik
if docker inspect traefik >/dev/null 2>&1; then
[[ "$(docker inspect traefik --format '{{ '{{' }}.State.Running{{ '}}' }}' 2>/dev/null)" == true ]] \
&& record ok 'traefik_running' || record critical 'traefik_not_running'
else
record critical 'traefik_container_missing'
fi
# adguardhome (hk2 variant: DoH 5443, DoT 853)
if docker inspect adguardhome >/dev/null 2>&1; then
[[ "$(docker inspect adguardhome --format '{{ '{{' }}.State.Running{{ '}}' }}' 2>/dev/null)" == true ]] \
&& record ok 'adguard_running' || record critical 'adguard_not_running'
else
record critical 'adguard_container_missing'
fi
# remark42
if docker inspect remark42 >/dev/null 2>&1; then
[[ "$(docker inspect remark42 --format '{{ '{{' }}.State.Running{{ '}}' }}' 2>/dev/null)" == true ]] \
&& record ok 'remark42_running' || record critical 'remark42_not_running'
else
record critical 'remark42_container_missing'
fi
# nginx-manager was removed 2026-08-12 (leftover config, never running).
# Warn if a container by that name ever reappears.
if docker inspect nginx-manager >/dev/null 2>&1; then
record warning 'nginx_manager_unexpectedly_running'
else
record ok 'nginx_manager_not_running'
fi
# Listening ports (Traefik 80/443/8080, AdGuard DoH 5443 / DoT 853).
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)80$' \
&& record ok 'traefik_http_80' || record critical 'traefik_http_80_missing'
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)443$' \
&& record ok 'traefik_https_443' || record critical 'traefik_https_443_missing'
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)8080$' \
&& record ok 'traefik_dashboard_8080' || record warning 'traefik_dashboard_8080_missing'
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)5443$' \
&& record ok 'adguard_doh_5443' || record critical 'adguard_doh_5443_missing'
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)853$' \
&& record ok 'adguard_dot_853' || record critical 'adguard_dot_853_missing'
emit_result
exit "$EXIT_CODE"
@@ -0,0 +1,25 @@
#!/usr/bin/env bash
set -uo pipefail
source '{{ healthcheck_install_root }}/health-common.sh'
require_command docker
require_command dig
# hbbs / hbbr must both be running (separate compose project at /opt/rustdesk).
output="$(docker compose --project-directory /opt/rustdesk ps --all 2>&1)"
if grep -qiE 'Exited|Restarting|[[:space:]]Dead[[:space:]]' <<<"$output"; then
record critical 'rustdesk_unhealthy_container'
else
record ok 'rustdesk_compose_ok'
fi
# hbbs must advertise the relay hostname that resolves to this host's public IP.
cmd="$(docker inspect hbbs --format '{{ '{{' }}json .Config.Cmd{{ '}}' }}' 2>/dev/null)" || record critical 'rustdesk_hbbs_missing'
grep -q 'hk2.chans.xyz:21117' <<<"$cmd" || record critical 'rustdesk_relay_misconfigured'
# The advertised relay hostname must resolve to this host's public IP.
resolved="$(dig +short hk2.chans.xyz A 2>/dev/null)"
grep -q '154.36.174.161' <<<"$resolved" || record critical 'rustdesk_relay_dns_missing'
emit_result
exit "$EXIT_CODE"
@@ -15,11 +15,12 @@ grep -Fq 'vw-db' <<<"$health" || record critical 'postgres_missing'
check_https 'https://auth.wsvc.info/' '^200$'
check_tls_days auth.wsvc.info 443
# Read effective config only inside the service and report booleans/fingerprints,
# never its SMTP password or other secret fields.
smtp_result="$(docker compose --project-directory '{{ compose_project_dir }}' exec -T vaultwarden python3 - <<'PY' 2>&1
# Read effective config from the mounted vw-data dir on the host and run the
# SMTP AUTH probe from the host (the vaultwarden image has no python3; the
# host does). Never print the SMTP password.
smtp_result="$(python3 - <<'PY' 2>&1
import json, pathlib, smtplib, ssl
cfg=json.loads(pathlib.Path('/data/config.json').read_text())
cfg=json.loads(pathlib.Path('{{ compose_project_dir }}/vw-data/config.json').read_text())
host=cfg.get('smtp_host'); port=int(cfg.get('smtp_port') or 0)
user=cfg.get('smtp_username')
smtp_secret=cfg.get('smtp_password')
+3
View File
@@ -1,5 +1,8 @@
---
restic_enabled: false
# Approved Restic source profile (a key of restic_sources) for the host. Set
# per-host in inventory; the role fails if it is not an approved source.
restic_backup_profile: ""
restic_binary: /usr/bin/restic
restic_config_path: /etc/vps-restic/repository.env
restic_state_dir: /var/lib/vps-restic
+2 -2
View File
@@ -10,8 +10,8 @@
- name: Validate supported Restic source profile
ansible.builtin.assert:
that:
- healthcheck_profile in restic_sources
fail_msg: "No approved Restic source profile for {{ healthcheck_profile }}."
- restic_backup_profile in restic_sources
fail_msg: "No approved Restic source profile for {{ restic_backup_profile }}."
- name: Verify Restic binary exists on target
ansible.builtin.stat:
+1 -1
View File
@@ -3,4 +3,4 @@ set -euo pipefail
# Repository and password credentials are host-local in {{ restic_config_path }}.
# shellcheck source=/dev/null
source '{{ restic_config_path }}'
exec '{{ restic_binary }}' backup --tag '{{ healthcheck_profile }}' --tag "$(hostname -s)" {% for source in restic_sources[healthcheck_profile] %}{{ source | quote }} {% endfor %}
exec '{{ restic_binary }}' backup --tag '{{ restic_backup_profile }}' --tag "$(hostname -s)" {% for source in restic_sources[restic_backup_profile] %}{{ source | quote }} {% endfor %}
@@ -2,4 +2,4 @@
set -euo pipefail
# shellcheck source=/dev/null
source '{{ restic_config_path }}'
exec '{{ restic_binary }}' forget --prune --keep-daily {{ restic_keep_daily }} --keep-weekly {{ restic_keep_weekly }} --keep-monthly {{ restic_keep_monthly }} --tag '{{ healthcheck_profile }}'
exec '{{ restic_binary }}' forget --prune --keep-daily {{ restic_keep_daily }} --keep-weekly {{ restic_keep_weekly }} --keep-monthly {{ restic_keep_monthly }} --tag '{{ restic_backup_profile }}'
@@ -1,5 +1,5 @@
[Unit]
Description=Restic backup for approved {{ healthcheck_profile }} sources
Description=Restic backup for approved {{ restic_backup_profile }} sources
After=network-online.target
Wants=network-online.target
+13
View File
@@ -0,0 +1,13 @@
---
# RustDesk server deployment (hbbs + hbbr) on hk2.
# Safe by default: without rustdesk_confirm=true the role only reports whether
# the declared compose file matches live state and refuses to recreate the stack.
rustdesk_confirm: false
# Compose project directory.
rustdesk_compose_dir: /opt/rustdesk
# Relay (hbbr) hostname:port advertised to every client via `hbbs -r`.
# MUST resolve to this host's public IP (154.36.174.161). The known-bad value
# 'hk2.wsvc.info' has no DNS record and must never be used.
rustdesk_relay: hk2.chans.xyz:21117
# Pinned server image (used for both hbbs and hbbr).
rustdesk_image: rustdesk/rustdesk-server:1.1.14
+85
View File
@@ -0,0 +1,85 @@
---
# Deploy/reconcile the self-hosted RustDesk server (hbbs + hbbr).
# Idempotent: deploys the declared compose file; only recreates the stack with
# explicit confirmation.
- name: Validate relay address is set and not the known-bad value
ansible.builtin.assert:
that:
- rustdesk_relay | length > 0
- "'hk2.wsvc.info' not in rustdesk_relay"
fail_msg: >-
rustdesk_relay must be a resolvable relay address. The known-bad
'hk2.wsvc.info' has no DNS record and must not be used.
- name: Ensure compose project directory exists
ansible.builtin.file:
path: "{{ rustdesk_compose_dir }}"
state: directory
owner: windy
group: root
mode: "0755"
- name: Deploy compose file
ansible.builtin.template:
src: compose.yml.j2
dest: "{{ rustdesk_compose_dir }}/compose.yml"
owner: windy
group: windy
mode: "0644"
register: rustdesk_compose_deployed
- name: Report no change needed
ansible.builtin.debug:
msg: "compose.yml already matches declared state; no change needed."
when: not rustdesk_compose_deployed.changed
- name: Refuse to recreate without explicit confirmation
ansible.builtin.fail:
msg: >-
compose.yml differs from declared state but rustdesk_confirm is not true.
Supply rustdesk_confirm=true to deploy the file and recreate the stack.
when:
- rustdesk_compose_deployed.changed
- not (rustdesk_confirm | bool)
- not ansible_check_mode
- name: Apply compose stack
ansible.builtin.command:
argv:
- docker
- compose
- --project-directory
- "{{ rustdesk_compose_dir }}"
- up
- -d
when:
- rustdesk_compose_deployed.changed
- rustdesk_confirm | bool
changed_when: true
register: rustdesk_apply
- name: Verify hbbs relay command
ansible.builtin.command:
argv:
- docker
- inspect
- hbbs
- --format
- '{{ "{{" }}json .Config.Cmd{{ "}}" }}'
register: rustdesk_hbbs_cmd
changed_when: false
when:
- rustdesk_compose_deployed.changed
- rustdesk_confirm | bool
- not ansible_check_mode
- name: Assert hbbs advertises the declared relay
ansible.builtin.assert:
that:
- "'{{ rustdesk_relay }}' in rustdesk_hbbs_cmd.stdout"
fail_msg: "hbbs is not advertising the declared relay {{ rustdesk_relay }}."
when:
- rustdesk_compose_deployed.changed
- rustdesk_confirm | bool
- not ansible_check_mode
@@ -0,0 +1,34 @@
networks:
rustdesk-net:
external: false
services:
hbbs:
container_name: hbbs
ports:
- 21115:21115
- 21116:21116
- 21116:21116/udp
- 21118:21118
image: {{ rustdesk_image }}
command: "hbbs -r {{ rustdesk_relay }}"
volumes:
- ./hbbs:/root
networks:
- rustdesk-net
depends_on:
- hbbr
restart: unless-stopped
hbbr:
container_name: hbbr
ports:
- 21117:21117
- 21119:21119
image: {{ rustdesk_image }}
command: hbbr
volumes:
- ./hbbr:/root
networks:
- rustdesk-net
restart: unless-stopped
+51
View File
@@ -0,0 +1,51 @@
# compose/ — repo-owned Compose declarations
Non-secret Compose sources for the Docker hosts. Secrets are **never** in these
files: every secret is a `${VAR}` reference resolved from the **server-local
`.env`** (docker compose reads `.env` from the project directory automatically).
## Source-of-truth matrix
| Project | Host | Compose source | Mechanism |
|---------|------|----------------|-----------|
| `vaultwarden` | us2 (`/opt/vaultwarden`) | `compose/vaultwarden/compose.yml` | static file + `compose-deploy.yml` |
| `pdns` | hk2 (`/opt/pdns`) | `compose/pdns/compose.yml` | static file + `compose-deploy.yml` |
| `pgdb` | pgdb (`/opt/database`, 无 ansible) | `compose/pgdb/compose.yml` | static file(手动部署:scp → `docker compose config -q``up -d`;服务器文件名 `docker-compose.yml` |
| `soft-serve` | us2 (`/opt/soft-serve`, 已退役停用) | `compose/soft-serve/compose.yml` (+ `Dockerfile.backup`, `scripts/`) | static file(参考镜像; 2026-09-18 被 gitea 替换 VPS-94, 数据保留作回滚) |
| `gitea` | us2 (`/opt/gitea`) | `compose/gitea/compose.yml` (+ `Dockerfile.backup`, `scripts/`) | static file(参考镜像, 未接入 compose-deploy; 服务器文件为准; 2026-09-18 替换 soft-serve, VPS-94 |
| `adguardhome` | dns.windy.lan (`/opt/adguardhome`) | — (待从 LAN 提取) | static file (pending) |
| `unifi` | ubnt (`/home/windy/unifi-9`) | — (待从 LAN 提取) | static file (pending) |
| `wireguard` | us4 (`/opt/wireguard`) | `ansible/templates/wireguard-compose.yml.j2` | role-rendered (inventory vars) |
| `rustdesk` | hk2 (`/opt/rustdesk`) | `ansible/roles/rustdesk/templates/compose.yml.j2` | role-rendered (inventory vars) |
| `mailcow` | mx2 (`/opt/mail`) | — (mailcow update generator owns it) | excluded by design |
Mechanism rule: **static** `compose/<project>/compose.yml` for declarations that
do not vary per host; **role-rendered j2** for declarations driven by inventory
vars (image pins, relay host). One mechanism per project; do not duplicate a
project in both.
## Deploying a static project
```bash
cd ansible
# Read-only diff + validation against the server .env (no writes)
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden --check --diff
# Apply: stage repo file → validate `docker compose config -q` → backup current
# file → promote → `docker compose up -d` (gated)
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden \
-e '{"compose_deploy_confirm": true}'
```
See [`../runbooks/ansible-operations.md`](../runbooks/ansible-operations.md).
## Adding a project
1. Sanitize the live compose so every secret is `${VAR}` from `.env`
(prefer `${VAR:?missing VAR}` for required keys).
2. Commit `compose/<project>/compose.yml` + `.env.example` (key names only).
3. Add `compose_repo_project` (+ `compose_remote_file` if not `compose.yml`) to
the host in `ansible/inventory/hosts.yml`, and allowlist the project in
`ansible/roles/compose_deploy/defaults/main.yml`.
4. Verify with `--check --diff` (zero diff) then a gated apply.
+4
View File
@@ -0,0 +1,4 @@
# compose/gitea — 秘密一律走服务器本地 .env, 不入库
# 迁移期一次性: Gitea 管理员生成的 token (mirror-migrate.sh 读取, 用后撤销)
GITEA_MIGRATE_USER=
GITEA_MIGRATE_TOKEN=
+3
View File
@@ -0,0 +1,3 @@
FROM alpine:3.20
RUN apk add --no-cache sqlite rsync tzdata
WORKDIR /scripts
+59
View File
@@ -0,0 +1,59 @@
# Gitea on us2 — reference compose (Plane VPS-94, 迁移完成 2026-09-18)
# 参考镜像, 服务器 /opt/gitea 文件为准 (同 soft-serve 约定, 未接入 compose-deploy)
# rootless 镜像: uid 1000 原生非 root; 数据 /var/lib/gitea (宿主 ./data), 配置 /etc/gitea (宿主 ./config)
# SSH: 容器内监听 2322 (非特权, SSH_LISTEN_PORT), 对外 repo.windy.me:2222 经 Traefik TCP entrypoint `ssh`
services:
gitea:
image: gitea/gitea@sha256:1c17ecaead42eb3b5391553d8708103a4beb0e86edf5b9ebc1eb269c318845f2 # 1.27.3-rootless
container_name: gitea
restart: unless-stopped
user: "1000:1000"
environment:
TZ: Asia/Shanghai
volumes:
- ./data:/var/lib/gitea
- ./config:/etc/gitea
- ./secrets:/secrets:ro # 复用的 soft-serve host key (SSH_SERVER_HOST_KEYS)
networks:
- traefik
labels:
- traefik.enable=true
# Web UI: repo.windy.me (2026-09-18 操作者决定复用现有域名, 免 DNS 变更)
- traefik.http.routers.gitea-web.rule=Host(`repo.windy.me`)
- traefik.http.routers.gitea-web.entrypoints=websecure
- traefik.http.routers.gitea-web.tls.certresolver=letsencrypt
- traefik.http.services.gitea-web.loadbalancer.server.port=3000
# SSH: 接管 :2222 (entrypoint 已存在, router 动态生效, 无需重启 Traefik)
- traefik.tcp.routers.gitea-ssh.entrypoints=ssh
- traefik.tcp.routers.gitea-ssh.rule=HostSNI(`*`)
- traefik.tcp.routers.gitea-ssh.tls=false
- traefik.tcp.services.gitea-ssh.loadbalancer.server.port=2322
gitea-backup:
build:
context: .
dockerfile: Dockerfile.backup
container_name: gitea-backup
restart: unless-stopped
volumes:
- ./data:/data:ro
- ./config:/config:ro
- ./backups:/backup
- ./scripts:/scripts
environment:
TZ: Asia/Shanghai
BACKUP_UID: 1000
BACKUP_GID: 1000
entrypoint: >
/bin/sh -ec "
umask 077 &&
touch /backup/backup.log &&
crontab /scripts/crontab.txt &&
echo '[INFO] gitea backup cron installed' &&
crond -f -l 8
"
networks:
traefik:
external: true
name: vw-net
+20
View File
@@ -0,0 +1,20 @@
#!/bin/sh
set -eu
umask 077
D() { date "+%Y-%m-%d %H:%M:%S"; }
TS=$(date +%Y%m%d_%H%M%S)
OUT="/backup/gitea_${TS}"
mkdir -p "$OUT"
echo "[$(D)] Starting gitea backup -> $OUT"
# rootless 布局: app.ini=/etc/gitea(宿主 ./config), db+repos=/var/lib/gitea/data(宿主 ./data/data)
# app.ini 含 SECRET_KEY/INTERNAL_TOKEN — 恢复 2FA/session/mirror 凭据必需
tar czf "$OUT/app.ini.tar.gz" -C /config app.ini
sqlite3 /data/data/gitea.db ".backup '$OUT/gitea.db'"
rsync -a /data/data/git/repositories/ "$OUT/repos/"
tar czf "$OUT/repos.tar.gz" -C "$OUT" repos
rm -rf "$OUT/repos"
chmod 600 "$OUT"/*
if [ -n "${BACKUP_UID:-}" ] && [ -n "${BACKUP_GID:-}" ]; then
chown -R "$BACKUP_UID:$BACKUP_GID" "$OUT" /backup/backup.log
fi
echo "[$(D)] Backup OK: $(du -sh "$OUT" | cut -f1)"
+4
View File
@@ -0,0 +1,4 @@
# Run gitea backup daily at 02:00
0 2 * * * /bin/sh /scripts/backup.sh >> /backup/backup.log 2>&1
# Prune backups older than 14 days daily at 03:00
0 3 * * * /bin/sh /scripts/prune.sh >> /backup/backup.log 2>&1
+41
View File
@@ -0,0 +1,41 @@
#!/bin/sh
# 一次性迁移辅助 (Plane VPS-94 Phase 2): 在 gitea 容器内执行。
# 已于 2026-09-18 执行完成 (16 仓), 留档备查; 复用时按 VPS-94 流程重生成一次性 token。
# 用法:
# GITEA_MIGRATE_USER=<user> GITEA_MIGRATE_TOKEN=<token> \
# docker exec -e GITEA_MIGRATE_USER -e GITEA_MIGRATE_TOKEN gitea \
# /scripts/mirror-migrate.sh [public_repo ...]
# 每仓: API 建仓 (默认 private, 参数中列出的为 public) -> push --mirror。
# default_branch 按源仓 symbolic-ref HEAD 设置, 避免非 main 源仓在 Gitea 显示为空。
# 结束后按 VPS-94 Phase 3 逐仓核对 git ls-remote ref 全集。
set -eu
MUSER="${GITEA_MIGRATE_USER:?need GITEA_MIGRATE_USER}"
TOKEN="${GITEA_MIGRATE_TOKEN:?need GITEA_MIGRATE_TOKEN}"
SRC="/migration-src"
API="http://localhost:3000/api/v1"
PUBLIC_REPOS=" $* "
migrate_one() {
dir="$1"
git -C "$dir" rev-parse --git-dir >/dev/null 2>&1 || { echo "[SKIP] $dir (not a git repo)"; return 0; }
name=$(basename "$dir"); name=${name%.git}
def_branch=$(git -C "$dir" symbolic-ref --short HEAD)
case "$PUBLIC_REPOS" in *" $name "*) private=false ;; *) private=true ;; esac
echo "[MIGRATE] $name (default=$def_branch private=$private)"
code=$(curl -s -o /dev/null -w '%{http_code}' -X POST "$API/user/repos" \
-H "Authorization: token $TOKEN" -H "Content-Type: application/json" \
-d "{\"name\":\"$name\",\"private\":$private,\"default_branch\":\"$def_branch\",\"auto_init\":false}")
case "$code" in
201) : ;;
409) echo " [WARN] $name 已存在, 直接补推" ;;
*) echo " [FAIL] create HTTP $code"; return 1 ;;
esac
git -C "$dir" push --mirror "http://$MUSER:$TOKEN@localhost:3000/$MUSER/$name.git"
echo " [OK] $name pushed"
}
for dir in "$SRC"/*.git "$SRC"/cdia; do
[ -d "$dir" ] || continue
migrate_one "$dir"
done
echo "[DONE] 全部处理完毕; 迁移后记得撤销一次性 token"
+5
View File
@@ -0,0 +1,5 @@
#!/bin/sh
set -eu
D() { date "+%Y-%m-%d %H:%M:%S"; }
ls -dt /backup/gitea_* 2>/dev/null | tail -n +15 | xargs -r rm -rf
echo "[$(D)] Pruned. Kept $(ls -d /backup/gitea_* 2>/dev/null | wc -l) backups (max 14)"
+39
View File
@@ -0,0 +1,39 @@
# .env.example — PowerDNS stack (hk2.chans.xyz, /opt/pdns)
#
# Non-secret key reference ONLY. Real values live in the server-local .env
# (never commit them). Compose requires the `:?`-marked keys to be present.
# Runtime
TZ=Asia/Shanghai
# Postgres superuser (db + backup + pgweb)
PGUSER=
PGPASSWORD=
DB_HOST=db
DB_PORT=5432
# Application database (auth / poweradmin / backup)
DB_NAME=pdns
DB_USER=pdns
DB_PASS=
ADMIN_DB=pdnsadmin
# Backups
CRON_SCHEDULE=0 3 * * *
RETENTION_DAYS=7
MAX_BACKUPS=7
DUMP_ROLES=true
# PowerDNS auth API
PDNS_API_KEY=
# Poweradmin (first-run admin + session)
PA_SESSION_KEY=
PA_ADMIN_USERNAME=
PA_ADMIN_PASSWORD=
PA_ADMIN_EMAIL=
PA_ADMIN_FULLNAME=
# pgweb debug profile
PGWEB_USER=
PGWEB_PASS=
+159
View File
@@ -0,0 +1,159 @@
networks:
frontend:
name: traefik
external: true
backend:
internal: true
edge:
services:
db:
image: postgres:16
container_name: pdns-db
environment:
POSTGRES_DB: postgres
POSTGRES_USER: ${PGUSER:?missing PGUSER}
POSTGRES_PASSWORD: ${PGPASSWORD:?missing PGPASSWORD}
TZ: ${TZ:-Asia/Shanghai}
PGTZ: ${TZ:-Asia/Shanghai}
volumes:
# Keep the existing mount path to avoid moving the current data directory.
- dbdata:/var/lib/postgresql
- ./db-init-generated:/docker-entrypoint-initdb.d:ro
- ./backup:/backup:ro
healthcheck:
test: ["CMD-SHELL", "pg_isready -U \"$${POSTGRES_USER}\" -d \"$${POSTGRES_DB}\""]
interval: 10s
timeout: 5s
retries: 10
restart: unless-stopped
networks: [backend, edge]
auth:
image: powerdns/pdns-auth-50:5.0.6
container_name: pdns-auth
depends_on:
db:
condition: service_healthy
ports:
- "53:53/udp"
- "53:53/tcp"
- "127.0.0.1:8081:8081"
environment:
PDNS_API_KEY: ${PDNS_API_KEY:?missing PDNS_API_KEY}
DB_NAME: ${DB_NAME:?missing DB_NAME}
DB_USER: ${DB_USER:?missing DB_USER}
DB_PASS: ${DB_PASS:?missing DB_PASS}
TEMPLATE_FILES: secrets
volumes:
- ./auth/pdns.conf:/etc/powerdns/pdns.conf:ro
- ./auth/templates.d:/etc/powerdns/templates.d:ro
- ./auth/keys:/var/lib/powerdns
- ./auth/import:/import
- ./auth/export:/export
- ./auth/logs:/var/log/pdns
healthcheck:
test:
[
"CMD-SHELL",
"python3 -c \"import json, os, urllib.request; req = urllib.request.Request('http://127.0.0.1:8081/api/v1/servers/localhost', headers={'X-API-Key': os.environ['PDNS_API_KEY']}); data = json.load(urllib.request.urlopen(req, timeout=3)); assert data['daemon_type'] == 'authoritative'\""
]
interval: 10s
timeout: 5s
retries: 12
restart: unless-stopped
networks: [backend, edge]
poweradmin:
image: poweradmin/poweradmin:stable
container_name: poweradmin
depends_on:
db:
condition: service_healthy
auth:
condition: service_healthy
environment:
DB_TYPE: pgsql
DB_HOST: ${DB_HOST:-db}
DB_PORT: ${DB_PORT:-5432}
DB_NAME: ${DB_NAME:?missing DB_NAME}
DB_USER: ${DB_USER:?missing DB_USER}
DB_PASS: ${DB_PASS:?missing DB_PASS}
PA_PDNS_API_URL: http://auth:8081
PA_PDNS_API_KEY: ${PDNS_API_KEY:?missing PDNS_API_KEY}
PA_DNS_BACKEND: sql
PDNS_VERSION: ${PDNS_VERSION:-50}
DNS_NS1: ${DNS_NS1:-ns1.wsvc.info}
DNS_NS2: ${DNS_NS2:-ns2.wsvc.info}
DNS_HOSTMASTER: ${DNS_HOSTMASTER:-hostmaster.wsvc.info}
PA_APP_TITLE: ${PA_APP_TITLE:-Poweradmin}
PA_TIMEZONE: ${TZ:-Asia/Shanghai}
PA_SESSION_KEY: ${PA_SESSION_KEY:?missing PA_SESSION_KEY}
PA_CREATE_ADMIN: ${PA_CREATE_ADMIN:-1}
PA_ADMIN_USERNAME: ${PA_ADMIN_USERNAME:?missing PA_ADMIN_USERNAME}
PA_ADMIN_PASSWORD: ${PA_ADMIN_PASSWORD:?missing PA_ADMIN_PASSWORD}
PA_ADMIN_EMAIL: ${PA_ADMIN_EMAIL:?missing PA_ADMIN_EMAIL}
PA_ADMIN_FULLNAME: ${PA_ADMIN_FULLNAME:?missing PA_ADMIN_FULLNAME}
TRUSTED_PROXIES: private_ranges
DEBUG: "false"
restart: unless-stopped
networks: [backend, frontend]
labels:
- "traefik.enable=true"
- "traefik.docker.network=traefik"
- "traefik.http.routers.poweradmin.rule=Host(`pdns.wsvc.info`)"
- "traefik.http.routers.poweradmin.entrypoints=websecure"
- "traefik.http.routers.poweradmin.tls.certresolver=letsencrypt"
- "traefik.http.services.poweradmin.loadbalancer.server.port=80"
backup:
# Use postgres:16 so bash/pg_dump/flock exist without runtime package installs.
# backend is internal:true — Alpine apk at start cannot reach mirrors.
image: postgres:16
container_name: pdns-backup
depends_on:
db:
condition: service_healthy
environment:
TZ: ${TZ:-Asia/Shanghai}
DB_HOST: ${DB_HOST:-db}
DB_PORT: ${DB_PORT:-5432}
DB_USER: ${PGUSER:?missing PGUSER}
DB_PASS: ${PGPASSWORD:?missing PGPASSWORD}
DB_NAME: ${DB_NAME:?missing DB_NAME}
RETENTION_DAYS: ${RETENTION_DAYS:-7}
MAX_BACKUPS: ${MAX_BACKUPS:-7}
DUMP_ROLES: ${DUMP_ROLES:-true}
CRON_SCHEDULE: ${CRON_SCHEDULE:?missing CRON_SCHEDULE}
volumes:
- ./backup:/backup
- ./scripts:/scripts:ro
entrypoint: ["/bin/bash", "/scripts/backup-scheduler.sh"]
restart: unless-stopped
networks: [backend]
pgweb:
image: sosedoff/pgweb:0.16.2
container_name: pdns_pgweb
restart: unless-stopped
environment:
PGWEB_DATABASE_URL: "postgres://${PGUSER:?missing PGUSER}:${PGPASSWORD:?missing PGPASSWORD}@${DB_HOST:-db}:${DB_PORT:-5432}/${DB_NAME:?missing DB_NAME}?sslmode=disable"
PGWEB_AUTH_USER: ${PGWEB_USER:?missing PGWEB_USER}
PGWEB_AUTH_PASS: ${PGWEB_PASS:?missing PGWEB_PASS}
TZ: ${TZ:-Asia/Shanghai}
depends_on:
db:
condition: service_healthy
networks: [backend, frontend]
labels:
- "traefik.enable=true"
- "traefik.docker.network=traefik"
- "traefik.http.routers.pgweb.rule=Host(`pgweb.wsvc.info`)"
- "traefik.http.routers.pgweb.entrypoints=websecure"
- "traefik.http.routers.pgweb.tls.certresolver=letsencrypt"
- "traefik.http.services.pgweb.loadbalancer.server.port=8081"
volumes:
dbdata: {}
+5
View File
@@ -0,0 +1,5 @@
# pgdb compose secrets — copy to /opt/database/.env on the host, chmod 600.
# NEVER commit the real values. Generate: openssl rand -hex 24
POSTGRES_PASSWORD=change-me-strong-hex
PGWEB_AUTH_USER=pgweb
PGWEB_AUTH_PASS=change-me-strong-hex
+65
View File
@@ -0,0 +1,65 @@
# pgdb (192.168.55.15) — TimescaleDB + pgweb GUI + nightly backup
#
# Deploy: copy this file to /opt/database/docker-compose.yml on pgdb,
# create /opt/database/.env (chmod 600) from .env.example, plus
# /opt/database/pgweb-bookmarks/{hass,scribe}.toml (chmod 600, contains DB password).
# Then: docker compose config --quiet && docker compose up -d
#
# Rollback: previous launch command is kept at /opt/database/run
# (container is stateless; data lives on /srv/pgdata).
services:
timescaledb:
image: timescale/timescaledb:latest-pg18
container_name: timescaledb
restart: unless-stopped
ports:
- "192.168.55.15:5432:5432" # bind VM IP only (no IPv6 wildcard)
environment:
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
volumes:
- /srv/pgdata:/var/lib/postgresql # data disk (ext4 /dev/sdb1)
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]
interval: 30s
timeout: 5s
retries: 5
start_period: 10s
pgweb:
image: sosedoff/pgweb:latest
container_name: pgweb
restart: unless-stopped
# bind/listen/readonly/sessions/bookmarks-only/bookmarks-dir are CLI flags (no env equivalent in v0.17.0)
command: ["pgweb", "--bind", "0.0.0.0", "--listen", "8081", "--readonly", "--sessions", "--bookmarks-only", "--bookmarks-dir", "/bookmarks"]
ports:
- "192.168.55.15:8081:8081" # LAN only + basic auth (see .env)
environment:
PGWEB_AUTH_USER: ${PGWEB_AUTH_USER}
PGWEB_AUTH_PASS: ${PGWEB_AUTH_PASS}
PGWEB_BOOKMARKS_DIR: /bookmarks
volumes:
- ./pgweb-bookmarks:/bookmarks:ro # bookmark .toml files (contain DB password, keep 0600)
depends_on:
timescaledb:
condition: service_healthy
pg-backup:
image: prodrigestivill/postgres-backup-local:latest # latest = postgres 18 base (pg_dump 18.x)
container_name: pg-backup
restart: unless-stopped
environment:
POSTGRES_HOST: timescaledb
POSTGRES_DB: "hass scribe postgres"
POSTGRES_USER: postgres
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
POSTGRES_EXTRA_OPTS: "-Fc" # custom-format dumps (pg_restore)
SCHEDULE: "0 2 * * *" # nightly 02:00 (TZ=Asia/Shanghai -> local 02:00)
BACKUP_ON_START: "TRUE" # immediate backup on first start
BACKUP_SUFFIX: ".dump"
HEALTHCHECK_PORT: "80" # go-cron health endpoint for the image healthcheck
TZ: "Asia/Shanghai" # match original host-cron 02:00 local (container default is UTC)
volumes:
- /opt/database/backups:/backups # POSIX fs required; root disk, separate from data disk
depends_on:
timescaledb:
condition: service_healthy
+24
View File
@@ -0,0 +1,24 @@
[Unit]
Description=Reconcile pgdb compose stack (timescaledb + pgweb + pg-backup) at boot
Documentation=file:///opt/database/docker-compose.yml
After=network-online.target docker.service
Wants=network-online.target
Requires=docker.service
[Service]
Type=oneshot
RemainAfterExit=yes
WorkingDirectory=/opt/database
# Idempotent boot-time reconcile. docker's own restore can fail to bind the
# published ports (192.168.55.15:5432/8081) when the VM IP is not yet usable
# right after boot (EADDRNOTAVAIL, observed 2026-08-30): timescaledb/pgweb
# then stay stopped until a manual `docker compose up`. This unit retries
# `docker compose up -d` (a no-op when the stack is healthy) until the port
# listens, and force-recreates as a last resort to recover a network-detached
# container. Data lives on bind mounts (/srv/pgdata, /opt/database/backups),
# so recreation is safe.
ExecStart=/bin/bash -c 'for i in $(seq 1 12); do docker compose up -d --remove-orphans; sleep 2; if ss -tln | grep -q "192.168.55.15:5432"; then exit 0; fi; sleep 3; done; echo "pgdb-compose: retries exhausted, force-recreating"; docker compose up -d --force-recreate; sleep 10; ss -tln | grep -q "192.168.55.15:5432"'
TimeoutStartSec=180
[Install]
WantedBy=multi-user.target
+2
View File
@@ -0,0 +1,2 @@
# Soft Serve initial admin public key (used only on first boot)
SOFT_SERVE_INITIAL_ADMIN_KEYS=ssh-ed25519 AAAA... # replace with admin public key
+3
View File
@@ -0,0 +1,3 @@
FROM alpine:3.20
RUN apk add --no-cache sqlite tzdata
WORKDIR /scripts
+59
View File
@@ -0,0 +1,59 @@
services:
soft-serve:
image: charmcli/soft-serve:v0.12.2
container_name: soft-serve
restart: unless-stopped
# non-root (uid 1000 = windy; 与 backup sidecar BACKUP_UID 一致)
user: "1000:1000"
environment:
SOFT_SERVE_DATA_PATH: /var/lib/soft-serve
SOFT_SERVE_INITIAL_ADMIN: windy
SOFT_SERVE_INITIAL_ADMIN_KEYS: ${SOFT_SERVE_INITIAL_ADMIN_KEYS}
volumes:
- ./data:/var/lib/soft-serve
- soft-serve-app:/soft-serve
networks:
- traefik
labels:
- traefik.enable=true
# SSH over TCP via Traefik (entryPoint ssh -> container port 23231)
- traefik.tcp.routers.softserve-ssh.entrypoints=ssh
- traefik.tcp.routers.softserve-ssh.rule=HostSNI(`*`)
- traefik.tcp.routers.softserve-ssh.tls=false
- traefik.tcp.services.softserve-ssh.loadbalancer.server.port=23231
soft-serve-backup:
build:
context: .
dockerfile: Dockerfile.backup
container_name: soft-serve-backup
restart: unless-stopped
volumes:
- ./data:/data:ro
- ./backups:/backup
- ./scripts:/scripts
environment:
TZ: Asia/Shanghai
BACKUP_UID: 1000
BACKUP_GID: 1000
entrypoint: >
/bin/sh -ec "
umask 077 &&
touch /backup/backup.log &&
crontab /scripts/crontab.txt &&
echo '[INFO] soft-serve backup cron installed' &&
crond -f -l 8
"
volumes:
soft-serve-app:
networks:
traefik:
external: true
name: vw-net
+15
View File
@@ -0,0 +1,15 @@
#!/bin/sh
set -eu
umask 077
D() { date "+%Y-%m-%d %H:%M:%S"; }
TS=$(date +%Y%m%d_%H%M%S)
OUT="/backup/soft-serve_${TS}"
mkdir -p "$OUT"
echo "[$(D)] Starting soft-serve backup -> $OUT"
tar czf "$OUT/repos-config.tar.gz" -C /data repos hooks config.yaml ssh
sqlite3 /data/soft-serve.db ".backup '$OUT/soft-serve.db'"
chmod 600 "$OUT/repos-config.tar.gz" "$OUT/soft-serve.db"
if [ -n "${BACKUP_UID:-}" ] && [ -n "${BACKUP_GID:-}" ]; then
chown -R "$BACKUP_UID:$BACKUP_GID" "$OUT" /backup/backup.log
fi
echo "[$(D)] Backup OK: $(du -sh "$OUT" | cut -f1)"
+4
View File
@@ -0,0 +1,4 @@
# Run soft-serve backup daily at 02:00
0 2 * * * /bin/sh /scripts/backup.sh >> /backup/backup.log 2>&1
# Prune backups older than 14 days daily at 03:00
0 3 * * * /bin/sh /scripts/prune.sh >> /backup/backup.log 2>&1
+5
View File
@@ -0,0 +1,5 @@
#!/bin/sh
set -eu
D() { date "+%Y-%m-%d %H:%M:%S"; }
ls -dt /backup/soft-serve_* 2>/dev/null | tail -n +15 | xargs -r rm -rf
echo "[$(D)] Pruned. Kept $(ls -d /backup/soft-serve_* 2>/dev/null | wc -l) backups (max 14)"
+38
View File
@@ -0,0 +1,38 @@
# .env.example — Vaultwarden (us2.wsvc.info, /opt/vaultwarden)
#
# Non-secret key reference ONLY. Real values live in the server-local .env
# (never commit them). Copy the keys below into the server .env if a key is
# missing; the compose file requires them via ${VAR} / env_file.
# Service identity
DOMAIN=https://auth.wsvc.info
TEMPLATES_FOLDER=
# Postgres (compose services vaultwarden / backup / pg / pgweb)
DB_HOST=pg
DB_PORT=5432
DB_NAME=vaultwarden
DB_USER=vaultwarden
DB_PASS=
# pgweb debug profile
PGWEB_USER=
PGWEB_PASS=
PGWEB_DATABASE_URL=
# SMTP (mailcow mx2.windy.me:587 starttls)
SMTP_HOST=mx2.windy.me
SMTP_PORT=587
SMTP_SECURITY=starttls
SMTP_USERNAME=
SMTP_PASSWORD=
SMTP_FROM=
HELO_NAME=
# Admin console
ADMIN_TOKEN=
# Runtime
UID=1000
GID=1000
IP_HEADER=X-Forwarded-For
+107
View File
@@ -0,0 +1,107 @@
services:
vaultwarden:
image: vaultwarden/server:1.37.2
container_name: vaultwarden
restart: unless-stopped
env_file: ".env"
environment:
DOMAIN: "https://auth.wsvc.info"
DATABASE_URL: "postgresql://${DB_USER}:${DB_PASS}@${DB_HOST}:${DB_PORT}/${DB_NAME}"
volumes:
- ./vw-data:/data
extra_hosts:
- "mx2.windy.me:194.163.160.244"
networks:
- net
depends_on:
pg:
condition: service_healthy
labels:
- "traefik.enable=true"
- "traefik.docker.network=vw-net"
- "traefik.http.routers.vaultwarden.rule=Host(`auth.wsvc.info`)"
- "traefik.http.routers.vaultwarden.entrypoints=websecure"
- "traefik.http.routers.vaultwarden.tls=true"
- "traefik.http.routers.vaultwarden.tls.certresolver=letsencrypt"
- "traefik.http.services.vaultwarden.loadbalancer.server.port=80"
backup:
build:
context: .
dockerfile: Dockerfile.backup
container_name: vaultwarden-backup
restart: unless-stopped
volumes:
- ./backups:/backup
- ./scripts:/scripts
#user: "${UID:-1000}:${GID:-1000}"
environment:
DB_HOST: ${DB_HOST}
DB_PORT: ${DB_PORT}
DB_USER: ${DB_USER}
DB_NAME: ${DB_NAME}
DB_PASS: ${DB_PASS}
BACKUP_UID: ${UID:-0}
BACKUP_GID: ${GID:-0}
TZ: Asia/Shanghai
entrypoint: >
/bin/sh -ec "
umask 077 &&
printf '%s:%s:*:%s:%s\n' \"$$DB_HOST\" \"$$DB_PORT\" \"$$DB_USER\" \"$$DB_PASS\" > /root/.pgpass &&
chmod 600 /root/.pgpass &&
touch /backup/backup.log &&
crontab /scripts/crontab.txt &&
echo '[INFO] Backup cron installed' &&
echo '[INFO] Starting crond...' &&
crond -f -l 8
"
networks: [net]
pg:
image: postgres:16
container_name: vw-db
restart: unless-stopped
environment:
POSTGRES_DB: ${DB_NAME}
POSTGRES_USER: ${DB_USER}
POSTGRES_PASSWORD: ${DB_PASS}
TZ: Asia/Shanghai
PGTZ: Asia/Shanghai
volumes:
- vwdata:/var/lib/postgresql/data
- ./backups:/backup # to import existing dump
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ${DB_USER} -d ${DB_NAME}"]
interval: 10s
timeout: 5s
retries: 10
networks: [net]
pgweb:
profiles: ["debug"]
image: sosedoff/pgweb:0.16.2
container_name: vaultwarden-pgweb
restart: unless-stopped
environment:
# 用 Vaultwarden 的数据库参数拼接连接串
#DATABASE_URL: "postgres://${DB_USER}:${DB_PASS}@${DB_HOST}:${DB_PORT}/${DB_NAME}?sslmode=disable"
PGWEB_AUTH_USER: ${PGWEB_USER}
PGWEB_AUTH_PASS: ${PGWEB_PASS}
TZ: Asia/Shanghai
#ports:
# - "8082:8081" # 本地访问 http://localhost:8082
depends_on:
pg:
condition: service_healthy
networks: [net]
networks:
net:
name: vw-net
external: true
volumes:
vwdata: {}
+51
View File
@@ -0,0 +1,51 @@
# Domain Docs
How the engineering skills should consume this repo's domain documentation when exploring the codebase.
## Before exploring, read these
- **`CONTEXT.md`** at the repo root, or
- **`CONTEXT-MAP.md`** at the repo root if it exists — it points at one `CONTEXT.md` per context. Read each one relevant to the topic.
- **`docs/adr/`** — read ADRs that touch the area you're about to work in. In multi-context repos, also check `src/<context>/docs/adr/` for context-scoped decisions.
If any of these files don't exist, **proceed silently**. Don't flag their absence; don't suggest creating them upfront. The `/domain-modeling` skill (reached via `/grill-with-docs` and `/improve-codebase-architecture`) creates them lazily when terms or decisions actually get resolved.
## File structure
Single-context repo (most repos):
```
/
├── CONTEXT.md
├── docs/adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
```
Multi-context repo (presence of `CONTEXT-MAP.md` at the root):
```
/
├── CONTEXT-MAP.md
├── docs/adr/ ← system-wide decisions
└── src/
├── ordering/
│ ├── CONTEXT.md
│ └── docs/adr/ ← context-specific decisions
└── billing/
├── CONTEXT.md
└── docs/adr/
```
## Use the glossary's vocabulary
When your output names a domain concept (in an issue title, a refactor proposal, a hypothesis, a test name), use the term as defined in `CONTEXT.md`. Don't drift to synonyms the glossary explicitly avoids.
If the concept you need isn't in the glossary yet, that's a signal — either you're inventing language the project doesn't use (reconsider) or there's a real gap (note it for `/domain-modeling`).
## Flag ADR conflicts
If your output contradicts an existing ADR, surface it explicitly rather than silently overriding:
> _Contradicts ADR-0007 (event-sourced orders) — but worth reopening because…_
+17
View File
@@ -0,0 +1,17 @@
# docs/archive — 归档文档
归档 = 单次调研、已过期,或与当前运维无行动指向的内容。恢复使用前先确认
内容仍与线上状态一致(本仓库原则:先证据后变更,live state 优先)。
## 归档清单
| 文件 | 归档日期 | 原位置 | 说明 |
|------|---------|--------|------|
| `lan-dns-alternatives.md` | 2026-08-17 | `docs/` | DNS 技术选型调研,零引用,无在途决策 |
| `agent-runbook-guide.md` | 2026-08-22 | `docs/` | 上游参考存档;仓库落地规范为 `RUNBOOKS.md` |
| `lan-core-switch-upgrade-plan.md` | 2026-08-22 | `docs/` | SE5420 历史规划参考;执行以 `docs/lan-se5420-deployment-guide.md` 为准 |
| `lan-rb5009-upgrade.md` | 2026-08-22 | `docs/` | 未采购的 ER-X→RB5009 休眠备选方案;其中 PVE 透传与 VLAN10 调研仍可参考 |
| `se5420-review-claim-verification-2026-08.md` | 2026-08-22 | `docs/` | 一次性评审复核调研(现场只读复核结论) |
> 购物类文档(打印机购买指南、交换机选型调研)已按整改计划移入 Obsidian
> vault`~/Documents/vault/my-vault/02_Areas/House/`),不在本目录。
+326
View File
@@ -0,0 +1,326 @@
# Agent Runbook 实用指南(v1
> **定位**:本指南用于把团队的重复性运维、交付与故障处理经验写成可由 Agent 安全执行的流程。它适用于以 Git 仓库为中心的工程协作模式,优先采用 **Markdown + Git 版本控制 + 明确的 Agent 路由规则**,而不是一开始引入复杂的自动化平台。
> 本文件为上游参考存档。仓库内落地规范见 [`RUNBOOKS.md`](../../RUNBOOKS.md),标准模板见 [`runbooks/_template.md`](../../runbooks/_template.md),索引见 [`runbooks/README.md`](../../runbooks/README.md)。
## 1. 什么是 Agent Runbook
Runbook 是预先设计的、可重复执行的操作流程,用于处理部署、告警、故障、配置变更、CI 修复等标准化工作。传统 Runbook 的主要读者是人;**Agent Runbook 则必须把人的隐性判断显式化**,使 Agent 能知道做什么、看到什么才算正常、下一步去哪里、何时停止以及如何撤销。
Google SRE 强调在事故发生前设计响应流程、系统化排障,并逐步将重复性运维工作自动化。[1] [2] AWS Systems Manager Automation 则把可执行 Runbook 建模为顺序步骤:每个步骤调用一个动作,前一步输出可以传递给后续步骤。[3] 这两种思路共同构成了 Agent Runbook 的实用基础。
| 层次 | 核心问题 | 应承担的职责 |
|---|---|---|
| `AGENTS.md` | **何时使用哪份流程?** | 工作路由、通用操作约束、无匹配流程时的默认行为 |
| `runbooks/*.md` | **这件事按什么流程做?** | 前置条件、分步操作、决策分支、验证、停止条件与回滚 |
| Skill / MCP / Tool | **有哪些可调用能力?** | 具体能力、参数、权限边界和使用说明 |
| Shell / GitHub / Linear / SSH 等 | **实际如何执行?** | 对系统、代码库或外部服务执行操作 |
## 2. 设计目标与适用边界
Agent Runbook 的目标不是让 Agent 在所有异常下“想办法修好”,而是在一个**已知、受控、可验证、可回退**的边界中提高执行一致性。它应当优先覆盖高频、后果明确、流程稳定的操作,例如 CI 失败定位、Issue 到合并请求、发布前检查、标准部署、回滚及网络变更。
| 适合纳入 Runbook | 暂不适合直接自动执行 |
|---|---|
| 明确输入、固定步骤、可观察结果的操作 | 目标或验收标准尚不清楚的探索性任务 |
| 可在每次修改后验证状态的变更 | 缺失关键参数、权限或上下文的任务 |
| 具有安全回滚路径的发布与配置调整 | 高破坏性、不可逆或影响面未知的操作 |
| 可由权限与审批规则约束的运维流程 | 与既有流程事实冲突、无法判断根因的异常场景 |
> **基本原则**:当实际状态与 Runbook 的假设冲突,Agent 应停止并呈报,而不是补全未知信息、绕过检查或继续试错。
## 3. Agent Runbook 的最小字段
与普通人工 Runbook 相比,Agent Runbook 必须显式包含以下六类控制信息。缺少其中任一项,都会增加盲目执行或错误恢复的风险。
| 字段 | 作用 | 写作要求 |
|---|---|---|
| **Action** | 定义当前要执行的动作 | 使用可观察、可执行的动词;避免“检查一下”“适当调整”等模糊表述 |
| **Expected** | 描述正常状态或预期输出 | 给出具体信号、阈值、状态码、测试结果或页面表现 |
| **Decision** | 定义分支与下一跳 | 用“条件 → 下一步”的形式;无法判断时指向 `STOP` |
| **Verification** | 确认变更真正生效 | 在每个有副作用的步骤后执行,不能被跳过 |
| **Stop condition** | 规定何时不得继续 | 明确列出信息缺失、状态冲突、权限不足、验证失败等条件 |
| **Rollback** | 描述如何恢复到变更前状态 | 标明触发条件、前提、撤销步骤及回滚后的验证方式 |
## 4. 推荐目录与路由机制
建议把流程与代码一起保存在 Git 仓库中。这样 Runbook 可以评审、版本化、随系统演进更新,也能与相关 Issue、PR 和配置建立可追溯关系。
```text
repo/
├── AGENTS.md
├── RUNBOOKS.md
├── runbooks/
│ ├── README.md
│ ├── issue-to-merge.md
│ ├── fix-ci.md
│ ├── release.md
│ ├── rollback.md
│ ├── network-change.md
│ └── network-recovery.md
└── ...
```
### `AGENTS.md`:只做路由与通用约束
`AGENTS.md` 不应重复流程细节。它只需要规定 Agent 在进行操作类工作前,先查找最具体且适用的 Runbook,并严格遵守其中的步骤、验证、停止和审批要求。
```markdown
# Operational Rules
Before performing operational work:
1. Inspect `runbooks/`.
2. Select the most specific applicable runbook.
3. Follow its steps in order.
4. Do not skip verification steps.
5. Respect STOP and approval conditions.
6. If no runbook applies, diagnose only; do not mutate production state.
## Routing
- CI failure → `runbooks/fix-ci.md`
- GitHub issue implementation → `runbooks/issue-to-merge.md`
- Deployment → `runbooks/release.md`
- Rollback → `runbooks/rollback.md`
- Network configuration → `runbooks/network-change.md`
- Network outage → `runbooks/network-recovery.md`
```
### `RUNBOOKS.md`:仓库级规范
`RUNBOOKS.md` 用于统一所有 Runbook 的字段、命名、评审要求和变更规则。每份 Runbook 只描述一种可识别的操作意图;如果流程已有明显分叉,应拆分为独立文件,而不是堆叠成长篇“万能流程”。
## 5. 规范模板
以下模板可直接保存为 `runbooks/_template.md` 使用。
```markdown
# Runbook: <名称>
## Purpose
说明本 Runbook 要解决的问题及成功结果。
## Scope
- 适用环境:<如 development / staging / production>
- 适用对象:<服务、仓库、组件或告警类型>
- 不适用情形:<需要改用其他 Runbook 或转人工的场景>
## Ownership
- Owner<团队或角色>
- Last reviewed<YYYY-MM-DD>
- Related systems<系统名称>
## Preconditions
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
## Inputs
| 输入 | 来源 | 是否必需 | 校验方法 |
|---|---|---:|---|
| <参数> | <来源> | 是/否 | <如何确认有效> |
## Safety
### Non-negotiable rules
- 先只读诊断,后执行变更。
- 不得把删除现有配置作为首次恢复动作。
- 不得猜测或编造缺失参数。
- 不得绕过失败的测试、检查或审批。
- 每次变更后必须完成对应验证。
- 破坏性操作必须获得明确批准。
### Stop conditions
- 实际状态与本文档的前提或预期结果冲突。
- 缺少必要输入、权限、审批或回滚能力。
- 验证失败且本文档没有明确的下一步。
- 影响范围超出 Scope。
### Approval gates
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|---|---|---:|---|
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
## Procedure
### Step 1 — Diagnose
**Action**
<执行只读诊断动作。>
**Expected**
<列出预期输出、状态或证据。>
**Decision**
- 若 <条件 A>,进入 Step 2。
- 若 <条件 B>,进入 Troubleshooting A。
- 若无法判断或状态冲突,`STOP` 并记录证据。
### Step 2 — Change
**Action**
<描述单一、可审计的变更动作。>
**Expected**
<变更后应出现的状态。>
**Verification**
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
**Rollback**
- 触发条件:<什么情况需要回滚>
- 回滚动作:<如何撤销>
- 回滚验证:<如何确认恢复成功>
## Troubleshooting
### Troubleshooting A — <异常名称>
- 证据收集:<日志、指标、命令输出、链接>
- 允许动作:<仅限已验证且低风险的动作>
- 下一步:<回到某步 / 转入另一 Runbook / STOP 并升级>
## Final Verification
只有同时满足以下标准,流程才算成功:
- <功能或服务状态>
- <自动化测试或健康检查>
- <监控指标或告警状态>
- <变更记录、PR 或 Issue 已更新>
## Failure Handling
若未能完成:
1. 停止进一步变更。
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
4. 按 <升级渠道> 交接,不继续猜测。
## References
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
```
## 6. 编写步骤的标准写法
每个步骤应只承担一个清晰目的,并使用“动作—预期—决策”的闭环表达。如下表所示,前者会导致 Agent 自主扩大操作范围,后者则为其提供安全边界。
| 不推荐写法 | 推荐写法 |
|---|---|
| “检查部署是否正常,不正常就修复。” | “读取部署状态与最近一次发布记录。若所有副本 `Ready` 且版本等于目标版本,进入 Final Verification;若副本未就绪,收集事件与日志并进入 Troubleshooting A;若版本不匹配且原因未知,`STOP`。” |
| “必要时修改配置。” | “仅当配置差异与变更单 `CHG-123` 完全一致且审批已记录时,应用指定键的值;应用后运行健康检查;失败则按 Rollback 回退。” |
| “测试失败时可先跳过。” | “任何必需测试失败均不得继续部署。记录失败测试、日志和提交版本;仅按 Troubleshooting B 处理。” |
## 7. 通用安全规则
以下规则适合在每份 Runbook 的 `Safety` 章节中复用。若某流程存在更严格要求,应以更严格要求为准。
```markdown
## Safety Rules
- Never delete an existing configuration as the first recovery action.
- Prefer read-only diagnosis before mutation.
- After every mutation, verify the expected state.
- If actual state conflicts with this runbook, STOP.
- Do not invent missing parameters.
- Do not bypass failed tests.
- Destructive actions require explicit approval.
```
这些约束体现了一个关键顺序:**先证据,后变更;先小范围,后扩大;先验证,后结束;不确定则停止。** 特别是停止条件必须可操作,例如“权限不足”“缺少变更单”“生产状态与前提不一致”“错误率超过 1%”等,而不应写成“情况复杂时停止”。
## 8. 运行与审计流程
Agent 执行 Runbook 时,应按照固定运行模型工作。每一步的输入、动作、输出和下一跳都应可追踪,这与 AWS 自动化 Runbook 的顺序步骤和输出传递思想一致。[3]
```text
输入与前置条件
只读诊断
确认预期状态或决策分支
获取审批(如需要)
执行最小变更
立即验证
成功收尾 / 回滚 / 停止并升级
```
| 阶段 | Agent 必须产出的证据 | 禁止行为 |
|---|---|---|
| 输入确认 | 参数来源、环境、目标资源、权限与审批状态 | 用猜测值补全必需参数 |
| 诊断 | 命令输出、日志、指标或页面状态 | 在未诊断前直接修改生产状态 |
| 变更 | 实际执行内容、变更范围、时间 | 将多个无关变更混在一起执行 |
| 验证 | 测试、健康检查、监控状态与预期对比 | 以“命令执行成功”代替业务验证 |
| 失败处理 | 已做步骤、异常证据、回滚状态和升级对象 | 无限制重试或绕过失败检查 |
## 9. 从人工操作到自动化的成熟路径
不建议在流程尚未稳定时先构建复杂 DSL 或全自动编排。应先积累真实案例,把可重复部分固化为 Markdown Runbook,再把已稳定、低歧义、可验证的操作迁移到脚本、CI、Skill 或自动化系统。Google SRE 将能够由机器替代的重复性人工工作视为应逐步消除的 toil。[4]
| 阶段 | 主要形式 | 人的角色 | 自动化边界 |
|---|---|---|---|
| 1. 人工处理 | 现场处置与复盘 | 执行、判断、记录 | 不自动化 |
| 2. Markdown Runbook | 固化步骤与证据要求 | 审核流程与异常判断 | Agent 可辅助诊断 |
| 3. Agent + Runbook | 严格按流程执行 | 审批高风险动作、处理例外 | 受停止条件约束的执行 |
| 4. Script / Skill / CI / Automation | 把稳定步骤程序化 | 处理异常和维护自动化 | 自动完成重复性操作 |
| 5. 人工审批 + 自动执行 | 常规流程端到端运行 | 决策、审计与治理 | 审批门控下的自动变更 |
## 10. 上线前检查清单
在将一份新 Runbook 交给 Agent 使用前,建议由流程所有者按以下清单审核。
| 检查项 | 合格标准 |
|---|---|
| 问题边界 | Purpose 与 Scope 清楚描述适用和不适用情形 |
| 输入 | 所有必需输入都有来源、格式和校验方法 |
| 步骤 | 每一步均有 Action、Expected 与明确的下一跳 |
| 变更控制 | 所有修改动作都有 Verification;关键动作有 Rollback |
| 安全控制 | Stop conditions、审批门槛和禁止行为已列明 |
| 异常处理 | 失败时知道收集什么证据、交给谁,而非继续猜测 |
| 可维护性 | 有 Owner、最近复审日期与关联文档;已在版本控制中评审 |
| 可演练性 | 已在安全环境或历史案例上走通至少一次 |
## 11. 建议的首批 Runbook
首次落地时,应优先选择频率较高、输入相对明确、变更可回退的场景。以下集合通常能覆盖大部分工程协作的基础需求。
| Runbook | 目的 | 关键安全控制 |
|---|---|---|
| `issue-to-merge.md` | 从已明确 Issue 到可评审变更 | Scope 锁定、测试门槛、PR 证据 |
| `fix-ci.md` | 诊断并修复 CI 失败 | 不跳过测试、不修改无关代码 |
| `release.md` | 执行标准发布 | 发布窗口、审批、健康检查、回滚点 |
| `rollback.md` | 恢复到已知稳定版本 | 明确触发条件、版本选择、回滚后验证 |
| `network-change.md` | 实施受控网络配置变更 | 影响评估、变更单、回退配置 |
| `network-recovery.md` | 处理网络异常与服务恢复 | 只读诊断优先、状态冲突即停止 |
## 12. 结论
Agent Runbook 的价值不在于把每一项运维工作立即自动化,而在于将团队的工程判断编码为**可路由、可验证、可停止、可回滚**的操作系统。对于多数团队,从仓库中的 `AGENTS.md``RUNBOOKS.md` 和一组 Markdown Runbook 起步,已经足够实用。
当某个流程经过多次执行、输入稳定、异常分支收敛且验证可靠后,再将其下沉为脚本、CI 或其他自动化能力。这样既能逐步降低重复性 toil,也能始终保留人类对高风险和例外情形的决策权。[4]
## References
[1]: https://sre.google/sre-book/managing-incidents/ "Google SRE Book — Managing Incidents"
[2]: https://sre.google/sre-book/effective-troubleshooting/ "Google SRE Book — Effective Troubleshooting"
[3]: https://docs.aws.amazon.com/systems-manager/latest/userguide/automation-documents.html "AWS Systems Manager — Creating your own runbooks"
[4]: https://sre.google/sre-book/eliminating-toil/ "Google SRE Book — Eliminating Toil"
[5]: https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-automation.html "AWS Systems Manager Automation"
[6]: https://docs.aws.amazon.com/systems-manager-automation-runbooks/latest/userguide/automation-runbook-reference.html "AWS Systems Manager Automation Runbook Reference"
[7]: https://learn.microsoft.com/en-us/azure/automation/manage-runbooks "Microsoft Learn — Manage runbooks in Azure Automation"
---
**来源**Manus AI《Agent Runbook 实用指南(v1.0)》,本仓库存档为规范参考。
@@ -1,13 +1,13 @@
# LAN 核心交换机升级计划(保留 ER-X)
**状态:** SE5420 **已采购**2026-08-09)。**实施与验证以 [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) 为准**
**状态:** SE5420 **已采购**2026-08-09)。**实施与验证以 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 为准**
本文为历史规划参考,**不得作为现场执行步骤**;所有实际操作均以部署指南为准。
**锁定硬件:** TP-Link **`TL-SE5420`**16 × 2.5GbE RJ45 + 4 × 10GbE SFP+)。
**目标:** SE5420 承接全部 LAN 物理接入与二层转发;ER-X 继续承担公网、NAT、防火墙、
LAN66/LAN55 网关与 DHCP。
**拓扑与流量的详细说明**(职责、逻辑网、流量路径、Wi-Fi 分工、验收边界)见:
[lan-erx-se5420-network.md](lan-erx-se5420-network.md)。
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)。
SE5420 官方资料:静态功耗 8 W、最大功耗 32 WVLAN、LACP、STP/RSTP/MSTP、ACL、
CLI/SNMP、配置导入导出与固件下载。无 PoE——AP 使用本地取电 + 普通网线。
@@ -65,7 +65,7 @@ VLAN tag。
```
完整端口表、流量路径与 Wi-Fi 分工见
[lan-erx-se5420-network.md](lan-erx-se5420-network.md)。
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)。
## VLAN 与端口设计
@@ -124,8 +124,8 @@ VLAN tag。
### 阶段 4VLAN10 升级专用 Wi-Fi(独立项目)
不与本次 Done 捆绑。仅 U6;网关 `gfw`;详见
[lan-erx-se5420-network.md](lan-erx-se5420-network.md) 第 6.5 / 8 节与
[unifi-network.md](unifi-network.md)。
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md) 第 6.5 / 8 节与
[unifi-network.md](../unifi-network.md)。
## 性能预期与不变瓶颈
@@ -148,9 +148,9 @@ VLAN tag。
## 参考
- [ER-X + SE5420 网络与拓扑说明](lan-erx-se5420-network.md)
- [LAN 概览](lan-overview.md)
- [ER-X 配置记录](edgerouter-x-configuration.md)
- [UniFi 网络](unifi-network.md)
- [`gfw`](../hosts/gfw.windy.lan.md)
- [ER-X + SE5420 网络与拓扑说明](../lan-erx-se5420-network.md)
- [LAN 概览](../lan-overview.md)
- [ER-X 配置记录](../edgerouter-x-configuration.md)
- [UniFi 网络](../unifi-network.md)
- [`gfw`](../../hosts/gfw.windy.lan.md)
- [TL-SE5420 官方规格](https://www.tp-link.com.cn/product_2899.html?v=specification)
+605
View File
@@ -0,0 +1,605 @@
# Home-LAN DNS alternatives for the windy LAN (research, 2026-08)
**Status: research only. No configuration was changed.** This page evaluates
resolvers/splitters that are genuinely better than — or meaningfully different
from — the current "AdGuard Home (AGH) + mosdns" setup on
[`dns.windy.lan`](../../hosts/dns.windy.lan.md) (`.36`), for a GFW-constrained
China home LAN. Claims are cited to primary sources (official repos, official
docs, upstream READMEs); anything not verified is flagged as such.
> 2026-08-12: facts in this page's scope recap were refreshed by W1N-56 live
> verification — mosdns on `.1` is **not idle**, it is clash's
> `nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution); the
> canonical decision record is
> [`lan-dns-architecture.md`](../lan-dns-architecture.md) (final verdict aligned,
> Phase 0 kill-test evidence incl. a measured upstream-blackhole degradation
> gap).
Scope recap (from [`lan-overview.md`](../lan-overview.md), verified 2026-08-06):
- Clients get DNS via EdgeRouter DHCP option 6 → AGH `192.168.66.36:53`.
- AGH upstreams: `dns.alidns.com` + `doh.pub` DoH (load-balanced), fallback
`https://adg.chans.xyz/dns-query`. **DNSSEC disabled** (known-bad-signature
check failed on the selected path). Rewrites: `hass.local` / `hass.windy.lan`.
- `gfw` OpenWrt (`.1`) runs OpenClash fake-ip + TPROXY; dnsmasq → clash DNS
`127.0.0.1#7874`. `mosdns` on `127.0.0.1:6052` is clash's
`nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution: domestic →
AGH `.36:53`, foreign → `223.5.5.5`/`119.29.29.29`); it is **not** in the LAN
client query path.
- No local authoritative PTR source yet; private reverse DNS is a known gap.
---
## 1. TL;DR / recommendation
**The current stack is already 80% of the answer.** AGH is a strong LAN DNS
front-end (filtering, rewrites, per-client upstreams, query log, web UI) and its
upstream layer — **per-domain upstreams** plus a **per-domain list loaded from a
file** (`upstream_dns_file`) — is exactly the mechanism the official docs
recommend for accelerating China CDN domains while keeping everything else on a
trusted path. [AGH configuration: upstreams](https://adguard-dns.io/kb/adguard-home/configuration/).
The genuinely worthwhile changes, in order of value:
1. **Add geo-split inside AGH** via `upstream_dns_file` fed by a converted
`accelerated-domains.china.conf` ([felixonmars/dnsmasq-china-list](https://github.com/felixonmars/dnsmasq-china-list)):
domestic CDN domains → `dns.alidns.com` / `doh.pub`; everything else →
the trusted foreign path (currently `adg.chans.xyz`). This is a documented
AGH use case, requires **no new daemon**, and removes the need for mosdns.
This is the top recommendation.
2. **Re-enable real DNSSEC** by putting validation behind AGH: AGH's
`enable_dnssec` only sets the DO bit — it does not validate
([AGH config: DNSSEC](https://adguard-dns.io/kb/adguard-home/configuration/)).
The two realistic ways are (a) point the foreign/trusted default upstream at
a validating resolver ([unbound](https://unbound.docs.nlnetlabs.nl/en/latest/),
[blocky](https://0xerr0r.github.io/blocky/latest/configuration/#dnssec-validation))
and re-test a known-bad-signature domain; or (b) insert a validating
resolver (blocky is the lightest) between AGH and the upstreams.
3. **mosdns on `.1` is resolved, not idle** — it is clash's
`nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution, verified
2026-08-12), so "delete it" is off the table; its role is documented in
[`lan-dns-architecture.md`](../lan-dns-architecture.md) §1. If a future change
moves this role to an AGH-side companion, keep in mind mosdns's cache strips
EDNS0 and it performs no DNSSEC validation
([mosdns v5 executable plugins](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)).
Top-3 alternatives worth pursuing (see §3 for detail):
| Rank | Option | Why |
|------|--------|-----|
| 1 | **AGH with China-list geo-split (`upstream_dns_file`)** | Documented AGH pattern; single box; no new service; keeps filtering/rewrites/UI. |
| 2 | **Blocky as validating backend behind AGH** | The only "new software" option that adds real in-process DNSSEC validation + conditional per-domain upstreams + ECS in one static binary ([blocky README](https://github.com/0xERR0R/blocky), [config](https://0xerr0r.github.io/blocky/latest/configuration/)). |
| 3 | **Unbound as validating recursive resolver** (replaces forwarders for the foreign path, or whole path) | True validation, full recursion (fail-open by nature), private `local-zone`s; heavier ops than AGH's file-driven split. |
Explicitly **not** recommended as replacements here: smartdns and chinadns-ng
(both excellent *splitters*, but neither validates DNSSEC and both lack AGH's
filtering/UI/query-log layer, so they add a daemon without closing the DNSSEC
gap); mihomo/sing-box DNS as the primary path (couples DNS to the proxy and is
fail-closed; keep for proxy-side concerns only); knot-resolver/dnsdist (overkill
for a single-operator home LAN).
---
## 2. Requirement matrix
Legend: **●** native/built-in · **◐** possible with config/lists · **○** absent/
not applicable. "Geo-split" = route domestic vs foreign names to different
upstreams. "Anti-pollution" = a mechanism to avoid/adjudicate poisoned answers
(IP-verdict or trusted-upstream routing). "DNSSEC" = performs validation
in-process (not just forwards DO).
| Candidate | Geo-split | Anti-pollution | DNSSEC (validate) | Cache | Private names / rewrites | Ops simplicity | License |
|---|---|---|---|---|---|---|---|
| AGH (current) | ◐ per-domain upstreams + list file | ◐ via trusted foreign upstream | ○ (DO bit only) | ● | ● rewrites, per-client, private-PTR | ● Docker + UI | GPL-3.0 |
| mosdns v5 | ● domain/ip list matchers | ◐ forward foreign→trusted | ○ | ● (strips EDNS0) | ● hosts/redirect/reverse_lookup | ◐ single binary, YAML, no UI | GPL-3.0 |
| smartdns | ● nameserver groups + domain lists | ● bogus-nxdomain / blacklist-ip / trusted groups | ○ (no option in config ref) | ● serve-expired | ● address / local-domain / lease file | ◐ single binary, optional WebUI plugin | GPL-3.0 |
| chinadns-ng | ● chnlist/gfwlist + tag:none IP-test | ● IP verdict via chnroute ipset/nftset | ○ | ● cache/stale/verdict | ◐ hosts / dns-rr-ip | ◐ single static binary, config file | AGPL-3.0 |
| dnsmasq-china-list | ◐ (data only) | ◐ (via host resolver) | ◐ via host | ◐ via host | ◐ via host | ◐ feed lists | WTFPL |
| unbound | ◐ forward-zones / RPZ / views | ◐ forward-zones + bogus-nxdomain | ● | ● serve-expired | ● local-zone / local-data | ◐ config daemon, no UI | BSD-style (NLnet) |
| blocky | ◐ conditional per-domain + client groups | ◐ blocking lists + conditional routing | ● | ● prefetch | ● customDNS / rewrite / hosts | ◐ single binary, YAML, REST (no full web UI) | Apache-2.0 |
| Technitium | ◐ conditional-forwarder zones / apps | ◐ blocked lists + forwarding | ● | ● persistent | ● zones, stub, split-horizon | ● .NET + web console | GPL-3.0 |
| sing-box | ● DNS rules (geoip/geosite) | ● rule-based servers + (proxy) sniffing | ○ | ● LRU + optimistic | ● hosts / local server | ◐ single binary, JSON | GPLv3-family (metadata "other") |
| mihomo | ● nameserver-policy + fallback-filter | ● geoip verdict + geosite | ○ | ● (cache-algorithm) | ● hosts; fake-ip-filter for `.lan` | ◐ single binary, YAML | not cleanly verifiable (repo obfuscated) |
| knot-resolver | ◐ policy modules | ◐ policy + RPZ | ● | ● persistent | ◐ hints / local data | ◐ systemd, Lua config | open source (CZ-NIC) |
| dnsdist | ◐ Lua rules (custom) | ◐ custom policies | ○ (balancer, not validator) | ○ (no cache of its own) | ○ | ○ power tool | GPL (PowerDNS) |
Notes:
- "Geo-split" for AGH/blocky/unbound/Technitium is real but requires feeding a
China domain list; chinadns-ng/mihomo additionally offer the **IP-verdict**
path for domains not in any list (query both, adopt CN result only if the
answer IP is mainland).
- mosdns v5's `cache` plugin ignores request EDNS0 and strips response EDNS0
([cache plugin](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)) —
relevant because AGH in front of it relies on the DO bit for DNSSEC-capable
upstreams.
- License for mihomo/sing-box/knot-resolver marked conservative: GitHub
metadata is "other"/custom or deliberately obfuscated; see §3 caveats.
---
## 3. Per-candidate evaluation
### 3.1 AdGuard Home — advanced upstream routing / built-ins
What it is: Go DNS proxy + adblock + DHCP, LAN DNS front-end
([official](https://adguard-dns.io/kb/adguard-home/overview/)).
Capabilities relevant here (all from the official configuration page):
- **Per-domain upstreams** dnsmasq-style: `[/domain/]upstream`, wildcards,
`#` = "default upstreams", empty `//` = unqualified names
([upstreams for domains](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-for-domains)).
- **List from file** `upstream_dns_file` — the docs *explicitly* call out China
CDN acceleration via dnsmasq lists, with the `server=/0-100.com/114.114.114.114`
`[/0-100.com/]114.114.114.114` conversion
([loading upstreams from file](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-from-file)).
- Upstream modes: `load_balance`, `parallel`, `fastest_addr`; plus `fallback_dns`
used only when primary upstreams fail
([config file: dns](https://adguard-dns.io/kb/adguard-home/configuration/)).
- Per-client upstreams (`clients.persistent[].upstreams`), rewrites
(`filtering.rewrites`, incl. wildcard), `local_ptr_upstreams` for private PTR,
ECS (`edns_client_subnet` with `use_custom` coarse prefix), optimistic cache
([same page](https://adguard-dns.io/kb/adguard-home/configuration/)).
- **DNSSEC is DO-bit only**: `enable_dnssec` "defines whether the proxy should
set the DO flag in the upstream requests" — validation must happen upstream
([same page](https://adguard-dns.io/kb/adguard-home/configuration/)).
- DoH/DoT/DoQ/DoH3 serving, `bind_hosts`/ACL guidance
([running securely](https://adguard-dns.io/kb/adguard-home/running-securely/)).
Verdict: **Already installed and capable of the geo-split itself.** The current
setup under-uses it: only a load-balanced CN pair + fallback, no per-domain
routing and no validating upstream. This is the cheapest "better" state — see §5.
### 3.2 mosdns v5 — installed, active as clash nameserver (gateway-side)
What it is: "一个 DNS 转发器" (a DNS forwarder) — plugin-based, sequence-driven
([README](https://github.com/IrineSistiana/mosdns), GPL-3.0, ~3.7k★).
What it does (verified from the v5 wiki and source tree):
- Servers: `udp_server`, `tcp_server` (TLS→DoT), `quic_server`, `http_server`
(DoH); upstreams in `forward` support `udp`, `tcp`, `tls`, `https`, `quic`,
HTTP/3, concurrent racing (`concurrent: n` picks the fastest) and socks5
([server plugins](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/fu-wu-qi-cha-jian.md),
[executable plugins](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)).
- Geo-split: v5 data providers are **`domain_set` / `ip_set` (text list files)**
plus `qname`/`resp_ip` matchers — verified from the current source tree
([plugin/data_provider](https://github.com/IrineSistiana/mosdns/tree/main/plugin/data_provider))
— and an `ipset`/`nftset` exec plugin to push answer IPs to kernel sets. The
old v4-style `geosite`/`geoip` `.dat` plugins are **not present** in the v5
tree; the v5 wiki's own matcher page currently states there are no matcher
plugins to document
([matcher page](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/pi-pei-qi-cha-jian.md)).
Plan on chnlist/gfwlist-style text lists, not `geosite.dat`.
- Cache: yes, incl. optional lazy cache and disk dump; **request EDNS0 is
ignored and response EDNS0 stripped** by the cache plugin
([cache](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)).
- Private names: `hosts` (domain-rules style, not OS /etc/hosts syntax),
`redirect`, `arbitrary` (zone records), `reverse_lookup` (PTR/HTTP lookup).
- Ops: single binary + YAML; `mosdns service install` ships a systemd/launchd
helper ([v5 overview](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5.md));
Docker image exists. No web UI of its own.
Verdict: capable splitter/forwarder, but **adds no DNSSEC and no filtering
layer**, and its cache interferes with EDNS0/DO handling. As a *back-end* splitter
behind AGH it is a legitimate choice only if DNSSEC stays off. Given AGH can do
the same per-domain split natively (3.1), mosdns's marginal value here is
concurrent upstream racing and ipset/nftset integration — neither is needed at
this LAN's scale. Either wire it up properly or remove it.
### 3.3 smartdns
What it is: local DNS server that queries multiple upstreams, **speed-tests the
answer IPs and returns the fastest**; DoH/DoT/DoQ/DoH3; GPL-3.0, ~11.2k★
([README](https://github.com/pymumu/smartdns)).
Capabilities (from the official config reference and FAQ):
- Multi upstream + "returns the fastest IP", unlike dnsmasq all-servers
([README](https://github.com/pymumu/smartdns)).
- Domain groups: `server ... -group <name>` + `nameserver /domain/group` routing,
per-`bind` port flags (`-group`, `-no-speed-check`…), client rules/MAC/IP
([config options](https://pymumu.github.io/smartdns/configuration/)).
- Anti-pollution tooling: `bogus-nxdomain` (return NXDOMAIN for poisoned IPs),
`blacklist-ip`, `whitelist-ip`, `ignore-ip`, `ipset`/`nftset` export
([same page](https://pymumu.github.io/smartdns/configuration/)).
- ECS: global `edns-client-subnet` and per-server `-subnet`
([same page](https://pymumu.github.io/smartdns/configuration/)).
- Cache: `cache-size`, `serve-expired` (RFC-like stale), `prefetch-domain`,
persistent cache file ([same page](https://pymumu.github.io/smartdns/configuration/)).
- Private names: `address`, `cname`, `local-domain`, `dnsmasq-lease-file`
([same page](https://pymumu.github.io/smartdns/configuration/)).
- **DNSSEC: no validation option appears anywhere in the official config
reference or FAQ** — its pollution model is blacklist/whitelist + trusted
groups + speed selection, not DNSSEC ([config options](https://pymumu.github.io/smartdns/configuration/),
[FAQ](https://pymumu.github.io/smartdns/faq/)). Flagged: verify on the version
you deploy before relying on it.
Verdict: the classic China-home "best-IP" resolver; good splitter, no DNSSEC,
speed-test model optimizes for latency rather than anti-pollution correctness.
Not better than AGH+China-list for this LAN; at most a back-end splitter behind
AGH, with the same DNSSEC caveat as mosdns.
### 3.4 chinadns-ng / chinadns2 / dnsmasq-china-list
**chinadns-ng** (the requested "china-dns-ng"; actual repo `zfl9/chinadns-ng`,
AGPL-3.0, Zig, ~1.4k★) is the maintained rewrite of shadowsocks/ChinaDNS:
- Two upstream groups (china / trust) + `chnlist.txt` / `gfwlist.txt` domain
lists; domains are tagged `chn`/`gfw`/`none`
([README](https://github.com/zfl9/chinadns-ng)).
- `tag:none` names are queried on **both** upstreams and the china answer is
adopted only if its A/AAAA is a mainland IP (tested against a `chnroute`
ipset/nftset loaded into the kernel); verdict caching avoids re-testing and
leaks ([README: 原理/verdict-cache](https://github.com/zfl9/chinadns-ng)).
- Cache with stale + pre-refresh + optional persistence; DoT upstream
(wolfssl build); `hosts` + `dns-rr-ip` local records; `nftset` add for
chn/gfw IPs; **no DoH by design** and **no DNSSEC** — the author's stated
philosophy is "one job, done well"
([README](https://github.com/zfl9/chinadns-ng)).
- Resource footprint is tiny: ~140 KB baseline, ~2.4 MB with 73k+ chnlist +
5.7k gfwlist entries ([README](https://github.com/zfl9/chinadns-ng)).
**chinadns2** (`zfl9/chinadns2`) is the older C predecessor; effectively
superseded by chinadns-ng for new deployments (README not directly fetched —
treat as legacy line).
**dnsmasq-china-list** (felixonmars, ~6.1k★) is data, not a daemon:
`accelerated-domains.china.conf`, `bogus-nxdomain.china.conf`,
`apple.china.conf`, `google.china.conf`, with generators for **dnsmasq,
unbound, bind, dnscrypt-proxy**
([README](https://github.com/felixonmars/dnsmasq-china-list), WTFPL per repo).
Verdict: chinadns-ng is the strongest *pure splitter* for GFW networks (IP
verdict beats pure list-based routing for unknown domains), but it cannot
validate DNSSEC and brings no filtering UI. As AGH's backend it duplicates what
AGH's per-domain upstreams already do; its IP-test mode requires shipping
`chnroute` ipset/nftset into the host. dnsmasq-china-list is best used as the
**data feed** for the AGH `upstream_dns_file` recommendation in §5.
### 3.5 unbound
What it is: validating, recursive, caching resolver from NLnet Labs
([docs](https://unbound.docs.nlnetlabs.nl/en/latest/)).
- **Real DNSSEC validation by default** (trust anchor, chain of trust); the
official home-network guide turns it on explicitly
([home resolver guide](https://unbound.docs.nlnetlabs.nl/en/latest/use-cases/home-resolver.html)).
- Full recursion → does not hard-depend on any upstream or proxy; serve-expired
(RFC 8767), aggressive NSEC, DoH/DoT/DoQ serving and TLS upstreams,
forward-zone/stub-zone/authority-zone, RPZ filtering, views, ECS module
([docs index](https://unbound.docs.nlnetlabs.nl/en/latest/)).
- Private names: `local-zone`/`local-data` for `*.windy.lan`-style names
([unbound.conf(5)](https://unbound.docs.nlnetlabs.nl/en/latest/manpages/unbound.conf.html)).
- No built-in China split: you assemble it with forward-zones fed by
dnsmasq-china-list (`make unbound` generator) + `bogus-nxdomain`; no UI, no
per-client grouping comparable to AGH.
Verdict: the gold standard for the **validation** half. Best used as (a) the
validating upstream behind AGH for the foreign/trusted path, or (b) a full
recursive resolver replacing the forwarders if you accept losing AGH-style
filtering/UI on top — keep AGH in front for that. System-package based, heavier
to operate than blocky but battle-tested.
### 3.6 blocky
What it is: Go DNS proxy + ad-blocker, "fast and lightweight", single static
binary, stateless, Apache-2.0, ~6.9k★
([README](https://github.com/0xERR0R/blocky)).
- **In-process DNSSEC validation**: `dnssec.validate` with DO bit, RRSIG
verification, chain-of-trust, NSEC/NSEC3, custom trust anchors, SERVFAIL on
bogus ([DNSSEC validation docs](https://0xerr0r.github.io/blocky/latest/configuration/#dnssec-validation)).
- Upstreams: `parallel_best` (2 random resolvers, fastest answer), `strict`,
`random`; per-client/per-subnet upstream **groups**; UDP/TCP/DoT/DoH/DoQ/DoH3;
DNS stamps; bootstrap DNS
([upstreams](https://0xerr0r.github.io/blocky/latest/configuration/#upstreams-configuration)).
- Conditional forwarding + `customDNS` mapping/rewrite (the AGH-rewrite
equivalent), hosts files, per-domain upstream routing
([custom DNS / conditional](https://0xerr0r.github.io/blocky/latest/configuration/#custom-dns)).
- ECS: `ecs.useAsClient` / `ecs.forward`
([ECS](https://0xerr0r.github.io/blocky/latest/configuration/#edns-client-subnet-options)).
- Cache with min/max TTL + **prefetching**; optional **Redis** cache/state sync
between instances; query log to SQLite/Postgres/CSV; Prometheus metrics; REST
API ([README](https://github.com/0xERR0R/blocky),
[config](https://0xerr0r.github.io/blocky/latest/configuration/)).
- No full web admin UI (metrics/REST/logs only) — an ops trade-off vs AGH's UI.
Verdict: the most attractive *new software* option for this LAN **as a backend
behind AGH**: it adds real DNSSEC validation + conditional upstream routing +
ECS with a single binary and YAML. It has no China-IP-verdict split built in —
feed it the China domain list via `conditional.mapping`/upstream groups, which
is fine at this scale. One caveat: no GUI means AGH stays the human-facing
front, so AGH→blocky is strictly additive.
### 3.7 Technitium DNS Server
What it is: self-hosted authoritative **and** recursive DNS server, .NET,
web console, GPL-3.0, ~9.5k★
([README](https://github.com/TechnitiumSoftware/DnsServer)).
- **DNSSEC validation** for recursive resolution, forwarders, and conditional
forwarders (RSA/ECDSA/EdDSA, NSEC/NSEC3); can also *serve* signed zones
([README](https://github.com/TechnitiumSoftware/DnsServer)).
- Conditional forwarder zones + bulk conditional forwarding app; blocked-domain
lists with regex support and per-client variants; split-horizon/geolocation
via DNS Apps; ECS; QNAME minimization
([README](https://github.com/TechnitiumSoftware/DnsServer)).
- Serving side: DoH/DoT/DoQ/DoH3 server, built-in DHCP, persistent cache,
caching with serve-stale/prefetch, clustering, HTTP/SOCKS5 proxy for DNS
(e.g. over Tor) ([README](https://github.com/TechnitiumSoftware/DnsServer)).
- Heavier footprint (needs .NET; Docker image available) and a full web console
with many features this LAN won't use.
Verdict: capable and genuinely feature-rich (a real AGH alternative in the
"everything in one box" sense — filtering, private zones, validation, DHCP), but
it's more moving parts than this LAN needs, and its geo-split still requires
manual conditional-forwarder lists. Not chosen over the lighter AGH+backend
approach.
### 3.8 sing-box / mihomo built-in DNS as the split resolver (fake-ip)
The "third option": let the proxy engine's DNS own resolution, AGH on top.
**sing-box** DNS object: multiple server types (local, udp, tcp, tls, https,
http3, quic, fakeip, hosts, dhcp, mdns…), rule-based server selection by
geoip/geosite, LRU cache + optimistic serving, per-query timeout, `client_subnet`
(ECS), `reverse_mapping`
([sing-box DNS docs](https://sing-box.sagernet.org/configuration/dns/)).
**mihomo** (Clash.Meta lineage; docs at
[wiki.metacubex.one](https://wiki.metacubex.one/en/config/dns/)):
`nameserver-policy` (geosite/rule-set/domain keys) routes specific domains to
specific resolvers; `fallback` + `fallback-filter` (geoip=CN, geosite=gfw,
ipcidr, domain) adjudicate pollution — a CN resolver's answer is adopted only if
the IP is mainland, otherwise the overseas fallback's answer is used;
`fake-ip`/`redir-host` enhanced mode, `fake-ip-filter` with e.g. `'*.lan'` to
keep local names on real-IP; per-DNS-server ECS; cache-algorithm
([mihomo DNS config](https://wiki.metacubex.one/en/config/dns/)).
Assessment for THIS LAN:
- The **pollution adjudication is strong** (geoip-verdict fallback, geosite
lists), and mihomo already runs on the gateway — so "clash DNS as splitter" is
tempting.
- But the DNS service is **coupled to the proxy**: foreign resolution rides the
proxy path, so when OpenClash/subscription is down, fake-ip mapping and
foreign lookups break (partial fail-open only if `direct-nameserver`/fallback
are carefully set). The LAN requirement says **must not hard-depend on the
proxy (fail-open)**.
- fake-ip adds an indirection layer for anything in front of it (AGH on top
resolves client IPs against fake-ip ranges; leaks/loops need careful rules).
- Neither engine **validates DNSSEC** (no RRSIG verification).
- sing-box repo license shows "other" in GitHub metadata (not cleanly
verifiable); mihomo's repo currently carries **deliberately obfuscated content**
("Void Terminal" parody) — treat `wiki.metacubex.one` as the authoritative
docs and expect the GitHub surface to change.
Verdict: keep clash/mihomo DNS exactly where it is (proxy-side, TPROXY/fake-ip),
do **not** make it the LAN resolver of record. If you ever want its IP-verdict
quality outside the proxy, chinadns-ng gives the same idea with zero proxy
dependency.
### 3.9 knot-resolver / dnsdist — power-resolver options
**knot-resolver** (CZ-NIC): minimal caching validating resolver, modular/Lua,
full DNSSEC validation, forwarding over TLS, query policies, RPZ, views/ACLs,
DNS64, persistent cache, serve-stale, even XDP fast-path
([docs](https://knot-resolver.readthedocs.io/en/stable/)). As powerful as
unbound but with more configuration surface (Lua); overkill for a one-operator
home LAN, though it would do the validating-resolver role well.
**dnsdist** (PowerDNS): "highly DNS-, DoS- and abuse-aware loadbalancer" —
routes traffic to backend servers, Lua/YAML config, runtime console, metrics
([overview](https://dnsdist.org/)). It is a **balancer, not a validator/cache**
— it fronts other resolvers. Overkill; only relevant if you wanted a
multi-backend DNS LB, which this LAN does not.
### 3.10 Emerging / also-considered options
- **AdGuard Home + dnsmasq-china-list** — covered in §3.1/§5; this is the
"emerging best practice" for China CDN splits on AGH and is officially
documented.
- **pi-hole** — adblock/dashboard equivalent of AGH but no per-domain upstream
routing worth choosing it over AGH here (not deeply verified for this write-up;
AGH already satisfies the role).
- **dnscrypt-proxy** — encrypted forwarder with stamp support; a transport
option, not a splitter/validator (not deeply verified for this write-up).
- **coredns** — plugin-based; geo-split is DIY via plugins; no DNSSEC
validation by default (not deeply verified for this write-up).
### 3.11 Other popular options (survey supplement, 2026-08-12)
Follow-up survey of additional popular solutions not covered above, evaluated
against this LAN's constraints (fail-open, keep DNS on `.36`, DNSSEC goal).
None of these change the §4/§5 recommendation.
**Encrypted-forwarder micro-tools (AGH downstream options, not replacements):**
- **dnscrypt-proxy** — the classic OpenWrt encrypted forwarder with
China-list support and DNS-stamp routing. No in-process DNSSEC validation and
no filtering UI; overlaps with AGH's own DoH upstream layer, so its marginal
value here is low.
- **dnsproxy** (AdGuardTeam) — lightweight DoH/DoT/DoQ forwarder/server.
Functionally a subset of AGH's upstream layer; only useful if forwarding logic
is deliberately split out of AGH.
- **Stubby** — dnsmasq→stubby→DoT (privacy-community pattern). Pure
forwarding, no split/filter/validation; adopting it alone would be a
downgrade from AGH.
**Managed / cloud DNS (zero-ops, not self-hosted):**
- **NextDNS / ControlD / AdGuard DNS / Cloudflare** — hosted filtering, logs,
per-device policies. This LAN already self-hosts AGH + a private
`adg.chans.xyz` fallback, so a cloud service would be a downgrade in control
(data leaves the LAN). Only realistic use: add one as an extra foreign-path
upstream inside AGH's `upstream_dns_file`.
**Heavier all-in-one resolvers:**
- **PowerDNS Recursor** — real DNSSEC validation + Lua policy, authoritative
and recursive in one. Capable but overlaps unbound; over-provisioned here.
- **BIND9** — classic authoritative/recursive; can validate DNSSEC and, more
interestingly, serve as a local **authoritative zone** that would close the
private-PTR gap. As a LAN resolver it lacks AGH's filtering/UI and is heavier
to operate; a small dnsmasq authoritative zone is a lighter way to achieve the
PTR goal (still deferred until a local authoritative source exists).
- **hickory-dns / trust-dns** (Rust) — emerging recursive resolver, DNSSEC
friendly, smaller ecosystem/ops track record than unbound/blocky; not yet
worth switching for this LAN.
**Popular stack patterns (structure, not new software):**
- **Pi-hole + unbound** — the most common global self-hosted combo
(filtering front-end + validating backend). AGH already occupies the
Pi-hole role here (and does more), so the equivalent is **AGH + unbound/
blocky** — exactly the report's recommendation #2.
- **dnsmasq + china-list + smartdns** (classic OpenWrt trio) — routes the
China list on the gateway itself. Equivalent to co-locating DNS with the
proxy host (`.1`), which violates the fail-open requirement; not recommended
for this LAN.
Verdict: the survey adds no better candidate. dnsproxy/dnscrypt-proxy duplicate
AGH's upstream layer, cloud DNS is a control downgrade, and the only genuinely
new capability (a local authoritative source for PTR) is better served by a
small dnsmasq authoritative zone than by replacing the resolver.
---
## 4. Architecture recommendation for this LAN
### 4.1 Preferred architecture (change is config-only)
```
clients (DHCP option 6 = .36)
│ UDP/TCP :53
AGH .36 (filtering, rewrites, query log, per-client upstreams)
│ upstream_dns_file:
│ [/cn-domain-list/] dns.alidns.com doh.pub ← CN CDN domains (China list)
│ default: https://adg.chans.xyz/dns-query … ← trusted/foreign path
└→ validating resolver (unbound OR blocky) for the foreign path (optional phase 2)
```
- Front = AGH stays the single LAN DNS box (filtering/rewrites/UI/query log
are its strong suit and are already operating).
- Split = AGH per-domain upstreams fed by a converted dnsmasq-china-list; no
new daemon. This is the documented AGH pattern
([upstreams from file](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-from-file)).
- Validation = add a validating resolver behind AGH for the trusted path
(blocky simplest; unbound most battle-tested) and re-run the known-bad-signature
check that failed before; then flip `enable_dnssec`.
### 4.2 Why not the alternatives as front-ends
- **smartdns / chinadns-ng as the LAN resolver**: they are pure splitters —
no adblock layer, no query log/UI, no DNSSEC. Replacing AGH with either is a
capability downgrade; behind AGH they duplicate AGH's built-in split while
adding a daemon and losing validation. Only chinadns-ng's IP-verdict mode is
genuinely beyond AGH, and it needs kernel ipset/nftset plumbing.
- **mosdns as the AGH backend**: viable splitter, but no validation and its
cache strips EDNS0/DO ([cache plugin](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)),
which fights the DNSSEC goal. It is already idle on the box — configure it
deliberately or remove it.
- **mihomo/sing-box DNS as the resolver of record**: fail-closed + proxy-coupled
+ no validation. Keep as proxy-side concern (§3.8).
- **knot-resolver / dnsdist / Technitium**: capable but over-provisioned;
Technitium is the only one that would *replace* AGH wholesale, and there's no
benefit worth the migration here.
### 4.3 Deployment location
- **Keep DNS on `dns.windy.lan` (.36)**. It is already the DHCP-advertised
resolver; it is a separate VM from the proxy host; DNS therefore stays
independent of OpenClash state (fail-open), which is an explicit requirement.
- **Do not move it to `gfw` (.1)**: the gateway is where OpenClash injects
TPROXY/fake-ip/DNS-hijack rules; co-locating LAN DNS there couples DNS to the
proxy and its restart/update lifecycle.
- A standalone resolver VM adds nothing: both current VMs already sit on the
same PVE hypervisor ([lan-overview.md](../lan-overview.md) §Positioning facts),
so a hypervisor outage takes out either placement equally; a second physical
host for HA is out of scope for a home LAN.
- If you ever run a validating resolver + AGH on `.36`, verify outbound from
`.36` to the foreign upstreams is not re-hijacked by OpenClash (loop check
already mandated in the [AGH review](../adguard-home-official-review-2026-08.md)).
### 4.4 Fail-open, DNSSEC, private names — by candidate
| Concern | How the recommended stack behaves |
|---|---|
| Fail-open when proxy/subscription down | AGH forwards directly to DoH upstreams; `.36`'s outbound is not forced through the proxy in normal ops (no TUN policy routing on `.36` — [dns host facts](../../hosts/dns.windy.lan.md)). With unbound/blocky behind, foreign resolution recurses/validates directly, independent of OpenClash. Avoid mihomo-DNS-as-resolver, which is proxy-coupled. |
| DNSSEC validation | Only unbound, blocky, knot-resolver, Technitium validate in-process. AGH sets DO only; mosdns/smartdns/chinadns-ng/mihomo/sing-box do not. Plan: validate behind AGH, or accept "validating public upstream" (confirm with `dig +dnssec`/known-bad test). |
| Private names / rewrites | AGH `rewrites` (already in use for `hass.windy.lan`) + `local_ptr_upstreams` once a local PTR source exists. blocky: `customDNS` mapping/rewrite + hosts. unbound: `local-zone`. All adequate. |
| Query log / visibility | AGH is the best at this of everything evaluated (14-day anonymized log already configured). |
---
## 5. What would make the current AGH + mosdns setup genuinely better
Concrete, in increasing effort:
1. **Implement the China-list geo-split in AGH itself**
(`upstream_dns_file` + converted `accelerated-domains.china.conf`, default
upstreams = trusted foreign path, `fallback_dns` kept). Official AGH docs
describe exactly this pattern
([loading upstreams from file](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-from-file));
list source: [dnsmasq-china-list](https://github.com/felixonmars/dnsmasq-china-list).
Wire a refresh path (cron/ansible) so the list stays current. Re-test CDN
resolution and the DNSSEC known-bad domain after.
2. **Put a validating resolver on the trusted path** (unbound or blocky), re-run
the known-bad-signature check, then enable AGH DNSSEC. Without this, AGH's
`enable_dnssec` is only a DO-flag — the exact reason it is currently off
([AGH DNSSEC semantics](https://adguard-dns.io/kb/adguard-home/configuration/),
[host facts](../../hosts/dns.windy.lan.md)).
3. **Either fully configure mosdns (systemd service, sequence, lists) or remove
it.** Leaving an idle `127.0.0.1:6052` listener documented as "not the active
path" is drift. If kept, plan around no-EDNS0 cache + no validation; if
removed, drop the listener and its config to reduce surface.
4. **Close the private-PTR gap**: once a local authoritative source exists (e.g.
dnsmasq on `gw`, or a tiny authoritative zone), point AGH
`local_ptr_upstreams` at it as the AGH review recommends
([AGH review](../adguard-home-official-review-2026-08.md));
don't set it before that source exists
([dns host facts](../../hosts/dns.windy.lan.md)).
5. **Optional: ECS** for CDN geo-accuracy — AGH `edns_client_subnet.use_custom`
with a coarse fixed prefix (or blocky `ecs.forward`) if measurements show a
benefit; note many CN resolvers ignore ECS
([AGH ECS](https://adguard-dns.io/kb/adguard-home/configuration/)).
If the DNS engineering budget is one afternoon, do #1 + #3. If the goal is
"real DNSSEC or nothing", do #1 + #2 + #3. Replacing the stack is only
justified if you want to abandon AGH's UI/filtering entirely — nothing evaluated
here beats it on that axis for this LAN.
---
## Caveats / not verified
- **Live behavior not tested**: all capability claims are from primary docs
reviewed 2026-08-12; DNSSEC behavior of `dns.alidns.com`/`doh.pub`/the
`adg.chans.xyz` path and mosdns's actual version on `.36` need on-box
`dig +dnssec` verification (per [adguard-home-health](../../runbooks/adguard-home-health.md)).
- **smartdns DNSSEC**: the official config reference lists no DNSSEC option;
if a newer version added one, it is not reflected here
([config options](https://pymumu.github.io/smartdns/configuration/)).
- **mosdns geosite/geoip**: v5 source tree (fetched 2026-08-12) contains only
`domain_set`/`ip_set` data providers; if a `geosite.dat` plugin exists in a
release branch, it is not in `main`
([plugin/data_provider](https://github.com/IrineSistiana/mosdns/tree/main/plugin/data_provider)).
- **mihomo**: the GitHub repo currently shows deliberately obfuscated metadata
(see §3.8); capabilities cited from
[wiki.metacubex.one](https://wiki.metacubex.one/en/config/dns/).
sing-box/knot-resolver/dnsdist license identifiers via GitHub metadata are
"other"/custom — treat the specific SPDX ids with caution.
- **chinadns2** README was not retrieved (404 on the raw URL); treated as the
legacy predecessor of chinadns-ng and not evaluated in depth.
- Obsidian/personal notes were not consulted; this is upstream-docs-only.
## Related docs
- [lan-overview.md](../lan-overview.md) — full topology (verified 2026-08-06)
- [hosts/dns.windy.lan.md](../../hosts/dns.windy.lan.md) — AGH host facts
- [hosts/gfw.windy.lan.md](../../hosts/gfw.windy.lan.md) — OpenClash facts
- [adguard-home-official-review-2026-08.md](../adguard-home-official-review-2026-08.md) — prior AGH config review
- [runbooks/adguard-home-health.md](../../runbooks/adguard-home-health.md)
@@ -2,7 +2,7 @@
**状态:** 规划文档(未采购、未接线、未改生产配置)。
**重要变更(2026-08-09):** **SE5420 已采购**,网络升级改为「保留 ER-X + SE5420 核心」路径——
实施与验证以 [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) 为准。
实施与验证以 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 为准。
本文保留为「ER-X 网关未来替换为 RB5009」的备选方案;其中 PVE 透传调研与 VLAN10 实现方法仍适用。
---
@@ -389,9 +389,9 @@ logread -e netifd
## 8. 参考
- 现网地图:[lan-overview.md](lan-overview.md)
- ER-X 现状:[edgerouter-x-configuration.md](edgerouter-x-configuration.md)、[hosts/gw.md](../hosts/gw.md)
- UniFi[unifi-network.md](unifi-network.md)、[hosts/ubnt.md](../hosts/ubnt.md)
- `gfw`[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md)
- 作废方案(**不再实施**):[lan-erx-se5420-network.md](lan-erx-se5420-network.md)、[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
- 现网地图:[lan-overview.md](../lan-overview.md)
- ER-X 现状:[edgerouter-x-configuration.md](../edgerouter-x-configuration.md)、[hosts/gw.md](../../hosts/gw.md)
- UniFi[unifi-network.md](../unifi-network.md)、[hosts/ubnt.md](../../hosts/ubnt.md)
- `gfw`[hosts/gfw.windy.lan.md](../../hosts/gfw.windy.lan.md)
- 作废方案(**不再实施**):[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)、[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
- MikroTik RB5009 官方:<https://mikrotik.com/product/rb5009ug_s_in>、RouterOS v7 手册
@@ -0,0 +1,74 @@
# SE5420 实施评审主张核实(2026-08-10
> **核对基准(历史快照):** 本文于 2026-08-10 针对 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 的**评审前版本**`35577d0`)撰写。该指南自 `ffb37a9`"finalize SE5420 deployment guide per review")起已按本评审修订,当前 `origin/main` 章节已重组:旧 §3.3 → §4.3、旧 §6(gfw)→ §11、旧 §7(SSID)→ §12、旧 §9(验收/IPv6)→ §13 + §11.4。文末「当前指南处理情况」列出各主张的现行状态;实施以部署指南现行为准。
**范围。** 本文核对对 `lan-se5420-deployment-guide.md` 的评审意见。结论分为
“已证实”(规范/一手资料直接支持)、“基本证实”(架构推论成立但仍须读取现场配置)和
“需现场核实”(不能仅由文档或产品手册断言)。这不是实施变更,也不替代维护窗前的
`uci show firewall`、交换机当前 VLAN 表和 PVE bridge 配置检查。
## 核实结论
| 评审主张 | 结论 | 依据与限定 |
| --- | --- | --- |
| `firewall.ubunt_upg.masq=1` 是错误方向,应在实际出站的 `wan` zone 做 IPv4 NAT | **已证实** | OpenWrt 明确规定 masquerade 是**按出站 zone/interface**控制;`masq` 通常在 `wan`。因此,对 `ubunt_upg → wan` 流量把 `masq` 放在源 zone 不是该需求的正确 zone 语义。若 `wan` 已 masq,不应重复开启;也可用 `masq_src` 只限 `192.168.10.0/24`。见 [OpenWrt firewall configuration](https://openwrt.org/docs/guide-user/firewall/firewall_configuration) 和 [fw4 masq 测试](https://lxr.openwrt.org/source/firewall4/tests/02_zones/02_masq)。 |
| `ubunt_upg → wan` 允许所有经 gfw `wan` 可路由的目的地,不等于只上互联网 | **已证实** | `forwarding``src`/`dest` 是 zone-to-zone 单向许可,未按“Internet”语义区分目标 IP;规则可用 `dest_ip` 限制。故若 gfw 的 `wan` 接在 LAN66 且 ER-X 可路由 LAN55,评审所列 LAN66/LAN55 风险成立。最终可达网段仍须以 gfw 路由表、ER-X 路由/防火墙现场检查为准。见 [OpenWrt forwarding/rule 参考](https://openwrt.org/docs/guide-user/firewall/firewall_configuration)。 |
| 不应把 `ubunt_upg.forward` 改为 `ACCEPT``forward_policy` 不是必要的标准 zone 选项;匿名 `uci add` 不可重复执行 | **基本证实** | OpenWrt zone 的标准项是 `forward`forwarding 是独立 section,参考页未定义 `forward_policy`。单独的具名 forwarding 足以允许跨 zone 路径,因此保持 zone 内 `forward=REJECT` 是较小权限配置。匿名 section 每执行一次都会新增一节,这是 UCI 的操作语义;实施应先读现场配置并使用具名 section。 |
| VLAN10 必须有显式 IPv6 策略,否则可能绕过仅 IPv4 的 NAT/隔离 | **已证实** | fw4 将 `masq`IPv4)和 `masq6`IPv6)分开;forwarding 默认 family 是 `any`。仅写 IPv4 DHCP/NAT/地址规则不能表达 VLAN10 的 IPv6 RA、DHCPv6、路由和过滤策略。是否已经存在可用 IPv6 前缀、以及 OpenClash 是否接管 IPv6,必须现场验证。见 [OpenWrt firewall configuration](https://openwrt.org/docs/guide-user/firewall/firewall_configuration)。 |
| 管理 SVI + 默认路由与“不开 SVI/静态路由/一切 L3”矛盾 | **已证实** | 指南(历史版)§3.3 同时要求 VLAN66 `192.168.66.253/24` 与默认路由,又要求不开 SVI/静态路由。TL-SE5420 官方称其为三层交换机,支持静态路由、RIP、DHCP server/relay。准确目标应是:仅保留 VLAN66 管理 L3 interface/默认网关,不给 VLAN55/10 建 L3 interface,且禁用不需要的 L3 服务和跨 VLAN routing。见 [TL-SE5420 官方页](https://www.tp-link.com.cn/product_2899.html?v=specification) 与 [官方安装手册](https://service.tp-link.com.cn/download/202310/TL-SE5420%20V1.0%E5%AE%89%E8%A3%85%E6%89%8B%E5%86%8C%201.0.2.pdf)。 |
| 必须明确移除 VLAN1 成员,PVID 变更本身不等于 access-port VLAN membership | **已证实** | PVID/native VLAN 只处理进入端口的未标记帧;access/trunk 的允许 VLAN 列表是独立概念。指南(历史版)§3.3 只列 VLAN66/55 member 和 PVID,未写移除 VLAN1 或 ingress filtering。验收应检查 VLAN1 member、VLAN1 管理 IP、端口允许 VLAN 和 tagged-frame ingress policy。见 [Ubiquiti 对 native/tagged/access/trunk 的定义](https://help.ui.com/hc/en-us/articles/26136855808919-Switch-Port-VLAN-Assignment-Trunk-Access-Ports)(术语与 802.1Q 语义)以及 [Linux bridge VLAN 配置示例](https://www.kernel.org/doc/html/v5.19/networking/dsa/b53.html)(显式 `bridge vlan del ... vid 1`)。TL-SE5420 具体 GUI/CLI 行为仍以其固件手册核验。 |
| NAS 不应在未先完成双端 LACP 时同时接两口;LACP 不使单 TCP 流自动达到 5G | **基本证实** | 这是标准二层环路/聚合变更控制结论:没有已协商的 LAG 时,两条同 VLAN 并行链路会构成潜在环路;STP 只能作为保护而非实施方法。官方产品页列出 LACP 相关资料,但本次未取得 TL-SE5420/TrueNAS 对端的精确配置与当前 NAS 连接状态,故“必然环路/双 IP”不能在桌面审阅中断言。单连接吞吐受链路散列限制是 802.3ad 的常见实现特性,应以 NAS 与交换机的 hash policy 和 `iperf3` 实测验收。 |
| 非 VLAN-aware 的 PVE `vmbr0` 不提供 VLAN10 的端口级隔离 | **已证实** | PVE 将 bridge 描述为虚拟交换机;VLAN-aware mode 才能给 guest NIC 赋 VLAN tag,或显式 trunk。Linux 内核说明:`vlan_filtering=0` 时 bridge 不考虑 VLAN tag,且默认关闭;开启后才按 MAC **和 VLAN tag**转发及进行严格 VID 检查。因此“共享非 VLAN-aware bridge 可让可控 guest 主动消费 VLAN10,不能作为严格隔离边界”成立。不能仅凭该结论断言每个 guest 必定收到每个单播帧:未知单播/广播会泛洪,已学习的单播会按 FDB 转发。见 [PVE 网络配置](https://pve.proxmox.com/wiki/Network_Configuration) 和 [Linux bridge 文档](https://docs.kernel.org/networking/switchdev.html)。 |
| AP VLAN10 tagged frame 在一个普通 untagged LAN66 access path 上会“自动去 tag 并泄漏到 LAN66” | **不成立/需改写** | 802.1Q 的 native VLAN 是对**未标记**流量的 VLANtagged VLAN 需被显式允许于 trunk。因此通常的正确表述是:若上游不允许 VLAN10 tagAP 到 VLAN10 网关/DHCP 的路径不存在,SSID 会成为不可用入口。实际设备的端口模式(包括是否错误地配置为 all/trunk、是否接受 tagged ingress)须现场查看,不能泛称必然去标签。见 [Ubiquiti VLAN 端口定义](https://help.ui.com/hc/en-us/articles/26136855808919-Switch-Port-VLAN-Assignment-Trunk-Access-Ports) 和 [Ubiquiti VLAN troubleshooting](https://help.ui.com/hc/en-us/articles/9592924981911-Virtual-Network-VLAN-Troubleshooting)。 |
| 仅保留 SSH 会话不是移动 PVE/ER-X 物理上联时的真正带外回滚路径;应全程保持 SE5420 Console | **已证实** | 这是直接的操作依赖判断:TCP SSH 的承载链路被拔除时会断,不能证明回滚可达。TL-SE5420 官方安装手册确认该机有 Type-C Console,且本仓库指南本身也把恢复出厂流程建立在 Console 上。故应在迁移前接通 Console、标注旧/新端口、逐根迁移并用 MAC 表与链路/错误计数验证。见 [官方安装手册](https://service.tp-link.com.cn/download/202310/TL-SE5420%20V1.0%E5%AE%89%E8%A3%85%E6%89%8B%E5%86%8C%201.0.2.pdf)。 |
| “所有设备均不得直连 ER-X”不是避免环路的必要条件 | **已证实** | 环路取决于同一 L2 广播域存在多条并行二层路径,不取决于是否还有一个独立终端直接接 ER-X。应禁止的是一个下级交换机/桥接主机同时形成平行路径。此项仍需以 ER-X switch0 VLAN/bridge 现场配置和实际接线图确认。 |
| 性能不应承诺全面 2.5G;同 VLAN 才可能在 SE5420 本地交换超 1G,跨 55/66 与 Internet 受 ER-X/宽带限制 | **已证实** | TL-SE5420 的 2.5G 端口仅提高经其本地二层转发的链路上限;跨子网必须由网关路由,Internet 另受 WAN/PPPoE 约束。产品页确认 16×2.5G + 4×10G SFP+,但 ER-X、NAS、PC、AP 的实际协商速率和 NIC/布线能力必须由 `ethtool`/端口状态及 `iperf3` 验证。见 [TL-SE5420 官方规格](https://www.tp-link.com.cn/product_2899.html?v=specification)。 |
## 已核对的文档内事实
现行指南的**历史版本**`35577d0`)确实包含评审指出的关键文字:旧 §3.3 的管理 IP/默认路由与“不开 SVI/静态路由”;旧 §6 的 `ubunt_upg.masq``forward=ACCEPT``forward_policy`、匿名 forwarding;旧 §7 只验“不可达 LAN66”;旧 §9 对新增 VLAN10 写“IPv6 行为与升级前一致”。因此上述评审不是对未出现内容的假设。
但这些内容在 `ffb37a9` 起的修订中已被修正或重组:`ubunt_upg.masq``forward_policy` 已删除,gfw 防火墙改为 §11(§11.3 第 2 步明确“不要给 `ubunt_upg` zone 加 masq”);VLAN10 IPv6 在 §11.4 显式写为“本阶段不提供”;SVI/L3 边界在 §4.3 第 16 步单列“L3 明确边界检查”。本文按历史快照保留评审结论,读者应以现行部署指南为准。
本仓库的 `hosts/gfw.windy.lan.md` 还记录 gfw 的 `eth0` 在 LAN66、`eth1` 在 LAN55,故 `ubunt_upg → wan` 的隔离结论应在执行前以当前 `ip route``uci show firewall``nft list ruleset` 复核,而不能从方案文字直接把规则写死。
## gfw 现场只读复核(2026-08-10
已通过 `ssh -4 root@192.168.66.1` 仅读取配置和运行规则,未修改设备。该结果会改变
评审中两项“当前状态”的表述:
| 现场事实 | 对评审的影响 |
| --- | --- |
| `wan` zone 已有 `masq='1'`;现有配置另有具名 `ubunt_upg_nat`,运行时渲染为 `oifname "eth0"` 且只匹配 `ip saddr 192.168.10.0/24 masquerade`。 | “必须在 wan 开 masq”的**方向原则**正确,但“当前无 masq”不正确。现有显式 SNAT 已在实际出 `eth0` 时执行;计划中再将 `masq` 加到 `ubunt_upg` 仍是多余且方向错误。 |
| 当前放行是具名 `ubunt_upg_to_lan`,不是 `ubunt_upg→wan`;其运行链先拒绝 `192.168.66.0/24`,再允许到 `lan`。gfw 的 IPv4 default route 是 `192.168.66.254`。 | 计划新增 `ubunt_upg→wan` 会是与当前设计不同、过宽的改动。现有 LAN66 阻断规则在该链中先匹配;但对经 ER-X 可达的 LAN55/其他内网仍没有显式拒绝,故隔离评审的**剩余风险成立**。应以明确内网前缀 deny + 所需外网 allow 重写,而不是加 WAN forwarding。 |
| `ubunt_upg` DHCPv6 和 RA 都是 `disabled`;运行路由表仅有各接口的 IPv6 link-local route,没有 IPv6 default route;全局 IPv6 forwarding 是 `1`。 | 评审“VLAN10 未明确 IPv6 策略”的表述对计划文本仍成立,但“IPv6 可能立即绕过”的事实判断在当前状态**未获证实**:现有 RA/DHCPv6 已关闭且无 IPv6 默认路由。实施文档仍应把这项显式写为“IPv6 不提供”,并在启用前复查。 |
| 系统是 ImmortalWrt **25.12.0**`/usr/bin/apk` 存在(apk-tools 3.0.5)。 | 评审中“ImmortalWrt 21.02.5 应使用 opkg”的版本判断错误/过时;在本机上 `apk add tcpdump` 是可用包管理器。仍应先检查软件包可用性,避免在维护文档中把两种命令并列为未经验证的替代方案。 |
这些命令输出未含凭据、令牌或私钥,故仅记录了安全相关的摘要;不将完整防火墙快照提交至仓库。
## 当前指南处理情况(2026-08-13 核对)
`origin/main``2fd354c`)逐项核对评审主张:
| 评审主张 | 现行状态 | 现行位置 |
| --- | --- | --- |
| `ubunt_upg.masq=1` 方向错误 | ✅ 已修复 | §11.3 第 2 步「不要给 `ubunt_upg` zone 加 masq」 |
| `ubunt_upg→wan` 不等于只上互联网 | ⚠️ 原则成立,指南已禁止新增宽泛 forwardingLAN55/RFC1918 显式 deny 仍为待办 | §11.1b「待补缺口」、§11.3 |
| 勿改 `forward=ACCEPT`;无 `forward_policy`;匿名 uci 不可重复 | ✅ 已修复 | §11.3 第 1 步保持 REJECT;指南已无 `forward_policy` |
| VLAN10 须显式 IPv6 策略 | ✅ 已修复 | §11.4「本阶段不提供 VLAN10 IPv6」 |
| 管理 SVI + 默认路由 vs「不开一切 L3」矛盾 | ✅ 已消解 | §4.3 第 16 步「L3 明确边界检查」 |
| 须移除 VLAN1 成员;PVID≠membership | ✅ 指南已加强;现网仍偏离(W1N-54) | §4.3 第 912 步;VLAN1 不可删说明 |
| NAS 双口未 LACP 前勿并行 | ✅ 已体现 | §7 第 7 步「仅口 8,口 12 断开」 |
| 非 VLAN-aware PVE bridge 不能作隔离边界 | ✅ 已体现 | §9.2 要求 VLAN-aware + `bridge-vids` |
| AP tagged 帧在 access 口自动去 tag | ✅ 本文已纠正(不成立) | — |
| SSH 非真正带外;须 Console | ✅ 已体现 | 开头第 2 条、§4.1、§16 |
| 「所有设备不得直连 ER-X」非必要 | ✅ 已体现 | 全程三条第 1 条、§7 |
| 不应承诺全面 2.5G | ✅ 已体现 | §8 第 7 条、§14 |
## 实施前的最低限度现场证据
1. gfw:保存并审阅 `uci show firewall``ip route``ip -6 route``nft list ruleset`;确认 wan 的 masq 与所有 WAN→内网、VLAN10→内网匹配次序。
2. SE5420 Console:导出/截图 VLAN1、55、66 member 和 PVID/ingress-filter 状态;确认唯一管理 L3 interface 和路由/relay/DHCP 状态。
3. PVE:记录 `/etc/network/interfaces`、VM NIC VLAN tags 和 `bridge vlan show`,再决定是否把 VLAN-aware 改造另开窗口。
4. AP:从实际设备 `info` 或控制器记录确认 Inform URL(本仓库目前记录 `http://192.168.66.46:9080/inform`),并验证 VLAN10 tag 只经 U6/PVE trunk。
5. NAS:单网口稳定后,另窗配置并验证两端 LACP,第二根线最后插入;用多流及单流 `iperf3` 分开验收。
+345
View File
@@ -0,0 +1,345 @@
# Home Assistant × Matrix integration
Reference for wiring the Home Assistant [Matrix integration](https://www.home-assistant.io/integrations/matrix)
to the self-hosted Matrix homeserver at [`synapse.chans.xyz`](../hosts/synapse.chans.xyz.md).
Deliberately contains no Matrix passwords, access tokens, or room encryption material.
> **Status (2026-08-15, W1N-139):** the built-in `matrix` integration has been
> **retired** on `hass.windy.lan` and replaced by the custom **`matrix_e2ee`**
> integration. The sections below on the built-in integration are kept for
> reference only. See [matrix_e2ee](#matrix-e2ee-custom-e2e-integration) for the
> active setup and [Device verification (SAS) model](#device-verification-sas-model)
> for how device trust works.
## Purpose
The integration lets Home Assistant send messages to Matrix rooms and react to
messages/reactions in Matrix rooms. "Reacting" is done by firing a
`matrix_command` event when one of the configured commands matches; automations
then trigger on that event. Sending is done through the `notify.matrix` platform
and the `matrix.send_message` / `matrix.react` actions.
## Environment mapping
| Integration setting | This deployment |
|---|---|
| `homeserver` | `https://synapse.chans.xyz` (client-server base URL) |
| `username` | full Matrix ID, e.g. `@ha_bot:chans.xyz` |
| `password` | MAS local-password account password (see below) |
| Room IDs / aliases | full forms with the identity domain, e.g. `!cUrbafjkfsMDVwdRDQ:chans.xyz` or `#room:chans.xyz` |
- Identity domain is `chans.xyz` (not `synapse.chans.xyz`); user IDs and room
aliases carry the `:chans.xyz` suffix.
- Authentication on this homeserver is MAS (Matrix Authentication Service) with
local-password accounts. The integration logs in with `m.login.password`
(username + password), so the bot account must be a local-password account —
same as the Hermes account documented in [`hermes-matrix.md`](hermes-matrix.md).
If MAS is later switched to OAuth2/OIDC-only (no legacy password login), the
integration's password login will stop working; keep that in mind before such
a change.
- Public registration is disabled. Create/reset the dedicated bot account via
MAS / Element Admin.
## Use a separate bot account (mandatory)
The docs are explicit: to prevent infinite loops when reacting to commands,
the integration **must** use a separate account from any account whose messages
it reacts to. Use a dedicated account such as `@ha_bot:chans.xyz`, not a human
account.
## configuration.yaml (example)
```yaml
# The Matrix integration
matrix:
homeserver: https://synapse.chans.xyz
username: "@ha_bot:chans.xyz"
password: supersecurepassword
rooms:
- "#hasstest:chans.xyz"
commands:
- word: my_command
name: my_command
```
After changing `configuration.yaml`, restart Home Assistant to apply the
changes. The integration then shows under **Settings → Devices & services**;
its entities are on the integration card and the Entities tab.
### Configuration variables
| Variable | Meaning |
|---|---|
| `username` | Full Matrix ID the bot logs in as, e.g. `@ha_bot:chans.xyz`. The `@` has a special YAML meaning, so always quote it. |
| `password` | The bot account's password (MAS local password). |
| `homeserver` | Full client-server URL of the homeserver. |
| `rooms` | Rooms the bot should join and listen in. List **all** rooms commands are to be received in, even if a command scopes itself to fewer rooms. Accepts internal room ID (`!…:chans.xyz`) or alias (`#room:chans.xyz`). |
| `commands` | Commands to listen for. Each fires a `matrix_command` event when triggered. |
### Command types
| Key | Triggers when |
|---|---|
| `word` | A message starts with `!<word>`. Arguments after the word are captured as a list in the event's `data`. |
| `expression` | A message matches the Python regexp. The regexp group dictionary is captured in the event's `data`. |
| `reaction` | A message is reacted to with the given emoji. |
| `name` | The command name, exposed as an attribute of the fired event. |
A command can be scoped to specific rooms with a per-command `rooms` list (the
room must still be listed under the top-level `rooms`).
## Event data
When a command triggers, a `matrix_command` event fires with:
- `name` — the command name.
- `data` — for `word` commands, a list of arguments (everything after the word,
split on spaces); for `expression` commands, the group dictionary of the
matching regexp.
- `event_id` — the received message's identifier.
- `thread_parent` — the root message ID of the thread; equals `event_id` when
the message is not inside a thread.
## Notifications (notify.matrix)
Deliver notifications from Home Assistant to a Matrix room (direct or group):
```yaml
notify:
- name: matrix_notify
platform: matrix
default_room: "#hasstest:chans.xyz"
```
- The target room must already exist; get its canonical ID from the room
settings dialog (`!<randomid>:chans.xyz`) or an alias (`#roomname:chans.xyz`).
Quote the room ID/alias in YAML to escape the `!` / `#` characters.
- The notifying account may need to be invited to the room, depending on room
policy.
Message formats (`data.format`): `text` (default) and `html`. Images can be
attached via `data.images` (list of file paths); files from outside allowed
folders require `homeassistant.allowlist_external_dirs` to list the source
folder.
Reply inside a thread by passing the root message ID into `data.thread_id`:
```yaml
action: notify.matrix_notify
data:
message: "Reply message goes here"
data:
thread_id: "{{ trigger.event.data.thread_parent }}"
```
## Actions
- `matrix.react` — send a reaction to a message in a Matrix room
(`reaction`, `room`, `message_id`).
- `matrix.send_message` — send a message to one or more Matrix rooms.
## Comprehensive example (adapted)
```yaml
matrix:
homeserver: https://synapse.chans.xyz
username: "@ha_bot:chans.xyz"
password: supersecurepassword
rooms:
- "#hasstest:chans.xyz"
- "#someothertest:chans.xyz"
commands:
- word: testword
name: testword
rooms:
- "#someothertest:chans.xyz"
- expression: "My name is (?P<name>.*)"
name: introduction
- reaction: 👍
name: thumbsup
notify:
- name: matrix_notify
platform: matrix
default_room: "#hasstest:chans.xyz"
automation:
- alias: "Respond to !testword"
triggers:
- trigger: event
event_type: matrix_command
event_data:
command: testword
actions:
- action: notify.matrix_notify
data:
message: "It looks like you wrote !testword"
```
## matrix_e2ee (custom E2E integration)
Custom integration [`windyboy/ha-matrix-e2ee`](https://github.com/windyboy/ha-matrix-e2ee),
release **v0.3.12** (Matrix activity events + push diagnostics), deployed on
`hass.windy.lan` 2026-08-20 (upgraded from v0.3.2, W1N-182/#34 emoji-wait
wizard fix; v0.3.9 brought the Connection health binary sensor, SAS/command
allowlist split, URL normalization and single-entry enforcement, W1N-156/W1N-190).
Runs a dedicated bot with a **persistent E2EE device identity**.
- Domain `matrix_e2ee`; Config Flow (UI) with YAML import migration, not in HACS. Does **not**
override the built-in `matrix` integration.
- Dependencies are declared **explicitly** in `manifest.json` to work around Home
Assistant's `is_installed` dropping the `[e2e]` extra (W1N-140):
`matrix-nio[e2e]==0.26.0` + `vodozemac` + `peewee` + `cachetools` + `atomicwrites`.
- **v0.2.0 migration:** YAML `matrix_e2ee:` block was auto-imported into a Config Entry
(`source: import`) on first startup, then removed. All settings now managed via
**Settings → Devices & Services → Matrix E2EE → Configure**.
See [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) for the deployed state.
### Services & events
- Services (all admin-only since v0.1.4):
- `send_message` (`message`, `room_id`)
- `start_verification` (`user_id`, `device_id`)
- `confirm_verification` (`transaction_id`)
- `cancel_verification` (`transaction_id`)
- `reauthenticate` (`password`) — soft-logout only
- `get_fingerprint` (no fields; returns bot's own `ed25519`/`curve25519` keys; added v0.1.3)
- `verify_device_by_fingerprint` (`user_id`, `device_id`, `ed25519`; added v0.1.3,
renamed from `verify_device` in v0.1.4; requires exact `ed25519` match)
- Events:
- `matrix_e2ee_command` (`room_id`, `sender`, `command`, `args` only —
never the raw body)
- `matrix_e2ee_error` (codes, no secrets)
- `matrix_e2ee_verification` (`stage`, `transaction_id`, `user_id`, `device_id`,
optional `emojis`, optional `expires_at`; `expires_at` added v0.1.3)
- `matrix_e2ee_fingerprint` (`user_id`, `device_id`, `ed25519`, `curve25519`
public keys only; added v0.1.3)
- `matrix_e2ee_message_received` (`room_id`, `sender`, `event_id`; added v0.3.12
activity events)
- `matrix_e2ee_verification_done` (`transaction_id`, `user_id`, `device_id`;
added v0.3.12)
- v0.3.12 also adds an `event.` platform entity (`Bot activity`,
`event_types: ["message", "command", "verification_done"]`) and a diagnostic
Connection binary sensor (`binary_sensor.*_connection`, CONNECTIVITY class).
- `notify.matrix_e2ee` is **not implemented** (upstream deferred) — notifications
must call `matrix_e2ee.send_message` (message + room_id).
- Commands fire Home Assistant events only; the integration never calls
`domain.service` itself. Map commands in automations.
- Encrypted rooms fail-closed on unverified devices.
- Since v0.1.4: `start_verification`, `confirm_verification`, `cancel_verification`,
`verify_device_by_fingerprint`, and `reauthenticate` are enforced as HA admin-only
via `async_register_admin_service`; non-admin users cannot call them.
### Storage & recovery
- `.storage/matrix_e2ee_session.json` (`user_id`, `device_id`, `access_token`,
`pickle_key`) and `.storage/matrix_e2ee_store/` (Olm/Megolm, device trust,
sync token). Both stay on the HA persistent volume and are in HA backups.
- Soft logout → `matrix_e2ee.reauthenticate` (keeps `device_id` + crypto store;
rejected outside soft-logout state since v0.1.3).
- Hard logout / store loss → delete session + store, restart with password, re-SAS
(a **new device**; old history not decryptable).
## Device verification (SAS + fingerprint) model
Researched 2026-08-15 (W1N-139 stage-6 pre-study), updated for v0.1.3/v0.1.4.
Sources: matrix.org
[cross-signing guide](https://matrix.org/docs/guides/implementing-more-advanced-e-2-ee-features-such-as-cross-signing/),
matrix-nio [examples](https://matrix-nio.readthedocs.io/en/latest/examples.html),
[element-android#6832](https://github.com/vector-im/element-android/issues/6832),
Element [device-verification](https://element.io/features/device-verification).
`matrix_e2ee` supports three verification paths (the wizard — v0.3.0
bot-initiated, reworked in v0.3.1/v0.3.2 to wait for a peer-initiated inbound
SAS from the user's Matrix client with emoji comparison — automates the SAS
flow):
### 1. SAS (mutual, manual confirmation since v0.1.4)
- SAS is device-to-device: exchange ephemeral keys → derive emojis → **a human on
each side compares and confirms** (`m.key.verification.mac`).
- Matrix distinguishes two cases (spec uses *should*, not *must*):
- **same user, two devices** → to-device messages (SAS);
- **two different users** → **in-room (DM) messages**, verifying the *user*
(cross-signing master key), not a specific device.
- Cross-signing: each user has master / self-signing / user-signing keys. A device
looks "verified" to another user via the chain
`my master → my user-signing → their master → their self-signing → their device`.
- Element's "Verify" button only starts **in-DM user verification**; it has no
"verify a specific device of another user via to-device" flow (matrix.org
recommends hiding per-device verification for other users).
- `matrix_e2ee` implements **raw to-device device SAS** (`start_verification`/
`confirm_verification`), **no cross-signing / in-room**. This is a non-standard
cross-user path: works with matrix-nio + Element Web/Desktop (reported in
element-android#6832), **not** on Element Android/X.
- **v0.1.3**: inbound SAS auto-complete was added; SAS events include `expires_at`.
- **v0.1.4 (breaking)**: auto-confirm was removed. **Every** device — including
another device of the bot's own account — requires explicit `confirm_verification`
after emoji comparison. Only the bot's own account or users in `allowed_users`
may initiate SAS (`verification_peer_denied` otherwise).
- **v0.2.1**: storage I/O moved off the event loop (`asyncio.to_thread`,
W1N-167); own-keys query on startup so inbound SAS can build a session (W1N-166).
- **v0.2.2** (not deployed): intermediate version.
- **v0.2.3**: sync loop runs as a background task (fixes bootstrap setup timeout,
W1N-168); SAS double-send of key and MAC fixed (W1N-169).
- **v0.2.6**: `_log_verification_state()` tracks SAS state transitions with
`async_write_ha_state` for diagnosis (W1N-174);
`_bridge_verification_request()` handles inbound
`m.key.verification.request``m.key.verification.ready` since nio lacks a
`request` framework (W1N-173).
- **v0.2.5**: bridge `m.key.verification.request``ready` (nio lacks
request framework, W1N-173).
- **v0.2.4**: `_patch_nio_sas_timeout()` works around nio 0.26.0
`_last_event_time` bug (SAS timed out at 60s regardless of activity — now uses
`_max_age` 5 min); `_repair_dropped_start()` recovers SAS `start` events nio
dropped when the peer device was unknown (W1N-170/W1N-172);
`VERIFICATION_TIMEOUT_SECONDS` 600→240 (fires before nio's `_max_age`).
- **v0.2.11**: `receive_mac_event` no longer overrides canceled state (W1N-179/#31).
- **v0.3.12**: Matrix activity events (`matrix_e2ee_message_received`,
`matrix_e2ee_verification_done`) + `event.` Bot activity entity + Connection
diagnostic binary sensor.
- **v0.3.9**: SAS driver gate split from the command allowlist — new
`verification_peer_users` option (W1N-156/#41); SAS/sync logs demoted
warning→info/debug (W1N-188/#38); Connection health binary sensor
(W1N-185/#40); URL normalization + single-entry enforcement (W1N-190/#42).
- **v0.3.8**: `m.key.verification.done` handshake completion for
request-based SAS (W1N-183/#35).
- **v0.3.2**: wizard waits for the inbound SAS to show emojis before moving
to the compare step (`_wait_for_inbound` requires `latest_sas_snapshot()` to
return `emojis`) — W1N-182/#34.
- **v0.3.1**: verification wizard now waits for a peer-initiated inbound SAS
(options flow no longer starts verification from the bot; `latest_sas_snapshot()`
skips verified/canceled transactions) — GitHub #33.
- **v0.3.0**: bot-initiated device verification wizard (W1N-180/#32).
- Inbound SAS is gated to `allowed_users` (v0.1.3); **since v0.3.9 (W1N-156)
the gate is the separate `verification_peer_users` allowlist**, which is
unset on hass.windy.lan — only the bot's own account may drive SAS until
`@zhiqiang:chans.xyz` is added there.
### 2. One-sided fingerprint (added v0.1.3, hardened v0.1.4)
- Call `matrix_e2ee.get_fingerprint` to get the bot's own `ed25519` device key
(read it from the `matrix_e2ee_fingerprint` event).
- In Element, open the bot user's sessions and use "Manually verify by text".
Compare the session key with the fingerprint.
- To trust another device from the bot's side, call
`matrix_e2ee.verify_device_by_fingerprint` with the peer's `user_id`, `device_id`,
and `ed25519` key. The match is exact (since v0.1.4's rename from `verify_device`).
Feed the **peer** key, not the bot's own key.
- This trusts from one side only; the peer still trusts the bot independently.
- Both `get_fingerprint` and `verify_device_by_fingerprint` are HA admin-only.
### Consequence
Whether `@zhiqiang`'s device can be verified depends on which Element client
they use. Open options recorded in W1N-139 (A: Web SAS test; B: upstream
in-room/cross-signing; C: unencrypted-room downgrade).
## References
- Home Assistant Matrix integration: <https://www.home-assistant.io/integrations/matrix>
- Matrix host facts: [`hosts/synapse.chans.xyz.md`](../hosts/synapse.chans.xyz.md)
- Matrix deployment and upstream index: [`matrix-upstream.md`](matrix-upstream.md)
- Hermes Agent Matrix channel (MAS local-password + access-token pattern): [`hermes-matrix.md`](hermes-matrix.md)
- HA host facts: [`hosts/hass.windy.lan.md`](../hosts/hass.windy.lan.md)
- HA maintenance runbook: [`runbooks/home-assistant-maintenance.md`](../runbooks/home-assistant-maintenance.md)
+246
View File
@@ -0,0 +1,246 @@
# 内网 DNS 架构调研与优化建议
> 状态:2026-08-12 调研定稿,Linear **W1N-56**。**Phase 0 验证(2026-08-12)全部完成,最终裁决(用户定稿)已对齐**;本期未改动任何生产 DNS 路径。
> 2026-08-12 现场核查修正:gfw 上 mosdns **不是闲置**——它是 OpenClash clash 的
> `nameserver`/`default-nameserver`(DIRECT 规则真实 IP 解析),处于活动链路,裁决第 5 条
> 的"删除闲置 mosdns"前提不成立,处置改为"正式纳管并文档化"。
> 相关:`docs/lan-overview.md`、`hosts/dns.windy.lan.md`、`hosts/gfw.windy.lan.md`、W1N-40。
> 2026-08-13 (W1N-62):gfw mosdns 国外分支已从"明文国内公网 DNS"改为**加密 DoH**
> (自建 `https://adg.chans.xyz/dns-query`,hk2),新增 `foreign_upstream`/`foreign_fallback`
> (primary=DoH, secondary=明文国内 DNS, threshold 1000ms),`bootstrap` 用现有国内公网 IP
> (防自举循环)。实测:mosdns 平面国外域名 A/AAAA 恢复(google AAAA `2607:f8b0…`)、
> `dup.baidustatic.com`→`0.0.0.0`(AGH 拦截保留)、clash 7874 fake-ip 平面不变。
> **重要修正(定稿)**:DoH 流量实测为 **gfw→hk2 直连,未经 clash 代理**——nft output 链
> (mangle mark/tcp redirect)计数为 0、`/proc/net/tcp` 存在到 hk2:443 的 established
> 连接,路由自身 TCP 输出当前并未被 OpenClash 重定向,故"经代理访问加密 DNS"的假设不成立。
> **最终决策:接受直连,不强行走代理**——`foreign_upstream` 以自建解析器
> `adg.chans.xyz`(hk2)为主力,自有 VPS 直连即可达、无被墙/污染问题、DoH/TLS 已加密、应答干净,
> 走代理毫无增益反而把 DNS 平面耦合进 clash;实测 kill clash 期间国外查询 0.01s 正常应答,
> watchdog 自动拉起,直连使 DNS 平面独立于代理(优于过代理)。
> **多上游冗余(同日)**:`concurrent: 3`,新增 `https://dns.quad9.net/dns-query` 与
> `https://dns.cloudflare.com/dns-query`(2026-08-13 本网络实测可达;`dns.quad101.net`
> TLS 握手失败已排除)——hk2 故障时仍由干净的国外 DoH 应答,最后才退化明文国内兜底。
> 验证:google AAAA 由 hk2 的 `2607:f8b0…` 变为新上游的 `2404:6800…`,国内外/拦截/代理平面无回归。
> 备份:`config.yaml.bak-foreign-doh-20260813-103746` / `config.yaml.bak-foreign-doh-20260813-103813` /
> `config.yaml.bak-multi-doh-20260813-105421`。
## 1. 现状(实测 2026-08-12)
| 角色 | 部署 | 职责 | 是否在活动路径 |
|------|------|------|----------------|
| **AdGuard Home** | `192.168.66.36`(PVE VM 120,Docker host 网络) | EdgeRouter DHCP 通告给 LAN55/66 客户端的唯一 DNS;广告/过滤、查询统计、Web 面板 | ✅ **是(LAN 客户端唯一入口)** |
| **mosdns** | `192.168.66.1`(gfw OpenWrt)监听 `127.0.0.1:6052`(仅本机) | clash 的 `nameserver`/`default-nameserver`:DIRECT 规则真实 IP 分流(国内 → AGH `.36:53`,国外 → `223.5.5.5`/`119.29.29.29`) | ✅ 网关侧(clash 消费,不面向客户端) |
| **OpenClash / clash(meta)** | `192.168.66.1`(gfw) | gateway 自身/被劫持流量的 fake-ip + 代理;DNS 走 dnsmasq→clash `#7874` | 仅网关侧与 VLAN10 |
**关键事实(全部实测):**
- LAN 客户端 DNS 直连 `66.36`,**不经过** gfw(EdgeRouter `service dns forwarding` cache 512,通告 `.36`)。
- AGH 上游:DoH `dns.alidns.com`(→`223.5.5.5`/`223.6.6.6`)+ `doh.pub`(→`120.53.53.53`/`1.12.12.12`),
`upstream_mode: load_balance``fastest_timeout: 1s``upstream_timeout: 10s`;bootstrap
`223.5.5.5`/`223.6.6.6`(公网 IP,无自举循环);兜底 DoH `adg.chans.xyz`(→`hk2.chans.xyz``154.36.174.161`)。
DNS 监听 UDP/TCP **v4 only(`0.0.0.0:53`)**,无 v6 监听;`enable_dnssec: false`;cache 4 MB;ratelimit 20。
- AGH 宿主出网:默认路由 `via 192.168.66.254`(EdgeRouter),**直连,不经 gfw**;DoH 实测可达
(223.5.5.5:443 → HTTP 400/0.05s,120.53.53.53 → 502/0.06s,adg.chans.xyz 冷连接 ~3.5s)。
- gfw 自身出网:OpenClash `openclash_mangle_output` 对非本地区域/国内 IP 流量统一
`mark 0x162 → tproxy 127.0.0.1:7895`(**gfw 自身流量默认走代理**,含 clash 的
nameserver-policy DoH);mosdns 的上游(AGH `.36` 本地区域、`223.5.5.5` 国内 IP)均被 bypass,保持直连。
- gfw DNS 链:`server=127.0.0.1#7874`(dnsmasq)→ clash:`nameserver: [127.0.0.1:6052]`(mosdns)、
`default-nameserver: [127.0.0.1:6052]``nameserver-policy` 国外域名 → DoH `https://1.1.1.1/dns-query`
`enhanced-mode: fake-ip`(198.18.0.1/16)、`ipv6: false`;nft 有 UDP/53 hijack → dnsmasq。
- **mosdns 配置缺陷(2026-08-12 发现并修复)**:`main` sequence 的国内分支
(`matches: qname $domestic_domains → exec: $domestic_upstream`)之后**缺少
`matches: has_resp → accept` 守卫**。mosdns v5 的 `sequence``forward` 成功后不会停止,
只有 `accept`/`reject`/`return` 或错误会终止——因此命中 `geosite_cn` 的查询会被转发**两次**
(AGH 与 223.5.5.5/119.29.29.29),最终应答来自最后一个 forward(国内公网 DNS),**AGH 的
拦截/rewrite 对 DIRECT 国内域名静默失效**。实测证据:`dup.baidustatic.com`(在 `geosite_cn`
且在 AGH 拦截表)经 mosdns 返回真实 IP `183.60.227.49` 而非 `0.0.0.0`。已修复(备份
`/etc/mosdns/config.yaml.bak-20260812`),修复后同一域名返回 `0.0.0.0`,taobao/google 解析
与 clash 链均无回归。**§6 的二期示例同款缺陷已一并修正。**
同日追加加固:`domestic_fallback`(fallback 插件:`primary: domestic_upstream`(AGH)、
`secondary: default_upstream`(223.5.5.5/119.29.29.29)、`threshold: 500ms`)使 AGH 宕机时
DIRECT 国内真实 IP 查询回退国内公网 DNS,不再直接报错;实测:AGH 停止时缓存未命中查询由
fallback 应答(NXDOMAIN/真实 IP),AGH 恢复后主路径即时应答且拦截(`0.0.0.0`)恢复。
备份:`/etc/mosdns/config.yaml.bak-fallback-20260812`
- AGH rewrites(实测):`hass.windy.lan`/`hass.local``192.168.55.11`;`dns.windy.lan``.36`;
`ubnt.windy.lan``.46`;`gfw.windy.lan``.1`;`nas.windy.local``.32`
- 拦截:仅启用 **AdGuard DNS filter**(filter_1);实测 `doubleclick.net`/`googleadservices.com``0.0.0.0`
**现状缺口(实测确认):**
1. AGH 上游是固定 DoH,**无"国内/国外分流"能力**;国外域名解析质量依赖唯一兜底路径。
2. **兜底失效**(kill-test 证实):主上游黑洞时,兜底 `adg.chans.xyz` 在客户端 15s 窗口内不生效
(`upstream_timeout: 10s` + TCP 重试行为),缓存未命中查询无有界降级——见 §8。
3. 国外域名 AAAA 经国内路径全部置空(见 §8),v6 解析缺位。
## 2. 最终裁决对齐(用户定稿 2026-08-12)
| # | 裁决 | 本issue处理 |
|---|------|------------|
| 1 | **保留 AGH `.36` 为唯一 LAN DNS 入口**(现有方案增强版),不改 EdgeRouter DHCP 通告 | ✅ 现状保持;本期零改动 |
| 2 | **否决"AGH 全局转发到 Clash fake-IP"**——DNS 平面必须与流量转发平面一致 | ✅ 分层方案(§4)明确 AGH 上游为**真实 IP** 解析路径,不与 fake-ip 混用 |
| 3 | **AGH → mosdns 仅为二期可选项**(经实测确有需求后启用) | ✅ §4 为二期方案;§8 kill-test 已给出"实测需求"证据(降级缺口) |
| 4 | 代理 VLAN10 将来用独立 OpenClash DNS 平面(fake-ip + TPROXY),不污染普通 LAN | ✅ 现状即此(dnsmasq→clash,非面向 LAN 客户端);文档记录 |
| 5 | **删除或明确禁用** `.1` 上未使用的 mosdns | ⚠️ 前提修正:mosdns 是 clash 的 nameserver,处于活动链路(§1)。处置改为**正式纳管并文档化**(本文件 + `hosts/gfw.windy.lan.md`),不删除 |
| 6 | 先完成验证再改动生产路径 | ✅ 本期完成全部 Phase 0 验证(§8),**未改任何生产 DNS 路径** |
| 7 | 建立 `home.arpa` 内部域(替代 `.local`) | ⏳ 后续任务:当前命名空间为 `.lan`(AGH rewrites + EdgeRouter DHCP domain),`hass.local` 兼容保留至迁移完成;home.arpa 需联动 AGH rewrites、DHCP domain、客户端,另行排期 |
| 8 | 高可用时增加第二个等价 AGH(独立物理故障域) | ⏳ 备用方案,记录不实施 |
## 3. 两个候选方案评估
### 方案 A:AGH 单独作为统一入口(现状演进)
- 优点:单解析点、面板/拦截/日志集中、维护简单。
- 缺点:AGH 对 geo 分流 + 防污染支持弱(官方定位是"过滤/家长控制")。固定 DoH 上游无法按域名
region 选路;且 §8 kill-test 显示主上游全挂时缓存未命中查询无有界降级。→ **不足以解决防污染/分流/降级问题。**
### 方案 B:mosdns 作为智能上游分流器
mosdns(v5)用 `sequence` 编排:`geosite/geoip` 匹配器 → 国内域名转发国内 DoH、国外域名转发加密 DoH(防污染),可加 `cache``reject`
- 优点:真正解决"国内快 / 国外不被污染"的分流;性能高。
- 缺点:纯转发器,无 Web 面板、无每客户端统计、拦截靠域名表。→ 单独当入口会退回原始体验。
**结论:两个方案互补,不是二选一。** 单用 A 无法分流防污染且降级无界;单用 B 失去 AGH 管理体验。
## 4. 二期可选项:分层架构(AGH 前端 + mosdns 后端,经实测需求后启用)
```
局域网客户端(DHCP DNS = 192.168.66.36)
AdGuard Home (66.36) ── 前端:广告/过滤、拦截表、每客户端统计、Web 面板
│ 上游 = mosdns
mosdns(66.36 伴生容器) ── 后端:智能分流 + 防污染
│ - geosite:cn → 国内 DoH/UDP(aliDNS / 腾讯 DNSPod)
│ - 其他 → 加密 DoH(自建 adg.chans.xyz 等)
上游 DoH
```
职责分离:
- **AGH = 策略/拦截/可观测**(拦截表、每客户端日志、面板)。
- **mosdns = 智能转发**(geo 分流 + 加密防污染 + 内置 cache,可显著缩短降级窗口)。
- **OpenClash(gfw)= 代理选路**(fake-ip + 规则)。与 DNS 解析选上游是两个独立决策,分开放最干净。
**部署位置:mosdns 与 AGH 同机(66.36 伴生容器),而非 gfw(.1)**:单点即 AGH 所在;不受网关重启/
OpenClash churn 影响;可纳入现有 compose/ansible 管理;不占用 OpenWrt 资源。放 gfw 会与 clash 的
DNS 处理互相干扰、耦合,且网关重启即断全 LAN DNS。**不推荐放 .1。**
> 启用条件(kill-test 实测需求,§8):主上游全挂时,当前 AGH 单入口对缓存未命中查询无有界降级。
> 分层方案(或下调 `upstream_timeout` + 改 failover 模式)可修;启用与否由用户在二期决定。
## 5. 更优替代方案(一并考虑)
1. **分层(推荐二期,见 §4)**:AGH(66.36)→ mosdns(66.36 伴生)→ 上游。体验最好、职责最清。
2. **纯 mosdns + 前端面板**:损失拦截/统计管理体验。**不推荐**用于替换。
3. **AGH 只挂一个带分流的上游(第三方 DoH 聚合)**:失去可控性且不可信。不推荐做主路径。
4. **全部交给 OpenClash fake-ip,关闭 AGH**:让"代理网关"成为全 LAN DNS 单点;且 AGH 拦截/日志也没了。**不推荐。**
## 6. mosdns 配置要点(mosdns v5,二期实施预留)
核心是 `sequence` + 上游拆分 + 缓存 + 屏蔽:
```yaml
plugins:
- tag: main
type: sequence
args:
- exec: cache 1024 # 缓存加速
- matches: has_resp
exec: accept
# 国内分流:命中 geosite:cn → 国内 DoH
- matches: [ qname &geosite:cn ]
exec: forward https://dns.alidns.com/dns-query
# 广告域名屏蔽交 AGH 前置,不重复维护
# 其余(国外)→ 加密 DoH 防污染
- exec: forward https://adg.chans.xyz/dns-query
- type: udp_server
args: { entry: main, listen: "127.0.0.1:5353" }
- type: tcp_server
args: { entry: main, listen: "127.0.0.1:5353" }
```
> **注意**:`sequence` 中每个 `forward` 分支之后必须跟 `matches: has_resp → accept`
> (或改用 `goto`/`jump` + `return` 结构),否则查询会继续执行后续规则被二次转发,
> 最终应答来自最后一个 forward——gfw 上 mosdns 的同类缺陷(2026-08-12)已实测并修复(见 §1)。
要点:
- 上游可加 `upstream``concurrent > 1` 与多地址故障切换;mosdns 自带 cache,能保证上游故障时
缓存命中仍即时应答(对应 §8 认定的降级缺口)。
- `geosite:cn` / `geoip:cn` 数据自动更新;国内 aliDNS/腾讯,国外可用自建 `adg.chans.xyz`(实测
唯一能返回国外 AAAA 的路径,§8)。
- 屏蔽交 AGH 前置,AGH 与 mosdns 不各自维护拦截表。
## 7. 迁移 / 实施顺序(二期,待用户确认启用)
1. 在 66.36 起 mosdns 伴生容器(`/opt/mosdns` + compose,**按 digest 固定镜像**,纳入 ansible)。
2. AGH「上游 DNS 服务器」改为指向 mosdns(`127.0.0.1:5353`,bootstrap 仍用公网 IP,避免
AGH → mosdns → AGH 死循环);只保留**一条语义一致的上游路径**,保持可回滚(备份 yaml + `--check-config`)。
3. 验证:国内域名、国外域名、被拦截域名、每客户端日志、AAA A 解析(§8 基线)。
4. 回归:EdgeRouter 通告不变(仍 `.36`),LAN 客户端无感;重启 AGH/mosdns 单点验证(§8 kill-test 模板)。
5. 上线后重跑 §8 kill-test,确认降级窗口有界。
## 8. Phase 0 验证证据(2026-08-12 全部实测)
### 8.1 基线与功能
| 项 | 结果 |
|----|------|
| rewrites:`hass.windy.lan` / `hass.local` | → `192.168.55.11` ✅(兼容保留) |
| rewrites:`dns.windy.lan` / `gfw.windy.lan` / `ubnt.windy.lan` / `nas.windy.local` | → `.36` / `.1` / `.46` / `.32` ✅ |
| 国内解析 `taobao.com`(经 AGH) | 真实 CN IP(59.82.x 等)✅ |
| 国外解析 `google.com` / `github.com`(经 AGH) | 真实 IP(142.250.x / 20.205.x),**无 fake-IP 泄漏** ✅ |
| 广告拦截 `doubleclick.net` / `googleadservices.com` | → `0.0.0.0` ✅ |
| clash 7874 `google.com` | `198.18.1.101`(fake-ip,仅网关/VLAN10 平面)✅ |
| clash 7874 `taobao.com` | 真实 IP(经 mosdns→AGH)✅ |
| dnsmasq :53(.1)`google.com` / `taobao.com` | fake-ip / 真实 IP ✅ |
| mosdns 6052 直连 `taobao.com` / `google.com` | 真实 IP(59.82.x / 142.250.73.78)✅ |
| DNSSEC:`dnssec-failed.org`(经 AGH) | 返回正常应答 `96.99.227.255`(非 SERVFAIL)→ 当前路径不校验,W1N-40 结论复现,**维持关闭** |
### 8.2 出口路径与 fake-IP 泄漏
- AGH 宿主默认路由 `via 192.168.66.254`(EdgeRouter),**直连出网,不经 gfw**;DoH 端点实测可达
(见 §1)。LAN 客户端经 AGH 的解析结果全部为真实 IP,无 `198.18/16` 泄漏。
- gfw 自身流量默认进代理(`openclash_mangle_output` mark 0x162 → tproxy :7895),
clash nameserver-policy 的 `1.1.1.1` DoH 实测 35ms 可达(走代理链路,不依赖直连)。
- mosdns 上游(AGH `.36``223.5.5.5`)命中本地/国内 bypass 规则,保持直连——设计意图达成。
### 8.3 IPv6 / RDNSS / AAAA
- LAN 有 IPv6 SLAAC(EdgeRouter dhcpv6-pd /60 → eth0 host-address + switch0,**仅 `service slaac`**,
**无 RDNSS/dns-server 通告**);`.36` 有全局 v6 地址 + RA 默认路由。
- **RDNSS 未通告** → v6 客户端无 v6 DNS,回退 v4 DNS(`.36`);AGH 仅监听 `0.0.0.0:53`(v4 only),无 v6 DNS 服务。
- **AAAA 解析实测**:`baidu.com` 公网本就无 AAAA(dns.google NOERROR/0,权威 NS 而已);
`taobao.com` AAAA 经国内路径正常(`2408:4001:f10::6f` 等);**国外域名(`google.com`)经
`223.5.5.5` UDP、alidns DoH、`8.8.8.8` UDP 全部返回空**,而 dns.google 与 `adg.chans.xyz`
DoH 均能返回 `2404:6800:4005:81a::200e` → **国内路径对国外域 AAAA 置空;`adg.chans.xyz`
兜底是当前唯一能返回国外 AAAA 的路径**。clash `ipv6: false` 亦不返回 AAAA。
### 8.4 自举(bootstrap)循环
- AGH `bootstrap_dns: [223.5.5.5, 223.6.6.6]`(公网 IP,非 AGH 自身)→ 无自举循环;AGH 解析
DoH 主机名不经过自身。二期方案要求 AGH→mosdns 时 bootstrap 仍用公网 IP(§7)。
### 8.5 Kill-test 矩阵(2026-08-12,全部实测)
| # | 场景 | 结果 |
|---|------|------|
| 1 | 重启 `.1` mosdns(init.d) | ✅ 直连 6052 与 clash 链恢复 |
| 2 | 重启 `.1` dnsmasq | ✅ 真实 IP 与 fake-ip 双路径恢复 |
| 3 | kill `.1` clash 核心 | ✅ LAN DNS(AGH)不受影响;gfw dnsmasq→clash **有界 3s 失败**(无卡死);OpenClash watchdog ~15s 自动拉起 |
| 4 | stop/start `.1` OpenClash | ✅ DNS 平面独立于代理;clash 与 tproxy 规则恢复 |
| 5 | 重启 `.36` AGH 容器 | ✅ 全量恢复:rewrites/拦截/国内外解析/DNSSEC 行为不变 |
| 6 | 黑洞 alidns DoH(223.5.5.5/223.6.6.6:443) | ✅ ~0.5s 内经 `doh.pub` 应答(load_balance 生效) |
| 7 | 黑洞全部主上游,兜底存活(adg.chans.xyz) | ⚠️ **客户端 15s 内无应答**——兜底未在窗口内生效 |
| 8 | 黑洞全部上游(含兜底),缓存未命中 | ⚠️ **25s 内无应答、无 SERVFAIL**——解析器对缓存未命中查询"卡死" |
| 9 | 黑洞全部上游,缓存命中 | ✅ 瞬时 NOERROR(cache 兜底) |
**结论(Phase 0 门禁):** 国内解析在代理停止/上游单点故障/组件重启下均维持可用;
但**"主上游全挂"时缓存未命中查询无有界降级**——`upstream_timeout: 10s` 与 TCP 重试行为使
兜底 `adg.chans.xyz` 在实践中无法在客户端期望窗口内生效。这是 §4 二期分层方案(或下调
`upstream_timeout` + failover 模式)的**实测需求依据**;按最终裁决,本期不改生产路径。
## 9. 风险与备注
- mosdns 仅监听 `127.0.0.1`(不对外),由 clash 消费;二期若启用,保持同样的边界,避免 LAN 出现两套入口。
- AGH 上游指向本机 mosdns 时务必配公网 bootstrap,否则自举死循环。
- 本方案不改 EdgeRouter DHCP/通告、不改 gfw OpenClash 代理规则,只动 66.36 上的 DNS 链路,风险可控。
- 已知降级缺口(§8.5 #7/#8):主上游全挂时缓存未命中查询无有界降级;启用二期前,LAN 客户端会感知
超时(约 10s+)。缓解:AGH cache 已覆盖高频域;根治需二期。
- 与 W1N-40「审查并修正 AdGuard Home」联动:该 issue 侧重 AGH 本身,本 issue 侧重整体 DNS 分层。
+1 -1
View File
@@ -307,7 +307,7 @@ VLAN10 / 升级 SSID / 客人 SSID **失败或未做,不否决**本次核心
## 13. 参考
- 实施阶段与清单:[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
- 实施阶段与清单:[lan-core-switch-upgrade-plan.md](archive/lan-core-switch-upgrade-plan.md)
- 现网地图:[lan-overview.md](lan-overview.md)
- ER-X[edgerouter-x-configuration.md](edgerouter-x-configuration.md)、[hosts/gw.md](../hosts/gw.md)
- UniFi / VLAN10 前置:[unifi-network.md](unifi-network.md)
+76 -11
View File
@@ -13,6 +13,11 @@ from each section below.
> **Verified live on 2026-08-06** by read-only SSH from the WSL client. No
> changes were made. `gfw.windy.lan` root SSH was re-verified the same day after
> the key was installed; its facts below are from the fresh probe.
>
> **IPv6 re-verified 2026-08-20** (read-only): UniFi controller `Default`
> network IPv6 enabled (SLAAC/RA), both APs hold global SLAAC addresses, and
> `zhiqiangf` key-only AP SSH re-confirmed. See
> [unifi-network.md](unifi-network.md).
---
@@ -35,10 +40,14 @@ from each section below.
│ ubnt — UniFi Network Controller (192.168.66.46)
```
> **SE5420 purchased (2026-08-09):** TP-Link `TL-SE5420` acquired; deployment plan is
> **SE5420 live (2026-08-22):** TP-Link `TL-SE5420` (purchased 2026-08-09) is
> online — management `192.168.66.253` reachable, web UI on :80/:443; LAN55
> 上联为 ER-X `switch0` **单口**`eth1` up、`eth2`/`eth3` down2026-08-22
> 只读核实)→ `switch0` 不再是 LAN55 全量抓包点(同段有线单播在 SE5420 本地
> 交换),全量点只能靠 SE5420 port mirroring。迁移状态见部署计划
> [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md). Design/planning refs:
> [lan-erx-se5420-network.md](lan-erx-se5420-network.md),
> [lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md).
> [lan-core-switch-upgrade-plan.md](archive/lan-core-switch-upgrade-plan.md).
---
@@ -47,19 +56,20 @@ from each section below.
| Host | Role | SSH | IPv4 | Facts |
|------|------|-----|------|-------|
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | [hosts/gw.md](../hosts/gw.md) |
| **PVE** | Proxmox host (`.66.26`/vmbr0 · `.55.26`/vmbr1) — hosts gfw/dns/ubnt/haos VMs | `ssh -4 root@192.168.66.26` | `192.168.66.26` | — |
| **PVE** | Proxmox host (`.66.26`/vmbr0 · `.55.26`/vmbr1) — hosts gfw/dns/ubnt VMs | `ssh -4 root@192.168.66.26` | `192.168.66.26` | — |
| **gfw.windy.lan** | OpenWrt LAN gateway / OpenClash — **PVE VM 140** | `ssh -4 root@192.168.66.1` | `192.168.66.1` | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy — **PVE VM 120** (`pihole`) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
| **ubnt** | UniFi Network Controller — **PVE VM 160** | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | [hosts/ubnt.md](../hosts/ubnt.md) |
| **haos** | Home Assistant (HAOS) — **PVE VM 180** (LAN55) | — | `192.168.55.11` | — |
| **hass.windy.lan** | Home Assistant (HAOS) — **x88 Pro physical box** (LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
| **pgdb** | TimescaleDB PG18 (Docker) — HA recorder 后端 — **PVE VM** (LAN55) | `ssh -4 windy@192.168.55.15` | `192.168.55.15` | [hosts/pgdb.md](../hosts/pgdb.md) |
| **NAS/FreeNAS** | NAS; `transmission` jail runs here (`.51`) | — | — | — |
| **U6 Lite** | UniFi AP (LAN66) | `ssh -4 zhiqiangf@192.168.66.6` | `192.168.66.6` | [docs/unifi-network.md](../docs/unifi-network.md) |
| **UAP-AC-Lite** | UniFi AP (LAN55) | `ssh -4 zhiqiangf@192.168.55.5` | `192.168.55.5` | [docs/unifi-network.md](../docs/unifi-network.md) |
> **Positioning facts (verified 2026-08-09):** `dns`/`ubnt`/`gfw`/`haos` are all VMs on PVE
> (no separate physical hosts); `transmission` is a FreeNAS/NAS jail. Only gw, PVE,
> NAS, U6, UAP-AC-Lite, and wired PCs/NAS are physical SE5420 ports. See
> [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) §1.
> **Positioning facts:** `dns`/`ubnt`/`gfw`/`pgdb` are VMs on PVE; `haos` is a **physical x88 Pro
> box** (HAOS bare-metal, `machine: green`), not a PVE VM (corrected 2026-08-15).
> `transmission` is a FreeNAS/NAS jail. Physical SE5420 ports: gw, PVE, haos, NAS,
> U6, UAP-AC-Lite, and wired PCs. See [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) §1.
---
@@ -76,7 +86,10 @@ from each section below.
| Port-forwards | `hass`→192.168.55.11:8123 · `transmission`→192.168.66.51:51413 · `ssh`→192.168.66.36:22 (orig 5822) · `openvpn`→192.168.66.32:1194 · WAN iface pppoe0 |
| Management | SSH TCP 22 · EdgeOS GUI HTTP 80 / HTTPS 443 |
**Static DHCP mappings (LAN66):** `OnePlus-12`=.37, `gfw`=.1, `hp-nas`=.32, `pihole`=.36, `pve`=.26, `transmission`=.51, `ubnt-6`=.6, `ubnt-app`=.46, `windy-pc`=.99. LAN55: `Aqara-Hub-M3-10CB`=.248.
**Static DHCP mappings (LAN66):** `OnePlus-12`=.37, `gfw`=.1, `hp-nas`=.32, `pihole`=.36, `pve`=.26, `transmission`=.51, `ubnt-6`=.6, `ubnt-app`=.46, `windy-pc`=.99. LAN55: `Aqara-Hub-M3-10CB`=.248, `SmartThings-Station`=.48, `espressif`=.47,
`hass`=.11, `hass-wifi`=.250, `ihost`=.12, `midea_ac_0418`=.10,
`midea_e3_0198`=.42, `roborock-wm-a141`=.43, `samsung-hub`=.251,
`matter`=.41 (added 2026-08-20).
> **Note:** `LAN_IN`/`LAN_OUT` are defined but not applied to an interface, so LAN55
> and LAN66 are bidirectionally reachable by default. Do not rely on those rules as
@@ -93,7 +106,7 @@ from each section below.
| SSH | `ssh -4 root@192.168.66.1` (key-only, verified 2026-08-06) |
| OpenClash | `/etc/openclash/clash` (clash_meta core) + config `/etc/openclash/pass-cat.yaml` |
| Mode | **fake-ip + TPROXY transparent proxy** (`operation_mode=fake-ip`, `en_mode=fake-ip`, `proxy_mode=rule`) |
| DNS | dnsmasq → clash DNS `127.0.0.1#7874`; `mosdns` also listens on `127.0.0.1:6052` (not the active path) |
| DNS | dnsmasq → clash DNS `127.0.0.1#7874`; clash `nameserver` = mosdns `127.0.0.1:6052` (DIRECT 规则真实 IP 解析,非客户端路径) |
| nft | `table inet fw4` with OpenClash TPROXY/redirect + DNS-hijack rules; residual `table inet passwall` (0 packets, unused) |
**OpenClash listeners:** HTTP `7890` · SOCKS `7891` · Redirect `7892` · Mixed `7893` · TPROXY `7895` · DNS `7874` · dashboard `9090`. `8443` is **not** an OpenClash listener (only in its TLS-sniffing port list).
@@ -145,6 +158,19 @@ See [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md).
---
## hass.windy.lan — Home Assistant (HAOS)
| Item | Value |
|------|-------|
| IPv4 | `192.168.55.11` (LAN55) |
| DNS | `hass.windy.lan` (AdGuard rewrite; legacy `hass.local` alias) |
| SSH | `ssh hassio@hass.windy.lan` (key-only, verified 2026-08-13) |
| Web UI | `http://hass.windy.lan:8123` |
| WAN | gw port-forward `hass``192.168.55.11:8123` |
| Platform | HAOS on physical x88 Pro box; kernel `6.1.115-haos` (aarch64), `machine: green` |
---
## Managed access points
| Name | Model | Mgmt IP | Firmware | Network | Inform |
@@ -155,6 +181,43 @@ See [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md).
Both reported **Connected** to `http://192.168.66.46:9080/inform` on 2026-08-06.
AP SSH account is `zhiqiangf` (key-only, verified). See [docs/unifi-network.md](../docs/unifi-network.md).
**IPv6 (verified 2026-08-20):** both APs hold global SLAAC IPv6 addresses on
`br0` — U6 Lite `240e:3bd:235:1fb1::/64` (LAN66), UAP-AC-Lite
`240e:3bd:235:1fb2::/64` (LAN55) — with RA default routes via `gw`; the
controller's `Default` network has IPv6 enabled (SLAAC). Prefixes are dynamic
(PPPoE PD), so they rotate on redial. Details:
[docs/unifi-network.md](../docs/unifi-network.md).
**SSID cleanup (2026-08-21, W1N-207):** the SmartThings Element/vWire provisioning
SSIDs (`element-8a0d5133c9438f12`, `vwire-8b2d67469e455785`, `vport-F09FC22004E9`)
were removed/disabled in the controller (`element_adopt` setting off, element wlanconf
deleted, connectivity `x_mesh_essid`/`x_mesh_psk` cleared, device `x_vwirekey` removed,
`vwire_enabled`/`mesh_sta_vap_enabled=false`) and cleared from both APs; all
vwire/vport/element flags on the remaining SSIDs are now `disabled`.
**Stable ULA on gw: not feasible (2026-08-21, W1N-207):** EdgeOS v3.0.1
`interfaces switch switch0` rejects a static `ipv6 address`, and an explicit
`router-advert` node *replaces* the DHCPv6-PD-slaac RA (drops the delegated GUA
prefix from radvd → LAN55 loses IPv6 egress after RA expiry). Attempted and rolled
back cleanly (no `save`; gw config unchanged). Consequence: after a PD rotation,
restart HA's matter-server (see [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md))
to clear stale IPv6 mDNS caches.
**LAN55 RA environment (observed 2026-08-21):** besides `gw`, the SmartThings
Station (.48) and Aqara M3 (.248) act as Thread border routers and advertise ULA
prefixes (`fd00:5a7:6415:1::/64`, `fd97:d580:16fe:1::/64`); several LAN55 hosts
(HA, PVE, UAP-AC-Lite) have IPv6 forwarding enabled and mark themselves as
routers in NDP. This is normal Thread-BDR behaviour and was not the Matter
failure cause.
**Matter 灯泡(2026-08-21 实测,W1N-207):** 两盏 ESP32-C2 Matter 灯泡
VP `0x4891/0x4100`OUI `34:98:7a`)——工作盏 MAC `34:98:7a:25:a1:f0`;故障盏
MAC `34:98:7a:27:7f:08`hostname `matter`,动态 .145)。故障盏已在 Aqara fabric
`4DF2B1455D19402D` 内、宣告 `CM=0`(不在配对模式)且缺 GUA → 找回需**恢复出厂**
后扫它自己的二维码。DHCP 保留 `matter`.45 → MAC `34:98:7a:27:10:bc`)与故障盏
MAC 不符,保留从未租出(待修,见 [hosts/gw.md](../hosts/gw.md))。完整排障知识:
[docs/matter-pairing-troubleshoot.md](matter-pairing-troubleshoot.md)。
---
## Quick orientation (who runs what)
@@ -165,6 +228,7 @@ AP SSH account is `zhiqiangf` (key-only, verified). See [docs/unifi-network.md](
| Transparent/explicit proxy (OpenClash) | gfw.windy.lan | `ssh -4 root@192.168.66.1` |
| LAN DNS (AdGuard Home) + Mihomo proxy | dns.windy.lan | `ssh -4 windy@192.168.66.36` |
| UniFi controller + dockge | ubnt | `ssh -4 windy@192.168.66.46` |
| Home Assistant | hass.windy.lan | `ssh hassio@hass.windy.lan` · UI `:8123` |
| Wi-Fi APs | U6 Lite / UAP-AC-Lite | via controller |
---
@@ -175,9 +239,10 @@ AP SSH account is `zhiqiangf` (key-only, verified). See [docs/unifi-network.md](
- [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) — OpenClash listeners
- [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) — AdGuard Home + Mihomo detail
- [hosts/ubnt.md](../hosts/ubnt.md) — UniFi controller + proxy contract
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — Home Assistant (HAOS) SSH + LAN access
- [docs/unifi-network.md](../docs/unifi-network.md) — APs, inform endpoint, recovery
- [docs/unifi-third-party-vlan10-dhcp.md](unifi-third-party-vlan10-dhcp.md) — VLAN Wi-Fi feasibility and DHCP boundary
- [docs/unifi-openwrt-vlan10-implementation-examples.md](unifi-openwrt-vlan10-implementation-examples.md) — supported topology and examples
- [docs/edgerouter-x-configuration.md](../docs/edgerouter-x-configuration.md) — effective gw config
- [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md) — proxy bypass
- [runbooks/adguard-home-health.md](../runbooks/adguard-home-health.md) — AGH health
- [runbooks/adguard-home-health.md](../runbooks/adguard-home-health.md) — AGH health
+6 -6
View File
@@ -106,7 +106,7 @@
- 第一步不建 VLAN10、不向 ER-X 送任何 tag、口 4/6PVE、U6)不做 trunk。
- NAS 只接口 8,口 12 断开(LACP 是独立维护窗)。
- 不占口的 VMdns(.36=VM120)、ubnt(.46=VM160)、gfw(.1=VM140)haos(.55.11=VM180)transmission(.51) 是 NAS jail。
- 不占口的 VMdns(.36=VM120)、ubnt(.46=VM160)、gfw(.1=VM140)haos(.55.11) 是物理 x88 Pro 盒子(非 VM)transmission(.51) 是 NAS jail。
## 3. 开箱与固件升级
@@ -207,7 +207,7 @@
4. **验证:**
- `ip -br addr``vmbr1` = `192.168.55.26/24`
- `ping -c3 192.168.55.254` → 通。
5. 逐台验证 VM(顺序:gfw → dns → ubnt → haos):
5. 逐台验证(顺序:gfw → dns → ubnt → haos;前三个是 VMhaos 是物理盒子):
```bash
ssh -4 root@192.168.66.26 'qm list'
```
@@ -215,7 +215,7 @@
- dns`ping -c3 192.168.66.36` → 通;
- ubnt`ping -c3 192.168.66.46` → 通;
- haos`ping -c3 192.168.55.11` → 通(注意是 55 网段)。
6. 每个 VM 再验业务:gfw 的 OpenClash 面板/DNS 正常、dns 的 AdGuard UI 能开、ubnt 控制器 Connected、haos 界面能开。不以"宿主开机"代替。
6. 每再验业务:gfw 的 OpenClash 面板/DNS 正常、dns 的 AdGuard UI 能开、ubnt 控制器 Connected、haos 界面能开。不以"宿主开机"代替。
## 7. 迁移 AP 与接入设备
@@ -462,7 +462,7 @@ ssh -4 root@192.168.66.1 'uci show network; uci show firewall; uci show dhcp; ip
- 每次实质变更后在 Linear `vps` 项目记录 scope / action / verification / 遗留 follow-up。
- 本仓库不记录 SE5420 口令、ER-X 配置快照(含 PPPoE/口令)、gfw 凭据。
- 实施前先读 `se5420-review-claim-verification-2026-08.md` 的现场只读复核结论。
- 实施前先读 `archive/se5420-review-claim-verification-2026-08.md` 的现场只读复核结论。
## 16. 回滚
@@ -486,10 +486,10 @@ ssh -4 root@192.168.66.1 'uci show network; uci show firewall; uci show dhcp; ip
## 参考
- 设计说明:[lan-erx-se5420-network.md](lan-erx-se5420-network.md)
- 评审核实:[se5420-review-claim-verification-2026-08.md](se5420-review-claim-verification-2026-08.md)
- 评审核实:[se5420-review-claim-verification-2026-08.md](archive/se5420-review-claim-verification-2026-08.md)
- 现网地图:[lan-overview.md](lan-overview.md)
- 官方安装手册(Markdown 版):[se5420-official-manuals/tl-se5420-install-manual.md](se5420-official-manuals/tl-se5420-install-manual.md)
- 官方 PDF<https://service.tp-link.com.cn/download/202310/TL-SE5420%20V1.0安装手册%201.0.2.pdf>
- 规格 / 固件:<https://www.tp-link.com.cn/product_2899.html?v=specification> · <https://www.tp-link.com.cn/product_2899.html?v=download>
- Omada VLAN 指南:<https://support.omadanetworks.com/en/document/12981/> · <https://support.omadanetworks.com/en/document/13135/>
- ER-X[edgerouter-x-configuration.md](edgerouter-x-configuration.md)gfw[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md)PVE VLAN10[lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09](lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09)
- ER-X[edgerouter-x-configuration.md](edgerouter-x-configuration.md)gfw[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md)PVE VLAN10[lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09](archive/lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09)
+1
View File
@@ -108,3 +108,4 @@ Steps:
- Matrix Authentication Service: <https://github.com/element-hq/matrix-authentication-service>
- Matrix spec: <https://spec.matrix.org/>
- Federation tester: <https://federationtester.matrix.org/>
- Home Assistant Matrix integration: [home-assistant-matrix.md](home-assistant-matrix.md)
+215
View File
@@ -0,0 +1,215 @@
# Matter 配网排障手册
> 基于 Matter 1.5.1 Core Spec §4.3.1 与本环境(EdgeRouter X + UniFi AP + Aqara M3 +
> Home Assistant2026-08-21 实测整理。配套 Linear W1N-207。
## 1. Matter 配网协议要点(发现即一切)
- **发现走 mDNSDNS-SD**UDP **5353**,组播 `224.0.0.251` / `ff02::fb`
**不经过单播 DNS(如 AdGuard .36)、不需要反向 DNS、不需要 DHCPv6**SLAAC 即满足 Matter
的 IPv6 要求)。
- 服务类型:
- `_matterc._udp` — 可配网设备(Commissionable),**配对模式才有效**
- `_matter._tcp` — 已配设备(Operational),TXT 里含 fabric 信息
- 子类型(配对方按此过滤):
- `_L<全12位 discriminator>`(如 `_L3266`)— 按二维码里的完整 discriminator 精确匹配
- `_S<高4位>`(如 `_S12`
- `_V<vendorId>``_T<deviceType>`(可选)
- `_CM`(仅真正处于配对模式时发布)
- TXT 关键键:`D=`discriminator,规范 **SHALL** 必填)、`VP=`vendor+product)、
**`CM=`**、`RI=`rotating id)、`PH=`/`PI=`(配对提示)。
- 配对端口:**TCP 5540**PASE/CASE)。部分生态(Aqara M3)为 Thread 中继节点用 **5552**
- 实例名:64 位随机 hex;**进入配对模式时更换**(可用作"是否重新进过配对"的信号)。
- 规范参考:[Matter 1.5.1 Core Spec §4.3.1](https://csa-iot.org/wp-content/uploads/2026/03/23-27349-010_Matter-1.5.1-Core-Specification.pdf)、
[Google Home: Commissionable and Operational Discovery](https://developers.home.google.com/matter/primer/commissionable-and-operational-discovery)、
[Matter Handbook: Discovery](https://handbook.buildwithmatter.com/how-it-works/discovery/)、
[connectedhomeip: IP commissioning](https://pigweed.googlesource.com/third_party/github/project-chip/connectedhomeip/+show/59edd2ff8506b1e3dabb7040d716f0e75a2312d1/docs/guides/ip_commissioning.md)。
## 2. 关键判据:CM=0 = 不在配对模式
规范 §4.3.1.2 / §4.3.1.7
- 设备可以长期宣告 `_matterc`**Extended Discovery**),但 **`CM=0` 表示"当前不接受配网"**。
- **已在 fabric 里的设备**(宣告里同时有 `_matter._tcp` + `_I<fabric>._sub` 运营记录)重配时
通常报 `CM=0` —— 它已配好,不是新设备。
- **配对方不能把已配设备当新设备加** → 重加/找回必须先**恢复出厂**(清 fabric,重启后以
`CM=1` 全新配对模式宣告),再用**它自己的二维码**添加。
- 常见误判:抓包看到 `_matterc` 宣告就以为"在配对模式"——**必须看 `CM=`**。
## 3. 本环境实测事实(2026-08-21W1N-207
| 事实 | 状态 |
|---|---|
| LAN55 IPv6/mDNS 链路 | ✅ 全正常(RA→交换机→AP→客户端;mDNS 双向通;igmp snooping off、mdns on、无客户端隔离、无组播增强、PMF off、WPA2、仅 2.4G |
| Matter 不依赖单播 DNS/.36、反向 DNS、DHCPv6 | ✅ 已排除(.36 健康且不在路径上) |
| HA matter-server 曾宣告两代前的旧 GUA | ✅ 已修复(重启 `core_matter_server`;宣告恢复当前前缀) |
| ISP PD /60 随重拨轮换 → Matter IPv6 缓存反复失效 | ⚠️ 环境性根因;对策 = 重拨后重启 matter-server + 重启 M3 |
| EdgeOS 上静态 ULA 不可行 | ✅ 已尝试并回滚(switch0 不支持静态 `ipv6 address`;显式 router-advert 会替换 PD-slaac RA |
| 在用的两盏 ESP32-C2 Matter 灯泡(VP `0x4891/0x4100`2026-08-23 复核) | 工作盏 MAC 已变为 `fc:e8:c0:25:a1:f0``.146`hostname `espressif`;原 `34:98:7a:25:a1:f0` 全网消失,疑固件更新后换 MAC——末 3 字节相同);新盏 `34:98:7a:27:10:bc``.148`hostname `matter`)。两盏各宣告 **3 个 fabric** 运营实例:Aqara `4DF2B1455D19402D``2F6E56020E1996E7`、HA `DCE86145C137AF0E`(见 §8 |
| 故障盏 `34:98:7a:27:7f:08`(曾 .145Aqara fabric`CM=0` 缺 GUA | 2026-08-23 复核:无租约、ARP incomplete、AP 无日志 = **已离网**(退役/退换) |
| **失败模式 C(2026-08-23 实测,两盏同时)**mDNS 活、5540 死 | 灯泡 ping 通(v4/v6)、DHCP 正常续租、mDNS 应答并宣告 `_matter._tcp`SRV :5540、TXT `T=1`、当前前缀 GUA),但 **TCP 5540 在 IPv4 与 IPv6fe80+GUA)均 RST 拒绝** → 配对方无法建立 CASEApp 显示离线;hass matter-server 侧无任何 established :5540 会话(详见 §8 |
| ISP PD 前缀再次轮换(2026-08-23 → `240e:3bd:238:4812::/64`08-22 为 `235:1fb2` | hass 与 `.148` 均持当前前缀 GUA;hass 残留 `.146` 旧前缀 GUA 的 **FAILED** 邻居项(旧地址缓存仍被某端尝试) |
| DHCP 保留 `matter`.45 → MAC `…10:bc` | ⚠️ 保留仍未生效:新灯泡(`…10:bc`)实际拿到动态 `.148` 而非保留的 `.45`(待修,见 hosts/gw.md |
| **新灯泡(2026-08-22 添加成功)**MAC `34:98:7a:27:10:bc`=DHCP 保留目标 MAC),hostname `matter`IP `.148`VP `4891/4100`D=`3377` | ✅ 已入 **Aqara fabric `4DF2B1455D19402D`****经 BLE 配网**Aqara Home App)——线上**无 TCP 5540** 属正常(BLE 会话对 AP/hass 抓包不可见) |
| ESP32-C2 灯泡 firmware 挂死模式(2026-08-22 实测) | 入网后宣告 `_matterc`CM=1、D=3377)约 **3 秒后网络栈完全静默**:STA 收发计数冻结、不掉线不重启、配对方(手机/M3 `_L3377` 查询)无应答 → 加不上。**对策=断电 10 秒重启**重新进配网模式(实例名更换:`3F4E2C66F2DA85CD``E5BA8E28E4DE23A0`),随即 App 添加即成功 |
| 遗留 SSIDelement/vwire/vport | ✅ 已清理 |
## 4. 抓包方法(BusyBox 兼容)
> 完整指令集(实时 / 落盘轮转 / 定向抓取 / Wireshark 解密)见
> [runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md)。
> 下面是最常用的两条。
**视角必须在 LAN55**。**HA matter-server 作配对方时推荐直接在 hass `end0` 抓**——配对方
必然参与配对流程的每一条通讯(mDNS 本段组播 + 自己的 TCP 5540 全程),覆盖最全;AP `br0`
能看到全部 mDNS 组播 + 无线客户端单播,但**看不到有线↔有线单播**(如 Thread 设备经有线 M3
配对时 HA↔M3 的 5540 在 AP 侧不可见)。66 网段电脑看不到 55 的组播。BusyBox 注意点仅适用
AP**不要用 `--line-buffered`**;引号外层双引号、内层单引号);hass 是 HAOS 全量 tcpdump。
完整抓取(跑配对时保持窗口开着,`Ctrl+C` 结束):
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
hass 侧(HA matter-server 作配对方,推荐;非交互 ssh 需显式 `sudo -n -i`):
```bash
ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
精简过滤(只看 Matter 信号):
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552' | grep -E '_matterc|_matter|_L[0-9]+|_S[0-9]+|_CM|_V[0-9]+|_T[0-9]+|\.5540|\.5552'"
```
存 pcap 供 Wireshark:把上面 `-w /tmp/matter.pcap` 追加到 tcpdump 参数(去掉 `-vvv`),
`scp zhiqiangf@192.168.55.5:/tmp/matter.pcap .` 拉回本地分析。
> **落盘务必轮转**AP `/tmp` 只有约 60MB。用
> `-C 5 -W 12 -w /tmp/matter.pcap`(每 5MB 轮转、最多 12 个文件)防止写满,
> 详见 runbook Step 3(落盘轮转)。
> **Matter 载荷是加密的**mDNS5353)明文可读;5540 上的 Matter 报文要看明文
> 需要 Wireshark matter-dissector + 会话密钥,详见 runbook Step 5(解密)。
### 阶段对照表
| 阶段 | 应该看到 | 对应问题 |
|---|---|---|
| 发现(设备侧) | `_matterc._udp` + `_L3266._sub` + `_S12._sub` + TXT `D=3266 CM=1` + SRV `:5540` + AAAA | **无宣告**=设备没入网/没进配对模式;**`CM=0`**=不在配对模式(已配设备);**无 `_L3266`**=固件子类型缺失 |
| 发现(配对方侧) | M3/手机查询 `_L3266._sub._matterc._udp` | 查询有、无应答 = 码/discriminator 不匹配或设备不在线 |
| 配对握手 | 到设备 IP **TCP 5540 SYN/SYN-ACK** 双向 | **SYN 无 ACK**=设备不可达/防火墙;**完全无 5540**=发现阶段没完成 |
| 配完后 | 设备宣告 `_matter._tcp` + `_I<fabric>._sub` | 出现 = 已入网成功 |
| BLE 配网(手机 App 直连设备 BLE,如 Aqara Home | 线上**无 TCP 5540**BLE 会话对 AP/hass 抓包不可见);设备入网后仍先 mDNS 宣告 `_matterc` | 成功判据=最终宣告 `_matter._tcp` + `_I<fabric>._sub`;无 5540 **不代表**失败 |
## 5. 排障决策树(按顺序)
1. 抓包看**有没有 `_matterc` 宣告**:没有 → 设备不通电 / 没连上 Wi-Fi / 没进配对模式
(先解决"设备在线",网络侧已反复验证正常)。
2. 有宣告但 **`CM=0`** → 设备已配 / 不在配对模式 → **恢复出厂**后重试(用它自己的二维码)。
3. 有宣告 `CM=1` 但**无 `_L<disc>` 子类型** → 固件 mDNS 缺陷 → 升固件或换通用发现配对方。
4. `CM=1` + 子类型齐全但**无 TCP 5540** → 配对方没匹配上(查码/discriminator)或设备不可达。
5. 有 5540 但配对中断 → 查 `CM` 源(码是否正确)、设备电源、fabric 状态(是否需先清)。
## 6. 相关文档
- [runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md) — Matter 抓包指令集(实时/落盘轮转/定向/解密)
- [docs/lan-overview.md](lan-overview.md) — LAN 拓扑、SSID 清理、ULA 不可行
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — matter-server 重拨运维规范
- [docs/unifi-network.md](unifi-network.md) — UniFi 网络/IPv6/SSID 记录
- [hosts/gw.md](../hosts/gw.md) — DHCP 保留 `matter` MAC 错位(待修)
## 7. 2026-08-22 实测记录:添加新 ESP32-C2 Matter 灯泡(成功 + 失败路径全记录)
> 场景:手机 AppAqara Home)添加一盏**新的** ESP32-C2 Matter 灯泡
> `34:98:7a:27:10:bc`hostname `matter`,最终 IP `.148`)。中途换了灯泡并断电重启,
> 共经历 **2 种失败模式** 和 **1 条成功路径**,全部抓包实证。
>
> 抓包点:UAP-AC-Lite `192.168.55.5` `br0`(轮转 `udp 5353 or tcp 5540 or tcp 5552`
> + 定向全量 `ether host 34:98:7a:27:10:bc`+ hostapd/stahtd 日志 + gw DHCP/ARP 交叉验证。
> 本地用 tshark 4.7.2 分析。
### 时间线(CST2026-08-22
| 时间 | 事件 | 判据 / 说明 |
|---|---|---|
| 10:50:28 | 启动轮转抓包 | — |
| 10:51:1219 | 手机 `.143`OnePlus,连 wifi0ap0=`ubnt-windy-2`)重新关联;查询 `_matter._tcp` | 运营查询(浏览已配设备),**不是**配网(配网应查 `_matterc._udp` |
| ~10:56 | 用户报「配置 wifi 后挂起,不能加入 wifi」 | 首次失败 |
| 11:00:15 | M3 `.248` 查询 `_L3266._sub._matterc` | 无应答(是另一台设备的 discriminator,无关) |
| 11:01:2442 | **失败模式 A**:故障盏 `…7f:08` 尝试关联 `wifi0ap1`(`ubnt-haas`):发 1 次 open-auth 帧(algorithm 0)→ AP 回 `status_code=0`**客户端不再发 assoc 请求** → 18s 后 `auth_failures=1` + disassociated | auth 阶段卡死(client 侧);非密码错——密码错会先 assoc 再 4-way 失败 |
| 11:04:0105 | **新灯泡 `…10:bc` 关联 `wifi0ap1` 成功**WPA2 4-way 完成,DHCP 拿 `.148`tracker `soft failure`: ip_delta 3.76savg_rssi -68);随即宣告 `_matterc`:实例 `3F4E2C66F2DA85CD`TXT `VP=4891+4100 D=3377 CM=1`SRV :5540**有 GUA** | 发现阶段判据全过 |
| 11:04:05 之后 | **失败模式 B**:灯泡网络栈完全静默——STA 收发计数冻结(rx=89/tx=5 持续 12s+ 不变)、**不掉线不重启** | firmware 挂死 |
| 11:04:5911:07:27 | 手机查 `_matterc` ×4、`.60` 解析实例、M3 查 `_L3377._sub._matterc` ×5discriminator 3377 正是新盏)——**全部无应答**TCP 5540/5552 全程 0 | 配对方找不到设备 → App 报「加不上」 |
| 11:09:3548 | **断电 10 秒重启**:灯泡重新关联 `wifi0ap1` ×2 | 对策生效 |
| 11:10:26 | DHCP 重新拿 `.148`;STA 计数恢复持续增长(活跃) | — |
| 11:1011:15 | 重新宣告 `_matterc`**新实例 `E5BA8E28E4DE23A0`**——入配对模式实例名更换,符合规范);App 走 **BLE 配网** | 线上无 TCP 5540BLE 对 AP 不可见,属正常) |
| 11:15:23 | 灯泡宣告 **`_matter._tcp`**`4DF2B1455D19402D-02EF2FF12DFAF10E`Aqara fabric+ SRV :5540 + GUA + A | ✅ **添加成功**(已入 Aqara fabric `4DF2B1455D19402D` |
### 结论与经验
1. **同族灯泡(VP 4891/4100OUI 34:98:7a)存在两种不同失败模式**
- 故障盏 `…7f:08`:auth 阶段卡死(auth 帧后不发 assoc);此前(08-21 21:27)成功关联后伴随
`ip_failures=1`(拿不到 IP)+ 缺 GUA —— 属更深层故障,需恢复出厂,本次未处理,仍离线。
- 新盏 `…10:bc`:入网 + 宣告 `_matterc`(CM=1)成功后约 3 秒固件挂死(全静默)。
**断电 10 秒重启即恢复**,是最简单有效的对策。
2. **AP 抓包看不到 BLE 配网**Aqara Home App 对 WiFi Matter 设备走 BLE 配网时,线上只有
mDNS/DHCP**无 TCP 5540 不代表失败**;成功判据 = 设备最终宣告 `_matter._tcp` + `_I<fabric>._sub`
3. **发现判据回顾**`_matterc` + TXT`CM=1``D=``VP=`+ SRV :5540 + AAAA(GUA) + A 全齐才算
设备真的在配对模式;配对方按 `_L<disc>._sub._matterc` 精确匹配 discriminator(本例 D=3377)。
4. 新盏 RSSI -68、DHCP 3.76s,射频偏弱,可能加剧 firmware 不稳定(待观察)。
5. DHCP 保留 `matter`.45→`…10:bc`**仍未生效**:新盏实际拿动态 `.148`(待修,见 hosts/gw.md)。
6. 识别「配对方在找但设备不答」的快速方法:抓包里配对方持续查 `_matterc`/`_L<disc>` 而目标 MAC
零应答 + STA 收发计数冻结 = 设备侧挂死;此时**先断电重启设备**,不要怀疑网络/AP。
### 后续:新盏 11:24 起离线循环(同一盏 `…10:bc`2026-08-22
配网成功后约 10 分钟(11:15–11:24 可控制),灯泡进入**持续性故障循环**:
| 时间 | 事件 | 模式 |
|---|---|---|
| 11:24:57 | `EVENT_STA_LEAVE`(真掉线) | 掉线 |
| 11:25:12 | 重连 `auth_failures=2` | **auth 卡死**(同故障盏 `…7f:08` 11:01 的模式) |
| 11:25:2325 | 重连成功,WPA2 完成,重新拿 `.148` | — |
| 11:25:38 | `soft failure`ip_delta 2.65s**avg_rssi -73**-68→-73 持续变差) | 射频偏弱 |
| 11:25 之后 | STA 计数冻结(rx=119/tx=64 不动);M3 持续查询其运营实例 `4DF2B1455D19402D-02EF2FF12DFAF10E._matter._tcp` **无应答** → App 显示「离线」 | 静默挂死 |
**结论**:三盏 ESP32-C2 灯泡中两盏(`…7f:08``…10:bc`)故障,表现覆盖 auth 卡死 / 静默挂死 /
随机掉线三种形态;工作盏 `…25:a1:f0` 正常。网络侧(AP、M3、DHCP、mDNS)均验证正常。
**疑似根因(按可能性)**:① ESP32-C2 Matter 灯泡 firmware 缺陷(同批次)② 射频偏弱
(RSSI -73,天线/距离/遮挡)加剧不稳定 ③ 供电不稳(brownout 造成 Wi-Fi 栈崩溃重启)。
**待办**:移近 AP 或改善供电后观察;App 内查固件更新;仍复发则考虑退换。
## 8. 2026-08-23 状态核查:两盏「半在线」——mDNS 宣告正常但 TCP 5540 无监听(失败模式 C)
> 全程**只读**核查(gw DHCP/ARP、AP hostapd 日志、hass matter-server 状态 + mDNS 抓包、
> 对灯泡 v4/v6 的 TCP 5540 探测,09:0x CST)。结论:**网络侧全部健康;两盏灯泡网络栈活着、
> mDNS 运营宣告正常,但 Matter 会话端点(TCP 5540)无监听**——配对方无法建立 CASE,
> App 内应显示离线/不可达。
| 对象 | 状态(2026-08-23 |
|---|---|
| 新盏 `34:98:7a:27:10:bc``.148`hostname `matter` | DHCP 04:40 续租;gw ARP 完整;ping 通(93122msESP32 省电时延);08-22 16:40 起稳定关联 `wifi0ap1`,关联时 `avg_rssi -70`。mDNS 宣告 3 实例:`4DF2B1455D19402D-02EF079EEB480D07`**新 node ID——08-22 之后被重新配网过**)、`2F6E56020E1996E7-137147AF27BE4EB6``DCE86145C137AF0E-0000000000000011`HA fabric);host 记录 A `.148` + fe80 + **当前前缀** GUA `240e:3bd:238:4812:*`。支持单播 legacy mDNS 查询(`dig -p 5353 @.148 _matter._tcp.local PTR` 可用) |
| 工作盏(MAC 已变)`fc:e8:c0:25:a1:f0``.146`hostname `espressif` | DHCP 07:17 续租;ping 通 v4/v6v6 fe80 3861ms)。mDNS 宣告 3 实例:`4DF2B1455D19402D-02EF4CA3F856B615``2F6E56020E1996E7-EE8F2E4F1A77BF05``DCE86145C137AF0E-000000000000000B`。原 MAC `34:98:7a:25:a1:f0` 全网消失(无租约/ARP/AP 日志)而新 MAC 末 3 字节相同 → 疑固件更新后改 MAC。**拒绝单播 5353**ICMP port unreachable),只应答组播查询——同族固件行为差异。hass 残留其旧前缀 GUA `240e:3bd:235:1fb2:fee8:c0ff:fe25:a1f0`**FAILED** 邻居项 |
| 故障盏 `34:98:7a:27:7f:08`(曾 `.145` | 无租约、ARP incomplete、AP 日志零事件 = 已离网 |
| **TCP 5540 探测(两盏)** | IPv4LAN66 与 hass 本段)、IPv6fe80%end0 + 当前 GUA)全部 **RSTConnection refused** —— SRV 宣告 :5540 且 TXT `T=1`,但实际无监听 |
| hass matter-server | `started`,v9.0.4,无更新;宣告自身运营实例 `DCE86145C137AF0E-…1B669`v4+v6,当前 GUA);**无任何 established :5540 会话**core/add-on 日志无 matter 错误 |
| 其他 Matter 控制器 | Aqara M3 `.248` 在线(有线 0.8ms),宣告含自身 fabric 节点 `4DF2B1455D19402D-11E158E46D24A000`SmartThings `.48` 在线并周期查询 `_matter._tcp.local`;手机(当前前缀 GUA)也在浏览。LAN55 共见 **5 个 fabric**`4DF2B1455D19402D`M3)、`DCE86145C137AF0E`HA)、`2F6E56020E1996E7``03BCFAEDD6153944``6A6FF80C2DB84DEE` |
**判定**:失败模式 C = TCP/IP 栈与 mDNS 守护进程活着(主动 RST、DHCP 续租、ping 通),
但 Matter 应用层监听不存在。与模式 A(auth 卡死)、模式 B(全静默挂死)同族不同形态;
**两盏同时处于同一状态**更指向共同诱因(固件缺陷,或 PD 轮换等共同事件后未恢复)。
**对策(推荐,未执行)**:逐盏断电 10 秒重启(模式 B 的已验证对策),重启后复测
TCP 5540 恢复监听即可确认。
**核查方法备忘**(只读,可复用):
- gw`show dhcp leases` / `show arp`(经 `/opt/vyatta/bin/vyatta-op-cmd-wrapper`)。
- AP`grep -i <mac> /var/log/messages`hostapd 关联事件 + stahtd RSSI/soft failure)。
- hass`sudo -n -i ha apps info core_matter_server``ip -6 neigh show dev end0`
(看灯泡 fe80/旧新前缀 GUA 与 FAILED 项);被动抓包
`sudo -n -i timeout 65 tcpdump -ni end0 -s 0 -tt 'udp port 5353'`——配对方周期查询
会自然引出灯泡宣告,无需主动发包。
- 5540 探测:hass 上 python3 对 v4 / fe80%end0 / GUA 各 connect 一次;RST=无监听,
超时=不可达(两者含义不同)。
+45
View File
@@ -0,0 +1,45 @@
# Plane CE 加固草稿(docs/plane-hardening/
> **状态:草稿,未应用、未提交。** 对应追踪:Plane vps 项目条目(2026-09-03**记录源**Linear W1N-277 已取消,Linear 自 2026-09-03 起不再作为记录源)。
> 线上实例:`plane.chans.xyz`synapse K3sns `plane`release `plane-app` = chart `plane-ce-1.8.0` / app `v1.4.1`)。
> 依据:2026-09-03 只读核查(13 条审查意见中 11 条属实、#3 基本属实、#9 指标归属错误)+ 上游 chart 模板逐条核对。
## 文件
| 文件 | 内容 |
|------|------|
| `values.hardened.yaml` | 可选硬化 valuesexternal secrets 引用、requireExplicitSecrets、minio pin、上传限额对齐);含 HTTP→HTTPS `extraObjects` 示例 |
| `secrets.yaml.example` | 6 组外部 Secret 结构占位(只含 key 名,真实值仅存宿主机) |
| `backup/plane-backup.yaml` | **PostgreSQL 备份 CronJob**pg_dump `-Fc`hostPath `/var/backups/plane`MinIO 已按实际用量剔除) |
| `backup/README.md` | 备份方案说明(排程/容量/保留/还原/阻塞) |
## 应用顺序(每步先 diff 后执行,全部需用户逐项确认)
### 现在就值得做:DB 备份(P0,见 backup/
`plane-backup.yaml` 部署 + 手动触发验证一次即可;88 MB 库每日快照几乎零成本。
### 可选(顺手做一次,不是必须)
- **Phase A 密钥外部化**(零行为变化、无停机,约 15 分钟):按 `secrets.yaml.example`
在宿主机建 6 个 Secret(值先复制当前集群),用 `values.hardened.yaml`
`helm diff upgrade``helm upgrade`;验证后删除 chart 生成的旧 Secret。
价值:默认密钥不再落在 chart 公开常量上,作为保险。
- **MCP API Key 轮换**:若审查对话出过你的环境,Plane 后台重生成 + 更新
`/home/windy/plane-k3s/mcp/mcp.env`0600+ 重启 Cursor MCP。
- **/god-mode IP 白名单**:若在意管理后台被公网爆破。chart 1.8.0 的 IngressRoute
不支持给单条路由追加 middleware → 需 post-renderer 或 upgrade 后 `kubectl patch`
(升级会覆盖,需固化);源 IP 清单待提供。
### 明确暂缓/跳过(个人单节点,等出现症状再处理)
- SECRET_KEY 等轮换(Phase B):等真要配 SMTP/OAuth 前再做(避免旧密文不可解)。
- NetworkPolicy、有状态组件 resources limitschart 无 values 开关,需 post-render/patch)、
HTTP→HTTPS(草稿已给 `extraObjects` 示例)、metrics-server/Sentry。
## 关键限制(chart 1.8.0 模板已核对)
- `external_secrets.*_existingSecret` 设置后,对应 Secret **必须**包含模板所需全部 key
(缺失不自动补),见 `secrets.yaml.example` 注释。
- `app_keys_existingSecret` 的 envFrom 在所有 workload 上**最后注入**(后置生效),
保证 app/live 共享密钥一致——不要在其后再放同名 key 的 Secret。
- `DATABASE_URL`/`AMQP_URL`/`REDIS_URL` 是 chart 生成的派生 URL,内嵌明文密码;
外部化后轮换 DB/队列密码时必须同步更新 `plane-app-env`
- minio 的 `MINIO_ROOT_*``AWS_*` 同源于一个 Secret;升级时 bucket Job 会重跑
(需 admin 权限凭据)——换 svcacct 前先确认权限覆盖该 Job。
+48
View File
@@ -0,0 +1,48 @@
# Plane CE 备份方案(DB-only)— DRAFT (2026-09-03), 未应用
> 关联:`plane-backup.yaml`CronJob);追踪:Plane vps 项目条目(记录源,2026-09-03 起不用 Linear)。
> 现状(实测):pg 全库 **88 MB**310 issues / 1 user);MinIO uploads **264 KB**(几乎空)。
## 范围决策(2026-09-03,实际角度)
- **做:PostgreSQL 逻辑备份** —— 覆盖现实故障(误删、升级失败、磁盘坏、重装),成本≈0。
- **不做:MinIO/附件备份** —— 桶仅 264 KB,个人实例附件可接受丢失;不为它付日常维护。
日后附件明显变多再按原完整版思路加 `mc mirror`(历史版本见本目录 git 历史/Plane 条目评论)。
- 异机同步暂不启用(见下"局限/阻塞")。
## 方案
集群内 CronJobns `plane`,每天 **01:30 UTC = 03:30 本地**,控制器按 UTC 跑):
1. 单容器 `postgres:15.7-alpine``pg_dump -Fc`(自定义压缩格式)打 `plane`
`/var/backups/plane/pg/plane-<UTC时间戳>.dump`hostPath `DirectoryOrCreate`
2. 保留 7 天(`find -mtime +7 -delete`),成功/失败历史各留 3/2
3. 凭据:现 chart Secret `plane-app-pgdb-secrets`Phase A 外部化后改 `plane-pgdb-credentials`
## 容量
- 库 88 MB → `-Fc` 快照约 10–40 MB/天 × 7 天 ≈ **<300 MB**,对 83 G 可用盘可忽略。
## 还原(未演练;应用前先做一次隔离测试)
```bash
# 目标 PG15 实例(临时起一个 postgres:15.7-alpine 容器或另一台机):
# 先建空库: createdb plane (user=plane)
pg_restore -h <target> -U plane -d plane --clean --if-exists /var/backups/plane/pg/plane-<TS>.dump
# 还原后确认 310 issues 量级一致;附件为空属预期(未备份 MinIO)
```
## 验收(应用前逐项过)
- [ ] CronJob 建立后手动触发一次:`kubectl -n plane create job --from=cronjob/plane-backup plane-backup-manual-1`Job `Completed`
- [ ] `/var/backups/plane/pg/plane-*.dump` 可被 `pg_restore -l` 列出
- [ ] 备份 Job 只依赖 pgdb 服务,不依赖 Plane 应用 Pod(应用故障期间也能出备份)
- [ ] 保留清理 dry-run`find ... -print`)正确;`df -h /` 前后对比记录
## 局限 / 阻塞
- **本地方案不是离机备份**:单节点磁盘/整机故障即丢。如日后要离机,纳入
[Restic 异机 repository 决策与存取隔离](https://plane.chans.xyz/space/projects/56874283-7e1d-43a8-afa4-631cf1c4ad5b/issues/7825d564-ae15-446b-bced-be26b648346b/)
(与 Matrix 备份同一决策);恢复演练纪律见
[服务级 restore runbook 与隔离复元演练](https://plane.chans.xyz/space/projects/56874283-7e1d-43a8-afa4-631cf1c4ad5b/issues/a9bea3ba-c958-4a74-b2f1-6bbb653f21d3/)。
- 提醒:同一节点 **Matrix 数据价值远高于 Plane 且同样无备份** —— 若投入备份精力,顺序上 Matrix 优先。
@@ -0,0 +1,63 @@
# Plane CE PostgreSQL backup CronJob — DRAFT (2026-09-03), NOT applied.
# ns: plane (synapse K3s single node). Output: hostPath /var/backups/plane (root disk, auto-created).
#
# Scope decision (2026-09-03, practical): DB-only. MinIO dropped — uploads bucket
# measured at 264 KB / 444 KB total; attachments are acceptable loss for this
# personal 1-user instance (310 issues / 88 MB DB). Revisit only if usage grows.
#
# Credentials: read from the CURRENT chart-generated Secret (works today). After the
# optional external-secrets migration (docs/plane-hardening/README.md Phase A) switch
# the secretKeyRef name to plane-pgdb-credentials.
#
# Apply:
# ssh windy@synapse.chans.xyz 'sudo k3s kubectl apply -n plane -f -' < plane-backup.yaml
# Manual run + verify:
# sudo k3s kubectl -n plane create job --from=cronjob/plane-backup plane-backup-manual-1
# sudo k3s kubectl -n plane get cronjob,job,pods | grep plane-backup
# sudo ls -lh /var/backups/plane/pg
# Restore steps + tuning: see backup/README.md
apiVersion: batch/v1
kind: CronJob
metadata:
name: plane-backup
namespace: plane
spec:
# 01:30 UTC daily = 03:30 local (CEST). CronJob controller runs in UTC.
schedule: "30 1 * * *"
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 2
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: OnFailure
volumes:
- name: backup
hostPath:
path: /var/backups/plane
type: DirectoryOrCreate
containers:
- name: pg-dump
image: postgres:15.7-alpine
env:
- name: PGPASSWORD
valueFrom:
secretKeyRef:
name: plane-app-pgdb-secrets # -> plane-pgdb-credentials after Phase A
key: POSTGRES_PASSWORD
command: ["/bin/sh", "-c"]
args:
- |
set -euo pipefail
TS=$(date -u +%Y%m%dT%H%M%SZ)
mkdir -p /backup/pg
pg_dump -h plane-app-pgdb.plane.svc.cluster.local -U plane -d plane \
-Fc -f "/backup/pg/plane-${TS}.dump"
find /backup/pg -type f -name 'plane-*.dump' -mtime +7 -delete
echo "pg_dump done: /backup/pg/plane-${TS}.dump ($(du -h /backup/pg/plane-${TS}.dump | cut -f1))"
volumeMounts:
- name: backup
mountPath: /backup
+91
View File
@@ -0,0 +1,91 @@
# External Secret structure for Plane CE hardening — EXAMPLE ONLY.
# No real values here; this file is safe to commit. Real values live only on the
# host (/home/windy/plane-k3s, 0600/0700) and in the cluster.
#
# Phase A — create each Secret with the CURRENT cluster values first (zero change):
# # current source Secrets (chart-generated):
# kubectl -n plane get secret plane-app-app-secrets -o jsonpath='{.data.SECRET_KEY}' | base64 -d
# kubectl -n plane get secret plane-app-live-secrets -o jsonpath='{.data.REDIS_URL}' | base64 -d
# kubectl -n plane get secret plane-app-pgdb-secrets -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d
# kubectl -n plane get secret plane-app-rabbitmq-secrets -o jsonpath='{.data.RABBITMQ_DEFAULT_PASS}' | base64 -d
# kubectl -n plane get secret plane-app-doc-store-secrets -o jsonpath='{.data}' | base64 -d
#
# e.g. kubectl -n plane create secret generic plane-app-keys \
# --from-literal=SECRET_KEY="$(<copy from above>)" \
# --from-literal=LIVE_SERVER_SECRET_KEY="$(<copy from above>)"
#
# All keys below are REQUIRED by chart templates/plane-ce-1.8.0 (verified 2026-09-03):
# missing keys are NOT auto-filled once an existingSecret is referenced.
---
apiVersion: v1
kind: Secret
metadata:
name: plane-app-keys # external_secrets.app_keys_existingSecret
namespace: plane
type: Opaque
stringData:
SECRET_KEY: "" # current: copy from plane-app-app-secrets; rotate only in Phase B
LIVE_SERVER_SECRET_KEY: "" # current: same value as above / plane-app-live-secrets
---
apiVersion: v1
kind: Secret
metadata:
name: plane-app-env # external_secrets.app_env_existingSecret
namespace: plane
type: Opaque
stringData:
REDIS_URL: "" # redis://plane-app-redis.plane.svc.cluster.local:6379/
DATABASE_URL: "" # postgresql://plane:plane@plane-app-pgdb.plane.svc.cluster.local/plane
AMQP_URL: "" # amqp://plane:plane@plane-app-rabbitmq.plane.svc.cluster.local/
---
apiVersion: v1
kind: Secret
metadata:
name: plane-live-env # external_secrets.live_env_existingSecret
namespace: plane
type: Opaque
stringData:
REDIS_URL: "" # redis://plane-app-redis.plane.svc.cluster.local:6379/
---
apiVersion: v1
kind: Secret
metadata:
name: plane-pgdb-credentials # external_secrets.pgdb_existingSecret
namespace: plane
type: Opaque
stringData:
POSTGRES_PASSWORD: "" # Phase A: keep current ('plane'); Phase B: ALTER USER first, then sync
POSTGRES_DB: "plane"
POSTGRES_USER: "plane"
---
apiVersion: v1
kind: Secret
metadata:
name: plane-rabbitmq-credentials # external_secrets.rabbitmq_existingSecret
namespace: plane
type: Opaque
stringData:
RABBITMQ_DEFAULT_USER: "plane"
RABBITMQ_DEFAULT_PASS: "" # Phase A: keep current; Phase B: rabbitmqctl change_password first
---
apiVersion: v1
kind: Secret
metadata:
name: plane-minio-credentials # external_secrets.doc_store_existingSecret
namespace: plane
type: Opaque
stringData:
FILE_SIZE_LIMIT: "20971520" # must match env.doc_upload_size_limit
AWS_S3_BUCKET_NAME: "uploads"
USE_MINIO: "1"
MINIO_ROOT_USER: "admin"
MINIO_ROOT_PASSWORD: "" # root creds take effect on first init only
AWS_ACCESS_KEY_ID: "admin"
AWS_SECRET_ACCESS_KEY: "" # == MINIO_ROOT_PASSWORD while minio.local_setup
AWS_S3_ENDPOINT_URL: "http://plane-app-minio:9000"
+109
View File
@@ -0,0 +1,109 @@
# Plane CE hardened values — DRAFT (2026-09-03), NOT applied.
# Target file on host: /home/windy/plane-k3s/values.yaml (synapse.chans.xyz)
# Reference release: plane-app, chart plane-ce-1.8.0 (values.yaml L1-362 + templates verified 2026-09-03).
# No secrets in this file. Secret *values* live only in k8s Secrets (see secrets.yaml.example).
#
# Two phases:
# Phase A: externalize secrets (reference names below) with CURRENT values copied -> zero change.
# Phase B: rotate credentials one by one (see README.md). SECRET_KEY rotation is cheap only while
# SMTP/OAuth are unconfigured (no encrypted config rows yet).
planeVersion: v1.4.1
ingress:
enabled: true
appHost: plane.chans.xyz
ingressClass: traefik
traefik:
# 20 MiB (chart default). Keep aligned with env.doc_upload_size_limit below.
maxRequestBodyBytes: 20971520
ssl:
createIssuer: true
issuer: http # HTTP-01; ssl_token_existingSecret not needed
email: admin@chans.xyz
generateCerts: true
postgres:
storageClass: local-path
volumeSize: 5Gi
# NOTE: chart 1.8.0 exposes NO resources knob for the bundled datastores
# (stateful templates render no resources block). Add limits via
# --post-renderer/kustomize or `kubectl -n plane patch sts ...` re-applied on
# every upgrade (P2 task; see README.md).
redis:
storageClass: local-path
# image: valkey/valkey:7.2.11-alpine # already pinned by chart default; uncomment to make explicit
minio:
# P2: pin. Digest of the currently running :latest (2026-09-03, pod plane-app-minio-wl-0).
image: minio/minio@sha256:14cea493d9a34af32f524e538b8346cf79f3321eff8e708c1e2960462bd8936e
# image_mc: minio/mc@sha256:... # optional: pin one-shot bucket-init client the same way
storageClass: local-path
volumeSize: 5Gi
rabbitmq:
storageClass: local-path
env:
# Fail the render instead of ever falling back to the chart's PUBLIC constants
# (values.yaml L340-341 in chart 1.8.0). Requires external_secrets below.
requireExplicitSecrets: true
# SECRET_KEY / LIVE_SERVER_SECRET_KEY are deliberately OMITTED here.
# They live in k8s Secret `plane-app-keys` (referenced below). With
# requireExplicitSecrets=true and app_keys_existingSecret set, the chart renders
# neither key itself and app+live workloads both envFrom `plane-app-keys` LAST
# (later envFrom wins), which keeps the shared signing key consistent.
pgdb_name: plane
docstore_bucket: uploads
# Align app-side upload cap with the Traefik body limit (was 5242880/5MiB).
# Keep both at 20MiB, or lower both together.
doc_upload_size_limit: "20971520"
external_secrets:
# Shared signing keys (used by app + live). REQUIRED keys: SECRET_KEY, LIVE_SERVER_SECRET_KEY.
app_keys_existingSecret: plane-app-keys
# REQUIRED keys: REDIS_URL, DATABASE_URL, AMQP_URL (chart-derived URLs; update on DB/queue rotation).
app_env_existingSecret: plane-app-env
# REQUIRED keys: REDIS_URL.
live_env_existingSecret: plane-live-env
# REQUIRED keys: POSTGRES_PASSWORD, POSTGRES_DB, POSTGRES_USER.
pgdb_existingSecret: plane-pgdb-credentials
# REQUIRED keys: RABBITMQ_DEFAULT_USER, RABBITMQ_DEFAULT_PASS.
rabbitmq_existingSecret: plane-rabbitmq-credentials
# REQUIRED keys: FILE_SIZE_LIMIT, AWS_S3_BUCKET_NAME, USE_MINIO, MINIO_ROOT_USER,
# MINIO_ROOT_PASSWORD, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_S3_ENDPOINT_URL.
doc_store_existingSecret: plane-minio-credentials
# ssl_token_existingSecret: '' # DNS-01 only (cloudflare/digitalocean); unused with HTTP-01
# Optional, P2: HTTP -> HTTPS 301. The chart's own IngressRoute binds only
# 'websecure' (http:// currently 404s). extraObjects is rendered verbatim (toYaml).
# Uncomment and `helm upgrade` once reviewed:
# extraObjects:
# - apiVersion: traefik.io/v1alpha1
# kind: Middleware
# metadata:
# name: plane-https-redirect
# namespace: plane
# spec:
# redirectScheme:
# scheme: https
# permanent: true
# - apiVersion: traefik.io/v1alpha1
# kind: IngressRoute
# metadata:
# name: plane-http-to-https
# namespace: plane
# spec:
# entryPoints: [web]
# routes:
# - match: Host(`plane.chans.xyz`)
# kind: Rule
# middlewares:
# - name: plane-https-redirect
# services:
# - name: plane-app-web
# port: 3000
@@ -1,158 +0,0 @@
# 希力威视 SR-S25G3218F 调查(2026-08-08
**结论:** 若需求是大量 2.5G 终端、少量 10G 光上联,`SR-S25G3218F` 的端口密度
更合适;厂商已公开该型号的固件页,但仍缺少完整规格书、管理手册与兼容矩阵。若需求是 8 条全部可协商
1/2.5/5/10G 的铜缆链路,且希望有可查的 L3 能力和固件入口,兮克
`SKS8300-8T` 是资料更完整、风险更低的选择;它的代价是主动风扇、外置 12 V 电源、
无 SFP+ 光口,且仍不应把消费级/SMB 设备当作安全边界或唯一核心。两者都应在
到货可退换期内完成实机验收。
本页为采购前资料调查,不代表已接入本地网络;检索日期为 2026-08-08。
## 已能核实的事项
| 项目 | 结论与证据强度 |
|---|---|
| 型号/端口 | 京东的希力威视商品标题称该 SKU 为 `SR-S25G3218F`,有 16 个 2.5G 电口和 2 个万兆光口,并宣传 VLAN、端口隔离与 LACP。该店铺被厂商官网列为可购买的「京东旗舰店」,因此可作为销售规格,非技术手册。[京东商品页](https://item.jd.com/100165071727.html)[厂商购买渠道说明](https://en.sirivision.com/contactus/) |
| 厂商身份 | 厂商官网为 Shenzhen/Guangdong Sirivision Communication;英文官网说明其自 2016 年起提供接入、汇聚和核心交换机方案。[厂商首页](https://en.sirivision.com/) |
| 公开的二手厂家资料 | 同一制造商名义的 Alibaba 出口页将精确型号写成 `16*2.5G+2*10G``120Gbps`,并列出 QoS、VLAN、SNMP、L3 与 stackable。这是制造商发布在平台上的销售资料,**不是**官网数据表;其中后五项不能据此视为已验收的功能承诺。[制造商平台页](https://www.alibaba.com/pla/SR-S25G3218F-QoS-Managed-SFP-Switch-1625G210G_1601494946214.html) |
| 固件入口 | 厂商已发布此精确型号的[固件页](https://www.sirivision.com/sr-s25g3218f%E5%9B%BA%E4%BB%B6/)。公开变更记录提到“光口自适应”和“增加 DAC 配置”;这证明厂商维护过该路径,**不**代表任意 SFP+/DAC/铜模块均兼容。 |
| 本机可计算的带宽 | 端口线速相加为单向 60 Gb/s16 × 2.5 + 2 × 10);若厂商所谓 `120Gbps` 是全双工交换容量,则数学上吻合。它**不**证明缓冲、PPS、表项规模或实际无阻塞性能。 |
## 网管/L2/L3 能力边界
京东标题足以支持把 VLAN、端口隔离、LACP 作为「卖家声称提供」的功能;不得由此推导出
ACL、IPv4/IPv6 静态路由、SVI 数量、DHCP relay、OSPF/RIP、VRRP、IGMP、ERPS、
802.1X、RADIUS/TACACS+、SSH/HTTPS 管理、SNMP 版本、日志/审计、配置备份或固件
安全维护一定存在。
尤其要注意:厂商官网把真正列出的 2.5G L3 产品标为
`SR-S25G3412F (8 × 2.5G + 4 × 10G SFP+)`;其 2.5G 类目只显示 7 个型号,
不含 `SR-S25G3218F`。官网也把 L2+、Web Smart、L3 分成不同产品类别。这个目录
差异**不是**证明 3218F 没有 L3,而是说明「三层」无法通过官网的精确型号文档确认。
[2.5G 产品目录](https://en.sirivision.com/product-category/products/2-5g-switches/)
[官网的 10G L3 目录](https://en.sirivision.com/product-category/products/10g-switches/10g-layer3-managed-switches/)
[官网的 L2+ 分类示例](https://en.sirivision.com/product-category/products/gigabit-switches/gigabit-layer2-managed-switches/)。
采购前请向京东/厂商索取**与机身 SKU、硬件 revision 和固件版本对应**的 PDF
数据表、管理手册和 release notes,并要求书面回答至少以下问题:
1. L3 是只有 VLAN Interface/IPv4 静态路由,还是另有 IPv6、ACL、动态路由、DHCP relay
等;每项的最大 VLAN、MAC、ARP、路由、ACL、LAG 数量分别是多少?
2. LACP 是否符合 802.3ad、一个 LAG 最多多少成员、能否跨两台设备(若销售页的
`stackable` 属实,堆叠的线缆/模块、最大成员、控制面和软件版本为何)?
3. 管理面是否支持 HTTPS/SSH、禁用 HTTP/Telnet、独立管理 VLAN、SNMPv3、syslog、NTP、
配置导出/回滚和已签名或可校验的固件;默认凭据首次登录是否强制修改?
## 供电、散热和光口:当前不能确认
针对该精确 SKU,厂商官网目录与公开搜索未找到说明书/数据表,所以以下均为**待确认,
不能猜测**
- 是否为内置 AC 电源、额定输入范围/最大功耗、是否带电源开关和接地端子;是否完全
不提供 PoE(本型号名和京东标题均未写 PoE,但这不足以替代规格书)。
- 风扇数量、常态/满载噪声、风向、环境温湿度、机架深度与安装耳;不要将「金属壳」
或产品照片等同于无风扇/静音。
- 两个槽是否均为 **10G SFP+**,是否可协商 1G SFP;支持的 SR/LR/BiDi 波长距离、
DAC/AOC 长度、第三方模块/EERPOM 兼容策略、10GBASE-T SFP+ 模块的功耗/温度限制,
以及是否支持 GPON/XPON ONU「猫棒」。
厂商确实单列「SFP Optical Modules」产品分类,但这不构成 3218F 的兼容清单。
[厂商产品导航](https://en.sirivision.com/)。购买光模块/直连线时,应要求厂商按这台
设备的硬件/固件 revision 出具兼容型号清单;没有书面清单时,先在可退换期实测两端的
链路、重启恢复、热插拔与长时间满载错误计数。
## 风险与建议验收
- **文档/生命周期风险(中到高):** 精确型号不在厂商当前官网 2.5G 目录,虽有固件下载页,
但未公开完整型号手册、明确 release notes 或兼容矩阵。官网的售后条款也要求按具体产品查询保修期,配件(含光纤头)
的保修条款与主机不同;不要把平台页的「3 年」当作中国零售 SKU 的已确认保修。
[厂商售后条款](https://en.sirivision.com/after-sale-protection/)
- **功能表述风险(高):** 页面将 L2 特性和「三层网管」并列;在命令/网页菜单、
手册和测试证明之前,将其当作 L2 VLAN/LACP 设备部署,跨 VLAN 路由仍由现有网关承担。
- **双 10G 上联约束(中):** 两个 SFP+ 可作双上联或一个二成员 LAG,但 LAG 增加的是
多流量总吞吐,单一 TCP/UDP 流通常仍受一条 10G 链路限制;上级设备也必须匹配 LACP
配置。
- **管理面风险(中到高):** 家用/低价网管设备常见明文管理、弱默认口令或不透明的固件
更新周期;采购后先置于受限管理 VLAN,改口令、升级已验证固件,且不将管理界面暴露
到 WAN/访客网。
最低验收应包括:逐口协商 100M/1G/2.5G、两只不同厂家 SFP+/DAC(仅在卖家承诺支持的
范围内)、VLAN trunk/access/PVID、STP/环路保护、LACP 故障切换、端口隔离、满载
双向 iperf3 与错误计数、冷启动后的配置保留,以及管理面的 HTTPS/SSH/SNMPv3/配置备份。
如无法提供与型号匹配的正式资料或其中任一关键项失败,应在退换期内退货,并选择公开
数据表、固件与兼容矩阵更完整的型号。
## 备选:兮克 SKS8300-8T 对比
### 已核实的厂商规格
兮克官网的精确型号页明确将 `SKS8300-8T` 定位为三层管理型 10G 全电口交换机,并列出:
- 8 × 1/2.5/5/10GBASE-T RJ45160 Gb/s 交换容量、119.05 Mpps、12 Mbit 缓存、
16K MAC、12 KB 巨帧、512 MB DRAM、32 MB Flash,尺寸 207 × 136 × 35 mm
- QoS、ACL、IP+MAC+端口绑定、流分类/优先级标记、多端口镜像、静态/灵活 QinQ、
sFlow,以及「基于策略的 IPv4/IPv6 单播路由」。
这些是厂商能力声明,并非对每一种路由协议或表项上限的承诺;但相对 3218F 的仅有
销售标题,它给出了精确型号、转发性能和 L3 范围。[兮克 SKS8300-8T
产品页](https://seekswan.com/user/custom-pages/SKS8300-8T.html)
独立的 OpenWrt 设备资料将其识别为 Realtek RTL9303、512 MB RAM,记录了原厂固件
下载入口和串口/TFTP 恢复路径;其硬件数据页列为 12 V / 4 A。这支持「可恢复、可替换
系统」的可操作性,但**不是**兮克对原厂功能的支持承诺。
[OpenWrt 设备页](https://openwrt.org/toh/xikestor/sks8300-8t)
[OpenWrt 硬件数据](https://openwrt.org/toh/hwdata/xikestor/xikestor_sks8300-8t)。
### 能力、物理与运维比较
| 维度 | 希力威视 SR-S25G3218F | 兮克 SKS8300-8T |
|---|---|---|
| 接口/典型用途 | 16 × 2.5G 电口 + 2 × 10G SFP+(销售规格);适合很多 2.5G 终端/NAS,以 10G 光或 DAC 上联。 | 8 × 1/2.5/5/10GBASE-T;适合 10G 铜缆设备、2.5/5G 多速率 NAS/主机。没有 SFP+,光纤上联必须经媒体转换或选另一型号。 |
| 可确认的三层范围 | 仅销售/平台资料称 L3;没有精确型号官方手册,不能确认静态路由以外的功能。 | 官网明确写策略型 IPv4/IPv6 单播路由、ACL/QoS/sFlow/QinQ;动态路由、VRRP、IPv6 ACL/SNMP/认证等仍须按当前固件手册确认。 |
| 冗余/二层 | 卖家声称 VLAN、端口隔离、LACP;STP/环网的实现与规格未知。 | 官网声明 L3 和多项转发特性,但未在产品页给出 STP/LACP/ERPS 的精确限制;购买前仍索取手册。 |
| 散热/噪声 | 无可核实的精确型号风扇、噪声、功耗或风向数据。 | 独立手册镜像和产品图均称智能温控风扇,但厂商产品页未给 dBA;应按「有风扇、可能听得见」规划,不能承诺静音。 |
| 供电 | 未找到精确型号官方输入/功耗资料。 | OpenWrt 硬件数据记录 12 V / 4 A;确认随附电源适配器的插头、余量和地区认证。官方产品页未给满载功耗。 |
| 固件/恢复 | 有精确型号官方固件页;公开记录包含光口自适应与 DAC 配置改动,但未找到完整 release notes、恢复步骤或兼容矩阵。 | 厂商产品页提供「相关下载」区,OpenWrt 还记录原厂固件入口、RJ45 串口和 U-Boot/TFTP 恢复;原厂镜像是否签名、漏洞修复 SLA、配置回退仍未知。 |
关于 8T 的风扇、满载功耗(常见转述为 ≤36 W)、温度范围、芯片型号等,本次未找到
相应的**厂商原始数据表**;不将第三方手册转录当作已核实规格。若噪声、UPS 容量或
机柜散热是购买约束,请先让卖家提供产品铭牌照片、适配器铭牌照片、额定/实测功耗和
dBA 测试条件。
### 选择与验收建议
-**3218F**:必须有 ≥12 个 2.5G 接入端、10G 光/DAC 上联、且 L3 留给现有路由器。
下单前先取得精确型号手册和 SFP+/DAC 兼容承诺;否则端口数量优势不足以抵消资料风险。
-**8T**:最多 8 个设备但需要多速率 10G RJ45、明确的 IPv4/IPv6 静态/策略路由和
以后自行维护/恢复的余地。不要把其 160 Gb/s 标称交换容量误解为 8 端口同时 10G
全双工的性能保证——该标称与端口总线速数学相等,但仍须以实测和厂商 PPS/缓冲说明为准。
- 两台都不应单独承担防火墙、访客/IoT 安全隔离或 WAN 暴露;VLAN 的跨网段策略和公网
边界留在受支持的网关/防火墙上。先为管理面创建专用 VLAN,仅从管理主机访问,禁用
未使用的远程管理协议,备份配置和原厂固件后再接入生产网络。
## 低功耗核心备选(8 × 2.5G + 2 × SFP+
如果核心只需接最多 8 台铜缆终端、上联/连接 NAS 使用 DAC 或光纤 10G,优先考虑没有
PoE 的以下两款。它们都满足 VLAN trunk、LACP 和至少两个 10G SFP+ 的需求;不要为
AP 选 PoE 版来承担核心,因为 PoE 预算、风扇和待机损耗都会明显增加。
| 型号 | 端口与管理能力(厂商声明) | 厂商功耗 / 噪声资料 | 对当前 LAN 的判断 |
|---|---|---|---|
| **TP-Link Omada SG3210X-M2** | 8 × 100M/1G/2.5G RJ45、2 × 10G SFP+,并有 RJ45 和 Micro-USB console。厂商规格列出 802.1Q VLAN、STP/RSTP/MSTP、静态 LAG 和 802.3ad LACP(最多 8 个聚合组、每组最多 8 端口);L3 是 32 个 IPv4/IPv6 接口、48 条静态路由。 | **无风扇**100240 V AC 内置电源。`UN 1.20` 数据表:待机最高 **6.0 W**220 V/50 Hz、25 °C),最高 **15.3 W**220 V)或 **15.0 W**110 V)。 | **首选低功耗方案。** 足以做 LAN66 核心、给 PVE/gfw 与 U6 Lite 做 VLAN 10 trunk,并以 SFP+ DAC/光口连接 10G NAS/主机;它不提供 5G/10G RJ4510G 铜缆需外置转换或 SFP+ 10GBASE-T 模块。 |
| **MikroTik CRS310-8G+2S+IN** | 8 × 2.5G RJ45、2 × 10G SFP+SFP+ 笼支持 1G/2.5G/10G。RouterOS v7(也可选 SwOS)支持 VLAN、链路聚合与 ACL。 | 18–57 V DC 外置供电;官方给出“无附件”最高 **21 W**、总体最高 **34 W**,且机内 **1 个风扇**。厂商没有在该页给出 dBA。 | 可用且软件/文档/恢复路径成熟,但不是本题的静音低功耗优先项:官方最大功耗显著高于 TP-Link,且有风扇。适合明确偏好 RouterOS/SwOS 与其可维护性时选。 |
功耗数字是各厂商的**上限/待机测试条件**,不是你实际墙插读数;SFP+ 光模块、DAC/AOC,尤其
10GBASE-T SFP+ 模块,会另增功耗和热量。对于本网络,用被动 DAC 或短距光模块连接 10G
设备,通常比全 RJ45 10G 核心更容易保持低温、低噪。
`SG3210X-M2` 的上表数据对应 TP-Link 的 `UN 1.20` 数据表;不同地区/硬件版本的包装、
认证和功耗标注可能不同,购买中国零售版本前应让卖家确认**准确硬件版本、保修渠道和固件地区**。
本次未找到 TP-Link 中国官网的该精确型号页,因此不能把海外官方页面当作大陆现货/售后承诺。
MikroTik 同样应通过其官方零售商查询渠道确认本地库存和保修。两台购买前还应确认所选
SFP+/DAC 的兼容清单。
来源:[TP-Link 产品规格](https://www.tp-link.com/uk/business-networking/omada-switch-access-pro/sg3210x-m2/)
[TP-Link `UN 1.20` 数据表](https://static.tp-link.com/upload/product-overview/2025/202512/20251224/SG3210X-M2%28UN%29%201.20_datasheet.pdf)
[MikroTik 产品页](https://mikrotik.com/product/crs310_8g_2s_in)
[MikroTik 用户手册](https://help.mikrotik.com/docs/spaces/UM/pages/214630429/CRS310-8G%2B2S%2BIN)。
+64 -3
View File
@@ -68,6 +68,66 @@ db.device.find(
).pretty()
```
## IPv6 status (verified 2026-08-20)
IPv6 is **enabled and live** on the main Wi-Fi networks. Read-only
verification, no changes made.
**Controller (`networkconf` in the `ace` DB):** the `Default` LAN network has
`ipv6_enabled: true`, `ipv6_client_address_assignment: slaac`,
`ipv6_ra_enabled: true`, `ipv6_ra_priority: high`, and
`dhcpdv6_allow_slaac: true`. `ipv6_interface_type: "none"` is expected: the
network's gateway is the third-party EdgeRouter (`gw`), so the controller does
not manage WAN-side IPv6 — RA/SLAAC is served by the router.
All active SSIDs map to the `Default` network: `ubnt-windy` (5G),
`ubnt-windy-2` (2.4G), `ubnt-haas` (2.4G) — clients on them receive SLAAC IPv6.
Exception: the dormant `ubnt-upg` VLAN 10 network (and its `ubnt-upg` SSID) has
no IPv6 configuration (default off). See
[Dedicated Wi-Fi through a third-party gateway](#dedicated-wi-fi-through-a-third-party-gateway).
**APs (live):** both managed APs hold global SLAAC addresses on `br0` with a
default route learned via RA from `gw`:
| AP | Global IPv6 on `br0` (at check time) | Default route |
|---|---|---|
| U6 Lite (`192.168.66.6`) | `240e:3bd:235:1fb1:...`/64 | `default via fe80::... dev br0 proto ra` |
| UAP-AC-Lite (`192.168.55.5`) | `240e:3bd:235:1fb2:...`/64 | `default via fe80::... dev br0 proto ra` |
The delegated prefixes are dynamic ISP allocations (PPPoE PD `/60`) and rotate
on redial; only the structure is stable.
**Gateway (`gw`):** the IPv6 routing table shows connected `/64`s on `eth0`
(LAN66) and `switch0` (LAN55) plus `::/0` via `pppoe0`.
Re-verify:
```bash
ssh -4 -o BatchMode=yes zhiqiangf@192.168.66.6 'ip -6 addr show br0; ip -6 route show'
ssh -4 -o BatchMode=yes zhiqiangf@192.168.55.5 'ip -6 addr show br0; ip -6 route show'
```
> **2026-08-21 (W1N-207):** SmartThings Element/vWire provisioning SSIDs
> (`element-8a0d5133c9438f12`, `vwire-8b2d67469e455785`, `vport-F09FC22004E9`) were
> removed (element_adopt setting disabled + element wlanconf deleted + device vwire
> fields cleared) and confirmed off on both APs (normal SSIDs unchanged: `ubnt-windy`,
> `ubnt-windy-2`, `ubnt-haas`, `ubnt-upg`). Root cause of Matter onboarding failure that
> day: HA's matter-server advertised a stale IPv6 GUA (two prefix generations old) in
> mDNS; fixed by restarting the add-on — see [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md).
>
> **SSID ↔ subnet split (Matter-relevant):** `ubnt-windy` (5G) is served only by the
> U6 Lite on LAN66; `ubnt-haas` / `ubnt-windy-2` (2.4G) only by the UAP-AC-Lite on
> LAN55. mDNS is link-local multicast and does **not** cross the routed 55/66
> boundary (no mDNS reflector). Matter commissioning therefore requires phone and
> device on the **same subnet (LAN55)**; a phone on 5G (LAN66) cannot discover a
> LAN55 Matter device.
>
> **Cleanup side-effects (left as-is, harmless):** after the direct-DB cleanup,
> `db.device.cfgversion` holds placeholder values (`0000000000000000` /
> `1111111111111111`) and UAP-AC-Lite has `mesh_sta_vap_enabled=false`; the
> controller has not reverted them and no functional impact was observed.
## Dedicated Wi-Fi through a third-party gateway
### Architecture boundary discovered on 2026-08-08
@@ -151,9 +211,10 @@ Use `ssh zhiqiangf@AP_IP` for the adopted-device account. Do not query or copy
the controller's `mgmt` database setting into logs or documentation: it can
contain the managed SSH password.
On 2026-08-06, key-only IPv4 SSH was verified for both managed APs using the
`zhiqiangf` account. Verify future access without permitting password or
keyboard-interactive fallback:
Key-only IPv4 SSH was verified for both managed APs using the `zhiqiangf`
account on 2026-08-06 and re-verified 2026-08-20 (BatchMode with password and
keyboard-interactive disabled; both APs still log in key-only). Verify future
access without permitting password or keyboard-interactive fallback:
```bash
ssh -4 -o BatchMode=yes -o PasswordAuthentication=no \
+13 -4
View File
@@ -32,18 +32,27 @@ Do not enable AdGuard Home DHCP unless the existing EdgeRouter DHCP service is
explicitly migrated and disabled first.
`agh-ui-access.service` loads `/etc/nftables-agh-ui-access.nft`. It permits
only `192.168.66.0/24` to TCP/80 and drops other TCP/80 input. It deliberately
`192.168.66.0/24` (LAN66) and `192.168.55.0/24` (LAN55, for Home Assistant
integration) to TCP/80 and drops other TCP/80 input. It deliberately
does **not** restrict DNS, SSH, Docker, or Mihomo ports. Keep it isolated from
Docker-managed nftables tables.
> 2026-08-12: added `192.168.55.0/24` allow so HAOS (`192.168.55.11`) can reach
the HTTP API on `:80` for the Home Assistant AdGuard Home integration; applied
via `sudo systemctl restart agh-ui-access.service` (file edited first, then
reloaded; syntax verified with `nft -c`). Other firewalls (EdgeRouter LAN_IN/
LAN_OUT inactive, PVE zero rules) were already open for LAN55->LAN66.
Current query-log policy is 14 days with anonymized client IPs. Check free
space before increasing retention. DNSSEC is disabled because the selected
upstream path did not pass the known-bad-signature validation check; do not
enable it without re-testing validated upstreams.
The compatible names `hass.windy.lan` and legacy `hass.local` currently point
to the same Home Assistant address. Migrate clients to `hass.windy.lan`; keep
the legacy rewrite until its planned retirement.
`hass.windy.lan` points to Home Assistant via this rewrite. The legacy
`hass.local` rewrite was removed on 2026-08-14; `hass.local` now resolves only
via HAOS mDNS/LLMNR (`hostname: hass`), not via AdGuard Home. The
`nas.windy.local` rewrite was likewise removed on 2026-08-14; `.local` names
are now left to mDNS only. Remaining rewrites all use `.windy.lan`.
## Mihomo and routing boundary
+58 -15
View File
@@ -8,7 +8,7 @@
| IPv4 | `192.168.66.1` |
| SSH | `ssh -4 root@192.168.66.1` (key-only, verified 2026-08-06) |
| OS | ImmortalWrt 25.12.0 (r37854), Linux `6.12.87`, x86/64 |
| **Host** | **PVE VM 140 (`gfw`)**dual NIC: `net0`→vmbr0(LAN66), `net1`→vmbr1(LAN55) (verified 2026-08-09) |
| **Host** | **PVE VM 140 (`gfw`)**3 NICs: `net0`→vmbr0(LAN66/eth0), `net1`→vmbr1(LAN55/eth1, up but unaddressed), `net2`→VLAN10/`ubunt_upg`(eth2, `192.168.10.1/24`) (topology 2026-08-09; eth2/VLAN10 live verified 2026-08-11) |
Do not store the root password in this repository.
@@ -25,8 +25,49 @@ OpenClash runs `/etc/openclash/clash` (clash_meta core) with configuration
- Mode: **fake-ip + TPROXY transparent proxy** (`operation_mode=fake-ip`,
`en_mode=fake-ip`, `proxy_mode=rule`); fake-ip network `198.18.0.0/16`
- DNS path: dnsmasq → clash DNS `127.0.0.1#7874` (`server=127.0.0.1#7874` in
dnsmasq config); `mosdns` also listens on `127.0.0.1:6052` but is not the
active resolver path
dnsmasq config); OpenClash custom DNS uses `mosdns` on `127.0.0.1:6052` as its
`nameserver`/`default-nameserver` for DIRECT-rule real-IP resolution
(`/etc/mosdns/config.yaml`): domestic domains → AGH `.36:53`, foreign →
`223.5.5.5`/`119.29.29.29` (Chinese public DNS). mosdns is **not** in the
client query path — LAN/VLAN10 clients receive fake-ip from clash :7874.
> 2026-08-12: fixed missing `has_resp → accept` guard after the domestic
> branch in `/etc/mosdns/config.yaml` (domestic queries were double-forwarded,
> final answer came from CN public DNS, bypassing AGH blocking/rewrites;
> verified via `dup.baidustatic.com` before/after); added `domestic_fallback`
> (fallback plugin: primary=AGH, secondary=CN public DNS, 500ms) so domestic
> DIRECT lookups survive an AGH outage. Backups:
> `config.yaml.bak-20260812` / `config.yaml.bak-fallback-20260812`. See
> [docs/lan-dns-architecture.md](../docs/lan-dns-architecture.md) §1.
> 2026-08-13 (W1N-62): foreign branch now uses encrypted DoH
> `https://adg.chans.xyz/dns-query` (self-hosted, hk2) via new
> `foreign_upstream` / `foreign_fallback` plugins; non-CN queries → DoH,
> falls back to CN public DNS after 1000ms. `bootstrap` = existing CN public
> DNS IPs (no self-loop). Live-verified: google/youtube real IP + AAAA
> restored (2607:f8b0…), `dup.baidustatic.com` → `0.0.0.0` (AGH intercept
> kept), clash 7874 fake-ip plane unchanged. **Final decision (2026-08-13):
> DoH goes DIRECT to hk2, not via clash proxy** — `foreign_upstream` points
> only at the self-hosted resolver `adg.chans.xyz` (hk2), which is directly
> reachable and already encrypted (DoH/TLS) with clean answers, so forcing
> the proxy adds nothing and would couple the DNS plane to clash (nft output
> chains also show OpenClash does not currently redirect router-own TCP).
> Kill-test: foreign queries answered during clash outage, watchdog
> auto-restarted. Backups: `config.yaml.bak-foreign-doh-20260813-103746` /
> `config.yaml.bak-foreign-doh-20260813-103813`.
> 2026-08-13 (W1N-62): added redundancy to `foreign_upstream` —
> `concurrent: 3`, upstreams = `adg.chans.xyz` (hk2) + `dns.quad9.net` +
> `dns.cloudflare.com` (both direct-reachable from CN, live-tested 2026-08-13;
> `dns.quad101.net` excluded — TLS handshake fails). Verified: google.com
> AAAA now `2404:6800…` (new upstream answering, was `2607:f8b0…` via hk2),
> taobao/intercept/clash-fake-ip all unchanged. Backup:
> `config.yaml.bak-multi-doh-20260813-105421`.
> 2026-09-01: removed `dns.quad9.net` from `foreign_upstream` — recurring
> `WARN foreign_upstream … unexpected EOF` bursts (481 log entries) against
> Quad9 DoH; endpoint answers on probe but gets intermittently
> connection-reset from this network (same failure class as the excluded
> `dns.quad101.net`). Remaining upstreams `adg.chans.xyz` (hk2) +
> `dns.cloudflare.com` both verified live; google.com A + youtube.com AAAA
> resolve through mosdns :6052 after restart. Backup:
> `config.yaml.bak-quad9-remove-20260901-201801`.
- nft: OpenClash injects TPROXY/redirect + DNS-hijack rules into
`table inet fw4`; a residual `table inet passwall` exists with 0 packets (unused)
@@ -43,20 +84,22 @@ OpenClash runs `/etc/openclash/clash` (clash_meta core) with configuration
`8443` is not an OpenClash listener and has no runtime nftables forwarding rule.
It is included only in OpenClash's common TLS-sniffing port list.
## VLAN 10 Wi-Fi feasibility
## VLAN 10 Wi-Fi
`gfw` is a VM attached to untagged LAN 66, rather than a physical VLAN-trunk
endpoint. The U6 Lite likewise reaches the ER-X over LAN 66, whose DHCP
service supplies client addresses. Therefore `gfw` cannot currently receive an
SSID's VLAN 10 traffic merely by creating an `eth0.10` interface inside the VM.
`gfw`'s third NIC `eth2` hosts the `ubunt_upg` interface at `192.168.10.1/24`,
serving the dedicated `ubnt-upg` SSID VLAN 10 (untagged access path from a
VLAN-capable switch/trunk; AP management stays untagged on LAN66). The
`ubunt_upg` zone runs the **only** DHCP server for `192.168.10.0/24` (UDP/67),
allows DNS (53), and applies `192.168.10.0/24 → eth0 masquerade` (NAT) for
Internet egress. `forward_ubunt_upg` isolates VLAN10 from LAN66/55 and RFC1918
(deny counters 0, `accept_to_lan` passes).
Do not treat a local `eth0.10`/`192.168.10.1` configuration as a deployable
Wi-Fi gateway unless the hypervisor/vSwitch and the complete physical path to
the AP have first been configured and verified to carry tagged VLAN 10. The
`ubnt-upg` test receiving a `192.168.66.x` lease was the expected consequence
of the available untagged path, not evidence that this host should compete
with ER-X DHCP. See
[UniFi Network: Dedicated Wi-Fi through a third-party gateway](../docs/unifi-network.md#dedicated-wi-fi-through-a-third-party-gateway).
Live-verified 2026-08-11: an `ubnt-upg` client received `192.168.10.168` (lease
in `/tmp/dhcp.leases`), the `192.168.10.0/24 masquerade` counter climbed
(215 pkts/42KB), and the LAN55/LAN66 deny counters stayed 0 → VLAN10→LAN
isolation holds. See
[docs/lan-se5420-deployment-guide.md](../docs/lan-se5420-deployment-guide.md),
[docs/unifi-openwrt-vlan10-implementation-examples.md](../docs/unifi-openwrt-vlan10-implementation-examples.md)
## Operational note
+61
View File
@@ -34,6 +34,18 @@ new SSH host key out of band before accepting it.
IPv6 prefix delegation assigns SLAAC-capable `/64` networks to both LANs.
`eth4` applies the WAN IPv4 and IPv6 firewall policies.
**SE5420 single-uplink topology (verified 2026-08-22):** the TP-Link `TL-SE5420`
core switch is deployed — management `192.168.66.253` (TP-Link OUI `f8:c9:03`,
web UI on :80/:443). The LAN55 uplink into `switch0` is a **single member
port**: `eth1` link up, `eth2`/`eth3` down. All LAN55 wired devices (hass
`.11`, Aqara M3 `.248`, SmartThings `.48`, UAP-AC-Lite `.5`) are reached via
`switch0` behind that one uplink, so same-segment wired↔wired unicast is
switched locally on the SE5420 and never reaches the ER-X. The switch FDB is
hardware-offloaded and not readable from the ER-X (`brctl showmacs switch0`
"Operation not supported"; `show mac-address-table` / `show ethernet-switch`
are not available on this EdgeOS build) — port link state (`show interfaces
ethernet`) plus ARP are the reliable topology checks.
Detailed effective configuration, including firewall binding and WAN exposure,
is recorded in [the EdgeRouter X configuration record](../docs/edgerouter-x-configuration.md).
@@ -77,6 +89,35 @@ relevant interface/direction to take effect. Use the operational `show
firewall` output—not merely the configured rule definitions—to determine the
effective policy.
## PPPoE redial
To force the `pppoe0` session to reconnect (e.g. to obtain a fresh WAN IP), use
the operational `disconnect` / `connect` commands — **not** `renew dhcp
interface`, which applies only to DHCP interfaces:
```bash
ssh -4 zhiqiang@192.168.66.254
/opt/vyatta/bin/vyatta-op-cmd-wrapper disconnect interface pppoe0
/opt/vyatta/bin/vyatta-op-cmd-wrapper connect interface pppoe0
```
`disconnect` tears down the PPP session; `connect` re-dials immediately. A
short pause between them (a few seconds, or minutes for cautious ISPs) lets the
old session finish teardown before redialing. This briefly drops the whole WAN
uplink and may change the public IPv4 and delegated IPv6 `/60`; in-flight
sessions and port-forwarded services are interrupted until the new session is
up.
The `zhiqiang` account logs into `vbash`, not the EdgeOS CLI, so operational
commands must be invoked through `/opt/vyatta/bin/vyatta-op-cmd-wrapper` and
depend on its passwordless `sudo`. The `ubnt` account lands directly in the
operational CLI, where the same commands are entered without the wrapper.
`show`/`configure` are interactive-only aliases (from
`/etc/bash_completion.d/vyatta-{op,cfg}`, loaded via `~/.bashrc`), so a
non-interactive `ssh ubnt@… 'show …'` also fails — from a script use the op
wrapper above, or `_vyatta_op_run` after sourcing `vyatta-op` with
`vyatta_op_templates=/opt/vyatta/share/vyatta-op/templates`.
## Maintenance notes
- EdgeOS writes persistent changes through its configuration tree: enter
@@ -100,3 +141,23 @@ from `192.168.55.254` reached the UniFi controller at `192.168.66.46` with
3/3 ICMP replies. This supports the AP Inform path to
`192.168.66.46:9080`; the controller listener and an online LAN55 AP provide
the corresponding application-level evidence. No firewall changes were made.
IPv6 was re-verified by read-only SSH on 2026-08-20 during the UniFi AP/AC
check: the IPv6 routing table shows connected `/64`s on `eth0` (LAN66) and
`switch0` (LAN55) plus `::/0` via `pppoe0`; both UniFi APs obtained SLAAC
addresses from the router's RAs. No configuration changes were made.
**DHCP 保留 `matter` 失效(2026-08-21 发现,2026-08-23 复核仍未生效,W1N-207):**
静态映射 `matter` → .45 / MAC `34:98:7a:27:10:bc`,但该灯泡一直以**动态租约**拿
`.148`hostname `matter`2026-08-23 09:02 时租约当日 04:40 已续租)。保留 .45 从未
被租出。2026-08-23 复核补充:另一盏工作灯泡的 MAC 已变为 `fc:e8:c0:25:a1:f0`
(动态 `.146`hostname `espressif`),原「把 MAC 改为 `34:98:7a:27:7f:08`」的修正
建议已过时(该灯泡已离网)。处置:删除该保留,或按现用 MAC(`.148`
`34:98:7a:27:10:bc` / `.146``fc:e8:c0:25:a1:f0`)重建,**未执行**。
**SE5420 部署 + switch0 单上联(2026-08-22 只读核实):** `switch0` 成员口
`eth1` link up、`eth2`/`eth3` down(单上联);SE5420 管理面 `192.168.66.253`
在线(TP-Link OUI `f8:c9:03`:80/:443);ARP 显示 LAN55 主机(hass `.11`
M3 `.248`、SmartThings `.48`、UAP-AC-Lite `.5`)全部经 switch0 可达。含义:
`switch0` 不再是 LAN55 的全量抓包点(同段有线单播在 SE5420 本地交换),详见
[runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md)。
+751
View File
@@ -0,0 +1,751 @@
[hosts/hass.windy.lan.md#8DF6]
# hass.windy.lan — Home Assistant (HAOS)
## Role and access
| Item | Value |
|---|---|
| Role | Home Assistant automation hub |
| IPv4 | `192.168.55.11` (LAN55) |
| DNS | `hass.windy.lan` (AdGuard rewrite on `dns.windy.lan`; legacy `hass.local` alias) |
| SSH | `ssh hassio@hass.windy.lan` |
| **Host** | **x88 Pro physical box** (HAOS bare-metal, `machine: green`; verified 2026-08-18) |
| Platform | Home Assistant OS; kernel `6.1.115-haos` (aarch64) |
| Web UI | `http://hass.windy.lan:8123` (LAN); WAN port-forward `hass` on gw → `:8123` |
Use `hassio` for routine SSH inspection. Key-only login was verified on
2026-08-13 from the WSL client (`BatchMode=yes`).
The `ha` supervisor CLI (`/usr/bin/ha`) authenticates with `SUPERVISOR_TOKEN`.
Interactive login works because `~hassio/.zprofile` runs `exec sudo -i`, which
loads a root environment carrying the supervisor API token. Non-interactive
`ssh hassio 'command'` does not source `.zprofile` and fails with
`unauthorized: missing or invalid API token`. Run `ha` non-interactively via:
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core info'
```
Verified 2026-08-13 that `sudo -n -i ha core info` works from the WSL client.
Never copy the supervisor token into this repository.
The current SSH ED25519 host-key fingerprint is
`SHA256:DMcMOgDzFsFTon1fndXowEP7jlyOK3/AX3PVK8BATvk` (verified 2026-08-13).
Verify a changed key out of band before accepting it.
Do not store Home Assistant long-lived tokens, integration credentials, or
recovery codes in this repository.
## Network
| Interface | Address / role |
|---|---|
| `end0` | IPv4 static `192.168.55.11/24` (gw `.254`, DNS `192.168.66.36`); IPv6 SLAAC `auto` with GUA on the current PD-derived /64 (`240e:3bd:235:1fb2:*` at 2026-08-22; rotates on PPPoE redial); primary LAN55 NIC (interface name verified live 2026-08-22 — `end1` does not exist) |
| `wlan0` | Supervisor **disabled** (verified 2026-08-14, W1N-104); IPv6 remains off on this RTL8821CS radio |
| `wg0` | `10.13.13.2/32`; WireGuard (add-on / integration tunnel) |
| `hassio` / `docker0` | internal HAOS Docker bridges (`172.30.32.0/23`, `172.30.232.0/23`) |
LAN55 clients reach the HTTP API on `dns.windy.lan:80` for the AdGuard Home
integration; see [hosts/dns.windy.lan.md](dns.windy.lan.md).
## API access
Home Assistant exposes a REST API at `http://hass.windy.lan:8123/api/` (same
as `http://192.168.55.11:8123/api/`). Authenticate with a **long-lived access
token** created under **Profile → Security → Long-lived access tokens**.
```bash
HA_URL="http://hass.windy.lan:8123"
HA_TOKEN="<long-lived-access-token>"
# Health check — expect {"message":"API running."} and HTTP:200
curl -sS -w "\nHTTP:%{http_code}\n" \
-H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/"
# Read one entity state
curl -sS -H "Authorization: Bearer $HA_TOKEN" \
"$HA_URL/api/states/sensor.csg_30d_max"
# List entities / recent errors
curl -sS -H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/states"
curl -sS -H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/error_log"
```
- `401` → token invalid or expired; create a new one.
- `404` on `/api/states/<id>` → entity does not exist.
- The token is a secret: never commit it here; keep it in the shell
environment or a secrets file outside the repo.
### HTTP proxy gotcha (verified 2026-08-13)
The WSL client had `http_proxy` set to Mihomo (`192.168.66.99:7890`). LAN
hostnames sent **through that proxy** returned empty `502`, even though DNS
resolved and the HA UI was up. Direct `192.168.55.11:8123` worked, and
`hass.windy.lan:8123` worked only after clearing the HTTP proxy.
Before debugging a "502" on a LAN URL, check `env | grep -i proxy` and bypass
the proxy:
```bash
unset http_proxy HTTP_PROXY all_proxy ALL_PROXY
curl -sS -w "\nHTTP:%{http_code}\n" \
-H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/"
```
For a persistent fix, add `.windy.lan` (leading dot) and the LAN ranges to
`NO_PROXY`, or add `*.windy.lan` to the proxy's own bypass/skip-proxy list.
See `~/.config/zsh/env/local/environment.env` for the client-side setting.
## Safe verification
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan 'hostname; ip -4 addr show end0'
```
From a LAN client, confirm DNS and UI reachability:
```bash
getent hosts hass.windy.lan
# expect 192.168.55.11
```
## Local patches (custom components)
### Manual custom-component install (this host)
Home Assistant loads custom integrations from
`<config>/custom_components/<domain>/` (HAOS: `/config``/homeassistant`).
A folder named after the integration domain, containing at least
`manifest.json` and `__init__.py`, is enough; Core must be restarted after
copying files. Official HA lookup order:
`<config>/custom_components/<domain>` then built-in
`homeassistant/components/<domain>`.
See [Integration file structure](https://developers.home-assistant.io/docs/creating_integration_file_structure).
This host **does not git-clone** custom components. The live tree is a file
copy. Do not `git pull` on HA.
**Official plugin path** (from
[windyboy/china_southern_power_grid_stat README](https://github.com/windyboy/china_southern_power_grid_stat)):
HACS **or** [手动下载安装](https://github.com/windyboy/china_southern_power_grid_stat/releases).
This host uses the latter. Releases here have no uploaded zip assets; use
GitHub's **Source code (zip)** / zipball of the tag.
**UI (Samba / File editor / Studio Code Server):**
1. Download Source code (zip) from the GitHub Release.
2. Extract. Copy only the inner
`custom_components/china_southern_power_grid_stat/` tree — not the repo
root, not a nested extra folder.
3. Place it at `/config/custom_components/china_southern_power_grid_stat/`.
4. Restart Core (**Settings → System → Restart**).
5. First install only: **Settings → Devices & services → Add integration**.
**SSH from the workstation** (verified 2026-08-14, W1N-107). Replace `v1.3.1`
with the tag being installed:
```bash
TAG=v1.3.1
STAGE=/tmp/csg-${TAG}-deploy
mkdir -p "$STAGE"
gh api "repos/windyboy/china_southern_power_grid_stat/zipball/${TAG}" \
> "$STAGE/src.zip"
unzip -q "$STAGE/src.zip" -d "$STAGE"
SRC=$(find "$STAGE" -type d -path '*/custom_components/china_southern_power_grid_stat' | head -1)
# expect .../custom_components/china_southern_power_grid_stat
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i mkdir -p /homeassistant/.csg-backups &&
sudo -n -i cp -a /homeassistant/custom_components/china_southern_power_grid_stat \
/homeassistant/.csg-backups/china_southern_power_grid_stat.bak-$(date +%Y%m%d)-manual'
rsync -a --delete \
-e 'ssh -o BatchMode=yes' \
"$SRC/" \
hassio@hass.windy.lan:/homeassistant/custom_components/china_southern_power_grid_stat/
# --delete cannot remove Core-owned __pycache__; wipe as root, then restart
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i rm -rf /homeassistant/custom_components/china_southern_power_grid_stat/__pycache__ \
/homeassistant/custom_components/china_southern_power_grid_stat/*/__pycache__ &&
sudo -n -i ha core restart'
```
Wait until Core is up (`ha core info` returns, typically 12 min; this CLI
build does not print a `state:` field).
Then:
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i cat /homeassistant/custom_components/china_southern_power_grid_stat/manifest.json'
# version must match the tag
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i ha core logs -n 2500' | grep -E 'china_southern_power_grid_stat|cannot pickle' || true
```
**Host constraints (do not skip):**
- Backups **must** live in `/homeassistant/.csg-backups/`. A `*.bak-*`
directory next to the live folder is scanned as the same domain and Core
fails with `No module named '...bak-YYYYMMDD-...'`.
- Do not install this fork via HACS on this host. HACS still tracks
`CubicPill/china_southern_power_grid_stat` `v1.2.0`; a HACS update would
overwrite the live copy.
- First poll after restart can time out to CSG over IPv4; if this-month
sensors stay `unknown` while last-month filled, reload the config entry
(UI: integration → Reload, or supervisor
`POST /core/api/config/config_entries/entry/<id>/reload`).
- `runbooks/scripts/ha-maintenance.sh --restart-core --yes` can print
nothing and exit 1 in under a second **without restarting Core**. The
wrapper's ssh line discards stderr (`2>/dev/null`); with `pipefail`,
an ssh failure yields empty stdout + exit 1 before any remote command
runs. Do not treat that as a completed restart. Confirm with elapsed
time (~2 min for a real restart) and `ha core info`. Prefer
`ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core restart'`.
Full command family: [runbooks/home-assistant-maintenance.md](../runbooks/home-assistant-maintenance.md).
### `china_southern_power_grid_stat` live tree
**v1.3.2** (`934f58c`, verified 2026-08-15, W1N-118): manual zipball of
GitHub release
[v1.3.2](https://github.com/windyboy/china_southern_power_grid_stat/releases/tag/v1.3.2)
copied to `/config/custom_components/china_southern_power_grid_stat`.
Earlier trees: v1.3.1/`55a293fc` (W1N-107), v1.3.0/`69f13c90` (W1N-106),
`a433e8c` (W1N-105), `de01914` (W1N-103), `eb8b174` (W1N-102). Backups:
`/homeassistant/.csg-backups/` (w1n102/104/105/106/107/118).
v1.3.0 crashed the coordinator on first refresh
(`TypeError: cannot pickle 'mappingproxy' object` in
`copy.deepcopy(self._config)` under Python 3.14 / HA 2026.8.1). v1.3.1
wraps those `deepcopy` calls with `dict(...)`. Post-restart 22:13 CST:
entry `loaded`, no pickle traceback. Native this-month sensors filled after
reloading entry `01KGCQDSZCF523A9X6SV3BZ1B9` (`ip_family: ipv4`). Native
cost/ladder sensors can stay `unknown` because CSG
`get_month_daily_cost_detail` returns a marketing-system SQL error; the
dashboard uses template ladder/cost entities instead. Do not change
`templates/csg_sensors.yaml` or the 电力监控 dashboard for an install.
**`templates/csg_sensors.yaml` hardened 2026-08-29 (W1N-239):** added
`availability` templates to all 12 `csg_*` sensors (numeric sensors can't
render `unknown`/`unavailable` in `state`; availability suppresses
rendering instead — native CSG down ⇒ derived sensors show `unavailable`,
no more fake zeros / "一档" / `0%`). `csg_yesterday_kwh` now falls back to
`last_month_by_day`'s last entry when `this_month_by_day` is empty (month
start); ladder constants (`t1/t2/p1/p2/p3`) deduped into per-block
`variables:` (Block B + Block D); `csg_mom_change` parses `date`
defensively. Backup:
`/homeassistant/.csg-backups/csg_sensors.yaml.bak-20260829-w1n239`.
**Verified:** `ha core check` OK; Core restart required (trigger-based
template blocks don't settle on `template.reload` — W1N-114 precedent);
post-restart all 12 entities numeric & consistent (302.47 kWh→180.28 元,
324.03 kWh→194.06 元, mom_change -3.6%, yesterday 7.66 kWh/2026-08-28),
no template errors in Core logs.
**`csg_sensors.yaml` off-by-one fixed 2026-08-29 (W1N-241):** CSG data
lags 1 day (`sum(this_month_by_day)` == `this_month_total_usage`, data
stops at yesterday), but templates used `now().day` as "days elapsed" →
`csg_predicted_usage` underestimated ~1 daily avg (~3%) and
`csg_mom_change` compared this-month 28 days vs last-month 29 days
(-3.6% vs true -0.3%). Both now derive the day number from
`this_month_by_day[-1].date` (fallback `now().day` when empty). Added
`sensor.csg_this_month_daily_avg` (month-to-date avg, 302.47/28=10.8) and
`sensor.csg_prediction_progress` (usage/predicted %, 90.3) in Block C
(trigger adds `csg_predicted_usage`). Backup:
`/homeassistant/.csg-backups/csg_sensors.yaml.bak-20260829-w1n241`.
**Verified (8/29):** predicted 324.03→334.81, mom_change -3.6→-0.3,
daily_avg 10.8, progress 90.3, predicted_cost 194.06→200.94 (334.81 kWh
ladder), ladder cost 180.28 unchanged, `ha core check` OK after restart,
no template errors; 14 csg_* entities total.
**电力监控面板(`lovelace.dashboard_unknown` / view `power-monitor`
updated 2026-08-29 (W1N-240 + W1N-242):** 「本月累计」gauge 对齐夏季阶梯:
`max:650`、segments `0/260/600`(绿/橙/红 = 一/二/三档;冬季 11-01 需切
`max:450``0/200/400`**seasonal switch point**,见下文)。「📊 统计
数据」卡新增本年/去年 4 行(原生传感器,口径标注「电费(账单)」、本年
「(至今)」)+ 本月日均/预测进度 2 行(`csg_this_month_daily_avg` /
`csg_prediction_progress`W1N-242);面板共引用 **20** 个实体。改前备份:
`/homeassistant/.lovelace-backups/dashboard-unknown-power-monitor-20260829-204845.json`
W1N-240)、`-20260829-210708.json`W1N-242
(改法:WS `lovelace/config/save`,参数 `url_path: dashboard-unknown` +
`config`;勿直改 `.storage/`)。验证:WS 读回 18→20 实体 diff ✓、gauge
配置一致 ✓、URL `http://hass.windy.lan:8123/dashboard-unknown/power-monitor`
**`csg_sensors.yaml` W1N-242:** `csg_predicted_usage` /
`csg_mom_change` / `csg_this_month_daily_avg` 三处取 `days[-1]` 前补
`sort(attribute='date')`(与 `csg_yesterday_kwh` 一致,防上游乱序取错
数据日)。备份 `csg_sensors.yaml.bak-20260829-w1n242`。验证:Core
restart 后回归值不变(334.81 / -0.3 / 10.8 / 90.3 / 200.94 / 180.28)。
**CSG 面板重构 2026-09-04VPS-90,先核对计价后展示层改动):** 核对
`power-monitor` 计价与 8 月账单一致(198.65 vs 账单 198.64,差 ≤0.01 元,
因模板用公众圆整价 0.589/0.639/0.889、账单用 6 位精确价),不改阶梯常量。
改动:① `csg_sensors.yaml` Block B 新增
`sensor.csg_this_month_avg_price`(本月阶梯电费÷本月用电,`元/kWh`
availability 照 W1N-239 惯例;**csg_* 实体 14→15**);② 面板改名「环比上月」
→「环比上月同期」;glance「本月/上月」grid 去重为单卡「上月」(本月用电/电费
行归 💰核心数据卡);⚡阶梯电价卡加「本月实际均价」行(当前档位/当前电价/
本月实际均价/档位剩余;面板唯一实体引用 20→21);③ `automations.yaml`
2 条提醒:`automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`10-25 09:00
`automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`4-25 09:00)经
`matrix_e2ee.send_message` 提醒切 gauge。④ 金额单位混排(原生 CNY vs 模板
元)**维持**`config/entity_registry/update` 拒绝自定义文本单位
`extra keys not allowed … Got '元'`),用户确认接受。备份:
`.lovelace-backups/dashboard-unknown-power-monitor-20260904-204757-pre-refactor.json`
`.csg-backups/csg_sensors.yaml.bak-20260904-204757-pre-refactor`(及
`-205301-pre-avgprice`)、`.automations-backups/automations.yaml.bak-*`
**WS 改法(2026.8,本机实测)**core/主机 python 无 ws 库、core 容器内经
supervisor 代理 WS 被拒(loop prevention),用
`docker run --rm --network host -e SUPERVISOR_TOKEN`supervisor 镜像
`aarch64-hassio-supervisor:2026.08.0`)连 `ws://172.30.32.2/core/websocket`
命令名 `lovelace/config`(读)+ `lovelace/config/save`(写),
`lovelace/config/get` 已不存在(unknown_command)。验证:新实体
0.589 元/kWh、15 个 csg_* 数值齐全、回归值不变(14.09/198.65/331.22/
304.99/181.89)、automations on、`ha core check` OK、日志无 template 错误。
**CSG 长期归档(W1N-243, 2026-08-29:** scribe 库新增 `csg_history`
表(逐日 usage/cost/ladder/balance + 逐月累计;2026-07-01 起回填,永久),
由 TimescaleDB 每日任务 **1008** `csg_daily_snapshot()`22:30
Asia/Shanghai**TS job 非 pg_cron**upsert 维护。日费用在原生
`latest_day_cost` 缺失时回退 = 昨日用电 × 当前档费率(模板
`csg_current_ladder_tariff` 0.639);月费用回退模板
`csg_this_month_ladder_cost`。**语义**day 行 usage/cost 为该日值,
ladder/balance 为 22:30 快照值。详见 [hosts/pgdb.md](../hosts/pgdb.md)。
> **Seasonal gauge switch (W1N-240 已知事项):** 每年 **11-01** 把
> `power-monitor` 视图「本月累计」gauge 切到冬季 `max:450` /
> `0/200/400`**5-01** 切回夏季 `max:650` / `0/260/600`(与模板
> `now().month` 季节逻辑对齐;模板常量在 Block B/D `variables`)。
> **提醒 automation2026-09-04 起,VPS-90:**
> `automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`10-25)与
> `automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`4-2509:00 经
> `matrix_e2ee.send_message` 发操作步骤提醒;gauge 的 max/segments 无法
> 模板化,仍需人工改卡配置。
Home PPPoE IPv4 to CSG is still blackholed (`curl -4` to `218.19.148.218:443`
times out). `end0` IPv6 is enabled (`ipv6.method: auto`); from HA,
`curl -6 https://95598.csg.cn` returns HTTP 200 via `240e:f9:8060::1:16`.
**`tianqi` weather recorder patch (verified 2026-08-13, W1N-75):**
`/config/custom_components/tianqi/weather.py` has a local patch adding
`_unrecorded_attributes = frozenset({"hourly_temperature", "hourly_skycon",
"hourly_cloudrate", "hourly_precipitation"})` to the `WeatherEntity` class.
Without it, weather.guangzhou's state attributes (~19 KB, dominated by the 4
hourly_* arrays of up to 48 entries) exceed the recorder 16384-byte limit, so
the recorder drops **all** attributes for the entity and logs
`Recorder.db_schema: State attributes for weather.guangzhou exceed maximum
size of 16384 bytes`. The patch excludes only the 4 arrays from recording
(live state unchanged; other attributes still stored; ~6.3 KB payload). Backup
at `weather.py.bak-w1n75`. **Re-apply after any `tianqi` component update.**
The `_unrecorded_attributes` mechanism exists in Core 2026.8.1
(`Entity.__init_subclass__``state_info["unrecorded_attributes"]`, consumed
by recorder `shared_attrs_bytes_from_event`).
### `matrix_e2ee` live tree (E2E Matrix bot, verified 2026-08-20)
**v0.3.12** (tag `v0.3.12`; feat — Matrix activity events
`matrix_e2ee_message_received` / `matrix_e2ee_verification_done` + push
diagnostics; v0.3.9 added Connection health binary sensor, SAS/command
allowlist split, URL normalization, single-entry enforcement):
source copy from `/home/windy/project/ha-matrix-e2ee` `ea421ed` (tag
`v0.3.12`) deployed 2026-08-20 via SSH rsync from workstation (upgraded
from v0.3.2, backup `matrix_e2ee.bak-20260820-v0.3.2`).
Custom **`matrix_e2ee`** integration — **Config Flow** (UI). See
[docs/home-assistant-matrix.md](../docs/home-assistant-matrix.md).
**Update runbook:** [runbooks/matrix-e2ee-update.md](../runbooks/matrix-e2ee-update.md).
Earlier: v0.3.2 (tag `v0.3.2`, W1N-182/#34: wizard waits for inbound SAS
emojis) deployed 2026-08-18 from `d35c484` (backup
`matrix_e2ee.bak-20260818-v0.3.1`); v0.3.1 (GitHub #33: peer-initiated
verification wizard fix) deployed 2026-08-18 from `d22e935` (backup
`matrix_e2ee.bak-20260818-v0.3.0`); v0.3.0 (W1N-180/#32: bot-initiated
verification wizard; W1N-179/#31 `receive_mac_event` cancel-state fix)
deployed 2026-08-18 from `216cc99` (backup
`matrix_e2ee.bak-20260818-v0.2.10`).
- Bot `@hass:chans.xyz` reused (E2EE device `rO1R915ncu`). Config Entry
`01M04D7C1M4T2GX5VPG7NVQ7GV` (`source: import`, `state: loaded`). All
settings via **Settings → Devices & Services → Matrix E2EE → Configure**.
- Config Entry options: `allowed_rooms` `["!gidvAzpDzwtzfEDrqu:chans.xyz", "!boxfylDSzOvrWkcsyY:chans.xyz"]`,
`allowed_users` `["@zhiqiang:chans.xyz"]`, `command_prefix` `"!"`.
**`verification_peer_users` not set** (v0.3.9+ SAS allowlist split from
`allowed_users`, W1N-156): defaults to empty → only the bot's own account
may drive SAS; `@zhiqiang` is denied until the option is added via
Settings → Devices & Services → Matrix E2EE → Configure.
- Storage: `/config/.storage/matrix_e2ee_session.json` +
`/config/.storage/matrix_e2ee_store/`. Backups:
`/homeassistant/.matrix-e2ee-backups/` (incl. `matrix_e2ee.bak-20260820-v0.3.2`,
`matrix_e2ee.bak-20260818-v0.3.1`,
`matrix_e2ee.bak-20260818-v0.3.0`,
`matrix_e2ee.bak-20260818-v0.2.10`,
`matrix_e2ee.bak-20260816-v0.2.9`, `matrix_e2ee.bak-20260816-v0.2.8`);
full HA backup slugs `3d9d36db` (pre-v0.1.4) + `9f223f35` (pre-v0.2.0).
- v0.3.12: Matrix activity events + push diagnostics
(`matrix_e2ee_message_received` / `matrix_e2ee_verification_done`).
v0.3.9: Connection health binary sensor (W1N-185/#40), config-entry
diagnostics (W1N-184/#39), SAS/command allowlist split
`verification_peer_users` (W1N-156/#41), SAS/sync logs demoted
warning→info/debug (W1N-188/#38), URL normalization + single-entry
enforcement (W1N-190/#42).
v0.3.8: `m.key.verification.done` handshake for request-based SAS
(W1N-183/#35).
v0.3.2: wizard waits for inbound SAS emojis before the compare step
(W1N-182/#34).
v0.3.1: verification wizard waits for a peer-initiated inbound SAS instead
of the bot starting SAS (GitHub #33).
v0.3.0: bot-initiated device verification wizard (W1N-180/#32).
v0.2.11: `receive_mac_event` no longer overrides canceled state (W1N-179/#31).
- v0.2.9: restore SAS emoji rendering after vodozemac migration (W1N-175/#29).
v0.2.8: SAS commitment unpadded base64 for Element interop (W1N-174/#28).
v0.2.7: SAS cancel code/reason logging. v0.2.6: verification state logging +
request→ready bridge. v0.2.4: `_patch_nio_sas_timeout()` +
`_repair_dropped_start()`; `VERIFICATION_TIMEOUT_SECONDS` 600→240.
- Automation `1761188403590`「Matrix 聊天关卫生间灯」: trigger
`matrix_e2ee_command` (command `关卫生间灯`), actions `light.turn_off` +
`matrix_e2ee.send_message` (room `!gidvAzpDzwtzfEDrqu`).
- **SAS not yet completed:** every device requires explicit `confirm_verification`.
Encrypted-room commands stay fail-closed until `@zhiqiang`'s device is verified.
Since v0.3.9 the SAS driver gate uses `verification_peer_users` (empty on
this host) instead of `allowed_users` — add `@zhiqiang:chans.xyz` there
before retrying the wizard. Three paths available: SAS manual confirm,
fingerprint, or the device verification wizard (v0.3.0 bot-initiated,
reworked in v0.3.1/v0.3.2 to wait for a peer-initiated inbound SAS from
Element with emoji comparison), see
[docs/home-assistant-matrix.md § Device verification](../docs/home-assistant-matrix.md).
### Scribe long-term history (3.8.0 setup 2026-08-29; 4.4.0 verified 2026-09-13)
- **Scribe 4.4.0** (`/homeassistant/custom_components/scribe/`, HACS repo
`jonathan-gtd/scribe`, = latest stable 2026-09-12; upgraded 2026-09-13 together
with Core 2026.9.1 / HAOS 18.2), configured from
`/homeassistant/scribe.yaml` — W1N-238 moved the block out of
`configuration.yaml` on 2026-08-29 (main config now carries
`scribe: !include scribe.yaml`; content moved verbatim; backup
`configuration.yaml.bak-20260829-201724-w1n238`). Config entry
`01KC2VFJWEQ3XDHY6TQKHPDVRB`, `source: import` — UI "Configure → Advanced"
edits are overridden by the YAML on restart; treat YAML as authoritative.
- TimescaleDB at `192.168.55.15:5432/scribe` (DB user `hass`; host in inventory,
see [hosts/pgdb.md](../hosts/pgdb.md)). Database re-initialized 2026-08-29 14:06 CST
(user-handled; earlier `relation "entities" does not exist` errors resolved).
Health: `binary_sensor.scribe_database_connection`.
- 2026-08-29 config applied (backup `/homeassistant/configuration.yaml.bak-20260829-scribe`):
- `record_events: true` with `include_events` whitelist: `automation_triggered`,
`matrix_e2ee_command`, `matrix_e2ee_message_received`,
`matrix_e2ee_verification_done`, `script_started`, `tag_scanned`,
`mobile_app_notification_action`, `homeassistant_start`, `homeassistant_stop`.
- State noise trimmed: `exclude_domains` update/button; glob
`sensor.zigbee2mqtt_bridge_*`; 4 hassio cpu/mem-percent entities.
- Global `exclude_attributes` drops tianqi `hourly_*` arrays (~19 KB/state —
the recorder-side `_unrecorded_attributes` patch does not apply to Scribe).
- `enable_stats_io` + `enable_stats_size` on → 14 `sensor.scribe_*` stats
entities (`scribe_states_written`, `scribe_events_written`, rates, sizes).
- Verified post-restart 14:23 CST: writer started, `scribe_events_written=1`
(homeassistant_start), states ~110/min, buffer 3, no scribe log errors.
- **4.x upgrade核对 2026-09-13(只读 + 一处配置变更)**live `manifest.json` =
4.4.0。两个 4.0 breaking change 在本机都不需要动作——数据库是 3.x 结构
`states_raw` PK `(metadata_id, time)` 在,4.2.0 的启动态去重因此可用),
TimescaleDB 2.29.2 已装。4.1.0 修了 `db_url` 优先级,YAML 里的
`!secret scribe_url` 现在是权威。`scribe.yaml` 现有键在 4.4.0 全部仍然合法
(未知键被忽略,`extra=vol.ALLOW_EXTRA`)。**配置优先级 YAML > entry
`options` > entry `data` > 默认值**,而 `_resolve_settings` 读的是
`hass.data[DOMAIN]["yaml_config"]`(只有 `async_setup` 会写),所以
**YAML 改动必须重启 Corereload config entry 不重读 YAML。**
- **`stats_io_interval: 300`2026-09-13 添加**,备份
`/homeassistant/scribe.yaml.bak-20260913-191558`)。4.4.0 不再让 HA 每 30s
轮询 I/O 统计传感器,改由集成自己每 60s 发布,间隔成为配置项。scribe 自己的
传感器此前是本机自写历史的主要来源(变更前 24h:11 019 / 87 461 行状态 =
12.6%),60s → 300s 把这部分降约 5 倍(每个 I/O 传感器约 1440 → 288 行/天)。
验证:`ha core check` OK;重启 88s`ScribeWriter started successfully`
无 scribe error/warningscribe Repairs 问题 0 条;传感器发布间隔实测正好
300s11:19:26 → 11:24:26 UTC)。
- **Retention 现在可用但刻意不设**:`retention_states` / `retention_events`
4.0.0)按间隔丢 chunk,留空 = 永久保留,符合本机定位(Scribe 是永久归档,
recorder 保留 365 天)。注意 retention 是**绕过** entry `data` 副本读取的
`from_entry_data=False`),所以删掉 YAML 行即撤销策略。`db_schema`
`enable_rollups``scribe.purge` 同样未用:图表走 `sensor_minute` +
`timescale_database_reader`(见 [hosts/pgdb.md](pgdb.md)),不吃 scribe 自己的
视图,配置里也没有任何 `scribe.query` 调用。`flush_interval` 仍是 entry
`data` 钉住的 5s——上游下一个版本把默认改成 30s,但 entry 值优先,要采用只能
在 YAML 显式写 `flush_interval: 30`
- Recorder stays external-Postgres with `purge_keep_days: 365` (W1N-243,
2026-08-29, raised from 30 — ~300 MB/yr, 1% of the 30G pgdb disk) for
native UI per-change history; Scribe is the permanent archive. Long-term
statistics stay permanent (not purged by `purge_keep_days`). Note:
extending retention does **not** recover pre-2026-08-29 raw history
(already purged); only `csg_history` day/month values cover that period.
### Config layout: scribe.yaml + templates/ merge (W1N-238, verified 2026-08-29)
- `configuration.yaml` line 29: `scribe: !include scribe.yaml`; line 9:
`template: !include_dir_merge_list templates`. No `packages/`.
- `scribe.yaml` (config root): the Scribe block, content identical to the
former inline one; import semantics unchanged.
- `templates/`: `csg_sensors.yaml` (12 template sensors, top-level **list**)
+ `quick_sensors.yaml` (scaffold for Quick-derived `quick_*` sensors, empty
list with convention header). **`!include_dir_merge_list` merges per-file
lists; non-list files are silently skipped** — every file in `templates/`
must be a top-level list (`- sensor:` blocks). Directory include only picks
up `*.yaml`, so the `.bak` / `.pre-*` backups in the dir are ignored. After
adding sensors, verify template-platform entity count = 12 + N (entity
registry `platform: template`).
- Convention (per review + W1N-233): pure sums/averages stay min_max helpers
(e.g. `sensor.dang_qian_zong_gong_lu`); only template-logic derivations
(ladder pricing, cross-entity conditions) go into `quick_sensors.yaml`.
- Post-change verification 20:19 CST: `ha core check` ok, 92 s restart
(2026.8.3), `binary_sensor.scribe_database_connection` on,
`scribe_states_written` 18581→19426 growing, template entities still 12,
csg sensors numeric, no scribe/template log errors.
### Timescale Plotly card + database reader (verified 2026-08-29)
Chart stack over the Scribe TimescaleDB archive. Upstream pair (no HACS;
manual copies): reader `remmob/timescale_database_reader` **v1.1.0** (main
`bb8776a`) + card `remmob/timescale-plotly-card` **2.2.0** (main `217961d`).
- **Reader integration**: `/homeassistant/custom_components/timescale_database_reader/`.
Config entry `01M165P77QT1FQEAVPNZHDT82W` ("Scribe", `source: user`): connects
`hass@192.168.55.15:5432/scribe` (credentials = `secrets.yaml` `scribe_url`),
`table: sensor_minute`. Exposes no entities/services — it serves WS command
`timescale/query` (window ≤ 365 d, ≤ 50 000 rows, `downsample` bucket seconds).
Benign startup warning `Error executing test query: column "time" does not
exist`: the self-test SQL assumes the LTSS column name; the scribe table uses
`minute` — real queries work (verified: 70 rows for a live power sensor).
- **Card**: `/homeassistant/www/community/timescale-plotly-card/timescale-plotly-card.js`
(root-owned, same convention as HACS dirs). Lovelace resource (storage)
id `2e360d17b5aa4ce59c2fd13c43b51215`
`/hacsfiles/timescale-plotly-card/timescale-plotly-card.js`, type `module`.
Card config matches the entry by `database: scribe` (name from the reader
entry). Updates: replace the file, resource URL unchanged — browsers need a
hard refresh or a bumped `?v=` query on the resource URL.
- **pgdb side** (`sensor_minute_aggregate` cagg + `sensor_minute` hypertable +
every-minute refresh job): see [hosts/pgdb.md](pgdb.md) § Databases.
- **Agent-side HA WebSocket without a long-lived token** (verified 2026-08-29):
connect `ws://supervisor/core/websocket` with header
`Authorization: Bearer $SUPERVISOR_TOKEN`, then send
`{"type":"auth","access_token":"$SUPERVISOR_TOKEN"}` — the Supervisor proxy
swaps it for a core token (works as the internal Supervisor admin user). Note
`lovelace/resources/create` in HA 2026.8 takes `res_type` (NOT
`resource_type`).
- Scribe stores numeric sensor values in `states_raw.value` with `state` NULL,
so `sensor_minute.state` shows `'0'` for numeric sensors; the card plots
`avg_state` (from `value`) — expected, not a bug.
- **Quick 仪表盘(`dashboard-quick`)图表套件**2026-08-29 创建,经 WS
`lovelace/config/save` 写入;W1N-230 修复 + W1N-231 round-2 改进):
5 张 timescale 卡——大功率电器/常驻负载功率(按量级拆图,避免尖峰压扁
<70 W 基线)、按插座用电量(`energy_mode` + cumulative/diff,数据质量前提
见 pgdb 的 refresh 过程补丁)、室内外温湿度(温度左轴/湿度右轴,4 位置同色
配对)、人体感应活动状态(3 个 `motion_state`banded `state_map`
none/small/medium/large → 0-11per-entity `line_color` 红/蓝/绿)。
空调实体引用为 `kong_diao_*``kong_tiao` 是笔误,W1N-230 修复;`grep -c
kong_tiao` 应为 0)。灯区:2×2 嵌套 grid(`grid_options: {columns: "full"}`
内层 `columns: 2`+ 4 卡统一 `mushroom-light-card`(显式 name、
`use_light_color: false`、内联亮度/色温控制),heading icon
`mdi:lightbulb-group`。heading badges:环境 4 温度(迷你/mini数显/数显/广州)、
大功率电器 空调/电脑当前功率、常驻负载 总功率
`sensor.dang_qian_zong_gong_lu`min_max **sum** helper`round_digits: 0`
任一源掉线 fail-closed → unknown)。常驻负载图卡级 `fill: 'tozeroy'` +
冰箱/主网络 per-entity `fill_color`(线色 20% 透明)+ 其余 5 条 `fill: false`
per-entity fill 逐系列退出,卡 JS `seriesConfig.fill !== false`)。
布局:视图 `type: sections` + `max_columns: 4`;灯/用电/环境/人体感应
`column_span: 4`,功率两图拆两个 `column_span: 2` 分区**并排**(等高 280px
桌面并排、手机回落堆叠;去卡内 title 省半宽图垂直空间)。
**分区/卡片是两套尺寸键,不可混用**:分区宽 = `column_span`
`hui-sections-view.ts` 缺省按 1 列渲染,绝不省略);卡片宽 =
`grid_options: {columns: <n|"full">}``hui-card.ts` 只读 `config.grid_options`
写在卡片上的 `column_span` 被静默忽略;缺省 12 列,分区内格 = 12 × 分区
span,故 span-4 分区里缺省卡片只有 1/4 宽)。
修改前备份:`/homeassistant/.lovelace-backups/dashboard-quick-*.json`
W1N-230 修复: `20260829-190256`round-2 改进: `20260829-194040`)。
- **Quick 时间范围扩容 (2026-09-13, VPS-92)**: 用户反馈「48 小时不够」。
各 timescale 卡可选档上调——大功率电器/常驻负载 `…,24h``+3d,7d`
环境 `6h,12h,24h,48h``+7d,14d,30d`;人体感应 `…,24h``+3d,7d`
用电量(按插座) `energy_time_ranges` `today,week,month,custom``+3mo`
**默认档未改**6h / 6h / today / 24h / 12h)。卡片 JS 只接受
`<n>m|<n>h|<n>d``parseDurationToMs` 正则 `/^(\d+)(m|h|d)$/`
仅 m/h/d,无 w)与命名档 `today|week|month|3mo|6mo|year|years|custom`
`energy_mode` 卡必须用后者。**数据下界注意**:scribe `sensor_minute`
目前最早只到 **2026-08-29**,所以 >15d 的档(14d 边缘、30d 明显)前半段
会是空白,等归档继续累积才好看。备份
`.lovelace-backups/dashboard-quick-20260913-190912-pre-timerange.json`
### 地图仪表盘:CARTO keyed tiles via `custom:map-card` (verified 2026-08-30, W1N-261)
- **背景:** CARTO 自 2026-08-26 起对无 key 栅格瓦片打 "API KEY REQUIRED"
水印,内置地图卡/zone 编辑器全部受影响。Core 2026.8.3 的 `MapCardConfig`
**没有任何瓦片配置项**frontend 20260729.7 源码核对:
`setup-leaflet-map.ts` 硬编码 CARTO voyager URL)。上游修复是 2026.9.0b1
起改用 OSMF 矢量瓦片(frontend PR #53816),stable 预计 2026-09-02 前后。
- **变更:** 「地图」仪表盘(url_path `map`storage)唯一 map 卡替换为
`custom:map-card`[nathan-gs/ha-map-card](https://github.com/nathan-gs/ha-map-card)
**v1.16.0**,手动安装非 HACS):`tile_layer_url` =
`https://{s}.basemaps.cartocdn.com/rastertiles/voyager/{z}/{x}/{y}.png?key=<CARTO_KEY>`
(配 `tile_layer_options: {subdomains: abcd, maxZoom: 20}` + OSM/CARTO
attribution)。实体不变:2 person + 4 zonezone 用 `display: icon` +
`circle: auto`circle 读实体 `radius` 属性画半径圈)。
- **CARTO key 是 secret**: 只存在于服务端 lovelace 存储(dashboard `map`
的卡片配置)和用户本人处;勿写入本仓库或 Linear。
- **文件/资源:** `/homeassistant/www/community/ha-map-card/map-card.js`
root:root 644678554 Bsha256
`f30dfb606e858d2216d5198d8cf758ce956d127006ebd7d66d4329153a247ec2`);
Lovelace resourcestorageid `9d2b50b52c60420d89ebd041f722cf60`
`/hacsfiles/ha-map-card/map-card.js`type moduleWS
`lovelace/resources/create`2026.8 参数名 `res_type`)。升级 = 手动替换
该文件(不在 HACS 管理下,浏览器需强刷)。
- **备份:** `/homeassistant/.lovelace-backups/dashboard-map-map-20260830-133714.json`
(还原 = 把备份里的 `views[0].cards[0]` 写回后再 WS `lovelace/config/save`
url_path `map`)。
- **验证 8/30:** 同瓦片无 key=水印 / 带 key=干净(256×256 PNG 视觉对比);
resource HTTP 200 text/javascriptWS 读回卡片配置(type/entities/key/
attribution/options)全部符合;HA 主机 `curl -4` 带 key 瓦片 200。
- **Follow-up:** Core 升 2026.9.0 stable 后内置地图/zone 编辑器自动切
OSMF 矢量瓦片;届时可保留 custom 卡(继续 keyed CARTO)或用备份还原
内置卡。zone 编辑器等其余内置地图的水印在 2026.9 前无解。
## Known issues
**Bluetooth hci0 instability — RTL8821CS (verified 2026-08-13, W1N-74):**
The local Bluetooth controller hci0 is an **RTL8821CS** combo chip on the
x88 Pro board. Kernel logs show recurring `hci0: hardware error 0x00`,
`Opcode 0x200c tx timeout` (HCI_LE_Set_Scan_Parameters), `Unable to disable
scanning: -110`, `Peer device has reset` — the chip hardware-stalls during
active scanning. HA's `bluetooth_auto_recovery` power-cycle then times out
after 5 s and retries every ~2 min:
`bluetooth_auto_recovery.recover: Could not reset the power state of the
Bluetooth adapter hci0 ... due to timeout after 5 seconds`. The HAOS image
already ships custom systemd units to cope (`x88-bt-hci-recovery.service` and
a "Patch HA Bluetooth scanner mode for x88 RTL8821CS" service, visible in host
journal). **No user impact:** there are **no BLE entities** in HA
(xiaomi_ble / bthome / led_ble / bluetooth / esphome domains are all empty;
platforms merely load from stray advertisements). Real IoT devices are Zigbee
(via Zigbee2MQTT) or WiFi/MQTT/cloud. An ESPHome Bluetooth-proxy ESP32
(`/config/esphome/bluetooth.yaml`, bluetooth_proxy: active, WiFi `ubnt-haas`)
is configured but currently offline (ESPHome add-on stopped, port 6053
unreachable) and produced no entities. Follow-up (optional): disable the
local adapter and rely on the ESPHome proxy, or stop the bluetooth
integration entirely.
**eMMC disk lifetime 10% (verified 2026-08-13, W1N-76):** `ha host info`
reports `disk_life_time: 10` — the boot eMMC (`/dev/mmcblk2`, CJTD4R
`0xacacc064`, 64 GB) has ~10% life left. `disk_free: 40.2/56.4 GB`. Full
backup `pre-maintenance-20260813` (slug `411a4ba5`, 144.26 MB) taken
2026-08-13 covers current config; monitor `disk_life_time` on each health
snapshot and plan a disk replacement / data-disk migration before the eMMC
fails.
## Matter Server (verified 2026-08-21)
- Add-on `core_matter_server` (`homeassistant/aarch64-addon-matter-server`) runs the Matter
commissioner on this host (host networking; add-on container `app_core_matter_server`).
- **After the ISP PD prefix rotates (PPPoE redial), the add-on can cache a stale IPv6 GUA
in its mDNS advertisement** — clients trying that dead address make Matter
commissioning/connection fail. Fix: restart the add-on so it re-enumerates addresses:
`ssh hassio@hass.windy.lan 'sudo -n -i ha apps restart core_matter_server'`
(`ha addons restart ...` also works; "addons" is deprecated in favor of "apps").
- Verified 2026-08-21 (W1N-207): stale `240e:3bd:234:2f22:*` AAAA in mDNS removed by
restart; advertisement now carries only current GUA `240e:3bd:235:1fb2:*` + link-local;
CASE sessions with Aqara M3 / SmartThings hubs resumed over IPv6 link-local.
> **Open items (2026-08-21, W1N-207):** a phone on LAN55 was querying five known
> `_matter._tcp` instances of which only HA answered — the other Matter nodes are
> offline / not announcing (device-side; user to confirm power/Wi-Fi). HA's IPv6
> default route via NetworkManager was observed missing once (curl -6 intermittent,
> while ping6 and `curl -6 --noproxy` work) — not the Matter root cause; re-check
> on the next health snapshot.
Verified 2026-08-23 (read-only, W1N-207): add-on `started`, version `9.0.4`, no
update pending; current GUA `240e:3bd:238:4812:*` (PD rotated again since 08-22)
advertised correctly over v4+v6. Both ESP32-C2 bulbs now announce `_matter._tcp`
(multi-fabric, including this host's fabric `DCE86145C137AF0E`) — but they
**refuse TCP 5540 on IPv4 and IPv6**, so matter-server holds **zero established
:5540 sessions** (device-side failure mode C; no errors logged — see
[docs/matter-pairing-troubleshoot.md §8](../docs/matter-pairing-troubleshoot.md)).
## 马桶换气电源(Matter 插座,半计量)+ 电量估算 (2026-09-13)
**设备**Matter `Smart Plug`SIXWGH`model_id 3596`hw 1.0 / sw 1.3.0),node 18
(0x12)`device_id 5ef1850953466d6e7a9c6b901fbebe1c`config entry
`01JF51VQ48PGJGXX3RNAG6MVAA`,区域**卫生间** (`wei_sheng_jian`)label `power`
2026-09-13 17:58 CST 配对。实体:
`switch.wei_sheng_jian_ma_tong_huan_qi_dian_yuan`(插座)、
`sensor.…_dian_yuan`(电源 W)、`sensor.…_dian_ya`(电压 V)、
`sensor.…_you_gong_dian_liu`(有功电流 A)、`sensor.…_dian_li`(电力 kWh
**永久 unknown**)。
**根因(实测 Matter 属性,node 18**:电量簇 0x0091 `FeatureMap = 13`
(IMPE|CUME|PERE,即**声明**支持导入/累计/周期电量),但
`CumulativeEnergyImported (0x0001)` 恒为 `null``PeriodicEnergyImported
(0x0003)` 带载也恒为 `{Energy: 0}``CumulativeEnergyExported (0x0002)`
不存在(EXPE 未声明,自洽)。HA 只用 `CumulativeEnergyImported` 建能量实体
`components/matter/sensor.py:1083``allow_none_value=True`)→ 该实体
**永远不会出数**。**功率计量本身正常**:0x0090 `FeatureMap = 2` (ALTC)
Voltage / ActiveCurrent / ActivePower 都随负载变化(实测 220.3 V / 118 mA /
24.7 WHA `电源` 0.0→24.9 W 有历史)。厂商 `update` 实体报无新固件。
**处理(方案 A:功率积分补电量)**
- 新建 **Integration (Riemann sum) 辅助元素**config entry
`01M2D53T188FW8WEC547ENHSVH`domain `integration`state `loaded`),
source `sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_dian_yuan`
`method: trapezoidal``unit_prefix: k``unit_time: h``round: 3`
`max_sub_interval: 60s`
- 实体 `sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_energy`(创建时 HA
自动生成 `…_dian_yuan_ma_tong_huan_qi_dian_yuan_dian_liang`,随后立即
`config/entity_registry/update` 改名为 `<插座>_energy` 以对齐约定;
该实体新建、无引用,改名安全),friendly name「马桶换气电源 电力」,
unit kWh、`device_class: energy`、**`state_class: total`**——能源仪表盘
允许 `TOTAL``TOTAL_INCREASING``components/energy/validate.py:279`)。
- **能源仪表盘** (`/energy`)grid 源 `[8]``…_dian_li` 改为 `…_energy`
其余 8 条插座源未动。注意这 9 条「插座」全部以 `type: grid` 注册,被当作
全屋用电代理;`switch` 卡片所在的 Grid 卡片此前第 9 行是空的,即本次修复点。
- **Quick 仪表盘**:「用电量(按插座)」图第 11 项由 `…_dian_li` 改为
`…_energy`;新增 `column_span: 2` 的「开关」区块(heading + tile
`switch.…` + `toggle` feature + 功率徽标)→ 视图 6→7 分区。
**口径警告**`…_energy` 是**估算值**Riemann 积分,只在 HA 运行期间累计、
非账单级),与另外 8 个原生计量插座的累计电量口径不同;功率传感器更新
间隔约 510 s(实测 24.9/24.8/25.0 W 抖动),加 `max_sub_interval: 60s`
保证静默时也继续累计。
**Agent 侧建辅助元素的方法(2026-09-13 实测)**HA 的 config flow 走
**REST**WS 只有 `config_entries/flow/progress|subscribe`,没有 start)。
经 supervisor 代理即可,无需 HA 长连接/长寿命 token:
```bash
# SUPERVISOR_TOKEN 由 sudo -n -i 提供
curl -s -X POST -H "Authorization: Bearer $SUPERVISOR_TOKEN" \
-H "Content-Type: application/json" -d '{"handler":"integration"}' \
http://supervisor/core/api/config/config_entries/flow # → {flow_id, step_id:"user", data_schema}
curl -s -X POST -H "Authorization: Bearer $SUPERVISOR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name":"…","source":"sensor.x","method":"trapezoidal","round":3,
"unit_prefix":"k","unit_time":"h","max_sub_interval":{"minutes":1}}' \
http://supervisor/core/api/config/config_entries/flow/<flow_id> # → create_entry
```
`auth/long_lived_access_token` 在 supervisor 代理身份下**失败**
`unknown_error`),故无法用长寿命 token 开浏览器会话;`DurationSelector`
的值是 `{"minutes":1}` 形式(`cv.time_period`)。
**备份/回滚**`.lovelace-backups/dashboard-quick-20260913-181251-pre-ma-tong-plug.json`
(改动前原件)、`…-20260913-183210-pre-repoint.json`(改名/换源前);
`.ha-backups/energy-20260913-183135-pre-ma-tong-repoint.json`(能源 prefs)。
回滚 = 把能源 prefs 的源 [8] 指回 `…_dian_li` + 还原 Quick 面板 JSON
如需彻底放弃估算电量 = 删除 config entry `01M2D53T188FW8WEC547ENHSVH`
**验证 (2026-09-13 18:3x)**`…_energy` 0.002→0.003 kWh 且随 24.6 W 负载
增长(换气扇关掉后回落 0.0 W,累计值保留);`recorder/list_statistic_ids`
已含该实体;Quick 面板 WS 读回 7 分区、用电量图 11 项指向新实体、旧
`_dian_li` 引用 0 处;能源 prefs 读回 9 源、第 9 条为新实体。
**`energy/validate` 已全绿**9 源 0 issue):创建后 ~5 min 内曾报
`statistics_not_defined`(recorder 的统计任务周期是 5 min,`statistics_meta`
行由该任务建立),18:39 复核时已自动消失——建辅助元素后**不要**把这条
瞬时告警当作失败。
## Related docs
- [runbooks/home-assistant-maintenance.md](../runbooks/home-assistant-maintenance.md) — `ha` CLI maintenance runbook + [script](../runbooks/scripts/ha-maintenance.sh); custom-component zip install is §7
- [docs/lan-overview.md](../docs/lan-overview.md) — LAN map and gw port-forward
- [hosts/dns.windy.lan.md](dns.windy.lan.md) — `hass.windy.lan` / `hass.local` rewrites
+73 -3
View File
@@ -56,7 +56,8 @@ See full shape in [docs/pdns-upstream.md](../docs/pdns-upstream.md). Live secret
| `poweradmin` | poweradmin | Up (healthy) | `poweradmin/poweradmin:stable` |
| `pdns_pgweb` | pgweb | Up | `sosedoff/pgweb:0.16.2` |
| `pdns-backup` | backup | Up | `postgres:16` (scheduler) |
| `powerdns-admin` | *(orphan)* | Exited | legacy PDA UI — not in active compose |
> Legacy PDA UI container `powerdns-admin` (orphan, Exited) was removed 2026-08-12 (W1N-59).
### Network model
@@ -100,9 +101,60 @@ See full shape in [docs/pdns-upstream.md](../docs/pdns-upstream.md). Live secret
**Quirk:** `backend` is internal — backup must not use Alpine + runtime `apk`/`crond`. Uses `postgres:16` + `backup-scheduler.sh` (fixed 2026-08-01).
## Other software on this host (stubs)
## RustDesk Server
`/opt/traefik`, `/opt/adguard`, `/opt/remark42`, `/opt/rustdesk`, `/opt/nginx-manager`, …
**Status: operational** (hbbs + hbbr Up; image pinned `1.1.14`; relay address fixed 2026-08-12, W1N-59).
| Item | Value |
|------|--------|
| Install path | `/opt/rustdesk` |
| Compose | `/opt/rustdesk/compose.yml` |
| Containers | `hbbs` (rendezvous), `hbbr` (relay) |
| Image | `rustdesk/rustdesk-server:1.1.14` (pinned) |
| Relay (hbbr) | `hk2.chans.xyz:21117` — advertised to clients via `hbbs -r` |
| Rendezvous (hbbs) | `21115/tcp` (NAT test), `21116/tcp+udp`, `21118/tcp` (ws) |
| Relay (hbbr) | `21117/tcp`, `21119/tcp` (ws) |
| Public IP | `154.36.174.161` |
| Health | [runbooks/rustdesk-health.md](../runbooks/rustdesk-health.md) |
**Note:** the relay hostname in `hbbs -r` must resolve to this host's public IP
(`154.36.174.161`). `hk2.chans.xyz` resolves correctly; the previously used
`hk2.wsvc.info` had **no DNS record** and broke relay connectivity for clients
(fixed 2026-08-12, W1N-59).
## Other software on this host (confirmed 2026-08-12)
Verified live via `docker ps` / port scan. Each runs as a separate compose
project under `/opt/<name>` and is fronted by Traefik where noted.
| Service | Path | Container(s) | Image | Ports / notes |
|---------|------|--------------|-------|---------------|
| Traefik | `/opt/traefik` | `traefik` | `traefik:v3.6.2` | `80`, `443` (TLS entry), `8080` (dashboard) |
| AdGuard Home | `/opt/adguard` | `adguardhome` | `adguard/adguardhome:latest` | DoH `5443`, DoT `853` (bridge; no LAN `:53`) |
| Remark42 | `/opt/remark42` | `remark42` | `ghcr.io/umputun/remark42:latest` | no host ports; via Traefik (in-container `8080`) |
### Traefik dashboard auth
| Item | Value |
|------|-------|
| Dashboard URL | `https://npm.chans.xyz` (Traefik `api@internal` router), also host `:8080` |
| Auth | HTTP Basic via Traefik `basicauth` middleware (label `dashboard-auth`) |
| User | `windy` — stored as a **bcrypt** hash (plaintext never stored) |
| Hash generator | `/opt/traefik/generate-dashboard-auth.sh` (bcrypt; auto `$``$$` compose escaping) |
| Config | `/opt/traefik/compose.yml` (label `traefik.http.middlewares.dashboard-auth.basicauth.users`) |
**Password rotated 2026-08-12** from apr1/MD5 to bcrypt via the generator script; the
plaintext lives only in the operator's password manager, never in this repo.
To rotate again: `cd /opt/traefik && ./generate-dashboard-auth.sh windy`, paste the
printed label into `compose.yml`, then `docker compose up -d --force-recreate traefik`.
`/opt/nginx-manager` was a leftover (compose + `data/` + `letsencrypt/`, no running
container) and was **removed 2026-08-12**; pre-deletion backup:
`/opt/backups/nginx-manager-20260812.tar.gz`.
Health coverage: these auxiliary services are checked by the `hk2aux`
health-check profile (`ansible/roles/healthcheck`). Run:
`cd ansible && ansible-playbook playbooks/health-report.yml --limit powerdns`.
## Ops / runbooks
@@ -133,6 +185,23 @@ dig @202.91.35.141 SOA wsvc.info +short
On-server docs: `/opt/pdns/README.md`, `CHANGELOG.md`.
## Disk / logging (VPS-81, 2026-09-02)
Root disk cleanup performed (runbook: [host-disk-cleanup](../runbooks/host-disk-cleanup.md)):
- Root `/` (20G vda1): 76% used → **38% used** (15G → 7.1G; free 4.7G → 12G).
- **AGH log flood root cause fixed**: `/opt/adguard/conf/AdGuardHome.yaml`
`log.verbose: true → false` (backup `AdGuardHome.yaml.bak-20260902-vps81`).
Verbose debug was streaming to stderr → container `json.log` (~120MB/day);
`log.file: ""` makes AGH's own rotation keys inert. Restart only (no recreate).
- Journald capped: `/etc/systemd/journald.conf.d/00-vps81.conf`
`SystemMaxUse=200M`; journal vacuumed to ~96M.
- Docker: engine **29.7.2**; 14 unused images removed (kept `pdns-auth-50:5.0.5`
rollback pin); 12 orphan anonymous volumes + build cache pruned. In-use
volumes intact (`pdns_dbdata`, `b594d738…` PG data, `e855d078…` backup).
- Follow-up: re-check AGH `json.log` growth **2026-09-09** (one-week checkpoint);
global docker log rotation only if still needed.
## Verified
Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded.
@@ -142,3 +211,4 @@ Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded.
- `only-notify=` + `also-notify=202.91.35.141`; MASTER `domains.master` cleared
- https://pdns.wsvc.info → **302**; https://pgweb.wsvc.info → **401**
- Hardening backlog: API/DB credential rotation + TSIG rotate (see upstream doc)
- 2026-09-02 (VPS-81): post-cleanup verified — 10 containers Up (adguardhome healthy), DNS SOA/NS + web endpoints OK; see Disk/logging section above.
+59
View File
@@ -0,0 +1,59 @@
# pgdb — TimescaleDB (PG18, Docker)
## Role and access
| Item | Value |
|---|---|
| Role | TimescaleDB PostgreSQL 18 (Docker) — Home Assistant recorder 后端(`hass`/`scribe` 库) |
| IPv4 | `192.168.55.15` (LAN55) |
| DNS | (none) |
| SSH | `ssh -4 windy@192.168.55.15`key auth 已验证可用 2026-08-29agent 沙箱用 `ssh -F /dev/null -o BatchMode=yes`password auth 亦可) |
| Host | PVE 管理的 QEMU VMi440FX**VMID 100**),Debian 13 (trixie),内核 6.12.105;宿主机 **pve2 `192.168.55.25`**Proxmox 9.2.2SSH `root@192.168.55.25``onboot: 1`QEMU guest agent 已装;2026-08-31 补记) |
| Resources | 3 GB RAM08-30 13:58 由 2G 上调、删除 balloon/ksm/shares 后重启生效)/ 30 GB disk26 G 空闲) |
| Docker | 29.7.2;容器 `timescaledb` = `timescale/timescaledb:latest-pg18`PG **18.6** + TimescaleDB **2.29.2**Apache-2.0 版) |
| Ports | `192.168.55.15:5432`PGIPv4 only);`192.168.55.15:8081`pgweb GUIbasic auth |
## Databases
| DB | Owner | Size | 用途 |
|---|---|---|---|
| `hass` | hass | ~406 MB2026-09-13 | HA recorderstates/events/statistics),客户端 HAOS `192.168.55.11` |
| `scribe` | postgres | ~2.6 GB2026-09-13 | HA scribe 集成(entities/areas/devices 注册表同步 + `states_raw`/`events` hypertable + `csg_history` 长期归档表);体积由 `sensor_minute` 图表管道主导(2.26 GB),见 Known issues |
| `postgres` | postgres | ~9 MB | 默认库 |
## Ops notes
- **Docker compose 管理**2026-08-29 改造):`/opt/database/docker-compose.yml`(源码在仓库 `compose/pgdb/`+ `/opt/database/.env`0600,密钥)+ `/opt/database/pgweb-bookmarks/`0600bookmark 含 DB 密码)。三个服务:
| 服务 | 镜像 | 端口 | 说明 |
|---|---|---|---|
| `timescaledb` | `timescale/timescaledb:latest-pg18` | `192.168.55.15:5432`IPv4 only | PG 18.6 + TS 2.29.2healthcheck pg_isready`restart: unless-stopped` |
| `pgweb` | `sosedoff/pgweb:latest`v0.17.0 | `192.168.55.15:8081` | Web GUIhttp://192.168.55.15:8081basic auth(用户名/密码见 .env `PGWEB_AUTH_USER/PASS`);`--readonly --sessions --bookmarks-only --bookmarks-dir /bookmarks`v0.17.0 不读 PGWEB_BOOKMARKS_DIR env,必须用 flag);bookmarks = hass/scribe |
| `pg-backup` | `prodrigestivill/postgres-backup-local:latest`(=PG18 客户端) | — | 每日 02:00(`TZ=Asia/Shanghai`,本地时区)`pg_dump -Fc` 三库 → `/opt/database/backups/{daily,weekly,monthly}`;保留 7 天/4 周/6 月;`BACKUP_ON_START` |
- **数据盘**`/dev/sdb1`32G ext4label `pgdata`)挂载 `/srv/pgdata`fstab 按 `UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86`defaults,noatime)持久化(2026-08-29 迁移)。容器 bind mount `/srv/pgdata:/var/lib/postgresql`
- 容器内 postgres 用户 uid/gid = **70**(Debian 系,非 999);迁移数据后需 `chown -R 70:70`
- **密码**:postgres 超级用户已换强密码(hex,存 `/opt/database/.env` 06002026-08-29)。HA 用 `hass` 角色不受影响。
- **备份**:由 `pg-backup` 容器接管(2026-08-29),宿主机 cron 与 `/opt/database/pg-backup.sh` 已退役。恢复用 `pg_restore`custom format)——2026-08-29 已实测还原 hass 库 dumpstates 10014 行)成功。
- **认证**:外部连接 scram-sha-256(密码必填,改密码有效);容器内 loopback 为 trust(官方镜像默认)。
- **回滚**:旧启动命令保留在 `/opt/database/run`(容器无状态,数据在 /srv/pgdata);旧匿名卷 `9375195843b950f4e04c34872409ca095e1136520dd019a8e86e2794be06c236`(根盘 ~82M)保留作兜底,确认稳定后可 `docker volume rm`
- **开机自愈**2026-08-30):新增 systemd oneshot `pgdb-compose.service`enabled,源码在仓库 `compose/pgdb/pgdb-compose.service`):`After=network-online.target docker.service`,开机后幂等执行 `docker compose up -d`,重试直到 `192.168.55.15:5432` 监听,重试耗尽 `--force-recreate` 兜底(数据在 bind mount,无损)。原因:2026-08-30 开机竞态——docker 恢复容器时 VM IP 尚未可绑(EADDRNOTAVAIL),timescaledb/pgweb 启动失败且 docker 不重试。手动重跑:`sudo systemctl restart pgdb-compose.service`
- 本机无防火墙(ufw/nft/iptables 均未装)——待办:如要彻底隔离可加 ufw 白名单 192.168.55.11。
- `/opt/database/backups/` 根下残留 `*-2026-08-29_1359.dump`(compose 化之前旧备份机制产物)与 `backup.log`——健康检查只看 `daily/`,残留可清理。
- **Runbooks**[pgdb-health](../runbooks/pgdb-health.md)(只读健康检查)、[pgdb-restore](../runbooks/pgdb-restore.md)pg_restore 还原)、[pgdb-update](../runbooks/pgdb-update.md)(镜像/compose 升级)。
- **CSG 长期归档(2026-08-29, W1N-243**`csg_history` 表(`period date / kind('day'|'month') / usage_kwh / cost / ladder / balance / updated_at`PK(period,kind)`GRANT SELECT TO hass`)保存南方电网有价值数据:day = 逐日(昨日用电/费用/阶梯/余额,2026-07-01 起),month = 当月累计(用电/费用,2025-01 起)。由 TimescaleDB 每日任务 **1008** `csg_daily_snapshot()`22:30 Asia/Shanghai**TS job 非 pg_cron**,本库未装 pg_cronupsert 维护:取「最新有值行」防瞬态 unknown 竞态;日费用缺原生 `latest_day_cost` 时回退 = 昨日用电 × 当前档费率(模板 `csg_current_ladder_tariff` 0.639);月费用回退模板 `csg_this_month_ladder_cost`。验证:day 08-28 = 7.66 / 4.89474 / 二档 / 0month 08 = 302.47 / 180.28。回填来源:集成 attributes `history_data`59 天)+ `by_month`(19 月)——08-29 前唯一残存历史。回滚:`DROP TABLE csg_history` + `SELECT delete_job(1008)`
## Known issues
- 2026-09-13**`sensor_minute` 体积构成与压缩窗口(只读诊断,暂不处理)**。`scribe` 库 2.6 GB = `sensor_minute` **2.26 GB**850 万行 / 16 天,约 5659 万行/天 = 331 实体 × 1440 分钟 LOCF+ `states_raw` 290 MB + `events` 1.5 MB`hass` 库另 406 MB。2.26 GB 中 1.50 GB 是 chunk `[09-03,09-10]`、0.76 GB 是 `[09-10,09-17]`,**都还没到压缩窗口**——TimescaleDB 的 `compress_after`**chunk 结束时间**判断(09-10 结束 + 7 天 = **09-17** 才合格),所以「7 天 chunk + 7 天 compress_after」的设计下限就是盘上常驻近 14 天原始数据;已压缩的 `[08-27,09-03]` 从 1.04 GB → **1.5 MB**LOCF 重复度极高,~700:1)。任务 1005 健康(30 成功 / 0 失败,最近 09-13 04:18 跑过但无合格 chunk);1002/1003/1006/1007 亦全 Success。稳态估算 ≈ 2 个未压缩 chunk(3–4.5 GB+ 已压缩归档(约 1.5 MB/周 ≈ 80 MB/年)≈ **45 GB 平台期**pgdata 卷 32 G 当前用 3.2 G,可用 27 G,**无需处理**。复查点 **2026-09-17 之后**`_hyper_4_6_chunk` 应转为 `compressed=true` 且库体积回落;若仍为 false 才需动手(手动 `compress_chunk()` 或调小 `compress_after`)。可选调优:chunk 间隔 7 天 → 1 天 + `compress_after` → 2 天,把常驻未压缩量压到 <1 GB(`set_chunk_time_interval` 只对新 chunk 生效,旧 chunk 不重切)。诊断命令:`select chunk_name, is_compressed from timescaledb_information.chunks where hypertable_name='sensor_minute';` + `pg_database_size('scribe')`。**注意:这是 pgdb 侧对象,HA/scribe 的 `retention_states` 管不到它;HA 侧唯一杠杆是少记/少画(等于砍图)。**
- 2026-08-29HA 侧 HACS 集成 `custom_components.scribe`YAML `scribe: db_url:`,连 `scribe` 库)建表被拒(`permission denied for schema public`hass 无 CREATE 权限),之后持续报 `relation "entities" does not exist`。**已解决**:① `GRANT CREATE ON SCHEMA public TO hass;`scribe 库)② 重启 HA Core 触发重跑建表。重启后自动创建 `entities`1591 行)/`users`/`areas`/`devices`/`integrations`/`states_raw` 表并启用 TimescaleDB 时间序列能力。报错已停止(最后一条 06:06 UTC),`states_raw` 持续写入。2026-08-29 复查:scribe 现有**两个** hypertable——`states_raw`segmentby `metadata_id`、orderby `time`)与 `events`segmentby `event_type`、orderby `time`),均 1 维 `time`;压缩已配置(`timescaledb_information.compression_settings` 可见对应行;2.29.x 该视图无 `compression_enabled` 列)。
- 2026-08-29**timescale reader 图表对象**(配套 hass 的 `timescale_database_reader` 集成 + `timescale-plotly-card`,上游 SQL `remmob/timescale_database_reader` `SQL/scribe/01+02` @ `bb8776a`,以 postgres 执行):`sensor_minute_aggregate` 连续聚合(1 分钟桶,last(state)/last(value),实时聚合开启)+ `sensor_minute_aggregate_entity` 视图(join `entities`+ `sensor_minute` hypertable`minute`/`entity_id`/`state`/`value`,LOCF 前向填充)。任务:1005 `sensor_minute` 压缩(7 天)、1006 `sensor_minute` 保留(10 年)、1007 `every_minute_refresh` 每分钟增量刷新(含 5 分钟回溯窗口修正)。授权:`GRANT SELECT ON sensor_minute_aggregate, sensor_minute_aggregate_entity, sensor_minute, entities TO hass`。种子 19529 行(331 实体,自首个数据点起)。**刻意跳过**了上游脚本对 `states_raw` 的 3 个月保留 + 压缩策略语句——与"`states_raw` 永久归档"定位冲突,如需磁盘回收属用户决策(scribe 自己的压缩任务 1000/1001 未动)。
- 2026-08-29**`sensor_minute_refresh` 本地补丁(类比 tianqi 补丁,重跑上游 02 SQL 后需重打)**:值 CASE 的 `ELSE 0``ELSE NULL`。原因:scribe 对 unavailable 分钟 value 为 NULL,上游刷新过程兜底写 0;对差分模式的用电图,0→计数器回升会把插座的**生命周期累计值**(最高 1588 kWh)算进掉线那一小时。同日一次性清理既有脏 0:头部占位行 DELETE 505 行(各实体首次非零分钟之前的 value=0);`sensor.%_energy` 与温湿度实体的 value=0 → NULL(10+16 行,物理上不可能的真 0,图表渲染为断点)。功率实体的中途 0 是真实待机读数,保留。
- `hass` 库的 recorder 表仍为普通表(无 hypertable);`scribe` 集成负责时间序列历史(`states_raw` + `events` hypertable)。
## Verification history
- 2026-08-31**13:58 重启根因确认,非停电**W1N-263):pve2`192.168.55.25`)任务日志显示 08-30 **13:58:00 `root@pam` 在 PVE Web UI 修改 VM 100 配置**`-delete allow-ksm,balloon,shares -memory 3072`),**13:58:06 点 Reboot**`qmreboot` → 客机 13:58:08 干净 ACPI 关机 → 13:58:13 自动重启)。宿主机全程在线(08-30 09:00 开机至今连续运行 1d12h+),`.66.26` PVE 及各 VM 均无重启——排除停电。HA recorder 在窗口(13:58:4647)报 2 次 `Connection refused`,DB 恢复后自动重连,**无数据丢失**(`hass.states`/`scribe.states_raw` 13:5514:02 逐分钟无缺口,recorder 内存队列吸收回写)。13:58:47 三容器已起,13:58:56 自愈单元 `pgdb-compose.service` 执行成功——本次自愈按设计工作。同日下午 12:54–12:55 另有一次**客机内自重启**(无 PVE 任务,工作站 SSH 会话相邻)。08-29 22:19→08-30 09:00 宿主机停机 10h41m 为**干净关机**(systemd 有序关闭,非停电)。
- 2026-08-30**开机竞态故障 + 修复**W1N-260):09:01 开机后 docker 恢复容器时绑定 `192.168.55.15:5432/8081` 失败(EADDRNOTAVAIL)→ timescaledb/pgweb 停摆至 12:16pg-backup 开机备份失败(解析不到 timescaledb)→ unhealthy。12:22 `docker compose up -d --force-recreate` 修复(三容器回 `database_default`、端口发布、今日备份、pgweb 恢复);用户重启 HA Core 后写入管道恢复。12:43 新增开机自愈 unit `pgdb-compose.service`enabled,已实测幂等 reconcile)。pgdb-health 8 项全绿。
- 2026-08-29:首次检查(只读)+ 修复 scribe 权限 + 安装夜间备份。见 Linear vps 项目登记。
- 2026-08-29**compose 改造完成**W1N-227,用户已验收):裸 `docker run``/opt/database/docker-compose.yml` 三服务(timescaledb + pgweb + pg-backup);superuser 换强密码;端口收紧 IPv4;备份容器化(TZ=Asia/Shanghaicron 02:00 本地);`pg_restore` 还原实测通过;pgweb UI 用户确认可查 hass/scribe 数据。源码在仓库 `compose/pgdb/`
- 2026-08-29**运维 runbook 落地**W1N-228,已验收):新增 `runbooks/pgdb-health.md`(只读,8 项诊断全绿)、`pgdb-restore.md`(流程式,temp-DB 安全还原 + 审批门)、`pgdb-update.md`(门控命令式,回滚=/opt/database/run + 旧卷);README 索引与 validate-repo.sh 分类同步更新;runbook 命令已对活主机逐条实测(含 `pg_restore -l` 校验当日 dump)。同日修正:SSH key auth 可用(facts 原记"密钥未安装"已过时);scribe 新增 `events` hypertable。
- 2026-08-29**CSG 长期归档 + recorder 365d**W1N-243):建 `csg_history` 表 + attributes 回填(逐日 59 + 逐月 19)+ 每日任务 1008(函数 v2:最新有值行读取、日费用阶梯回退);hass `purge_keep_days` 30→365(备份 `configuration.yaml.bak-20260829-purge365`)。见 Linear vps W1N-243。
+45 -1
View File
@@ -24,6 +24,7 @@ ssh -4 windy@synapse.chans.xyz
| DB | ESS embedded PostgreSQL 17 (PVC 20Gi, local-path) |
| Cache | ESS embedded Redis (PVC 2Gi) |
| Chart | `oci://ghcr.io/element-hq/ess-helm/matrix-stack`, version `26.7.2` |
| Plane | Helm `plane-ce-1.8.0` (app `v1.4.1`), namespace `plane` — self-hosted Plane project management |
### Matrix service endpoints
@@ -51,8 +52,50 @@ All other ports internal only (no K3s API, no database, no Redis exposed).
- `ess` — all ESS workloads (Synapse, MAS, Element, Postgres, Redis, HAProxy)
- `matrix-system` — cluster base resources (ResourceQuota, LimitRange, mrtc-placeholder)
- `plane` — Plane project management (Helm release `plane-app`)
- `cert-manager` — cert-manager
## Plane (project management)
Self-hosted [Plane](https://github.com/makeplane/plane) on the same K3s node, deployed via the official `plane-ce` Helm chart.
| Item | Detail |
|------|--------|
| Release | `plane-app` (ns `plane`), chart `plane-ce-1.8.0`, app `v1.4.1`, revision 1 |
| URL | https://plane.chans.xyz |
| Install date | 2026-09-01 |
| Values source | `/home/windy/plane-k3s/values.yaml` (plain file, not a git repo) |
| Images | `artifacts.plane.so/makeplane/*` (`plane-frontend`, `plane-backend`, `plane-admin`, `plane-live`), pullPolicy `Always` |
| Ingress | Traefik `IngressRoute` `plane-app-ingress``/`→web, `/api` `/auth`→api, `/spaces`→space, `/god-mode`→admin, `/live`→live, `/uploads`→minio; `maxRequestBodyBytes` 20Mi |
| TLS | Own namespace `Issuer` `plane-app-cert-issuer` (HTTP-01, LE prod, `admin@chans.xyz`); cert `plane-app-ssl-cert` (CN `plane.chans.xyz`) |
| DB | Bundled Postgres `15.7-alpine` (PVC 5Gi, local-path) |
| Cache/queue | Bundled Redis (PVC 100Mi), RabbitMQ `3.13.6-management-alpine` (PVC 100Mi) |
| Storage | Bundled MinIO (`minio/minio:latest`, root user `admin`, PVC 5Gi) — S3 for uploads/docs |
| Resources | Every workload: cpu 50m/500m, mem 50Mi/1000Mi, replicas 1 |
| SMTP | Not configured (no `smtp` values) — Plane invites/password resets won't email yet |
Workloads (all 1/1 Running): 7 Deployments (`plane-app-{admin,api,beat-worker,live,space,web,worker}-wl`) + 4 StatefulSets (`plane-app-{minio,pgdb,rabbitmq,redis}-wl`); init Jobs `api-migrate-1` / `minio-bucket-1` Completed. All PVCs Bound on `local-path` (root disk).
### Plane configuration notes
- **`planeVersion: v1.4.1`** pinned in values.yaml; chart tracks Plane's own tags.
- **Secrets**: Helm-generated Opaque secrets (`plane-app-app-secrets`, `-doc-store-secrets`, `-pgdb-secrets`, `-rabbitmq-secrets`, `-live-secrets`); `requireExplicitSecrets: false`. Values live in `$SECRET_KEY`, `DATABASE_URL`, `AMQP_URL`, `REDIS_URL` etc.
- **Sentry / CORS**: `sentry_dsn` and `cors_allowed_origins` empty (defaults fine for single-host).
- **MinIO is `latest` tag** — pin a version for reproducibility.
- **Backup**: NOT covered by `/var/backups/matrix` (which is paused anyway) — Plane Postgres/MinIO PVCs have no backup tier yet.
### Plane verification
```bash
# Release + workloads
sudo helm list -A
sudo k3s kubectl -n plane get deploy,sts,pods -o wide
# Cert + ingress
sudo k3s kubectl -n plane get certificate,ingressroute
# Endpoint
curl -4 -s -o /dev/null -w '%{http_code}\n' https://plane.chans.xyz/
```
## Local backup
| Item | Detail |
@@ -62,7 +105,7 @@ All other ports internal only (no K3s API, no database, no Redis exposed).
| Retention | 7 days |
| Disk warning | 80% (healthcheck), 90% (backup stops) |
| Content | Planned: PostgreSQL `synapse` + `mas` logical dumps, media store archive, `/etc/matrix-bootstrap` |
| Status | **Not operational** — no current Matrix backup or recovery tier |
| Status | **Not operational** — no current Matrix backup or recovery tier. **Plane data (its own Postgres + MinIO PVCs in ns `plane`) is also not covered by any backup.** |
## Health checks
@@ -101,5 +144,6 @@ diagnosis and imperative recovery work.
- MatrixRTC / Element Call / LiveKit / Coturn not deployed (`mrtc.chans.xyz` reserved only)
- SMTP email not yet configured (requires manual secret bootstrap followed by a
reviewed Ansible stack deployment)
- Plane `minio` image uses `latest` tag (pin a version)
- No off-site Restic backup
- Single-node K3s (no HA for control plane)
+5
View File
@@ -53,6 +53,11 @@ of `8080`. During adoption or recovery, use the documented `:9080/inform` URL;
an AP left on `:8080` can remain reachable by ping and SSH while showing
offline in the controller.
IPv6 is enabled on the controller's `Default` network (`ipv6_enabled: true`,
client assignment SLAAC; RA is served by `gw`, so `ipv6_interface_type` is
`none`); both managed APs hold global SLAAC addresses — verified 2026-08-20.
See [docs/unifi-network.md](../docs/unifi-network.md).
## Safe reconciliation and verification
```bash
+26 -12
View File
@@ -5,12 +5,12 @@
| Role | Multi-service VPS (Vaultwarden, Traefik, Soft Serve, …) |
| SSH | `ssh -4 windy@us2.wsvc.info` (prefer IPv4 from WSL) |
| IPv4 | `193.9.44.165` |
| Also DNS | `auth.wsvc.info` → this host; `repo.windy.me` → this host (Soft Serve) |
| Also DNS | `auth.wsvc.info` → this host; `repo.windy.me` → this host (Gitea) |
| Public HTTPS | Traefik on `:80` / `:443` (`/opt/traefik`) |
## Vaultwarden (Bitwarden-compatible)
**Status: operational** (Postgres live, HTTPS 200, healthy containers, SMTP AUTH OK — last probe 2026-08-01 18:55 CST).
**Status: operational** (Postgres live, HTTPS 200, healthy containers, SMTP AUTH OK — last probe 2026-08-29).
Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
@@ -21,9 +21,9 @@ Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
| Env file | `/opt/vaultwarden/.env` |
| Admin overrides | `/opt/vaultwarden/vw-data/config.json` (**wins over env**) |
| Public URL / `DOMAIN` | `https://auth.wsvc.info` |
| Image | `vaultwarden/server:1.37.1` (pinned) |
| Image | `vaultwarden/server:1.37.2` (pinned) |
| Live DB | **Postgres 16** (`vw-db` / service `pg`) via compose `DATABASE_URL` |
| Data (probe) | users=1, ciphers=1327 |
| Data (probe) | users=1, ciphers=1360 |
| Cold SQLite | `backups/sqlite-cold/db.sqlite3.pre-pg-20260801` (not used live) |
| Pre-migrate backup | `backups/pre-pg-migrate-20260801_161204/` |
| Data dir | `./vw-data``/data` (attachments, rsa keys, `config.json`) |
@@ -50,7 +50,7 @@ Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
| Container | Status |
|-----------|--------|
| `vaultwarden` | Up (healthy), `vaultwarden/server:1.37.1` |
| `vaultwarden` | Up (healthy), `vaultwarden/server:1.37.2` |
| `vw-db` | Up (healthy) — **live** Postgres |
| `vaultwarden-backup` | Up (`pg_dump`) |
| `vaultwarden-pgweb` | Exited (profile `debug`) |
@@ -80,19 +80,33 @@ ansible-playbook playbooks/compose-reconcile.yml --limit vaultwarden \
| Container | Status | Image / notes |
|-----------|--------|---------------|
| `soft-serve` | Up | `ghcr.io/charmbracelet/soft-serve:latest` (`repo.windy.me:2222`) |
| `gitea` | Up | `gitea/gitea@sha256:1c17ecaead42e…` (1.27.3-rootless) — SSH `repo.windy.me:2222`, web `https://repo.windy.me` |
| `gitea-backup` | Up | alpine + sqlite3/rsync sidecar (daily backup 02:00 / prune 03:00, crond) |
### Gitea (replaced Soft Serve 2026-09-18; [Plane VPS-94](https://plane.chans.xyz))
- `/opt/gitea/compose.yml` (+ `Dockerfile.backup`, `scripts/`, `config/app.ini`, `data/`, `secrets/`, `backups/`); 镜像: `compose/gitea/`(参考, 服务器文件为准)
- **rootless 镜像** uid 1000:1000; SQLite `/opt/gitea/data/data/gitea.db`; repos `/opt/gitea/data/data/git/repositories/`; app.ini `/opt/gitea/config/app.ini`(600, 含 SECRET_KEY)
- SSH: 内置 server 容器内 `:2322`(`SSH_LISTEN_PORT` 非特权), Traefik TCP entrypoint `ssh`(`:2222``gitea:2322`, `HostSNI(*)`, `tls=false`) on `vw-net`; clone URL `ssh://git@repo.windy.me:2222/windy/<repo>.git`(owner 段 `windy`)
- **host key 复用 soft-serve**(`SSH_SERVER_HOST_KEYS=/secrets/soft_serve_host_ed25519`, ed25519, 指纹 `SHA256:PdxZRe74…`): 客户端 known_hosts 零变更; 仅公钥认证(密码认证未启用)
- Web: `https://repo.windy.me`(Traefik websecure + letsencrypt); `DISABLE_REGISTRATION=true`, Actions 关闭; 管理员 `windy`(凭据仅存服务器 `/opt/gitea/.admin-credentials`, 勿入库/入 Plane)
- 仓库: 16 个(顶层 11 + `cdia/` 4 + `windyboy/go-caatsm`), 2026-09-18 自 soft-serve `push --mirror` 迁移, 逐仓 `ls-remote` ref 全集 + HEAD symref 两端一致; 可见性仅 `dotfiles-personal` private, 其余 public(与 soft-serve 现状一致)
- 备份: sidecar 每日 02:00 → `backups/gitea_<TS>/{app.ini.tar.gz, gitea.db, repos.tar.gz}`(app.ini 含恢复必需 SECRET_KEY), 03:00 prune 保留 14 份; 已验证手动备份产物 109.9M
- 回滚: `/opt/soft-serve` 未删(compose stop + sidecar 停, 数据与旧备份冻结保留), 回滚 = Traefik `:2222` 指回 `soft-serve:23231` + 客户端 remote 回改旧无 owner 段路径; 观察 24 周后清理(历史: W1N-244~248)
| `traefik` | Up | `traefik:v3.6.2` (`/opt/traefik`, public `:80`/`:443`) |
| `nghttpx-proxy` + `squid-backend` | Up | HTTP forward-proxy stack (`/opt/nghttpx`), network `nghttpx_internal-net`; details TBD |
Directories for `authelia`, `conduit`, `dendrite`, `mastodon`, `rustdesk`, `zitadel`, etc. exist under `/opt` but have no running containers; treat them as dormant, not documented services.
**Disk cleanup 2026-09-18** ([Plane vps VPS-93](https://plane.chans.xyz)): root 71% → **23%** (~33G freed) keeping soft-serve / vaultwarden / traefik (nghttpx kept running per operator choice). Removed: unused Docker images + orphan volumes (incl. `zitadel_data` 801M), dormant `/opt` dirs (dendrite + its disabled `dendrite.service` unit, mastodon, dailysync, keycloak, media-repo, authelia, conduit, npm, manager, fusion, zitadel, rustdesk), rootless podman storage (6.4G stale goauthentik), home dev caches, apt cache, journal 3.8G→162M (+`SystemMaxUse=200M` drop-in, active next boot), truncated container logs (nghttpx 550M / traefik / squid). Follow-up: nghttpx-proxy logs grow ~25M/day (INFO per-connection); root-cause log-level/rotation fix still open (needs container restart approval).
Remaining running services on this host: `gitea`, `vaultwarden` stack, `traefik`, `nghttpx-proxy` + `squid-backend` (undocumented forward proxy, `/opt/nghttpx`). `/opt/soft-serve` kept stopped as rollback (24 weeks, data intact). `/home/windy/authelia` (76M) left in place — outside approved cleanup scope.
## Verified
Last checked: **2026-08-01 18:55 CST** — operational.
Last checked: **2026-09-18** — operational; disk cleanup done (see note above, Plane vps VPS-93). Prior full probe: 2026-08-29.
- `vaultwarden` + `vw-db` healthy; `DATABASE_URL``pg:5432/vaultwarden`
- `https://auth.wsvc.info/` **200**, `/admin` **200**, `/api/config` OK (`disableUserRegistration: true`)
- Identity wrong-password → **400** business error (DB readable, not 500)
- SMTP: container → `mx2:587` OK; STARTTLS cert CN=`mx2.windy.me`; **AUTH OK** with effective `config.json` password (synced with `.env` / `.smtp-credentials`)
- LE cert CN=`auth.wsvc.info`
- PG counts: users=1, ciphers=1327
- SMTP: container → `mx2:587` OK; **AUTH OK** with effective `config.json` password (synced with `.env` / `.smtp-credentials`, fingerprint match)
- PG counts: users=1, ciphers=1360
- Image `vaultwarden/server:1.37.2` (**upgraded 2026-08-29** from 1.37.1; required for Bitwarden clients 2026.8.0+); post-upgrade 404 fixed by Traefik restart, then 200
- vps-health local check **installed 2026-08-29** (`vps-healthcheck.timer` daily 06:15 + `/usr/local/lib/vps-health/run`); `health-report.yml --limit vaultwarden` now passes (**ok**, was failing due to missing check infra + script bugs fixed: trim_blocks render, pgweb debug-profile false positive, SMTP probe moved host-side since image lacks python3)
+133 -1
View File
@@ -9,10 +9,74 @@
| Compose file | `/opt/wireguard/compose.yml` |
| Container | `wireguard` |
| Image policy | Immutable digest, updated only in an approved maintenance window |
| Public port | UDP `51820` on IPv4 and IPv6 |
| Public endpoint | `us4.wsvc.info:51820/udp`; DNS publishes only A `185.201.226.122` (no native AAAA) |
| Tunnel subnet | `10.13.13.0/24` |
| Routing policy | IPv4-only full tunnel (`ALLOWEDIPS=0.0.0.0/0`); IPv6 traffic is not guaranteed to use the VPN |
Upstream image documentation:
[LinuxServer.io WireGuard](https://docs.linuxserver.io/images/docker-wireguard/).
## Deployment configuration
The repository-owned, non-secret Compose declaration is rendered from
`ansible/templates/wireguard-compose.yml.j2`. The live declaration was verified
on 2026-08-12 with these core settings:
| Setting | Live value / intent |
|---------|---------------------|
| Image | `lscr.io/linuxserver/wireguard@sha256:ac43e1226878d2611315172d6ea357a95cb326ee73124b91108118efc8666889` |
| Image version | `1.0.20260223-r0-ls119` (build 2026-07-30) |
| Required capability | `NET_ADMIN` only; host kernel already supplies WireGuard/iptables, so `SYS_MODULE` and `/lib/modules` are not granted |
| Filesystem | Read-only container root; executable tmpfs at `/run`; writable bind mount `/opt/wireguard/config:/config` |
| Restart | `unless-stopped` |
| Server mode | Named peers `ha`, `phone`, `mbp`; runtime and configured peer counts both `3` |
| Client DNS | `1.1.1.1` |
| Tunnel routing | IPv4 full tunnel, `0.0.0.0/0`; no client IPv6 tunnel |
| Runtime interface | `wg0`, server address `10.13.13.1/32`, listen port `51820` |
| Forwarding/NAT | IPv4 forwarding enabled in the container namespace; `wg0` forwarding allowed and egress masqueraded on `eth+`; IPv6 forwarding disabled |
Docker binds UDP `51820` on both host socket families, but the public hostname
has no AAAA record. Clients using `us4.wsvc.info` therefore reach the server over
IPv4.
## Other host services and firewall (2026-08-12)
This host also carries the `windy.me` secondary MX and several web applications;
do not build its firewall allowlist from the WireGuard role alone.
| Port | Owner / purpose | Effective public state |
|------|-----------------|------------------------|
| TCP `22` | SSH management | Open |
| TCP `25` | Postfix, `mx.windy.me` (MX priority 30) | Open; retain until the secondary-MX role is explicitly retired |
| TCP `80`, `443` | Traefik for `update.wsvc.info`, `us4-gate.wsvc.info`, and `trlm.wsvc.info` | Open |
| TCP `3000` | Semaphore UI direct Docker publish | Open; redundant with the Traefik route and should be removed or bound to loopback |
| TCP `8080` | Traefik direct Docker publish | Open; redundant with the authenticated dashboard route and should be removed or bound to loopback |
| UDP `51820` | WireGuard | Required public endpoint |
| TCP `9443` | Host nghttpx-to-Squid proxy | Listening but blocked by the current firewall |
| UDP `123` | ntpsec | Listening but blocked by the current firewall |
PostgreSQL (`5433`/`5434`/`5435`), MariaDB (`3306`), and the host Squid TCP
listener (`3128`) are loopback-only. Squid also owns wildcard UDP sockets, which
are not allowed by the current public zone.
UFW is not installed. Firewalld `2.3.1` is active with nftables. On 2026-08-12,
the reviewed `ansible/playbooks/us4-firewalld.yml` reconciliation removed the
stale `imap`, `imaps`, `smtp-submission`, and `smtps` services plus TCP `24`,
`6443`, and `8443` without reloading or restarting firewalld. Runtime and
permanent public-zone state now match exactly: services `dhcpv6-client`, `http`,
`https`, `smtp`, and `ssh`, with no explicit ports.
Docker-published ports are accepted through Docker's DNAT/FORWARD chains, so
the public-zone cleanup does not close `3000` or `8080`. Their Compose bindings
remain a separate, staged follow-up after the required observation window.
Firewalld logged Docker chain/policy conflicts during the 2026-08-10 boots;
treat any firewall reload or service restart as a maintenance-window operation
and reverify Docker routing. Tracking: Linear `W1N-60`.
`mx.windy.me` also publishes AAAA `2602:f9f3:0:2::878`, while the host currently
has no global IPv6 address or IPv6 default route. Treat that as a separate
secondary-MX reachability issue.
## Safety
- Private keys, preshared keys, peer configuration files, and QR codes remain
@@ -21,6 +85,17 @@
- Local rollback archives are stored in `/opt/wireguard/backups` (directory
mode `0700`, archives mode `0600`). They contain private keys, are not an
off-host disaster-recovery backup, and must never leave the server.
- Live private keys, preshared keys, generated peer configs, QR images, and
`wg0.conf` are mode `0600`. Template-only `peer.conf` and `server.conf` files
are mode `0644` and do not contain generated key material.
- `/opt/wireguard/config` is mode `0755`, but its sensitive files are `0600`.
The current files are owned by the image's numeric UID/GID rather than the
declared `PUID=1000` / `PGID=1000`; the root-run WireGuard processes can use
them, but reconcile ownership only after a protected backup and maintenance
review.
- `LOG_CONFS` is currently unset and the inspected container log contained no
QR-code/config banners. Do not enable config logging; generated QR images are
credentials.
- Do not delete, move, or regenerate `/opt/wireguard/config` during
maintenance.
- Before a container recreation, validate `docker compose config` and retain a
@@ -35,6 +110,23 @@ cd ansible
ansible-playbook playbooks/health-report.yml --limit wireguard
```
Preview the narrow, fail-closed public-zone reconciliation:
```bash
ansible-galaxy collection install -r requirements.yml
ansible-playbook playbooks/us4-firewalld.yml --limit us4 --check --diff
```
Apply it only after testing the provider console and keeping an independent SSH
rollback session open. The playbook creates a protected server-local backup and
a 15-minute automatic rollback before changing rules; it cancels that rollback
only after SSH, HTTPS, SMTP, Docker, Fail2ban, and WireGuard checks pass:
```bash
ansible-playbook playbooks/us4-firewalld.yml --limit us4 \
-e '{"us4_firewalld_confirm": true, "us4_console_confirm": true}'
```
The image update and recreate procedure is deliberately separate and requires
an immutable image digest in the server-side Compose file plus an explicit
maintenance-window confirmation:
@@ -59,3 +151,43 @@ ansible-playbook playbooks/wireguard-harden.yml --limit wireguard \
- Validate a known client can handshake and sends IPv4 traffic through the VPN.
- Do not treat inactive mobile peers as a failure solely because their latest
handshake is old.
## Live audit snapshot (2026-08-12)
The WireGuard service itself is healthy and its installation is broadly
reasonable:
- The sanitized Ansible health report returned `status=ok`; Compose is valid,
the container is running with zero restarts, `wg0` exists, and UDP `51820` is
listening.
- One of three peers had a current handshake during the audit. Two peers had
not handshaken since the current container/interface start; confirm those
clients only if they are expected to be active.
- The image is immutable-digest pinned, key-bearing files are protected, the
container root is read-only, and the container has `NET_ADMIN` without the
broader `SYS_MODULE` capability.
- Debian `13.6`, kernel `6.12.101+deb13-amd64`, Docker Engine `29.7.2`, and
Docker Compose `v5.4.0` were observed. No Debian package updates or reboot
requirement were pending.
Open host-level follow-up (do not conflate these with a WireGuard outage):
1. **Disk capacity:** `/` was 90% used with about 3.4 GiB free. Docker reported
about 2.48 GB of reclaimable images and the system journal used about 1.9
GB, but do not prune or vacuum without reviewing retention and rollback
needs first.
2. **Docker exposure:** the firewalld public-zone cleanup is complete, but
Docker still publishes `3000` and `8080` outside the ordinary host INPUT
path. Remove those redundant Compose bindings in separate maintenance units
after the observation window, and confirm provider firewall rules first.
3. **Image maintenance:** the upstream `latest` amd64 image had advanced to
`1.0.20260223-r0-ls120` (build 2026-08-06). Review and pin its immutable
digest in a maintenance window rather than updating unattended.
4. **Host hygiene:** `apache2.service`, `certbot.service`, and
`postgresql@9.6-main.service` were in a failed state while unrelated Docker
workloads remained active. Establish ownership and remove or repair stale
units separately.
5. **Resource/log limits:** the WireGuard container has no memory, CPU, or PID
limit and uses Docker's `json-file` log driver without a per-container
rotation setting. Current log size was small, but limits/rotation should be
considered during a reviewed Compose update.
+30 -19
View File
@@ -8,28 +8,38 @@ diagnosis and procedures that are deliberately interactive or destructive; see
For a live-verified map of the **internal LAN** (gw, gfw, dns, ubnt, APs) and
the software deployed there, see [the LAN overview](../docs/lan-overview.md).
| Host | Role | SSH | IPv4 | Status | Facts |
|------|------|-----|------|--------|-------|
| mx2.windy.me | mailcow (primary MX prio 20) | `ssh -4 windy@mx2.windy.me` | 194.163.160.244 | active | [hosts/mx2.windy.me.md](../hosts/mx2.windy.me.md) |
| us2.wsvc.info | Vaultwarden/Postgres (+ Traefik, Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | 193.9.44.165 | active | [hosts/us2.wsvc.info.md](../hosts/us2.wsvc.info.md) |
| mx.windy.me | mail (secondary MX prio 30) | TBD | see AAAA/A | stub | |
| repo.windy.me | Soft Serve git (on us2) | `ssh -p 2222 windy@repo.windy.me` | 193.9.44.165 | stub | see us2 |
| auth.wsvc.info | Vaultwarden public hostname | — (HTTPS) | → us2 | active | see us2 |
| us1.wsvc.info | PowerDNS secondary (ns2 host) | TBD | 202.91.35.141 | stub | Auth 5.0.5; see hk2 |
| us4.wsvc.info | WireGuard VPN | `ssh -4 windy@us4.wsvc.info` | 185.201.226.122 | active | [hosts/us4.wsvc.info.md](../hosts/us4.wsvc.info.md) |
| hk2.chans.xyz | PowerDNS auth (ns1) | `ssh -4 windy@hk2.chans.xyz` | 154.36.174.161 | active | [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) |
| ns1.wsvc.info | PowerDNS public NS name | — (DNS) | → hk2 `154.36.174.161` | active | see hk2 |
| ns2.wsvc.info | Secondary NS (AXFR/NOTIFY peer) | — (DNS) | → us1 `202.91.35.141` | active | see hk2 |
| pdns.wsvc.info | Poweradmin UI | — (HTTPS) | → hk2 | active | see hk2 |
| pgweb.wsvc.info | PowerDNS Postgres UI | — (HTTPS) | → hk2 | active | see hk2 |
| **synapse.chans.xyz** | Matrix homeserver (ESS: Synapse + MAS + Element) | `ssh -4 windy@synapse.chans.xyz` | `169.58.86.13` | **active** | [hosts/synapse.chans.xyz.md](../hosts/synapse.chans.xyz.md) |
| **gfw.windy.lan** | OpenWrt (ImmortalWrt) LAN gateway / OpenClash (PVE VM 140) | `ssh -4 root@192.168.66.1` | `192.168.66.1` | **active** | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy (PVE VM 120) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | **active** | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | **active** | [hosts/gw.md](../hosts/gw.md) |
| **ubnt** | UniFi Network Controller (PVE VM 160) | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | **active** | [hosts/ubnt.md](../hosts/ubnt.md) |
**Ansible 列**`✓` = 该主机在 [`ansible/inventory/hosts.yml`](../ansible/inventory/hosts.yml)
(执行真相),用其 inventory key(见括号注)跑 playbook`—` = 不由 Ansible 管理,
原因是该平台无 ansible 覆盖或仅是公网别名/服务端点。
| Host | Role | SSH | IPv4 | Ansible | Status | Facts |
|------|------|-----|------|---------|--------|-------|
| mx2.windy.me | mailcow (primary MX prio 20) | `ssh -4 windy@mx2.windy.me` | 194.163.160.244 | ✓ (mx2) | active | [hosts/mx2.windy.me.md](../hosts/mx2.windy.me.md) |
| us2.wsvc.info | Vaultwarden/Postgres (+ Traefik, Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | 193.9.44.165 | ✓ (us2) | active | [hosts/us2.wsvc.info.md](../hosts/us2.wsvc.info.md) |
| mx.windy.me | mail (secondary MX prio 30) | TBD | see AAAA/A | — (stub) | stub | — |
| repo.windy.me | Gitea git (on us2) | `ssh -p 2222 windy@repo.windy.me` | 193.9.44.165 | — (service on us2) | stub | see us2 |
| auth.wsvc.info | Vaultwarden public hostname | — (HTTPS) | → us2 | — (alias) | active | see us2 |
| us1.wsvc.info | PowerDNS secondary (ns2 host) | TBD | 202.91.35.141 | — (stub) | stub | Auth 5.0.5; see hk2 |
| us4.wsvc.info | WireGuard VPN | `ssh -4 windy@us4.wsvc.info` | 185.201.226.122 | ✓ (us4) | active | [hosts/us4.wsvc.info.md](../hosts/us4.wsvc.info.md) |
| hk2.chans.xyz | PowerDNS auth (ns1) | `ssh -4 windy@hk2.chans.xyz` | 154.36.174.161 | ✓ (hk2) | active | [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) |
| ns1.wsvc.info | PowerDNS public NS name | — (DNS) | → hk2 `154.36.174.161` | — (alias) | active | see hk2 |
| ns2.wsvc.info | Secondary NS (AXFR/NOTIFY peer) | — (DNS) | → us1 `202.91.35.141` | — (alias) | active | see hk2 |
| pdns.wsvc.info | Poweradmin UI | — (HTTPS) | → hk2 | — (alias) | active | see hk2 |
| pgweb.wsvc.info | PowerDNS Postgres UI | — (HTTPS) | → hk2 | — (alias) | active | see hk2 |
| **synapse.chans.xyz** | Matrix homeserver (ESS: Synapse + MAS + Element) | `ssh -4 windy@synapse.chans.xyz` | `169.58.86.13` | ✓ (matrix_vps) | **active** | [hosts/synapse.chans.xyz.md](../hosts/synapse.chans.xyz.md) |
| **gfw.windy.lan** | OpenWrt (ImmortalWrt) LAN gateway / OpenClash (PVE VM 140) | `ssh -4 root@192.168.66.1` | `192.168.66.1` | — (OpenWrt, no ansible) | **active** | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy (PVE VM 120) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | ✓ (dns_windy_lan) | **active** | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | — (EdgeOS, no ansible) | **active** | [hosts/gw.md](../hosts/gw.md) |
| **ubnt** | UniFi Network Controller (PVE VM 160) | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | ✓ (ubnt) | **active** | [hosts/ubnt.md](../hosts/ubnt.md) |
| **hass.windy.lan** | Home Assistant (HAOS, x88 Pro physical box, LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | — (HAOS, no ansible) | **active** | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
| **pgdb** | TimescaleDB PG18 (Docker) — HA recorder backend (PVE VM, LAN55) | `ssh -4 windy@192.168.55.15` | `192.168.55.15` | — (no ansible) | **active** | [hosts/pgdb.md](../hosts/pgdb.md) |
`status: stub` = known to exist; fill `hosts/<name>.md` when next touched.
**命名映射**ansible inventory key ↔ 本表主机名 —— `matrix_vps``synapse.chans.xyz`
`dns_windy_lan``dns.windy.lan`。inventory key 不随主机名改(防止破坏 `--limit` 用法),
通过 inventory 内的 `display_name` 变量与文档交叉引用。
### Matrix services (synapse.chans.xyz)
| URL | Service | Notes |
@@ -38,4 +48,5 @@ the software deployed there, see [the LAN overview](../docs/lan-overview.md).
| https://synapse.chans.xyz | Synapse API | Client-Server + Federation API |
| https://account.chans.xyz | MAS | Matrix Authentication Service (local passwords) |
| https://admin.chans.xyz | Element Admin | Admin console (MAS admin auth) |
| https://plane.chans.xyz | Plane | Project management (Helm `plane-ce` v1.4.1, ns `plane`) |
| `mrtc.chans.xyz` | MatrixRTC | **Reserved** not deployed |
+42
View File
@@ -0,0 +1,42 @@
# Runbook index
Entry point for all runbooks. Before operational work, read the repo entry
[`AGENTS.md`](../AGENTS.md) and the spec [`RUNBOOKS.md`](../RUNBOOKS.md). New
runbooks start from [`_template.md`](_template.md).
## Route by intent
| Intent | Runbook | Type |
|---|---|---|
| mailcow health check | [mailcow-health.md](mailcow-health.md) | read-only |
| mailcow update | [mailcow-update.md](mailcow-update.md) | change (gated) |
| mailcow SMTP/IMAP client | [mailcow-smtp-client.md](mailcow-smtp-client.md) | reference |
| Vaultwarden health check | [vaultwarden-health.md](vaultwarden-health.md) | read-only |
| Vaultwarden SQLite→PG migrate | [vaultwarden-sqlite-to-postgres.md](vaultwarden-sqlite-to-postgres.md) | change (destructive) |
| PowerDNS health check | [pdns-health.md](pdns-health.md) | read-only |
| RustDesk health check | [rustdesk-health.md](rustdesk-health.md) | read-only |
| Matrix health check | [matrix-health.md](matrix-health.md) | read-only |
| Plane health check | [plane-health.md](plane-health.md) | read-only |
| pgdb health check | [pgdb-health.md](pgdb-health.md) | read-only |
| pgdb DB restore (pg_restore) | [pgdb-restore.md](pgdb-restore.md) | change (procedure) |
| pgdb image/compose update | [pgdb-update.md](pgdb-update.md) | change (gated) |
| AdGuard Home health check | [adguard-home-health.md](adguard-home-health.md) | read-only |
| Host disk cleanup (logs/apt/docker) | [host-disk-cleanup.md](host-disk-cleanup.md) | change (gated) |
| Matter packet capture | [matter-packet-capture.md](matter-packet-capture.md) | read-only |
| Home Assistant maintenance | [home-assistant-maintenance.md](home-assistant-maintenance.md) | change (gated) |
| matrix_e2ee integration update | [matrix-e2ee-update.md](matrix-e2ee-update.md) | change (gated) |
| Routine Ansible operations | [ansible-operations.md](ansible-operations.md) | change (allowlisted) |
| Linear issue → mergeable change | [issue-to-merge.md](issue-to-merge.md) | delivery |
| Failing health/playbook run | [fix-ci.md](fix-ci.md) | change |
| Release a reviewed change to production | [release.md](release.md) | change (gated) |
| Roll back a change | [rollback.md](rollback.md) | change (gated) |
| Controlled network configuration | [network-change.md](network-change.md) | change (gated) |
| Network outage / service recovery | [network-recovery.md](network-recovery.md) | recovery |
## Notes
- `fix-ci.md`, `release.md`, `rollback.md`, `network-change.md`, `network-recovery.md`
are adapted from the upstream guide to this repo's VPS-ops context (execution
layer is Ansible + SSH + Linear, not a software CI/CD pipeline).
- Health runbooks are read-only; they stop (`STOP`) when live state conflicts
with the expected state instead of mutating production.
+119
View File
@@ -0,0 +1,119 @@
# Runbook: <名称>
## Purpose
<说明本 Runbook 要解决的问题及成功结果,1–2 行。>
## Scope
- 适用环境:<production / staging / LAN …>
- 适用对象:<服务、主机、组件或告警类型>
- 不适用情形:<需要改用其他 runbook 或转人工的场景>
## Ownership
- Owner<团队或角色>
- Last reviewed<YYYY-MM-DD>
- Related systems<主机名 / 服务名>
## Preconditions
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
## Inputs
| 输入 | 来源 | 是否必需 | 校验方法 |
|---|---|---:|---|
| <参数> | <来源> | 是/否 | <如何确认有效> |
## Safety
### Non-negotiable rules
- 先只读诊断,后执行变更。
- 不得把删除现有配置作为首次恢复动作。
- 不得猜测或编造缺失参数。
- 不得绕过失败的测试、检查或审批。
- 每次变更后必须完成对应验证。
- 破坏性操作必须获得明确批准。
### Stop conditions
- 实际状态与本文档的前提或预期结果冲突。
- 缺少必要输入、权限、审批或回滚能力。
- 验证失败且本文档没有明确的下一步。
- 影响范围超出 Scope。
### Approval gates
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|---|---|---:|---|
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
## Procedure
### Step 1 — Diagnose
**Action**
<执行只读诊断动作。>
**Expected**
<列出预期输出、状态或证据。>
**Decision**
- 若 <条件 A>,进入 Step 2。
- 若 <条件 B>,进入 Troubleshooting A。
- 若无法判断或状态冲突,`STOP` 并记录证据。
### Step 2 — Change
**Action**
<描述单一、可审计的变更动作。>
**Expected**
<变更后应出现的状态。>
**Verification**
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
**Rollback**
- 触发条件:<什么情况需要回滚>
- 回滚动作:<如何撤销>
- 回滚验证:<如何确认恢复成功>
## Troubleshooting
### Troubleshooting A — <异常名称>
- 证据收集:<日志、指标、命令输出、链接>
- 允许动作:<仅限已验证且低风险的动作>
- 下一步:<回到某步 / 转入另一 runbook / STOP 并升级>
## Final Verification
只有同时满足以下标准,流程才算成功:
- <功能或服务状态>
- <自动化测试或健康检查>
- <监控指标或告警状态>
- <变更记录、PR 或 Issue 已更新>
## Failure Handling
若未能完成:
1. 停止进一步变更。
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
4. 按 <升级渠道> 交接,不继续猜测。
## References
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
+21
View File
@@ -1,5 +1,20 @@
# AdGuard Home health — dns.windy.lan
## Purpose
Read-only health check of the AdGuard Home LAN DNS service.
## Scope
- Applicable: [dns.windy.lan](../hosts/dns.windy.lan.md) (`192.168.66.36`).
- Read-only: does not expose query-log contents or secrets; does not change configuration.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: dns.windy.lan (`/opt/adguardhome`)
This runbook is read-only. It does not expose query-log contents or secrets.
Routine checks run through Ansible on demand:
@@ -58,3 +73,9 @@ a known-bad-signature test; an enabled DO bit alone is not validation.
Private PTR forwarding is intentionally absent because the EdgeRouter does
not currently answer private PTR requests.
## Safety
- Read-only: never change the DNS policy or the `agh-ui-access.service` nftables rule during this check.
- Do not infer a broken DNS policy from an empty `allowed_clients`.
- If live state conflicts with an expected value, `STOP` and report.
+66
View File
@@ -1,8 +1,29 @@
# Runbook: routine operations through Ansible
## Purpose
Routine operations (health, reconcile, maintenance) through the Ansible playbooks.
## Scope
- Applicable: every inventory host, run from `ansible/`.
- Not applicable: arbitrary remote commands — the reconcile playbook is allowlisted and gated.
Run commands from `ansible/`. The inventory forces IPv4 and uses the `windy`
account with sudo. Do a read-only health pass before any reconciliation.
## Safety
- Read-only health pass before any reconciliation.
- Mutating playbooks require explicit confirmation variables; do not bypass them.
- If a reconcile target or service name is not allowlisted, `STOP` — do not invent one.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: Ansible control-plane + all inventory hosts
## Health report (read-only)
```bash
@@ -48,6 +69,27 @@ Do not use this playbook for a Mailcow update, database migration, DNS record
change, or secret rotation. Those operations require their dedicated reviewed
and, where appropriate, interactive procedures.
## Deploy repo-owned Compose (static projects)
Repo source: `compose/<project>/compose.yml` (non-secret; secrets come from the
server-local `.env` via `${VAR}`). Mechanism and per-project status:
[`compose/README.md`](../compose/README.md).
```bash
# Read-only: staged-file diff + allowlist/confirmation asserts, no writes
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden --check --diff
ansible-playbook playbooks/compose-deploy.yml --limit powerdns --check --diff
# Apply: stage repo file → validate `docker compose config -q` against the
# server .env → backup current file (*.bak-<ts>) → promote → `up -d` (gated)
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden \
-e '{"compose_deploy_confirm": true}'
```
The playbook never writes, reads, or transfers the server `.env`. A failed
validation never touches the live compose file. Hosts without an allowlisted
`compose_repo_project` fail the assert — do not invent targets.
## Host-level maintenance
These playbooks cover every inventory host, including the Matrix K3s node:
@@ -60,6 +102,30 @@ ansible-playbook playbooks/maintenance-preview.yml
ansible-playbook playbooks/baseline.yml
```
## us4 firewalld reconciliation
The us4 playbook owns only the audited `public` zone allowlist. It fails closed
on unknown services or ports, never reloads/restarts firewalld, and does not
manage Docker-published ports.
```bash
cd ansible
ansible-galaxy collection install -r requirements.yml
# Read-only preview
ansible-playbook playbooks/us4-firewalld.yml --limit us4 --check --diff
# Apply only after testing the provider console and retaining an independent
# SSH rollback session.
ansible-playbook playbooks/us4-firewalld.yml --limit us4 \
-e '{"us4_firewalld_confirm": true, "us4_console_confirm": true}'
```
Apply creates a protected server-local backup and schedules a 15-minute
automatic rollback before changing rules. The rollback is cancelled only after
the playbook verifies fresh SSH/sudo access, public HTTPS routes, SMTP, Docker,
Fail2ban, and WireGuard. Do not bypass either confirmation variable.
## UniFi SSO login setting (mutating)
Reconciles `super_sdn.sso_login_enabled` on the UniFi controller (host `ubnt`,
+76
View File
@@ -0,0 +1,76 @@
# Runbook: fix a failing health/playbook run
> Adapted from the upstream guide's `fix-ci`. This repo has no software CI; the
> equivalent "pipeline" is the Ansible **health report** and the gated playbooks.
> This runbook covers diagnosing and fixing a failed or warning/critical run.
## Purpose
Diagnose and fix a failing Ansible health-report or playbook run without
skipping checks or changing unrelated code.
## Scope
- Applicable: `ansible-playbook playbooks/health-report.yml` and the gated playbooks under `ansible/playbooks/`.
- Not applicable: production changes beyond fixing the run; network/DNS changes → `network-change.md`.
## Safety
- Do not skip or weaken a failing check to make it pass.
- Do not change unrelated hosts or services.
- Prefer read-only diagnosis before mutation; destructive fixes require approval.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: Ansible health report / gated playbooks
## Procedure
### Step 1 — Reproduce and read
**Action** — re-run the failing playbook with `--limit <host>` and capture the task that failed.
```bash
cd ansible
ansible-playbook playbooks/health-report.yml --limit <host> -v
```
**Expected** — a specific failed task, host, and message (warning vs critical).
**Decision** — clear failure → Step 2; ambiguous → `STOP` and collect `-vvv` output + the relevant `latest.json`.
### Step 2 — Diagnose
**Action** — inspect the corresponding service on the host using the matching health runbook (`mailcow-health.md`, `vaultwarden-health.md`, `pdns-health.md`, etc.).
**Expected** — a root cause (container down, cert expired, queue backlog, drift).
**Decision** — root cause found → Step 3; live state conflicts with the runbook's assumptions → `STOP`.
### Step 3 — Fix within scope
**Action** — apply the minimal fix the service runbook prescribes (e.g. `compose-reconcile` for a config drift, or a documented update). Use only allowlisted/gated playbooks.
**Verification** — re-run the health report and confirm it passes.
**Rollback** — revert to the prior config/state and re-run; see `rollback.md` for the general procedure.
## Troubleshooting
### Troubleshooting A — Intermittent/flaky failure
- Evidence: timing, DNS stub flakiness (use `1.1.1.1`/`8.8.8.8` for probes).
- Allowed: re-run once with the documented resolver workaround.
- Next: still failing → `STOP` and escalate.
## Final Verification
- Health report passes for the affected host.
- No checks were skipped or weakened; the fix is committed/documented.
## References
- [`ansible-operations.md`](ansible-operations.md)
- Per-service health runbooks under [`runbooks/`](.)
+420
View File
@@ -0,0 +1,420 @@
# Runbook: Home Assistant maintenance (hass.windy.lan)
Target: [hass.windy.lan](../hosts/hass.windy.lan.md) (physical x88 Pro box, HAOS `machine: green`)
Upstream: HAOS 18.2 / Supervisor 2026.09.0 / Core 2026.9.1 (verified 2026-09-13)
This runbook covers routine Home Assistant maintenance through the **`ha`
supervisor CLI**. All commands are wrapped by a single script
[`scripts/ha-maintenance.sh`](scripts/ha-maintenance.sh); the sections below
document the exact commands it runs, for manual/agent use.
## Purpose
Run routine Home Assistant maintenance on `hass.windy.lan` (health snapshot,
config validation, log inspection, updates, and recovery) through the `ha`
supervisor CLI.
## Scope
Applies to `hass.windy.lan` only (HAOS, `machine: green`). Covers both
read-only checks and gated mutating operations; the "Command families
intentionally NOT scripted" table below lists what is deliberately out of
scope.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: hass.windy.lan (HAOS, `machine: green`)
## Safety
- Prefer read-only checks first; the health snapshot mutates nothing.
- Every mutating mode (update / restart / rebuild / rollback / reboot /
backup / restore / add-on lifecycle) refuses to run without `--yes`.
- `--restore` overwrites the current installation; `--rollback-os`,
`--reboot`, and `--rebuild-core` are disruptive. Run them only from a
planned recovery with the backup verified.
- Never commit `SUPERVISOR_TOKEN` or a long-lived `HA_TOKEN`; read entity
state via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`).
- The `--restart-core` wrapper exits 1 silently on ssh failure — treat an
empty/exit-1 result as failure and confirm with `ha core info`.
- If live state conflicts with a documented expectation, `STOP` and report;
do not improvise command families outside this script.
## Access pattern
`ha` authenticates to the Supervisor with `SUPERVISOR_TOKEN`. Interactive SSH
login works because `~hassio/.zprofile` runs `exec sudo -i`; the root login
environment carries the token. Non-interactive use must be:
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha <cmd>'
```
Running `ha` as `hassio` directly (or `sudo -n` without `-i`) returns
`unauthorized: missing or invalid API token`.
**MOTD:** every `sudo -n -i` login prints the SSH & Web Terminal MOTD banner.
The script runs its whole procedure in one remote login (`sudo -n -i bash -s`)
so the banner appears once, then strips it with `awk` up to the
`System is ready! Use browser or app to configure.` line.
**Restart wrapper (verified 2026-08-14, W1N-107):**
`./ha-maintenance.sh --restart-core --yes` exited 1 with no output in <1s
and **did not restart Core**. The wrapper pipes a remote script through
`ssh … 2>/dev/null | awk …`; with `set -uo pipefail`, an ssh failure is
silent and the pipeline returns empty/exit 1 **before any remote command
runs**. That is not a MOTD-strip artifact after a successful restart.
The working restart was
`ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core restart'`
(~131s, `Command completed successfully.`). Treat empty/exit 1 as
failure; confirm with elapsed time and `ha core info`.
## Script usage
```bash
cd runbooks/scripts
./ha-maintenance.sh # read-only health snapshot
./ha-maintenance.sh --check-config # validate core configuration
./ha-maintenance.sh --logs core 2500 # tail core logs (use 2500 after a restart)
./ha-maintenance.sh --logs supervisor # tail supervisor logs (default 100)
./ha-maintenance.sh --logs host 50 # tail host journald logs
./ha-maintenance.sh --logs apps:<slug> # tail an add-on log
# Mutating — refuse to run without --yes:
./ha-maintenance.sh --update --yes # refresh + update core(--backup)/supervisor/os
./ha-maintenance.sh --restart-core --yes # restart Core
./ha-maintenance.sh --restart-core --safe-mode --yes # restart Core in safe mode
./ha-maintenance.sh --rebuild-core --yes # rebuild Core image (after options change)
./ha-maintenance.sh --rollback-os --yes # boot previous OS slot (A/B rollback)
./ha-maintenance.sh --reboot --yes # reboot the HAOS host
./ha-maintenance.sh --backup [NAME] --yes # full backup (optionally named)
./ha-maintenance.sh --restore <slug> --yes # restore a backup (DESTRUCTIVE)
./ha-maintenance.sh --app restart core_mosquitto --yes # add-on lifecycle
```
- `--app` action is one of `start|stop|restart|update`; needs an add-on slug.
- `HA_HOST` / `HA_SSH_USER` override the defaults (`hass.windy.lan` / `hassio`).
- `--restore` overwrites the current installation — run only from a planned
recovery, with the backup verified.
## Command reference (verified 2026-08-14)
All verified against the live host. MOTD prepends each command's output; strip
with the `awk` pattern above or read the last block.
### Routine / read-only
| Purpose | Command |
|---|---|
| General overview | `ha info` |
| Core version/status | `ha core info` (this CLI build has no `state:` field; success is a normal info dump) |
| Core config validation | `ha core check` |
| Core stats | `ha core stats` |
| Supervisor status | `ha supervisor info` (incl. add-on list) |
| Supervisor stats | `ha supervisor stats` |
| OS status | `ha os info` (boot slots A/B) |
| Host status | `ha host info` (disk free/total, kernel) |
| Network | `ha network info` (`supervisor_internet`) |
| Hardware | `ha hardware info` |
| Pending updates | `ha available-updates` |
| Reload stores/versions | `ha refresh-updates` |
| Job manager | `ha jobs info` |
| Resolution center | `ha resolution info` |
| Core logs | `ha core logs -n 100` (`-f` follow, `-b` boot id). Default 100 misses setup; use `-n 2500` after a custom-component restart. `/config/home-assistant.log` may be missing — `ha core logs` is the source of truth. |
| Supervisor logs | `ha supervisor logs -n 100` |
| Host journald logs | `ha host logs -n 100` |
| Add-on logs | `ha apps logs <slug> -n 100` |
| Add-on list | `ha supervisor info``addons:` (started/stopped/error) |
| Security integrity | `ha security integrity` |
### Mutating (require --yes)
| Purpose | Command |
|---|---|
| Update core (with partial backup) | `ha core update --backup` |
| Update supervisor | `ha supervisor update` |
| Update OS | `ha os update` |
| Update add-on | `ha apps update <slug>` |
| Restart core | `ha core restart` / `ha core restart --safe-mode` |
| Rebuild core | `ha core rebuild` |
| OS rollback | `ha os boot-slot other` |
| Reboot host | `ha host reboot` |
| Full backup | `ha backups new [--name NAME]` |
| Restore backup | `ha backups restore <slug>` |
| Add-on start/stop/restart | `ha apps start\|stop\|restart <slug>` |
## Procedure
### 1. Health snapshot (read-only)
```bash
./ha-maintenance.sh
```
Review: supervisor `healthy: true`/`supported: true`; core/OS `update_available`;
add-on states (any `state: error`?); `resolution info` issues; disk free.
### 2. Validate config after any `configuration.yaml` change
```bash
./ha-maintenance.sh --check-config
```
Expect `Command completed successfully.` before a Core restart.
`ha core check` / a YAML reload is **not** enough after copying Python
custom-component files — restart Core.
### 3. Inspect logs
```bash
./ha-maintenance.sh --logs core 2500 # after a Core restart / custom-component copy
./ha-maintenance.sh --logs supervisor
./ha-maintenance.sh --logs apps:core_mosquitto
```
Default `--logs core` (100 lines) is too short to catch coordinator pickle /
setup errors. `/config/home-assistant.log` may be absent while
`ha core logs` still has history.
### 4. Apply updates (mutating)
```bash
./ha-maintenance.sh --update --yes
```
Runs `refresh-updates``core update --backup` (partial backup first) →
`supervisor update``os update`, then re-prints pending updates. Prefer the
web UI (**Settings → System → Updates**) for a human-supervised pass.
### 5. Recovery operations (mutating, only when needed)
```bash
./ha-maintenance.sh --restart-core --safe-mode --yes # start Core without custom integrations
./ha-maintenance.sh --rollback-os --yes # OS update broke boot? go back one slot
./ha-maintenance.sh --restore <slug> --yes # full restore; overwrites current install
```
OS update policy: HAOS uses two boot slots (A/B); `ha os info` shows which slot
booted. After a bad OS update, `ha os boot-slot other` boots the previous slot.
### 6. Backup before major changes
```bash
./ha-maintenance.sh --backup pre-migration --yes # named backup
```
### 7. Install or update a custom component (manual zip)
Home Assistant loads custom integrations from
`/config/custom_components/<domain>/` (on this HAOS host `/config`
`/homeassistant`). Official lookup:
`<config>/custom_components/<domain>` then built-in
`homeassistant/components/<domain>`
([Integration file structure](https://developers.home-assistant.io/docs/creating_integration_file_structure)).
A folder named after the domain, with at least `manifest.json` and
`__init__.py`, is enough. **Restart Core** after copying — `ha core check`
and a YAML reload do not pick up new Python packages.
This host's live trees are **file copies**, not git clones. Do not
`git pull` inside `custom_components/`.
#### Official plugin paths (CSG)
[windyboy/china_southern_power_grid_stat README](https://github.com/windyboy/china_southern_power_grid_stat):
[HACS](https://hacs.xyz/) **or**
[手动下载安装](https://github.com/windyboy/china_southern_power_grid_stat/releases).
This host uses the zip path. **Do not HACS-update this integration here.**
HACS still tracks upstream `CubicPill/china_southern_power_grid_stat`
`v1.2.0` and would overwrite the fork. Releases have no uploaded zip
assets — use GitHub **Source code (zip)** / zipball of the tag.
Worked SSH example (tag, backup, `rsync`, `__pycache__`, restart):
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § Manual
custom-component install.
#### Procedure
1. **Backup the live tree off `custom_components/`.** HA scans every
directory under `custom_components/` whose `manifest.json` `domain`
matches. A `*.bak-*` folder next to the live tree makes Core import
the backup (`No module named '...bak-YYYYMMDD-...'`, W1N-106). CSG
backups: `/homeassistant/.csg-backups/`.
2. **Copy only the inner `custom_components/<domain>/` tree**, not the
repo root and not an extra nested folder.
3. **Wipe `__pycache__` as root.** `rsync --delete` as `hassio` cannot
unlink Core-owned `.pyc` (permission denied, exit 23); stale
`cpython-314` bytecode can keep the old coordinator in memory until
restart. Then restart:
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i rm -rf /homeassistant/custom_components/<domain>/__pycache__ \
/homeassistant/custom_components/<domain>/*/__pycache__ &&
sudo -n -i ha core restart'
```
4. **Wait 12 min**, then `ha core info` (this CLI build has no `state:`
field; success is a normal info dump). Confirm `manifest.json`
`version` matches the tag.
5. **Read enough Core logs.** Default `ha core logs` is too short to
catch setup. Use `-n 2500` (or `--logs core 2500`) and look for
`Setting up <domain>` plus the first coordinator errors.
6. **First poll can time out.** If last-month sensors have numbers but
this-month stay `unknown`/`unavailable`, reload the config entry
(UI: integration → Reload). Supervisor:
```bash
# entry id from .storage/core.config_entries (CSG: 01KGCQDSZCF523A9X6SV3BZ1B9)
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i python3 -c "
import os, urllib.request
req = urllib.request.Request(
\"http://supervisor/core/api/config/config_entries/entry/<ENTRY_ID>/reload\",
method=\"POST\",
headers={\"Authorization\": \"Bearer \" + os.environ[\"SUPERVISOR_TOKEN\"]},
)
print(urllib.request.urlopen(req, timeout=60).status)
"'
```
7. **Do not edit the dashboard or `templates/csg_sensors.yaml` for an
install.** Entity IDs did not change across v1.3.0/v1.3.1/v1.3.2.
(The old `| float(0)` fake-zero follow-up was resolved 2026-08-29 by
W1N-239: template sensors now carry `availability` templates and show
`unavailable` instead of fake zeros when native CSG sensors are down.
Template edits go through that issue, not the install path.)
#### Verify (CSG, after v1.3.2 / W1N-118)
| Check | Expect |
|---|---|
| `manifest.json` `version` | `1.3.2` |
| `ha core logs` after this restart | `Setting up china_southern_power_grid_stat`; **no** `cannot pickle 'mappingproxy'` |
| Config entry | `state: loaded` |
| `sensor.0800041935246530_balance` | numeric (may be `0.0`) |
| `sensor.0800041935246530_this_month_total_usage` | numeric after reload if first poll timed out |
| Native `*_total_cost` / `current_ladder` | may stay `unknown` (CSG marketing calendar SQL error); dashboard uses W1N-114 `csg_*` ladder/cost templates |
`monetary` + `total_increasing` warnings on this-month/year cost sensors
are a remaining plugin issue, not an install failure.
There is no long-lived `HA_TOKEN` in the agent environment. Read entity
states via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`) at
`http://supervisor/core/api/states/<entity_id>`.
#### CSG display refactor 2026-09-04 (VPS-90)
Template/dashboard changes made **after** pricing cross-check (8月账单
198.64 元 vs 模板 198.65 元,≤0.01 元;阶梯常量 0.589/0.639/0.889、
260/600 夏档未动):
- `templates/csg_sensors.yaml` Block B 新增
`sensor.csg_this_month_avg_price`(本月阶梯电费÷本月用电,`元/kWh`);
**csg_* template sensors = 15**
- Panel `power-monitor``lovelace.dashboard_unknown`):环比行改名
「环比上月同期」;glance「本月/上月」去重为单卡「上月」(本月行归
💰核心数据卡);⚡阶梯电价卡加「本月实际均价」行。实体引用 20→21。
- `automations.yaml` +2 提醒:`automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`
10-25/ `automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`4-25
09:00 Matrix 提醒人工切「本月累计」gauge 季节档(max/segments 不可模板化)。
- 金额单位混排(原生 CNY vs 模板 元)**保留**`config/entity_registry/update`
拒绝自定义文本单位(`extra keys not allowed … Got '元'`),已定案接受。
**WS 改面板(2026.8,本机实测,后续沿用)**: core/主机 python 无 ws 库、
core 容器内经 supervisor 代理 WS 被拒(loop prevention)。用
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i sh -c "docker run --rm -i --network host -e SUPERVISOR_TOKEN \
--entrypoint python3 r.hassbus.com/home-assistant/aarch64-hassio-supervisor:2026.08.0 \
- < /tmp/x.py"'
```
`ws://172.30.32.2/core/websocket`aiohttpheader `Authorization: Bearer
$SUPERVISOR_TOKEN`,随后 auth 帧同 token)。命令名 **`lovelace/config`**(读)
+ **`lovelace/config/save`**(写,url_path + 全量 config);`lovelace/config/get`
已不存在(unknown_command)。备份与细节见
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § CSG 面板重构 2026-09-04。
## Command families intentionally NOT scripted
These exist in `ha` but are either rare, dangerous, or better done in the web
UI; documented here so nothing is a surprise. Use `ha <family> --help` on the
host for exact syntax.
| Family | Notes |
|---|---|
| `ha audio` | Audio device management; peripheral. |
| `ha authentication` | `auth list/reset/cache`; user password ops — do in web UI. `auth list` is local-terminal only. |
| `ha cli` | Internal CLI backend info/update; self-maintained. |
| `ha dns` | Internal DNS server; only relevant if Supervisor DNS add-on in use. |
| `ha docker` | Host Docker backend info/options/registries; HAOS-managed. |
| `ha mounts` | Network storage (NFS/CIFS) mounts — configure in **Settings → System → Storage**. |
| `ha multicast` / `ha observer` | Internal services; self-maintained. |
| `ha network scan/update/vlan` | WiFi AP scan & interface config — prefer web UI networking. |
| `ha host disks/options/shutdown/reload` | Disk ops / host options; `shutdown` is equivalent to `--reboot` but off. |
| `ha os datadisk list/move/wipe` | Data-disk migration; `wipe` is **local-terminal only** and erases all data. |
| `ha os import` | Import config from USB stick. |
| `ha os boards` / `os config` | Board / OS settings. |
| `ha core options` / `supervisor options` | Core/OS config options (e.g. `--duplicate-log-file`); changes need `ha core rebuild` + restart. |
| `ha backups freeze/thaw/remove/options` | Freeze/thaw for external backup tools; removal is destructive. |
| `ha jobs options/reset` | Job-manager tuning. |
| `ha resolution check/healthcheck/issue/suggestion` | Resolution center management; `healthcheck` runs fixups. |
| `ha store add/delete/repair` | Repository management — add repos in web UI app store. |
| `ha security info/options` | Security backend options. |
## Docs vs actual CLI discrepancies
The [official HAOS common-tasks docs](https://www.home-assistant.io/common-tasks/os/)
also mention `ha host update`, which **does not exist** in this CLI
(2026-08-14). Docs' `ha backups list` is not a subcommand either:
`ha backups --help` lists freeze/info/new/options/reload/remove/restore/thaw;
extra positional args (`list`, `nonsense`, ...) are ignored and the default
list still prints with exit 0. The list command is plain `ha backups`.
Per-backup: `ha backups info <slug>` (slug required). Trust the server CLI
(`ha <cmd> --help`) over the docs.
This CLI's `ha core info` also has no `state:` field (verified 2026-08-14).
Wait for a successful info dump after restart, not a `state: running` line.
## Known issues on hass.windy.lan (2026-08-14)
2026-08-13 snapshot items were resolved same day (W1N-70/71/72/73/74/75/76):
OTBR and the duplicate SSH add-on uninstalled, resolution-center empty,
full backup `pre-maintenance-20260813` (slug `411a4ba5`). Remaining:
- **Bluetooth hci0 instability (RTL8821CS)**: `bluetooth_auto_recovery`
power-reset times out every ~2 min; kernel `hci0 hardware error`. No BLE
entities exist, so no user impact. HAOS image ships `x88-bt-hci-recovery`
workaround units.
- `host info` reports `disk_life_time: 10` (boot eMMC ~10% life left) —
monitor on each snapshot; plan disk replacement / data-disk migration.
- **Home PPPoE IPv4 to CSG is blackholed** (`curl -4` to
`218.19.148.218:443` times out). `end0` IPv6 works (`curl -6
https://95598.csg.cn` → HTTP 200). Entry `ip_family: ipv4` still
matches the stored option; first post-restart poll can still time out
— reload the config entry rather than reinstalling.
- **WSL HTTP proxy**: LAN `hass.windy.lan:8123` through Mihomo returns
empty `502`. Bypass proxy or add `.windy.lan` to `NO_PROXY` before
debugging UI/API from the workstation
([hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § HTTP proxy
gotcha).
- **No long-lived HA token in the agent environment.** Read entity
states via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`) at
`http://supervisor/core/api/states/...`, not a committed `HA_TOKEN`.
## Pass criteria
- Health snapshot completes; supervisor `healthy`/`supported: true`
- Mutating modes refuse to run without `--yes` (incl. `--restore`, `--app`)
- `--check-config` returns success
- Update / rollback / restore / reboot confirmed only after explicit `--yes`
- Custom-component zip install: live `manifest.json` version matches the
tag; backups not under `custom_components/`; Core restarted; logs show
`Setting up <domain>` without import / pickle errors
- Update the **Verified** line on [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md)
+276
View File
@@ -0,0 +1,276 @@
# Runbook: Host disk cleanup (unbounded container logs / apt cache / docker artifacts)
## Purpose
Reclaim space on a root filesystem that is filling up (≥70% used) on a Docker
Compose host, by fixing unbounded container log growth at the source, clearing
apt/journal caches, and removing unused Docker images/volumes. Success: root
usage drops to a safe band (≤55% used, or per acceptance in the tracking issue)
and log growth stays bounded afterwards.
## Scope
- 适用环境: production single-root-fs hosts running Docker Compose stacks
(first application: `hk2.chans.xyz`; reusable for `mx2.windy.me` / `us2.wsvc.info`
which run the same unbounded-`json.log` pattern).
- 适用对象: root filesystem usage; container stdout/stderr log files
(`/var/lib/docker/containers/*/*-json.log`); `/var/cache/apt`; systemd journal;
unused Docker images / anonymous volumes / build cache.
- 不适用情形: hosts without systemd-journald or without Docker; LAN/HAOS hosts
(use their own runbooks); cases needing disk *growth* (provider resize) rather
than cleanup; anything touching service data volumes or `/opt/*` configs
(STOP and use the service-specific runbook instead).
## Ownership
- Owner: windy (operator) + agent executing per approval
- Last reviewed: 2026-09-02
- Related systems: hk2.chans.xyz (PowerDNS auth / AdGuard Home / Traefik / RustDesk compose stacks)
## Preconditions
- SSH access to the target host with **passwordless sudo** (`sudo -n true` must succeed).
- A recorded `df -h` baseline and `docker system df` baseline.
- **Explicit user approval** for every service touch listed in Approval gates
(recorded in the tracking issue, e.g. Plane `vps` VPS-81).
- No open incident on the target host.
- Container log growth root cause identified in Diagnose before mutating.
## Inputs
| Input | Source | Required | Validation |
|---|---:|---|
| Target host | inventory/hosts.md | yes | SSH login + `uname -r` |
| df/docker baseline | live read-only probe | yes | recorded before first mutation |
| Approved service touches | user confirmation in tracking issue | yes | issue comment states approval |
| Image keep-list (rollback pins) | operator decision in issue | yes | review `docker image ls` before rmi |
| Backup of any config edited | local copy with timestamp | yes | exists before edit |
## Safety
### Non-negotiable rules
- Prefer read-only diagnosis before mutation (never mutate on an unmeasured disk).
- Never use `rm` on a live container log — use `truncate -s 0` (keeps the fd valid).
- Never run `docker image prune -a` when a keep-list is intended — no keep-list
exists; delete explicitly with `docker rmi`.
- Never run `docker volume prune -a` — plain `docker volume prune` (no `-a`)
removes only unused anonymous volumes; named/in-use volumes stay.
- After every mutation, verify the expected state (`df -h`, container status).
- Destructive actions require explicit approval (Approval gates).
### Stop conditions
- Live state conflicts with this runbook's preconditions or expectations (e.g.
root usage differs wildly from baseline, or a container is unhealthy).
- Missing approval, missing backup, or missing rollback ability.
- A verification step fails with no documented next step.
- Any step would touch a volume/container/mount that is not on the approved list.
### Approval gates
| Action | Risk | Explicit approval | Approval record |
|---|---:|---|---|
| `docker restart <chatty container>` | low (sec-level blip of that service only) | yes | tracking issue (VPS-81 T1) |
| `systemctl restart systemd-journald` | low (sec-level, no state loss) | yes | tracking issue (VPS-81 T2) |
| `apt-get clean` | low (re-downloadable) | no | — |
| `journalctl --vacuum-*` / journald drop-in | low | no (restart above is gated) | — |
| `docker rmi` of unused images | medium (rollback pin removed unless kept) | yes (keep-list) | tracking issue (VPS-81 T3) |
| `docker volume prune` | medium (data in anonymous volumes lost) | yes | tracking issue (VPS-81 T4) |
| `docker builder prune` | low | no | — |
## Procedure
### Step 1 — Diagnose
**Action**
Read-only: `df -h`, `df -i`, `sudo du -x -h --max-depth=1 /`, `docker system df`,
and locate oversized container logs:
`sudo ls -la /var/lib/docker/containers/*/*-json.log`. Map a big log to its
container (`docker inspect -f '{{.Name}} {{.LogPath}}' <id>`), then inspect what
it logs (`sudo tail -c 400000 <logpath>`; count `[debug]` lines) and find the
config flag driving it (e.g. AGH `log.verbose` in its YAML; note `log.file: ""`
means the app's own rotation keys are inert and output goes to the container log).
**Expected**
A full accounting of root usage and identification of: (a) any unbounded
container log and its root-cause flag; (b) reclaimable apt cache; (c) journal
size and journald limits; (d) unused images (0 dangling expected) and unused
anonymous volumes.
**Decision**
- If root is ≥70% used or any container log is unbounded → Step 2.
- If root is healthy and logs are bounded → STOP (no change needed; record evidence).
- If state conflicts with expectations (e.g. missing sudo, unexpected mount) → STOP.
### Step 2 — Fix noisy container logging at the source, then truncate
**Action**
1. Back up the app config: `sudo cp <config> <config>.bak-YYYYMMDD-<issue>`.
2. Disable the debug/verbose flag (e.g. `log.verbose: true → false` in the AGH YAML).
3. Apply config with a container restart: `docker restart <container>` (config-level
change; **no recreate** needed and daemon.json rotation would not apply anyway).
4. Truncate the accumulated logs: `sudo truncate -s 0 <json.log>` for the chatty
container(s) (and any other oversized ones, e.g. traefik).
5. Record `df -h` before/after.
**Expected**
`docker logs <container>` no longer shows the `[debug]` flood; the `*-json.log`
stops growing; several GB reclaimed.
**Verification**
- `sudo tail -c 200000 <json.log>` after ≥1 minute → no new debug lines.
- `df -h` improvement recorded.
- Container still `Up (healthy)`.
**Rollback**
- Trigger: log volume unchanged, service degraded, or debug output is actually needed.
- Action: restore the config backup and `docker restart <container>`.
- Verify: original verbose behaviour back; container healthy.
### Step 3 — Clear apt cache and cap journald
**Action**
1. `sudo apt-get clean` (clears only `/var/cache/apt/archives`; `/var/lib/apt/lists`
is not cleared by it and regenerates on `apt update` — optional/low value, skip).
2. `sudo journalctl --vacuum-size=100M`.
3. Write drop-in `/etc/systemd/journald.conf.d/00-disk-<issue>.conf`:
`[Journal]` + `SystemMaxUse=200M`.
4. `sudo systemctl restart systemd-journald` (approved service touch).
5. Record `df -h` before/after.
**Expected**
Archives cleared (~1.4G on hk2), journal ≤100M, future journal capped at 200M.
**Verification**
- `du -sh /var/cache/apt/archives` → ~0.
- `journalctl --disk-usage` → ≤100M.
- `systemctl show systemd-journald -p ...` or restart log confirms new limit;
`journalctl -b` still readable.
**Rollback**
- Trigger: journald fails to start or logs lost unexpectedly.
- Action: remove the drop-in, `sudo systemctl restart systemd-journald`.
- Verify: journald active, prior journal entries still listed.
### Step 4 — Remove unused Docker images (explicit keep-list)
**Action**
1. Enumerate unused images: `docker image ls` cross-checked against the images of
running containers (`docker ps --format '{{.Image}}'`). Re-enumerate at
execution time — the list drifts.
2. Present the exact removal list to the operator; keep the agreed rollback pin(s)
(e.g. `powerdns/pdns-auth-50:5.0.5`) and delete the rest explicitly:
`docker rmi <repo:tag> ...` (per image).
3. Record `df -h` before/after.
**Expected**
Only in-use images + kept pins remain; ~12.5G reclaimed (reclaim is an upper
bound — layers shared with kept images are not freed; measure with `df`, do not
promise the estimate).
**Verification**
- `docker image ls` shows only the expected set.
- `docker system df` images reclaimable ≈ 0 for the removed set.
- All containers still `Up`.
**Rollback**
- Trigger: an image that was actually needed was removed.
- Action: re-pull it from the registry (`docker pull <repo:tag>`); if a kept pin
must change, update the compose pin and `up -d`.
- Verify: image present; affected service healthy.
### Step 5 — Remove unused anonymous volumes and build cache
**Action**
1. Enumerate volumes: `docker volume ls`, and confirm which are referenced by
containers (`docker inspect` Mounts). Expected targets: anonymous volumes with
no container reference.
2. `docker volume prune` (**no `-a`**) — engine ≥ v23 removes only unused
anonymous volumes; in-use volumes (e.g. PG data) are protected by container
references in every version.
3. `docker builder prune -f`.
4. Record `df -h` before/after.
**Expected**
Unused anonymous volumes (~1.2G on hk2) and build cache gone; in-use volumes intact.
**Verification**
- `docker volume ls` shows only in-use volumes.
- Services that own volumes (e.g. postgres) report healthy and data present.
- `df -h` improvement recorded.
**Rollback**
- Trigger: data loss suspected in a removed volume.
- Action: restore from backup if the volume ever contained data; verify against
the pre-prune enumeration (targets must be anonymous + unreferenced before prune).
- Note: this is why target enumeration is recorded before pruning.
## Troubleshooting
### Troubleshooting A — Log still grows after disabling verbose
- Evidence: `sudo tail -c 200000 <json.log>` still shows new lines; app config re-checked.
- Allowed actions: check for a second verbose source (container entrypoint flags,
other apps in the same log); check `docker inspect <c> --format '{{.HostConfig.LogConfig}}'`.
- Next step: back to Step 2 or STOP if a container-level log-opts change (recreate)
would be needed — that is a separate approval.
### Troubleshooting B — `docker rmi` fails (image in use)
- Evidence: `image is being used by stopped container ...`.
- Allowed actions: identify the stopped container (`docker ps -a`); confirm it is
not needed; remove it only with explicit approval.
- Next step: re-run rmi for the remaining images; never force-delete blindly.
### Troubleshooting C — `docker volume prune` would remove more than expected
- Evidence: prune dry-run/listing includes a named or referenced volume.
- Allowed actions: abort; do not add `-a`; re-check references.
- Next step: STOP and report to the operator with the enumeration.
## Final Verification
The flow is successful only when all of the following hold:
- `df -h` root usage is in the agreed band (VPS-81: 76% → ≤55% used; measure, do not assume).
- `docker system df` shows reclaimable ≈ 0 for images/volumes targeted.
- All containers `Up` (health checks pass); public services verified
(`dig @<host-ip> SOA <zone>` for DNS hosts; service URLs reachable).
- Tracking issue updated with before/after `df`, actions, and the one-week
observation checkpoint for log growth.
## Failure Handling
If the flow cannot complete:
1. Stop further mutation.
2. Collect command output, timestamps, and the exact step that failed.
3. Record completed steps, actual results, unmet expectations, and whether a
rollback ran.
4. Hand over per the tracking issue with evidence; do not guess further.
## References
- Plane `vps` issue VPS-81 "hk2: 释放根盘空间" (+ subtasks VPS-82…88) — plan, review findings, approvals.
- [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) — host facts.
- [RUNBOOKS.md](../RUNBOOKS.md) — runbook spec; [runbooks/README.md](README.md) — index.
+93
View File
@@ -0,0 +1,93 @@
# Runbook: issue → mergeable change
## Purpose
Turn an approved Linear `vps` issue into a reviewed, mergeable change in this
repo (docs, runbooks, hosts facts, or Ansible playbooks).
## Scope
- Applicable: repo content under `docs/`, `runbooks/`, `hosts/`, `inventory/`, `ansible/`.
- Not applicable: mutating production state directly — that goes through `release.md` / `ansible-operations.md`.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: Linear MCP (`vps` project), git
## Inputs
| Input | Source | Required | Validation |
|---|---|---:|---|
| Issue identifier | Linear (`vps` project) | Yes | `linear_get_issue <id>` returns a description |
| Current repo state | `git status` / `git log` | Yes | Clean or intended worktree |
## Safety
- Scope is locked to the issue: do not bundle unrelated changes.
- Never commit secrets (see `AGENTS.md` §Safety).
- Verify every change; do not merge a change whose verification was skipped.
## Procedure
### Step 1 — Read the issue
**Action** — `linear_get_issue <id>`, read description and acceptance criteria.
**Expected** — clear scope, action, and verification for the change.
**Decision** — if the issue is ambiguous or lacks verification criteria, `STOP`
and ask for clarification (add a `needs-info` label if applicable). Otherwise go to Step 2.
### Step 2 — Inspect and change
**Action** — read the relevant files, then make the minimal change the issue asks for.
**Expected** — diff is scoped to the issue.
**Decision** — if the change needs production mutation, `STOP` and route to
`release.md`. Otherwise go to Step 3.
### Step 3 — Verify
**Action** — run `scripts/validate-repo.sh` from the repo root (covers secret
scan, inventory cross-check, markdown link check, runbook-spec check, and
Ansible `--syntax-check`); for changes that alter playbook behavior, also run
a read-only `ansible-playbook --check` where possible.
**Verification** — `scripts/validate-repo.sh` exits 0; the concrete checks
must match the change type.
**Decision** — verification passed → Step 4; failed → Troubleshooting A.
### Step 4 — Commit and link
**Action** — commit with a message containing the full issue ID (e.g. `W1N-123: …`); open a PR if the change is substantial; link the issue via `linear_save_comment`.
**Verification** — `git log -1` shows the issue ID; the issue has the commit/PR pointer.
**Rollback** — `git revert <sha>` or `git checkout <branch>` to drop the change; re-verify after.
## Troubleshooting
### Troubleshooting A — Verification failed
- Evidence: command output, failing check.
- Allowed: fix the change within scope; re-run verification.
- Next: still failing → `STOP` and report in the issue.
## Final Verification
- Change matches the issue scope.
- Verification passed and the issue is updated with evidence.
## Failure Handling
If unfinished: stop, collect the failed check output, record completed steps, and
hand back to the issue — do not guess.
## References
- [`docs/agents/issue-tracker.md`](../docs/agents/issue-tracker.md)
- [`RUNBOOKS.md`](../RUNBOOKS.md)
+20
View File
@@ -1,5 +1,20 @@
# Runbook: mailcow health (mx2)
## Purpose
Read-only health check of the mailcow stack on mx2.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
- Read-only: does not change mailcow configuration or service state.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: mx2.windy.me (`/opt/mail`)
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Path: `/opt/mail`
Prefer: the Ansible health report (`ansible/playbooks/health-report.yml --limit mailcow`),
@@ -71,6 +86,11 @@ The sanitized Ansible health profile is `mailcow` (`ansible/playbooks/healthchec
The server-local timer emits a sanitized result at `/var/lib/vps-health/latest.json`.
It does not change Mailcow configuration or service state.
## Safety
- Read-only: never mutate configuration or service state during this check.
- If live state conflicts with an expected value below, `STOP` and report; do not "fix" on the fly.
## Pass criteria
- Compose stack up; watchdog ~100%
+21
View File
@@ -1,5 +1,20 @@
# Runbook: use mailcow SMTP / IMAP (client)
## Purpose
Reference for configuring mail clients against the mailcow SMTP/IMAP endpoints.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md) client submission (587/465) and IMAP/POP (993/995).
- Not applicable: server-side mailcow configuration or administration.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: mx2.windy.me (SMTP/IMAP client endpoints)
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Prerequisite: a mailbox on `windy.me` (password from mailcow UI, not the admin account unless it is that mailbox).
@@ -52,3 +67,9 @@ Do not commit or paste real passwords into this repo.
- Port/TLS mode mismatch (587 vs 465)
- Account active in mailcow; not rate-limited / fail2banned after bad attempts
- Apps that store SMTP in their own config (e.g. Vaultwarden `config.json`) may keep a **stale** password even when `.env` is correct — verify AUTH against the effective config ([vaultwarden-health](vaultwarden-health.md) §5)
## Safety
- Do not commit or paste real passwords into this repo or chat.
- Use submission (587/465) for client sending; never use port 25 as a desktop/app outbound port.
- If live state conflicts with the endpoint values above, `STOP` and report; do not change server-side settings during this reference check.
+28
View File
@@ -1,9 +1,37 @@
# Runbook: mailcow update (mx2)
## Purpose
Update the mailcow stack on mx2 to the latest supported release.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
- Not applicable: config changes beyond the update, DB migration, secret rotation.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: mx2.windy.me (`/opt/mail`)
## Approval gates
| Action | Risk | Explicit approval |
|---|---|---|
| Run `./update.sh` (recreates containers, brief mail interruption) | Medium | Yes — user confirmation required |
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Path: `/opt/mail`
**Confirm with the user before running an update.**
## Safety
- Never run the update without explicit user confirmation.
- Never pass secrets into the chat log; do not commit `mailcow.conf`.
- If a step fails, capture `docker compose ps` and logs and stop before further changes.
- If live state conflicts with this runbook's assumptions (e.g. unexpected `mailcow.conf` values), `STOP` and report.
## Before
1. Run [mailcow-health](mailcow-health.md) (Ansible health report). Record baseline.
+171
View File
@@ -0,0 +1,171 @@
# matrix_e2ee update (hass.windy.lan)
## Purpose
Update the custom **`matrix_e2ee`** integration on `hass.windy.lan` while
preserving a verified rollback point and confirming that Home Assistant loads it.
## Scope
- Applicable: deploying a reviewed `matrix_e2ee` source revision to
`hass.windy.lan`.
- Not applicable: Home Assistant Core upgrades, integration configuration
changes, or recovery without a usable live-tree backup.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-23
- Related systems: hass.windy.lan (HAOS, `machine: green`)
## Approval gates
| Action | Risk | Explicit approval |
|---|---|---|
| Replace the live integration tree and restart Home Assistant Core | Medium | Yes — user confirmation required |
## Safety
- Do not replace the live tree or restart Core without explicit user confirmation.
- If any precondition or verification fails, `STOP` and record evidence before continuing.
## Preconditions
- The source repo at `/home/windy/project/ha-matrix-e2ee` is on the **target
state**: either a release tag (`git tag -l 'v*'`) or a commit whose
`manifest.json` `version` is the target. Note v0.3.0 was deployed from an
**untagged** `main` HEAD (`216cc99`), so the tag check alone is not enough —
confirm the working-tree `custom_components/matrix_e2ee/manifest.json`.
- The working tree matches HEAD: `git status --short` clean (only ignorables)
and `git diff HEAD -- custom_components/` empty. Record
`git rev-parse HEAD` for the docs/Linear record — HEAD can move during a
session, so re-check right before rsync (verified 2026-08-18: HEAD moved
from a `w1n-180` branch merge to `main` mid-deploy).
- The remote host is reachable and `sudo -n -i ha core info` succeeds.
- The workstation HTTP proxy does not interfere — LAN hosts must be reachable
without proxying (unset `http_proxy` / `HTTP_PROXY` if needed).
- If the agent sandbox hits `Bad owner or permissions on /etc/ssh/ssh_config.d/20-systemd-ssh-proxy.conf`, add `-F /dev/null` to the `ssh` / `rsync` commands below.
- Domain is **`matrix_e2ee`** (double-e). Older notes may say `matrix_e2e`;
paths, events, and services all use `matrix_e2ee`.
## Procedure
### 1. Backup the live tree
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i mkdir -p /homeassistant/.matrix-e2ee-backups &&
sudo -n -i cp -a /homeassistant/custom_components/matrix_e2ee \
/homeassistant/.matrix-e2ee-backups/matrix_e2ee.bak-$(date +%Y%m%d)-v<OLD_VERSION>'
```
The backup lives in `/homeassistant/.matrix-e2ee-backups/` — a directory
separated from `custom_components/` to avoid HA scanning it as a custom
component domain.
### 2. Rsync the new source
```bash
rsync -a --delete -e 'ssh -o BatchMode=yes' \
/home/windy/project/ha-matrix-e2ee/custom_components/matrix_e2ee/ \
hassio@hass.windy.lan:/homeassistant/custom_components/matrix_e2ee/
```
The `--delete` cannot remove Core-owned `__pycache__` — that is handled
in the next step. Source `.py` files and `manifest.json` are transferred
correctly even with the `__pycache__` errors, but **rsync exits with code 23
(`some files/attrs were not transferred`)** — that is expected, not a failure.
Confirm the transfer by checking the manifest on the host before restarting.
### 3. Wipe `__pycache__` (as root) and restart Core
Quote the nested `__pycache__` glob — remote login shell is zsh and will
fail with `no matches found` if left unquoted.
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
"sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee/__pycache__ \
'/homeassistant/custom_components/matrix_e2ee/*/__pycache__' &&
sudo -n -i ha core restart"
```
Stale `cpython-314` bytecode in Core-owned `__pycache__` keeps the old
coordinator in memory until restart. Wipe before restart.
Wait for `Command completed successfully.` (typically 12 min).
### 4. Verify the deployment
#### 4a. Confirm manifest version
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i cat /homeassistant/custom_components/matrix_e2ee/manifest.json'
```
Expect `"version": "<NEW_VERSION>"`.
#### 4b. Check Core logs for matrix_e2ee
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i ha core logs -n 2500' | grep -E 'matrix_e2ee|Setting up matrix' | head -20
```
Expect:
- `Setup of domain matrix_e2ee took ...` (older wording `Setting up matrix_e2ee` may appear)
- `matrix_e2ee restored existing device; user=@hass:chans.xyz device=rO1R915ncu`
- No `ERROR` level messages from `custom_components.matrix_e2ee`
- Blocking-call WARNINGs from `_patch_nio_sas_timeout` / nio store I/O are expected
#### 4c. Verify the entry is loaded (optional, via Supervisor API)
No trailing slash on the entries URL (trailing `/` returns 404 on Core 2026.8.1).
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
"sudo -n -i python3 - <<'PY'
import os, json, urllib.request
req = urllib.request.Request(
'http://supervisor/core/api/config/config_entries/entry',
headers={'Authorization': 'Bearer ' + os.environ['SUPERVISOR_TOKEN']},
)
entries = json.loads(urllib.request.urlopen(req, timeout=30).read())
for e in entries:
if e['domain'] == 'matrix_e2ee':
print(f\"{e['domain']}: state={e['state']} source={e['source']}\")
PY"
```
Expect `state: loaded`.
### 5. Record the deployment
- Update the `matrix_e2ee` live-tree section in
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md): new version, source
commit (`git rev-parse HEAD`), backup name, and any new feature notes.
- Record the operation in the Linear `vps` project (scope, action,
verification, follow-up); see [docs/agents/issue-tracker.md](../docs/agents/issue-tracker.md).
## Rollback
If Core fails to start after the update:
```bash
# Restore the backup
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee &&
sudo -n -i cp -a /homeassistant/.matrix-e2ee-backups/matrix_e2ee.bak-<DATE>-v<OLD_VERSION> \
/homeassistant/custom_components/matrix_e2ee &&
sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee/__pycache__ &&
sudo -n -i ha core restart'
```
If a full HA backup exists (pre-update), restore via `ha backups restore <slug>`.
## References
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — current live version and config
- [docs/home-assistant-matrix.md](../docs/home-assistant-matrix.md) — integration architecture and verification model
- [home-assistant-maintenance.md](home-assistant-maintenance.md) — general HA maintenance procedures
- [ha-matrix-e2ee source](https://github.com/windyboy/ha-matrix-e2ee) — GitHub repo
+21
View File
@@ -1,5 +1,20 @@
# Matrix Health Check
## Purpose
Read-only health check of the Matrix homeserver (ESS on K3s).
## Scope
- Applicable: [synapse.chans.xyz](../hosts/synapse.chans.xyz.md), namespace `ess`.
- Read-only: does not change pods, ingress, certificates, or configuration.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: synapse.chans.xyz (ESS chart `26.7.2`, K3s)
Monitor the Matrix homeserver running on `synapse.chans.xyz` (ESS chart `26.7.2`, K3s node).
Prefer `cd ansible && ansible-playbook playbooks/health-report.yml --limit matrix`
@@ -90,3 +105,9 @@ Backup automation is currently paused. `/var/backups/matrix/` is retained for a
| Well-known returns 404/redirect | Root `chans.xyz` ingress missing or misconfigured |
| 502 Bad Gateway | Synapse pod restarting or DB down |
| SMTP emails not sent | MAS SMTP config incomplete; TCP reachable but AUTH failing — see `runbooks/vaultwarden-health.md` |
## Safety
- Read-only: never mutate pods, ingress, certificates, or configuration during this check.
- Backup automation is paused; do not treat `/var/backups/matrix/` as a recovery source.
- If live state conflicts with an expected value, `STOP` and report.
+312
View File
@@ -0,0 +1,312 @@
# Matter packet capture (read-only)
## Purpose
Capture Matter-related traffic on LAN55 (mDNS discovery + PASE/CASE commissioning +
operational traffic) to determine whether a device is on the network, is in
commissioning mode, and whether the commissioning handshake completes. Capture
is read-only and changes no device or network state.
## Scope
- Environment: LAN55 (`hass.windy.lan`, Aqara M3, ESP32-C2 Matter bulbs,
phone / HA matter-server all on the 55 subnet).
- Subject: Matter over Wi-Fi and Thread relay nodes. The Thread 802.15.4 air
side itself is not capturable — only IPv6 forwarding by a Thread relay such
as the M3 is visible.
- Not applicable: BLE commissioning, Thread 802.15.4 frames, cross-subnet
multicast (66-subnet hosts cannot see the 55 subnet's mDNS — link-local
multicast does not cross the routed 55/66 boundary, there is no reflector).
- Read-only: no AP/device/network config is modified; state returns to normal
when tcpdump exits.
### Capture-point selection
Matter commissioning is a two-party conversation and the commissioner
participates in every message of it, so capturing on the commissioner host
equals capturing the whole flow.
| Capture point | Sees | Blind spot | Notes |
|---|---|---|---|
| **hass `end0` — commissioner side (recommended)** | The full HA-driven commissioning conversation: all mDNS queries/announcements (segment multicast) + the complete TCP 5540 PASE/CASE session | Phone-as-commissioner flows (the phone's session to the device does not pass through hass) | `core_matter_server` uses **host networking**, so tcpdump on `end0` sees the add-on's traffic directly; `/` is overlay with ~42 GB free — no 60 MB tmpfs rotation needed |
| **UAP-AC-Lite `br0` (192.168.55.5)** | All mDNS multicast (flooded; igmp snooping off) + all wireless-client unicast + unicast to/from the AP | Wired↔wired unicast — e.g. HA↔M3 TCP 5540 while a Thread device commissions via the M3 (wired, observed) — is switched locally and never traverses the AP | AP `/tmp` is a ~60 MB tmpfs → rotating capture is **mandatory** |
For the common "add device" case with HA matter-server as the commissioner,
capture on hass `end0`. Use the AP `br0` point for wireless-device or
phone-driven flows (a wireless client's unicast to/from its AP is only visible
there).
A third point, `gw` `switch0`, is **verified as a limited capture point**
(cross-subnet/gateway/mDNS flows only — not a full mirror of LAN55) — see
[Capture point: gw switch0](#capture-point-gw-switch0).
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-22
- Related systems: UAP-AC-Lite AP `192.168.55.5` (br0), `core_matter_server` on
`hass.windy.lan` (`end0`), Aqara M3, ESP32-C2 Matter bulbs
## Preconditions
- SSH to the capture point:
- AP: `ssh zhiqiangf@192.168.55.5` (key-only, `BatchMode=yes` verified).
- hass: `ssh hassio@hass.windy.lan`. Non-interactive SSH does **not** source
`.zprofile`, so run tcpdump as `sudo -n -i tcpdump …` (verified 2026-08-22).
- tcpdump available:
- AP: full 4.9.2 / libpcap 1.8.1 (verified 2026-08-22).
- hass: `/usr/bin/tcpdump` via `sudo -n -i` (verified 2026-08-22).
- Trigger source ready: put the Matter device into commissioning mode, or have
HA/phone perform discovery/commissioning — otherwise no relevant packets.
- AP `/tmp` is a ~60 MB tmpfs (61.3 M total, 60.4 M free): rotating capture
(`-C`/`-W`) is mandatory on the AP. hass `/` is overlay — rotation optional
but keep the habit for long captures.
## Safety
### Non-negotiable rules
- Read-only diagnosis: no installs, config changes, or service restarts on the
AP, hass, devices, or network.
- pcap files are limited to `/tmp`; pull them off and delete them afterwards
(mandatory on the AP; same hygiene on hass).
- Never write captured content (including any plaintext key material) into this
repository or Linear.
### Stop conditions
- Capture point unreachable (ssh fails) → `STOP`, fix the network first.
- tcpdump reports "Permission denied" or cannot listen → `STOP` (admin needed;
on hass verify `sudo -n -i` works).
- Filter expression syntax error → `STOP`, use only expressions verified in
this document.
- AP `/tmp` nearly full (rotation file count × single-file size ≈ 60 MB) →
`STOP` and clean old pcaps.
## Procedure
### Step 1 — Choose the capture point
**Action**
- HA matter-server is the commissioner (the "add device" case) → hass `end0`.
- Wireless device or phone-driven flow → AP `br0`.
**Expected**
- The chosen point is reachable and tcpdump starts listening.
**Decision**
- Capture point chosen and reachable → Step 2.
- Neither applies or the choice is unclear → `STOP` and record why.
### Step 2 — Realtime observation (quick confirm traffic appears)
**Action**
AP:
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
hass (commissioner side):
```bash
ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 -s 0 -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
Keep the window open, trigger the device behavior (enter commissioning mode /
start commissioning / send a command), `Ctrl+C` to stop.
**Expected**
- `_matterc._udp` / `_matter._tcp` mDNS announcements (UDP 5353, multicast
`224.0.0.251` / `ff02::fb`).
- During commissioning: TCP **5540** (PASE/CASE) SYN/SYN-ACK between the device
IP and HA/M3.
- If the target device's MAC is known, add `and ether host <mac>` to keep only
that device (see variants).
- `5552` is not a standard Matter port; it is an observed port for the Aqara M3
Thread-relay node (see `docs/matter-pairing-troubleshoot.md`).
**Decision**
- Expected packets present → Step 3 to save evidence, or judge directly against
the stage table (`docs/matter-pairing-troubleshoot.md` §4).
- No packets at all → `STOP`: fix device online / commissioning-mode first; the
network side is repeatedly verified healthy (see troubleshooting doc).
- mDNS present but no 5540 → see troubleshooting doc decision tree, item 4
(§5).
### Step 3 — Rotating capture + pull to WSL
**Action** (`-C 5` = rotate every 5 MB, `-W 12` = max 12 files, ≈ 60 MB ≤ AP tmpfs)
AP:
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -C 5 -W 12 -w /tmp/matter.pcap 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
hass (rotation optional — overlay disk):
```bash
ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 -s 0 -C 5 -W 12 -w /tmp/matter.pcap 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
Trigger the traffic, then `Ctrl+C`. Files are `/tmp/matter.pcap`,
`/tmp/matter.pcap1`, …
**Expected**
- tcpdump prints capture statistics (`N packets captured`).
- `ls -la /tmp/matter.pcap*` shows the files; total stays < 60 MB on the AP.
**Verification**
```bash
ssh zhiqiangf@192.168.55.5 "ls -la /tmp/matter.pcap*"
# or
ssh hassio@hass.windy.lan "ls -la /tmp/matter.pcap*"
```
**Pull to WSL for analysis and clean up afterwards**
```bash
scp zhiqiangf@192.168.55.5:/tmp/matter.pcap* .
# or
scp hassio@hass.windy.lan:/tmp/matter.pcap* .
# clean up on the capture point
ssh zhiqiangf@192.168.55.5 "rm -f /tmp/matter.pcap*"
ssh hassio@hass.windy.lan "sudo -n -i rm -f /tmp/matter.pcap*"
```
### Step 4 — Wireshark analysis (optional)
**Action**
Open the pcap in Wireshark. mDNS (UDP 5353) is plaintext and directly
readable; Matter payloads on TCP/UDP 5540 show only the handshake by default —
plaintext needs the dissector plus session keys (Step 5).
**Expected**
- `mDNS` filter shows all discovery records; `tcp.port==5540` shows the
commissioning handshake.
### Step 5 — Decrypt Matter plaintext (optional, needs session keys)
Matter payloads are encrypted (AES-CCM); mDNS plaintext contains no keys. To
decrypt, one of:
1. **Capture-side key leak with a chip tool (most common)**: the commissioner
(HA matter-server / chip-tool) prints or exports session keys during
commissioning; enter them in Wireshark → Preferences → Protocols → Matter.
See [matter-dissector README](https://github.com/project-chip/matter-dissector#security-features).
2. **well-known CASE keys**: both sides compiled with
`MATTER_CONFIG_SECURITY_TEST_MODE` / `CASEUseKnownECDHKey`; not enabled in
this environment (ESP32-C2 + HA official matter-server).
**Expected**
- Matter dissector expands protocol headers, IM commands, and cluster content.
**Stop condition (decryption)**: with no session keys or test keys obtainable,
do not fabricate keys to force a decrypt — plaintext mDNS + TCP handshake
still resolves most troubleshooting; for plaintext payloads, upgrade to
exporting keys on the commissioner side, then return to this runbook.
## Targeted capture variants
### One device only (known MAC)
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -tt 'ether host 34:98:7a:27:7f:08 and (udp port 5353 or tcp port 5540 or tcp port 5552)'"
```
MACs from `docs/matter-pairing-troubleshoot.md` §3 (working bulb
`34:98:7a:25:a1:f0`, broken bulb `34:98:7a:27:7f:08`).
### mDNS announcements only (no 5540 noise)
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -tt 'udp port 5353'"
```
### Rotating capture with timestamped filename (multiple runs)
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -C 5 -W 12 -w /tmp/matter-\$(date +%H%M%S).pcap 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
> The nested `$(date)` must be escaped as `\$(...)` inside the double-quoted ssh
> command so the remote shell expands it.
For the hass point, prefix the same commands with
`ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 …"`.
## Capture point: gw switch0
**Status: verified 2026-08-22 — limited capture point; NOT a full mirror of
LAN55.**
- `gw` `switch0` (`eth1``eth3`, `192.168.55.254/24`) is LAN55's L2 aggregation
only while devices plug directly into the ER-X. EdgeOS ships tcpdump;
`tcpdump -ni switch0` follows Linux bridge semantics.
- **Live topology (verified 2026-08-22): the SE5420 core switch is deployed**
(management `192.168.66.253` up — TP-Link OUI `f8:c9:03`, web UI on
:80/:443) and the ER-X uplink is a **single switch0 member port**: `eth1`
link up, `eth2`/`eth3` down. All LAN55 wired devices (hass `.11`, Aqara M3
`.248`, SmartThings `.48`, UAP-AC-Lite `.5`) are reached via `switch0`
behind that one uplink. Same-segment wired↔wired unicast switches locally on
the SE5420 and never reaches `switch0`.
- **What `switch0` still sees:** cross-subnet (66↔55) unicast, traffic to/from
the gateway itself (DHCP, DNS forwarding, port-forwards), and LAN55 mDNS
multicast (flooded up the uplink). Use it only for those flows; for a full
commissioning conversation use the hass `end0` or AP `br0` point instead.
- **Full mirror:** only via SE5420 port mirroring (the switch cannot run
tcpdump). Not configured; out of scope here.
- **Verification commands (EdgeOS v3.0.1 build 5862409):**
- Interactive: `ssh ubnt@192.168.66.254` (or `zhiqiang`), then
`show interfaces ethernet` — port link states are the decisive check
(`eth1` up + `eth2`/`eth3` down = single uplink). `configure` (config
mode) also accepts `show ...`.
- Non-interactive (agent/script): `show`/`configure` are interactive-only
aliases on this build; use the op wrapper:
```bash
ssh ubnt@192.168.66.254 '/opt/vyatta/bin/vyatta-op-cmd-wrapper show interfaces ethernet'
```
- `show ethernet-switch port all` and `show mac-address-table` are NOT
available on this build; the switch FDB is hardware-offloaded
(`brctl showmacs switch0` → "Operation not supported"). Port link state
+ ARP (`show arp`) are the reliable checks.
- SE5420 liveness: `ping 192.168.66.253` and `:80/:443`.
- Sample capture at this point (cross-segment/gateway/mDNS flows only;
tcpdump needs root — `zhiqiang` has passwordless sudo):
```bash
ssh zhiqiang@192.168.66.254 "sudo -n tcpdump -ni switch0 -s 0 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
## Pass criteria
- Realtime capture consistently shows the target device's mDNS announcements
(`_matterc` / `_matter._tcp`) on the chosen point.
- Commissioning shows the TCP 5540 handshake (SYN/SYN-ACK/ACK); on the hass
`end0` point this includes wired Thread-relay commissioning (HA↔M3), which
the AP point cannot see.
- Saved pcap opens in Wireshark and filters by `mDNS` / `tcp.port==5540`.
## References
- [docs/matter-pairing-troubleshoot.md](../docs/matter-pairing-troubleshoot.md) —
troubleshooting decision tree, stage table, device MAC/fabric facts
- [docs/unifi-network.md](../docs/unifi-network.md) — UniFi network/IPv6/SSID records
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — matter-server host
networking + `sudo -n -i` non-interactive note
- [hosts/gw.md](../hosts/gw.md) — DHCP `matter` reservation MAC mismatch (pending, W1N-207)
- [matter-dissector](https://github.com/project-chip/matter-dissector) —
Wireshark Matter dissector (incl. decryption)
- [Silabs: Using Wireshark to Capture Network Traffic in Matter](https://docs.silabs.com/matter/2.9.1/matter-references/matter-wireshark)
+74
View File
@@ -0,0 +1,74 @@
# Runbook: controlled network change
## Purpose
Apply a controlled network configuration change (DNS records, firewall, LAN
gateway, VLAN) with impact assessment, approval, and a rollback path.
## Scope
- Applicable: PowerDNS zone records, `us4` firewalld allowlist, LAN gateway/VLAN/DNS changes, WireGuard.
- Not applicable: SSH access-policy changes (see `AGENTS.md` §SSH access safety — mandatory lockout-risk procedure).
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: PowerDNS / us4 firewalld / LAN gateway / WireGuard
## Preconditions
- A change record (Linear `vps` issue) describes the change, its reason, and rollback.
- Read-only impact assessment done (current config captured, blast radius known).
## Safety
- Never change DNS or network config without a change record and approval.
- Capture the current config first; never delete existing config as the first action.
- For DNS: record the current record values and TTL before editing.
- For firewall: retain an independent SSH rollback session before applying (see `ansible-operations.md` §us4).
## Procedure
### Step 1 — Assess and capture
**Action** — capture the current state (e.g. `dig` for DNS, `--check --diff` for firewall, `show` for gateway).
**Expected** — a baseline of current config and an identified blast radius.
**Decision** — change fully specified with rollback → Step 2; missing → `STOP`.
### Step 2 — Approve
**Action** — confirm approval is recorded in the issue/change record.
**Decision** — approved → Step 3; not approved → `STOP`.
### Step 3 — Change
**Action** — apply the single change (edit the record, run the gated playbook, or change gateway config) and only that change.
**Expected** — the new value/state is in effect.
**Verification** — re-query/verify the new state and confirm dependent services still pass health.
**Rollback** — restore the captured prior config and re-verify.
## Troubleshooting
### Troubleshooting A — Change broke dependent service
- Evidence: health report / endpoint failure.
- Allowed: roll back to the captured prior config.
- Next: verify; if still broken, escalate.
## Final Verification
- New state verified; dependent services healthy.
- Change and outcome recorded in the issue.
## References
- [`ansible-operations.md`](ansible-operations.md)
- [`rollback.md`](rollback.md)
- [`network-recovery.md`](network-recovery.md)
+67
View File
@@ -0,0 +1,67 @@
# Runbook: network outage / service recovery
## Purpose
Recover from a network outage or service failure, starting from read-only
diagnosis and mutating only when the root cause is confirmed.
## Scope
- Applicable: unreachable VPS services, LAN gateway/DNS failures, DNS resolution failures.
- Not applicable: planned changes (→ `network-change.md`), SSH access recovery (→ `AGENTS.md` §SSH access safety).
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: VPS services / LAN gateway / DNS
## Safety
- Read-only diagnosis first; do not mutate while the root cause is unknown.
- If live state conflicts with a runbook's assumptions, `STOP` and report.
- Keep the current verified management session open as the recovery path.
## Procedure
### Step 1 — Diagnose (read-only)
**Action** — gather evidence without changing anything:
```bash
# From laptop, pin DNS to a public resolver if the stub is flaky
dig @1.1.1.1 +short <host> A
curl -4 -sS -I --max-time 10 https://<host>/
# From a reachable host, inspect the service
ssh -4 windy@<host> 'docker compose ps -a; df -h /; tail -n 50 /var/lib/vps-health/latest.json'
```
**Expected** — a clear picture: is it DNS, connectivity, host, or service?
**Decision** — root cause localized → Step 2; ambiguous or conflicting → `STOP` and escalate (provider console if host is unreachable).
### Step 2 — Confirm and route
**Action** — match the failure to the owning runbook (`mailcow-health.md`, `pdns-health.md`, `matrix-health.md`, etc.) or `network-change.md` for a config fix.
**Expected** — an applicable runbook with a recovery action.
**Decision** — applicable → follow it; none → `STOP` (diagnose only, do not mutate).
### Step 3 — Recover (gated)
**Action** — apply only the runbook's documented recovery, with approval.
**Verification** — re-run the health report / endpoint check and confirm recovery.
**Rollback** — if recovery makes it worse, revert per `rollback.md`.
## Final Verification
- Service reachable and health report green.
- Incident and recovery recorded in the Linear `vps` issue.
## References
- [`network-change.md`](network-change.md)
- Per-service health runbooks under [`runbooks/`](.)
+22 -1
View File
@@ -1,5 +1,20 @@
# PowerDNS health (hk2)
## Purpose
Read-only health check of the `/opt/pdns` PowerDNS stack.
## Scope
- Applicable: [hk2.chans.xyz](../hosts/hk2.chans.xyz.md), `/opt/pdns`.
- Read-only: does not change PowerDNS, DNS records, or secrets.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: hk2.chans.xyz (`/opt/pdns`)
Read-only checks for the `/opt/pdns` stack on **hk2.chans.xyz** (`ns1.wsvc.info`).
Facts: [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) · Upstream: [docs/pdns-upstream.md](../docs/pdns-upstream.md)
@@ -18,7 +33,7 @@ Use these only after the Ansible health report needs investigation.
ssh -4 windy@hk2.chans.xyz 'cd /opt/pdns && docker compose ps -a'
```
Expect `auth`, `db`, `poweradmin` healthy; `backup` Up; `pgweb` Up. Ignore stopped orphan `powerdns-admin` unless cleaning orphans.
Expect `auth`, `db`, `poweradmin` healthy; `backup` Up; `pgweb` Up. The legacy PDA orphan `powerdns-admin` was removed 2026-08-12 (W1N-59).
### Version / security poll
@@ -87,6 +102,12 @@ Expect: `primary=yes`, `also-notify=202.91.35.141`, `only-notify=` empty, `gpgsq
The sanitized Ansible health profile is `pdns` (`ansible/playbooks/healthchecks.yml`). It runs locally through `vps-healthcheck.timer`, writes a sanitized JSON result to `/var/lib/vps-health/latest.json`, and uses the API key only inside the PowerDNS container. It does not modify PowerDNS, DNS records, or secrets.
## Safety
- Read-only: never mutate PowerDNS configuration or DNS records during this check.
- Do not paste the API key into chat/logs.
- If live state conflicts with an expected value, `STOP` and report.
## After config changes
- `auth/pdns.conf`, `auth/templates.d/secrets.j2`, or auth-related `.env` → use the Ansible Compose reconcile playbook with target `auth`
+181
View File
@@ -0,0 +1,181 @@
# Runbook: pgdb health (TimescaleDB + pgweb + pg-backup)
## Purpose
Read-only health check of the pgdb TimescaleDB compose stack (PG18 + pgweb GUI + nightly custom-format backups). Confirms the stack is serving Home Assistant (hass/scribe) and that backups are current and restorable.
## Scope
- Applicable: [pgdb](../hosts/pgdb.md) (`192.168.55.15`), `/opt/database` compose stack.
- Read-only: never mutates containers, databases, backups, or secrets.
- Not applicable: restoring data (use [pgdb-restore](pgdb-restore.md)), upgrading images (use [pgdb-update](pgdb-update.md)), HA-side changes (see `hosts/hass.windy.lan.md`).
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-29
- Related systems: pgdb (`/opt/database`, TimescaleDB 18.6 / TS 2.29.2), HA `192.168.55.11` (hass/scribe clients)
## Access
SSH to pgdb (key auth works from the WSL agent shell as of 2026-08-29):
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15
```
From the agent sandbox, always use `-F /dev/null` (system ssh config is unreadable there) and prefer IPv4. The compose project lives at `/opt/database` — prefix every `docker compose` call with `cd /opt/database`. Never print `.env` values or bookmarks (they contain DB passwords); compare or use them only inside commands that output non-secret signals (status codes, counts, names).
## Safety
### Non-negotiable rules
- Read-only diagnosis only; never "fix while checking".
- Never print passwords or secrets — redact/consume them inside commands.
- If live state conflicts with an expected value, `STOP` and record evidence; do not invent parameters or bypass a failed check.
- Restore/update work belongs to the change runbooks, not this one.
### Stop conditions
- Any container `Exited`, `Restarting`, or not `healthy` where expected.
- Expected database/table/hypertable missing or a key query errors.
- Write-activity sample does not increase (scribe `states_raw` static).
- Latest daily backup older than today, not custom format, or `pg_restore -l` fails.
- Disk usage near full on `/srv/pgdata` or `/`.
- New `ERROR`/`FATAL` lines in the timescaledb log or backup failures in the pg-backup log.
## Pass criteria
- `docker compose ps -a`: `timescaledb` + `pg-backup` **Up (healthy)**, `pgweb` **Up**; ports bound to `192.168.55.15:5432` and `:8081`.
- PG 18.x; databases `hass`, `scribe`, `postgres` present; HA (`192.168.55.11`) connected as `hass` to both `hass` and `scribe`.
- `hass.states` and `scribe.states_raw` row counts grow between two samples (scribe writes continuously).
- TimescaleDB extension 2.29.x; scribe hypertables `states_raw` + `events` (1-dim, `time`); compression configured (segmentby/orderby rows in `timescaledb_information.compression_settings`); `entities` table exists.
- pgweb: no credentials → HTTP 401; with credentials → HTTP 200; `/api/bookmarks``["hass","scribe"]`.
- `daily/*-latest.dump` symlinks point to today's dumps; `file -L` reports `PostgreSQL custom database dump`.
- `/srv/pgdata` (`/dev/sdb1`, 32G) and `/` not near full; fstab mounts `/srv/pgdata` by `UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86` with `defaults,noatime`.
- timescaledb log: no new `ERROR`/`FATAL`; pg-backup log: recent successful backup.
## Checks
### 1. Containers
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose ps -a --format "table {{.Name}}\t{{.Status}}\t{{.Ports}}"'
```
**Expected**
- `timescaledb` Up (healthy), `pg-backup` Up (healthy), `pgweb` Up.
- Ports: `192.168.55.15:5432->5432/tcp` (timescaledb), `192.168.55.15:8081->8081/tcp` (pgweb).
**Stop** if any container is `Exited`/`Restarting`/`unhealthy`, or a port binding changed.
### 2. PG core and HA clients
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select version();" | head -1'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select datname from pg_database where datistemplate=false order by 1;"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select datname, usename, client_addr from pg_stat_activity where client_addr is not null group by 1,2,3 order by 1;"'
```
**Expected**
- `PostgreSQL 18.x` (verified: 18.6).
- Databases: `hass`, `postgres`, `scribe`.
- HA sessions: `hass|hass|192.168.55.11` and `scribe|hass|192.168.55.11` (the HAOS recorder/scribe clients from `192.168.55.11`).
**Stop** if a database is missing, the version is not 18.x, or HA has no live sessions (scribe connectivity is part of the HA pipeline).
### 3. Write activity
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database
A=$(docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select count(*) from states_raw;")
sleep 30
B=$(docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select count(*) from states_raw;")
echo "states_raw $A -> $B"'
```
Also sample `hass.states` once (recorder table, bulk-writes on HA restart): `docker compose exec -T timescaledb psql -U postgres -d hass -Atc "select count(*) from states;"`.
**Expected** — `states_raw` increases between samples (verified: 2665 → 2693 in 30 s). `states` count is sane (thousands).
**Stop** if `states_raw` is static across samples while HA is up — writes have stalled.
### 4. TimescaleDB (extension, hypertables, compression)
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select extversion from pg_extension where extname='"'"'timescaledb'"'"';"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select hypertable_name, num_dimensions from timescaledb_information.hypertables order by 1;"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select hypertable_name, attname, segmentby_column_index, orderby_column_index from timescaledb_information.compression_settings order by 1,3,4;"'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select to_regclass('"'"'public.entities'"'"');"'
```
**Expected**
- Extension version `2.29.x` (verified: 2.29.2).
- Scribe hypertables: `states_raw` and `events`, both `1` dimension.
- Compression configured for `states_raw` (segmentby `metadata_id` idx 1, orderby `time` idx 1) and `events` (segmentby `event_type`, orderby `time`). Note: TimescaleDB 2.29.x has **no** `compression_enabled` column in this view — row presence is the enabled signal.
- `entities` resolves (scribe registry table).
**Stop** if the extension version differs from the pinned 2.29.x line, a hypertable is missing, compression rows vanish, or `entities` is absent (scribe schema broke — see Known issues in [hosts/pgdb.md](../hosts/pgdb.md)).
### 5. pgweb GUI
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database
curl -s -o /dev/null -w "no-auth:%{http_code}\n" --max-time 8 http://192.168.55.15:8081/
U=$(grep -E "^PGWEB_AUTH_USER=" .env | cut -d= -f2-); P=$(grep -E "^PGWEB_AUTH_PASS=" .env | cut -d= -f2-)
curl -s -o /dev/null -w "auth:%{http_code}\n" --max-time 8 -u "$U:$P" http://192.168.55.15:8081/
echo -n "bookmarks:"; curl -s --max-time 8 -u "$U:$P" http://192.168.55.15:8081/api/bookmarks; echo'
```
Use `192.168.55.15:8081` (pgweb binds the VM IP only — loopback is not bound). Credentials are read from `.env` on the host and never printed.
**Expected** — `no-auth:401`, `auth:200`, `bookmarks:["hass","scribe"]`.
**Stop** if pgweb is unreachable, unauthenticated access is not 401, or bookmarks diverge from `["hass","scribe"]`.
### 6. Backups
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'ls -l --time-style=long-iso /opt/database/backups/daily/'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'file -L /opt/database/backups/daily/hass-latest.dump'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T pg-backup pg_restore -l /backups/daily/hass-latest.dump | head -4'
```
**Expected**
- `daily/*-latest.dump` symlinks point to **today's** `*-YYYYMMDD.dump` (nightly 02:00 local `Asia/Shanghai`; a fresh container also fires `BACKUP_ON_START`).
- `file -L` reports `PostgreSQL custom database dump` (pg_restore format; verified v1.16-0).
- `pg_restore -l` from the **pg-backup** container lists the archive TOC without error (timescaledb does not mount `/backups`).
**Stop** if the latest dump is not from today, is not custom format, or `pg_restore -l` fails. Stray non-`daily/` dumps at the `/opt/database/backups/` root are pre-compose leftovers — ignore for health, flag for cleanup.
### 7. Disk and mount
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'df -h /srv/pgdata /'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'grep -E "srv/pgdata" /etc/fstab'
```
**Expected**
- `/srv/pgdata` = `/dev/sdb1` 32G (verified: 88M used / 30G avail) and `/` with comfortable headroom.
- fstab: `UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86 /srv/pgdata ext4 defaults,noatime 0 2`.
**Stop** if either filesystem is near full (define threshold before acting) or the fstab entry is missing/changed.
### 8. Logs
```bash
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose logs --since 24h timescaledb 2>&1 | grep -E "ERROR|FATAL" | tail -10'
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose logs --since 24h pg-backup 2>&1 | tail -5'
```
**Expected**
- timescaledb: no new `ERROR`/`FATAL`. Known benign history: an old `relation "hass.states" does not exist` from a wrong-schema probe and `compression_enabled` column errors from an outdated query — neither recurs with the commands above.
- pg-backup: recent successful run (`SQL backup created successfully` for each database, no restore/cleanup errors).
**Stop** if repeated `ERROR`/`FATAL` appear or a backup run failed.

Some files were not shown because too many files have changed in this diff Show More