Author SHA1 Message Date
windyboy ff1a92110c docs(us2): Soft Serve → Gitea 迁移事实与参考镜像 (Plane VPS-94)
- hosts/us2: Gitea 1.27.3-rootless 部署实况 (repo.windy.me SSH:2222/Web), 16 仓迁移核对, 备份/回滚; soft-serve 停用保留作回滚
- compose/gitea: 参考镜像 (rootless compose + 备份 sidecar + 一次性迁移脚本留档)
- AGENTS/inventory/compose README: 服务表与索引同步
2026-09-18 17:24:35 +08:00
windyboy cc3fb99c14 docs(us2): 根盘清理 71%→23% 事实记录 (Plane VPS-93) 2026-09-18 11:46:17 +08:00
windyboy 6879d79cc6 docs(hass,pgdb): scribe 4.4.0 升级核对 + stats_io_interval 300 + sensor_minute 体积诊断
- hass: Scribe 3.8.0 → 4.4.0(HACS jonathan-gtd/scribe,2026-09-13 随 Core
  2026.9.1 / HAOS 18.2 升级)。记录 4.0 两个 breaking change 在本机均无需动作
  (3.x 结构 DB + PK 在、TimescaleDB 2.29.2 已装)、配置优先级
  YAML > options > entry data > 默认值、以及「YAML 改动必须重启 Core」。
- hass: 新增 stats_io_interval: 300(备份 scribe.yaml.bak-20260913-191558)。
  变更前 24h scribe 自写 11 019/87 461 行(12.6%);实测发布间隔 300s。
  retention_states/retention_events 可用但刻意不设;flush_interval 仍被 entry
  data 钉在 5s(采用新默认 30s 需显式写 YAML)。
- pgdb: sensor_minute 2.26 GB 是压缩窗口内的正常暂存(compress_after 按 chunk
  结束时间判断,09-17 才合格),稳态 4–5 GB 平台期,暂不处理;复查点 09-17
  之后。顺带修正 hass/scribe 库体积事实。

校验:scripts/validate-repo.sh PASS (0 warnings)。
2026-09-13 19:34:04 +08:00
windyboy c445c5f512 Merge remote-tracking branch 'origin/main' into main
AGENTS.md 冲突(两侧都改了 Linear→Plane 记录源):取本地更完整的表述
(self-hosted plane.chans.xyz + mcp__plane__* + 2026-09-03 停用日期 +
issue-tracker.md 已过时),并吸收远端的 `plane-workflow` skill 指引。
2026-09-13 19:13:06 +08:00
windyboy 3c83246f24 docs(pgdb,gfw): pgdb VM 重启根因 (W1N-263) + gfw Quad9 上游移除记录 2026-09-13 19:12:32 +08:00
windyboy c0c975584a docs(plane): 自托管 Plane 落地事实入仓库 + plane-health runbook + hardening 草稿
记录源 Linear→Plane (2026-09-03 起, Plane MCP) + plane.chans.xyz 服务行/upstream 段;
inventory + hosts/synapse.chans.xyz.md 补 Plane 部署事实 (Helm plane-ce-1.8.0 / app v1.4.1,
ns plane, IngressRoute/自有证书 issuer/PVC 5+5Gi local-path/无备份层);
新增 runbooks/plane-health.md (只读健康检查) 与 docs/plane-hardening/ 草稿
(values.hardened.yaml、secrets.yaml.example 占位、backup/ CronJob), 均为未应用设计稿;
.gitignore 增加 .tmp-* agent 临时文件。
2026-09-13 19:12:32 +08:00
windyboy 0034cec925 Merge remote-tracking branch 'origin/main' into HEAD 2026-09-13 19:11:13 +08:00
windyboy c9dcde1274 docs(hass): Quick 面板时间范围扩容 — 功率/环境/人体感应加 3d-30d 档,用电量加 3mo (VPS-92)
卡片 JS 只接受 <n>m|<n>h|<n>d 与命名档(today/week/month/3mo/6mo/year/custom);
记录 scribe sensor_minute 数据下界 2026-08-29 对 >15d 档的影响。
2026-09-13 19:09:54 +08:00
windyboy 037c4ccaa5 docs(hass): 马桶换气电源 Matter 插座接入 + 功率积分补电量 (VPS-92)
Matter Smart Plug (SIXWGH model 3596) 的 cluster 0x0091 声明 IMPE+CUME+PERE 但
CumulativeEnergyImported 恒为 null -> HA 能量实体永久 unknown。功率计量正常
(实测 24.7 W),故用 Integration (Riemann sum) 辅助元素补电量实体
sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_energy:
能源仪表盘 grid 源 [8] 与 Quick「用电量(按插座)」图换到该实体,Quick 另加
span-2「开关」区块。记录 config-flow 经 supervisor 代理走 REST 的 Agent 方法、
备份路径与回滚方式;energy/validate 已全绿。
2026-09-13 18:39:22 +08:00
windyboy 17bb171578 docs(hass): CSG 面板重构入维护 runbook — avg price 传感器 + 季节提醒 automation + WS 改法 (VPS-90) 2026-09-04 21:55:40 +08:00
windyboy cecf7e6331 docs(hass): CSG 电力监控面板重构 — avg price 传感器 + 去重/改名 + 季节切换提醒 (VPS-90) 2026-09-04 21:44:43 +08:00
windyboy e957bc2bb1 docs: add host-disk-cleanup runbook + hk2 disk facts; record source -> Plane vps (VPS-81) 2026-09-02 18:02:22 +08:00
windyboy de52cb8b57 docs(hass): 地图卡 CARTO 水印修复 — custom:map-card v1.16.0 + keyed tiles (W1N-261) 2026-08-30 13:39:42 +08:00
windyboy b15e19bce9 docs(pgdb): compose 开机竞态故障修复 + 自愈 unit (W1N-260)
- 根因:开机时 docker 恢复容器绑定 192.168.55.15:5432/8081 失败(EADDRNOTAVAIL,IP 尚未可绑)→ timescaledb/pgweb 启动失败且不重试,停摆 3h15m;pg-backup 开机备份失败 → unhealthy
- 处置:docker compose up -d --force-recreate(三容器回 database_default、端口发布、备份恢复、pgweb 恢复);用户重启 HA Core 后写入管道恢复
- 防复发:新增开机自愈 systemd oneshot pgdb-compose.service(enabled),源码 compose/pgdb/pgdb-compose.service
2026-08-30 13:19:05 +08:00
windyboy bee54a6858 docs(hass): CSG 长期归档 csg_history + recorder 365d 补录 (W1N-243) 2026-08-30 13:19:05 +08:00
windyboy d6747028b4 docs(soft-serve): 镜像固定 v0.12.2 + 备份 sidecar + 非 root 运行 (W1N-244..248)
- 镜像 pinned charmcli/soft-serve:v0.12.2(GHCR 为 dev/nightly 源,无 v0.12.x tag)
- soft-serve-backup sidecar:每日 02:00 sqlite .backup + repos-config 打包,03:00 prune 保留 14 份
- 非 root 运行(user 1000:1000),data chown;ssh.public_url 修复
- compose 源码参考:compose/soft-serve/(服务器文件为准)
2026-08-30 13:19:05 +08:00
windyboy f174aa1219 docs(hass): CSG 复核遗留修复 — 模板 days[-1] 补排序 + 本月日均/预测进度上屏 (W1N-242) 2026-08-29 21:08:13 +08:00
windyboy 5b5f6042e6 docs(hass): CSG 模板/面板修复记录 — availability 硬化 (W1N-239)、off-by-one + 新传感器 (W1N-241)、gauge 阶梯对齐 + 年度统计 (W1N-240)
- hosts/hass.windy.lan.md: csg_sensors 三次变更记录 + 季节性 gauge 切换已知事项 (11-01/5-01)
- runbooks/home-assistant-maintenance.md: float(0) fake-zero follow-up 标记已解决
2026-08-29 21:02:33 +08:00
windyboy 50136b2ffd docs(hass): W1N-238 — scribe config split to scribe.yaml + templates/ merge include
- Scribe 3.8.0 block moved verbatim from configuration.yaml to
  /homeassistant/scribe.yaml (scribe: !include scribe.yaml); YAML stays
  authoritative, import semantics unchanged.
- template: switched to !include_dir_merge_list templates; new
  quick_sensors.yaml scaffold (top-level list, quick_ prefix, unique_id
  required; pure sums stay min_max per W1N-233).
- Verified post-restart 2026-08-29 20:19 CST: core check ok, scribe
  connection on, states_written 18581→19426, template entities = 12,
  no scribe/template log errors. Backup configuration.yaml.bak-20260829-201724-w1n238.
2026-08-29 20:25:53 +08:00
windyboy 6707cebc88 chore: gitignore .agent-work/ agent scratch dir
Untracked agent scratch made validate-repo.sh link scan fail; same
category as the already-ignored .agents/ and .claude/ dirs.
2026-08-29 20:25:53 +08:00
windyboy 908ff5412a docs(hass): Quick dashboard round-2 state — badges, total-power helper, span-2 power pair, fill/colors (W1N-231)
Sync Quick dashboard section with live config: 2x2 mushroom light grid,
kong_diao AC entity fix, heading badges (env temps / AC / PC / total power),
min_max sum helper sensor.dang_qian_zong_gong_lu (fail-closed), selective
tozeroy fill on base-load chart, power charts paired as span-2 sections with
card titles removed, motion per-entity colors.
2026-08-29 19:54:35 +08:00
windyboy e58283210a chore(vaultwarden,healthcheck): upgrade 1.37.2 (Bitwarden 2026.8+); fix vps-health checks
- vaultwarden/server:1.37.1 -> 1.37.2 (required for Bitwarden clients 2026.8.0+)
- compose probe: flag only active services (config --services) so debug-profile
  pgweb 'Exited' no longer false-positives
- runner: build aggregate args line-by-line (robust vs Jinja trim_blocks)
- SMTP AUTH probe moved host-side (vaultwarden image has no python3); never
  prints the SMTP password
- us2 facts: probe refresh 2026-08-29, image/version, vps-health install
2026-08-29 19:54:32 +08:00
windyboy 8c73d1f894 docs(hass,pgdb): timescale-plotly-card chart stack — reader+card install, sensor_minute pipeline, Quick dashboard
- reader timescale_database_reader v1.1.0 (bb8776a) + card timescale-plotly-card
  2.2.0 (217961d), manual installs; config entry, Lovelace resource id recorded
- pgdb scribe: sensor_minute_aggregate cagg + sensor_minute hypertable + jobs
  1005/1006/1007; states_raw 3-month retention/compression statements deliberately
  skipped (permanent archive per host doc)
- sensor_minute_refresh local patch ELSE 0 → ELSE NULL (unavailable-minute zeros
  poison diff-mode energy charts) + one-time cleanup (505 head rows, 26 impossible
  zeros); re-apply after re-running upstream 02 SQL
- Quick dashboard: 5 chart cards via WS lovelace/config/save; documented section
  column_span (absent → span 1) vs card grid_options sizing rules
2026-08-29 15:54:40 +08:00
windyboy bc0a86245d docs: Scribe retention v4.x status (user declined RCs, 2026-08-29); HA Core 2026.8.3 verified 2026-08-29 14:56:05 +08:00
windyboy 13032fd0bb Merge origin/main (production Makefile) into pgdb runbooks delivery 2026-08-29 14:52:10 +08:00
windyboy d2063e7496 Merge pgdb ops runbooks — health/restore/update (W1N-228) 2026-08-29 14:51:31 +08:00
windyboy 2d95f87897 docs(runbooks): pgdb ops runbooks — health / restore / update + facts refresh (W1N-228)
- runbooks/pgdb-health.md: read-only health check (8 diagnostics) — containers,
  PG core + HA clients, write activity, TimescaleDB hypertables/compression,
  pgweb auth/bookmarks, daily custom-format backups, disk/fstab, logs
- runbooks/pgdb-restore.md: procedure-type restore (pg_restore -Fc, temp-DB swap,
  approval gates, rollback) — precondition command verified live
- runbooks/pgdb-update.md: gated command reference (pull -> config -q -> up -> verify;
  rollback = /opt/database/run + old volumes)
- index + validate-repo.sh classification updated; hosts/pgdb.md refreshed
  (SSH key auth works, scribe events hypertable, runbook cross-refs)
2026-08-29 14:51:20 +08:00
windyboy b61513c93e chore(pgdb): land W1N-227 compose-化 leftovers (compose source, host facts, inventory, scribe notes) 2026-08-29 14:51:20 +08:00
windyboy ab808088b8 Add production Makefile for routine VPS ops
Wrap validate-repo.sh and routine Ansible playbooks with safe-by-default
targets: read-only health/audit flows, CONFIRM=1 gates for mutating work,
and LIMIT/TARGETS guards. Document entry point in AGENTS.md.
2026-08-26 11:27:02 +08:00
windyboy c7dc4fd25c chore: remove duplicate hook setup
Keep pre-commit configuration as the single validation path and tolerate deleted tracked Markdown during link validation.
2026-08-23 17:58:19 +08:00
windyboy 7bc7d3f99b docs: align runbooks and validation structure 2026-08-23 17:42:31 +08:00
windyboy fea9a6560f chore: gitignore .opencode/ and .zcode/ local tool caches 2026-08-23 17:10:07 +08:00
windyboy a95b626636 docs: Matter bulbs failure mode C — both bulbs announce mDNS but refuse TCP 5540 (08-23 read-only verification); record 08-22 add/loop saga, working bulb MAC change, 3-fabric map, PD rotation to 238:4812; refresh stale DHCP-reservation note on gw (W1N-207) 2026-08-23 11:06:35 +08:00
windyboy 27fe9c078e docs: archive historical planning documents and fix references
Move self-described historical/upstream docs to docs/archive/:
- agent-runbook-guide.md
- lan-core-switch-upgrade-plan.md
- lan-rb5009-upgrade.md
- se5420-review-claim-verification-2026-08.md

Update archive/README.md manifest and fix relative links in active docs
and archived docs. Update AGENTS.md docs/ layout description.
2026-08-22 19:34:05 +08:00
windyboy aaa4ee312e docs: verify gw switch0 as limited capture point — SE5420 single-uplink (eth1 up, eth2/3 down), LAN55 wired hosts behind SE5420; record EdgeOS 3 CLI/access quirks (W1N-207) 2026-08-22 10:44:20 +08:00
windyboy 6b298491a8 docs: Matter packet-capture runbook v2 — hass end0 commissioner capture point, corrected AP-point scope (no wired↔wired unicast), UAP-AC-Lite model fix, SE5420 live status; fix hass interface end1→end0 (W1N-207) 2026-08-22 10:14:29 +08:00
windyboy 32631e2996 docs: add Matter pairing troubleshooting handbook; record bulb state + DHCP reservation mismatch (W1N-207) 2026-08-22 08:18:23 +08:00
windyboy 1426b4ecfe docs: matrix_e2ee v0.3.12/v0.3.9 notes; gw/ubnt IPv6 re-verification; agent sandbox SSH quirk (2026-08-20) 2026-08-21 08:57:53 +08:00
windyboy 1f5e58bf17 docs: record Matter/IPv6 findings — stale matter-server mDNS address, SSID cleanup, ER-X ULA infeasibility (W1N-207) 2026-08-21 08:56:26 +08:00
windyboy e5819eeba3 docs: correct hass hardware to x88 Pro physical box; sync CSG v1.3.2
- hass.windy.lan is a physical x88 Pro box (HAOS bare-metal, machine: green,
  CPE x88pro20, virtualization empty) — not PVE VM 180 (verified live 2026-08-18)
- Record CSG v1.3.2 (934f58c, W1N-118) deploy in maintenance runbook verify
  section and hosts live-tree section (backups now include w1n118)
2026-08-18 13:32:33 +08:00
windyboy 079332e082 docs: record matrix_e2ee v0.3.0 deploy; fix and rename matrix-e2ee update runbook
- Deploy v0.3.0 (main 216cc99, W1N-180 bot-initiated device verification
  wizard) on hass.windy.lan; backup matrix_e2ee.bak-20260818-v0.2.10
- Fix runbook: tag-only prerequisite (v0.3.0 was untagged), working-tree
  HEAD check, rsync exit-23 note, actual setup log line, post-deploy record
  step, ssh_config.d -F /dev/null gotcha
- Rename runbook matrix-e2e-update.md -> matrix-e2ee-update.md and update
  AGENTS.md/hosts references (domain is matrix_e2ee, double-e)
- Unify matrix_e2ee naming and update version history in
  docs/home-assistant-matrix.md
2026-08-18 13:31:15 +08:00
windyboy 7cedba7f51 docs: rename integration name from matrix_e2ee to matrix_e2e
- Rename runbook: matrix-e2ee-update.md -> matrix-e2e-update.md
- Update all references in AGENTS.md, hass.windy.lan.md,
  home-assistant-matrix.md to use the short name matrix_e2e
- The code domain stays matrix_e2ee (E2EE) in source; all
  doc prose and command references now use matrix_e2e
2026-08-18 13:31:15 +08:00
windyboyandCursor 343c5db415 feat: add gated Compose deploy and make inventory the host source of truth
Keep sanitized Compose sources in-repo with a confirmation-gated Ansible
playbook, add repo-wide validation, tighten runbook ownership/STOP/review
metadata, and archive stale research docs.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-17 17:36:39 +08:00
windyboy 885d977531 docs(runbooks): light-enhance home-assistant-maintenance and index it
Rebase onto origin/main surfaced runbooks/home-assistant-maintenance.md
(W1N-69) which predated the runbook reorg. Add Purpose/Scope/Safety
headers and add it to the README routing index (now 17/17 consistent).
2026-08-17 16:02:33 +08:00
windyboy b0c01b2551 docs(runbooks): add runbook spec, template, index and 6 first-batch runbooks; light-enhance existing 10
- RUNBOOKS.md: repo-level spec (six-field model, naming, safety, maturity path)
- runbooks/_template.md + README.md: standard template and 16-entry routing index
- new: issue-to-merge, fix-ci, release, rollback, network-change, network-recovery
- light-enhance 10 existing runbooks with Purpose/Scope/Safety headers
- AGENTS.md: point step 3 at index/spec, add runbook execution rules
- docs/agent-runbook-guide.md: archive of Manus AI guide
2026-08-17 15:59:46 +08:00
windyboy 047ac03346 chore(skills): remove vendored encrypted-dns-skill (installed globally via skills CLI) 2026-08-15 12:35:55 +08:00
windyboy 1ec9246156 docs: treat ha backups list as ignored arg, not a subcommand
ha backups --help has no list; extra positional args still print
the default backup list with exit 0. Keep the ha host update fix.
2026-08-14 22:44:02 +08:00
windyboy 1936b8f5fe docs: correct ha backups list CLI note in HA runbook
ha backups list exists on this host; only ha host update is missing.
2026-08-14 22:42:44 +08:00
windyboy eda6536ddb docs: record CSG v1.3.1 zip install and correct ha-maintenance restart failure
ha-maintenance.sh --restart-core --yes exited 1 in <1s without restarting
Core. Empty output is ssh failure hidden by 2>/dev/null + pipefail, not a
MOTD-strip after a successful restart. Direct `ha core restart` is the
working path.
2026-08-14 22:31:31 +08:00
windyboy 1dc880362d docs: add HA Matrix integration notes and record .local rewrite removals 2026-08-14 18:17:26 +08:00
windyboy 70aea6cd72 feat(ednsdiag): add DoQ/DoH3/DNSCrypt transports, proxy support, probe & compare 2026-08-14 18:17:26 +08:00
windyboy 8303d78caf Record W1N-105 CSG network step, P1, and fork master→main rename on hass.windy.lan. 2026-08-14 18:15:28 +08:00
windyboy ebfe7b8488 docs(gw): document EdgeOS PPPoE redial procedure 2026-08-14 16:26:23 +08:00
windyboy 88eaefda33 Record W1N-104 CSG auto dual-stack deploy on hass.windy.lan.
end1 IPv6 is on, wlan0 stays off, and the live custom component is de01914
with ip_family=auto after the IPv4 blackhole.
2026-08-14 16:11:03 +08:00
windyboy fcb76d3d5a Record W1N-102 CSG deploy and IPv4 blackhole on hass.windy.lan.
The host note now states the aiohttp IPv4 client is live, Core loaded it,
and home PPPoE IPv4 to 95598.csg.cn is currently blackholed while IPv6
works on gw. HA still has IPv6 disabled (W1N-85).
2026-08-14 15:47:08 +08:00
windyboy 6ae835037b hass.windy.lan: resolve health snapshot issues (W1N-70..76) and document findings
- Add tianqi weather recorder patch notes (W1N-75: _unrecorded_attributes)
- Document Bluetooth hci0 RTL8821CS instability (W1N-74) and eMMC lifetime
  10% (W1N-76) as known issues
- Rewrite runbook known-issues section: all snapshot items resolved; link
  hosts doc for the two remaining known issues
2026-08-13 19:48:18 +08:00
115 changed files with 6120 additions and 2473 deletions
@@ -1,52 +0,0 @@
name: CI
on:
push:
branches: [main]
pull_request:
permissions:
contents: read
jobs:
quality:
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v4
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache: true
- name: Check formatting
run: |
files="$(gofmt -l .)"
if [ -n "$files" ]; then
echo "Unformatted Go files:"
echo "$files"
exit 1
fi
- name: Check module files
run: go mod tidy && git diff --exit-code
test:
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
runs-on: ${{ matrix.os }}
steps:
- name: Check out repository
uses: actions/checkout@v4
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache: true
- name: Resolve modules
run: go mod download
- name: Test
run: go test ./...
- name: Vet
run: go vet ./...
@@ -1,10 +0,0 @@
# Go build outputs
/bin/
/dist/
*.exe
*.test
*.out
# Local development
.env
.DS_Store
-202
View File
@@ -1,202 +0,0 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [yyyy] [name of copyright owner]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
@@ -1,159 +0,0 @@
# Encrypted DNS Skill
[![CI](https://github.com/windyboy/encrypted-dns-skill/actions/workflows/ci.yml/badge.svg)](https://github.com/windyboy/encrypted-dns-skill/actions/workflows/ci.yml)
An [Agent Skill](https://agentskills.io/specification) and deterministic Go CLI
for querying and diagnosing encrypted DNS resolvers.
The skill tells an agent when and how to perform encrypted DNS diagnostics;
`ednsdiag` performs the protocol exchange. Agents do not need to construct DoH
URLs, TLS sessions, or DNS wire messages themselves.
## Status
| Protocol | Status | Standard |
| --- | --- | --- |
| DNS over HTTPS (DoH) | Available (GET and POST) | [RFC 8484](https://www.rfc-editor.org/rfc/rfc8484.html) |
| DNS over TLS (DoT) | Available (strict authentication) | [RFC 7858](https://www.rfc-editor.org/rfc/rfc7858.html), [RFC 8310](https://www.rfc-editor.org/rfc/rfc8310.html) |
| DNS over QUIC (DoQ) | Planned | [RFC 9250](https://www.rfc-editor.org/rfc/rfc9250.html) |
| DoH over HTTP/3 (DoH3) | Planned | RFC 8484 over HTTP/3 |
| DNSCrypt | Planned | [DNSCrypt protocol specification](https://github.com/DNSCrypt/dnscrypt-protocol) |
| Oblivious DoH (ODoH) | Research | [RFC 9230](https://www.rfc-editor.org/rfc/rfc9230.html) |
| Anonymized DNSCrypt | Research | [Anonymized DNSCrypt specification](https://github.com/DNSCrypt/dnscrypt-protocol/blob/master/ANONYMIZED-DNSCRYPT.txt) |
Run `ednsdiag capabilities` instead of assuming a protocol is implemented.
## Why a Skill and a CLI?
- `SKILL.md` provides compact instructions, safety boundaries, and result
interpretation for an AI agent.
- `ednsdiag` provides repeatable wire-format DNS, HTTP, TLS, input validation,
and structured JSON output.
- Reference files keep protocol, provider, and security details grounded in
authoritative sources without bloating the agent's active context.
The CLI never silently downgrades to plaintext DNS, and it does not connect to
addresses returned in DNS answers.
## Requirements
- Go 1.26 or later when running or building from source
- Network access to the selected encrypted DNS resolver
- A host that supports the [Agent Skills package format](https://agentskills.io/specification) when using the repository as a Skill
## Install the Skill
Clone or copy this repository into a skill discovery directory supported by
your agent host. Keep the repository layout intact so `SKILL.md`, `references/`,
`schemas/`, and the Go source remain together.
For example, in a host that discovers project-local skills from `.agents/skills`:
```bash
git clone https://github.com/windyboy/encrypted-dns-skill.git \
.agents/skills/encrypted-dns-skill
```
Discovery paths differ between hosts. Follow the host's documentation rather
than moving only `SKILL.md`.
## Run from Source
No precompiled executable is required. Go can compile and run the command from
the repository root:
```bash
go run ./cmd/ednsdiag capabilities
go run ./cmd/ednsdiag query example.com A --protocol doh --provider cloudflare
go run ./cmd/ednsdiag query gmail.com MX --protocol dot --provider google --timeout 5s
```
The first run may download the modules pinned in `go.mod` and `go.sum`.
To build a reusable local executable:
```bash
go build -o ./bin/ednsdiag ./cmd/ednsdiag
./bin/ednsdiag capabilities
```
Do not download or execute an unverified third-party binary. This repository
does not currently publish release binaries.
## Usage
```text
ednsdiag capabilities
ednsdiag version
ednsdiag query <domain> [type] \
[--protocol doh|dot] \
[--provider cloudflare|google|quad9|adguard] \
[--method post|get] \
[--timeout 5s]
```
Defaults are `A`, `doh`, `cloudflare`, `post`, and `5s`. `--method` applies
only to DoH. The timeout must be between `250ms` and `30s`.
Built-in resolver profiles:
| Provider | Profile |
| --- | --- |
| Cloudflare | Unfiltered |
| Google | Unfiltered |
| Quad9 | Security-filtered |
| AdGuard | Ad- and security-filtered |
Filtering policies can affect DNS answers. Results always identify the
provider and profile used.
## Result Semantics
Every query returns structured JSON compatible with
[`schemas/result-v1.schema.json`](schemas/result-v1.schema.json).
- `completed: true` means the encrypted protocol exchange completed; it does
not mean the DNS response was `NOERROR`.
- `dns.rcode` is the DNS result. `NXDOMAIN`, `SERVFAIL`, and `REFUSED` are DNS
outcomes, not transport failures.
- Empty `dns.answers` with `NOERROR` means NODATA.
- `transport.server_authenticated` reports resolver endpoint authentication.
- `dns.resolver_reports_dnssec_authenticated` reflects the resolver's AD bit;
it is not local DNSSEC validation.
- `transport.bootstrap: system_resolver` means the operating system resolver
was used to locate the encrypted resolver endpoint.
## Security Model
- DoH uses standard `application/dns-message` wire messages.
- DoT verifies the PKIX certificate chain and configured authentication domain.
- Plaintext fallback is prohibited.
- DNS errors are not retried through another protocol as transport failures.
- Provider and protocol results remain separate.
- DNS answers are data only; the tool does not make application connections to
returned addresses.
See [`references/security.md`](references/security.md) for the complete threat
model and privacy boundaries.
## Development
```bash
go test ./...
go vet ./...
```
Protocol behavior must remain aligned with
[`references/standards.md`](references/standards.md), provider changes with
[`references/providers.md`](references/providers.md), and output with the v1
JSON schema.
## Scope
This project targets client-to-recursive encrypted DNS diagnostics. It is not a
system stub resolver, an authoritative DNS server, a hosted DNS record manager,
or a zone-transfer tool.
## License
Licensed under the [Apache License 2.0](LICENSE).
@@ -1,90 +0,0 @@
---
name: encrypted-dns-skill
description: Query, probe, and compare DNS resolution through supported encrypted transports. Use for encrypted DNS record lookups, resolver connectivity tests, TLS and QUIC diagnostics, protocol comparisons, DNSSEC status inspection, and troubleshooting DoH, DoT, DoQ, DoH3, or DNSCrypt resolver endpoints.
---
# Encrypted DNS Diagnostics
Use `ednsdiag` for encrypted DNS work. Do not assemble protocol requests with
`curl`, `openssl`, or ad-hoc scripts when `ednsdiag` supports the operation.
The executable requires network access.
Prefer an installed `ednsdiag` executable. When it is unavailable and Go 1.26+
is installed, run the source from the skill root with:
```bash
go run ./cmd/ednsdiag <command> [arguments]
```
Do not download or execute an unverified binary automatically. Building from
source may require permission to download pinned Go modules.
## Check capabilities
Before attempting an operation, run:
```bash
ednsdiag capabilities
```
Only use a protocol when its reported status is `available`. Never describe a
`planned` or `experimental` capability as implemented.
## Commands
```bash
ednsdiag query example.com A --protocol doh --provider cloudflare
ednsdiag query gmail.com MX --protocol dot --provider google --timeout 5s
ednsdiag capabilities
ednsdiag version
```
Use `--method get` or `--method post` only with DoH. The default is POST.
Built-in providers are `cloudflare`, `google`, `quad9`, and `adguard`. Provider
filtering policies differ and are included in the result. `probe` and `compare`
remain reserved until their capabilities are implemented.
## Required behavior
- Use standard DNS wire messages for DoH, not provider-specific JSON APIs.
- Apply strict certificate and authentication-domain validation.
- Never silently downgrade to plaintext DNS.
- Do not retry `NXDOMAIN`, `NODATA`, `SERVFAIL`, or `REFUSED` through another
protocol as though they were transport failures.
- Keep results from different providers and protocols separate.
- Report every fallback attempt and its reason.
- Treat the DNS `AD` bit as validation reported by the selected resolver, not
as local DNSSEC validation.
- Do not connect to addresses returned in DNS answers.
## Result interpretation
- `completed: true` means a protocol exchange completed. It does not imply
`NOERROR`.
- Read `dns.rcode` for the DNS outcome.
- Read `transport.server_authenticated` separately from DNSSEC fields.
- Read `transport.bootstrap`; `system_resolver` means resolving the encrypted
resolver endpoint itself used the operating system resolver.
- Empty answers with `NOERROR` represent NODATA.
- A filtering resolver may synthesize `NXDOMAIN`; disclose the provider.
## References
- Read [references/standards.md](references/standards.md) before changing
protocol behavior.
- Read [references/security.md](references/security.md) before changing TLS,
bootstrap, fallback, endpoint, or privacy behavior.
- Read [references/providers.md](references/providers.md) before adding or
modifying a built-in provider.
- Keep output compatible with
[schemas/result-v1.schema.json](schemas/result-v1.schema.json).
## Scope
The target scope is widely deployed client-to-recursive encrypted DNS:
DoH, DoT, DoQ, DoH3, and DNSCrypt. ODoH and Anonymized DNSCrypt remain
research capabilities until explicitly marked available.
Do not use this skill for DNS-over-DTLS, zone transfers, authoritative-server
operation, changing hosted DNS records, or replacing the operating system's
stub resolver.
@@ -1,184 +0,0 @@
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"os"
"strings"
"time"
"github.com/windyboy/encrypted-dns-skill/internal/edns"
)
const version = "0.1.0-dev"
type capability struct {
Protocol string `json:"protocol"`
Status string `json:"status"`
Standard string `json:"standard,omitempty"`
Note string `json:"note,omitempty"`
}
type capabilitiesResult struct {
SchemaVersion int `json:"schema_version"`
Command string `json:"command"`
Version string `json:"version"`
Capabilities []capability `json:"capabilities"`
}
func main() {
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
}
func run(args []string, stdout, stderr io.Writer) int {
if len(args) == 0 {
writeUsage(stderr)
return 2
}
switch args[0] {
case "capabilities":
if len(args) != 1 {
fmt.Fprintln(stderr, "capabilities does not accept arguments")
return 2
}
result := capabilitiesResult{
SchemaVersion: 1,
Command: "capabilities",
Version: version,
Capabilities: []capability{
{Protocol: "doh", Status: "available", Standard: "RFC 8484", Note: "RFC wire format over HTTP GET or POST"},
{Protocol: "dot", Status: "available", Standard: "RFC 7858 and RFC 8310", Note: "strict PKIX and authentication-domain validation"},
{Protocol: "doq", Status: "planned", Standard: "RFC 9250"},
{Protocol: "doh3", Status: "planned", Standard: "RFC 8484 over HTTP/3"},
{Protocol: "dnscrypt", Status: "planned", Standard: "DNSCrypt protocol specification"},
{Protocol: "odoh", Status: "research", Standard: "RFC 9230", Note: "No maintained Go dependency has been selected."},
{Protocol: "anonymized-dnscrypt", Status: "research", Standard: "Anonymized DNSCrypt specification"},
},
}
return writeJSON(stdout, stderr, result)
case "version":
if len(args) != 1 {
fmt.Fprintln(stderr, "version does not accept arguments")
return 2
}
fmt.Fprintln(stdout, version)
return 0
case "query":
options, timeout, err := parseQueryArgs(args[1:])
if err != nil {
fmt.Fprintln(stderr, err)
writeQueryUsage(stderr)
return 2
}
ctx, cancel := context.WithTimeout(context.Background(), timeout)
defer cancel()
result := edns.Query(ctx, options)
if code := writeJSON(stdout, stderr, result); code != 0 {
return code
}
if result.Completed {
return 0
}
if result.Error != nil && result.Error.Class == "input" {
return 2
}
return 3
case "probe", "compare":
fmt.Fprintf(stderr, "%s is not implemented in %s; run ednsdiag capabilities\n", args[0], version)
return 4
default:
fmt.Fprintf(stderr, "unknown command %q\n", args[0])
writeUsage(stderr)
return 2
}
}
func parseQueryArgs(args []string) (edns.QueryOptions, time.Duration, error) {
options := edns.QueryOptions{
RecordType: "A",
Protocol: "doh",
Provider: "cloudflare",
Method: "post",
}
timeout := 5 * time.Second
positionals := make([]string, 0, 2)
for index := 0; index < len(args); index++ {
argument := args[index]
if !strings.HasPrefix(argument, "--") {
positionals = append(positionals, argument)
continue
}
key, value, found := strings.Cut(strings.TrimPrefix(argument, "--"), "=")
if !found {
index++
if index >= len(args) {
return options, 0, fmt.Errorf("--%s requires a value", key)
}
value = args[index]
}
switch key {
case "protocol":
options.Protocol = strings.ToLower(value)
case "provider":
options.Provider = strings.ToLower(value)
case "url":
options.EndpointURL = value
case "method":
options.Method = strings.ToLower(value)
case "timeout":
parsed, err := time.ParseDuration(value)
if err != nil {
return options, 0, fmt.Errorf("invalid timeout %q: %w", value, err)
}
timeout = parsed
default:
return options, 0, fmt.Errorf("unknown query option --%s", key)
}
}
if len(positionals) < 1 || len(positionals) > 2 {
return options, 0, fmt.Errorf("query requires a domain and optional record type")
}
if timeout < 250*time.Millisecond || timeout > 30*time.Second {
return options, 0, fmt.Errorf("timeout must be between 250ms and 30s")
}
options.Name = positionals[0]
if len(positionals) == 2 {
options.RecordType = strings.ToUpper(positionals[1])
}
if options.Protocol != "doh" && options.Protocol != "dot" {
return options, 0, fmt.Errorf("protocol %q is not available", options.Protocol)
}
if options.Method != "get" && options.Method != "post" {
return options, 0, fmt.Errorf("DoH method must be get or post")
}
if options.Protocol == "dot" && options.Method != "post" {
return options, 0, fmt.Errorf("--method applies only to DoH")
}
return options, timeout, nil
}
func writeJSON(stdout, stderr io.Writer, value any) int {
encoder := json.NewEncoder(stdout)
encoder.SetIndent("", " ")
if err := encoder.Encode(value); err != nil {
fmt.Fprintf(stderr, "encode JSON result: %v\n", err)
return 1
}
return 0
}
func writeUsage(writer io.Writer) {
fmt.Fprintln(writer, "usage: ednsdiag <capabilities|version|query|probe|compare>")
}
func writeQueryUsage(writer io.Writer) {
fmt.Fprintln(writer, "usage: ednsdiag query <domain> [type] [--protocol doh|dot] [--provider cloudflare|google|quad9|adguard] [--url https://host/dns-query] [--method post|get] [--timeout 5s]")
}
@@ -1,82 +0,0 @@
package main
import (
"bytes"
"encoding/json"
"strings"
"testing"
"time"
)
func TestCapabilities(t *testing.T) {
var stdout bytes.Buffer
var stderr bytes.Buffer
code := run([]string{"capabilities"}, &stdout, &stderr)
if code != 0 {
t.Fatalf("run capabilities returned %d; stderr=%q", code, stderr.String())
}
var result capabilitiesResult
if err := json.Unmarshal(stdout.Bytes(), &result); err != nil {
t.Fatalf("decode capabilities: %v", err)
}
if result.SchemaVersion != 1 {
t.Fatalf("schema version = %d, want 1", result.SchemaVersion)
}
if result.Command != "capabilities" {
t.Fatalf("command = %q, want capabilities", result.Command)
}
if len(result.Capabilities) == 0 {
t.Fatal("capabilities list is empty")
}
available := map[string]bool{}
for _, item := range result.Capabilities {
available[item.Protocol] = item.Status == "available"
}
if !available["doh"] || !available["dot"] {
t.Fatalf("DoH and DoT must be available: %#v", available)
}
if available["doq"] || available["doh3"] || available["dnscrypt"] {
t.Fatalf("planned transports must not be available: %#v", available)
}
}
func TestReservedCommandIsNotImplemented(t *testing.T) {
var stdout bytes.Buffer
var stderr bytes.Buffer
code := run([]string{"probe"}, &stdout, &stderr)
if code != 4 {
t.Fatalf("run probe returned %d, want 4", code)
}
if !strings.Contains(stderr.String(), "not implemented") {
t.Fatalf("stderr = %q, want not implemented message", stderr.String())
}
}
func TestParseQueryArgsAllowsInterspersedOptions(t *testing.T) {
options, timeout, err := parseQueryArgs([]string{"example.com", "MX", "--protocol", "dot", "--provider=quad9", "--timeout", "3s"})
if err != nil {
t.Fatalf("parse query args: %v", err)
}
if options.Name != "example.com" || options.RecordType != "MX" || options.Protocol != "dot" || options.Provider != "quad9" {
t.Fatalf("unexpected options: %#v", options)
}
if timeout != 3*time.Second {
t.Fatalf("timeout = %v, want 3s", timeout)
}
}
func TestUnknownCommand(t *testing.T) {
var stdout bytes.Buffer
var stderr bytes.Buffer
code := run([]string{"unknown"}, &stdout, &stderr)
if code != 2 {
t.Fatalf("run unknown returned %d, want 2", code)
}
if !strings.Contains(stderr.String(), "unknown command") {
t.Fatalf("stderr = %q, want unknown command message", stderr.String())
}
}
@@ -1,7 +0,0 @@
module github.com/windyboy/encrypted-dns-skill
go 1.26.0
require golang.org/x/net v0.58.0
require golang.org/x/text v0.41.0 // indirect
@@ -1,4 +0,0 @@
golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To=
golang.org/x/net v0.58.0/go.mod h1:YwCddHnFlT7eLQqVprV19OnhLGtc5xOKgE0RyqgfWAU=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
@@ -1,229 +0,0 @@
package edns
import (
"crypto/rand"
"encoding/base64"
"encoding/binary"
"fmt"
"net"
"strings"
"golang.org/x/net/dns/dnsmessage"
"golang.org/x/net/idna"
)
var recordTypes = map[string]dnsmessage.Type{
"A": dnsmessage.TypeA,
"AAAA": dnsmessage.TypeAAAA,
"CNAME": dnsmessage.TypeCNAME,
"MX": dnsmessage.TypeMX,
"TXT": dnsmessage.TypeTXT,
"NS": dnsmessage.TypeNS,
"SOA": dnsmessage.TypeSOA,
"CAA": dnsmessage.Type(257),
"SRV": dnsmessage.TypeSRV,
"SVCB": dnsmessage.TypeSVCB,
"HTTPS": dnsmessage.TypeHTTPS,
}
func BuildQuery(name, recordType string) ([]byte, QueryInfo, uint16, error) {
canonical, err := canonicalName(name)
if err != nil {
return nil, QueryInfo{}, 0, err
}
typeName := strings.ToUpper(recordType)
qtype, ok := recordTypes[typeName]
if !ok {
return nil, QueryInfo{}, 0, fmt.Errorf("unsupported record type %q", recordType)
}
dnsName, err := dnsmessage.NewName(canonical + ".")
if err != nil {
return nil, QueryInfo{}, 0, fmt.Errorf("encode domain name: %w", err)
}
var randomID [2]byte
if _, err := rand.Read(randomID[:]); err != nil {
return nil, QueryInfo{}, 0, fmt.Errorf("generate DNS transaction ID: %w", err)
}
id := binary.BigEndian.Uint16(randomID[:])
message := dnsmessage.Message{
Header: dnsmessage.Header{ID: id, RecursionDesired: true},
Questions: []dnsmessage.Question{{
Name: dnsName,
Type: qtype,
Class: dnsmessage.ClassINET,
}},
}
wire, err := message.Pack()
if err != nil {
return nil, QueryInfo{}, 0, fmt.Errorf("pack DNS query: %w", err)
}
return wire, QueryInfo{Name: canonical, Type: typeName}, id, nil
}
func ParseResponse(wire []byte, expectedID uint16, query QueryInfo) (DNSInfo, error) {
var message dnsmessage.Message
if err := message.Unpack(wire); err != nil {
return DNSInfo{}, fmt.Errorf("unpack DNS response: %w", err)
}
if !message.Header.Response {
return DNSInfo{}, fmt.Errorf("received a DNS query instead of a response")
}
if message.Header.ID != expectedID {
return DNSInfo{}, fmt.Errorf("DNS transaction ID mismatch")
}
if len(message.Questions) != 1 {
return DNSInfo{}, fmt.Errorf("DNS response contains %d questions, want 1", len(message.Questions))
}
wantType := recordTypes[query.Type]
question := message.Questions[0]
if trimRoot(question.Name.String()) != query.Name || question.Type != wantType {
return DNSInfo{}, fmt.Errorf("DNS response question does not match request")
}
answers := make([]AnswerRecord, 0, len(message.Answers))
for _, resource := range message.Answers {
answers = append(answers, normalizeAnswer(resource))
}
return DNSInfo{
RCode: rcodeName(message.Header.RCode),
RCodeValue: int(message.Header.RCode),
ResolverReportsDNSSECAuthenticated: message.Header.AuthenticData,
ClientValidatedDNSSEC: false,
Answers: answers,
}, nil
}
func canonicalName(input string) (string, error) {
name := strings.TrimSuffix(strings.TrimSpace(input), ".")
if name == "" {
return "", fmt.Errorf("domain name is empty")
}
if net.ParseIP(name) != nil {
return "", fmt.Errorf("IP literals are not accepted as domain names")
}
ascii, err := idna.Lookup.ToASCII(name)
if err != nil {
return "", fmt.Errorf("convert domain name to IDNA ASCII: %w", err)
}
ascii = strings.ToLower(ascii)
if len(ascii) > 253 {
return "", fmt.Errorf("domain name exceeds 253 bytes")
}
for _, label := range strings.Split(ascii, ".") {
if label == "" || len(label) > 63 {
return "", fmt.Errorf("domain name contains an invalid label")
}
}
blocked := []string{"localhost", ".local", ".internal", ".lan", ".arpa"}
for _, suffix := range blocked {
if ascii == strings.TrimPrefix(suffix, ".") || strings.HasSuffix(ascii, suffix) {
return "", fmt.Errorf("domain name is blocked by the local-name policy")
}
}
return ascii, nil
}
func normalizeAnswer(resource dnsmessage.Resource) AnswerRecord {
record := AnswerRecord{
"name": trimRoot(resource.Header.Name.String()),
"type": typeName(resource.Header.Type),
"ttl": resource.Header.TTL,
}
switch body := resource.Body.(type) {
case *dnsmessage.AResource:
record["address"] = net.IP(body.A[:]).String()
case *dnsmessage.AAAAResource:
record["address"] = net.IP(body.AAAA[:]).String()
case *dnsmessage.CNAMEResource:
record["target"] = trimRoot(body.CNAME.String())
case *dnsmessage.MXResource:
record["priority"] = body.Pref
record["exchange"] = trimRoot(body.MX.String())
case *dnsmessage.TXTResource:
record["strings"] = body.TXT
case *dnsmessage.NSResource:
record["host"] = trimRoot(body.NS.String())
case *dnsmessage.PTRResource:
record["target"] = trimRoot(body.PTR.String())
case *dnsmessage.SOAResource:
record["primary_ns"] = trimRoot(body.NS.String())
record["responsible_mailbox"] = trimRoot(body.MBox.String())
record["serial"] = body.Serial
record["refresh"] = body.Refresh
record["retry"] = body.Retry
record["expire"] = body.Expire
record["minimum_ttl"] = body.MinTTL
case *dnsmessage.SRVResource:
record["priority"] = body.Priority
record["weight"] = body.Weight
record["port"] = body.Port
record["target"] = trimRoot(body.Target.String())
case *dnsmessage.SVCBResource:
addSVCBFields(record, body.Priority, body.Target, body.Params)
case *dnsmessage.HTTPSResource:
addSVCBFields(record, body.Priority, body.Target, body.Params)
case *dnsmessage.UnknownResource:
if resource.Header.Type == dnsmessage.Type(257) && len(body.Data) >= 2 {
record["flags"] = body.Data[0]
tagLength := int(body.Data[1])
if 2+tagLength <= len(body.Data) {
record["tag"] = string(body.Data[2 : 2+tagLength])
record["value"] = string(body.Data[2+tagLength:])
} else {
record["rdata_base64"] = base64.StdEncoding.EncodeToString(body.Data)
}
} else {
record["rdata_base64"] = base64.StdEncoding.EncodeToString(body.Data)
}
}
return record
}
func addSVCBFields(record AnswerRecord, priority uint16, target dnsmessage.Name, params []dnsmessage.SVCParam) {
record["priority"] = priority
record["target"] = trimRoot(target.String())
values := make([]map[string]any, 0, len(params))
for _, param := range params {
values = append(values, map[string]any{
"key": param.Key.String(),
"key_value": uint16(param.Key),
"value_base64": base64.StdEncoding.EncodeToString(param.Value),
})
}
record["params"] = values
}
func trimRoot(name string) string {
return strings.TrimSuffix(strings.ToLower(name), ".")
}
func typeName(recordType dnsmessage.Type) string {
for name, value := range recordTypes {
if value == recordType {
return name
}
}
return fmt.Sprintf("TYPE%d", recordType)
}
func rcodeName(rcode dnsmessage.RCode) string {
names := map[dnsmessage.RCode]string{
dnsmessage.RCodeSuccess: "NOERROR",
dnsmessage.RCodeFormatError: "FORMERR",
dnsmessage.RCodeServerFailure: "SERVFAIL",
dnsmessage.RCodeNameError: "NXDOMAIN",
dnsmessage.RCodeNotImplemented: "NOTIMP",
dnsmessage.RCodeRefused: "REFUSED",
}
if name, ok := names[rcode]; ok {
return name
}
return fmt.Sprintf("RCODE%d", rcode)
}
@@ -1,100 +0,0 @@
package edns
import (
"encoding/binary"
"testing"
"golang.org/x/net/dns/dnsmessage"
)
func TestBuildAndParseResponse(t *testing.T) {
queryWire, query, transactionID, err := BuildQuery("Example.COM.", "A")
if err != nil {
t.Fatalf("build query: %v", err)
}
if query.Name != "example.com" || query.Type != "A" {
t.Fatalf("canonical query = %#v", query)
}
var request dnsmessage.Message
if err := request.Unpack(queryWire); err != nil {
t.Fatalf("unpack query: %v", err)
}
response := dnsmessage.Message{
Header: dnsmessage.Header{
ID: transactionID,
Response: true,
RecursionDesired: true,
RecursionAvailable: true,
AuthenticData: true,
},
Questions: request.Questions,
Answers: []dnsmessage.Resource{{
Header: dnsmessage.ResourceHeader{Name: request.Questions[0].Name, Class: dnsmessage.ClassINET, TTL: 60},
Body: &dnsmessage.AResource{A: [4]byte{192, 0, 2, 1}},
}},
}
responseWire, err := response.Pack()
if err != nil {
t.Fatalf("pack response: %v", err)
}
dnsResult, err := ParseResponse(responseWire, transactionID, query)
if err != nil {
t.Fatalf("parse response: %v", err)
}
if dnsResult.RCode != "NOERROR" || !dnsResult.ResolverReportsDNSSECAuthenticated {
t.Fatalf("unexpected DNS result: %#v", dnsResult)
}
if got := dnsResult.Answers[0]["address"]; got != "192.0.2.1" {
t.Fatalf("address = %v, want 192.0.2.1", got)
}
}
func TestBuildQueryIDNAAndBlockedNames(t *testing.T) {
_, query, _, err := BuildQuery("bücher.example", "AAAA")
if err != nil {
t.Fatalf("build IDNA query: %v", err)
}
if query.Name != "xn--bcher-kva.example" {
t.Fatalf("IDNA name = %q", query.Name)
}
blocked := []string{"localhost", "router.local", "service.internal", "host.lan", "1.0.0.127.in-addr.arpa", "127.0.0.1"}
for _, name := range blocked {
if _, _, _, err := BuildQuery(name, "A"); err == nil {
t.Errorf("BuildQuery(%q) succeeded, want policy error", name)
}
}
}
func TestNormalizeCAA(t *testing.T) {
name := dnsmessage.MustNewName("example.com.")
data := append([]byte{0, 5}, []byte("issueletsencrypt.org")...)
record := normalizeAnswer(dnsmessage.Resource{
Header: dnsmessage.ResourceHeader{Name: name, Type: dnsmessage.Type(257), Class: dnsmessage.ClassINET, TTL: 300},
Body: &dnsmessage.UnknownResource{Type: dnsmessage.Type(257), Data: data},
})
if record["tag"] != "issue" || record["value"] != "letsencrypt.org" {
t.Fatalf("unexpected CAA normalization: %#v", record)
}
}
func TestParseResponseRejectsTransactionMismatch(t *testing.T) {
name := dnsmessage.MustNewName("example.com.")
message := dnsmessage.Message{
Header: dnsmessage.Header{ID: 2, Response: true},
Questions: []dnsmessage.Question{{Name: name, Type: dnsmessage.TypeA, Class: dnsmessage.ClassINET}},
}
wire, err := message.Pack()
if err != nil {
t.Fatalf("pack response: %v", err)
}
if _, err := ParseResponse(wire, 1, QueryInfo{Name: "example.com", Type: "A"}); err == nil {
t.Fatal("transaction mismatch was accepted")
}
if binary.BigEndian.Uint16(wire[:2]) != 2 {
t.Fatal("test response ID was not encoded")
}
}
@@ -1,131 +0,0 @@
package edns
import (
"bytes"
"context"
"crypto/tls"
"encoding/base64"
"fmt"
"io"
"mime"
"net"
"net/http"
"net/url"
"strings"
"time"
)
const maxDNSMessageSize = 65535
func exchangeDoH(ctx context.Context, provider Provider, wire []byte, method string) ([]byte, TransportInfo, error) {
client := newDoHClient(provider.DoHURL)
return exchangeDoHWithClient(ctx, client, provider.DoHURL, wire, method)
}
func newDoHClient(endpoint string) *http.Client {
origin, _ := url.Parse(endpoint)
transport := &http.Transport{
ForceAttemptHTTP2: true,
DialContext: (&net.Dialer{Timeout: 5 * time.Second, KeepAlive: 30 * time.Second}).DialContext,
TLSClientConfig: &tls.Config{
MinVersion: tls.VersionTLS12,
},
TLSHandshakeTimeout: 5 * time.Second,
}
return &http.Client{
Transport: transport,
CheckRedirect: func(request *http.Request, via []*http.Request) error {
if len(via) >= 3 {
return fmt.Errorf("too many DoH redirects")
}
if request.URL.Scheme != "https" {
return fmt.Errorf("DoH redirect changed to a non-HTTPS scheme")
}
if !strings.EqualFold(request.URL.Hostname(), origin.Hostname()) {
return fmt.Errorf("DoH redirect changed authentication domain")
}
return nil
},
}
}
func exchangeDoHWithClient(ctx context.Context, client *http.Client, endpoint string, wire []byte, method string) ([]byte, TransportInfo, error) {
started := time.Now()
info := TransportInfo{
Protocol: "doh",
Encrypted: true,
Bootstrap: "system_resolver",
}
requestURL := endpoint
var body io.Reader
switch strings.ToLower(method) {
case "get":
parsed, err := url.Parse(endpoint)
if err != nil {
return nil, info, fmt.Errorf("parse DoH endpoint: %w", err)
}
query := parsed.Query()
query.Set("dns", base64.RawURLEncoding.EncodeToString(wire))
parsed.RawQuery = query.Encode()
requestURL = parsed.String()
case "post", "":
method = "post"
body = bytes.NewReader(wire)
default:
return nil, info, fmt.Errorf("unsupported DoH method %q", method)
}
request, err := http.NewRequestWithContext(ctx, strings.ToUpper(method), requestURL, body)
if err != nil {
return nil, info, fmt.Errorf("create DoH request: %w", err)
}
request.Header.Set("Accept", "application/dns-message")
if strings.EqualFold(method, "post") {
request.Header.Set("Content-Type", "application/dns-message")
}
request.Header.Set("User-Agent", "ednsdiag/0.1.0-dev")
response, err := client.Do(request)
info.ElapsedMS = time.Since(started).Milliseconds()
if err != nil {
return nil, info, fmt.Errorf("perform DoH exchange: %w", err)
}
defer response.Body.Close()
info.HTTPVersion = response.Proto
if response.TLS == nil || len(response.TLS.VerifiedChains) == 0 {
return nil, info, fmt.Errorf("DoH server TLS identity was not verified")
}
info.ServerAuthenticated = true
info.TLSVersion = tlsVersionName(response.TLS.Version)
info.ALPN = response.TLS.NegotiatedProtocol
if response.StatusCode < 200 || response.StatusCode > 299 {
return nil, info, fmt.Errorf("DoH server returned HTTP status %d", response.StatusCode)
}
mediaType, _, err := mime.ParseMediaType(response.Header.Get("Content-Type"))
if err != nil || !strings.EqualFold(mediaType, "application/dns-message") {
return nil, info, fmt.Errorf("DoH server returned unsupported content type %q", response.Header.Get("Content-Type"))
}
payload, err := io.ReadAll(io.LimitReader(response.Body, maxDNSMessageSize+1))
if err != nil {
return nil, info, fmt.Errorf("read DoH response: %w", err)
}
if len(payload) > maxDNSMessageSize {
return nil, info, fmt.Errorf("DoH response exceeds %d bytes", maxDNSMessageSize)
}
return payload, info, nil
}
func tlsVersionName(version uint16) string {
switch version {
case tls.VersionTLS13:
return "TLS1.3"
case tls.VersionTLS12:
return "TLS1.2"
default:
return fmt.Sprintf("0x%04x", version)
}
}
@@ -1,75 +0,0 @@
package edns
import (
"io"
"net/http"
"net/http/httptest"
"testing"
"golang.org/x/net/dns/dnsmessage"
)
func TestExchangeDoHGETAndPOST(t *testing.T) {
server := httptest.NewTLSServer(http.HandlerFunc(func(writer http.ResponseWriter, request *http.Request) {
var payload []byte
var err error
if request.Method == http.MethodGet {
payload, err = decodeGETQuery(request.URL.Query().Get("dns"))
} else {
payload, err = io.ReadAll(request.Body)
}
if err != nil {
http.Error(writer, err.Error(), http.StatusBadRequest)
return
}
if request.Header.Get("Accept") != "application/dns-message" {
http.Error(writer, "missing accept", http.StatusNotAcceptable)
return
}
var query dnsmessage.Message
if err := query.Unpack(payload); err != nil {
http.Error(writer, err.Error(), http.StatusBadRequest)
return
}
response := dnsmessage.Message{
Header: dnsmessage.Header{ID: query.Header.ID, Response: true, RecursionAvailable: true},
Questions: query.Questions,
}
responseWire, err := response.Pack()
if err != nil {
http.Error(writer, err.Error(), http.StatusInternalServerError)
return
}
writer.Header().Set("Content-Type", "application/dns-message")
_, _ = writer.Write(responseWire)
}))
defer server.Close()
wire, _, _, err := BuildQuery("example.com", "A")
if err != nil {
t.Fatalf("build query: %v", err)
}
for _, method := range []string{"get", "post"} {
t.Run(method, func(t *testing.T) {
response, info, err := exchangeDoHWithClient(t.Context(), server.Client(), server.URL, wire, method)
if err != nil {
t.Fatalf("exchange DoH: %v", err)
}
if len(response) == 0 || !info.Encrypted || !info.ServerAuthenticated {
t.Fatalf("unexpected result: response=%d info=%#v", len(response), info)
}
})
}
}
func TestExchangeDoHRejectsHTTPError(t *testing.T) {
server := httptest.NewTLSServer(http.HandlerFunc(func(writer http.ResponseWriter, _ *http.Request) {
http.Error(writer, "unavailable", http.StatusServiceUnavailable)
}))
defer server.Close()
if _, _, err := exchangeDoHWithClient(t.Context(), server.Client(), server.URL, []byte{1}, "post"); err == nil {
t.Fatal("HTTP error was accepted")
}
}
@@ -1,106 +0,0 @@
package edns
import (
"context"
"crypto/tls"
"encoding/binary"
"fmt"
"io"
"net"
"time"
)
func exchangeDoT(ctx context.Context, provider Provider, wire []byte) ([]byte, TransportInfo, error) {
return exchangeDoTWithTLSConfig(ctx, provider, wire, &tls.Config{
ServerName: provider.DoTName,
MinVersion: tls.VersionTLS12,
NextProtos: []string{"dot"},
})
}
func exchangeDoTWithTLSConfig(ctx context.Context, provider Provider, wire []byte, tlsConfig *tls.Config) ([]byte, TransportInfo, error) {
started := time.Now()
info := TransportInfo{
Protocol: "dot",
Encrypted: true,
Bootstrap: "system_resolver",
}
rawConnection, err := (&net.Dialer{}).DialContext(ctx, "tcp", provider.DoTAddr)
if err != nil {
info.ElapsedMS = time.Since(started).Milliseconds()
return nil, info, fmt.Errorf("connect to DoT server: %w", err)
}
defer rawConnection.Close()
if deadline, ok := ctx.Deadline(); ok {
if err := rawConnection.SetDeadline(deadline); err != nil {
return nil, info, fmt.Errorf("set DoT deadline: %w", err)
}
}
tlsConfig = tlsConfig.Clone()
tlsConfig.ServerName = provider.DoTName
tlsConnection := tls.Client(rawConnection, tlsConfig)
if err := tlsConnection.HandshakeContext(ctx); err != nil {
info.ElapsedMS = time.Since(started).Milliseconds()
return nil, info, fmt.Errorf("authenticate DoT server: %w", err)
}
state := tlsConnection.ConnectionState()
if len(state.VerifiedChains) == 0 {
return nil, info, fmt.Errorf("DoT server TLS identity was not verified")
}
info.ServerAuthenticated = true
info.TLSVersion = tlsVersionName(state.Version)
info.ALPN = state.NegotiatedProtocol
if info.ALPN != "" && info.ALPN != "dot" {
info.ElapsedMS = time.Since(started).Milliseconds()
return nil, info, fmt.Errorf("DoT server negotiated unexpected ALPN protocol %q", info.ALPN)
}
response, err := exchangeTCPFrame(tlsConnection, wire)
info.ElapsedMS = time.Since(started).Milliseconds()
if err != nil {
return nil, info, fmt.Errorf("perform DoT exchange: %w", err)
}
return response, info, nil
}
func exchangeTCPFrame(connection io.ReadWriter, wire []byte) ([]byte, error) {
if len(wire) == 0 || len(wire) > maxDNSMessageSize {
return nil, fmt.Errorf("invalid DNS message length %d", len(wire))
}
frame := make([]byte, 2+len(wire))
binary.BigEndian.PutUint16(frame[:2], uint16(len(wire)))
copy(frame[2:], wire)
if err := writeAll(connection, frame); err != nil {
return nil, fmt.Errorf("write framed DNS query: %w", err)
}
var lengthBytes [2]byte
if _, err := io.ReadFull(connection, lengthBytes[:]); err != nil {
return nil, fmt.Errorf("read DNS response length: %w", err)
}
length := int(binary.BigEndian.Uint16(lengthBytes[:]))
if length == 0 {
return nil, fmt.Errorf("DoT server returned an empty DNS message")
}
response := make([]byte, length)
if _, err := io.ReadFull(connection, response); err != nil {
return nil, fmt.Errorf("read DNS response: %w", err)
}
return response, nil
}
func writeAll(writer io.Writer, payload []byte) error {
for len(payload) > 0 {
written, err := writer.Write(payload)
if err != nil {
return err
}
if written == 0 {
return io.ErrShortWrite
}
payload = payload[written:]
}
return nil
}
@@ -1,254 +0,0 @@
package edns
import (
"bytes"
"context"
"crypto/ed25519"
"crypto/rand"
"crypto/tls"
"crypto/x509"
"encoding/binary"
"io"
"math/big"
"net"
"strings"
"testing"
"time"
"golang.org/x/net/dns/dnsmessage"
)
type scriptedReadWriter struct {
read *bytes.Reader
written bytes.Buffer
}
func (stream *scriptedReadWriter) Read(payload []byte) (int, error) {
return stream.read.Read(payload)
}
func (stream *scriptedReadWriter) Write(payload []byte) (int, error) {
return stream.written.Write(payload)
}
func TestExchangeTCPFrame(t *testing.T) {
responsePayload := []byte{9, 8, 7}
framedResponse := make([]byte, 2+len(responsePayload))
binary.BigEndian.PutUint16(framedResponse[:2], uint16(len(responsePayload)))
copy(framedResponse[2:], responsePayload)
stream := &scriptedReadWriter{read: bytes.NewReader(framedResponse)}
query := []byte{1, 2, 3, 4}
response, err := exchangeTCPFrame(stream, query)
if err != nil {
t.Fatalf("exchange TCP frame: %v", err)
}
if !bytes.Equal(response, responsePayload) {
t.Fatalf("response = %v, want %v", response, responsePayload)
}
written := stream.written.Bytes()
if int(binary.BigEndian.Uint16(written[:2])) != len(query) || !bytes.Equal(written[2:], query) {
t.Fatalf("invalid query frame: %v", written)
}
}
func TestExchangeDoTAuthenticatesServer(t *testing.T) {
certificate, roots := newTestCertificate(t, "resolver.test")
listener, err := tls.Listen("tcp", "127.0.0.1:0", &tls.Config{
Certificates: []tls.Certificate{certificate},
MinVersion: tls.VersionTLS12,
NextProtos: []string{"dot"},
})
if err != nil {
t.Fatalf("listen for DoT: %v", err)
}
defer listener.Close()
serverError := make(chan error, 1)
go func() {
connection, err := listener.Accept()
if err != nil {
serverError <- err
return
}
defer connection.Close()
response, err := serveOneDoTQuery(connection)
if err == nil {
err = writeAll(connection, response)
}
serverError <- err
}()
queryWire, query, transactionID, err := BuildQuery("example.com", "A")
if err != nil {
t.Fatalf("build query: %v", err)
}
ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
defer cancel()
response, info, err := exchangeDoTWithTLSConfig(ctx, Provider{DoTAddr: listener.Addr().String(), DoTName: "resolver.test"}, queryWire, &tls.Config{
RootCAs: roots,
MinVersion: tls.VersionTLS12,
NextProtos: []string{"dot"},
})
if err != nil {
t.Fatalf("exchange DoT: %v", err)
}
if err := <-serverError; err != nil {
t.Fatalf("serve DoT: %v", err)
}
if !info.ServerAuthenticated || info.ALPN != "dot" {
t.Fatalf("unexpected transport info: %#v", info)
}
if _, err := ParseResponse(response, transactionID, query); err != nil {
t.Fatalf("parse response: %v", err)
}
}
func TestExchangeDoTAllowsMissingALPN(t *testing.T) {
certificate, roots := newTestCertificate(t, "resolver.test")
listener, err := tls.Listen("tcp", "127.0.0.1:0", &tls.Config{
Certificates: []tls.Certificate{certificate},
MinVersion: tls.VersionTLS12,
})
if err != nil {
t.Fatalf("listen for DoT: %v", err)
}
defer listener.Close()
serverError := make(chan error, 1)
go func() {
connection, err := listener.Accept()
if err != nil {
serverError <- err
return
}
defer connection.Close()
response, err := serveOneDoTQuery(connection)
if err == nil {
err = writeAll(connection, response)
}
serverError <- err
}()
queryWire, _, _, err := BuildQuery("example.com", "A")
if err != nil {
t.Fatalf("build query: %v", err)
}
ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
defer cancel()
_, info, err := exchangeDoTWithTLSConfig(ctx, Provider{DoTAddr: listener.Addr().String(), DoTName: "resolver.test"}, queryWire, &tls.Config{
RootCAs: roots,
MinVersion: tls.VersionTLS12,
NextProtos: []string{"dot"},
})
if err != nil {
t.Fatalf("exchange DoT without server ALPN: %v", err)
}
if err := <-serverError; err != nil {
t.Fatalf("serve DoT: %v", err)
}
if !info.ServerAuthenticated || info.ALPN != "" {
t.Fatalf("unexpected transport info: %#v", info)
}
}
func TestExchangeDoTRejectsUnexpectedALPN(t *testing.T) {
certificate, roots := newTestCertificate(t, "resolver.test")
listener, err := tls.Listen("tcp", "127.0.0.1:0", &tls.Config{
Certificates: []tls.Certificate{certificate},
MinVersion: tls.VersionTLS12,
NextProtos: []string{"http/1.1"},
})
if err != nil {
t.Fatalf("listen for TLS: %v", err)
}
defer listener.Close()
serverError := make(chan error, 1)
go func() {
connection, err := listener.Accept()
if err != nil {
serverError <- err
return
}
defer connection.Close()
serverError <- connection.(*tls.Conn).Handshake()
}()
queryWire, _, _, err := BuildQuery("example.com", "A")
if err != nil {
t.Fatalf("build query: %v", err)
}
ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
defer cancel()
_, info, err := exchangeDoTWithTLSConfig(ctx, Provider{DoTAddr: listener.Addr().String(), DoTName: "resolver.test"}, queryWire, &tls.Config{
RootCAs: roots,
MinVersion: tls.VersionTLS12,
NextProtos: []string{"http/1.1"},
})
if err == nil || !strings.Contains(err.Error(), "unexpected ALPN protocol") {
t.Fatalf("exchange DoT error = %v, want unexpected ALPN error", err)
}
if err := <-serverError; err != nil {
t.Fatalf("complete TLS handshake: %v", err)
}
if !info.ServerAuthenticated || info.ALPN != "http/1.1" {
t.Fatalf("unexpected transport info: %#v", info)
}
}
func newTestCertificate(t *testing.T, name string) (tls.Certificate, *x509.CertPool) {
t.Helper()
publicKey, privateKey, err := ed25519.GenerateKey(rand.Reader)
if err != nil {
t.Fatalf("generate key: %v", err)
}
template := &x509.Certificate{
SerialNumber: big.NewInt(1),
DNSNames: []string{name},
NotBefore: time.Now().Add(-time.Hour),
NotAfter: time.Now().Add(time.Hour),
KeyUsage: x509.KeyUsageDigitalSignature | x509.KeyUsageCertSign,
ExtKeyUsage: []x509.ExtKeyUsage{x509.ExtKeyUsageServerAuth},
IsCA: true,
BasicConstraintsValid: true,
}
der, err := x509.CreateCertificate(rand.Reader, template, template, publicKey, privateKey)
if err != nil {
t.Fatalf("create certificate: %v", err)
}
parsed, err := x509.ParseCertificate(der)
if err != nil {
t.Fatalf("parse certificate: %v", err)
}
roots := x509.NewCertPool()
roots.AddCert(parsed)
return tls.Certificate{Certificate: [][]byte{der}, PrivateKey: privateKey}, roots
}
func serveOneDoTQuery(connection net.Conn) ([]byte, error) {
var lengthBytes [2]byte
if _, err := io.ReadFull(connection, lengthBytes[:]); err != nil {
return nil, err
}
wire := make([]byte, int(binary.BigEndian.Uint16(lengthBytes[:])))
if _, err := io.ReadFull(connection, wire); err != nil {
return nil, err
}
var query dnsmessage.Message
if err := query.Unpack(wire); err != nil {
return nil, err
}
response := dnsmessage.Message{
Header: dnsmessage.Header{ID: query.Header.ID, Response: true, RecursionAvailable: true},
Questions: query.Questions,
}
responseWire, err := response.Pack()
if err != nil {
return nil, err
}
framed := make([]byte, 2+len(responseWire))
binary.BigEndian.PutUint16(framed[:2], uint16(len(responseWire)))
copy(framed[2:], responseWire)
return framed, nil
}
@@ -1,59 +0,0 @@
package edns
type QueryOptions struct {
Name string
RecordType string
Protocol string
Provider string
Method string
EndpointURL string
}
type Result struct {
SchemaVersion int `json:"schema_version"`
Operation string `json:"operation"`
Completed bool `json:"completed"`
Query QueryInfo `json:"query"`
Resolver ResolverInfo `json:"resolver"`
Transport TransportInfo `json:"transport"`
DNS DNSInfo `json:"dns"`
Warnings []string `json:"warnings,omitempty"`
Error *ErrorInfo `json:"error,omitempty"`
}
type QueryInfo struct {
Name string `json:"name"`
Type string `json:"type"`
}
type ResolverInfo struct {
Provider string `json:"provider"`
Endpoint string `json:"endpoint"`
Profile string `json:"profile"`
}
type TransportInfo struct {
Protocol string `json:"protocol"`
Encrypted bool `json:"encrypted"`
ServerAuthenticated bool `json:"server_authenticated"`
ElapsedMS int64 `json:"elapsed_ms"`
Bootstrap string `json:"bootstrap"`
TLSVersion string `json:"tls_version,omitempty"`
ALPN string `json:"alpn,omitempty"`
HTTPVersion string `json:"http_version,omitempty"`
}
type DNSInfo struct {
RCode string `json:"rcode"`
RCodeValue int `json:"rcode_value"`
ResolverReportsDNSSECAuthenticated bool `json:"resolver_reports_dnssec_authenticated"`
ClientValidatedDNSSEC bool `json:"client_validated_dnssec"`
Answers []AnswerRecord `json:"answers"`
}
type AnswerRecord map[string]any
type ErrorInfo struct {
Class string `json:"class"`
Message string `json:"message"`
}
@@ -1,53 +0,0 @@
package edns
import (
"fmt"
"strings"
)
type Provider struct {
ID string
Profile string
DoHURL string
DoTAddr string
DoTName string
}
var providers = map[string]Provider{
"cloudflare": {
ID: "cloudflare",
Profile: "unfiltered",
DoHURL: "https://cloudflare-dns.com/dns-query",
DoTAddr: "one.one.one.one:853",
DoTName: "one.one.one.one",
},
"google": {
ID: "google",
Profile: "unfiltered",
DoHURL: "https://dns.google/dns-query",
DoTAddr: "dns.google:853",
DoTName: "dns.google",
},
"quad9": {
ID: "quad9",
Profile: "security-filtered",
DoHURL: "https://dns.quad9.net/dns-query",
DoTAddr: "dns.quad9.net:853",
DoTName: "dns.quad9.net",
},
"adguard": {
ID: "adguard",
Profile: "ad-and-security-filtered",
DoHURL: "https://dns.adguard-dns.com/dns-query",
DoTAddr: "dns.adguard-dns.com:853",
DoTName: "dns.adguard-dns.com",
},
}
func FindProvider(name string) (Provider, error) {
provider, ok := providers[strings.ToLower(name)]
if !ok {
return Provider{}, fmt.Errorf("unknown provider %q", name)
}
return provider, nil
}
@@ -1,18 +0,0 @@
package edns
import "testing"
func TestBuiltInProvidersHaveStrictEndpoints(t *testing.T) {
for _, name := range []string{"cloudflare", "google", "quad9", "adguard"} {
provider, err := FindProvider(name)
if err != nil {
t.Fatalf("find provider %s: %v", name, err)
}
if provider.DoHURL == "" || provider.DoTAddr == "" || provider.DoTName == "" {
t.Fatalf("provider %s is incomplete: %#v", name, provider)
}
}
if _, err := FindProvider("custom"); err == nil {
t.Fatal("unapproved custom provider was accepted")
}
}
@@ -1,70 +0,0 @@
package edns
import (
"context"
"fmt"
"net/url"
)
func Query(ctx context.Context, options QueryOptions) Result {
wire, query, transactionID, err := BuildQuery(options.Name, options.RecordType)
result := Result{
SchemaVersion: 1,
Operation: "query",
Query: query,
Transport: TransportInfo{
Protocol: options.Protocol,
Encrypted: true,
Bootstrap: "system_resolver",
},
DNS: DNSInfo{Answers: []AnswerRecord{}},
}
if err != nil {
result.Query = QueryInfo{Name: options.Name, Type: options.RecordType}
result.Error = &ErrorInfo{Class: "input", Message: err.Error()}
return result
}
provider, err := FindProvider(options.Provider)
if err != nil {
result.Error = &ErrorInfo{Class: "input", Message: err.Error()}
return result
}
if options.EndpointURL != "" {
if options.Protocol != "doh" {
result.Error = &ErrorInfo{Class: "input", Message: "custom --url applies only to DoH"}
return result
}
parsed, err := url.Parse(options.EndpointURL)
if err != nil || parsed.Scheme != "https" || parsed.Hostname() == "" {
result.Error = &ErrorInfo{Class: "input", Message: fmt.Sprintf("invalid DoH endpoint URL %q", options.EndpointURL)}
return result
}
provider = Provider{ID: "custom", Profile: "custom", DoHURL: options.EndpointURL}
}
result.Resolver = ResolverInfo{Provider: provider.ID, Profile: provider.Profile}
var response []byte
switch options.Protocol {
case "doh":
result.Resolver.Endpoint = provider.DoHURL
response, result.Transport, err = exchangeDoH(ctx, provider, wire, options.Method)
case "dot":
result.Resolver.Endpoint = provider.DoTAddr
response, result.Transport, err = exchangeDoT(ctx, provider, wire)
default:
err = fmt.Errorf("protocol %q is not available; run ednsdiag capabilities", options.Protocol)
}
if err != nil {
result.Error = &ErrorInfo{Class: "transport", Message: err.Error()}
return result
}
result.DNS, err = ParseResponse(response, transactionID, query)
if err != nil {
result.Error = &ErrorInfo{Class: "protocol", Message: err.Error()}
return result
}
result.Completed = true
return result
}
@@ -1,7 +0,0 @@
package edns
import "encoding/base64"
func decodeGETQuery(value string) ([]byte, error) {
return base64.RawURLEncoding.DecodeString(value)
}
@@ -1,30 +0,0 @@
# Built-in provider policy
Provider endpoints and capabilities must be verified against the provider's
official documentation before they are added or changed. The built-in entries
below were verified on 2026-08-13.
## Candidate providers
| Provider | DoH endpoint | DoT endpoint / authentication name | Official documentation | Profile |
| --- | --- | --- | --- | --- |
| Cloudflare | `https://cloudflare-dns.com/dns-query` | `one.one.one.one:853` | [DoH](https://developers.cloudflare.com/1.1.1.1/encryption/dns-over-https/make-api-requests/) / [DoT](https://developers.cloudflare.com/1.1.1.1/encryption/dns-over-tls/) | Unfiltered |
| Google | `https://dns.google/dns-query` | `dns.google:853` | [DoH](https://developers.google.com/speed/public-dns/docs/doh) / [DoT](https://developers.google.com/speed/public-dns/docs/dns-over-tls) | Unfiltered |
| Quad9 | `https://dns.quad9.net/dns-query` | `dns.quad9.net:853` | [Quad9 services](https://docs.quad9.net/services/) | Security filtered; HTTP/2 required |
| AdGuard | `https://dns.adguard-dns.com/dns-query` | `dns.adguard-dns.com:853` | [AdGuard providers](https://adguard-dns.io/kb/general/dns-providers/) | Ads, tracking, and security filtered |
## Registry requirements
Each built-in provider entry must include:
- stable provider identifier;
- protocol and endpoint;
- authentication domain name;
- bootstrap addresses only when officially published;
- filtering/ECS profile;
- official source URL;
- last verification date.
Do not infer one protocol endpoint from another. Do not treat filtering and
non-filtering services as interchangeable. Provider comparison results must
remain separate.
@@ -1,58 +0,0 @@
# Security and privacy requirements
Read this file before changing transports, bootstrap behavior, endpoint
validation, fallback, or result claims.
## Non-negotiable rules
1. Never silently downgrade to plaintext DNS.
2. Validate certificates and authentication domain names. DoT follows the
strict privacy profile in [RFC 8310](https://www.rfc-editor.org/rfc/rfc8310.html).
3. Treat certificate, hostname, SNI, and negotiated ALPN mismatches as hard
failures, not fallback opportunities.
4. Bound response sizes, per-attempt timeouts, total time, redirects, and the
number of attempts.
5. Do not expose an unrestricted endpoint parameter to an Agent. Built-in
providers are allowlisted; private or custom endpoints require explicit
user intent and policy approval.
6. Do not connect to addresses returned in DNS answers.
7. Do not enable AXFR, IXFR, or ANY queries.
8. Do not persist full query names or client identifiers by default.
## DoT ALPN policy
The client advertises the `dot` ALPN identifier. RFC 7858 and RFC 8310 do not
require a DoT server on its dedicated port to select an ALPN protocol, so an
empty negotiated ALPN is permitted and reported as empty. If a server selects
a non-empty protocol other than `dot`, abort before sending the DNS query.
## Bootstrap transparency
Connecting to a resolver hostname may require an initial DNS lookup. Report
whether the endpoint was reached using a configured bootstrap address, the
system resolver, or an already-known IP. Do not claim that a query avoided the
system resolver when bootstrap used it.
## DNS status and fallback
An HTTP, TLS, or QUIC exchange can succeed while DNS returns `NXDOMAIN`,
`SERVFAIL`, or `REFUSED`. Those are DNS outcomes and must not be converted into
transport errors. Cross-provider or cross-protocol fallback is permitted only
for explicitly classified transport failures and must be disclosed.
## DNSSEC language
The AD bit means the selected recursive resolver reports authenticated data.
It is not proof that this client validated the DNSSEC chain. Use separate
fields for resolver-reported and locally validated DNSSEC state.
## Privacy language
Encrypted transport protects the path between the client and the selected
resolver. The resolver can still observe the query. Provider policy, logging,
filtering, ECS behavior, and jurisdiction remain relevant. See
[RFC 8932](https://www.rfc-editor.org/rfc/rfc8932.html).
ODoH and Anonymized DNSCrypt add relay models but do not justify claims of
absolute anonymity. Their proxy, relay, and target roles must be reported
separately.
@@ -1,40 +0,0 @@
# Standards and authoritative sources
Verified on 2026-08-13. Protocol behavior must be based on the published
standard, not on summaries or provider-specific JSON APIs.
| Capability | Authority | Project scope |
| --- | --- | --- |
| Agent Skills package | [Agent Skills specification](https://agentskills.io/specification) | Required package format |
| OMP discovery | [OMP Skills documentation](https://github.com/can1357/oh-my-pi/blob/main/docs/skills.md) | Supported host |
| DoH | [RFC 8484](https://www.rfc-editor.org/rfc/rfc8484.html) | Planned |
| DoT | [RFC 7858](https://www.rfc-editor.org/rfc/rfc7858.html) | Planned |
| DoT authentication profiles | [RFC 8310](https://www.rfc-editor.org/rfc/rfc8310.html) | Strict privacy only |
| DoQ | [RFC 9250](https://www.rfc-editor.org/rfc/rfc9250.html) | Planned |
| ODoH | [RFC 9230](https://www.rfc-editor.org/rfc/rfc9230.html) | Research until a maintained implementation is selected |
| DNS privacy operations | [RFC 8932](https://www.rfc-editor.org/rfc/rfc8932.html) | Security and privacy guidance |
| EDNS(0) padding | [RFC 7830](https://www.rfc-editor.org/rfc/rfc7830.html) and [RFC 8467](https://www.rfc-editor.org/rfc/rfc8467.html) | Evaluate per transport |
| DNSCrypt | [DNSCrypt protocol specification](https://github.com/DNSCrypt/dnscrypt-protocol) | Planned, non-IETF |
| Anonymized DNSCrypt | [Anonymized DNSCrypt specification](https://github.com/DNSCrypt/dnscrypt-protocol/blob/master/ANONYMIZED-DNSCRYPT.txt) | Research |
| Go DNS wire and IDNA support | [Go x/net module](https://pkg.go.dev/golang.org/x/net) | Pinned to v0.58.0; use `dnsmessage` and `idna` |
## Deliberate exclusions
- DNS-over-DTLS ([RFC 8094](https://www.rfc-editor.org/rfc/rfc8094.html))
is experimental and is not a target transport.
- DNS zone transfer over TLS
([RFC 9103](https://www.rfc-editor.org/rfc/rfc9103.html)) is outside the
client-to-recursive diagnostic scope.
- Recursive-to-authoritative encryption and resolver/server operation are
outside the initial scope.
## Terminology
DoH3 means RFC 8484 semantics carried over HTTP/3. It is not a separate DNS
message format. DNSCrypt is an encrypted DNS protocol with its own
specification; do not label it as an IETF RFC.
For DoH, accept and send `application/dns-message`. Keep HTTP status separate
from the DNS RCODE: a valid NXDOMAIN or SERVFAIL response still uses HTTP 2xx.
For DoT, use the strict privacy profile and verify both the PKIX chain and the
configured authentication domain name.
@@ -1,82 +0,0 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://github.com/windyboy/encrypted-dns-skill/schemas/result-v1.schema.json",
"title": "Encrypted DNS diagnostic result",
"type": "object",
"required": ["schema_version", "operation", "completed", "query", "transport", "dns"],
"properties": {
"schema_version": { "const": 1 },
"operation": { "enum": ["query", "probe", "compare"] },
"completed": { "type": "boolean" },
"query": {
"type": "object",
"required": ["name", "type"],
"properties": {
"name": { "type": "string" },
"type": { "type": "string" }
},
"additionalProperties": false
},
"resolver": {
"type": "object",
"properties": {
"provider": { "type": "string" },
"endpoint": { "type": "string" },
"profile": { "type": "string" }
},
"additionalProperties": false
},
"transport": {
"type": "object",
"required": ["protocol", "encrypted", "server_authenticated"],
"properties": {
"protocol": { "enum": ["doh", "dot", "doq", "doh3", "dnscrypt", "odoh", "anonymized-dnscrypt"] },
"encrypted": { "type": "boolean" },
"server_authenticated": { "type": "boolean" },
"elapsed_ms": { "type": "integer", "minimum": 0 },
"bootstrap": { "type": "string" },
"tls_version": { "type": "string" },
"alpn": { "type": "string" },
"http_version": { "type": "string" }
},
"additionalProperties": true
},
"dns": {
"type": "object",
"required": ["rcode", "rcode_value", "answers"],
"properties": {
"rcode": { "type": "string" },
"rcode_value": { "type": "integer", "minimum": 0 },
"resolver_reports_dnssec_authenticated": { "type": "boolean" },
"client_validated_dnssec": { "type": "boolean" },
"answers": {
"type": "array",
"items": {
"type": "object",
"required": ["name", "type", "ttl"],
"properties": {
"name": { "type": "string" },
"type": { "type": "string" },
"ttl": { "type": "integer", "minimum": 0 }
},
"additionalProperties": true
}
}
},
"additionalProperties": true
},
"fallback_used": { "type": "boolean" },
"attempts": { "type": "array", "items": { "type": "object" } },
"warnings": { "type": "array", "items": { "type": "string" } },
"error": {
"type": "object",
"required": ["class", "message"],
"properties": {
"class": { "enum": ["input", "transport", "protocol", "internal"] },
"message": { "type": "string" }
},
"additionalProperties": false
}
},
"additionalProperties": true
}
+5
View File
@@ -19,6 +19,8 @@ facts/
.agents/
.claude/
.omp/
.opencode/
.zcode/
.mcp.json
WATCHDOG.yml
skills-lock.json
@@ -28,3 +30,6 @@ skills-lock.json
.vscode/
.idea/
*~
# Agent working scratch (not repo content).
.agent-work/
.tmp-*
+8
View File
@@ -0,0 +1,8 @@
repos:
- repo: local
hooks:
- id: validate-repo
name: validate repository
entry: scripts/validate-repo.sh
language: system
pass_filenames: false
+72 -25
View File
@@ -2,37 +2,58 @@
This repo is the **agent ops handbook + fact source** for maintaining personal VPS hosts. Prefer verifying live state over assuming docs are complete.
Also readable as `agent.md` (symlink → this file).
## How to work
1. Read [`inventory/hosts.md`](inventory/hosts.md) for the machine list.
2. Open the matching [`hosts/<name>.md`](hosts/) for SSH, roles, paths, and quirks.
3. For common tasks, follow a runbook under [`runbooks/`](runbooks/).
3. For common tasks, follow a runbook under [`runbooks/`](runbooks/). Pick the
most specific applicable one from [`runbooks/README.md`](runbooks/README.md);
the spec is [`RUNBOOKS.md`](RUNBOOKS.md) and new runbooks start from
[`runbooks/_template.md`](runbooks/_template.md).
4. Prefer read-only checks first; change only after confirming current state.
5. For routine checks and approved service reconciliation, run the matching
Ansible playbook from `ansible/`; see [routine Ansible operations](runbooks/ansible-operations.md).
6. Default SSH access (`ssh -4 windy@<host>`) is for focused diagnostics,
imperative upstream procedures, and incident work. Prefer **IPv4** from this
WSL client (AAAA often exists but IPv6 route does not).
> **Agent sandbox SSH quirk (verified 2026-08-20):** the agent shell runs in
> a sandboxed user namespace — system files such as
> `/etc/ssh/ssh_config.d/20-systemd-ssh-proxy.conf` appear owned by `nobody`,
> so plain `ssh` aborts with `Bad owner or permissions on ...`. Always use
> `ssh -F /dev/null` from the agent shell and pass options explicitly
> (`~/.ssh/config` is skipped; e.g. `ssh -F /dev/null -p 2222
> -i ~/.ssh/id_ed25519 windy@repo.windy.me`). `sudo` never works in the
> sandbox (`NoNewPrivs`, no capabilities, `/` read-only). The host itself is
> healthy — to inspect or act on the real host from the sandbox use
> `/mnt/c/WINDOWS/system32/wsl.exe -u root -- <cmd>` (real root: keep
> read-only unless a change is approved).
7. Record each material VPS operation, incident, configuration change, or
verification outcome in the corresponding **Linear `vps` project**. Include
scope, action, verification, and remaining follow-up; never put passwords,
tokens, private keys, recovery keys, or private room IDs in Linear.
verification outcome in the corresponding **Plane `vps` project**
(self-hosted `plane.chans.xyz`, Plane MCP `mcp__plane__*`, following the
`plane-workflow` skill). **Linear is retired as a record source (2026-09-03)
— do not create Linear issues;** existing W1N-* entries are read-only
history. Include scope, action, verification, and remaining follow-up; never
put passwords, tokens, private keys, recovery keys, or private room IDs in
Plane or Linear.
## Active hosts (quick map)
### Runbook execution rules
| Host | Role | SSH | Facts |
|------|------|-----|--------|
| **mx2.windy.me** | mailcow (`/opt/mail`, project `cow`) | `ssh -4 windy@mx2.windy.me` | [hosts/mx2.windy.me.md](hosts/mx2.windy.me.md) |
| **us2.wsvc.info** | Vaultwarden + Traefik (+ Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | [hosts/us2.wsvc.info.md](hosts/us2.wsvc.info.md) |
| **hk2.chans.xyz** | PowerDNS auth ns1 (`/opt/pdns`) | `ssh -4 windy@hk2.chans.xyz` | [hosts/hk2.chans.xyz.md](hosts/hk2.chans.xyz.md) |
| **synapse.chans.xyz** | Matrix ESS (Synapse + MAS + Element) on K3s | `ssh -4 windy@synapse.chans.xyz` | [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy | `ssh -4 windy@192.168.66.36` | [hosts/dns.windy.lan.md](hosts/dns.windy.lan.md) |
| **gfw.windy.lan** | OpenWrt LAN gateway / OpenClash | `ssh -4 root@192.168.66.1` | [hosts/gfw.windy.lan.md](hosts/gfw.windy.lan.md) |
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | [hosts/gw.md](hosts/gw.md) |
| **ubnt** | UniFi Network Controller | `ssh -4 windy@192.168.66.46` | [hosts/ubnt.md](hosts/ubnt.md) |
| **hass.windy.lan** | Home Assistant (HAOS, LAN55) | `ssh hassio@hass.windy.lan` | [hosts/hass.windy.lan.md](hosts/hass.windy.lan.md) |
Before operational work: inspect `runbooks/`, select the most specific
applicable runbook, follow its steps in order, do not skip verification steps,
and respect its STOP and approval conditions. If no runbook applies, diagnose
only — do not mutate production state. When live state conflicts with a
runbook's assumptions, `STOP` and report; never invent missing parameters or
bypass failed checks. The spec is [`RUNBOOKS.md`](RUNBOOKS.md).
## Active hosts
The canonical machine list (roles, SSH endpoints, Ansible coverage, status) is
[`inventory/hosts.md`](inventory/hosts.md) — the single human-readable source
of truth. Per-host facts live in [`hosts/`](hosts/). The Ansible execution
inventory is [`ansible/inventory/hosts.yml`](ansible/inventory/hosts.yml). Do
not maintain a second copy of the machine table here.
### Public services
@@ -42,7 +63,7 @@ Also readable as `agent.md` (symlink → this file).
| SMTP `mx2.windy.me:587` (STARTTLS) or `:465` | mx2 | client submission; full email + mailbox password — [runbook](runbooks/mailcow-smtp-client.md) |
| IMAP `mx2.windy.me:993` | mx2 | same mailbox credentials |
| https://auth.wsvc.info | us2 (`/opt/vaultwarden`) | Vaultwarden (Postgres, **operational**) — client Server URL |
| `repo.windy.me:2222` | us2 (`/opt/soft-serve`) | Soft Serve (stub details) |
| `repo.windy.me` (git SSH `:2222` / web HTTPS) | us2 (`/opt/gitea`) | Gitea — 1.27.3-rootless pinned, backup sidecar; details in [hosts/us2.wsvc.info.md](hosts/us2.wsvc.info.md) |
| DNS `ns1.wsvc.info:53` | hk2 (`/opt/pdns`, Auth **5.0.6**) | PowerDNS auth — zones `windy.me`, `wsvc.info`, `chans.xyz` |
| https://pdns.wsvc.info | hk2 (`poweradmin`) | Poweradmin UI |
| https://pgweb.wsvc.info | hk2 (`pgweb`) | PowerDNS Postgres browser |
@@ -50,6 +71,7 @@ Also readable as `agent.md` (symlink → this file).
| https://synapse.chans.xyz | synapse | Synapse Client-Server + Federation API |
| https://account.chans.xyz | synapse | Matrix Authentication Service (local passwords) |
| https://admin.chans.xyz | synapse | Element Admin console (MAS admin auth) |
| https://plane.chans.xyz | synapse (`plane`, Helm `plane-ce` 1.8.0 / v1.4.1) | Plane project management (self-hosted, K3s) |
### Upstream docs
@@ -59,6 +81,8 @@ Also readable as `agent.md` (symlink → this file).
**Matrix (ESS on synapse):** Matrix homeserver running on `synapse.chans.xyz` via the official ESS (Element Server Suite) Helm chart with Synapse + MAS + Element Web + Admin. DNS zone `chans.xyz` managed by hk2 PowerDNS. Before changing config, read [docs/matrix-upstream.md](docs/matrix-upstream.md) and [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md). K3s cluster on this node has hostPort 80/443 for Traefik (no ServiceLB). Health: [matrix-health](runbooks/matrix-health.md).
**Plane (on synapse):** Self-hosted Plane project management at `plane.chans.xyz`, Helm release `plane-app` (chart `plane-ce-1.8.0`, app `v1.4.1`) in ns `plane` on the same K3s node as Matrix. Config from `/home/windy/plane-k3s/values.yaml`; workload/cert/ingress details in [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md). Its Postgres/MinIO PVCs are **not** backed up.
**RustDesk:** Self-hosted RustDesk server on `hk2.chans.xyz` (`/opt/rustdesk`, containers `hbbs`/`hbbr`, image pinned `1.1.14`). The `hbbs -r` relay hostname must resolve to the host's public IP `154.36.174.161` — use `hk2.chans.xyz` (never `hk2.wsvc.info`, which has no DNS record). Health: [rustdesk-health](runbooks/rustdesk-health.md).
## Runbooks & scripts
@@ -74,14 +98,26 @@ Also readable as `agent.md` (symlink → this file).
| PowerDNS health (hk2) | [runbooks/pdns-health.md](runbooks/pdns-health.md) |
| PowerDNS upstream refs | [docs/pdns-upstream.md](docs/pdns-upstream.md) |
| Matrix health | [runbooks/matrix-health.md](runbooks/matrix-health.md) |
| Plane health | [runbooks/plane-health.md](runbooks/plane-health.md) |
| RustDesk health (hk2) | [runbooks/rustdesk-health.md](runbooks/rustdesk-health.md) |
| AdGuard Home health | [runbooks/adguard-home-health.md](runbooks/adguard-home-health.md) |
| Host disk cleanup | [runbooks/host-disk-cleanup.md](runbooks/host-disk-cleanup.md) |
| Home Assistant maintenance | [runbooks/home-assistant-maintenance.md](runbooks/home-assistant-maintenance.md) + [scripts/ha-maintenance.sh](runbooks/scripts/ha-maintenance.sh) |
| matrix_e2ee update (hass.windy.lan) | [runbooks/matrix-e2ee-update.md](runbooks/matrix-e2ee-update.md) |
| Matrix upstream refs | [docs/matrix-upstream.md](docs/matrix-upstream.md) |
| Hermes Agent Matrix channel | [docs/hermes-matrix.md](docs/hermes-matrix.md) |
| UniFi local-service proxy bypass | [docs/unifi-openclash-localhost.md](docs/unifi-openclash-localhost.md) |
| UniFi SSO login setting (Ansible) | `cd ansible && ansible-playbook playbooks/unifi-sso.yml --limit unifi` |
| Routine Ansible operations | [runbooks/ansible-operations.md](runbooks/ansible-operations.md) |
| Routine make commands | `make help` (wraps `ansible-operations.md` read-only + gated flows) |
| Issue → mergeable change | [runbooks/issue-to-merge.md](runbooks/issue-to-merge.md) |
| Fix failing health/playbook run | [runbooks/fix-ci.md](runbooks/fix-ci.md) |
| Release a reviewed change | [runbooks/release.md](runbooks/release.md) |
| Roll back a change | [runbooks/rollback.md](runbooks/rollback.md) |
| Controlled network change | [runbooks/network-change.md](runbooks/network-change.md) |
| Network outage recovery | [runbooks/network-recovery.md](runbooks/network-recovery.md) |
Full index: [runbooks/README.md](runbooks/README.md). Spec: [RUNBOOKS.md](RUNBOOKS.md).
Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.yml --limit mailcow`. The local stub resolver is flaky; DNS probes use `1.1.1.1` / `8.8.8.8`.
@@ -89,7 +125,12 @@ Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.
### Issue tracker
Issues are tracked in Linear and created/updated via the Linear MCP (`vps` project). See `docs/agents/issue-tracker.md`.
Issues are tracked in **Plane** — self-hosted at `plane.chans.xyz`, project
`vps` — and created/updated via the Plane MCP (`mcp__plane__*`), following the
`plane-workflow` skill. **Linear is retired as a record source (2026-09-03); do
not create Linear issues.** Existing W1N-* entries are read-only history.
`docs/agents/issue-tracker.md` documents the retired Linear workflow and is
stale; treat this section as authoritative.
### Triage labels
@@ -97,7 +138,8 @@ Default triage labels: needs-triage, needs-info, ready-for-agent, ready-for-huma
### Domain docs
Single-context layout: `CONTEXT.md` + `docs/adr/` at the repo root. See `docs/agents/domain.md`.
Domain-documentation conventions, including lazily created `CONTEXT.md` and
`docs/adr/` entries when needed, are described in [`docs/agents/domain.md`](docs/agents/domain.md).
## Safety
@@ -130,9 +172,14 @@ Bills, rough notes, and personal clutter stay in the Obsidian vault. This repo h
## Layout
```
AGENTS.md / agent.md # this entry (agent.md → AGENTS.md)
inventory/hosts.md # machine index
AGENTS.md # this entry
RUNBOOKS.md # runbook spec (six-field model, naming, review rules)
inventory/hosts.md # machine index (human-readable source of truth)
ansible/ # playbooks, roles, sanitized control-plane inventory
compose/ # repo-owned non-secret Compose sources (+ .env.example)
hosts/ # per-host facts
runbooks/ # step-by-step ops
docs/ # upstream doc indexes / design notes
runbooks/ # step-by-step ops (README.md = index, _template.md = template)
docs/ # upstream refs / design notes / research records (active + archive/)
scripts/validate-repo.sh # repo-wide validation (run before merging)
Makefile # routine validate / health / gated ansible wrappers
```
+183
View File
@@ -0,0 +1,183 @@
# VPS ops hub — routine validate / health / gated Ansible wrappers.
# See runbooks/ansible-operations.md for playbook semantics.
SHELL := /usr/bin/env bash
.SHELLFLAGS := -eu -o pipefail -c
.DEFAULT_GOAL := help
REPO_ROOT := $(CURDIR)
ANSIBLE_DIR := $(REPO_ROOT)/ansible
export ANSIBLE_LOCAL_TEMP := $(REPO_ROOT)/.ansible/tmp
export ANSIBLE_HOME := $(REPO_ROOT)/.ansible
LIMIT ?=
EXTRA ?=
VERBOSE ?= 0
CONFIRM ?= 0
TARGETS ?=
TRAEFIK ?= 0
LIMIT_FLAG := $(if $(LIMIT),--limit $(LIMIT),)
VERBOSE_FLAG := $(if $(filter 1,$(VERBOSE)),-v,$(if $(filter 2,$(VERBOSE)),-vvv,))
.PHONY: help validate check deps galaxy syntax ansible-prep \
ping inventory audit health health-mailcow health-matrix \
maint-preview baseline compose-check \
install-healthchecks install-matrix-healthchecks compose-deploy reconcile
help:
@printf '%s\n' \
'VPS ops hub — make targets (run from repo root)' \
'' \
'Variables: LIMIT=<group|host> CONFIRM=1 TARGETS=<svc[,svc]> TRAEFIK=1 VERBOSE=0|1|2 EXTRA=...' \
'' \
'Local / repo:' \
' validate, check scripts/validate-repo.sh (pre-merge gate)' \
' deps, galaxy ansible-galaxy collection install' \
' syntax ansible-playbook --syntax-check all playbooks' \
'' \
'Read-only remote (ansible):' \
' ping ansible managed -m ping' \
' inventory ansible-inventory --graph' \
' audit playbooks/audit.yml' \
' health [LIMIT=…] playbooks/health-report.yml' \
' health-mailcow health --limit mailcow' \
' health-matrix health --limit matrix' \
' maint-preview playbooks/maintenance-preview.yml' \
' baseline playbooks/baseline.yml' \
' compose-check compose-deploy --check --diff (requires LIMIT=)' \
'' \
'Mutating (require CONFIRM=1; host-scoped targets require LIMIT=):' \
' install-healthchecks playbooks/healthchecks.yml' \
' install-matrix-healthchecks playbooks/matrix-healthchecks.yml' \
' compose-deploy playbooks/compose-deploy.yml' \
' reconcile playbooks/compose-reconcile.yml (requires TARGETS=)' \
'' \
'Examples:' \
' make validate' \
' make health LIMIT=mailcow' \
' make compose-check LIMIT=vaultwarden' \
' make compose-deploy LIMIT=vaultwarden CONFIRM=1' \
' make reconcile LIMIT=powerdns TARGETS=auth CONFIRM=1' \
' make reconcile LIMIT=vaultwarden TARGETS=vaultwarden TRAEFIK=1 CONFIRM=1' \
'' \
'Advanced (not wrapped — use ansible-playbook directly):' \
' us4-firewalld, unifi-sso, k3s-server, matrix-stack, wireguard-harden,' \
' restic, rustdesk, email-alerts, mailcow update runbook'
validate check:
@bash "$(REPO_ROOT)/scripts/validate-repo.sh"
deps galaxy: ansible-prep
@command -v ansible-galaxy >/dev/null 2>&1 || { echo "ansible-galaxy not found; install Ansible first." >&2; exit 1; }
@cd "$(ANSIBLE_DIR)" && ansible-galaxy collection install -r requirements.yml
syntax: ansible-prep
@if ! command -v ansible-playbook >/dev/null 2>&1; then \
echo "ansible-playbook not found; syntax check skipped." >&2; \
exit 0; \
fi
@fail=0; \
for p in "$(ANSIBLE_DIR)"/playbooks/*.yml; do \
if ! (cd "$(ANSIBLE_DIR)" && ansible-playbook --syntax-check "playbooks/$$(basename "$$p")" >/dev/null 2>&1); then \
echo "syntax-check failed: $$p" >&2; \
fail=1; \
fi; \
done; \
exit $$fail
ansible-prep:
@mkdir -p "$(ANSIBLE_HOME)/tmp" "$(ANSIBLE_HOME)/ssh-control"
define require_ansible
@command -v ansible-playbook >/dev/null 2>&1 || { echo "ansible-playbook not found; install Ansible first." >&2; exit 1; }
endef
define require_limit
@if [ -z "$(LIMIT)" ]; then \
echo "LIMIT is required (e.g. LIMIT=mailcow, LIMIT=vaultwarden, LIMIT=powerdns)." >&2; \
exit 1; \
fi
endef
define require_confirm
@if [ "$(CONFIRM)" != "1" ]; then \
echo "Mutating operation blocked. Re-run with CONFIRM=1" >&2; \
exit 1; \
fi
endef
define require_targets
@if [ -z "$(TARGETS)" ]; then \
echo "TARGETS is required (comma-separated service names, e.g. TARGETS=auth or TARGETS=vaultwarden)." >&2; \
exit 1; \
fi
endef
ping: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible managed -m ping $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
inventory: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-inventory --graph $(EXTRA)
audit: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/audit.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
health: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
health-mailcow: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml --limit mailcow $(VERBOSE_FLAG) $(EXTRA)
health-matrix: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml --limit matrix $(VERBOSE_FLAG) $(EXTRA)
maint-preview: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/maintenance-preview.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
baseline: ansible-prep
$(require_ansible)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/baseline.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
compose-check: ansible-prep
$(require_ansible)
$(require_limit)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-deploy.yml --check --diff --limit $(LIMIT) $(VERBOSE_FLAG) $(EXTRA)
install-healthchecks: ansible-prep
$(require_ansible)
$(require_confirm)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/healthchecks.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
install-matrix-healthchecks: ansible-prep
$(require_ansible)
$(require_confirm)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/matrix-healthchecks.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
compose-deploy: ansible-prep
$(require_ansible)
$(require_limit)
$(require_confirm)
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-deploy.yml --limit $(LIMIT) \
-e '{"compose_deploy_confirm": true}' $(VERBOSE_FLAG) $(EXTRA)
reconcile: ansible-prep
$(require_ansible)
$(require_limit)
$(require_targets)
$(require_confirm)
@json=$$(python3 -c 'import json,sys; t=[x.strip() for x in sys.argv[1].split(",") if x.strip()]; \
(not t) and sys.exit("TARGETS must contain at least one non-empty service name"); \
d={"service_reconcile_confirm": True, "service_reconcile_targets": t}; \
(sys.argv[2]=="1") and d.update({"service_reconcile_restart_traefik": True}); \
print(json.dumps(d))' "$(TARGETS)" "$(TRAEFIK)"); \
cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-reconcile.yml --limit $(LIMIT) \
-e "$$json" $(VERBOSE_FLAG) $(EXTRA)
+89
View File
@@ -0,0 +1,89 @@
# RUNBOOKS — 仓库级规范
本文件统一所有 Runbook 的字段、命名、评审与变更规则。上游参考:[docs/archive/agent-runbook-guide.md](docs/archive/agent-runbook-guide.md)。
## 目录结构
```text
runbooks/
├── README.md # 意图 → 文件 路由索引(本目录的入口)
├── _template.md # 新建 runbook 的标准模板(复制后填写)
├── <intent>.md # 每份 runbook 只描述一种可识别的操作意图
└── ...
```
## 最小字段模型
每份 runbook 必须显式包含以下控制信息,否则盲目执行或错误恢复的风险会升高:
| 字段 | 作用 | 写作要求 |
|---|---|---|
| **Action** | 定义当前要执行的动作 | 可观察、可执行的动词;避免“检查一下”“适当调整” |
| **Expected** | 描述正常状态或预期输出 | 具体信号、阈值、状态码、测试结果或页面表现 |
| **Decision** | 定义分支与下一跳 | “条件 → 下一步”;无法判断时指向 `STOP` |
| **Verification** | 确认变更真正生效 | 每个有副作用的步骤后执行,不可跳过 |
| **Stop condition** | 规定何时不得继续 | 列出信息缺失、状态冲突、权限不足、验证失败等 |
| **Rollback** | 如何恢复到变更前状态 | 触发条件、前提、撤销步骤、回滚后验证 |
> 只读类 runbook 不产生副作用,可省略 Rollback;但必须保留 Stop condition(状态与预期冲突即 `STOP` 并记录证据)。
**只读类变体(read-only variant**:只读 runbookhealth 类、参考类)不强制
六字段模型,但必须包含以下最小结构,否则不视为达标:
- `## Purpose`12 行)+ `## Scope`(适用/不适用)
- `## Safety` 或等效章节,其中**必须**含显式 Stop condition(状态与预期冲突即
`STOP` 并记录证据;不得在执行中自行"顺手修复")
- 只读健康类另含可观察的 `## Pass criteria`(或等效的 Expected 信号)
- 每份 runbook 顶部/元信息区必须标注 `Last reviewed: <YYYY-MM-DD>`
## 命名与拆分规则
- 文件名采用小写连字符,反映**操作意图**而非目标主机,例如 `mailcow-health.md``release.md`
- 一份文件只描述一种意图。流程出现明显分叉时拆分为独立文件,不堆叠“万能流程”。
- 只读诊断与变更操作应分离:health 类 runbook 保持只读,变更走 `ansible-operations.md``release.md``rollback.md` 或对应 gated playbook。
## 章节约定
- 每份 runbook 顶部含 `## Purpose`12 行)与 `## Scope`(适用/不适用情形)。
- 变更型 runbook 必须记录明确的审批门:门控命令式使用 `## Approval gates` 表;
流程式在步骤中记录审批动作、证据位置和未批准时的 `STOP`。破坏性/不可逆操作必须获得明确批准。
- 语言约定:**runbook 正文统一使用英文**(由 agent 逐字执行,降低二义性);
元规范文件(AGENTS.md / RUNBOOKS.md / 模板注释)可保留中文。
- 变更型 runbook 的两种形态:
- **流程式(Procedure 型)**:使用六字段模型,适用多分支/多步骤变更
(现有:`fix-ci.md``issue-to-merge.md``network-change.md`
`network-recovery.md``release.md``rollback.md`)。
- **门控命令式(gated command reference**:已稳定、低歧义、可验证的
操作以命令集 + 门控呈现(现有:`mailcow-update.md`
`ansible-operations.md``home-assistant-maintenance.md`
`matrix-e2ee-update.md``vaultwarden-sqlite-to-postgres.md`),必须含 Approval gates 或确认变量
要求 + 显式 STOP,不替代流程式形态。新写的变更 runbook 默认用流程式。
- 统一在 `## Safety` 或正文中复用以下通用安全规则(更严格要求优先)。
```markdown
## Safety Rules
- Never delete an existing configuration as the first recovery action.
- Prefer read-only diagnosis before mutation.
- After every mutation, verify the expected state.
- If actual state conflicts with this runbook, STOP.
- Do not invent missing parameters.
- Do not bypass failed tests.
- Destructive actions require explicit approval.
```
## 评审与变更规则
- 新建/修改 runbook 与代码同仓评审,随系统演进更新。
- 每份 runbook 标注 `Last reviewed`;流程执行过程中发现的偏差记入对应的 Linear `vps` 项目 issue。
- 破坏性流程(迁移、删除、DNS 变更、网络变更)保持人工审批,不自动下沉。
## 成熟路径
1. **人工处理** → 现场处置与复盘,记录证据。
2. **Markdown runbook** → 固化步骤与证据要求,Agent 可辅助诊断。
3. **Agent + runbook** → 严格按流程执行,受 Stop/Approval 约束。
4. **Script / Ansible / Skill** → 把已稳定、低歧义、可验证的操作程序化(本仓库的执行层是 Ansible playbook)。
5. **人工审批 + 自动执行** → 审批门控下的自动变更(如 gated playbook + 确认变量)。
原则:先证据后变更,先小范围后扩大,先验证后结束,不确定则停止。
+5
View File
@@ -11,3 +11,8 @@ host_key_checking = True
become = True
become_method = sudo
become_ask_pass = False
[ssh_connection]
# Keep SSH control sockets inside the repo (gitignored .ansible/) so playbook
# runs work in sandboxed/CI environments without touching ~/.ansible.
ssh_args = -C -o ControlMaster=auto -o ControlPersist=60s -o ControlPath=.ansible/ssh-control/%h-%p-%r
+11
View File
@@ -15,6 +15,7 @@ all:
mx2:
ansible_host: mx2.windy.me
ansible_host_ipv4: 194.163.160.244
display_name: mx2.windy.me
service_role: mailcow
compose_project_dir: /opt/mail
healthcheck_profiles: [mailcow]
@@ -24,8 +25,11 @@ all:
us2:
ansible_host: us2.wsvc.info
ansible_host_ipv4: 193.9.44.165
display_name: us2.wsvc.info
service_role: vaultwarden
compose_project_dir: /opt/vaultwarden
compose_repo_project: vaultwarden
compose_remote_file: docker-compose.yml
healthcheck_profiles: [vaultwarden]
restic_backup_profile: vaultwarden
service_reconcile_services:
@@ -35,8 +39,11 @@ all:
hk2:
ansible_host: hk2.chans.xyz
ansible_host_ipv4: 154.36.174.161
display_name: hk2.chans.xyz
service_role: powerdns
compose_project_dir: /opt/pdns
compose_repo_project: pdns
compose_remote_file: compose.yml
healthcheck_profiles: [pdns, rustdesk, hk2aux]
restic_backup_profile: pdns
service_reconcile_services:
@@ -54,6 +61,7 @@ all:
us4:
ansible_host: us4.wsvc.info
ansible_host_ipv4: 185.201.226.122
display_name: us4.wsvc.info
service_role: wireguard
compose_project_dir: /opt/wireguard
healthcheck_profiles: [wireguard]
@@ -65,6 +73,7 @@ all:
dns_windy_lan:
ansible_host: 192.168.66.36
ansible_host_ipv4: 192.168.66.36
display_name: dns.windy.lan
service_role: adguardhome
compose_project_dir: /opt/adguardhome
healthcheck_profiles: [adguardhome]
@@ -94,6 +103,7 @@ all:
ubnt:
ansible_host: 192.168.66.46
ansible_host_ipv4: 192.168.66.46
display_name: ubnt
vars:
service_role: unifi
compose_project_dir: /home/windy/unifi-9
@@ -114,6 +124,7 @@ all:
matrix_vps:
ansible_host: 169.58.86.13
ansible_host_ipv4: 169.58.86.13
display_name: synapse.chans.xyz
service_role: matrix_k3s
matrix_server_name: chans.xyz
matrix_synapse_host: synapse.chans.xyz
+13
View File
@@ -0,0 +1,13 @@
---
# Deploy repo-owned Compose declarations (compose/<project>/compose.yml) to
# inventory hosts. Non-secret source; server-local .env provides the values.
# Gated: apply requires compose_deploy_confirm=true; --check is a read-only
# diff + validation. See runbooks/ansible-operations.md.
- name: Deploy repo-owned Compose declarations
hosts: docker_hosts
become: true
gather_facts: false
serial: 1
roles:
- role: compose_deploy
tags: [compose, deploy, mutating]
@@ -0,0 +1,8 @@
---
# Allowlist of compose/ projects this playbook may deploy. A host may only
# reference a project listed here (see tasks: "Require a repo compose project").
compose_repo_projects:
- vaultwarden
- pdns
- adguardhome
- unifi
@@ -0,0 +1,84 @@
---
# Deploy the repo-owned, sanitized Compose declaration to the host.
#
# Safety model:
# - Only hosts with an inventory `compose_repo_project` (allowlisted) are valid.
# - The repo file is staged to `<file>.dsh-new` and validated with
# `docker compose config --quiet` against the server-local .env BEFORE it
# replaces anything. A failed validation never touches the live file.
# - The current file is kept as `*.bak-<timestamp>` before promotion.
# - Apply mode requires `compose_deploy_confirm=true`; `--check` gives a
# read-only diff + validation without writes.
# - The playbook never writes, reads, or transfers the server .env.
- name: Require an allowlisted repo compose project for this host
ansible.builtin.assert:
that:
- compose_repo_project is defined
- compose_repo_project in compose_repo_projects
fail_msg: >-
No allowlisted compose_repo_project for {{ inventory_hostname }}.
Supported: {{ compose_repo_projects | join(', ') }}.
- name: Require explicit confirmation for apply mode
ansible.builtin.assert:
that:
- ansible_check_mode or (compose_deploy_confirm | bool)
fail_msg: >-
This playbook replaces the server compose file and may recreate
containers. Run with --check for a read-only diff, or supply
compose_deploy_confirm=true to apply.
- name: Stage the repo compose file next to the live one
ansible.builtin.copy:
src: "{{ playbook_dir }}/../../compose/{{ compose_repo_project }}/compose.yml"
dest: "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
mode: "0644"
diff: true
register: compose_stage
- name: Validate staged compose against the server .env (read-only)
ansible.builtin.command:
argv:
- docker
- compose
- -f
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
- --project-directory
- "{{ compose_project_dir }}"
- config
- --quiet
register: compose_validate
changed_when: false
failed_when: compose_validate.rc != 0
- name: Show staged-vs-live difference
ansible.builtin.debug:
msg: "{{ compose_stage.diff | default('(no change)') }}"
when: ansible_check_mode
- name: Back up the current compose file (apply mode)
ansible.builtin.shell:
cmd: >-
cp -a '{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}'
'{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.bak-$(date +%Y%m%d-%H%M%S)'
when: not ansible_check_mode
- name: Promote the validated compose file (apply mode)
ansible.builtin.command:
argv:
- mv
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}"
when: not ansible_check_mode
- name: Apply the compose declaration (apply mode)
ansible.builtin.command:
argv:
- docker
- compose
- --project-directory
- "{{ compose_project_dir }}"
- up
- -d
when: not ansible_check_mode
@@ -31,9 +31,27 @@ compose_ps() {
}
check_compose() {
local output
output="$(compose_ps)" || { record critical 'compose_ps_failed'; return; }
if grep -qiE 'Exited|Restarting|[[:space:]]Dead[[:space:]]' <<<"$output"; then
local output services bad
# Only flag containers of *active* services (config --services excludes
# debug/profile-gated services such as vaultwarden's pgweb, which is
# intentionally stopped unless started with --profile debug).
services="$(docker compose --project-directory '{{ compose_project_dir }}' config --services 2>/dev/null)" || { record critical 'compose_ps_failed'; return; }
output="$(docker compose --project-directory '{{ compose_project_dir }}' ps --all --format json 2>&1)" || { record critical 'compose_ps_failed'; return; }
bad="$(printf '%s\n' "$output" | python3 -c '
import json, sys
services = set(sys.argv[1].split())
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
c = json.loads(line)
except Exception:
continue
if c.get("Service") in services and c.get("State") in ("exited", "restarting", "dead"):
print(c.get("Service"))
' "$services")"
if [[ -n "$bad" ]]; then
record critical 'compose_unhealthy_container'
else
record ok 'compose_ok'
@@ -14,5 +14,11 @@ rm -f '{{ healthcheck_state_dir }}/latest-{{ healthcheck_profile_scripts[profile
this_rc="${PIPESTATUS[0]}"
[ "$this_rc" -gt "$rc" ] && rc="$this_rc"
{% endfor %}
aggregate_result{% for profile in healthcheck_profiles %} {{ healthcheck_profile_scripts[profile] | replace('.sh', '') }}{% endfor %}
# Collect profile check names line-by-line (robust against Jinja trim_blocks
# whitespace control, which would otherwise merge this into one line).
aggregate_args=""
{% for profile in healthcheck_profiles %}
aggregate_args="$aggregate_args {{ healthcheck_profile_scripts[profile] | replace('.sh', '') }}"
{% endfor %}
aggregate_result $aggregate_args
exit "$rc"
@@ -15,11 +15,12 @@ grep -Fq 'vw-db' <<<"$health" || record critical 'postgres_missing'
check_https 'https://auth.wsvc.info/' '^200$'
check_tls_days auth.wsvc.info 443
# Read effective config only inside the service and report booleans/fingerprints,
# never its SMTP password or other secret fields.
smtp_result="$(docker compose --project-directory '{{ compose_project_dir }}' exec -T vaultwarden python3 - <<'PY' 2>&1
# Read effective config from the mounted vw-data dir on the host and run the
# SMTP AUTH probe from the host (the vaultwarden image has no python3; the
# host does). Never print the SMTP password.
smtp_result="$(python3 - <<'PY' 2>&1
import json, pathlib, smtplib, ssl
cfg=json.loads(pathlib.Path('/data/config.json').read_text())
cfg=json.loads(pathlib.Path('{{ compose_project_dir }}/vw-data/config.json').read_text())
host=cfg.get('smtp_host'); port=int(cfg.get('smtp_port') or 0)
user=cfg.get('smtp_username')
smtp_secret=cfg.get('smtp_password')
+51
View File
@@ -0,0 +1,51 @@
# compose/ — repo-owned Compose declarations
Non-secret Compose sources for the Docker hosts. Secrets are **never** in these
files: every secret is a `${VAR}` reference resolved from the **server-local
`.env`** (docker compose reads `.env` from the project directory automatically).
## Source-of-truth matrix
| Project | Host | Compose source | Mechanism |
|---------|------|----------------|-----------|
| `vaultwarden` | us2 (`/opt/vaultwarden`) | `compose/vaultwarden/compose.yml` | static file + `compose-deploy.yml` |
| `pdns` | hk2 (`/opt/pdns`) | `compose/pdns/compose.yml` | static file + `compose-deploy.yml` |
| `pgdb` | pgdb (`/opt/database`, 无 ansible) | `compose/pgdb/compose.yml` | static file(手动部署:scp → `docker compose config -q``up -d`;服务器文件名 `docker-compose.yml` |
| `soft-serve` | us2 (`/opt/soft-serve`, 已退役停用) | `compose/soft-serve/compose.yml` (+ `Dockerfile.backup`, `scripts/`) | static file(参考镜像; 2026-09-18 被 gitea 替换 VPS-94, 数据保留作回滚) |
| `gitea` | us2 (`/opt/gitea`) | `compose/gitea/compose.yml` (+ `Dockerfile.backup`, `scripts/`) | static file(参考镜像, 未接入 compose-deploy; 服务器文件为准; 2026-09-18 替换 soft-serve, VPS-94 |
| `adguardhome` | dns.windy.lan (`/opt/adguardhome`) | — (待从 LAN 提取) | static file (pending) |
| `unifi` | ubnt (`/home/windy/unifi-9`) | — (待从 LAN 提取) | static file (pending) |
| `wireguard` | us4 (`/opt/wireguard`) | `ansible/templates/wireguard-compose.yml.j2` | role-rendered (inventory vars) |
| `rustdesk` | hk2 (`/opt/rustdesk`) | `ansible/roles/rustdesk/templates/compose.yml.j2` | role-rendered (inventory vars) |
| `mailcow` | mx2 (`/opt/mail`) | — (mailcow update generator owns it) | excluded by design |
Mechanism rule: **static** `compose/<project>/compose.yml` for declarations that
do not vary per host; **role-rendered j2** for declarations driven by inventory
vars (image pins, relay host). One mechanism per project; do not duplicate a
project in both.
## Deploying a static project
```bash
cd ansible
# Read-only diff + validation against the server .env (no writes)
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden --check --diff
# Apply: stage repo file → validate `docker compose config -q` → backup current
# file → promote → `docker compose up -d` (gated)
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden \
-e '{"compose_deploy_confirm": true}'
```
See [`../runbooks/ansible-operations.md`](../runbooks/ansible-operations.md).
## Adding a project
1. Sanitize the live compose so every secret is `${VAR}` from `.env`
(prefer `${VAR:?missing VAR}` for required keys).
2. Commit `compose/<project>/compose.yml` + `.env.example` (key names only).
3. Add `compose_repo_project` (+ `compose_remote_file` if not `compose.yml`) to
the host in `ansible/inventory/hosts.yml`, and allowlist the project in
`ansible/roles/compose_deploy/defaults/main.yml`.
4. Verify with `--check --diff` (zero diff) then a gated apply.
+4
View File
@@ -0,0 +1,4 @@
# compose/gitea — 秘密一律走服务器本地 .env, 不入库
# 迁移期一次性: Gitea 管理员生成的 token (mirror-migrate.sh 读取, 用后撤销)
GITEA_MIGRATE_USER=
GITEA_MIGRATE_TOKEN=
+3
View File
@@ -0,0 +1,3 @@
FROM alpine:3.20
RUN apk add --no-cache sqlite rsync tzdata
WORKDIR /scripts
+59
View File
@@ -0,0 +1,59 @@
# Gitea on us2 — reference compose (Plane VPS-94, 迁移完成 2026-09-18)
# 参考镜像, 服务器 /opt/gitea 文件为准 (同 soft-serve 约定, 未接入 compose-deploy)
# rootless 镜像: uid 1000 原生非 root; 数据 /var/lib/gitea (宿主 ./data), 配置 /etc/gitea (宿主 ./config)
# SSH: 容器内监听 2322 (非特权, SSH_LISTEN_PORT), 对外 repo.windy.me:2222 经 Traefik TCP entrypoint `ssh`
services:
gitea:
image: gitea/gitea@sha256:1c17ecaead42eb3b5391553d8708103a4beb0e86edf5b9ebc1eb269c318845f2 # 1.27.3-rootless
container_name: gitea
restart: unless-stopped
user: "1000:1000"
environment:
TZ: Asia/Shanghai
volumes:
- ./data:/var/lib/gitea
- ./config:/etc/gitea
- ./secrets:/secrets:ro # 复用的 soft-serve host key (SSH_SERVER_HOST_KEYS)
networks:
- traefik
labels:
- traefik.enable=true
# Web UI: repo.windy.me (2026-09-18 操作者决定复用现有域名, 免 DNS 变更)
- traefik.http.routers.gitea-web.rule=Host(`repo.windy.me`)
- traefik.http.routers.gitea-web.entrypoints=websecure
- traefik.http.routers.gitea-web.tls.certresolver=letsencrypt
- traefik.http.services.gitea-web.loadbalancer.server.port=3000
# SSH: 接管 :2222 (entrypoint 已存在, router 动态生效, 无需重启 Traefik)
- traefik.tcp.routers.gitea-ssh.entrypoints=ssh
- traefik.tcp.routers.gitea-ssh.rule=HostSNI(`*`)
- traefik.tcp.routers.gitea-ssh.tls=false
- traefik.tcp.services.gitea-ssh.loadbalancer.server.port=2322
gitea-backup:
build:
context: .
dockerfile: Dockerfile.backup
container_name: gitea-backup
restart: unless-stopped
volumes:
- ./data:/data:ro
- ./config:/config:ro
- ./backups:/backup
- ./scripts:/scripts
environment:
TZ: Asia/Shanghai
BACKUP_UID: 1000
BACKUP_GID: 1000
entrypoint: >
/bin/sh -ec "
umask 077 &&
touch /backup/backup.log &&
crontab /scripts/crontab.txt &&
echo '[INFO] gitea backup cron installed' &&
crond -f -l 8
"
networks:
traefik:
external: true
name: vw-net
+20
View File
@@ -0,0 +1,20 @@
#!/bin/sh
set -eu
umask 077
D() { date "+%Y-%m-%d %H:%M:%S"; }
TS=$(date +%Y%m%d_%H%M%S)
OUT="/backup/gitea_${TS}"
mkdir -p "$OUT"
echo "[$(D)] Starting gitea backup -> $OUT"
# rootless 布局: app.ini=/etc/gitea(宿主 ./config), db+repos=/var/lib/gitea/data(宿主 ./data/data)
# app.ini 含 SECRET_KEY/INTERNAL_TOKEN — 恢复 2FA/session/mirror 凭据必需
tar czf "$OUT/app.ini.tar.gz" -C /config app.ini
sqlite3 /data/data/gitea.db ".backup '$OUT/gitea.db'"
rsync -a /data/data/git/repositories/ "$OUT/repos/"
tar czf "$OUT/repos.tar.gz" -C "$OUT" repos
rm -rf "$OUT/repos"
chmod 600 "$OUT"/*
if [ -n "${BACKUP_UID:-}" ] && [ -n "${BACKUP_GID:-}" ]; then
chown -R "$BACKUP_UID:$BACKUP_GID" "$OUT" /backup/backup.log
fi
echo "[$(D)] Backup OK: $(du -sh "$OUT" | cut -f1)"
+4
View File
@@ -0,0 +1,4 @@
# Run gitea backup daily at 02:00
0 2 * * * /bin/sh /scripts/backup.sh >> /backup/backup.log 2>&1
# Prune backups older than 14 days daily at 03:00
0 3 * * * /bin/sh /scripts/prune.sh >> /backup/backup.log 2>&1
+41
View File
@@ -0,0 +1,41 @@
#!/bin/sh
# 一次性迁移辅助 (Plane VPS-94 Phase 2): 在 gitea 容器内执行。
# 已于 2026-09-18 执行完成 (16 仓), 留档备查; 复用时按 VPS-94 流程重生成一次性 token。
# 用法:
# GITEA_MIGRATE_USER=<user> GITEA_MIGRATE_TOKEN=<token> \
# docker exec -e GITEA_MIGRATE_USER -e GITEA_MIGRATE_TOKEN gitea \
# /scripts/mirror-migrate.sh [public_repo ...]
# 每仓: API 建仓 (默认 private, 参数中列出的为 public) -> push --mirror。
# default_branch 按源仓 symbolic-ref HEAD 设置, 避免非 main 源仓在 Gitea 显示为空。
# 结束后按 VPS-94 Phase 3 逐仓核对 git ls-remote ref 全集。
set -eu
MUSER="${GITEA_MIGRATE_USER:?need GITEA_MIGRATE_USER}"
TOKEN="${GITEA_MIGRATE_TOKEN:?need GITEA_MIGRATE_TOKEN}"
SRC="/migration-src"
API="http://localhost:3000/api/v1"
PUBLIC_REPOS=" $* "
migrate_one() {
dir="$1"
git -C "$dir" rev-parse --git-dir >/dev/null 2>&1 || { echo "[SKIP] $dir (not a git repo)"; return 0; }
name=$(basename "$dir"); name=${name%.git}
def_branch=$(git -C "$dir" symbolic-ref --short HEAD)
case "$PUBLIC_REPOS" in *" $name "*) private=false ;; *) private=true ;; esac
echo "[MIGRATE] $name (default=$def_branch private=$private)"
code=$(curl -s -o /dev/null -w '%{http_code}' -X POST "$API/user/repos" \
-H "Authorization: token $TOKEN" -H "Content-Type: application/json" \
-d "{\"name\":\"$name\",\"private\":$private,\"default_branch\":\"$def_branch\",\"auto_init\":false}")
case "$code" in
201) : ;;
409) echo " [WARN] $name 已存在, 直接补推" ;;
*) echo " [FAIL] create HTTP $code"; return 1 ;;
esac
git -C "$dir" push --mirror "http://$MUSER:$TOKEN@localhost:3000/$MUSER/$name.git"
echo " [OK] $name pushed"
}
for dir in "$SRC"/*.git "$SRC"/cdia; do
[ -d "$dir" ] || continue
migrate_one "$dir"
done
echo "[DONE] 全部处理完毕; 迁移后记得撤销一次性 token"
+5
View File
@@ -0,0 +1,5 @@
#!/bin/sh
set -eu
D() { date "+%Y-%m-%d %H:%M:%S"; }
ls -dt /backup/gitea_* 2>/dev/null | tail -n +15 | xargs -r rm -rf
echo "[$(D)] Pruned. Kept $(ls -d /backup/gitea_* 2>/dev/null | wc -l) backups (max 14)"
+39
View File
@@ -0,0 +1,39 @@
# .env.example — PowerDNS stack (hk2.chans.xyz, /opt/pdns)
#
# Non-secret key reference ONLY. Real values live in the server-local .env
# (never commit them). Compose requires the `:?`-marked keys to be present.
# Runtime
TZ=Asia/Shanghai
# Postgres superuser (db + backup + pgweb)
PGUSER=
PGPASSWORD=
DB_HOST=db
DB_PORT=5432
# Application database (auth / poweradmin / backup)
DB_NAME=pdns
DB_USER=pdns
DB_PASS=
ADMIN_DB=pdnsadmin
# Backups
CRON_SCHEDULE=0 3 * * *
RETENTION_DAYS=7
MAX_BACKUPS=7
DUMP_ROLES=true
# PowerDNS auth API
PDNS_API_KEY=
# Poweradmin (first-run admin + session)
PA_SESSION_KEY=
PA_ADMIN_USERNAME=
PA_ADMIN_PASSWORD=
PA_ADMIN_EMAIL=
PA_ADMIN_FULLNAME=
# pgweb debug profile
PGWEB_USER=
PGWEB_PASS=
+159
View File
@@ -0,0 +1,159 @@
networks:
frontend:
name: traefik
external: true
backend:
internal: true
edge:
services:
db:
image: postgres:16
container_name: pdns-db
environment:
POSTGRES_DB: postgres
POSTGRES_USER: ${PGUSER:?missing PGUSER}
POSTGRES_PASSWORD: ${PGPASSWORD:?missing PGPASSWORD}
TZ: ${TZ:-Asia/Shanghai}
PGTZ: ${TZ:-Asia/Shanghai}
volumes:
# Keep the existing mount path to avoid moving the current data directory.
- dbdata:/var/lib/postgresql
- ./db-init-generated:/docker-entrypoint-initdb.d:ro
- ./backup:/backup:ro
healthcheck:
test: ["CMD-SHELL", "pg_isready -U \"$${POSTGRES_USER}\" -d \"$${POSTGRES_DB}\""]
interval: 10s
timeout: 5s
retries: 10
restart: unless-stopped
networks: [backend, edge]
auth:
image: powerdns/pdns-auth-50:5.0.6
container_name: pdns-auth
depends_on:
db:
condition: service_healthy
ports:
- "53:53/udp"
- "53:53/tcp"
- "127.0.0.1:8081:8081"
environment:
PDNS_API_KEY: ${PDNS_API_KEY:?missing PDNS_API_KEY}
DB_NAME: ${DB_NAME:?missing DB_NAME}
DB_USER: ${DB_USER:?missing DB_USER}
DB_PASS: ${DB_PASS:?missing DB_PASS}
TEMPLATE_FILES: secrets
volumes:
- ./auth/pdns.conf:/etc/powerdns/pdns.conf:ro
- ./auth/templates.d:/etc/powerdns/templates.d:ro
- ./auth/keys:/var/lib/powerdns
- ./auth/import:/import
- ./auth/export:/export
- ./auth/logs:/var/log/pdns
healthcheck:
test:
[
"CMD-SHELL",
"python3 -c \"import json, os, urllib.request; req = urllib.request.Request('http://127.0.0.1:8081/api/v1/servers/localhost', headers={'X-API-Key': os.environ['PDNS_API_KEY']}); data = json.load(urllib.request.urlopen(req, timeout=3)); assert data['daemon_type'] == 'authoritative'\""
]
interval: 10s
timeout: 5s
retries: 12
restart: unless-stopped
networks: [backend, edge]
poweradmin:
image: poweradmin/poweradmin:stable
container_name: poweradmin
depends_on:
db:
condition: service_healthy
auth:
condition: service_healthy
environment:
DB_TYPE: pgsql
DB_HOST: ${DB_HOST:-db}
DB_PORT: ${DB_PORT:-5432}
DB_NAME: ${DB_NAME:?missing DB_NAME}
DB_USER: ${DB_USER:?missing DB_USER}
DB_PASS: ${DB_PASS:?missing DB_PASS}
PA_PDNS_API_URL: http://auth:8081
PA_PDNS_API_KEY: ${PDNS_API_KEY:?missing PDNS_API_KEY}
PA_DNS_BACKEND: sql
PDNS_VERSION: ${PDNS_VERSION:-50}
DNS_NS1: ${DNS_NS1:-ns1.wsvc.info}
DNS_NS2: ${DNS_NS2:-ns2.wsvc.info}
DNS_HOSTMASTER: ${DNS_HOSTMASTER:-hostmaster.wsvc.info}
PA_APP_TITLE: ${PA_APP_TITLE:-Poweradmin}
PA_TIMEZONE: ${TZ:-Asia/Shanghai}
PA_SESSION_KEY: ${PA_SESSION_KEY:?missing PA_SESSION_KEY}
PA_CREATE_ADMIN: ${PA_CREATE_ADMIN:-1}
PA_ADMIN_USERNAME: ${PA_ADMIN_USERNAME:?missing PA_ADMIN_USERNAME}
PA_ADMIN_PASSWORD: ${PA_ADMIN_PASSWORD:?missing PA_ADMIN_PASSWORD}
PA_ADMIN_EMAIL: ${PA_ADMIN_EMAIL:?missing PA_ADMIN_EMAIL}
PA_ADMIN_FULLNAME: ${PA_ADMIN_FULLNAME:?missing PA_ADMIN_FULLNAME}
TRUSTED_PROXIES: private_ranges
DEBUG: "false"
restart: unless-stopped
networks: [backend, frontend]
labels:
- "traefik.enable=true"
- "traefik.docker.network=traefik"
- "traefik.http.routers.poweradmin.rule=Host(`pdns.wsvc.info`)"
- "traefik.http.routers.poweradmin.entrypoints=websecure"
- "traefik.http.routers.poweradmin.tls.certresolver=letsencrypt"
- "traefik.http.services.poweradmin.loadbalancer.server.port=80"
backup:
# Use postgres:16 so bash/pg_dump/flock exist without runtime package installs.
# backend is internal:true — Alpine apk at start cannot reach mirrors.
image: postgres:16
container_name: pdns-backup
depends_on:
db:
condition: service_healthy
environment:
TZ: ${TZ:-Asia/Shanghai}
DB_HOST: ${DB_HOST:-db}
DB_PORT: ${DB_PORT:-5432}
DB_USER: ${PGUSER:?missing PGUSER}
DB_PASS: ${PGPASSWORD:?missing PGPASSWORD}
DB_NAME: ${DB_NAME:?missing DB_NAME}
RETENTION_DAYS: ${RETENTION_DAYS:-7}
MAX_BACKUPS: ${MAX_BACKUPS:-7}
DUMP_ROLES: ${DUMP_ROLES:-true}
CRON_SCHEDULE: ${CRON_SCHEDULE:?missing CRON_SCHEDULE}
volumes:
- ./backup:/backup
- ./scripts:/scripts:ro
entrypoint: ["/bin/bash", "/scripts/backup-scheduler.sh"]
restart: unless-stopped
networks: [backend]
pgweb:
image: sosedoff/pgweb:0.16.2
container_name: pdns_pgweb
restart: unless-stopped
environment:
PGWEB_DATABASE_URL: "postgres://${PGUSER:?missing PGUSER}:${PGPASSWORD:?missing PGPASSWORD}@${DB_HOST:-db}:${DB_PORT:-5432}/${DB_NAME:?missing DB_NAME}?sslmode=disable"
PGWEB_AUTH_USER: ${PGWEB_USER:?missing PGWEB_USER}
PGWEB_AUTH_PASS: ${PGWEB_PASS:?missing PGWEB_PASS}
TZ: ${TZ:-Asia/Shanghai}
depends_on:
db:
condition: service_healthy
networks: [backend, frontend]
labels:
- "traefik.enable=true"
- "traefik.docker.network=traefik"
- "traefik.http.routers.pgweb.rule=Host(`pgweb.wsvc.info`)"
- "traefik.http.routers.pgweb.entrypoints=websecure"
- "traefik.http.routers.pgweb.tls.certresolver=letsencrypt"
- "traefik.http.services.pgweb.loadbalancer.server.port=8081"
volumes:
dbdata: {}
+5
View File
@@ -0,0 +1,5 @@
# pgdb compose secrets — copy to /opt/database/.env on the host, chmod 600.
# NEVER commit the real values. Generate: openssl rand -hex 24
POSTGRES_PASSWORD=change-me-strong-hex
PGWEB_AUTH_USER=pgweb
PGWEB_AUTH_PASS=change-me-strong-hex
+65
View File
@@ -0,0 +1,65 @@
# pgdb (192.168.55.15) — TimescaleDB + pgweb GUI + nightly backup
#
# Deploy: copy this file to /opt/database/docker-compose.yml on pgdb,
# create /opt/database/.env (chmod 600) from .env.example, plus
# /opt/database/pgweb-bookmarks/{hass,scribe}.toml (chmod 600, contains DB password).
# Then: docker compose config --quiet && docker compose up -d
#
# Rollback: previous launch command is kept at /opt/database/run
# (container is stateless; data lives on /srv/pgdata).
services:
timescaledb:
image: timescale/timescaledb:latest-pg18
container_name: timescaledb
restart: unless-stopped
ports:
- "192.168.55.15:5432:5432" # bind VM IP only (no IPv6 wildcard)
environment:
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
volumes:
- /srv/pgdata:/var/lib/postgresql # data disk (ext4 /dev/sdb1)
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]
interval: 30s
timeout: 5s
retries: 5
start_period: 10s
pgweb:
image: sosedoff/pgweb:latest
container_name: pgweb
restart: unless-stopped
# bind/listen/readonly/sessions/bookmarks-only/bookmarks-dir are CLI flags (no env equivalent in v0.17.0)
command: ["pgweb", "--bind", "0.0.0.0", "--listen", "8081", "--readonly", "--sessions", "--bookmarks-only", "--bookmarks-dir", "/bookmarks"]
ports:
- "192.168.55.15:8081:8081" # LAN only + basic auth (see .env)
environment:
PGWEB_AUTH_USER: ${PGWEB_AUTH_USER}
PGWEB_AUTH_PASS: ${PGWEB_AUTH_PASS}
PGWEB_BOOKMARKS_DIR: /bookmarks
volumes:
- ./pgweb-bookmarks:/bookmarks:ro # bookmark .toml files (contain DB password, keep 0600)
depends_on:
timescaledb:
condition: service_healthy
pg-backup:
image: prodrigestivill/postgres-backup-local:latest # latest = postgres 18 base (pg_dump 18.x)
container_name: pg-backup
restart: unless-stopped
environment:
POSTGRES_HOST: timescaledb
POSTGRES_DB: "hass scribe postgres"
POSTGRES_USER: postgres
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
POSTGRES_EXTRA_OPTS: "-Fc" # custom-format dumps (pg_restore)
SCHEDULE: "0 2 * * *" # nightly 02:00 (TZ=Asia/Shanghai -> local 02:00)
BACKUP_ON_START: "TRUE" # immediate backup on first start
BACKUP_SUFFIX: ".dump"
HEALTHCHECK_PORT: "80" # go-cron health endpoint for the image healthcheck
TZ: "Asia/Shanghai" # match original host-cron 02:00 local (container default is UTC)
volumes:
- /opt/database/backups:/backups # POSIX fs required; root disk, separate from data disk
depends_on:
timescaledb:
condition: service_healthy
+24
View File
@@ -0,0 +1,24 @@
[Unit]
Description=Reconcile pgdb compose stack (timescaledb + pgweb + pg-backup) at boot
Documentation=file:///opt/database/docker-compose.yml
After=network-online.target docker.service
Wants=network-online.target
Requires=docker.service
[Service]
Type=oneshot
RemainAfterExit=yes
WorkingDirectory=/opt/database
# Idempotent boot-time reconcile. docker's own restore can fail to bind the
# published ports (192.168.55.15:5432/8081) when the VM IP is not yet usable
# right after boot (EADDRNOTAVAIL, observed 2026-08-30): timescaledb/pgweb
# then stay stopped until a manual `docker compose up`. This unit retries
# `docker compose up -d` (a no-op when the stack is healthy) until the port
# listens, and force-recreates as a last resort to recover a network-detached
# container. Data lives on bind mounts (/srv/pgdata, /opt/database/backups),
# so recreation is safe.
ExecStart=/bin/bash -c 'for i in $(seq 1 12); do docker compose up -d --remove-orphans; sleep 2; if ss -tln | grep -q "192.168.55.15:5432"; then exit 0; fi; sleep 3; done; echo "pgdb-compose: retries exhausted, force-recreating"; docker compose up -d --force-recreate; sleep 10; ss -tln | grep -q "192.168.55.15:5432"'
TimeoutStartSec=180
[Install]
WantedBy=multi-user.target
+2
View File
@@ -0,0 +1,2 @@
# Soft Serve initial admin public key (used only on first boot)
SOFT_SERVE_INITIAL_ADMIN_KEYS=ssh-ed25519 AAAA... # replace with admin public key
+3
View File
@@ -0,0 +1,3 @@
FROM alpine:3.20
RUN apk add --no-cache sqlite tzdata
WORKDIR /scripts
+59
View File
@@ -0,0 +1,59 @@
services:
soft-serve:
image: charmcli/soft-serve:v0.12.2
container_name: soft-serve
restart: unless-stopped
# non-root (uid 1000 = windy; 与 backup sidecar BACKUP_UID 一致)
user: "1000:1000"
environment:
SOFT_SERVE_DATA_PATH: /var/lib/soft-serve
SOFT_SERVE_INITIAL_ADMIN: windy
SOFT_SERVE_INITIAL_ADMIN_KEYS: ${SOFT_SERVE_INITIAL_ADMIN_KEYS}
volumes:
- ./data:/var/lib/soft-serve
- soft-serve-app:/soft-serve
networks:
- traefik
labels:
- traefik.enable=true
# SSH over TCP via Traefik (entryPoint ssh -> container port 23231)
- traefik.tcp.routers.softserve-ssh.entrypoints=ssh
- traefik.tcp.routers.softserve-ssh.rule=HostSNI(`*`)
- traefik.tcp.routers.softserve-ssh.tls=false
- traefik.tcp.services.softserve-ssh.loadbalancer.server.port=23231
soft-serve-backup:
build:
context: .
dockerfile: Dockerfile.backup
container_name: soft-serve-backup
restart: unless-stopped
volumes:
- ./data:/data:ro
- ./backups:/backup
- ./scripts:/scripts
environment:
TZ: Asia/Shanghai
BACKUP_UID: 1000
BACKUP_GID: 1000
entrypoint: >
/bin/sh -ec "
umask 077 &&
touch /backup/backup.log &&
crontab /scripts/crontab.txt &&
echo '[INFO] soft-serve backup cron installed' &&
crond -f -l 8
"
volumes:
soft-serve-app:
networks:
traefik:
external: true
name: vw-net
+15
View File
@@ -0,0 +1,15 @@
#!/bin/sh
set -eu
umask 077
D() { date "+%Y-%m-%d %H:%M:%S"; }
TS=$(date +%Y%m%d_%H%M%S)
OUT="/backup/soft-serve_${TS}"
mkdir -p "$OUT"
echo "[$(D)] Starting soft-serve backup -> $OUT"
tar czf "$OUT/repos-config.tar.gz" -C /data repos hooks config.yaml ssh
sqlite3 /data/soft-serve.db ".backup '$OUT/soft-serve.db'"
chmod 600 "$OUT/repos-config.tar.gz" "$OUT/soft-serve.db"
if [ -n "${BACKUP_UID:-}" ] && [ -n "${BACKUP_GID:-}" ]; then
chown -R "$BACKUP_UID:$BACKUP_GID" "$OUT" /backup/backup.log
fi
echo "[$(D)] Backup OK: $(du -sh "$OUT" | cut -f1)"
+4
View File
@@ -0,0 +1,4 @@
# Run soft-serve backup daily at 02:00
0 2 * * * /bin/sh /scripts/backup.sh >> /backup/backup.log 2>&1
# Prune backups older than 14 days daily at 03:00
0 3 * * * /bin/sh /scripts/prune.sh >> /backup/backup.log 2>&1
+5
View File
@@ -0,0 +1,5 @@
#!/bin/sh
set -eu
D() { date "+%Y-%m-%d %H:%M:%S"; }
ls -dt /backup/soft-serve_* 2>/dev/null | tail -n +15 | xargs -r rm -rf
echo "[$(D)] Pruned. Kept $(ls -d /backup/soft-serve_* 2>/dev/null | wc -l) backups (max 14)"
+38
View File
@@ -0,0 +1,38 @@
# .env.example — Vaultwarden (us2.wsvc.info, /opt/vaultwarden)
#
# Non-secret key reference ONLY. Real values live in the server-local .env
# (never commit them). Copy the keys below into the server .env if a key is
# missing; the compose file requires them via ${VAR} / env_file.
# Service identity
DOMAIN=https://auth.wsvc.info
TEMPLATES_FOLDER=
# Postgres (compose services vaultwarden / backup / pg / pgweb)
DB_HOST=pg
DB_PORT=5432
DB_NAME=vaultwarden
DB_USER=vaultwarden
DB_PASS=
# pgweb debug profile
PGWEB_USER=
PGWEB_PASS=
PGWEB_DATABASE_URL=
# SMTP (mailcow mx2.windy.me:587 starttls)
SMTP_HOST=mx2.windy.me
SMTP_PORT=587
SMTP_SECURITY=starttls
SMTP_USERNAME=
SMTP_PASSWORD=
SMTP_FROM=
HELO_NAME=
# Admin console
ADMIN_TOKEN=
# Runtime
UID=1000
GID=1000
IP_HEADER=X-Forwarded-For
+107
View File
@@ -0,0 +1,107 @@
services:
vaultwarden:
image: vaultwarden/server:1.37.2
container_name: vaultwarden
restart: unless-stopped
env_file: ".env"
environment:
DOMAIN: "https://auth.wsvc.info"
DATABASE_URL: "postgresql://${DB_USER}:${DB_PASS}@${DB_HOST}:${DB_PORT}/${DB_NAME}"
volumes:
- ./vw-data:/data
extra_hosts:
- "mx2.windy.me:194.163.160.244"
networks:
- net
depends_on:
pg:
condition: service_healthy
labels:
- "traefik.enable=true"
- "traefik.docker.network=vw-net"
- "traefik.http.routers.vaultwarden.rule=Host(`auth.wsvc.info`)"
- "traefik.http.routers.vaultwarden.entrypoints=websecure"
- "traefik.http.routers.vaultwarden.tls=true"
- "traefik.http.routers.vaultwarden.tls.certresolver=letsencrypt"
- "traefik.http.services.vaultwarden.loadbalancer.server.port=80"
backup:
build:
context: .
dockerfile: Dockerfile.backup
container_name: vaultwarden-backup
restart: unless-stopped
volumes:
- ./backups:/backup
- ./scripts:/scripts
#user: "${UID:-1000}:${GID:-1000}"
environment:
DB_HOST: ${DB_HOST}
DB_PORT: ${DB_PORT}
DB_USER: ${DB_USER}
DB_NAME: ${DB_NAME}
DB_PASS: ${DB_PASS}
BACKUP_UID: ${UID:-0}
BACKUP_GID: ${GID:-0}
TZ: Asia/Shanghai
entrypoint: >
/bin/sh -ec "
umask 077 &&
printf '%s:%s:*:%s:%s\n' \"$$DB_HOST\" \"$$DB_PORT\" \"$$DB_USER\" \"$$DB_PASS\" > /root/.pgpass &&
chmod 600 /root/.pgpass &&
touch /backup/backup.log &&
crontab /scripts/crontab.txt &&
echo '[INFO] Backup cron installed' &&
echo '[INFO] Starting crond...' &&
crond -f -l 8
"
networks: [net]
pg:
image: postgres:16
container_name: vw-db
restart: unless-stopped
environment:
POSTGRES_DB: ${DB_NAME}
POSTGRES_USER: ${DB_USER}
POSTGRES_PASSWORD: ${DB_PASS}
TZ: Asia/Shanghai
PGTZ: Asia/Shanghai
volumes:
- vwdata:/var/lib/postgresql/data
- ./backups:/backup # to import existing dump
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ${DB_USER} -d ${DB_NAME}"]
interval: 10s
timeout: 5s
retries: 10
networks: [net]
pgweb:
profiles: ["debug"]
image: sosedoff/pgweb:0.16.2
container_name: vaultwarden-pgweb
restart: unless-stopped
environment:
# 用 Vaultwarden 的数据库参数拼接连接串
#DATABASE_URL: "postgres://${DB_USER}:${DB_PASS}@${DB_HOST}:${DB_PORT}/${DB_NAME}?sslmode=disable"
PGWEB_AUTH_USER: ${PGWEB_USER}
PGWEB_AUTH_PASS: ${PGWEB_PASS}
TZ: Asia/Shanghai
#ports:
# - "8082:8081" # 本地访问 http://localhost:8082
depends_on:
pg:
condition: service_healthy
networks: [net]
networks:
net:
name: vw-net
external: true
volumes:
vwdata: {}
+17
View File
@@ -0,0 +1,17 @@
# docs/archive — 归档文档
归档 = 单次调研、已过期,或与当前运维无行动指向的内容。恢复使用前先确认
内容仍与线上状态一致(本仓库原则:先证据后变更,live state 优先)。
## 归档清单
| 文件 | 归档日期 | 原位置 | 说明 |
|------|---------|--------|------|
| `lan-dns-alternatives.md` | 2026-08-17 | `docs/` | DNS 技术选型调研,零引用,无在途决策 |
| `agent-runbook-guide.md` | 2026-08-22 | `docs/` | 上游参考存档;仓库落地规范为 `RUNBOOKS.md` |
| `lan-core-switch-upgrade-plan.md` | 2026-08-22 | `docs/` | SE5420 历史规划参考;执行以 `docs/lan-se5420-deployment-guide.md` 为准 |
| `lan-rb5009-upgrade.md` | 2026-08-22 | `docs/` | 未采购的 ER-X→RB5009 休眠备选方案;其中 PVE 透传与 VLAN10 调研仍可参考 |
| `se5420-review-claim-verification-2026-08.md` | 2026-08-22 | `docs/` | 一次性评审复核调研(现场只读复核结论) |
> 购物类文档(打印机购买指南、交换机选型调研)已按整改计划移入 Obsidian
> vault`~/Documents/vault/my-vault/02_Areas/House/`),不在本目录。
+326
View File
@@ -0,0 +1,326 @@
# Agent Runbook 实用指南(v1
> **定位**:本指南用于把团队的重复性运维、交付与故障处理经验写成可由 Agent 安全执行的流程。它适用于以 Git 仓库为中心的工程协作模式,优先采用 **Markdown + Git 版本控制 + 明确的 Agent 路由规则**,而不是一开始引入复杂的自动化平台。
> 本文件为上游参考存档。仓库内落地规范见 [`RUNBOOKS.md`](../../RUNBOOKS.md),标准模板见 [`runbooks/_template.md`](../../runbooks/_template.md),索引见 [`runbooks/README.md`](../../runbooks/README.md)。
## 1. 什么是 Agent Runbook
Runbook 是预先设计的、可重复执行的操作流程,用于处理部署、告警、故障、配置变更、CI 修复等标准化工作。传统 Runbook 的主要读者是人;**Agent Runbook 则必须把人的隐性判断显式化**,使 Agent 能知道做什么、看到什么才算正常、下一步去哪里、何时停止以及如何撤销。
Google SRE 强调在事故发生前设计响应流程、系统化排障,并逐步将重复性运维工作自动化。[1] [2] AWS Systems Manager Automation 则把可执行 Runbook 建模为顺序步骤:每个步骤调用一个动作,前一步输出可以传递给后续步骤。[3] 这两种思路共同构成了 Agent Runbook 的实用基础。
| 层次 | 核心问题 | 应承担的职责 |
|---|---|---|
| `AGENTS.md` | **何时使用哪份流程?** | 工作路由、通用操作约束、无匹配流程时的默认行为 |
| `runbooks/*.md` | **这件事按什么流程做?** | 前置条件、分步操作、决策分支、验证、停止条件与回滚 |
| Skill / MCP / Tool | **有哪些可调用能力?** | 具体能力、参数、权限边界和使用说明 |
| Shell / GitHub / Linear / SSH 等 | **实际如何执行?** | 对系统、代码库或外部服务执行操作 |
## 2. 设计目标与适用边界
Agent Runbook 的目标不是让 Agent 在所有异常下“想办法修好”,而是在一个**已知、受控、可验证、可回退**的边界中提高执行一致性。它应当优先覆盖高频、后果明确、流程稳定的操作,例如 CI 失败定位、Issue 到合并请求、发布前检查、标准部署、回滚及网络变更。
| 适合纳入 Runbook | 暂不适合直接自动执行 |
|---|---|
| 明确输入、固定步骤、可观察结果的操作 | 目标或验收标准尚不清楚的探索性任务 |
| 可在每次修改后验证状态的变更 | 缺失关键参数、权限或上下文的任务 |
| 具有安全回滚路径的发布与配置调整 | 高破坏性、不可逆或影响面未知的操作 |
| 可由权限与审批规则约束的运维流程 | 与既有流程事实冲突、无法判断根因的异常场景 |
> **基本原则**:当实际状态与 Runbook 的假设冲突,Agent 应停止并呈报,而不是补全未知信息、绕过检查或继续试错。
## 3. Agent Runbook 的最小字段
与普通人工 Runbook 相比,Agent Runbook 必须显式包含以下六类控制信息。缺少其中任一项,都会增加盲目执行或错误恢复的风险。
| 字段 | 作用 | 写作要求 |
|---|---|---|
| **Action** | 定义当前要执行的动作 | 使用可观察、可执行的动词;避免“检查一下”“适当调整”等模糊表述 |
| **Expected** | 描述正常状态或预期输出 | 给出具体信号、阈值、状态码、测试结果或页面表现 |
| **Decision** | 定义分支与下一跳 | 用“条件 → 下一步”的形式;无法判断时指向 `STOP` |
| **Verification** | 确认变更真正生效 | 在每个有副作用的步骤后执行,不能被跳过 |
| **Stop condition** | 规定何时不得继续 | 明确列出信息缺失、状态冲突、权限不足、验证失败等条件 |
| **Rollback** | 描述如何恢复到变更前状态 | 标明触发条件、前提、撤销步骤及回滚后的验证方式 |
## 4. 推荐目录与路由机制
建议把流程与代码一起保存在 Git 仓库中。这样 Runbook 可以评审、版本化、随系统演进更新,也能与相关 Issue、PR 和配置建立可追溯关系。
```text
repo/
├── AGENTS.md
├── RUNBOOKS.md
├── runbooks/
│ ├── README.md
│ ├── issue-to-merge.md
│ ├── fix-ci.md
│ ├── release.md
│ ├── rollback.md
│ ├── network-change.md
│ └── network-recovery.md
└── ...
```
### `AGENTS.md`:只做路由与通用约束
`AGENTS.md` 不应重复流程细节。它只需要规定 Agent 在进行操作类工作前,先查找最具体且适用的 Runbook,并严格遵守其中的步骤、验证、停止和审批要求。
```markdown
# Operational Rules
Before performing operational work:
1. Inspect `runbooks/`.
2. Select the most specific applicable runbook.
3. Follow its steps in order.
4. Do not skip verification steps.
5. Respect STOP and approval conditions.
6. If no runbook applies, diagnose only; do not mutate production state.
## Routing
- CI failure → `runbooks/fix-ci.md`
- GitHub issue implementation → `runbooks/issue-to-merge.md`
- Deployment → `runbooks/release.md`
- Rollback → `runbooks/rollback.md`
- Network configuration → `runbooks/network-change.md`
- Network outage → `runbooks/network-recovery.md`
```
### `RUNBOOKS.md`:仓库级规范
`RUNBOOKS.md` 用于统一所有 Runbook 的字段、命名、评审要求和变更规则。每份 Runbook 只描述一种可识别的操作意图;如果流程已有明显分叉,应拆分为独立文件,而不是堆叠成长篇“万能流程”。
## 5. 规范模板
以下模板可直接保存为 `runbooks/_template.md` 使用。
```markdown
# Runbook: <名称>
## Purpose
说明本 Runbook 要解决的问题及成功结果。
## Scope
- 适用环境:<如 development / staging / production>
- 适用对象:<服务、仓库、组件或告警类型>
- 不适用情形:<需要改用其他 Runbook 或转人工的场景>
## Ownership
- Owner<团队或角色>
- Last reviewed<YYYY-MM-DD>
- Related systems<系统名称>
## Preconditions
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
## Inputs
| 输入 | 来源 | 是否必需 | 校验方法 |
|---|---|---:|---|
| <参数> | <来源> | 是/否 | <如何确认有效> |
## Safety
### Non-negotiable rules
- 先只读诊断,后执行变更。
- 不得把删除现有配置作为首次恢复动作。
- 不得猜测或编造缺失参数。
- 不得绕过失败的测试、检查或审批。
- 每次变更后必须完成对应验证。
- 破坏性操作必须获得明确批准。
### Stop conditions
- 实际状态与本文档的前提或预期结果冲突。
- 缺少必要输入、权限、审批或回滚能力。
- 验证失败且本文档没有明确的下一步。
- 影响范围超出 Scope。
### Approval gates
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|---|---|---:|---|
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
## Procedure
### Step 1 — Diagnose
**Action**
<执行只读诊断动作。>
**Expected**
<列出预期输出、状态或证据。>
**Decision**
- 若 <条件 A>,进入 Step 2。
- 若 <条件 B>,进入 Troubleshooting A。
- 若无法判断或状态冲突,`STOP` 并记录证据。
### Step 2 — Change
**Action**
<描述单一、可审计的变更动作。>
**Expected**
<变更后应出现的状态。>
**Verification**
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
**Rollback**
- 触发条件:<什么情况需要回滚>
- 回滚动作:<如何撤销>
- 回滚验证:<如何确认恢复成功>
## Troubleshooting
### Troubleshooting A — <异常名称>
- 证据收集:<日志、指标、命令输出、链接>
- 允许动作:<仅限已验证且低风险的动作>
- 下一步:<回到某步 / 转入另一 Runbook / STOP 并升级>
## Final Verification
只有同时满足以下标准,流程才算成功:
- <功能或服务状态>
- <自动化测试或健康检查>
- <监控指标或告警状态>
- <变更记录、PR 或 Issue 已更新>
## Failure Handling
若未能完成:
1. 停止进一步变更。
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
4. 按 <升级渠道> 交接,不继续猜测。
## References
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
```
## 6. 编写步骤的标准写法
每个步骤应只承担一个清晰目的,并使用“动作—预期—决策”的闭环表达。如下表所示,前者会导致 Agent 自主扩大操作范围,后者则为其提供安全边界。
| 不推荐写法 | 推荐写法 |
|---|---|
| “检查部署是否正常,不正常就修复。” | “读取部署状态与最近一次发布记录。若所有副本 `Ready` 且版本等于目标版本,进入 Final Verification;若副本未就绪,收集事件与日志并进入 Troubleshooting A;若版本不匹配且原因未知,`STOP`。” |
| “必要时修改配置。” | “仅当配置差异与变更单 `CHG-123` 完全一致且审批已记录时,应用指定键的值;应用后运行健康检查;失败则按 Rollback 回退。” |
| “测试失败时可先跳过。” | “任何必需测试失败均不得继续部署。记录失败测试、日志和提交版本;仅按 Troubleshooting B 处理。” |
## 7. 通用安全规则
以下规则适合在每份 Runbook 的 `Safety` 章节中复用。若某流程存在更严格要求,应以更严格要求为准。
```markdown
## Safety Rules
- Never delete an existing configuration as the first recovery action.
- Prefer read-only diagnosis before mutation.
- After every mutation, verify the expected state.
- If actual state conflicts with this runbook, STOP.
- Do not invent missing parameters.
- Do not bypass failed tests.
- Destructive actions require explicit approval.
```
这些约束体现了一个关键顺序:**先证据,后变更;先小范围,后扩大;先验证,后结束;不确定则停止。** 特别是停止条件必须可操作,例如“权限不足”“缺少变更单”“生产状态与前提不一致”“错误率超过 1%”等,而不应写成“情况复杂时停止”。
## 8. 运行与审计流程
Agent 执行 Runbook 时,应按照固定运行模型工作。每一步的输入、动作、输出和下一跳都应可追踪,这与 AWS 自动化 Runbook 的顺序步骤和输出传递思想一致。[3]
```text
输入与前置条件
只读诊断
确认预期状态或决策分支
获取审批(如需要)
执行最小变更
立即验证
成功收尾 / 回滚 / 停止并升级
```
| 阶段 | Agent 必须产出的证据 | 禁止行为 |
|---|---|---|
| 输入确认 | 参数来源、环境、目标资源、权限与审批状态 | 用猜测值补全必需参数 |
| 诊断 | 命令输出、日志、指标或页面状态 | 在未诊断前直接修改生产状态 |
| 变更 | 实际执行内容、变更范围、时间 | 将多个无关变更混在一起执行 |
| 验证 | 测试、健康检查、监控状态与预期对比 | 以“命令执行成功”代替业务验证 |
| 失败处理 | 已做步骤、异常证据、回滚状态和升级对象 | 无限制重试或绕过失败检查 |
## 9. 从人工操作到自动化的成熟路径
不建议在流程尚未稳定时先构建复杂 DSL 或全自动编排。应先积累真实案例,把可重复部分固化为 Markdown Runbook,再把已稳定、低歧义、可验证的操作迁移到脚本、CI、Skill 或自动化系统。Google SRE 将能够由机器替代的重复性人工工作视为应逐步消除的 toil。[4]
| 阶段 | 主要形式 | 人的角色 | 自动化边界 |
|---|---|---|---|
| 1. 人工处理 | 现场处置与复盘 | 执行、判断、记录 | 不自动化 |
| 2. Markdown Runbook | 固化步骤与证据要求 | 审核流程与异常判断 | Agent 可辅助诊断 |
| 3. Agent + Runbook | 严格按流程执行 | 审批高风险动作、处理例外 | 受停止条件约束的执行 |
| 4. Script / Skill / CI / Automation | 把稳定步骤程序化 | 处理异常和维护自动化 | 自动完成重复性操作 |
| 5. 人工审批 + 自动执行 | 常规流程端到端运行 | 决策、审计与治理 | 审批门控下的自动变更 |
## 10. 上线前检查清单
在将一份新 Runbook 交给 Agent 使用前,建议由流程所有者按以下清单审核。
| 检查项 | 合格标准 |
|---|---|
| 问题边界 | Purpose 与 Scope 清楚描述适用和不适用情形 |
| 输入 | 所有必需输入都有来源、格式和校验方法 |
| 步骤 | 每一步均有 Action、Expected 与明确的下一跳 |
| 变更控制 | 所有修改动作都有 Verification;关键动作有 Rollback |
| 安全控制 | Stop conditions、审批门槛和禁止行为已列明 |
| 异常处理 | 失败时知道收集什么证据、交给谁,而非继续猜测 |
| 可维护性 | 有 Owner、最近复审日期与关联文档;已在版本控制中评审 |
| 可演练性 | 已在安全环境或历史案例上走通至少一次 |
## 11. 建议的首批 Runbook
首次落地时,应优先选择频率较高、输入相对明确、变更可回退的场景。以下集合通常能覆盖大部分工程协作的基础需求。
| Runbook | 目的 | 关键安全控制 |
|---|---|---|
| `issue-to-merge.md` | 从已明确 Issue 到可评审变更 | Scope 锁定、测试门槛、PR 证据 |
| `fix-ci.md` | 诊断并修复 CI 失败 | 不跳过测试、不修改无关代码 |
| `release.md` | 执行标准发布 | 发布窗口、审批、健康检查、回滚点 |
| `rollback.md` | 恢复到已知稳定版本 | 明确触发条件、版本选择、回滚后验证 |
| `network-change.md` | 实施受控网络配置变更 | 影响评估、变更单、回退配置 |
| `network-recovery.md` | 处理网络异常与服务恢复 | 只读诊断优先、状态冲突即停止 |
## 12. 结论
Agent Runbook 的价值不在于把每一项运维工作立即自动化,而在于将团队的工程判断编码为**可路由、可验证、可停止、可回滚**的操作系统。对于多数团队,从仓库中的 `AGENTS.md``RUNBOOKS.md` 和一组 Markdown Runbook 起步,已经足够实用。
当某个流程经过多次执行、输入稳定、异常分支收敛且验证可靠后,再将其下沉为脚本、CI 或其他自动化能力。这样既能逐步降低重复性 toil,也能始终保留人类对高风险和例外情形的决策权。[4]
## References
[1]: https://sre.google/sre-book/managing-incidents/ "Google SRE Book — Managing Incidents"
[2]: https://sre.google/sre-book/effective-troubleshooting/ "Google SRE Book — Effective Troubleshooting"
[3]: https://docs.aws.amazon.com/systems-manager/latest/userguide/automation-documents.html "AWS Systems Manager — Creating your own runbooks"
[4]: https://sre.google/sre-book/eliminating-toil/ "Google SRE Book — Eliminating Toil"
[5]: https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-automation.html "AWS Systems Manager Automation"
[6]: https://docs.aws.amazon.com/systems-manager-automation-runbooks/latest/userguide/automation-runbook-reference.html "AWS Systems Manager Automation Runbook Reference"
[7]: https://learn.microsoft.com/en-us/azure/automation/manage-runbooks "Microsoft Learn — Manage runbooks in Azure Automation"
---
**来源**Manus AI《Agent Runbook 实用指南(v1.0)》,本仓库存档为规范参考。
@@ -1,13 +1,13 @@
# LAN 核心交换机升级计划(保留 ER-X)
**状态:** SE5420 **已采购**2026-08-09)。**实施与验证以 [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) 为准**
**状态:** SE5420 **已采购**2026-08-09)。**实施与验证以 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 为准**
本文为历史规划参考,**不得作为现场执行步骤**;所有实际操作均以部署指南为准。
**锁定硬件:** TP-Link **`TL-SE5420`**16 × 2.5GbE RJ45 + 4 × 10GbE SFP+)。
**目标:** SE5420 承接全部 LAN 物理接入与二层转发;ER-X 继续承担公网、NAT、防火墙、
LAN66/LAN55 网关与 DHCP。
**拓扑与流量的详细说明**(职责、逻辑网、流量路径、Wi-Fi 分工、验收边界)见:
[lan-erx-se5420-network.md](lan-erx-se5420-network.md)。
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)。
SE5420 官方资料:静态功耗 8 W、最大功耗 32 WVLAN、LACP、STP/RSTP/MSTP、ACL、
CLI/SNMP、配置导入导出与固件下载。无 PoE——AP 使用本地取电 + 普通网线。
@@ -65,7 +65,7 @@ VLAN tag。
```
完整端口表、流量路径与 Wi-Fi 分工见
[lan-erx-se5420-network.md](lan-erx-se5420-network.md)。
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)。
## VLAN 与端口设计
@@ -124,8 +124,8 @@ VLAN tag。
### 阶段 4VLAN10 升级专用 Wi-Fi(独立项目)
不与本次 Done 捆绑。仅 U6;网关 `gfw`;详见
[lan-erx-se5420-network.md](lan-erx-se5420-network.md) 第 6.5 / 8 节与
[unifi-network.md](unifi-network.md)。
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md) 第 6.5 / 8 节与
[unifi-network.md](../unifi-network.md)。
## 性能预期与不变瓶颈
@@ -148,9 +148,9 @@ VLAN tag。
## 参考
- [ER-X + SE5420 网络与拓扑说明](lan-erx-se5420-network.md)
- [LAN 概览](lan-overview.md)
- [ER-X 配置记录](edgerouter-x-configuration.md)
- [UniFi 网络](unifi-network.md)
- [`gfw`](../hosts/gfw.windy.lan.md)
- [ER-X + SE5420 网络与拓扑说明](../lan-erx-se5420-network.md)
- [LAN 概览](../lan-overview.md)
- [ER-X 配置记录](../edgerouter-x-configuration.md)
- [UniFi 网络](../unifi-network.md)
- [`gfw`](../../hosts/gfw.windy.lan.md)
- [TL-SE5420 官方规格](https://www.tp-link.com.cn/product_2899.html?v=specification)
@@ -3,7 +3,7 @@
**Status: research only. No configuration was changed.** This page evaluates
resolvers/splitters that are genuinely better than — or meaningfully different
from — the current "AdGuard Home (AGH) + mosdns" setup on
[`dns.windy.lan`](../hosts/dns.windy.lan.md) (`.36`), for a GFW-constrained
[`dns.windy.lan`](../../hosts/dns.windy.lan.md) (`.36`), for a GFW-constrained
China home LAN. Claims are cited to primary sources (official repos, official
docs, upstream READMEs); anything not verified is flagged as such.
@@ -11,11 +11,11 @@ docs, upstream READMEs); anything not verified is flagged as such.
> verification — mosdns on `.1` is **not idle**, it is clash's
> `nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution); the
> canonical decision record is
> [`lan-dns-architecture.md`](lan-dns-architecture.md) (final verdict aligned,
> [`lan-dns-architecture.md`](../lan-dns-architecture.md) (final verdict aligned,
> Phase 0 kill-test evidence incl. a measured upstream-blackhole degradation
> gap).
Scope recap (from [`lan-overview.md`](lan-overview.md), verified 2026-08-06):
Scope recap (from [`lan-overview.md`](../lan-overview.md), verified 2026-08-06):
- Clients get DNS via EdgeRouter DHCP option 6 → AGH `192.168.66.36:53`.
- AGH upstreams: `dns.alidns.com` + `doh.pub` DoH (load-balanced), fallback
@@ -58,7 +58,7 @@ The genuinely worthwhile changes, in order of value:
3. **mosdns on `.1` is resolved, not idle** — it is clash's
`nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution, verified
2026-08-12), so "delete it" is off the table; its role is documented in
[`lan-dns-architecture.md`](lan-dns-architecture.md) §1. If a future change
[`lan-dns-architecture.md`](../lan-dns-architecture.md) §1. If a future change
moves this role to an AGH-side companion, keep in mind mosdns's cache strips
EDNS0 and it performs no DNSSEC validation
([mosdns v5 executable plugins](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)).
@@ -517,18 +517,18 @@ AGH .36 (filtering, rewrites, query log, per-client upstreams)
TPROXY/fake-ip/DNS-hijack rules; co-locating LAN DNS there couples DNS to the
proxy and its restart/update lifecycle.
- A standalone resolver VM adds nothing: both current VMs already sit on the
same PVE hypervisor ([lan-overview.md](lan-overview.md) §Positioning facts),
same PVE hypervisor ([lan-overview.md](../lan-overview.md) §Positioning facts),
so a hypervisor outage takes out either placement equally; a second physical
host for HA is out of scope for a home LAN.
- If you ever run a validating resolver + AGH on `.36`, verify outbound from
`.36` to the foreign upstreams is not re-hijacked by OpenClash (loop check
already mandated in the [AGH review](adguard-home-official-review-2026-08.md)).
already mandated in the [AGH review](../adguard-home-official-review-2026-08.md)).
### 4.4 Fail-open, DNSSEC, private names — by candidate
| Concern | How the recommended stack behaves |
|---|---|
| Fail-open when proxy/subscription down | AGH forwards directly to DoH upstreams; `.36`'s outbound is not forced through the proxy in normal ops (no TUN policy routing on `.36` — [dns host facts](../hosts/dns.windy.lan.md)). With unbound/blocky behind, foreign resolution recurses/validates directly, independent of OpenClash. Avoid mihomo-DNS-as-resolver, which is proxy-coupled. |
| Fail-open when proxy/subscription down | AGH forwards directly to DoH upstreams; `.36`'s outbound is not forced through the proxy in normal ops (no TUN policy routing on `.36` — [dns host facts](../../hosts/dns.windy.lan.md)). With unbound/blocky behind, foreign resolution recurses/validates directly, independent of OpenClash. Avoid mihomo-DNS-as-resolver, which is proxy-coupled. |
| DNSSEC validation | Only unbound, blocky, knot-resolver, Technitium validate in-process. AGH sets DO only; mosdns/smartdns/chinadns-ng/mihomo/sing-box do not. Plan: validate behind AGH, or accept "validating public upstream" (confirm with `dig +dnssec`/known-bad test). |
| Private names / rewrites | AGH `rewrites` (already in use for `hass.windy.lan`) + `local_ptr_upstreams` once a local PTR source exists. blocky: `customDNS` mapping/rewrite + hosts. unbound: `local-zone`. All adequate. |
| Query log / visibility | AGH is the best at this of everything evaluated (14-day anonymized log already configured). |
@@ -551,7 +551,7 @@ Concrete, in increasing effort:
the known-bad-signature check, then enable AGH DNSSEC. Without this, AGH's
`enable_dnssec` is only a DO-flag — the exact reason it is currently off
([AGH DNSSEC semantics](https://adguard-dns.io/kb/adguard-home/configuration/),
[host facts](../hosts/dns.windy.lan.md)).
[host facts](../../hosts/dns.windy.lan.md)).
3. **Either fully configure mosdns (systemd service, sequence, lists) or remove
it.** Leaving an idle `127.0.0.1:6052` listener documented as "not the active
path" is drift. If kept, plan around no-EDNS0 cache + no validation; if
@@ -559,9 +559,9 @@ Concrete, in increasing effort:
4. **Close the private-PTR gap**: once a local authoritative source exists (e.g.
dnsmasq on `gw`, or a tiny authoritative zone), point AGH
`local_ptr_upstreams` at it as the AGH review recommends
([AGH review](adguard-home-official-review-2026-08.md));
([AGH review](../adguard-home-official-review-2026-08.md));
don't set it before that source exists
([dns host facts](../hosts/dns.windy.lan.md)).
([dns host facts](../../hosts/dns.windy.lan.md)).
5. **Optional: ECS** for CDN geo-accuracy — AGH `edns_client_subnet.use_custom`
with a coarse fixed prefix (or blocky `ecs.forward`) if measurements show a
benefit; note many CN resolvers ignore ECS
@@ -579,7 +579,7 @@ here beats it on that axis for this LAN.
- **Live behavior not tested**: all capability claims are from primary docs
reviewed 2026-08-12; DNSSEC behavior of `dns.alidns.com`/`doh.pub`/the
`adg.chans.xyz` path and mosdns's actual version on `.36` need on-box
`dig +dnssec` verification (per [adguard-home-health](../runbooks/adguard-home-health.md)).
`dig +dnssec` verification (per [adguard-home-health](../../runbooks/adguard-home-health.md)).
- **smartdns DNSSEC**: the official config reference lists no DNSSEC option;
if a newer version added one, it is not reflected here
([config options](https://pymumu.github.io/smartdns/configuration/)).
@@ -598,8 +598,8 @@ here beats it on that axis for this LAN.
## Related docs
- [lan-overview.md](lan-overview.md) — full topology (verified 2026-08-06)
- [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) — AGH host facts
- [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) — OpenClash facts
- [adguard-home-official-review-2026-08.md](adguard-home-official-review-2026-08.md) — prior AGH config review
- [runbooks/adguard-home-health.md](../runbooks/adguard-home-health.md)
- [lan-overview.md](../lan-overview.md) — full topology (verified 2026-08-06)
- [hosts/dns.windy.lan.md](../../hosts/dns.windy.lan.md) — AGH host facts
- [hosts/gfw.windy.lan.md](../../hosts/gfw.windy.lan.md) — OpenClash facts
- [adguard-home-official-review-2026-08.md](../adguard-home-official-review-2026-08.md) — prior AGH config review
- [runbooks/adguard-home-health.md](../../runbooks/adguard-home-health.md)
@@ -2,7 +2,7 @@
**状态:** 规划文档(未采购、未接线、未改生产配置)。
**重要变更(2026-08-09):** **SE5420 已采购**,网络升级改为「保留 ER-X + SE5420 核心」路径——
实施与验证以 [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) 为准。
实施与验证以 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 为准。
本文保留为「ER-X 网关未来替换为 RB5009」的备选方案;其中 PVE 透传调研与 VLAN10 实现方法仍适用。
---
@@ -389,9 +389,9 @@ logread -e netifd
## 8. 参考
- 现网地图:[lan-overview.md](lan-overview.md)
- ER-X 现状:[edgerouter-x-configuration.md](edgerouter-x-configuration.md)、[hosts/gw.md](../hosts/gw.md)
- UniFi[unifi-network.md](unifi-network.md)、[hosts/ubnt.md](../hosts/ubnt.md)
- `gfw`[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md)
- 作废方案(**不再实施**):[lan-erx-se5420-network.md](lan-erx-se5420-network.md)、[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
- 现网地图:[lan-overview.md](../lan-overview.md)
- ER-X 现状:[edgerouter-x-configuration.md](../edgerouter-x-configuration.md)、[hosts/gw.md](../../hosts/gw.md)
- UniFi[unifi-network.md](../unifi-network.md)、[hosts/ubnt.md](../../hosts/ubnt.md)
- `gfw`[hosts/gfw.windy.lan.md](../../hosts/gfw.windy.lan.md)
- 作废方案(**不再实施**):[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)、[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
- MikroTik RB5009 官方:<https://mikrotik.com/product/rb5009ug_s_in>、RouterOS v7 手册
@@ -1,6 +1,6 @@
# SE5420 实施评审主张核实(2026-08-10)
> **核对基准(历史快照):** 本文于 2026-08-10 针对 [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) 的**评审前版本**`35577d0`)撰写。该指南自 `ffb37a9`"finalize SE5420 deployment guide per review")起已按本评审修订,当前 `origin/main` 章节已重组:旧 §3.3 → §4.3、旧 §6(gfw)→ §11、旧 §7(SSID)→ §12、旧 §9(验收/IPv6)→ §13 + §11.4。文末「当前指南处理情况」列出各主张的现行状态;实施以部署指南现行为准。
> **核对基准(历史快照):** 本文于 2026-08-10 针对 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 的**评审前版本**`35577d0`)撰写。该指南自 `ffb37a9`"finalize SE5420 deployment guide per review")起已按本评审修订,当前 `origin/main` 章节已重组:旧 §3.3 → §4.3、旧 §6(gfw)→ §11、旧 §7(SSID)→ §12、旧 §9(验收/IPv6)→ §13 + §11.4。文末「当前指南处理情况」列出各主张的现行状态;实施以部署指南现行为准。
**范围。** 本文核对对 `lan-se5420-deployment-guide.md` 的评审意见。结论分为
“已证实”(规范/一手资料直接支持)、“基本证实”(架构推论成立但仍须读取现场配置)和
+345
View File
@@ -0,0 +1,345 @@
# Home Assistant × Matrix integration
Reference for wiring the Home Assistant [Matrix integration](https://www.home-assistant.io/integrations/matrix)
to the self-hosted Matrix homeserver at [`synapse.chans.xyz`](../hosts/synapse.chans.xyz.md).
Deliberately contains no Matrix passwords, access tokens, or room encryption material.
> **Status (2026-08-15, W1N-139):** the built-in `matrix` integration has been
> **retired** on `hass.windy.lan` and replaced by the custom **`matrix_e2ee`**
> integration. The sections below on the built-in integration are kept for
> reference only. See [matrix_e2ee](#matrix-e2ee-custom-e2e-integration) for the
> active setup and [Device verification (SAS) model](#device-verification-sas-model)
> for how device trust works.
## Purpose
The integration lets Home Assistant send messages to Matrix rooms and react to
messages/reactions in Matrix rooms. "Reacting" is done by firing a
`matrix_command` event when one of the configured commands matches; automations
then trigger on that event. Sending is done through the `notify.matrix` platform
and the `matrix.send_message` / `matrix.react` actions.
## Environment mapping
| Integration setting | This deployment |
|---|---|
| `homeserver` | `https://synapse.chans.xyz` (client-server base URL) |
| `username` | full Matrix ID, e.g. `@ha_bot:chans.xyz` |
| `password` | MAS local-password account password (see below) |
| Room IDs / aliases | full forms with the identity domain, e.g. `!cUrbafjkfsMDVwdRDQ:chans.xyz` or `#room:chans.xyz` |
- Identity domain is `chans.xyz` (not `synapse.chans.xyz`); user IDs and room
aliases carry the `:chans.xyz` suffix.
- Authentication on this homeserver is MAS (Matrix Authentication Service) with
local-password accounts. The integration logs in with `m.login.password`
(username + password), so the bot account must be a local-password account —
same as the Hermes account documented in [`hermes-matrix.md`](hermes-matrix.md).
If MAS is later switched to OAuth2/OIDC-only (no legacy password login), the
integration's password login will stop working; keep that in mind before such
a change.
- Public registration is disabled. Create/reset the dedicated bot account via
MAS / Element Admin.
## Use a separate bot account (mandatory)
The docs are explicit: to prevent infinite loops when reacting to commands,
the integration **must** use a separate account from any account whose messages
it reacts to. Use a dedicated account such as `@ha_bot:chans.xyz`, not a human
account.
## configuration.yaml (example)
```yaml
# The Matrix integration
matrix:
homeserver: https://synapse.chans.xyz
username: "@ha_bot:chans.xyz"
password: supersecurepassword
rooms:
- "#hasstest:chans.xyz"
commands:
- word: my_command
name: my_command
```
After changing `configuration.yaml`, restart Home Assistant to apply the
changes. The integration then shows under **Settings → Devices & services**;
its entities are on the integration card and the Entities tab.
### Configuration variables
| Variable | Meaning |
|---|---|
| `username` | Full Matrix ID the bot logs in as, e.g. `@ha_bot:chans.xyz`. The `@` has a special YAML meaning, so always quote it. |
| `password` | The bot account's password (MAS local password). |
| `homeserver` | Full client-server URL of the homeserver. |
| `rooms` | Rooms the bot should join and listen in. List **all** rooms commands are to be received in, even if a command scopes itself to fewer rooms. Accepts internal room ID (`!…:chans.xyz`) or alias (`#room:chans.xyz`). |
| `commands` | Commands to listen for. Each fires a `matrix_command` event when triggered. |
### Command types
| Key | Triggers when |
|---|---|
| `word` | A message starts with `!<word>`. Arguments after the word are captured as a list in the event's `data`. |
| `expression` | A message matches the Python regexp. The regexp group dictionary is captured in the event's `data`. |
| `reaction` | A message is reacted to with the given emoji. |
| `name` | The command name, exposed as an attribute of the fired event. |
A command can be scoped to specific rooms with a per-command `rooms` list (the
room must still be listed under the top-level `rooms`).
## Event data
When a command triggers, a `matrix_command` event fires with:
- `name` — the command name.
- `data` — for `word` commands, a list of arguments (everything after the word,
split on spaces); for `expression` commands, the group dictionary of the
matching regexp.
- `event_id` — the received message's identifier.
- `thread_parent` — the root message ID of the thread; equals `event_id` when
the message is not inside a thread.
## Notifications (notify.matrix)
Deliver notifications from Home Assistant to a Matrix room (direct or group):
```yaml
notify:
- name: matrix_notify
platform: matrix
default_room: "#hasstest:chans.xyz"
```
- The target room must already exist; get its canonical ID from the room
settings dialog (`!<randomid>:chans.xyz`) or an alias (`#roomname:chans.xyz`).
Quote the room ID/alias in YAML to escape the `!` / `#` characters.
- The notifying account may need to be invited to the room, depending on room
policy.
Message formats (`data.format`): `text` (default) and `html`. Images can be
attached via `data.images` (list of file paths); files from outside allowed
folders require `homeassistant.allowlist_external_dirs` to list the source
folder.
Reply inside a thread by passing the root message ID into `data.thread_id`:
```yaml
action: notify.matrix_notify
data:
message: "Reply message goes here"
data:
thread_id: "{{ trigger.event.data.thread_parent }}"
```
## Actions
- `matrix.react` — send a reaction to a message in a Matrix room
(`reaction`, `room`, `message_id`).
- `matrix.send_message` — send a message to one or more Matrix rooms.
## Comprehensive example (adapted)
```yaml
matrix:
homeserver: https://synapse.chans.xyz
username: "@ha_bot:chans.xyz"
password: supersecurepassword
rooms:
- "#hasstest:chans.xyz"
- "#someothertest:chans.xyz"
commands:
- word: testword
name: testword
rooms:
- "#someothertest:chans.xyz"
- expression: "My name is (?P<name>.*)"
name: introduction
- reaction: 👍
name: thumbsup
notify:
- name: matrix_notify
platform: matrix
default_room: "#hasstest:chans.xyz"
automation:
- alias: "Respond to !testword"
triggers:
- trigger: event
event_type: matrix_command
event_data:
command: testword
actions:
- action: notify.matrix_notify
data:
message: "It looks like you wrote !testword"
```
## matrix_e2ee (custom E2E integration)
Custom integration [`windyboy/ha-matrix-e2ee`](https://github.com/windyboy/ha-matrix-e2ee),
release **v0.3.12** (Matrix activity events + push diagnostics), deployed on
`hass.windy.lan` 2026-08-20 (upgraded from v0.3.2, W1N-182/#34 emoji-wait
wizard fix; v0.3.9 brought the Connection health binary sensor, SAS/command
allowlist split, URL normalization and single-entry enforcement, W1N-156/W1N-190).
Runs a dedicated bot with a **persistent E2EE device identity**.
- Domain `matrix_e2ee`; Config Flow (UI) with YAML import migration, not in HACS. Does **not**
override the built-in `matrix` integration.
- Dependencies are declared **explicitly** in `manifest.json` to work around Home
Assistant's `is_installed` dropping the `[e2e]` extra (W1N-140):
`matrix-nio[e2e]==0.26.0` + `vodozemac` + `peewee` + `cachetools` + `atomicwrites`.
- **v0.2.0 migration:** YAML `matrix_e2ee:` block was auto-imported into a Config Entry
(`source: import`) on first startup, then removed. All settings now managed via
**Settings → Devices & Services → Matrix E2EE → Configure**.
See [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) for the deployed state.
### Services & events
- Services (all admin-only since v0.1.4):
- `send_message` (`message`, `room_id`)
- `start_verification` (`user_id`, `device_id`)
- `confirm_verification` (`transaction_id`)
- `cancel_verification` (`transaction_id`)
- `reauthenticate` (`password`) — soft-logout only
- `get_fingerprint` (no fields; returns bot's own `ed25519`/`curve25519` keys; added v0.1.3)
- `verify_device_by_fingerprint` (`user_id`, `device_id`, `ed25519`; added v0.1.3,
renamed from `verify_device` in v0.1.4; requires exact `ed25519` match)
- Events:
- `matrix_e2ee_command` (`room_id`, `sender`, `command`, `args` only —
never the raw body)
- `matrix_e2ee_error` (codes, no secrets)
- `matrix_e2ee_verification` (`stage`, `transaction_id`, `user_id`, `device_id`,
optional `emojis`, optional `expires_at`; `expires_at` added v0.1.3)
- `matrix_e2ee_fingerprint` (`user_id`, `device_id`, `ed25519`, `curve25519`
public keys only; added v0.1.3)
- `matrix_e2ee_message_received` (`room_id`, `sender`, `event_id`; added v0.3.12
activity events)
- `matrix_e2ee_verification_done` (`transaction_id`, `user_id`, `device_id`;
added v0.3.12)
- v0.3.12 also adds an `event.` platform entity (`Bot activity`,
`event_types: ["message", "command", "verification_done"]`) and a diagnostic
Connection binary sensor (`binary_sensor.*_connection`, CONNECTIVITY class).
- `notify.matrix_e2ee` is **not implemented** (upstream deferred) — notifications
must call `matrix_e2ee.send_message` (message + room_id).
- Commands fire Home Assistant events only; the integration never calls
`domain.service` itself. Map commands in automations.
- Encrypted rooms fail-closed on unverified devices.
- Since v0.1.4: `start_verification`, `confirm_verification`, `cancel_verification`,
`verify_device_by_fingerprint`, and `reauthenticate` are enforced as HA admin-only
via `async_register_admin_service`; non-admin users cannot call them.
### Storage & recovery
- `.storage/matrix_e2ee_session.json` (`user_id`, `device_id`, `access_token`,
`pickle_key`) and `.storage/matrix_e2ee_store/` (Olm/Megolm, device trust,
sync token). Both stay on the HA persistent volume and are in HA backups.
- Soft logout → `matrix_e2ee.reauthenticate` (keeps `device_id` + crypto store;
rejected outside soft-logout state since v0.1.3).
- Hard logout / store loss → delete session + store, restart with password, re-SAS
(a **new device**; old history not decryptable).
## Device verification (SAS + fingerprint) model
Researched 2026-08-15 (W1N-139 stage-6 pre-study), updated for v0.1.3/v0.1.4.
Sources: matrix.org
[cross-signing guide](https://matrix.org/docs/guides/implementing-more-advanced-e-2-ee-features-such-as-cross-signing/),
matrix-nio [examples](https://matrix-nio.readthedocs.io/en/latest/examples.html),
[element-android#6832](https://github.com/vector-im/element-android/issues/6832),
Element [device-verification](https://element.io/features/device-verification).
`matrix_e2ee` supports three verification paths (the wizard — v0.3.0
bot-initiated, reworked in v0.3.1/v0.3.2 to wait for a peer-initiated inbound
SAS from the user's Matrix client with emoji comparison — automates the SAS
flow):
### 1. SAS (mutual, manual confirmation since v0.1.4)
- SAS is device-to-device: exchange ephemeral keys → derive emojis → **a human on
each side compares and confirms** (`m.key.verification.mac`).
- Matrix distinguishes two cases (spec uses *should*, not *must*):
- **same user, two devices** → to-device messages (SAS);
- **two different users** → **in-room (DM) messages**, verifying the *user*
(cross-signing master key), not a specific device.
- Cross-signing: each user has master / self-signing / user-signing keys. A device
looks "verified" to another user via the chain
`my master → my user-signing → their master → their self-signing → their device`.
- Element's "Verify" button only starts **in-DM user verification**; it has no
"verify a specific device of another user via to-device" flow (matrix.org
recommends hiding per-device verification for other users).
- `matrix_e2ee` implements **raw to-device device SAS** (`start_verification`/
`confirm_verification`), **no cross-signing / in-room**. This is a non-standard
cross-user path: works with matrix-nio + Element Web/Desktop (reported in
element-android#6832), **not** on Element Android/X.
- **v0.1.3**: inbound SAS auto-complete was added; SAS events include `expires_at`.
- **v0.1.4 (breaking)**: auto-confirm was removed. **Every** device — including
another device of the bot's own account — requires explicit `confirm_verification`
after emoji comparison. Only the bot's own account or users in `allowed_users`
may initiate SAS (`verification_peer_denied` otherwise).
- **v0.2.1**: storage I/O moved off the event loop (`asyncio.to_thread`,
W1N-167); own-keys query on startup so inbound SAS can build a session (W1N-166).
- **v0.2.2** (not deployed): intermediate version.
- **v0.2.3**: sync loop runs as a background task (fixes bootstrap setup timeout,
W1N-168); SAS double-send of key and MAC fixed (W1N-169).
- **v0.2.6**: `_log_verification_state()` tracks SAS state transitions with
`async_write_ha_state` for diagnosis (W1N-174);
`_bridge_verification_request()` handles inbound
`m.key.verification.request``m.key.verification.ready` since nio lacks a
`request` framework (W1N-173).
- **v0.2.5**: bridge `m.key.verification.request``ready` (nio lacks
request framework, W1N-173).
- **v0.2.4**: `_patch_nio_sas_timeout()` works around nio 0.26.0
`_last_event_time` bug (SAS timed out at 60s regardless of activity — now uses
`_max_age` 5 min); `_repair_dropped_start()` recovers SAS `start` events nio
dropped when the peer device was unknown (W1N-170/W1N-172);
`VERIFICATION_TIMEOUT_SECONDS` 600→240 (fires before nio's `_max_age`).
- **v0.2.11**: `receive_mac_event` no longer overrides canceled state (W1N-179/#31).
- **v0.3.12**: Matrix activity events (`matrix_e2ee_message_received`,
`matrix_e2ee_verification_done`) + `event.` Bot activity entity + Connection
diagnostic binary sensor.
- **v0.3.9**: SAS driver gate split from the command allowlist — new
`verification_peer_users` option (W1N-156/#41); SAS/sync logs demoted
warning→info/debug (W1N-188/#38); Connection health binary sensor
(W1N-185/#40); URL normalization + single-entry enforcement (W1N-190/#42).
- **v0.3.8**: `m.key.verification.done` handshake completion for
request-based SAS (W1N-183/#35).
- **v0.3.2**: wizard waits for the inbound SAS to show emojis before moving
to the compare step (`_wait_for_inbound` requires `latest_sas_snapshot()` to
return `emojis`) — W1N-182/#34.
- **v0.3.1**: verification wizard now waits for a peer-initiated inbound SAS
(options flow no longer starts verification from the bot; `latest_sas_snapshot()`
skips verified/canceled transactions) — GitHub #33.
- **v0.3.0**: bot-initiated device verification wizard (W1N-180/#32).
- Inbound SAS is gated to `allowed_users` (v0.1.3); **since v0.3.9 (W1N-156)
the gate is the separate `verification_peer_users` allowlist**, which is
unset on hass.windy.lan — only the bot's own account may drive SAS until
`@zhiqiang:chans.xyz` is added there.
### 2. One-sided fingerprint (added v0.1.3, hardened v0.1.4)
- Call `matrix_e2ee.get_fingerprint` to get the bot's own `ed25519` device key
(read it from the `matrix_e2ee_fingerprint` event).
- In Element, open the bot user's sessions and use "Manually verify by text".
Compare the session key with the fingerprint.
- To trust another device from the bot's side, call
`matrix_e2ee.verify_device_by_fingerprint` with the peer's `user_id`, `device_id`,
and `ed25519` key. The match is exact (since v0.1.4's rename from `verify_device`).
Feed the **peer** key, not the bot's own key.
- This trusts from one side only; the peer still trusts the bot independently.
- Both `get_fingerprint` and `verify_device_by_fingerprint` are HA admin-only.
### Consequence
Whether `@zhiqiang`'s device can be verified depends on which Element client
they use. Open options recorded in W1N-139 (A: Web SAS test; B: upstream
in-room/cross-signing; C: unencrypted-room downgrade).
## References
- Home Assistant Matrix integration: <https://www.home-assistant.io/integrations/matrix>
- Matrix host facts: [`hosts/synapse.chans.xyz.md`](../hosts/synapse.chans.xyz.md)
- Matrix deployment and upstream index: [`matrix-upstream.md`](matrix-upstream.md)
- Hermes Agent Matrix channel (MAS local-password + access-token pattern): [`hermes-matrix.md`](hermes-matrix.md)
- HA host facts: [`hosts/hass.windy.lan.md`](../hosts/hass.windy.lan.md)
- HA maintenance runbook: [`runbooks/home-assistant-maintenance.md`](../runbooks/home-assistant-maintenance.md)
+1 -1
View File
@@ -307,7 +307,7 @@ VLAN10 / 升级 SSID / 客人 SSID **失败或未做,不否决**本次核心
## 13. 参考
- 实施阶段与清单:[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
- 实施阶段与清单:[lan-core-switch-upgrade-plan.md](archive/lan-core-switch-upgrade-plan.md)
- 现网地图:[lan-overview.md](lan-overview.md)
- ER-X[edgerouter-x-configuration.md](edgerouter-x-configuration.md)、[hosts/gw.md](../hosts/gw.md)
- UniFi / VLAN10 前置:[unifi-network.md](unifi-network.md)
+60 -10
View File
@@ -13,6 +13,11 @@ from each section below.
> **Verified live on 2026-08-06** by read-only SSH from the WSL client. No
> changes were made. `gfw.windy.lan` root SSH was re-verified the same day after
> the key was installed; its facts below are from the fresh probe.
>
> **IPv6 re-verified 2026-08-20** (read-only): UniFi controller `Default`
> network IPv6 enabled (SLAAC/RA), both APs hold global SLAAC addresses, and
> `zhiqiangf` key-only AP SSH re-confirmed. See
> [unifi-network.md](unifi-network.md).
---
@@ -35,10 +40,14 @@ from each section below.
│ ubnt — UniFi Network Controller (192.168.66.46)
```
> **SE5420 purchased (2026-08-09):** TP-Link `TL-SE5420` acquired; deployment plan is
> **SE5420 live (2026-08-22):** TP-Link `TL-SE5420` (purchased 2026-08-09) is
> online — management `192.168.66.253` reachable, web UI on :80/:443; LAN55
> 上联为 ER-X `switch0` **单口**`eth1` up、`eth2`/`eth3` down2026-08-22
> 只读核实)→ `switch0` 不再是 LAN55 全量抓包点(同段有线单播在 SE5420 本地
> 交换),全量点只能靠 SE5420 port mirroring。迁移状态见部署计划
> [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md). Design/planning refs:
> [lan-erx-se5420-network.md](lan-erx-se5420-network.md),
> [lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md).
> [lan-core-switch-upgrade-plan.md](archive/lan-core-switch-upgrade-plan.md).
---
@@ -47,19 +56,20 @@ from each section below.
| Host | Role | SSH | IPv4 | Facts |
|------|------|-----|------|-------|
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | [hosts/gw.md](../hosts/gw.md) |
| **PVE** | Proxmox host (`.66.26`/vmbr0 · `.55.26`/vmbr1) — hosts gfw/dns/ubnt/haos VMs | `ssh -4 root@192.168.66.26` | `192.168.66.26` | — |
| **PVE** | Proxmox host (`.66.26`/vmbr0 · `.55.26`/vmbr1) — hosts gfw/dns/ubnt VMs | `ssh -4 root@192.168.66.26` | `192.168.66.26` | — |
| **gfw.windy.lan** | OpenWrt LAN gateway / OpenClash — **PVE VM 140** | `ssh -4 root@192.168.66.1` | `192.168.66.1` | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy — **PVE VM 120** (`pihole`) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
| **ubnt** | UniFi Network Controller — **PVE VM 160** | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | [hosts/ubnt.md](../hosts/ubnt.md) |
| **hass.windy.lan** | Home Assistant (HAOS) — **PVE VM 180** (LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
| **hass.windy.lan** | Home Assistant (HAOS) — **x88 Pro physical box** (LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
| **pgdb** | TimescaleDB PG18 (Docker) — HA recorder 后端 — **PVE VM** (LAN55) | `ssh -4 windy@192.168.55.15` | `192.168.55.15` | [hosts/pgdb.md](../hosts/pgdb.md) |
| **NAS/FreeNAS** | NAS; `transmission` jail runs here (`.51`) | — | — | — |
| **U6 Lite** | UniFi AP (LAN66) | `ssh -4 zhiqiangf@192.168.66.6` | `192.168.66.6` | [docs/unifi-network.md](../docs/unifi-network.md) |
| **UAP-AC-Lite** | UniFi AP (LAN55) | `ssh -4 zhiqiangf@192.168.55.5` | `192.168.55.5` | [docs/unifi-network.md](../docs/unifi-network.md) |
> **Positioning facts (verified 2026-08-09):** `dns`/`ubnt`/`gfw`/`haos` are all VMs on PVE
> (no separate physical hosts); `transmission` is a FreeNAS/NAS jail. Only gw, PVE,
> NAS, U6, UAP-AC-Lite, and wired PCs/NAS are physical SE5420 ports. See
> [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) §1.
> **Positioning facts:** `dns`/`ubnt`/`gfw`/`pgdb` are VMs on PVE; `haos` is a **physical x88 Pro
> box** (HAOS bare-metal, `machine: green`), not a PVE VM (corrected 2026-08-15).
> `transmission` is a FreeNAS/NAS jail. Physical SE5420 ports: gw, PVE, haos, NAS,
> U6, UAP-AC-Lite, and wired PCs. See [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) §1.
---
@@ -76,7 +86,10 @@ from each section below.
| Port-forwards | `hass`→192.168.55.11:8123 · `transmission`→192.168.66.51:51413 · `ssh`→192.168.66.36:22 (orig 5822) · `openvpn`→192.168.66.32:1194 · WAN iface pppoe0 |
| Management | SSH TCP 22 · EdgeOS GUI HTTP 80 / HTTPS 443 |
**Static DHCP mappings (LAN66):** `OnePlus-12`=.37, `gfw`=.1, `hp-nas`=.32, `pihole`=.36, `pve`=.26, `transmission`=.51, `ubnt-6`=.6, `ubnt-app`=.46, `windy-pc`=.99. LAN55: `Aqara-Hub-M3-10CB`=.248.
**Static DHCP mappings (LAN66):** `OnePlus-12`=.37, `gfw`=.1, `hp-nas`=.32, `pihole`=.36, `pve`=.26, `transmission`=.51, `ubnt-6`=.6, `ubnt-app`=.46, `windy-pc`=.99. LAN55: `Aqara-Hub-M3-10CB`=.248, `SmartThings-Station`=.48, `espressif`=.47,
`hass`=.11, `hass-wifi`=.250, `ihost`=.12, `midea_ac_0418`=.10,
`midea_e3_0198`=.42, `roborock-wm-a141`=.43, `samsung-hub`=.251,
`matter`=.41 (added 2026-08-20).
> **Note:** `LAN_IN`/`LAN_OUT` are defined but not applied to an interface, so LAN55
> and LAN66 are bidirectionally reachable by default. Do not rely on those rules as
@@ -154,7 +167,7 @@ See [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md).
| SSH | `ssh hassio@hass.windy.lan` (key-only, verified 2026-08-13) |
| Web UI | `http://hass.windy.lan:8123` |
| WAN | gw port-forward `hass``192.168.55.11:8123` |
| Platform | HAOS; kernel `6.1.115-haos` (aarch64) |
| Platform | HAOS on physical x88 Pro box; kernel `6.1.115-haos` (aarch64), `machine: green` |
---
@@ -168,6 +181,43 @@ See [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md).
Both reported **Connected** to `http://192.168.66.46:9080/inform` on 2026-08-06.
AP SSH account is `zhiqiangf` (key-only, verified). See [docs/unifi-network.md](../docs/unifi-network.md).
**IPv6 (verified 2026-08-20):** both APs hold global SLAAC IPv6 addresses on
`br0` — U6 Lite `240e:3bd:235:1fb1::/64` (LAN66), UAP-AC-Lite
`240e:3bd:235:1fb2::/64` (LAN55) — with RA default routes via `gw`; the
controller's `Default` network has IPv6 enabled (SLAAC). Prefixes are dynamic
(PPPoE PD), so they rotate on redial. Details:
[docs/unifi-network.md](../docs/unifi-network.md).
**SSID cleanup (2026-08-21, W1N-207):** the SmartThings Element/vWire provisioning
SSIDs (`element-8a0d5133c9438f12`, `vwire-8b2d67469e455785`, `vport-F09FC22004E9`)
were removed/disabled in the controller (`element_adopt` setting off, element wlanconf
deleted, connectivity `x_mesh_essid`/`x_mesh_psk` cleared, device `x_vwirekey` removed,
`vwire_enabled`/`mesh_sta_vap_enabled=false`) and cleared from both APs; all
vwire/vport/element flags on the remaining SSIDs are now `disabled`.
**Stable ULA on gw: not feasible (2026-08-21, W1N-207):** EdgeOS v3.0.1
`interfaces switch switch0` rejects a static `ipv6 address`, and an explicit
`router-advert` node *replaces* the DHCPv6-PD-slaac RA (drops the delegated GUA
prefix from radvd → LAN55 loses IPv6 egress after RA expiry). Attempted and rolled
back cleanly (no `save`; gw config unchanged). Consequence: after a PD rotation,
restart HA's matter-server (see [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md))
to clear stale IPv6 mDNS caches.
**LAN55 RA environment (observed 2026-08-21):** besides `gw`, the SmartThings
Station (.48) and Aqara M3 (.248) act as Thread border routers and advertise ULA
prefixes (`fd00:5a7:6415:1::/64`, `fd97:d580:16fe:1::/64`); several LAN55 hosts
(HA, PVE, UAP-AC-Lite) have IPv6 forwarding enabled and mark themselves as
routers in NDP. This is normal Thread-BDR behaviour and was not the Matter
failure cause.
**Matter 灯泡(2026-08-21 实测,W1N-207):** 两盏 ESP32-C2 Matter 灯泡
VP `0x4891/0x4100`OUI `34:98:7a`)——工作盏 MAC `34:98:7a:25:a1:f0`;故障盏
MAC `34:98:7a:27:7f:08`hostname `matter`,动态 .145)。故障盏已在 Aqara fabric
`4DF2B1455D19402D` 内、宣告 `CM=0`(不在配对模式)且缺 GUA → 找回需**恢复出厂**
后扫它自己的二维码。DHCP 保留 `matter`.45 → MAC `34:98:7a:27:10:bc`)与故障盏
MAC 不符,保留从未租出(待修,见 [hosts/gw.md](../hosts/gw.md))。完整排障知识:
[docs/matter-pairing-troubleshoot.md](matter-pairing-troubleshoot.md)。
---
## Quick orientation (who runs what)
+6 -6
View File
@@ -106,7 +106,7 @@
- 第一步不建 VLAN10、不向 ER-X 送任何 tag、口 4/6PVE、U6)不做 trunk。
- NAS 只接口 8,口 12 断开(LACP 是独立维护窗)。
- 不占口的 VMdns(.36=VM120)、ubnt(.46=VM160)、gfw(.1=VM140)haos(.55.11=VM180)transmission(.51) 是 NAS jail。
- 不占口的 VMdns(.36=VM120)、ubnt(.46=VM160)、gfw(.1=VM140)haos(.55.11) 是物理 x88 Pro 盒子(非 VM)transmission(.51) 是 NAS jail。
## 3. 开箱与固件升级
@@ -207,7 +207,7 @@
4. **验证:**
- `ip -br addr``vmbr1` = `192.168.55.26/24`
- `ping -c3 192.168.55.254` → 通。
5. 逐台验证 VM(顺序:gfw → dns → ubnt → haos):
5. 逐台验证(顺序:gfw → dns → ubnt → haos;前三个是 VMhaos 是物理盒子):
```bash
ssh -4 root@192.168.66.26 'qm list'
```
@@ -215,7 +215,7 @@
- dns`ping -c3 192.168.66.36` → 通;
- ubnt`ping -c3 192.168.66.46` → 通;
- haos`ping -c3 192.168.55.11` → 通(注意是 55 网段)。
6. 每个 VM 再验业务:gfw 的 OpenClash 面板/DNS 正常、dns 的 AdGuard UI 能开、ubnt 控制器 Connected、haos 界面能开。不以"宿主开机"代替。
6. 每再验业务:gfw 的 OpenClash 面板/DNS 正常、dns 的 AdGuard UI 能开、ubnt 控制器 Connected、haos 界面能开。不以"宿主开机"代替。
## 7. 迁移 AP 与接入设备
@@ -462,7 +462,7 @@ ssh -4 root@192.168.66.1 'uci show network; uci show firewall; uci show dhcp; ip
- 每次实质变更后在 Linear `vps` 项目记录 scope / action / verification / 遗留 follow-up。
- 本仓库不记录 SE5420 口令、ER-X 配置快照(含 PPPoE/口令)、gfw 凭据。
- 实施前先读 `se5420-review-claim-verification-2026-08.md` 的现场只读复核结论。
- 实施前先读 `archive/se5420-review-claim-verification-2026-08.md` 的现场只读复核结论。
## 16. 回滚
@@ -486,10 +486,10 @@ ssh -4 root@192.168.66.1 'uci show network; uci show firewall; uci show dhcp; ip
## 参考
- 设计说明:[lan-erx-se5420-network.md](lan-erx-se5420-network.md)
- 评审核实:[se5420-review-claim-verification-2026-08.md](se5420-review-claim-verification-2026-08.md)
- 评审核实:[se5420-review-claim-verification-2026-08.md](archive/se5420-review-claim-verification-2026-08.md)
- 现网地图:[lan-overview.md](lan-overview.md)
- 官方安装手册(Markdown 版):[se5420-official-manuals/tl-se5420-install-manual.md](se5420-official-manuals/tl-se5420-install-manual.md)
- 官方 PDF<https://service.tp-link.com.cn/download/202310/TL-SE5420%20V1.0安装手册%201.0.2.pdf>
- 规格 / 固件:<https://www.tp-link.com.cn/product_2899.html?v=specification> · <https://www.tp-link.com.cn/product_2899.html?v=download>
- Omada VLAN 指南:<https://support.omadanetworks.com/en/document/12981/> · <https://support.omadanetworks.com/en/document/13135/>
- ER-X[edgerouter-x-configuration.md](edgerouter-x-configuration.md)gfw[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md)PVE VLAN10[lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09](lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09)
- ER-X[edgerouter-x-configuration.md](edgerouter-x-configuration.md)gfw[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md)PVE VLAN10[lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09](archive/lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09)
@@ -1,67 +0,0 @@
# 低打印量黑白激光一体机:采购决策树(2026-08)
## 结论
对于“打印量很小、20 页/分钟足够、不需要 ADF”的需求,不应为 34 页/分钟、250 页纸盒或双面 ADF 付费。它们解决的是连续文档处理,而不是偶尔打印。优先级应为:**不被耗材/账号绑定、稳定的局域网打印、平板扫描、紧凑尺寸和可获得的原装耗材**。
首选是 **Brother DCP-L1638W 或 DCP-L1848W**:两者都是传统鼓粉分离路线,20 ppm、150 页纸盒、百兆有线网口及 2.4/5 GHz Wi-Fi;官方资料不能证实 1848 相比 1638 有实质功能升级,因此按正规渠道的含税到手价和保修选较便宜、在售的一台即可。[L1638W 官方参数表](https://www.brother.cn/-/media/ap/cn/products/pdf-file/prt/ESL-done/DCP-L1628L1638W.ashx)[L1848W 官方参数表](https://www.brother.cn/-/media/ap/cn/products/pdf-file/prt/ESL_DCP-L1848W.ashx)
不把 Brother DCP-11W 作为默认首选。它是 Brother 当前标为新品的同级机器,却是“云充”按页模式:激活后含 700 页,页数用完要通过绑定的微信账户购买套餐才能继续打印;粉仓和硒鼓由官方免费提供。这适合愿意换取三年保修和明确按页预算的人,不适合希望长期离线、自主选择耗材的人。[新品发布](https://www.brother.cn/info/news/20250613)[云充规则](https://www.brother.cn/minisite/sppackage/esl/)
## 先按需求分流,而不是按品牌或 ppm
```text
需要批量扫描/复印多页原稿?
├─ 是 → 本指南不适用;选择带 ADF 的 L2648DW 等级,双面原稿高频则看双 CIS 机型。
└─ 否
└─ 需要自动双面打印?
├─ 是 → 选择 B7628DW 等带 duplex 的型号;这是功能升级,不是速度升级。
└─ 否
└─ 必须有有线网口或 5 GHz Wi-Fi
├─ 是 → Brother L1638W/L1848W 为基准选择。
└─ 否 / 接受仅 2.4 GHz Wi-Fi
└─ 是否接受按页充值并绑定微信?
├─ 是 → Brother DCP-11W(先比较套餐与传统耗材总价)。
└─ 否 → L1638W/L1848W;或比较下列 Canon/HP/Pantum。
```
### 采购前的一票否决项
- 需要 macOS、iPhone/iPad:确认 AirPrint;不要假定“Wi-Fi”就等于无需驱动。
- 打印机准备接交换机:确认具体 SKU 有 **Ethernet**,不能把 Wi-Fi Direct 当作局域网网口。HP MFP 1188w 的中国规格仅列 USB 和 2.4 GHz Wi-Fi,不列有线网口。
- 偶尔使用也建议选激光,但纸张要长期放在干燥、封闭处;低使用量时,受潮纸和旧粉盒造成的底灰、掉粉或卡纸比 20/22 ppm 差异更常见。
- 需要自动双面打印、ADF、双面扫描时,直接跳级;入门平板机硬凑这些功能没有性价比。
## 候选机型:仅保留与需求相符者
| 机型 | 联网 / 移动打印 | 核心纸路与扫描 | 耗材结构 | 面向本需求的判断 |
|---|---|---|---|---|
| **Brother DCP-L1638W / L1848W** | USB、100M Ethernet、2.4/5 GHz Wi-FiAirPrint、Mopria、Wireless Direct、Brother Mobile Connect | 20 ppm、150 页进/50 页出、平板 CIS;无 ADF、无自动双面 | TN118 约 1,500 页 + DR118 约 10,000 页,鼓粉分离 | **首选**:功能刚好够用,网络规格最好,自主耗材路径最清晰。两台按价格/现货择一。 |
| **Brother DCP-11W** | USB、100M Ethernet、2.4/5 GHz Wi-Fi / Wi-Fi Direct | 20 ppm、150 页纸盒、平板扫描;未列 ADF 或自动双面 | 云充按页;耗材由官方供给 | 只有明确接受微信充值、看重三年保修时才选;不是“便宜传统激光机”。 |
| **Pantum M6509NW** | USB、100M Ethernet、2.4 GHz 802.11b/g/n Wi-Fi;自带热点 | 22 ppm、150 页进/100 页出、平板扫描;手动双面 | 鼓粉一体 PD-219,官方标称 1,600 页 | **价格明显更低时的可比替代**。接口齐全,支持扫 PC/邮件/FTP/移动端,但机身较宽、仅 2.4 GHz,且鼓粉一体意味着每次换耗材同时更换感光组件。 |
| **HP Laser MFP 1188w** | USB、2.4 GHz 802.11b/g/n、Wi-Fi Direct、AirPrint/Mopria/HP 应用;**无 Ethernet** | 22 ppm、150 页进/100 页出、平板扫描;手动双面 | 一体式黑色硒鼓;随机约 1,500 页 | 仅无线且 2.4 GHz 能满足时再比价。优点是 AirPrint/Mopria 明确、首张页快;不符合“网口或双频”优先条件。 |
| **Canon iC MF232w** | Ethernet、2.4 GHz Wi-Fi、AirPrint、网络/移动扫描 | 23 ppm、250 页纸盒、平板扫描;无自动双面 | CRG337 一体式硒鼓 2,400 页 | 纸盒需求确实较大才考虑。官方建议价较高,且产品规格呈现的是较老的 IPv4 / 2.4 GHz 组合,不是本需求下的优先解。 |
| **Canon iC MF272dw** | Ethernet、2.4 GHz Wi-Fi、AirPrint/Mopria | 29 ppm、150 页纸盒、平板扫描、自动双面打印 | CRG071700 页随机、1,200/2,500 页商品硒鼓 | 功能不错但属于为自动双面打印升级;官方建议价 ¥3,838,不适合低量、单面为主时以性价比为目标的采购。 |
Brother 规格与耗材页数以官方参数表为准;Pantum 的接口、PD-219 和建议月印量 2502,000 页见[官方产品页](https://www.pantum.cn/product-center/1487019260672548865.html)。HP 的 22 ppm、150 页、仅 Wi-Fi/USB、手动双面和一年保修见[中国官方规格](https://support.hp.com/cn-zh/product/product-specs/hp/2101513893)。Canon MF232w 的 23 ppm、250 页、CRG337、IPv4 和接口见[官方规格](https://www.canon.com.cn/product/icmf232w/spec.html)MF272dw 的 29 ppm、自动双面、接口及 CRG071 页数见[官方规格](https://www.canon.com.cn/product/icmf272dw/spec.html)。
## 耗材与锁定:应怎样理解
**不要只用“每页成本”决定低量用户。** 一年只打印几十到几百页时,机器差价、过期/存放不当的耗材风险和购买便利性,通常超过高容量粉盒带来的单位页优势。页产量也是 ISO 覆盖率下的额定值,不等于实际能稳定打印的页数。
- **传统耗材(Brother L1638W/L1848W**:粉盒和硒鼓分开,硒鼓寿命远高于单盒粉量;这是长期低量使用中最可预测的结构。原装 TN118 / DR118 的料号和页数已由 Brother 公布。第三方粉盒或灌粉可以降低成本,但不属于厂商性能/保修承诺;低量用户省下的钱很有限,反而更容易把故障归因变复杂。建议首个生命周期使用原装或可靠授权渠道耗材。
- **云充(DCP-11W)**:这里的锁定不是“第三方粉盒风险”,而是服务依赖:打印资格、套餐和耗材供给都依赖绑定的微信/官方流程。购买前应把预计三年页数代入套餐,确认账号更换、迁移、停服或转让场景的处理规则;并接受双面一张按两页计。
- **一体式硒鼓(Pantum、HP、Canon)**:换粉即换鼓,维护动作简单;缺点是无法像鼓粉分离机那样只更换粉盒。不要据此推断“第三方一定不能用”或“必然会被固件锁死”——厂商公开资料通常只承诺原装耗材效果/保修,兼容耗材的芯片兼容性、质量和售后由销售方承担,应按批次验证。
## 最终推荐与购买动作
1. **默认买 Brother DCP-L1638W 或 DCP-L1848W**:选到手价更低、可开票、有本地退换/保修的那个;功能层面无需为 1848 付溢价。
2. 若二者断货或溢价过大,**Pantum M6509NW** 是功能不降级的对照品;要求 5 GHz Wi-Fi 时排除它。
3. 若只用手机/2.4 GHz Wi-Fi,且 HP 的即时价格有明显优势,才纳入 **HP 1188w**;它没有网口,不能接入现有有线网络。
4. **不要因为“最新”买 DCP-11W**,除非云充模式本身是主动选择。对低量家庭用户,耗材自主权通常比三年保修更重要。
到货后先完成一次有线或基础 Wi-Fi 配网、AirPrint/Windows/macOS 实测、扫描为 PDF、睡眠唤醒和一张双面手动测试;保留试机页与发票。将设备放在受信任 LAN;如果启用 Wi-Fi Direct,设置强口令,平时不需要则关闭。
## 调研边界
本表只比较中国市场仍可由厂商官方页面/支持页核实的代表 SKU,价格、实际库存和促销会实时变化,未把电商标价写入结论。所谓“最新”以 Brother 中国目录/公告为准,而非“功能最强”或“最适合”。资料核查日期:2026-08-11。
+1
View File
@@ -108,3 +108,4 @@ Steps:
- Matrix Authentication Service: <https://github.com/element-hq/matrix-authentication-service>
- Matrix spec: <https://spec.matrix.org/>
- Federation tester: <https://federationtester.matrix.org/>
- Home Assistant Matrix integration: [home-assistant-matrix.md](home-assistant-matrix.md)
+215
View File
@@ -0,0 +1,215 @@
# Matter 配网排障手册
> 基于 Matter 1.5.1 Core Spec §4.3.1 与本环境(EdgeRouter X + UniFi AP + Aqara M3 +
> Home Assistant2026-08-21 实测整理。配套 Linear W1N-207。
## 1. Matter 配网协议要点(发现即一切)
- **发现走 mDNSDNS-SD**UDP **5353**,组播 `224.0.0.251` / `ff02::fb`
**不经过单播 DNS(如 AdGuard .36)、不需要反向 DNS、不需要 DHCPv6**SLAAC 即满足 Matter
的 IPv6 要求)。
- 服务类型:
- `_matterc._udp` — 可配网设备(Commissionable),**配对模式才有效**
- `_matter._tcp` — 已配设备(Operational),TXT 里含 fabric 信息
- 子类型(配对方按此过滤):
- `_L<全12位 discriminator>`(如 `_L3266`)— 按二维码里的完整 discriminator 精确匹配
- `_S<高4位>`(如 `_S12`
- `_V<vendorId>``_T<deviceType>`(可选)
- `_CM`(仅真正处于配对模式时发布)
- TXT 关键键:`D=`discriminator,规范 **SHALL** 必填)、`VP=`vendor+product)、
**`CM=`**、`RI=`rotating id)、`PH=`/`PI=`(配对提示)。
- 配对端口:**TCP 5540**PASE/CASE)。部分生态(Aqara M3)为 Thread 中继节点用 **5552**
- 实例名:64 位随机 hex;**进入配对模式时更换**(可用作"是否重新进过配对"的信号)。
- 规范参考:[Matter 1.5.1 Core Spec §4.3.1](https://csa-iot.org/wp-content/uploads/2026/03/23-27349-010_Matter-1.5.1-Core-Specification.pdf)、
[Google Home: Commissionable and Operational Discovery](https://developers.home.google.com/matter/primer/commissionable-and-operational-discovery)、
[Matter Handbook: Discovery](https://handbook.buildwithmatter.com/how-it-works/discovery/)、
[connectedhomeip: IP commissioning](https://pigweed.googlesource.com/third_party/github/project-chip/connectedhomeip/+show/59edd2ff8506b1e3dabb7040d716f0e75a2312d1/docs/guides/ip_commissioning.md)。
## 2. 关键判据:CM=0 = 不在配对模式
规范 §4.3.1.2 / §4.3.1.7
- 设备可以长期宣告 `_matterc`**Extended Discovery**),但 **`CM=0` 表示"当前不接受配网"**。
- **已在 fabric 里的设备**(宣告里同时有 `_matter._tcp` + `_I<fabric>._sub` 运营记录)重配时
通常报 `CM=0` —— 它已配好,不是新设备。
- **配对方不能把已配设备当新设备加** → 重加/找回必须先**恢复出厂**(清 fabric,重启后以
`CM=1` 全新配对模式宣告),再用**它自己的二维码**添加。
- 常见误判:抓包看到 `_matterc` 宣告就以为"在配对模式"——**必须看 `CM=`**。
## 3. 本环境实测事实(2026-08-21W1N-207
| 事实 | 状态 |
|---|---|
| LAN55 IPv6/mDNS 链路 | ✅ 全正常(RA→交换机→AP→客户端;mDNS 双向通;igmp snooping off、mdns on、无客户端隔离、无组播增强、PMF off、WPA2、仅 2.4G |
| Matter 不依赖单播 DNS/.36、反向 DNS、DHCPv6 | ✅ 已排除(.36 健康且不在路径上) |
| HA matter-server 曾宣告两代前的旧 GUA | ✅ 已修复(重启 `core_matter_server`;宣告恢复当前前缀) |
| ISP PD /60 随重拨轮换 → Matter IPv6 缓存反复失效 | ⚠️ 环境性根因;对策 = 重拨后重启 matter-server + 重启 M3 |
| EdgeOS 上静态 ULA 不可行 | ✅ 已尝试并回滚(switch0 不支持静态 `ipv6 address`;显式 router-advert 会替换 PD-slaac RA |
| 在用的两盏 ESP32-C2 Matter 灯泡(VP `0x4891/0x4100`2026-08-23 复核) | 工作盏 MAC 已变为 `fc:e8:c0:25:a1:f0``.146`hostname `espressif`;原 `34:98:7a:25:a1:f0` 全网消失,疑固件更新后换 MAC——末 3 字节相同);新盏 `34:98:7a:27:10:bc``.148`hostname `matter`)。两盏各宣告 **3 个 fabric** 运营实例:Aqara `4DF2B1455D19402D``2F6E56020E1996E7`、HA `DCE86145C137AF0E`(见 §8 |
| 故障盏 `34:98:7a:27:7f:08`(曾 .145Aqara fabric`CM=0` 缺 GUA | 2026-08-23 复核:无租约、ARP incomplete、AP 无日志 = **已离网**(退役/退换) |
| **失败模式 C(2026-08-23 实测,两盏同时)**mDNS 活、5540 死 | 灯泡 ping 通(v4/v6)、DHCP 正常续租、mDNS 应答并宣告 `_matter._tcp`SRV :5540、TXT `T=1`、当前前缀 GUA),但 **TCP 5540 在 IPv4 与 IPv6fe80+GUA)均 RST 拒绝** → 配对方无法建立 CASEApp 显示离线;hass matter-server 侧无任何 established :5540 会话(详见 §8 |
| ISP PD 前缀再次轮换(2026-08-23 → `240e:3bd:238:4812::/64`08-22 为 `235:1fb2` | hass 与 `.148` 均持当前前缀 GUA;hass 残留 `.146` 旧前缀 GUA 的 **FAILED** 邻居项(旧地址缓存仍被某端尝试) |
| DHCP 保留 `matter`.45 → MAC `…10:bc` | ⚠️ 保留仍未生效:新灯泡(`…10:bc`)实际拿到动态 `.148` 而非保留的 `.45`(待修,见 hosts/gw.md |
| **新灯泡(2026-08-22 添加成功)**MAC `34:98:7a:27:10:bc`=DHCP 保留目标 MAC),hostname `matter`IP `.148`VP `4891/4100`D=`3377` | ✅ 已入 **Aqara fabric `4DF2B1455D19402D`****经 BLE 配网**Aqara Home App)——线上**无 TCP 5540** 属正常(BLE 会话对 AP/hass 抓包不可见) |
| ESP32-C2 灯泡 firmware 挂死模式(2026-08-22 实测) | 入网后宣告 `_matterc`CM=1、D=3377)约 **3 秒后网络栈完全静默**:STA 收发计数冻结、不掉线不重启、配对方(手机/M3 `_L3377` 查询)无应答 → 加不上。**对策=断电 10 秒重启**重新进配网模式(实例名更换:`3F4E2C66F2DA85CD``E5BA8E28E4DE23A0`),随即 App 添加即成功 |
| 遗留 SSIDelement/vwire/vport | ✅ 已清理 |
## 4. 抓包方法(BusyBox 兼容)
> 完整指令集(实时 / 落盘轮转 / 定向抓取 / Wireshark 解密)见
> [runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md)。
> 下面是最常用的两条。
**视角必须在 LAN55**。**HA matter-server 作配对方时推荐直接在 hass `end0` 抓**——配对方
必然参与配对流程的每一条通讯(mDNS 本段组播 + 自己的 TCP 5540 全程),覆盖最全;AP `br0`
能看到全部 mDNS 组播 + 无线客户端单播,但**看不到有线↔有线单播**(如 Thread 设备经有线 M3
配对时 HA↔M3 的 5540 在 AP 侧不可见)。66 网段电脑看不到 55 的组播。BusyBox 注意点仅适用
AP**不要用 `--line-buffered`**;引号外层双引号、内层单引号);hass 是 HAOS 全量 tcpdump。
完整抓取(跑配对时保持窗口开着,`Ctrl+C` 结束):
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
hass 侧(HA matter-server 作配对方,推荐;非交互 ssh 需显式 `sudo -n -i`):
```bash
ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
```
精简过滤(只看 Matter 信号):
```bash
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552' | grep -E '_matterc|_matter|_L[0-9]+|_S[0-9]+|_CM|_V[0-9]+|_T[0-9]+|\.5540|\.5552'"
```
存 pcap 供 Wireshark:把上面 `-w /tmp/matter.pcap` 追加到 tcpdump 参数(去掉 `-vvv`),
`scp zhiqiangf@192.168.55.5:/tmp/matter.pcap .` 拉回本地分析。
> **落盘务必轮转**AP `/tmp` 只有约 60MB。用
> `-C 5 -W 12 -w /tmp/matter.pcap`(每 5MB 轮转、最多 12 个文件)防止写满,
> 详见 runbook Step 3(落盘轮转)。
> **Matter 载荷是加密的**mDNS5353)明文可读;5540 上的 Matter 报文要看明文
> 需要 Wireshark matter-dissector + 会话密钥,详见 runbook Step 5(解密)。
### 阶段对照表
| 阶段 | 应该看到 | 对应问题 |
|---|---|---|
| 发现(设备侧) | `_matterc._udp` + `_L3266._sub` + `_S12._sub` + TXT `D=3266 CM=1` + SRV `:5540` + AAAA | **无宣告**=设备没入网/没进配对模式;**`CM=0`**=不在配对模式(已配设备);**无 `_L3266`**=固件子类型缺失 |
| 发现(配对方侧) | M3/手机查询 `_L3266._sub._matterc._udp` | 查询有、无应答 = 码/discriminator 不匹配或设备不在线 |
| 配对握手 | 到设备 IP **TCP 5540 SYN/SYN-ACK** 双向 | **SYN 无 ACK**=设备不可达/防火墙;**完全无 5540**=发现阶段没完成 |
| 配完后 | 设备宣告 `_matter._tcp` + `_I<fabric>._sub` | 出现 = 已入网成功 |
| BLE 配网(手机 App 直连设备 BLE,如 Aqara Home | 线上**无 TCP 5540**BLE 会话对 AP/hass 抓包不可见);设备入网后仍先 mDNS 宣告 `_matterc` | 成功判据=最终宣告 `_matter._tcp` + `_I<fabric>._sub`;无 5540 **不代表**失败 |
## 5. 排障决策树(按顺序)
1. 抓包看**有没有 `_matterc` 宣告**:没有 → 设备不通电 / 没连上 Wi-Fi / 没进配对模式
(先解决"设备在线",网络侧已反复验证正常)。
2. 有宣告但 **`CM=0`** → 设备已配 / 不在配对模式 → **恢复出厂**后重试(用它自己的二维码)。
3. 有宣告 `CM=1` 但**无 `_L<disc>` 子类型** → 固件 mDNS 缺陷 → 升固件或换通用发现配对方。
4. `CM=1` + 子类型齐全但**无 TCP 5540** → 配对方没匹配上(查码/discriminator)或设备不可达。
5. 有 5540 但配对中断 → 查 `CM` 源(码是否正确)、设备电源、fabric 状态(是否需先清)。
## 6. 相关文档
- [runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md) — Matter 抓包指令集(实时/落盘轮转/定向/解密)
- [docs/lan-overview.md](lan-overview.md) — LAN 拓扑、SSID 清理、ULA 不可行
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — matter-server 重拨运维规范
- [docs/unifi-network.md](unifi-network.md) — UniFi 网络/IPv6/SSID 记录
- [hosts/gw.md](../hosts/gw.md) — DHCP 保留 `matter` MAC 错位(待修)
## 7. 2026-08-22 实测记录:添加新 ESP32-C2 Matter 灯泡(成功 + 失败路径全记录)
> 场景:手机 AppAqara Home)添加一盏**新的** ESP32-C2 Matter 灯泡
> `34:98:7a:27:10:bc`hostname `matter`,最终 IP `.148`)。中途换了灯泡并断电重启,
> 共经历 **2 种失败模式** 和 **1 条成功路径**,全部抓包实证。
>
> 抓包点:UAP-AC-Lite `192.168.55.5` `br0`(轮转 `udp 5353 or tcp 5540 or tcp 5552`
> + 定向全量 `ether host 34:98:7a:27:10:bc`+ hostapd/stahtd 日志 + gw DHCP/ARP 交叉验证。
> 本地用 tshark 4.7.2 分析。
### 时间线(CST2026-08-22
| 时间 | 事件 | 判据 / 说明 |
|---|---|---|
| 10:50:28 | 启动轮转抓包 | — |
| 10:51:1219 | 手机 `.143`OnePlus,连 wifi0ap0=`ubnt-windy-2`)重新关联;查询 `_matter._tcp` | 运营查询(浏览已配设备),**不是**配网(配网应查 `_matterc._udp` |
| ~10:56 | 用户报「配置 wifi 后挂起,不能加入 wifi」 | 首次失败 |
| 11:00:15 | M3 `.248` 查询 `_L3266._sub._matterc` | 无应答(是另一台设备的 discriminator,无关) |
| 11:01:2442 | **失败模式 A**:故障盏 `…7f:08` 尝试关联 `wifi0ap1`(`ubnt-haas`):发 1 次 open-auth 帧(algorithm 0)→ AP 回 `status_code=0`**客户端不再发 assoc 请求** → 18s 后 `auth_failures=1` + disassociated | auth 阶段卡死(client 侧);非密码错——密码错会先 assoc 再 4-way 失败 |
| 11:04:0105 | **新灯泡 `…10:bc` 关联 `wifi0ap1` 成功**WPA2 4-way 完成,DHCP 拿 `.148`tracker `soft failure`: ip_delta 3.76savg_rssi -68);随即宣告 `_matterc`:实例 `3F4E2C66F2DA85CD`TXT `VP=4891+4100 D=3377 CM=1`SRV :5540**有 GUA** | 发现阶段判据全过 |
| 11:04:05 之后 | **失败模式 B**:灯泡网络栈完全静默——STA 收发计数冻结(rx=89/tx=5 持续 12s+ 不变)、**不掉线不重启** | firmware 挂死 |
| 11:04:5911:07:27 | 手机查 `_matterc` ×4、`.60` 解析实例、M3 查 `_L3377._sub._matterc` ×5discriminator 3377 正是新盏)——**全部无应答**TCP 5540/5552 全程 0 | 配对方找不到设备 → App 报「加不上」 |
| 11:09:3548 | **断电 10 秒重启**:灯泡重新关联 `wifi0ap1` ×2 | 对策生效 |
| 11:10:26 | DHCP 重新拿 `.148`;STA 计数恢复持续增长(活跃) | — |
| 11:1011:15 | 重新宣告 `_matterc`**新实例 `E5BA8E28E4DE23A0`**——入配对模式实例名更换,符合规范);App 走 **BLE 配网** | 线上无 TCP 5540BLE 对 AP 不可见,属正常) |
| 11:15:23 | 灯泡宣告 **`_matter._tcp`**`4DF2B1455D19402D-02EF2FF12DFAF10E`Aqara fabric+ SRV :5540 + GUA + A | ✅ **添加成功**(已入 Aqara fabric `4DF2B1455D19402D` |
### 结论与经验
1. **同族灯泡(VP 4891/4100OUI 34:98:7a)存在两种不同失败模式**
- 故障盏 `…7f:08`:auth 阶段卡死(auth 帧后不发 assoc);此前(08-21 21:27)成功关联后伴随
`ip_failures=1`(拿不到 IP)+ 缺 GUA —— 属更深层故障,需恢复出厂,本次未处理,仍离线。
- 新盏 `…10:bc`:入网 + 宣告 `_matterc`(CM=1)成功后约 3 秒固件挂死(全静默)。
**断电 10 秒重启即恢复**,是最简单有效的对策。
2. **AP 抓包看不到 BLE 配网**Aqara Home App 对 WiFi Matter 设备走 BLE 配网时,线上只有
mDNS/DHCP**无 TCP 5540 不代表失败**;成功判据 = 设备最终宣告 `_matter._tcp` + `_I<fabric>._sub`
3. **发现判据回顾**`_matterc` + TXT`CM=1``D=``VP=`+ SRV :5540 + AAAA(GUA) + A 全齐才算
设备真的在配对模式;配对方按 `_L<disc>._sub._matterc` 精确匹配 discriminator(本例 D=3377)。
4. 新盏 RSSI -68、DHCP 3.76s,射频偏弱,可能加剧 firmware 不稳定(待观察)。
5. DHCP 保留 `matter`.45→`…10:bc`**仍未生效**:新盏实际拿动态 `.148`(待修,见 hosts/gw.md)。
6. 识别「配对方在找但设备不答」的快速方法:抓包里配对方持续查 `_matterc`/`_L<disc>` 而目标 MAC
零应答 + STA 收发计数冻结 = 设备侧挂死;此时**先断电重启设备**,不要怀疑网络/AP。
### 后续:新盏 11:24 起离线循环(同一盏 `…10:bc`2026-08-22
配网成功后约 10 分钟(11:15–11:24 可控制),灯泡进入**持续性故障循环**:
| 时间 | 事件 | 模式 |
|---|---|---|
| 11:24:57 | `EVENT_STA_LEAVE`(真掉线) | 掉线 |
| 11:25:12 | 重连 `auth_failures=2` | **auth 卡死**(同故障盏 `…7f:08` 11:01 的模式) |
| 11:25:2325 | 重连成功,WPA2 完成,重新拿 `.148` | — |
| 11:25:38 | `soft failure`ip_delta 2.65s**avg_rssi -73**-68→-73 持续变差) | 射频偏弱 |
| 11:25 之后 | STA 计数冻结(rx=119/tx=64 不动);M3 持续查询其运营实例 `4DF2B1455D19402D-02EF2FF12DFAF10E._matter._tcp` **无应答** → App 显示「离线」 | 静默挂死 |
**结论**:三盏 ESP32-C2 灯泡中两盏(`…7f:08``…10:bc`)故障,表现覆盖 auth 卡死 / 静默挂死 /
随机掉线三种形态;工作盏 `…25:a1:f0` 正常。网络侧(AP、M3、DHCP、mDNS)均验证正常。
**疑似根因(按可能性)**:① ESP32-C2 Matter 灯泡 firmware 缺陷(同批次)② 射频偏弱
(RSSI -73,天线/距离/遮挡)加剧不稳定 ③ 供电不稳(brownout 造成 Wi-Fi 栈崩溃重启)。
**待办**:移近 AP 或改善供电后观察;App 内查固件更新;仍复发则考虑退换。
## 8. 2026-08-23 状态核查:两盏「半在线」——mDNS 宣告正常但 TCP 5540 无监听(失败模式 C)
> 全程**只读**核查(gw DHCP/ARP、AP hostapd 日志、hass matter-server 状态 + mDNS 抓包、
> 对灯泡 v4/v6 的 TCP 5540 探测,09:0x CST)。结论:**网络侧全部健康;两盏灯泡网络栈活着、
> mDNS 运营宣告正常,但 Matter 会话端点(TCP 5540)无监听**——配对方无法建立 CASE,
> App 内应显示离线/不可达。
| 对象 | 状态(2026-08-23 |
|---|---|
| 新盏 `34:98:7a:27:10:bc``.148`hostname `matter` | DHCP 04:40 续租;gw ARP 完整;ping 通(93122msESP32 省电时延);08-22 16:40 起稳定关联 `wifi0ap1`,关联时 `avg_rssi -70`。mDNS 宣告 3 实例:`4DF2B1455D19402D-02EF079EEB480D07`**新 node ID——08-22 之后被重新配网过**)、`2F6E56020E1996E7-137147AF27BE4EB6``DCE86145C137AF0E-0000000000000011`HA fabric);host 记录 A `.148` + fe80 + **当前前缀** GUA `240e:3bd:238:4812:*`。支持单播 legacy mDNS 查询(`dig -p 5353 @.148 _matter._tcp.local PTR` 可用) |
| 工作盏(MAC 已变)`fc:e8:c0:25:a1:f0``.146`hostname `espressif` | DHCP 07:17 续租;ping 通 v4/v6v6 fe80 3861ms)。mDNS 宣告 3 实例:`4DF2B1455D19402D-02EF4CA3F856B615``2F6E56020E1996E7-EE8F2E4F1A77BF05``DCE86145C137AF0E-000000000000000B`。原 MAC `34:98:7a:25:a1:f0` 全网消失(无租约/ARP/AP 日志)而新 MAC 末 3 字节相同 → 疑固件更新后改 MAC。**拒绝单播 5353**ICMP port unreachable),只应答组播查询——同族固件行为差异。hass 残留其旧前缀 GUA `240e:3bd:235:1fb2:fee8:c0ff:fe25:a1f0`**FAILED** 邻居项 |
| 故障盏 `34:98:7a:27:7f:08`(曾 `.145` | 无租约、ARP incomplete、AP 日志零事件 = 已离网 |
| **TCP 5540 探测(两盏)** | IPv4LAN66 与 hass 本段)、IPv6fe80%end0 + 当前 GUA)全部 **RSTConnection refused** —— SRV 宣告 :5540 且 TXT `T=1`,但实际无监听 |
| hass matter-server | `started`,v9.0.4,无更新;宣告自身运营实例 `DCE86145C137AF0E-…1B669`v4+v6,当前 GUA);**无任何 established :5540 会话**core/add-on 日志无 matter 错误 |
| 其他 Matter 控制器 | Aqara M3 `.248` 在线(有线 0.8ms),宣告含自身 fabric 节点 `4DF2B1455D19402D-11E158E46D24A000`SmartThings `.48` 在线并周期查询 `_matter._tcp.local`;手机(当前前缀 GUA)也在浏览。LAN55 共见 **5 个 fabric**`4DF2B1455D19402D`M3)、`DCE86145C137AF0E`HA)、`2F6E56020E1996E7``03BCFAEDD6153944``6A6FF80C2DB84DEE` |
**判定**:失败模式 C = TCP/IP 栈与 mDNS 守护进程活着(主动 RST、DHCP 续租、ping 通),
但 Matter 应用层监听不存在。与模式 A(auth 卡死)、模式 B(全静默挂死)同族不同形态;
**两盏同时处于同一状态**更指向共同诱因(固件缺陷,或 PD 轮换等共同事件后未恢复)。
**对策(推荐,未执行)**:逐盏断电 10 秒重启(模式 B 的已验证对策),重启后复测
TCP 5540 恢复监听即可确认。
**核查方法备忘**(只读,可复用):
- gw`show dhcp leases` / `show arp`(经 `/opt/vyatta/bin/vyatta-op-cmd-wrapper`)。
- AP`grep -i <mac> /var/log/messages`hostapd 关联事件 + stahtd RSSI/soft failure)。
- hass`sudo -n -i ha apps info core_matter_server``ip -6 neigh show dev end0`
(看灯泡 fe80/旧新前缀 GUA 与 FAILED 项);被动抓包
`sudo -n -i timeout 65 tcpdump -ni end0 -s 0 -tt 'udp port 5353'`——配对方周期查询
会自然引出灯泡宣告,无需主动发包。
- 5540 探测:hass 上 python3 对 v4 / fe80%end0 / GUA 各 connect 一次;RST=无监听,
超时=不可达(两者含义不同)。
+45
View File
@@ -0,0 +1,45 @@
# Plane CE 加固草稿(docs/plane-hardening/
> **状态:草稿,未应用、未提交。** 对应追踪:Plane vps 项目条目(2026-09-03**记录源**Linear W1N-277 已取消,Linear 自 2026-09-03 起不再作为记录源)。
> 线上实例:`plane.chans.xyz`synapse K3sns `plane`release `plane-app` = chart `plane-ce-1.8.0` / app `v1.4.1`)。
> 依据:2026-09-03 只读核查(13 条审查意见中 11 条属实、#3 基本属实、#9 指标归属错误)+ 上游 chart 模板逐条核对。
## 文件
| 文件 | 内容 |
|------|------|
| `values.hardened.yaml` | 可选硬化 valuesexternal secrets 引用、requireExplicitSecrets、minio pin、上传限额对齐);含 HTTP→HTTPS `extraObjects` 示例 |
| `secrets.yaml.example` | 6 组外部 Secret 结构占位(只含 key 名,真实值仅存宿主机) |
| `backup/plane-backup.yaml` | **PostgreSQL 备份 CronJob**pg_dump `-Fc`hostPath `/var/backups/plane`MinIO 已按实际用量剔除) |
| `backup/README.md` | 备份方案说明(排程/容量/保留/还原/阻塞) |
## 应用顺序(每步先 diff 后执行,全部需用户逐项确认)
### 现在就值得做:DB 备份(P0,见 backup/
`plane-backup.yaml` 部署 + 手动触发验证一次即可;88 MB 库每日快照几乎零成本。
### 可选(顺手做一次,不是必须)
- **Phase A 密钥外部化**(零行为变化、无停机,约 15 分钟):按 `secrets.yaml.example`
在宿主机建 6 个 Secret(值先复制当前集群),用 `values.hardened.yaml`
`helm diff upgrade``helm upgrade`;验证后删除 chart 生成的旧 Secret。
价值:默认密钥不再落在 chart 公开常量上,作为保险。
- **MCP API Key 轮换**:若审查对话出过你的环境,Plane 后台重生成 + 更新
`/home/windy/plane-k3s/mcp/mcp.env`0600+ 重启 Cursor MCP。
- **/god-mode IP 白名单**:若在意管理后台被公网爆破。chart 1.8.0 的 IngressRoute
不支持给单条路由追加 middleware → 需 post-renderer 或 upgrade 后 `kubectl patch`
(升级会覆盖,需固化);源 IP 清单待提供。
### 明确暂缓/跳过(个人单节点,等出现症状再处理)
- SECRET_KEY 等轮换(Phase B):等真要配 SMTP/OAuth 前再做(避免旧密文不可解)。
- NetworkPolicy、有状态组件 resources limitschart 无 values 开关,需 post-render/patch)、
HTTP→HTTPS(草稿已给 `extraObjects` 示例)、metrics-server/Sentry。
## 关键限制(chart 1.8.0 模板已核对)
- `external_secrets.*_existingSecret` 设置后,对应 Secret **必须**包含模板所需全部 key
(缺失不自动补),见 `secrets.yaml.example` 注释。
- `app_keys_existingSecret` 的 envFrom 在所有 workload 上**最后注入**(后置生效),
保证 app/live 共享密钥一致——不要在其后再放同名 key 的 Secret。
- `DATABASE_URL`/`AMQP_URL`/`REDIS_URL` 是 chart 生成的派生 URL,内嵌明文密码;
外部化后轮换 DB/队列密码时必须同步更新 `plane-app-env`
- minio 的 `MINIO_ROOT_*``AWS_*` 同源于一个 Secret;升级时 bucket Job 会重跑
(需 admin 权限凭据)——换 svcacct 前先确认权限覆盖该 Job。
+48
View File
@@ -0,0 +1,48 @@
# Plane CE 备份方案(DB-only)— DRAFT (2026-09-03), 未应用
> 关联:`plane-backup.yaml`CronJob);追踪:Plane vps 项目条目(记录源,2026-09-03 起不用 Linear)。
> 现状(实测):pg 全库 **88 MB**310 issues / 1 user);MinIO uploads **264 KB**(几乎空)。
## 范围决策(2026-09-03,实际角度)
- **做:PostgreSQL 逻辑备份** —— 覆盖现实故障(误删、升级失败、磁盘坏、重装),成本≈0。
- **不做:MinIO/附件备份** —— 桶仅 264 KB,个人实例附件可接受丢失;不为它付日常维护。
日后附件明显变多再按原完整版思路加 `mc mirror`(历史版本见本目录 git 历史/Plane 条目评论)。
- 异机同步暂不启用(见下"局限/阻塞")。
## 方案
集群内 CronJobns `plane`,每天 **01:30 UTC = 03:30 本地**,控制器按 UTC 跑):
1. 单容器 `postgres:15.7-alpine``pg_dump -Fc`(自定义压缩格式)打 `plane`
`/var/backups/plane/pg/plane-<UTC时间戳>.dump`hostPath `DirectoryOrCreate`
2. 保留 7 天(`find -mtime +7 -delete`),成功/失败历史各留 3/2
3. 凭据:现 chart Secret `plane-app-pgdb-secrets`Phase A 外部化后改 `plane-pgdb-credentials`
## 容量
- 库 88 MB → `-Fc` 快照约 10–40 MB/天 × 7 天 ≈ **<300 MB**,对 83 G 可用盘可忽略。
## 还原(未演练;应用前先做一次隔离测试)
```bash
# 目标 PG15 实例(临时起一个 postgres:15.7-alpine 容器或另一台机):
# 先建空库: createdb plane (user=plane)
pg_restore -h <target> -U plane -d plane --clean --if-exists /var/backups/plane/pg/plane-<TS>.dump
# 还原后确认 310 issues 量级一致;附件为空属预期(未备份 MinIO)
```
## 验收(应用前逐项过)
- [ ] CronJob 建立后手动触发一次:`kubectl -n plane create job --from=cronjob/plane-backup plane-backup-manual-1`Job `Completed`
- [ ] `/var/backups/plane/pg/plane-*.dump` 可被 `pg_restore -l` 列出
- [ ] 备份 Job 只依赖 pgdb 服务,不依赖 Plane 应用 Pod(应用故障期间也能出备份)
- [ ] 保留清理 dry-run`find ... -print`)正确;`df -h /` 前后对比记录
## 局限 / 阻塞
- **本地方案不是离机备份**:单节点磁盘/整机故障即丢。如日后要离机,纳入
[Restic 异机 repository 决策与存取隔离](https://plane.chans.xyz/space/projects/56874283-7e1d-43a8-afa4-631cf1c4ad5b/issues/7825d564-ae15-446b-bced-be26b648346b/)
(与 Matrix 备份同一决策);恢复演练纪律见
[服务级 restore runbook 与隔离复元演练](https://plane.chans.xyz/space/projects/56874283-7e1d-43a8-afa4-631cf1c4ad5b/issues/a9bea3ba-c958-4a74-b2f1-6bbb653f21d3/)。
- 提醒:同一节点 **Matrix 数据价值远高于 Plane 且同样无备份** —— 若投入备份精力,顺序上 Matrix 优先。
@@ -0,0 +1,63 @@
# Plane CE PostgreSQL backup CronJob — DRAFT (2026-09-03), NOT applied.
# ns: plane (synapse K3s single node). Output: hostPath /var/backups/plane (root disk, auto-created).
#
# Scope decision (2026-09-03, practical): DB-only. MinIO dropped — uploads bucket
# measured at 264 KB / 444 KB total; attachments are acceptable loss for this
# personal 1-user instance (310 issues / 88 MB DB). Revisit only if usage grows.
#
# Credentials: read from the CURRENT chart-generated Secret (works today). After the
# optional external-secrets migration (docs/plane-hardening/README.md Phase A) switch
# the secretKeyRef name to plane-pgdb-credentials.
#
# Apply:
# ssh windy@synapse.chans.xyz 'sudo k3s kubectl apply -n plane -f -' < plane-backup.yaml
# Manual run + verify:
# sudo k3s kubectl -n plane create job --from=cronjob/plane-backup plane-backup-manual-1
# sudo k3s kubectl -n plane get cronjob,job,pods | grep plane-backup
# sudo ls -lh /var/backups/plane/pg
# Restore steps + tuning: see backup/README.md
apiVersion: batch/v1
kind: CronJob
metadata:
name: plane-backup
namespace: plane
spec:
# 01:30 UTC daily = 03:30 local (CEST). CronJob controller runs in UTC.
schedule: "30 1 * * *"
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 2
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: OnFailure
volumes:
- name: backup
hostPath:
path: /var/backups/plane
type: DirectoryOrCreate
containers:
- name: pg-dump
image: postgres:15.7-alpine
env:
- name: PGPASSWORD
valueFrom:
secretKeyRef:
name: plane-app-pgdb-secrets # -> plane-pgdb-credentials after Phase A
key: POSTGRES_PASSWORD
command: ["/bin/sh", "-c"]
args:
- |
set -euo pipefail
TS=$(date -u +%Y%m%dT%H%M%SZ)
mkdir -p /backup/pg
pg_dump -h plane-app-pgdb.plane.svc.cluster.local -U plane -d plane \
-Fc -f "/backup/pg/plane-${TS}.dump"
find /backup/pg -type f -name 'plane-*.dump' -mtime +7 -delete
echo "pg_dump done: /backup/pg/plane-${TS}.dump ($(du -h /backup/pg/plane-${TS}.dump | cut -f1))"
volumeMounts:
- name: backup
mountPath: /backup
+91
View File
@@ -0,0 +1,91 @@
# External Secret structure for Plane CE hardening — EXAMPLE ONLY.
# No real values here; this file is safe to commit. Real values live only on the
# host (/home/windy/plane-k3s, 0600/0700) and in the cluster.
#
# Phase A — create each Secret with the CURRENT cluster values first (zero change):
# # current source Secrets (chart-generated):
# kubectl -n plane get secret plane-app-app-secrets -o jsonpath='{.data.SECRET_KEY}' | base64 -d
# kubectl -n plane get secret plane-app-live-secrets -o jsonpath='{.data.REDIS_URL}' | base64 -d
# kubectl -n plane get secret plane-app-pgdb-secrets -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d
# kubectl -n plane get secret plane-app-rabbitmq-secrets -o jsonpath='{.data.RABBITMQ_DEFAULT_PASS}' | base64 -d
# kubectl -n plane get secret plane-app-doc-store-secrets -o jsonpath='{.data}' | base64 -d
#
# e.g. kubectl -n plane create secret generic plane-app-keys \
# --from-literal=SECRET_KEY="$(<copy from above>)" \
# --from-literal=LIVE_SERVER_SECRET_KEY="$(<copy from above>)"
#
# All keys below are REQUIRED by chart templates/plane-ce-1.8.0 (verified 2026-09-03):
# missing keys are NOT auto-filled once an existingSecret is referenced.
---
apiVersion: v1
kind: Secret
metadata:
name: plane-app-keys # external_secrets.app_keys_existingSecret
namespace: plane
type: Opaque
stringData:
SECRET_KEY: "" # current: copy from plane-app-app-secrets; rotate only in Phase B
LIVE_SERVER_SECRET_KEY: "" # current: same value as above / plane-app-live-secrets
---
apiVersion: v1
kind: Secret
metadata:
name: plane-app-env # external_secrets.app_env_existingSecret
namespace: plane
type: Opaque
stringData:
REDIS_URL: "" # redis://plane-app-redis.plane.svc.cluster.local:6379/
DATABASE_URL: "" # postgresql://plane:plane@plane-app-pgdb.plane.svc.cluster.local/plane
AMQP_URL: "" # amqp://plane:plane@plane-app-rabbitmq.plane.svc.cluster.local/
---
apiVersion: v1
kind: Secret
metadata:
name: plane-live-env # external_secrets.live_env_existingSecret
namespace: plane
type: Opaque
stringData:
REDIS_URL: "" # redis://plane-app-redis.plane.svc.cluster.local:6379/
---
apiVersion: v1
kind: Secret
metadata:
name: plane-pgdb-credentials # external_secrets.pgdb_existingSecret
namespace: plane
type: Opaque
stringData:
POSTGRES_PASSWORD: "" # Phase A: keep current ('plane'); Phase B: ALTER USER first, then sync
POSTGRES_DB: "plane"
POSTGRES_USER: "plane"
---
apiVersion: v1
kind: Secret
metadata:
name: plane-rabbitmq-credentials # external_secrets.rabbitmq_existingSecret
namespace: plane
type: Opaque
stringData:
RABBITMQ_DEFAULT_USER: "plane"
RABBITMQ_DEFAULT_PASS: "" # Phase A: keep current; Phase B: rabbitmqctl change_password first
---
apiVersion: v1
kind: Secret
metadata:
name: plane-minio-credentials # external_secrets.doc_store_existingSecret
namespace: plane
type: Opaque
stringData:
FILE_SIZE_LIMIT: "20971520" # must match env.doc_upload_size_limit
AWS_S3_BUCKET_NAME: "uploads"
USE_MINIO: "1"
MINIO_ROOT_USER: "admin"
MINIO_ROOT_PASSWORD: "" # root creds take effect on first init only
AWS_ACCESS_KEY_ID: "admin"
AWS_SECRET_ACCESS_KEY: "" # == MINIO_ROOT_PASSWORD while minio.local_setup
AWS_S3_ENDPOINT_URL: "http://plane-app-minio:9000"
+109
View File
@@ -0,0 +1,109 @@
# Plane CE hardened values — DRAFT (2026-09-03), NOT applied.
# Target file on host: /home/windy/plane-k3s/values.yaml (synapse.chans.xyz)
# Reference release: plane-app, chart plane-ce-1.8.0 (values.yaml L1-362 + templates verified 2026-09-03).
# No secrets in this file. Secret *values* live only in k8s Secrets (see secrets.yaml.example).
#
# Two phases:
# Phase A: externalize secrets (reference names below) with CURRENT values copied -> zero change.
# Phase B: rotate credentials one by one (see README.md). SECRET_KEY rotation is cheap only while
# SMTP/OAuth are unconfigured (no encrypted config rows yet).
planeVersion: v1.4.1
ingress:
enabled: true
appHost: plane.chans.xyz
ingressClass: traefik
traefik:
# 20 MiB (chart default). Keep aligned with env.doc_upload_size_limit below.
maxRequestBodyBytes: 20971520
ssl:
createIssuer: true
issuer: http # HTTP-01; ssl_token_existingSecret not needed
email: admin@chans.xyz
generateCerts: true
postgres:
storageClass: local-path
volumeSize: 5Gi
# NOTE: chart 1.8.0 exposes NO resources knob for the bundled datastores
# (stateful templates render no resources block). Add limits via
# --post-renderer/kustomize or `kubectl -n plane patch sts ...` re-applied on
# every upgrade (P2 task; see README.md).
redis:
storageClass: local-path
# image: valkey/valkey:7.2.11-alpine # already pinned by chart default; uncomment to make explicit
minio:
# P2: pin. Digest of the currently running :latest (2026-09-03, pod plane-app-minio-wl-0).
image: minio/minio@sha256:14cea493d9a34af32f524e538b8346cf79f3321eff8e708c1e2960462bd8936e
# image_mc: minio/mc@sha256:... # optional: pin one-shot bucket-init client the same way
storageClass: local-path
volumeSize: 5Gi
rabbitmq:
storageClass: local-path
env:
# Fail the render instead of ever falling back to the chart's PUBLIC constants
# (values.yaml L340-341 in chart 1.8.0). Requires external_secrets below.
requireExplicitSecrets: true
# SECRET_KEY / LIVE_SERVER_SECRET_KEY are deliberately OMITTED here.
# They live in k8s Secret `plane-app-keys` (referenced below). With
# requireExplicitSecrets=true and app_keys_existingSecret set, the chart renders
# neither key itself and app+live workloads both envFrom `plane-app-keys` LAST
# (later envFrom wins), which keeps the shared signing key consistent.
pgdb_name: plane
docstore_bucket: uploads
# Align app-side upload cap with the Traefik body limit (was 5242880/5MiB).
# Keep both at 20MiB, or lower both together.
doc_upload_size_limit: "20971520"
external_secrets:
# Shared signing keys (used by app + live). REQUIRED keys: SECRET_KEY, LIVE_SERVER_SECRET_KEY.
app_keys_existingSecret: plane-app-keys
# REQUIRED keys: REDIS_URL, DATABASE_URL, AMQP_URL (chart-derived URLs; update on DB/queue rotation).
app_env_existingSecret: plane-app-env
# REQUIRED keys: REDIS_URL.
live_env_existingSecret: plane-live-env
# REQUIRED keys: POSTGRES_PASSWORD, POSTGRES_DB, POSTGRES_USER.
pgdb_existingSecret: plane-pgdb-credentials
# REQUIRED keys: RABBITMQ_DEFAULT_USER, RABBITMQ_DEFAULT_PASS.
rabbitmq_existingSecret: plane-rabbitmq-credentials
# REQUIRED keys: FILE_SIZE_LIMIT, AWS_S3_BUCKET_NAME, USE_MINIO, MINIO_ROOT_USER,
# MINIO_ROOT_PASSWORD, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_S3_ENDPOINT_URL.
doc_store_existingSecret: plane-minio-credentials
# ssl_token_existingSecret: '' # DNS-01 only (cloudflare/digitalocean); unused with HTTP-01
# Optional, P2: HTTP -> HTTPS 301. The chart's own IngressRoute binds only
# 'websecure' (http:// currently 404s). extraObjects is rendered verbatim (toYaml).
# Uncomment and `helm upgrade` once reviewed:
# extraObjects:
# - apiVersion: traefik.io/v1alpha1
# kind: Middleware
# metadata:
# name: plane-https-redirect
# namespace: plane
# spec:
# redirectScheme:
# scheme: https
# permanent: true
# - apiVersion: traefik.io/v1alpha1
# kind: IngressRoute
# metadata:
# name: plane-http-to-https
# namespace: plane
# spec:
# entryPoints: [web]
# routes:
# - match: Host(`plane.chans.xyz`)
# kind: Rule
# middlewares:
# - name: plane-https-redirect
# services:
# - name: plane-app-web
# port: 3000
@@ -1,158 +0,0 @@
# 希力威视 SR-S25G3218F 调查(2026-08-08
**结论:** 若需求是大量 2.5G 终端、少量 10G 光上联,`SR-S25G3218F` 的端口密度
更合适;厂商已公开该型号的固件页,但仍缺少完整规格书、管理手册与兼容矩阵。若需求是 8 条全部可协商
1/2.5/5/10G 的铜缆链路,且希望有可查的 L3 能力和固件入口,兮克
`SKS8300-8T` 是资料更完整、风险更低的选择;它的代价是主动风扇、外置 12 V 电源、
无 SFP+ 光口,且仍不应把消费级/SMB 设备当作安全边界或唯一核心。两者都应在
到货可退换期内完成实机验收。
本页为采购前资料调查,不代表已接入本地网络;检索日期为 2026-08-08。
## 已能核实的事项
| 项目 | 结论与证据强度 |
|---|---|
| 型号/端口 | 京东的希力威视商品标题称该 SKU 为 `SR-S25G3218F`,有 16 个 2.5G 电口和 2 个万兆光口,并宣传 VLAN、端口隔离与 LACP。该店铺被厂商官网列为可购买的「京东旗舰店」,因此可作为销售规格,非技术手册。[京东商品页](https://item.jd.com/100165071727.html)[厂商购买渠道说明](https://en.sirivision.com/contactus/) |
| 厂商身份 | 厂商官网为 Shenzhen/Guangdong Sirivision Communication;英文官网说明其自 2016 年起提供接入、汇聚和核心交换机方案。[厂商首页](https://en.sirivision.com/) |
| 公开的二手厂家资料 | 同一制造商名义的 Alibaba 出口页将精确型号写成 `16*2.5G+2*10G``120Gbps`,并列出 QoS、VLAN、SNMP、L3 与 stackable。这是制造商发布在平台上的销售资料,**不是**官网数据表;其中后五项不能据此视为已验收的功能承诺。[制造商平台页](https://www.alibaba.com/pla/SR-S25G3218F-QoS-Managed-SFP-Switch-1625G210G_1601494946214.html) |
| 固件入口 | 厂商已发布此精确型号的[固件页](https://www.sirivision.com/sr-s25g3218f%E5%9B%BA%E4%BB%B6/)。公开变更记录提到“光口自适应”和“增加 DAC 配置”;这证明厂商维护过该路径,**不**代表任意 SFP+/DAC/铜模块均兼容。 |
| 本机可计算的带宽 | 端口线速相加为单向 60 Gb/s16 × 2.5 + 2 × 10);若厂商所谓 `120Gbps` 是全双工交换容量,则数学上吻合。它**不**证明缓冲、PPS、表项规模或实际无阻塞性能。 |
## 网管/L2/L3 能力边界
京东标题足以支持把 VLAN、端口隔离、LACP 作为「卖家声称提供」的功能;不得由此推导出
ACL、IPv4/IPv6 静态路由、SVI 数量、DHCP relay、OSPF/RIP、VRRP、IGMP、ERPS、
802.1X、RADIUS/TACACS+、SSH/HTTPS 管理、SNMP 版本、日志/审计、配置备份或固件
安全维护一定存在。
尤其要注意:厂商官网把真正列出的 2.5G L3 产品标为
`SR-S25G3412F (8 × 2.5G + 4 × 10G SFP+)`;其 2.5G 类目只显示 7 个型号,
不含 `SR-S25G3218F`。官网也把 L2+、Web Smart、L3 分成不同产品类别。这个目录
差异**不是**证明 3218F 没有 L3,而是说明「三层」无法通过官网的精确型号文档确认。
[2.5G 产品目录](https://en.sirivision.com/product-category/products/2-5g-switches/)
[官网的 10G L3 目录](https://en.sirivision.com/product-category/products/10g-switches/10g-layer3-managed-switches/)
[官网的 L2+ 分类示例](https://en.sirivision.com/product-category/products/gigabit-switches/gigabit-layer2-managed-switches/)。
采购前请向京东/厂商索取**与机身 SKU、硬件 revision 和固件版本对应**的 PDF
数据表、管理手册和 release notes,并要求书面回答至少以下问题:
1. L3 是只有 VLAN Interface/IPv4 静态路由,还是另有 IPv6、ACL、动态路由、DHCP relay
等;每项的最大 VLAN、MAC、ARP、路由、ACL、LAG 数量分别是多少?
2. LACP 是否符合 802.3ad、一个 LAG 最多多少成员、能否跨两台设备(若销售页的
`stackable` 属实,堆叠的线缆/模块、最大成员、控制面和软件版本为何)?
3. 管理面是否支持 HTTPS/SSH、禁用 HTTP/Telnet、独立管理 VLAN、SNMPv3、syslog、NTP、
配置导出/回滚和已签名或可校验的固件;默认凭据首次登录是否强制修改?
## 供电、散热和光口:当前不能确认
针对该精确 SKU,厂商官网目录与公开搜索未找到说明书/数据表,所以以下均为**待确认,
不能猜测**
- 是否为内置 AC 电源、额定输入范围/最大功耗、是否带电源开关和接地端子;是否完全
不提供 PoE(本型号名和京东标题均未写 PoE,但这不足以替代规格书)。
- 风扇数量、常态/满载噪声、风向、环境温湿度、机架深度与安装耳;不要将「金属壳」
或产品照片等同于无风扇/静音。
- 两个槽是否均为 **10G SFP+**,是否可协商 1G SFP;支持的 SR/LR/BiDi 波长距离、
DAC/AOC 长度、第三方模块/EERPOM 兼容策略、10GBASE-T SFP+ 模块的功耗/温度限制,
以及是否支持 GPON/XPON ONU「猫棒」。
厂商确实单列「SFP Optical Modules」产品分类,但这不构成 3218F 的兼容清单。
[厂商产品导航](https://en.sirivision.com/)。购买光模块/直连线时,应要求厂商按这台
设备的硬件/固件 revision 出具兼容型号清单;没有书面清单时,先在可退换期实测两端的
链路、重启恢复、热插拔与长时间满载错误计数。
## 风险与建议验收
- **文档/生命周期风险(中到高):** 精确型号不在厂商当前官网 2.5G 目录,虽有固件下载页,
但未公开完整型号手册、明确 release notes 或兼容矩阵。官网的售后条款也要求按具体产品查询保修期,配件(含光纤头)
的保修条款与主机不同;不要把平台页的「3 年」当作中国零售 SKU 的已确认保修。
[厂商售后条款](https://en.sirivision.com/after-sale-protection/)
- **功能表述风险(高):** 页面将 L2 特性和「三层网管」并列;在命令/网页菜单、
手册和测试证明之前,将其当作 L2 VLAN/LACP 设备部署,跨 VLAN 路由仍由现有网关承担。
- **双 10G 上联约束(中):** 两个 SFP+ 可作双上联或一个二成员 LAG,但 LAG 增加的是
多流量总吞吐,单一 TCP/UDP 流通常仍受一条 10G 链路限制;上级设备也必须匹配 LACP
配置。
- **管理面风险(中到高):** 家用/低价网管设备常见明文管理、弱默认口令或不透明的固件
更新周期;采购后先置于受限管理 VLAN,改口令、升级已验证固件,且不将管理界面暴露
到 WAN/访客网。
最低验收应包括:逐口协商 100M/1G/2.5G、两只不同厂家 SFP+/DAC(仅在卖家承诺支持的
范围内)、VLAN trunk/access/PVID、STP/环路保护、LACP 故障切换、端口隔离、满载
双向 iperf3 与错误计数、冷启动后的配置保留,以及管理面的 HTTPS/SSH/SNMPv3/配置备份。
如无法提供与型号匹配的正式资料或其中任一关键项失败,应在退换期内退货,并选择公开
数据表、固件与兼容矩阵更完整的型号。
## 备选:兮克 SKS8300-8T 对比
### 已核实的厂商规格
兮克官网的精确型号页明确将 `SKS8300-8T` 定位为三层管理型 10G 全电口交换机,并列出:
- 8 × 1/2.5/5/10GBASE-T RJ45160 Gb/s 交换容量、119.05 Mpps、12 Mbit 缓存、
16K MAC、12 KB 巨帧、512 MB DRAM、32 MB Flash,尺寸 207 × 136 × 35 mm
- QoS、ACL、IP+MAC+端口绑定、流分类/优先级标记、多端口镜像、静态/灵活 QinQ、
sFlow,以及「基于策略的 IPv4/IPv6 单播路由」。
这些是厂商能力声明,并非对每一种路由协议或表项上限的承诺;但相对 3218F 的仅有
销售标题,它给出了精确型号、转发性能和 L3 范围。[兮克 SKS8300-8T
产品页](https://seekswan.com/user/custom-pages/SKS8300-8T.html)
独立的 OpenWrt 设备资料将其识别为 Realtek RTL9303、512 MB RAM,记录了原厂固件
下载入口和串口/TFTP 恢复路径;其硬件数据页列为 12 V / 4 A。这支持「可恢复、可替换
系统」的可操作性,但**不是**兮克对原厂功能的支持承诺。
[OpenWrt 设备页](https://openwrt.org/toh/xikestor/sks8300-8t)
[OpenWrt 硬件数据](https://openwrt.org/toh/hwdata/xikestor/xikestor_sks8300-8t)。
### 能力、物理与运维比较
| 维度 | 希力威视 SR-S25G3218F | 兮克 SKS8300-8T |
|---|---|---|
| 接口/典型用途 | 16 × 2.5G 电口 + 2 × 10G SFP+(销售规格);适合很多 2.5G 终端/NAS,以 10G 光或 DAC 上联。 | 8 × 1/2.5/5/10GBASE-T;适合 10G 铜缆设备、2.5/5G 多速率 NAS/主机。没有 SFP+,光纤上联必须经媒体转换或选另一型号。 |
| 可确认的三层范围 | 仅销售/平台资料称 L3;没有精确型号官方手册,不能确认静态路由以外的功能。 | 官网明确写策略型 IPv4/IPv6 单播路由、ACL/QoS/sFlow/QinQ;动态路由、VRRP、IPv6 ACL/SNMP/认证等仍须按当前固件手册确认。 |
| 冗余/二层 | 卖家声称 VLAN、端口隔离、LACP;STP/环网的实现与规格未知。 | 官网声明 L3 和多项转发特性,但未在产品页给出 STP/LACP/ERPS 的精确限制;购买前仍索取手册。 |
| 散热/噪声 | 无可核实的精确型号风扇、噪声、功耗或风向数据。 | 独立手册镜像和产品图均称智能温控风扇,但厂商产品页未给 dBA;应按「有风扇、可能听得见」规划,不能承诺静音。 |
| 供电 | 未找到精确型号官方输入/功耗资料。 | OpenWrt 硬件数据记录 12 V / 4 A;确认随附电源适配器的插头、余量和地区认证。官方产品页未给满载功耗。 |
| 固件/恢复 | 有精确型号官方固件页;公开记录包含光口自适应与 DAC 配置改动,但未找到完整 release notes、恢复步骤或兼容矩阵。 | 厂商产品页提供「相关下载」区,OpenWrt 还记录原厂固件入口、RJ45 串口和 U-Boot/TFTP 恢复;原厂镜像是否签名、漏洞修复 SLA、配置回退仍未知。 |
关于 8T 的风扇、满载功耗(常见转述为 ≤36 W)、温度范围、芯片型号等,本次未找到
相应的**厂商原始数据表**;不将第三方手册转录当作已核实规格。若噪声、UPS 容量或
机柜散热是购买约束,请先让卖家提供产品铭牌照片、适配器铭牌照片、额定/实测功耗和
dBA 测试条件。
### 选择与验收建议
-**3218F**:必须有 ≥12 个 2.5G 接入端、10G 光/DAC 上联、且 L3 留给现有路由器。
下单前先取得精确型号手册和 SFP+/DAC 兼容承诺;否则端口数量优势不足以抵消资料风险。
-**8T**:最多 8 个设备但需要多速率 10G RJ45、明确的 IPv4/IPv6 静态/策略路由和
以后自行维护/恢复的余地。不要把其 160 Gb/s 标称交换容量误解为 8 端口同时 10G
全双工的性能保证——该标称与端口总线速数学相等,但仍须以实测和厂商 PPS/缓冲说明为准。
- 两台都不应单独承担防火墙、访客/IoT 安全隔离或 WAN 暴露;VLAN 的跨网段策略和公网
边界留在受支持的网关/防火墙上。先为管理面创建专用 VLAN,仅从管理主机访问,禁用
未使用的远程管理协议,备份配置和原厂固件后再接入生产网络。
## 低功耗核心备选(8 × 2.5G + 2 × SFP+
如果核心只需接最多 8 台铜缆终端、上联/连接 NAS 使用 DAC 或光纤 10G,优先考虑没有
PoE 的以下两款。它们都满足 VLAN trunk、LACP 和至少两个 10G SFP+ 的需求;不要为
AP 选 PoE 版来承担核心,因为 PoE 预算、风扇和待机损耗都会明显增加。
| 型号 | 端口与管理能力(厂商声明) | 厂商功耗 / 噪声资料 | 对当前 LAN 的判断 |
|---|---|---|---|
| **TP-Link Omada SG3210X-M2** | 8 × 100M/1G/2.5G RJ45、2 × 10G SFP+,并有 RJ45 和 Micro-USB console。厂商规格列出 802.1Q VLAN、STP/RSTP/MSTP、静态 LAG 和 802.3ad LACP(最多 8 个聚合组、每组最多 8 端口);L3 是 32 个 IPv4/IPv6 接口、48 条静态路由。 | **无风扇**100240 V AC 内置电源。`UN 1.20` 数据表:待机最高 **6.0 W**220 V/50 Hz、25 °C),最高 **15.3 W**220 V)或 **15.0 W**110 V)。 | **首选低功耗方案。** 足以做 LAN66 核心、给 PVE/gfw 与 U6 Lite 做 VLAN 10 trunk,并以 SFP+ DAC/光口连接 10G NAS/主机;它不提供 5G/10G RJ4510G 铜缆需外置转换或 SFP+ 10GBASE-T 模块。 |
| **MikroTik CRS310-8G+2S+IN** | 8 × 2.5G RJ45、2 × 10G SFP+SFP+ 笼支持 1G/2.5G/10G。RouterOS v7(也可选 SwOS)支持 VLAN、链路聚合与 ACL。 | 18–57 V DC 外置供电;官方给出“无附件”最高 **21 W**、总体最高 **34 W**,且机内 **1 个风扇**。厂商没有在该页给出 dBA。 | 可用且软件/文档/恢复路径成熟,但不是本题的静音低功耗优先项:官方最大功耗显著高于 TP-Link,且有风扇。适合明确偏好 RouterOS/SwOS 与其可维护性时选。 |
功耗数字是各厂商的**上限/待机测试条件**,不是你实际墙插读数;SFP+ 光模块、DAC/AOC,尤其
10GBASE-T SFP+ 模块,会另增功耗和热量。对于本网络,用被动 DAC 或短距光模块连接 10G
设备,通常比全 RJ45 10G 核心更容易保持低温、低噪。
`SG3210X-M2` 的上表数据对应 TP-Link 的 `UN 1.20` 数据表;不同地区/硬件版本的包装、
认证和功耗标注可能不同,购买中国零售版本前应让卖家确认**准确硬件版本、保修渠道和固件地区**。
本次未找到 TP-Link 中国官网的该精确型号页,因此不能把海外官方页面当作大陆现货/售后承诺。
MikroTik 同样应通过其官方零售商查询渠道确认本地库存和保修。两台购买前还应确认所选
SFP+/DAC 的兼容清单。
来源:[TP-Link 产品规格](https://www.tp-link.com/uk/business-networking/omada-switch-access-pro/sg3210x-m2/)
[TP-Link `UN 1.20` 数据表](https://static.tp-link.com/upload/product-overview/2025/202512/20251224/SG3210X-M2%28UN%29%201.20_datasheet.pdf)
[MikroTik 产品页](https://mikrotik.com/product/crs310_8g_2s_in)
[MikroTik 用户手册](https://help.mikrotik.com/docs/spaces/UM/pages/214630429/CRS310-8G%2B2S%2BIN)。
+64 -3
View File
@@ -68,6 +68,66 @@ db.device.find(
).pretty()
```
## IPv6 status (verified 2026-08-20)
IPv6 is **enabled and live** on the main Wi-Fi networks. Read-only
verification, no changes made.
**Controller (`networkconf` in the `ace` DB):** the `Default` LAN network has
`ipv6_enabled: true`, `ipv6_client_address_assignment: slaac`,
`ipv6_ra_enabled: true`, `ipv6_ra_priority: high`, and
`dhcpdv6_allow_slaac: true`. `ipv6_interface_type: "none"` is expected: the
network's gateway is the third-party EdgeRouter (`gw`), so the controller does
not manage WAN-side IPv6 — RA/SLAAC is served by the router.
All active SSIDs map to the `Default` network: `ubnt-windy` (5G),
`ubnt-windy-2` (2.4G), `ubnt-haas` (2.4G) — clients on them receive SLAAC IPv6.
Exception: the dormant `ubnt-upg` VLAN 10 network (and its `ubnt-upg` SSID) has
no IPv6 configuration (default off). See
[Dedicated Wi-Fi through a third-party gateway](#dedicated-wi-fi-through-a-third-party-gateway).
**APs (live):** both managed APs hold global SLAAC addresses on `br0` with a
default route learned via RA from `gw`:
| AP | Global IPv6 on `br0` (at check time) | Default route |
|---|---|---|
| U6 Lite (`192.168.66.6`) | `240e:3bd:235:1fb1:...`/64 | `default via fe80::... dev br0 proto ra` |
| UAP-AC-Lite (`192.168.55.5`) | `240e:3bd:235:1fb2:...`/64 | `default via fe80::... dev br0 proto ra` |
The delegated prefixes are dynamic ISP allocations (PPPoE PD `/60`) and rotate
on redial; only the structure is stable.
**Gateway (`gw`):** the IPv6 routing table shows connected `/64`s on `eth0`
(LAN66) and `switch0` (LAN55) plus `::/0` via `pppoe0`.
Re-verify:
```bash
ssh -4 -o BatchMode=yes zhiqiangf@192.168.66.6 'ip -6 addr show br0; ip -6 route show'
ssh -4 -o BatchMode=yes zhiqiangf@192.168.55.5 'ip -6 addr show br0; ip -6 route show'
```
> **2026-08-21 (W1N-207):** SmartThings Element/vWire provisioning SSIDs
> (`element-8a0d5133c9438f12`, `vwire-8b2d67469e455785`, `vport-F09FC22004E9`) were
> removed (element_adopt setting disabled + element wlanconf deleted + device vwire
> fields cleared) and confirmed off on both APs (normal SSIDs unchanged: `ubnt-windy`,
> `ubnt-windy-2`, `ubnt-haas`, `ubnt-upg`). Root cause of Matter onboarding failure that
> day: HA's matter-server advertised a stale IPv6 GUA (two prefix generations old) in
> mDNS; fixed by restarting the add-on — see [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md).
>
> **SSID ↔ subnet split (Matter-relevant):** `ubnt-windy` (5G) is served only by the
> U6 Lite on LAN66; `ubnt-haas` / `ubnt-windy-2` (2.4G) only by the UAP-AC-Lite on
> LAN55. mDNS is link-local multicast and does **not** cross the routed 55/66
> boundary (no mDNS reflector). Matter commissioning therefore requires phone and
> device on the **same subnet (LAN55)**; a phone on 5G (LAN66) cannot discover a
> LAN55 Matter device.
>
> **Cleanup side-effects (left as-is, harmless):** after the direct-DB cleanup,
> `db.device.cfgversion` holds placeholder values (`0000000000000000` /
> `1111111111111111`) and UAP-AC-Lite has `mesh_sta_vap_enabled=false`; the
> controller has not reverted them and no functional impact was observed.
## Dedicated Wi-Fi through a third-party gateway
### Architecture boundary discovered on 2026-08-08
@@ -151,9 +211,10 @@ Use `ssh zhiqiangf@AP_IP` for the adopted-device account. Do not query or copy
the controller's `mgmt` database setting into logs or documentation: it can
contain the managed SSH password.
On 2026-08-06, key-only IPv4 SSH was verified for both managed APs using the
`zhiqiangf` account. Verify future access without permitting password or
keyboard-interactive fallback:
Key-only IPv4 SSH was verified for both managed APs using the `zhiqiangf`
account on 2026-08-06 and re-verified 2026-08-20 (BatchMode with password and
keyboard-interactive disabled; both APs still log in key-only). Verify future
access without permitting password or keyboard-interactive fallback:
```bash
ssh -4 -o BatchMode=yes -o PasswordAuthentication=no \
+5 -3
View File
@@ -48,9 +48,11 @@ space before increasing retention. DNSSEC is disabled because the selected
upstream path did not pass the known-bad-signature validation check; do not
enable it without re-testing validated upstreams.
The compatible names `hass.windy.lan` and legacy `hass.local` currently point
to the same Home Assistant address. Migrate clients to `hass.windy.lan`; keep
the legacy rewrite until its planned retirement.
`hass.windy.lan` points to Home Assistant via this rewrite. The legacy
`hass.local` rewrite was removed on 2026-08-14; `hass.local` now resolves only
via HAOS mDNS/LLMNR (`hostname: hass`), not via AdGuard Home. The
`nas.windy.local` rewrite was likewise removed on 2026-08-14; `.local` names
are now left to mDNS only. Remaining rewrites all use `.windy.lan`.
## Mihomo and routing boundary
+8
View File
@@ -60,6 +60,14 @@ OpenClash runs `/etc/openclash/clash` (clash_meta core) with configuration
> AAAA now `2404:6800…` (new upstream answering, was `2607:f8b0…` via hk2),
> taobao/intercept/clash-fake-ip all unchanged. Backup:
> `config.yaml.bak-multi-doh-20260813-105421`.
> 2026-09-01: removed `dns.quad9.net` from `foreign_upstream` — recurring
> `WARN foreign_upstream … unexpected EOF` bursts (481 log entries) against
> Quad9 DoH; endpoint answers on probe but gets intermittently
> connection-reset from this network (same failure class as the excluded
> `dns.quad101.net`). Remaining upstreams `adg.chans.xyz` (hk2) +
> `dns.cloudflare.com` both verified live; google.com A + youtube.com AAAA
> resolve through mosdns :6052 after restart. Backup:
> `config.yaml.bak-quad9-remove-20260901-201801`.
- nft: OpenClash injects TPROXY/redirect + DNS-hijack rules into
`table inet fw4`; a residual `table inet passwall` exists with 0 packets (unused)
+61
View File
@@ -34,6 +34,18 @@ new SSH host key out of band before accepting it.
IPv6 prefix delegation assigns SLAAC-capable `/64` networks to both LANs.
`eth4` applies the WAN IPv4 and IPv6 firewall policies.
**SE5420 single-uplink topology (verified 2026-08-22):** the TP-Link `TL-SE5420`
core switch is deployed — management `192.168.66.253` (TP-Link OUI `f8:c9:03`,
web UI on :80/:443). The LAN55 uplink into `switch0` is a **single member
port**: `eth1` link up, `eth2`/`eth3` down. All LAN55 wired devices (hass
`.11`, Aqara M3 `.248`, SmartThings `.48`, UAP-AC-Lite `.5`) are reached via
`switch0` behind that one uplink, so same-segment wired↔wired unicast is
switched locally on the SE5420 and never reaches the ER-X. The switch FDB is
hardware-offloaded and not readable from the ER-X (`brctl showmacs switch0`
"Operation not supported"; `show mac-address-table` / `show ethernet-switch`
are not available on this EdgeOS build) — port link state (`show interfaces
ethernet`) plus ARP are the reliable topology checks.
Detailed effective configuration, including firewall binding and WAN exposure,
is recorded in [the EdgeRouter X configuration record](../docs/edgerouter-x-configuration.md).
@@ -77,6 +89,35 @@ relevant interface/direction to take effect. Use the operational `show
firewall` output—not merely the configured rule definitions—to determine the
effective policy.
## PPPoE redial
To force the `pppoe0` session to reconnect (e.g. to obtain a fresh WAN IP), use
the operational `disconnect` / `connect` commands — **not** `renew dhcp
interface`, which applies only to DHCP interfaces:
```bash
ssh -4 zhiqiang@192.168.66.254
/opt/vyatta/bin/vyatta-op-cmd-wrapper disconnect interface pppoe0
/opt/vyatta/bin/vyatta-op-cmd-wrapper connect interface pppoe0
```
`disconnect` tears down the PPP session; `connect` re-dials immediately. A
short pause between them (a few seconds, or minutes for cautious ISPs) lets the
old session finish teardown before redialing. This briefly drops the whole WAN
uplink and may change the public IPv4 and delegated IPv6 `/60`; in-flight
sessions and port-forwarded services are interrupted until the new session is
up.
The `zhiqiang` account logs into `vbash`, not the EdgeOS CLI, so operational
commands must be invoked through `/opt/vyatta/bin/vyatta-op-cmd-wrapper` and
depend on its passwordless `sudo`. The `ubnt` account lands directly in the
operational CLI, where the same commands are entered without the wrapper.
`show`/`configure` are interactive-only aliases (from
`/etc/bash_completion.d/vyatta-{op,cfg}`, loaded via `~/.bashrc`), so a
non-interactive `ssh ubnt@… 'show …'` also fails — from a script use the op
wrapper above, or `_vyatta_op_run` after sourcing `vyatta-op` with
`vyatta_op_templates=/opt/vyatta/share/vyatta-op/templates`.
## Maintenance notes
- EdgeOS writes persistent changes through its configuration tree: enter
@@ -100,3 +141,23 @@ from `192.168.55.254` reached the UniFi controller at `192.168.66.46` with
3/3 ICMP replies. This supports the AP Inform path to
`192.168.66.46:9080`; the controller listener and an online LAN55 AP provide
the corresponding application-level evidence. No firewall changes were made.
IPv6 was re-verified by read-only SSH on 2026-08-20 during the UniFi AP/AC
check: the IPv6 routing table shows connected `/64`s on `eth0` (LAN66) and
`switch0` (LAN55) plus `::/0` via `pppoe0`; both UniFi APs obtained SLAAC
addresses from the router's RAs. No configuration changes were made.
**DHCP 保留 `matter` 失效(2026-08-21 发现,2026-08-23 复核仍未生效,W1N-207):**
静态映射 `matter` → .45 / MAC `34:98:7a:27:10:bc`,但该灯泡一直以**动态租约**拿
`.148`hostname `matter`2026-08-23 09:02 时租约当日 04:40 已续租)。保留 .45 从未
被租出。2026-08-23 复核补充:另一盏工作灯泡的 MAC 已变为 `fc:e8:c0:25:a1:f0`
(动态 `.146`hostname `espressif`),原「把 MAC 改为 `34:98:7a:27:7f:08`」的修正
建议已过时(该灯泡已离网)。处置:删除该保留,或按现用 MAC(`.148`
`34:98:7a:27:10:bc` / `.146``fc:e8:c0:25:a1:f0`)重建,**未执行**。
**SE5420 部署 + switch0 单上联(2026-08-22 只读核实):** `switch0` 成员口
`eth1` link up、`eth2`/`eth3` down(单上联);SE5420 管理面 `192.168.66.253`
在线(TP-Link OUI `f8:c9:03`:80/:443);ARP 显示 LAN55 主机(hass `.11`
M3 `.248`、SmartThings `.48`、UAP-AC-Lite `.5`)全部经 switch0 可达。含义:
`switch0` 不再是 LAN55 的全量抓包点(同段有线单播在 SE5420 本地交换),详见
[runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md)。
+641 -4
View File
@@ -1,3 +1,4 @@
[hosts/hass.windy.lan.md#8DF6]
# hass.windy.lan — Home Assistant (HAOS)
## Role and access
@@ -8,7 +9,7 @@
| IPv4 | `192.168.55.11` (LAN55) |
| DNS | `hass.windy.lan` (AdGuard rewrite on `dns.windy.lan`; legacy `hass.local` alias) |
| SSH | `ssh hassio@hass.windy.lan` |
| **Host** | **PVE VM 180 (`haos`)** — not a separate physical host (verified 2026-08-09) |
| **Host** | **x88 Pro physical box** (HAOS bare-metal, `machine: green`; verified 2026-08-18) |
| Platform | Home Assistant OS; kernel `6.1.115-haos` (aarch64) |
| Web UI | `http://hass.windy.lan:8123` (LAN); WAN port-forward `hass` on gw → `:8123` |
@@ -39,7 +40,8 @@ recovery codes in this repository.
| Interface | Address / role |
|---|---|
| `end1` | `192.168.55.11/24`; primary LAN55 address |
| `end0` | IPv4 static `192.168.55.11/24` (gw `.254`, DNS `192.168.66.36`); IPv6 SLAAC `auto` with GUA on the current PD-derived /64 (`240e:3bd:235:1fb2:*` at 2026-08-22; rotates on PPPoE redial); primary LAN55 NIC (interface name verified live 2026-08-22 — `end1` does not exist) |
| `wlan0` | Supervisor **disabled** (verified 2026-08-14, W1N-104); IPv6 remains off on this RTL8821CS radio |
| `wg0` | `10.13.13.2/32`; WireGuard (add-on / integration tunnel) |
| `hassio` / `docker0` | internal HAOS Docker bridges (`172.30.32.0/23`, `172.30.232.0/23`) |
@@ -97,7 +99,7 @@ See `~/.config/zsh/env/local/environment.env` for the client-side setting.
## Safe verification
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan 'hostname; ip -4 addr show end1'
ssh -o BatchMode=yes hassio@hass.windy.lan 'hostname; ip -4 addr show end0'
```
From a LAN client, confirm DNS and UI reachability:
@@ -107,8 +109,643 @@ getent hosts hass.windy.lan
# expect 192.168.55.11
```
## Local patches (custom components)
### Manual custom-component install (this host)
Home Assistant loads custom integrations from
`<config>/custom_components/<domain>/` (HAOS: `/config``/homeassistant`).
A folder named after the integration domain, containing at least
`manifest.json` and `__init__.py`, is enough; Core must be restarted after
copying files. Official HA lookup order:
`<config>/custom_components/<domain>` then built-in
`homeassistant/components/<domain>`.
See [Integration file structure](https://developers.home-assistant.io/docs/creating_integration_file_structure).
This host **does not git-clone** custom components. The live tree is a file
copy. Do not `git pull` on HA.
**Official plugin path** (from
[windyboy/china_southern_power_grid_stat README](https://github.com/windyboy/china_southern_power_grid_stat)):
HACS **or** [手动下载安装](https://github.com/windyboy/china_southern_power_grid_stat/releases).
This host uses the latter. Releases here have no uploaded zip assets; use
GitHub's **Source code (zip)** / zipball of the tag.
**UI (Samba / File editor / Studio Code Server):**
1. Download Source code (zip) from the GitHub Release.
2. Extract. Copy only the inner
`custom_components/china_southern_power_grid_stat/` tree — not the repo
root, not a nested extra folder.
3. Place it at `/config/custom_components/china_southern_power_grid_stat/`.
4. Restart Core (**Settings → System → Restart**).
5. First install only: **Settings → Devices & services → Add integration**.
**SSH from the workstation** (verified 2026-08-14, W1N-107). Replace `v1.3.1`
with the tag being installed:
```bash
TAG=v1.3.1
STAGE=/tmp/csg-${TAG}-deploy
mkdir -p "$STAGE"
gh api "repos/windyboy/china_southern_power_grid_stat/zipball/${TAG}" \
> "$STAGE/src.zip"
unzip -q "$STAGE/src.zip" -d "$STAGE"
SRC=$(find "$STAGE" -type d -path '*/custom_components/china_southern_power_grid_stat' | head -1)
# expect .../custom_components/china_southern_power_grid_stat
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i mkdir -p /homeassistant/.csg-backups &&
sudo -n -i cp -a /homeassistant/custom_components/china_southern_power_grid_stat \
/homeassistant/.csg-backups/china_southern_power_grid_stat.bak-$(date +%Y%m%d)-manual'
rsync -a --delete \
-e 'ssh -o BatchMode=yes' \
"$SRC/" \
hassio@hass.windy.lan:/homeassistant/custom_components/china_southern_power_grid_stat/
# --delete cannot remove Core-owned __pycache__; wipe as root, then restart
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i rm -rf /homeassistant/custom_components/china_southern_power_grid_stat/__pycache__ \
/homeassistant/custom_components/china_southern_power_grid_stat/*/__pycache__ &&
sudo -n -i ha core restart'
```
Wait until Core is up (`ha core info` returns, typically 12 min; this CLI
build does not print a `state:` field).
Then:
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i cat /homeassistant/custom_components/china_southern_power_grid_stat/manifest.json'
# version must match the tag
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i ha core logs -n 2500' | grep -E 'china_southern_power_grid_stat|cannot pickle' || true
```
**Host constraints (do not skip):**
- Backups **must** live in `/homeassistant/.csg-backups/`. A `*.bak-*`
directory next to the live folder is scanned as the same domain and Core
fails with `No module named '...bak-YYYYMMDD-...'`.
- Do not install this fork via HACS on this host. HACS still tracks
`CubicPill/china_southern_power_grid_stat` `v1.2.0`; a HACS update would
overwrite the live copy.
- First poll after restart can time out to CSG over IPv4; if this-month
sensors stay `unknown` while last-month filled, reload the config entry
(UI: integration → Reload, or supervisor
`POST /core/api/config/config_entries/entry/<id>/reload`).
- `runbooks/scripts/ha-maintenance.sh --restart-core --yes` can print
nothing and exit 1 in under a second **without restarting Core**. The
wrapper's ssh line discards stderr (`2>/dev/null`); with `pipefail`,
an ssh failure yields empty stdout + exit 1 before any remote command
runs. Do not treat that as a completed restart. Confirm with elapsed
time (~2 min for a real restart) and `ha core info`. Prefer
`ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core restart'`.
Full command family: [runbooks/home-assistant-maintenance.md](../runbooks/home-assistant-maintenance.md).
### `china_southern_power_grid_stat` live tree
**v1.3.2** (`934f58c`, verified 2026-08-15, W1N-118): manual zipball of
GitHub release
[v1.3.2](https://github.com/windyboy/china_southern_power_grid_stat/releases/tag/v1.3.2)
copied to `/config/custom_components/china_southern_power_grid_stat`.
Earlier trees: v1.3.1/`55a293fc` (W1N-107), v1.3.0/`69f13c90` (W1N-106),
`a433e8c` (W1N-105), `de01914` (W1N-103), `eb8b174` (W1N-102). Backups:
`/homeassistant/.csg-backups/` (w1n102/104/105/106/107/118).
v1.3.0 crashed the coordinator on first refresh
(`TypeError: cannot pickle 'mappingproxy' object` in
`copy.deepcopy(self._config)` under Python 3.14 / HA 2026.8.1). v1.3.1
wraps those `deepcopy` calls with `dict(...)`. Post-restart 22:13 CST:
entry `loaded`, no pickle traceback. Native this-month sensors filled after
reloading entry `01KGCQDSZCF523A9X6SV3BZ1B9` (`ip_family: ipv4`). Native
cost/ladder sensors can stay `unknown` because CSG
`get_month_daily_cost_detail` returns a marketing-system SQL error; the
dashboard uses template ladder/cost entities instead. Do not change
`templates/csg_sensors.yaml` or the 电力监控 dashboard for an install.
**`templates/csg_sensors.yaml` hardened 2026-08-29 (W1N-239):** added
`availability` templates to all 12 `csg_*` sensors (numeric sensors can't
render `unknown`/`unavailable` in `state`; availability suppresses
rendering instead — native CSG down ⇒ derived sensors show `unavailable`,
no more fake zeros / "一档" / `0%`). `csg_yesterday_kwh` now falls back to
`last_month_by_day`'s last entry when `this_month_by_day` is empty (month
start); ladder constants (`t1/t2/p1/p2/p3`) deduped into per-block
`variables:` (Block B + Block D); `csg_mom_change` parses `date`
defensively. Backup:
`/homeassistant/.csg-backups/csg_sensors.yaml.bak-20260829-w1n239`.
**Verified:** `ha core check` OK; Core restart required (trigger-based
template blocks don't settle on `template.reload` — W1N-114 precedent);
post-restart all 12 entities numeric & consistent (302.47 kWh→180.28 元,
324.03 kWh→194.06 元, mom_change -3.6%, yesterday 7.66 kWh/2026-08-28),
no template errors in Core logs.
**`csg_sensors.yaml` off-by-one fixed 2026-08-29 (W1N-241):** CSG data
lags 1 day (`sum(this_month_by_day)` == `this_month_total_usage`, data
stops at yesterday), but templates used `now().day` as "days elapsed" →
`csg_predicted_usage` underestimated ~1 daily avg (~3%) and
`csg_mom_change` compared this-month 28 days vs last-month 29 days
(-3.6% vs true -0.3%). Both now derive the day number from
`this_month_by_day[-1].date` (fallback `now().day` when empty). Added
`sensor.csg_this_month_daily_avg` (month-to-date avg, 302.47/28=10.8) and
`sensor.csg_prediction_progress` (usage/predicted %, 90.3) in Block C
(trigger adds `csg_predicted_usage`). Backup:
`/homeassistant/.csg-backups/csg_sensors.yaml.bak-20260829-w1n241`.
**Verified (8/29):** predicted 324.03→334.81, mom_change -3.6→-0.3,
daily_avg 10.8, progress 90.3, predicted_cost 194.06→200.94 (334.81 kWh
ladder), ladder cost 180.28 unchanged, `ha core check` OK after restart,
no template errors; 14 csg_* entities total.
**电力监控面板(`lovelace.dashboard_unknown` / view `power-monitor`
updated 2026-08-29 (W1N-240 + W1N-242):** 「本月累计」gauge 对齐夏季阶梯:
`max:650`、segments `0/260/600`(绿/橙/红 = 一/二/三档;冬季 11-01 需切
`max:450``0/200/400`**seasonal switch point**,见下文)。「📊 统计
数据」卡新增本年/去年 4 行(原生传感器,口径标注「电费(账单)」、本年
「(至今)」)+ 本月日均/预测进度 2 行(`csg_this_month_daily_avg` /
`csg_prediction_progress`W1N-242);面板共引用 **20** 个实体。改前备份:
`/homeassistant/.lovelace-backups/dashboard-unknown-power-monitor-20260829-204845.json`
W1N-240)、`-20260829-210708.json`W1N-242
(改法:WS `lovelace/config/save`,参数 `url_path: dashboard-unknown` +
`config`;勿直改 `.storage/`)。验证:WS 读回 18→20 实体 diff ✓、gauge
配置一致 ✓、URL `http://hass.windy.lan:8123/dashboard-unknown/power-monitor`
**`csg_sensors.yaml` W1N-242:** `csg_predicted_usage` /
`csg_mom_change` / `csg_this_month_daily_avg` 三处取 `days[-1]` 前补
`sort(attribute='date')`(与 `csg_yesterday_kwh` 一致,防上游乱序取错
数据日)。备份 `csg_sensors.yaml.bak-20260829-w1n242`。验证:Core
restart 后回归值不变(334.81 / -0.3 / 10.8 / 90.3 / 200.94 / 180.28)。
**CSG 面板重构 2026-09-04VPS-90,先核对计价后展示层改动):** 核对
`power-monitor` 计价与 8 月账单一致(198.65 vs 账单 198.64,差 ≤0.01 元,
因模板用公众圆整价 0.589/0.639/0.889、账单用 6 位精确价),不改阶梯常量。
改动:① `csg_sensors.yaml` Block B 新增
`sensor.csg_this_month_avg_price`(本月阶梯电费÷本月用电,`元/kWh`
availability 照 W1N-239 惯例;**csg_* 实体 14→15**);② 面板改名「环比上月」
→「环比上月同期」;glance「本月/上月」grid 去重为单卡「上月」(本月用电/电费
行归 💰核心数据卡);⚡阶梯电价卡加「本月实际均价」行(当前档位/当前电价/
本月实际均价/档位剩余;面板唯一实体引用 20→21);③ `automations.yaml`
2 条提醒:`automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`10-25 09:00
`automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`4-25 09:00)经
`matrix_e2ee.send_message` 提醒切 gauge。④ 金额单位混排(原生 CNY vs 模板
元)**维持**`config/entity_registry/update` 拒绝自定义文本单位
`extra keys not allowed … Got '元'`),用户确认接受。备份:
`.lovelace-backups/dashboard-unknown-power-monitor-20260904-204757-pre-refactor.json`
`.csg-backups/csg_sensors.yaml.bak-20260904-204757-pre-refactor`(及
`-205301-pre-avgprice`)、`.automations-backups/automations.yaml.bak-*`
**WS 改法(2026.8,本机实测)**core/主机 python 无 ws 库、core 容器内经
supervisor 代理 WS 被拒(loop prevention),用
`docker run --rm --network host -e SUPERVISOR_TOKEN`supervisor 镜像
`aarch64-hassio-supervisor:2026.08.0`)连 `ws://172.30.32.2/core/websocket`
命令名 `lovelace/config`(读)+ `lovelace/config/save`(写),
`lovelace/config/get` 已不存在(unknown_command)。验证:新实体
0.589 元/kWh、15 个 csg_* 数值齐全、回归值不变(14.09/198.65/331.22/
304.99/181.89)、automations on、`ha core check` OK、日志无 template 错误。
**CSG 长期归档(W1N-243, 2026-08-29:** scribe 库新增 `csg_history`
表(逐日 usage/cost/ladder/balance + 逐月累计;2026-07-01 起回填,永久),
由 TimescaleDB 每日任务 **1008** `csg_daily_snapshot()`22:30
Asia/Shanghai**TS job 非 pg_cron**upsert 维护。日费用在原生
`latest_day_cost` 缺失时回退 = 昨日用电 × 当前档费率(模板
`csg_current_ladder_tariff` 0.639);月费用回退模板
`csg_this_month_ladder_cost`。**语义**day 行 usage/cost 为该日值,
ladder/balance 为 22:30 快照值。详见 [hosts/pgdb.md](../hosts/pgdb.md)。
> **Seasonal gauge switch (W1N-240 已知事项):** 每年 **11-01**
> `power-monitor` 视图「本月累计」gauge 切到冬季 `max:450` /
> `0/200/400`**5-01** 切回夏季 `max:650` / `0/260/600`(与模板
> `now().month` 季节逻辑对齐;模板常量在 Block B/D `variables`)。
> **提醒 automation2026-09-04 起,VPS-90:**
> `automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`10-25)与
> `automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`4-2509:00 经
> `matrix_e2ee.send_message` 发操作步骤提醒;gauge 的 max/segments 无法
> 模板化,仍需人工改卡配置。
Home PPPoE IPv4 to CSG is still blackholed (`curl -4` to `218.19.148.218:443`
times out). `end0` IPv6 is enabled (`ipv6.method: auto`); from HA,
`curl -6 https://95598.csg.cn` returns HTTP 200 via `240e:f9:8060::1:16`.
**`tianqi` weather recorder patch (verified 2026-08-13, W1N-75):**
`/config/custom_components/tianqi/weather.py` has a local patch adding
`_unrecorded_attributes = frozenset({"hourly_temperature", "hourly_skycon",
"hourly_cloudrate", "hourly_precipitation"})` to the `WeatherEntity` class.
Without it, weather.guangzhou's state attributes (~19 KB, dominated by the 4
hourly_* arrays of up to 48 entries) exceed the recorder 16384-byte limit, so
the recorder drops **all** attributes for the entity and logs
`Recorder.db_schema: State attributes for weather.guangzhou exceed maximum
size of 16384 bytes`. The patch excludes only the 4 arrays from recording
(live state unchanged; other attributes still stored; ~6.3 KB payload). Backup
at `weather.py.bak-w1n75`. **Re-apply after any `tianqi` component update.**
The `_unrecorded_attributes` mechanism exists in Core 2026.8.1
(`Entity.__init_subclass__``state_info["unrecorded_attributes"]`, consumed
by recorder `shared_attrs_bytes_from_event`).
### `matrix_e2ee` live tree (E2E Matrix bot, verified 2026-08-20)
**v0.3.12** (tag `v0.3.12`; feat — Matrix activity events
`matrix_e2ee_message_received` / `matrix_e2ee_verification_done` + push
diagnostics; v0.3.9 added Connection health binary sensor, SAS/command
allowlist split, URL normalization, single-entry enforcement):
source copy from `/home/windy/project/ha-matrix-e2ee` `ea421ed` (tag
`v0.3.12`) deployed 2026-08-20 via SSH rsync from workstation (upgraded
from v0.3.2, backup `matrix_e2ee.bak-20260820-v0.3.2`).
Custom **`matrix_e2ee`** integration — **Config Flow** (UI). See
[docs/home-assistant-matrix.md](../docs/home-assistant-matrix.md).
**Update runbook:** [runbooks/matrix-e2ee-update.md](../runbooks/matrix-e2ee-update.md).
Earlier: v0.3.2 (tag `v0.3.2`, W1N-182/#34: wizard waits for inbound SAS
emojis) deployed 2026-08-18 from `d35c484` (backup
`matrix_e2ee.bak-20260818-v0.3.1`); v0.3.1 (GitHub #33: peer-initiated
verification wizard fix) deployed 2026-08-18 from `d22e935` (backup
`matrix_e2ee.bak-20260818-v0.3.0`); v0.3.0 (W1N-180/#32: bot-initiated
verification wizard; W1N-179/#31 `receive_mac_event` cancel-state fix)
deployed 2026-08-18 from `216cc99` (backup
`matrix_e2ee.bak-20260818-v0.2.10`).
- Bot `@hass:chans.xyz` reused (E2EE device `rO1R915ncu`). Config Entry
`01M04D7C1M4T2GX5VPG7NVQ7GV` (`source: import`, `state: loaded`). All
settings via **Settings → Devices & Services → Matrix E2EE → Configure**.
- Config Entry options: `allowed_rooms` `["!gidvAzpDzwtzfEDrqu:chans.xyz", "!boxfylDSzOvrWkcsyY:chans.xyz"]`,
`allowed_users` `["@zhiqiang:chans.xyz"]`, `command_prefix` `"!"`.
**`verification_peer_users` not set** (v0.3.9+ SAS allowlist split from
`allowed_users`, W1N-156): defaults to empty → only the bot's own account
may drive SAS; `@zhiqiang` is denied until the option is added via
Settings → Devices & Services → Matrix E2EE → Configure.
- Storage: `/config/.storage/matrix_e2ee_session.json` +
`/config/.storage/matrix_e2ee_store/`. Backups:
`/homeassistant/.matrix-e2ee-backups/` (incl. `matrix_e2ee.bak-20260820-v0.3.2`,
`matrix_e2ee.bak-20260818-v0.3.1`,
`matrix_e2ee.bak-20260818-v0.3.0`,
`matrix_e2ee.bak-20260818-v0.2.10`,
`matrix_e2ee.bak-20260816-v0.2.9`, `matrix_e2ee.bak-20260816-v0.2.8`);
full HA backup slugs `3d9d36db` (pre-v0.1.4) + `9f223f35` (pre-v0.2.0).
- v0.3.12: Matrix activity events + push diagnostics
(`matrix_e2ee_message_received` / `matrix_e2ee_verification_done`).
v0.3.9: Connection health binary sensor (W1N-185/#40), config-entry
diagnostics (W1N-184/#39), SAS/command allowlist split
`verification_peer_users` (W1N-156/#41), SAS/sync logs demoted
warning→info/debug (W1N-188/#38), URL normalization + single-entry
enforcement (W1N-190/#42).
v0.3.8: `m.key.verification.done` handshake for request-based SAS
(W1N-183/#35).
v0.3.2: wizard waits for inbound SAS emojis before the compare step
(W1N-182/#34).
v0.3.1: verification wizard waits for a peer-initiated inbound SAS instead
of the bot starting SAS (GitHub #33).
v0.3.0: bot-initiated device verification wizard (W1N-180/#32).
v0.2.11: `receive_mac_event` no longer overrides canceled state (W1N-179/#31).
- v0.2.9: restore SAS emoji rendering after vodozemac migration (W1N-175/#29).
v0.2.8: SAS commitment unpadded base64 for Element interop (W1N-174/#28).
v0.2.7: SAS cancel code/reason logging. v0.2.6: verification state logging +
request→ready bridge. v0.2.4: `_patch_nio_sas_timeout()` +
`_repair_dropped_start()`; `VERIFICATION_TIMEOUT_SECONDS` 600→240.
- Automation `1761188403590`「Matrix 聊天关卫生间灯」: trigger
`matrix_e2ee_command` (command `关卫生间灯`), actions `light.turn_off` +
`matrix_e2ee.send_message` (room `!gidvAzpDzwtzfEDrqu`).
- **SAS not yet completed:** every device requires explicit `confirm_verification`.
Encrypted-room commands stay fail-closed until `@zhiqiang`'s device is verified.
Since v0.3.9 the SAS driver gate uses `verification_peer_users` (empty on
this host) instead of `allowed_users` — add `@zhiqiang:chans.xyz` there
before retrying the wizard. Three paths available: SAS manual confirm,
fingerprint, or the device verification wizard (v0.3.0 bot-initiated,
reworked in v0.3.1/v0.3.2 to wait for a peer-initiated inbound SAS from
Element with emoji comparison), see
[docs/home-assistant-matrix.md § Device verification](../docs/home-assistant-matrix.md).
### Scribe long-term history (3.8.0 setup 2026-08-29; 4.4.0 verified 2026-09-13)
- **Scribe 4.4.0** (`/homeassistant/custom_components/scribe/`, HACS repo
`jonathan-gtd/scribe`, = latest stable 2026-09-12; upgraded 2026-09-13 together
with Core 2026.9.1 / HAOS 18.2), configured from
`/homeassistant/scribe.yaml` — W1N-238 moved the block out of
`configuration.yaml` on 2026-08-29 (main config now carries
`scribe: !include scribe.yaml`; content moved verbatim; backup
`configuration.yaml.bak-20260829-201724-w1n238`). Config entry
`01KC2VFJWEQ3XDHY6TQKHPDVRB`, `source: import` — UI "Configure → Advanced"
edits are overridden by the YAML on restart; treat YAML as authoritative.
- TimescaleDB at `192.168.55.15:5432/scribe` (DB user `hass`; host in inventory,
see [hosts/pgdb.md](../hosts/pgdb.md)). Database re-initialized 2026-08-29 14:06 CST
(user-handled; earlier `relation "entities" does not exist` errors resolved).
Health: `binary_sensor.scribe_database_connection`.
- 2026-08-29 config applied (backup `/homeassistant/configuration.yaml.bak-20260829-scribe`):
- `record_events: true` with `include_events` whitelist: `automation_triggered`,
`matrix_e2ee_command`, `matrix_e2ee_message_received`,
`matrix_e2ee_verification_done`, `script_started`, `tag_scanned`,
`mobile_app_notification_action`, `homeassistant_start`, `homeassistant_stop`.
- State noise trimmed: `exclude_domains` update/button; glob
`sensor.zigbee2mqtt_bridge_*`; 4 hassio cpu/mem-percent entities.
- Global `exclude_attributes` drops tianqi `hourly_*` arrays (~19 KB/state —
the recorder-side `_unrecorded_attributes` patch does not apply to Scribe).
- `enable_stats_io` + `enable_stats_size` on → 14 `sensor.scribe_*` stats
entities (`scribe_states_written`, `scribe_events_written`, rates, sizes).
- Verified post-restart 14:23 CST: writer started, `scribe_events_written=1`
(homeassistant_start), states ~110/min, buffer 3, no scribe log errors.
- **4.x upgrade核对 2026-09-13(只读 + 一处配置变更)**live `manifest.json` =
4.4.0。两个 4.0 breaking change 在本机都不需要动作——数据库是 3.x 结构
`states_raw` PK `(metadata_id, time)` 在,4.2.0 的启动态去重因此可用),
TimescaleDB 2.29.2 已装。4.1.0 修了 `db_url` 优先级,YAML 里的
`!secret scribe_url` 现在是权威。`scribe.yaml` 现有键在 4.4.0 全部仍然合法
(未知键被忽略,`extra=vol.ALLOW_EXTRA`)。**配置优先级 YAML > entry
`options` > entry `data` > 默认值**,而 `_resolve_settings` 读的是
`hass.data[DOMAIN]["yaml_config"]`(只有 `async_setup` 会写),所以
**YAML 改动必须重启 Corereload config entry 不重读 YAML。**
- **`stats_io_interval: 300`2026-09-13 添加**,备份
`/homeassistant/scribe.yaml.bak-20260913-191558`)。4.4.0 不再让 HA 每 30s
轮询 I/O 统计传感器,改由集成自己每 60s 发布,间隔成为配置项。scribe 自己的
传感器此前是本机自写历史的主要来源(变更前 24h:11 019 / 87 461 行状态 =
12.6%),60s → 300s 把这部分降约 5 倍(每个 I/O 传感器约 1440 → 288 行/天)。
验证:`ha core check` OK;重启 88s`ScribeWriter started successfully`
无 scribe error/warningscribe Repairs 问题 0 条;传感器发布间隔实测正好
300s11:19:26 → 11:24:26 UTC)。
- **Retention 现在可用但刻意不设**`retention_states` / `retention_events`
4.0.0)按间隔丢 chunk,留空 = 永久保留,符合本机定位(Scribe 是永久归档,
recorder 保留 365 天)。注意 retention 是**绕过** entry `data` 副本读取的
`from_entry_data=False`),所以删掉 YAML 行即撤销策略。`db_schema`
`enable_rollups``scribe.purge` 同样未用:图表走 `sensor_minute` +
`timescale_database_reader`(见 [hosts/pgdb.md](pgdb.md)),不吃 scribe 自己的
视图,配置里也没有任何 `scribe.query` 调用。`flush_interval` 仍是 entry
`data` 钉住的 5s——上游下一个版本把默认改成 30s,但 entry 值优先,要采用只能
在 YAML 显式写 `flush_interval: 30`
- Recorder stays external-Postgres with `purge_keep_days: 365` (W1N-243,
2026-08-29, raised from 30 — ~300 MB/yr, 1% of the 30G pgdb disk) for
native UI per-change history; Scribe is the permanent archive. Long-term
statistics stay permanent (not purged by `purge_keep_days`). Note:
extending retention does **not** recover pre-2026-08-29 raw history
(already purged); only `csg_history` day/month values cover that period.
### Config layout: scribe.yaml + templates/ merge (W1N-238, verified 2026-08-29)
- `configuration.yaml` line 29: `scribe: !include scribe.yaml`; line 9:
`template: !include_dir_merge_list templates`. No `packages/`.
- `scribe.yaml` (config root): the Scribe block, content identical to the
former inline one; import semantics unchanged.
- `templates/`: `csg_sensors.yaml` (12 template sensors, top-level **list**)
+ `quick_sensors.yaml` (scaffold for Quick-derived `quick_*` sensors, empty
list with convention header). **`!include_dir_merge_list` merges per-file
lists; non-list files are silently skipped** — every file in `templates/`
must be a top-level list (`- sensor:` blocks). Directory include only picks
up `*.yaml`, so the `.bak` / `.pre-*` backups in the dir are ignored. After
adding sensors, verify template-platform entity count = 12 + N (entity
registry `platform: template`).
- Convention (per review + W1N-233): pure sums/averages stay min_max helpers
(e.g. `sensor.dang_qian_zong_gong_lu`); only template-logic derivations
(ladder pricing, cross-entity conditions) go into `quick_sensors.yaml`.
- Post-change verification 20:19 CST: `ha core check` ok, 92 s restart
(2026.8.3), `binary_sensor.scribe_database_connection` on,
`scribe_states_written` 18581→19426 growing, template entities still 12,
csg sensors numeric, no scribe/template log errors.
### Timescale Plotly card + database reader (verified 2026-08-29)
Chart stack over the Scribe TimescaleDB archive. Upstream pair (no HACS;
manual copies): reader `remmob/timescale_database_reader` **v1.1.0** (main
`bb8776a`) + card `remmob/timescale-plotly-card` **2.2.0** (main `217961d`).
- **Reader integration**: `/homeassistant/custom_components/timescale_database_reader/`.
Config entry `01M165P77QT1FQEAVPNZHDT82W` ("Scribe", `source: user`): connects
`hass@192.168.55.15:5432/scribe` (credentials = `secrets.yaml` `scribe_url`),
`table: sensor_minute`. Exposes no entities/services — it serves WS command
`timescale/query` (window ≤ 365 d, ≤ 50 000 rows, `downsample` bucket seconds).
Benign startup warning `Error executing test query: column "time" does not
exist`: the self-test SQL assumes the LTSS column name; the scribe table uses
`minute` — real queries work (verified: 70 rows for a live power sensor).
- **Card**: `/homeassistant/www/community/timescale-plotly-card/timescale-plotly-card.js`
(root-owned, same convention as HACS dirs). Lovelace resource (storage)
id `2e360d17b5aa4ce59c2fd13c43b51215`
`/hacsfiles/timescale-plotly-card/timescale-plotly-card.js`, type `module`.
Card config matches the entry by `database: scribe` (name from the reader
entry). Updates: replace the file, resource URL unchanged — browsers need a
hard refresh or a bumped `?v=` query on the resource URL.
- **pgdb side** (`sensor_minute_aggregate` cagg + `sensor_minute` hypertable +
every-minute refresh job): see [hosts/pgdb.md](pgdb.md) § Databases.
- **Agent-side HA WebSocket without a long-lived token** (verified 2026-08-29):
connect `ws://supervisor/core/websocket` with header
`Authorization: Bearer $SUPERVISOR_TOKEN`, then send
`{"type":"auth","access_token":"$SUPERVISOR_TOKEN"}` — the Supervisor proxy
swaps it for a core token (works as the internal Supervisor admin user). Note
`lovelace/resources/create` in HA 2026.8 takes `res_type` (NOT
`resource_type`).
- Scribe stores numeric sensor values in `states_raw.value` with `state` NULL,
so `sensor_minute.state` shows `'0'` for numeric sensors; the card plots
`avg_state` (from `value`) — expected, not a bug.
- **Quick 仪表盘(`dashboard-quick`)图表套件**2026-08-29 创建,经 WS
`lovelace/config/save` 写入;W1N-230 修复 + W1N-231 round-2 改进):
5 张 timescale 卡——大功率电器/常驻负载功率(按量级拆图,避免尖峰压扁
<70 W 基线)、按插座用电量(`energy_mode` + cumulative/diff,数据质量前提
见 pgdb 的 refresh 过程补丁)、室内外温湿度(温度左轴/湿度右轴,4 位置同色
配对)、人体感应活动状态(3 个 `motion_state`banded `state_map`
none/small/medium/large → 0-11per-entity `line_color` 红/蓝/绿)。
空调实体引用为 `kong_diao_*``kong_tiao` 是笔误,W1N-230 修复;`grep -c
kong_tiao` 应为 0)。灯区:2×2 嵌套 grid(`grid_options: {columns: "full"}`
内层 `columns: 2`+ 4 卡统一 `mushroom-light-card`(显式 name、
`use_light_color: false`、内联亮度/色温控制),heading icon
`mdi:lightbulb-group`。heading badges:环境 4 温度(迷你/mini数显/数显/广州)、
大功率电器 空调/电脑当前功率、常驻负载 总功率
`sensor.dang_qian_zong_gong_lu`min_max **sum** helper`round_digits: 0`
任一源掉线 fail-closed → unknown)。常驻负载图卡级 `fill: 'tozeroy'` +
冰箱/主网络 per-entity `fill_color`(线色 20% 透明)+ 其余 5 条 `fill: false`
per-entity fill 逐系列退出,卡 JS `seriesConfig.fill !== false`)。
布局:视图 `type: sections` + `max_columns: 4`;灯/用电/环境/人体感应
`column_span: 4`,功率两图拆两个 `column_span: 2` 分区**并排**(等高 280px
桌面并排、手机回落堆叠;去卡内 title 省半宽图垂直空间)。
**分区/卡片是两套尺寸键,不可混用**:分区宽 = `column_span`
`hui-sections-view.ts` 缺省按 1 列渲染,绝不省略);卡片宽 =
`grid_options: {columns: <n|"full">}``hui-card.ts` 只读 `config.grid_options`
写在卡片上的 `column_span` 被静默忽略;缺省 12 列,分区内格 = 12 × 分区
span,故 span-4 分区里缺省卡片只有 1/4 宽)。
修改前备份:`/homeassistant/.lovelace-backups/dashboard-quick-*.json`
W1N-230 修复: `20260829-190256`round-2 改进: `20260829-194040`)。
- **Quick 时间范围扩容 (2026-09-13, VPS-92)**: 用户反馈「48 小时不够」。
各 timescale 卡可选档上调——大功率电器/常驻负载 `…,24h``+3d,7d`
环境 `6h,12h,24h,48h``+7d,14d,30d`;人体感应 `…,24h``+3d,7d`
用电量(按插座) `energy_time_ranges` `today,week,month,custom``+3mo`
**默认档未改**6h / 6h / today / 24h / 12h)。卡片 JS 只接受
`<n>m|<n>h|<n>d``parseDurationToMs` 正则 `/^(\d+)(m|h|d)$/`
仅 m/h/d,无 w)与命名档 `today|week|month|3mo|6mo|year|years|custom`
`energy_mode` 卡必须用后者。**数据下界注意**:scribe `sensor_minute`
目前最早只到 **2026-08-29**,所以 >15d 的档(14d 边缘、30d 明显)前半段
会是空白,等归档继续累积才好看。备份
`.lovelace-backups/dashboard-quick-20260913-190912-pre-timerange.json`
### 地图仪表盘:CARTO keyed tiles via `custom:map-card` (verified 2026-08-30, W1N-261)
- **背景:** CARTO 自 2026-08-26 起对无 key 栅格瓦片打 "API KEY REQUIRED"
水印,内置地图卡/zone 编辑器全部受影响。Core 2026.8.3 的 `MapCardConfig`
**没有任何瓦片配置项**frontend 20260729.7 源码核对:
`setup-leaflet-map.ts` 硬编码 CARTO voyager URL)。上游修复是 2026.9.0b1
起改用 OSMF 矢量瓦片(frontend PR #53816),stable 预计 2026-09-02 前后。
- **变更:** 「地图」仪表盘(url_path `map`storage)唯一 map 卡替换为
`custom:map-card`[nathan-gs/ha-map-card](https://github.com/nathan-gs/ha-map-card)
**v1.16.0**,手动安装非 HACS):`tile_layer_url` =
`https://{s}.basemaps.cartocdn.com/rastertiles/voyager/{z}/{x}/{y}.png?key=<CARTO_KEY>`
(配 `tile_layer_options: {subdomains: abcd, maxZoom: 20}` + OSM/CARTO
attribution)。实体不变:2 person + 4 zonezone 用 `display: icon` +
`circle: auto`circle 读实体 `radius` 属性画半径圈)。
- **CARTO key 是 secret**: 只存在于服务端 lovelace 存储(dashboard `map`
的卡片配置)和用户本人处;勿写入本仓库或 Linear。
- **文件/资源:** `/homeassistant/www/community/ha-map-card/map-card.js`
root:root 644678554 Bsha256
`f30dfb606e858d2216d5198d8cf758ce956d127006ebd7d66d4329153a247ec2`);
Lovelace resourcestorageid `9d2b50b52c60420d89ebd041f722cf60`
`/hacsfiles/ha-map-card/map-card.js`type moduleWS
`lovelace/resources/create`2026.8 参数名 `res_type`)。升级 = 手动替换
该文件(不在 HACS 管理下,浏览器需强刷)。
- **备份:** `/homeassistant/.lovelace-backups/dashboard-map-map-20260830-133714.json`
(还原 = 把备份里的 `views[0].cards[0]` 写回后再 WS `lovelace/config/save`
url_path `map`)。
- **验证 8/30:** 同瓦片无 key=水印 / 带 key=干净(256×256 PNG 视觉对比);
resource HTTP 200 text/javascriptWS 读回卡片配置(type/entities/key/
attribution/options)全部符合;HA 主机 `curl -4` 带 key 瓦片 200。
- **Follow-up:** Core 升 2026.9.0 stable 后内置地图/zone 编辑器自动切
OSMF 矢量瓦片;届时可保留 custom 卡(继续 keyed CARTO)或用备份还原
内置卡。zone 编辑器等其余内置地图的水印在 2026.9 前无解。
## Known issues
**Bluetooth hci0 instability — RTL8821CS (verified 2026-08-13, W1N-74):**
The local Bluetooth controller hci0 is an **RTL8821CS** combo chip on the
x88 Pro board. Kernel logs show recurring `hci0: hardware error 0x00`,
`Opcode 0x200c tx timeout` (HCI_LE_Set_Scan_Parameters), `Unable to disable
scanning: -110`, `Peer device has reset` — the chip hardware-stalls during
active scanning. HA's `bluetooth_auto_recovery` power-cycle then times out
after 5 s and retries every ~2 min:
`bluetooth_auto_recovery.recover: Could not reset the power state of the
Bluetooth adapter hci0 ... due to timeout after 5 seconds`. The HAOS image
already ships custom systemd units to cope (`x88-bt-hci-recovery.service` and
a "Patch HA Bluetooth scanner mode for x88 RTL8821CS" service, visible in host
journal). **No user impact:** there are **no BLE entities** in HA
(xiaomi_ble / bthome / led_ble / bluetooth / esphome domains are all empty;
platforms merely load from stray advertisements). Real IoT devices are Zigbee
(via Zigbee2MQTT) or WiFi/MQTT/cloud. An ESPHome Bluetooth-proxy ESP32
(`/config/esphome/bluetooth.yaml`, bluetooth_proxy: active, WiFi `ubnt-haas`)
is configured but currently offline (ESPHome add-on stopped, port 6053
unreachable) and produced no entities. Follow-up (optional): disable the
local adapter and rely on the ESPHome proxy, or stop the bluetooth
integration entirely.
**eMMC disk lifetime 10% (verified 2026-08-13, W1N-76):** `ha host info`
reports `disk_life_time: 10` — the boot eMMC (`/dev/mmcblk2`, CJTD4R
`0xacacc064`, 64 GB) has ~10% life left. `disk_free: 40.2/56.4 GB`. Full
backup `pre-maintenance-20260813` (slug `411a4ba5`, 144.26 MB) taken
2026-08-13 covers current config; monitor `disk_life_time` on each health
snapshot and plan a disk replacement / data-disk migration before the eMMC
fails.
## Matter Server (verified 2026-08-21)
- Add-on `core_matter_server` (`homeassistant/aarch64-addon-matter-server`) runs the Matter
commissioner on this host (host networking; add-on container `app_core_matter_server`).
- **After the ISP PD prefix rotates (PPPoE redial), the add-on can cache a stale IPv6 GUA
in its mDNS advertisement** — clients trying that dead address make Matter
commissioning/connection fail. Fix: restart the add-on so it re-enumerates addresses:
`ssh hassio@hass.windy.lan 'sudo -n -i ha apps restart core_matter_server'`
(`ha addons restart ...` also works; "addons" is deprecated in favor of "apps").
- Verified 2026-08-21 (W1N-207): stale `240e:3bd:234:2f22:*` AAAA in mDNS removed by
restart; advertisement now carries only current GUA `240e:3bd:235:1fb2:*` + link-local;
CASE sessions with Aqara M3 / SmartThings hubs resumed over IPv6 link-local.
> **Open items (2026-08-21, W1N-207):** a phone on LAN55 was querying five known
> `_matter._tcp` instances of which only HA answered — the other Matter nodes are
> offline / not announcing (device-side; user to confirm power/Wi-Fi). HA's IPv6
> default route via NetworkManager was observed missing once (curl -6 intermittent,
> while ping6 and `curl -6 --noproxy` work) — not the Matter root cause; re-check
> on the next health snapshot.
Verified 2026-08-23 (read-only, W1N-207): add-on `started`, version `9.0.4`, no
update pending; current GUA `240e:3bd:238:4812:*` (PD rotated again since 08-22)
advertised correctly over v4+v6. Both ESP32-C2 bulbs now announce `_matter._tcp`
(multi-fabric, including this host's fabric `DCE86145C137AF0E`) — but they
**refuse TCP 5540 on IPv4 and IPv6**, so matter-server holds **zero established
:5540 sessions** (device-side failure mode C; no errors logged — see
[docs/matter-pairing-troubleshoot.md §8](../docs/matter-pairing-troubleshoot.md)).
## 马桶换气电源(Matter 插座,半计量)+ 电量估算 (2026-09-13)
**设备**Matter `Smart Plug`SIXWGH`model_id 3596`hw 1.0 / sw 1.3.0),node 18
(0x12)`device_id 5ef1850953466d6e7a9c6b901fbebe1c`config entry
`01JF51VQ48PGJGXX3RNAG6MVAA`,区域**卫生间** (`wei_sheng_jian`)label `power`
2026-09-13 17:58 CST 配对。实体:
`switch.wei_sheng_jian_ma_tong_huan_qi_dian_yuan`(插座)、
`sensor.…_dian_yuan`(电源 W)、`sensor.…_dian_ya`(电压 V)、
`sensor.…_you_gong_dian_liu`(有功电流 A)、`sensor.…_dian_li`(电力 kWh
**永久 unknown**)。
**根因(实测 Matter 属性,node 18**:电量簇 0x0091 `FeatureMap = 13`
(IMPE|CUME|PERE,即**声明**支持导入/累计/周期电量),但
`CumulativeEnergyImported (0x0001)` 恒为 `null``PeriodicEnergyImported
(0x0003)` 带载也恒为 `{Energy: 0}``CumulativeEnergyExported (0x0002)`
不存在(EXPE 未声明,自洽)。HA 只用 `CumulativeEnergyImported` 建能量实体
`components/matter/sensor.py:1083``allow_none_value=True`)→ 该实体
**永远不会出数**。**功率计量本身正常**:0x0090 `FeatureMap = 2` (ALTC)
Voltage / ActiveCurrent / ActivePower 都随负载变化(实测 220.3 V / 118 mA /
24.7 WHA `电源` 0.0→24.9 W 有历史)。厂商 `update` 实体报无新固件。
**处理(方案 A:功率积分补电量)**:
- 新建 **Integration (Riemann sum) 辅助元素**config entry
`01M2D53T188FW8WEC547ENHSVH`domain `integration`state `loaded`),
source `sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_dian_yuan`
`method: trapezoidal``unit_prefix: k``unit_time: h``round: 3`
`max_sub_interval: 60s`
- 实体 `sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_energy`(创建时 HA
自动生成 `…_dian_yuan_ma_tong_huan_qi_dian_yuan_dian_liang`,随后立即
`config/entity_registry/update` 改名为 `<插座>_energy` 以对齐约定;
该实体新建、无引用,改名安全),friendly name「马桶换气电源 电力」,
unit kWh、`device_class: energy`、**`state_class: total`**——能源仪表盘
允许 `TOTAL``TOTAL_INCREASING``components/energy/validate.py:279`)。
- **能源仪表盘** (`/energy`)grid 源 `[8]``…_dian_li` 改为 `…_energy`
其余 8 条插座源未动。注意这 9 条「插座」全部以 `type: grid` 注册,被当作
全屋用电代理;`switch` 卡片所在的 Grid 卡片此前第 9 行是空的,即本次修复点。
- **Quick 仪表盘**:「用电量(按插座)」图第 11 项由 `…_dian_li` 改为
`…_energy`;新增 `column_span: 2` 的「开关」区块(heading + tile
`switch.…` + `toggle` feature + 功率徽标)→ 视图 6→7 分区。
**口径警告**`…_energy` 是**估算值**Riemann 积分,只在 HA 运行期间累计、
非账单级),与另外 8 个原生计量插座的累计电量口径不同;功率传感器更新
间隔约 510 s(实测 24.9/24.8/25.0 W 抖动),加 `max_sub_interval: 60s`
保证静默时也继续累计。
**Agent 侧建辅助元素的方法(2026-09-13 实测)**HA 的 config flow 走
**REST**WS 只有 `config_entries/flow/progress|subscribe`,没有 start)。
经 supervisor 代理即可,无需 HA 长连接/长寿命 token:
```bash
# SUPERVISOR_TOKEN 由 sudo -n -i 提供
curl -s -X POST -H "Authorization: Bearer $SUPERVISOR_TOKEN" \
-H "Content-Type: application/json" -d '{"handler":"integration"}' \
http://supervisor/core/api/config/config_entries/flow # → {flow_id, step_id:"user", data_schema}
curl -s -X POST -H "Authorization: Bearer $SUPERVISOR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name":"…","source":"sensor.x","method":"trapezoidal","round":3,
"unit_prefix":"k","unit_time":"h","max_sub_interval":{"minutes":1}}' \
http://supervisor/core/api/config/config_entries/flow/<flow_id> # → create_entry
```
`auth/long_lived_access_token` 在 supervisor 代理身份下**失败**
`unknown_error`),故无法用长寿命 token 开浏览器会话;`DurationSelector`
的值是 `{"minutes":1}` 形式(`cv.time_period`)。
**备份/回滚**`.lovelace-backups/dashboard-quick-20260913-181251-pre-ma-tong-plug.json`
(改动前原件)、`…-20260913-183210-pre-repoint.json`(改名/换源前);
`.ha-backups/energy-20260913-183135-pre-ma-tong-repoint.json`(能源 prefs)。
回滚 = 把能源 prefs 的源 [8] 指回 `…_dian_li` + 还原 Quick 面板 JSON
如需彻底放弃估算电量 = 删除 config entry `01M2D53T188FW8WEC547ENHSVH`
**验证 (2026-09-13 18:3x)**`…_energy` 0.002→0.003 kWh 且随 24.6 W 负载
增长(换气扇关掉后回落 0.0 W,累计值保留);`recorder/list_statistic_ids`
已含该实体;Quick 面板 WS 读回 7 分区、用电量图 11 项指向新实体、旧
`_dian_li` 引用 0 处;能源 prefs 读回 9 源、第 9 条为新实体。
**`energy/validate` 已全绿**9 源 0 issue):创建后 ~5 min 内曾报
`statistics_not_defined`(recorder 的统计任务周期是 5 min,`statistics_meta`
行由该任务建立),18:39 复核时已自动消失——建辅助元素后**不要**把这条
瞬时告警当作失败。
## Related docs
- [runbooks/home-assistant-maintenance.md](../runbooks/home-assistant-maintenance.md) — `ha` CLI maintenance runbook + [script](../runbooks/scripts/ha-maintenance.sh)
- [runbooks/home-assistant-maintenance.md](../runbooks/home-assistant-maintenance.md) — `ha` CLI maintenance runbook + [script](../runbooks/scripts/ha-maintenance.sh); custom-component zip install is §7
- [docs/lan-overview.md](../docs/lan-overview.md) — LAN map and gw port-forward
- [hosts/dns.windy.lan.md](dns.windy.lan.md) — `hass.windy.lan` / `hass.local` rewrites
+18
View File
@@ -185,6 +185,23 @@ dig @202.91.35.141 SOA wsvc.info +short
On-server docs: `/opt/pdns/README.md`, `CHANGELOG.md`.
## Disk / logging (VPS-81, 2026-09-02)
Root disk cleanup performed (runbook: [host-disk-cleanup](../runbooks/host-disk-cleanup.md)):
- Root `/` (20G vda1): 76% used → **38% used** (15G → 7.1G; free 4.7G → 12G).
- **AGH log flood root cause fixed**: `/opt/adguard/conf/AdGuardHome.yaml`
`log.verbose: true → false` (backup `AdGuardHome.yaml.bak-20260902-vps81`).
Verbose debug was streaming to stderr → container `json.log` (~120MB/day);
`log.file: ""` makes AGH's own rotation keys inert. Restart only (no recreate).
- Journald capped: `/etc/systemd/journald.conf.d/00-vps81.conf`
`SystemMaxUse=200M`; journal vacuumed to ~96M.
- Docker: engine **29.7.2**; 14 unused images removed (kept `pdns-auth-50:5.0.5`
rollback pin); 12 orphan anonymous volumes + build cache pruned. In-use
volumes intact (`pdns_dbdata`, `b594d738…` PG data, `e855d078…` backup).
- Follow-up: re-check AGH `json.log` growth **2026-09-09** (one-week checkpoint);
global docker log rotation only if still needed.
## Verified
Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded.
@@ -194,3 +211,4 @@ Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded.
- `only-notify=` + `also-notify=202.91.35.141`; MASTER `domains.master` cleared
- https://pdns.wsvc.info → **302**; https://pgweb.wsvc.info → **401**
- Hardening backlog: API/DB credential rotation + TSIG rotate (see upstream doc)
- 2026-09-02 (VPS-81): post-cleanup verified — 10 containers Up (adguardhome healthy), DNS SOA/NS + web endpoints OK; see Disk/logging section above.
+59
View File
@@ -0,0 +1,59 @@
# pgdb — TimescaleDB (PG18, Docker)
## Role and access
| Item | Value |
|---|---|
| Role | TimescaleDB PostgreSQL 18 (Docker) — Home Assistant recorder 后端(`hass`/`scribe` 库) |
| IPv4 | `192.168.55.15` (LAN55) |
| DNS | (none) |
| SSH | `ssh -4 windy@192.168.55.15`key auth 已验证可用 2026-08-29agent 沙箱用 `ssh -F /dev/null -o BatchMode=yes`password auth 亦可) |
| Host | PVE 管理的 QEMU VMi440FX**VMID 100**),Debian 13 (trixie),内核 6.12.105;宿主机 **pve2 `192.168.55.25`**Proxmox 9.2.2SSH `root@192.168.55.25``onboot: 1`QEMU guest agent 已装;2026-08-31 补记) |
| Resources | 3 GB RAM08-30 13:58 由 2G 上调、删除 balloon/ksm/shares 后重启生效)/ 30 GB disk26 G 空闲) |
| Docker | 29.7.2;容器 `timescaledb` = `timescale/timescaledb:latest-pg18`PG **18.6** + TimescaleDB **2.29.2**Apache-2.0 版) |
| Ports | `192.168.55.15:5432`PGIPv4 only);`192.168.55.15:8081`pgweb GUIbasic auth |
## Databases
| DB | Owner | Size | 用途 |
|---|---|---|---|
| `hass` | hass | ~406 MB2026-09-13 | HA recorderstates/events/statistics),客户端 HAOS `192.168.55.11` |
| `scribe` | postgres | ~2.6 GB2026-09-13 | HA scribe 集成(entities/areas/devices 注册表同步 + `states_raw`/`events` hypertable + `csg_history` 长期归档表);体积由 `sensor_minute` 图表管道主导(2.26 GB),见 Known issues |
| `postgres` | postgres | ~9 MB | 默认库 |
## Ops notes
- **Docker compose 管理**2026-08-29 改造):`/opt/database/docker-compose.yml`(源码在仓库 `compose/pgdb/`+ `/opt/database/.env`0600,密钥)+ `/opt/database/pgweb-bookmarks/`0600bookmark 含 DB 密码)。三个服务:
| 服务 | 镜像 | 端口 | 说明 |
|---|---|---|---|
| `timescaledb` | `timescale/timescaledb:latest-pg18` | `192.168.55.15:5432`IPv4 only | PG 18.6 + TS 2.29.2healthcheck pg_isready`restart: unless-stopped` |
| `pgweb` | `sosedoff/pgweb:latest`v0.17.0 | `192.168.55.15:8081` | Web GUIhttp://192.168.55.15:8081basic auth(用户名/密码见 .env `PGWEB_AUTH_USER/PASS`);`--readonly --sessions --bookmarks-only --bookmarks-dir /bookmarks`v0.17.0 不读 PGWEB_BOOKMARKS_DIR env,必须用 flag);bookmarks = hass/scribe |
| `pg-backup` | `prodrigestivill/postgres-backup-local:latest`(=PG18 客户端) | — | 每日 02:00(`TZ=Asia/Shanghai`,本地时区)`pg_dump -Fc` 三库 → `/opt/database/backups/{daily,weekly,monthly}`;保留 7 天/4 周/6 月;`BACKUP_ON_START` |
- **数据盘**`/dev/sdb1`32G ext4label `pgdata`)挂载 `/srv/pgdata`fstab 按 `UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86`defaults,noatime)持久化(2026-08-29 迁移)。容器 bind mount `/srv/pgdata:/var/lib/postgresql`
- 容器内 postgres 用户 uid/gid = **70**(Debian 系,非 999);迁移数据后需 `chown -R 70:70`
- **密码**:postgres 超级用户已换强密码(hex,存 `/opt/database/.env` 06002026-08-29)。HA 用 `hass` 角色不受影响。
- **备份**:由 `pg-backup` 容器接管(2026-08-29),宿主机 cron 与 `/opt/database/pg-backup.sh` 已退役。恢复用 `pg_restore`custom format)——2026-08-29 已实测还原 hass 库 dumpstates 10014 行)成功。
- **认证**:外部连接 scram-sha-256(密码必填,改密码有效);容器内 loopback 为 trust(官方镜像默认)。
- **回滚**:旧启动命令保留在 `/opt/database/run`(容器无状态,数据在 /srv/pgdata);旧匿名卷 `9375195843b950f4e04c34872409ca095e1136520dd019a8e86e2794be06c236`(根盘 ~82M)保留作兜底,确认稳定后可 `docker volume rm`
- **开机自愈**2026-08-30):新增 systemd oneshot `pgdb-compose.service`enabled,源码在仓库 `compose/pgdb/pgdb-compose.service`):`After=network-online.target docker.service`,开机后幂等执行 `docker compose up -d`,重试直到 `192.168.55.15:5432` 监听,重试耗尽 `--force-recreate` 兜底(数据在 bind mount,无损)。原因:2026-08-30 开机竞态——docker 恢复容器时 VM IP 尚未可绑(EADDRNOTAVAIL),timescaledb/pgweb 启动失败且 docker 不重试。手动重跑:`sudo systemctl restart pgdb-compose.service`
- 本机无防火墙(ufw/nft/iptables 均未装)——待办:如要彻底隔离可加 ufw 白名单 192.168.55.11。
- `/opt/database/backups/` 根下残留 `*-2026-08-29_1359.dump`(compose 化之前旧备份机制产物)与 `backup.log`——健康检查只看 `daily/`,残留可清理。
- **Runbooks**[pgdb-health](../runbooks/pgdb-health.md)(只读健康检查)、[pgdb-restore](../runbooks/pgdb-restore.md)pg_restore 还原)、[pgdb-update](../runbooks/pgdb-update.md)(镜像/compose 升级)。
- **CSG 长期归档(2026-08-29, W1N-243**`csg_history` 表(`period date / kind('day'|'month') / usage_kwh / cost / ladder / balance / updated_at`PK(period,kind)`GRANT SELECT TO hass`)保存南方电网有价值数据:day = 逐日(昨日用电/费用/阶梯/余额,2026-07-01 起),month = 当月累计(用电/费用,2025-01 起)。由 TimescaleDB 每日任务 **1008** `csg_daily_snapshot()`22:30 Asia/Shanghai**TS job 非 pg_cron**,本库未装 pg_cronupsert 维护:取「最新有值行」防瞬态 unknown 竞态;日费用缺原生 `latest_day_cost` 时回退 = 昨日用电 × 当前档费率(模板 `csg_current_ladder_tariff` 0.639);月费用回退模板 `csg_this_month_ladder_cost`。验证:day 08-28 = 7.66 / 4.89474 / 二档 / 0month 08 = 302.47 / 180.28。回填来源:集成 attributes `history_data`59 天)+ `by_month`(19 月)——08-29 前唯一残存历史。回滚:`DROP TABLE csg_history` + `SELECT delete_job(1008)`
## Known issues
- 2026-09-13**`sensor_minute` 体积构成与压缩窗口(只读诊断,暂不处理)**。`scribe` 库 2.6 GB = `sensor_minute` **2.26 GB**850 万行 / 16 天,约 5659 万行/天 = 331 实体 × 1440 分钟 LOCF+ `states_raw` 290 MB + `events` 1.5 MB`hass` 库另 406 MB。2.26 GB 中 1.50 GB 是 chunk `[09-03,09-10]`、0.76 GB 是 `[09-10,09-17]`,**都还没到压缩窗口**——TimescaleDB 的 `compress_after`**chunk 结束时间**判断(09-10 结束 + 7 天 = **09-17** 才合格),所以「7 天 chunk + 7 天 compress_after」的设计下限就是盘上常驻近 14 天原始数据;已压缩的 `[08-27,09-03]` 从 1.04 GB → **1.5 MB**LOCF 重复度极高,~700:1)。任务 1005 健康(30 成功 / 0 失败,最近 09-13 04:18 跑过但无合格 chunk);1002/1003/1006/1007 亦全 Success。稳态估算 ≈ 2 个未压缩 chunk(3–4.5 GB+ 已压缩归档(约 1.5 MB/周 ≈ 80 MB/年)≈ **45 GB 平台期**pgdata 卷 32 G 当前用 3.2 G,可用 27 G,**无需处理**。复查点 **2026-09-17 之后**`_hyper_4_6_chunk` 应转为 `compressed=true` 且库体积回落;若仍为 false 才需动手(手动 `compress_chunk()` 或调小 `compress_after`)。可选调优:chunk 间隔 7 天 → 1 天 + `compress_after` → 2 天,把常驻未压缩量压到 <1 GB(`set_chunk_time_interval` 只对新 chunk 生效,旧 chunk 不重切)。诊断命令:`select chunk_name, is_compressed from timescaledb_information.chunks where hypertable_name='sensor_minute';` + `pg_database_size('scribe')`。**注意:这是 pgdb 侧对象,HA/scribe 的 `retention_states` 管不到它;HA 侧唯一杠杆是少记/少画(等于砍图)。**
- 2026-08-29HA 侧 HACS 集成 `custom_components.scribe`YAML `scribe: db_url:`,连 `scribe` 库)建表被拒(`permission denied for schema public`hass 无 CREATE 权限),之后持续报 `relation "entities" does not exist`。**已解决**:① `GRANT CREATE ON SCHEMA public TO hass;`scribe 库)② 重启 HA Core 触发重跑建表。重启后自动创建 `entities`1591 行)/`users`/`areas`/`devices`/`integrations`/`states_raw` 表并启用 TimescaleDB 时间序列能力。报错已停止(最后一条 06:06 UTC),`states_raw` 持续写入。2026-08-29 复查:scribe 现有**两个** hypertable——`states_raw`segmentby `metadata_id`、orderby `time`)与 `events`segmentby `event_type`、orderby `time`),均 1 维 `time`;压缩已配置(`timescaledb_information.compression_settings` 可见对应行;2.29.x 该视图无 `compression_enabled` 列)。
- 2026-08-29**timescale reader 图表对象**(配套 hass 的 `timescale_database_reader` 集成 + `timescale-plotly-card`,上游 SQL `remmob/timescale_database_reader` `SQL/scribe/01+02` @ `bb8776a`,以 postgres 执行):`sensor_minute_aggregate` 连续聚合(1 分钟桶,last(state)/last(value),实时聚合开启)+ `sensor_minute_aggregate_entity` 视图(join `entities`+ `sensor_minute` hypertable`minute`/`entity_id`/`state`/`value`,LOCF 前向填充)。任务:1005 `sensor_minute` 压缩(7 天)、1006 `sensor_minute` 保留(10 年)、1007 `every_minute_refresh` 每分钟增量刷新(含 5 分钟回溯窗口修正)。授权:`GRANT SELECT ON sensor_minute_aggregate, sensor_minute_aggregate_entity, sensor_minute, entities TO hass`。种子 19529 行(331 实体,自首个数据点起)。**刻意跳过**了上游脚本对 `states_raw` 的 3 个月保留 + 压缩策略语句——与"`states_raw` 永久归档"定位冲突,如需磁盘回收属用户决策(scribe 自己的压缩任务 1000/1001 未动)。
- 2026-08-29**`sensor_minute_refresh` 本地补丁(类比 tianqi 补丁,重跑上游 02 SQL 后需重打)**:值 CASE 的 `ELSE 0``ELSE NULL`。原因:scribe 对 unavailable 分钟 value 为 NULL,上游刷新过程兜底写 0;对差分模式的用电图,0→计数器回升会把插座的**生命周期累计值**(最高 1588 kWh)算进掉线那一小时。同日一次性清理既有脏 0:头部占位行 DELETE 505 行(各实体首次非零分钟之前的 value=0);`sensor.%_energy` 与温湿度实体的 value=0 → NULL(10+16 行,物理上不可能的真 0,图表渲染为断点)。功率实体的中途 0 是真实待机读数,保留。
- `hass` 库的 recorder 表仍为普通表(无 hypertable);`scribe` 集成负责时间序列历史(`states_raw` + `events` hypertable)。
## Verification history
- 2026-08-31**13:58 重启根因确认,非停电**W1N-263):pve2`192.168.55.25`)任务日志显示 08-30 **13:58:00 `root@pam` 在 PVE Web UI 修改 VM 100 配置**`-delete allow-ksm,balloon,shares -memory 3072`),**13:58:06 点 Reboot**`qmreboot` → 客机 13:58:08 干净 ACPI 关机 → 13:58:13 自动重启)。宿主机全程在线(08-30 09:00 开机至今连续运行 1d12h+),`.66.26` PVE 及各 VM 均无重启——排除停电。HA recorder 在窗口(13:58:4647)报 2 次 `Connection refused`,DB 恢复后自动重连,**无数据丢失**(`hass.states`/`scribe.states_raw` 13:5514:02 逐分钟无缺口,recorder 内存队列吸收回写)。13:58:47 三容器已起,13:58:56 自愈单元 `pgdb-compose.service` 执行成功——本次自愈按设计工作。同日下午 12:54–12:55 另有一次**客机内自重启**(无 PVE 任务,工作站 SSH 会话相邻)。08-29 22:19→08-30 09:00 宿主机停机 10h41m 为**干净关机**(systemd 有序关闭,非停电)。
- 2026-08-30**开机竞态故障 + 修复**W1N-260):09:01 开机后 docker 恢复容器时绑定 `192.168.55.15:5432/8081` 失败(EADDRNOTAVAIL)→ timescaledb/pgweb 停摆至 12:16pg-backup 开机备份失败(解析不到 timescaledb)→ unhealthy。12:22 `docker compose up -d --force-recreate` 修复(三容器回 `database_default`、端口发布、今日备份、pgweb 恢复);用户重启 HA Core 后写入管道恢复。12:43 新增开机自愈 unit `pgdb-compose.service`enabled,已实测幂等 reconcile)。pgdb-health 8 项全绿。
- 2026-08-29:首次检查(只读)+ 修复 scribe 权限 + 安装夜间备份。见 Linear vps 项目登记。
- 2026-08-29**compose 改造完成**W1N-227,用户已验收):裸 `docker run``/opt/database/docker-compose.yml` 三服务(timescaledb + pgweb + pg-backup);superuser 换强密码;端口收紧 IPv4;备份容器化(TZ=Asia/Shanghaicron 02:00 本地);`pg_restore` 还原实测通过;pgweb UI 用户确认可查 hass/scribe 数据。源码在仓库 `compose/pgdb/`
- 2026-08-29**运维 runbook 落地**W1N-228,已验收):新增 `runbooks/pgdb-health.md`(只读,8 项诊断全绿)、`pgdb-restore.md`(流程式,temp-DB 安全还原 + 审批门)、`pgdb-update.md`(门控命令式,回滚=/opt/database/run + 旧卷);README 索引与 validate-repo.sh 分类同步更新;runbook 命令已对活主机逐条实测(含 `pg_restore -l` 校验当日 dump)。同日修正:SSH key auth 可用(facts 原记"密钥未安装"已过时);scribe 新增 `events` hypertable。
- 2026-08-29**CSG 长期归档 + recorder 365d**W1N-243):建 `csg_history` 表 + attributes 回填(逐日 59 + 逐月 19)+ 每日任务 1008(函数 v2:最新有值行读取、日费用阶梯回退);hass `purge_keep_days` 30→365(备份 `configuration.yaml.bak-20260829-purge365`)。见 Linear vps W1N-243。
+45 -1
View File
@@ -24,6 +24,7 @@ ssh -4 windy@synapse.chans.xyz
| DB | ESS embedded PostgreSQL 17 (PVC 20Gi, local-path) |
| Cache | ESS embedded Redis (PVC 2Gi) |
| Chart | `oci://ghcr.io/element-hq/ess-helm/matrix-stack`, version `26.7.2` |
| Plane | Helm `plane-ce-1.8.0` (app `v1.4.1`), namespace `plane` — self-hosted Plane project management |
### Matrix service endpoints
@@ -51,8 +52,50 @@ All other ports internal only (no K3s API, no database, no Redis exposed).
- `ess` — all ESS workloads (Synapse, MAS, Element, Postgres, Redis, HAProxy)
- `matrix-system` — cluster base resources (ResourceQuota, LimitRange, mrtc-placeholder)
- `plane` — Plane project management (Helm release `plane-app`)
- `cert-manager` — cert-manager
## Plane (project management)
Self-hosted [Plane](https://github.com/makeplane/plane) on the same K3s node, deployed via the official `plane-ce` Helm chart.
| Item | Detail |
|------|--------|
| Release | `plane-app` (ns `plane`), chart `plane-ce-1.8.0`, app `v1.4.1`, revision 1 |
| URL | https://plane.chans.xyz |
| Install date | 2026-09-01 |
| Values source | `/home/windy/plane-k3s/values.yaml` (plain file, not a git repo) |
| Images | `artifacts.plane.so/makeplane/*` (`plane-frontend`, `plane-backend`, `plane-admin`, `plane-live`), pullPolicy `Always` |
| Ingress | Traefik `IngressRoute` `plane-app-ingress``/`→web, `/api` `/auth`→api, `/spaces`→space, `/god-mode`→admin, `/live`→live, `/uploads`→minio; `maxRequestBodyBytes` 20Mi |
| TLS | Own namespace `Issuer` `plane-app-cert-issuer` (HTTP-01, LE prod, `admin@chans.xyz`); cert `plane-app-ssl-cert` (CN `plane.chans.xyz`) |
| DB | Bundled Postgres `15.7-alpine` (PVC 5Gi, local-path) |
| Cache/queue | Bundled Redis (PVC 100Mi), RabbitMQ `3.13.6-management-alpine` (PVC 100Mi) |
| Storage | Bundled MinIO (`minio/minio:latest`, root user `admin`, PVC 5Gi) — S3 for uploads/docs |
| Resources | Every workload: cpu 50m/500m, mem 50Mi/1000Mi, replicas 1 |
| SMTP | Not configured (no `smtp` values) — Plane invites/password resets won't email yet |
Workloads (all 1/1 Running): 7 Deployments (`plane-app-{admin,api,beat-worker,live,space,web,worker}-wl`) + 4 StatefulSets (`plane-app-{minio,pgdb,rabbitmq,redis}-wl`); init Jobs `api-migrate-1` / `minio-bucket-1` Completed. All PVCs Bound on `local-path` (root disk).
### Plane configuration notes
- **`planeVersion: v1.4.1`** pinned in values.yaml; chart tracks Plane's own tags.
- **Secrets**: Helm-generated Opaque secrets (`plane-app-app-secrets`, `-doc-store-secrets`, `-pgdb-secrets`, `-rabbitmq-secrets`, `-live-secrets`); `requireExplicitSecrets: false`. Values live in `$SECRET_KEY`, `DATABASE_URL`, `AMQP_URL`, `REDIS_URL` etc.
- **Sentry / CORS**: `sentry_dsn` and `cors_allowed_origins` empty (defaults fine for single-host).
- **MinIO is `latest` tag** — pin a version for reproducibility.
- **Backup**: NOT covered by `/var/backups/matrix` (which is paused anyway) — Plane Postgres/MinIO PVCs have no backup tier yet.
### Plane verification
```bash
# Release + workloads
sudo helm list -A
sudo k3s kubectl -n plane get deploy,sts,pods -o wide
# Cert + ingress
sudo k3s kubectl -n plane get certificate,ingressroute
# Endpoint
curl -4 -s -o /dev/null -w '%{http_code}\n' https://plane.chans.xyz/
```
## Local backup
| Item | Detail |
@@ -62,7 +105,7 @@ All other ports internal only (no K3s API, no database, no Redis exposed).
| Retention | 7 days |
| Disk warning | 80% (healthcheck), 90% (backup stops) |
| Content | Planned: PostgreSQL `synapse` + `mas` logical dumps, media store archive, `/etc/matrix-bootstrap` |
| Status | **Not operational** — no current Matrix backup or recovery tier |
| Status | **Not operational** — no current Matrix backup or recovery tier. **Plane data (its own Postgres + MinIO PVCs in ns `plane`) is also not covered by any backup.** |
## Health checks
@@ -101,5 +144,6 @@ diagnosis and imperative recovery work.
- MatrixRTC / Element Call / LiveKit / Coturn not deployed (`mrtc.chans.xyz` reserved only)
- SMTP email not yet configured (requires manual secret bootstrap followed by a
reviewed Ansible stack deployment)
- Plane `minio` image uses `latest` tag (pin a version)
- No off-site Restic backup
- Single-node K3s (no HA for control plane)
+5
View File
@@ -53,6 +53,11 @@ of `8080`. During adoption or recovery, use the documented `:9080/inform` URL;
an AP left on `:8080` can remain reachable by ping and SSH while showing
offline in the controller.
IPv6 is enabled on the controller's `Default` network (`ipv6_enabled: true`,
client assignment SLAAC; RA is served by `gw`, so `ipv6_interface_type` is
`none`); both managed APs hold global SLAAC addresses — verified 2026-08-20.
See [docs/unifi-network.md](../docs/unifi-network.md).
## Safe reconciliation and verification
```bash
+26 -12
View File
@@ -5,12 +5,12 @@
| Role | Multi-service VPS (Vaultwarden, Traefik, Soft Serve, …) |
| SSH | `ssh -4 windy@us2.wsvc.info` (prefer IPv4 from WSL) |
| IPv4 | `193.9.44.165` |
| Also DNS | `auth.wsvc.info` → this host; `repo.windy.me` → this host (Soft Serve) |
| Also DNS | `auth.wsvc.info` → this host; `repo.windy.me` → this host (Gitea) |
| Public HTTPS | Traefik on `:80` / `:443` (`/opt/traefik`) |
## Vaultwarden (Bitwarden-compatible)
**Status: operational** (Postgres live, HTTPS 200, healthy containers, SMTP AUTH OK — last probe 2026-08-01 18:55 CST).
**Status: operational** (Postgres live, HTTPS 200, healthy containers, SMTP AUTH OK — last probe 2026-08-29).
Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
@@ -21,9 +21,9 @@ Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
| Env file | `/opt/vaultwarden/.env` |
| Admin overrides | `/opt/vaultwarden/vw-data/config.json` (**wins over env**) |
| Public URL / `DOMAIN` | `https://auth.wsvc.info` |
| Image | `vaultwarden/server:1.37.1` (pinned) |
| Image | `vaultwarden/server:1.37.2` (pinned) |
| Live DB | **Postgres 16** (`vw-db` / service `pg`) via compose `DATABASE_URL` |
| Data (probe) | users=1, ciphers=1327 |
| Data (probe) | users=1, ciphers=1360 |
| Cold SQLite | `backups/sqlite-cold/db.sqlite3.pre-pg-20260801` (not used live) |
| Pre-migrate backup | `backups/pre-pg-migrate-20260801_161204/` |
| Data dir | `./vw-data``/data` (attachments, rsa keys, `config.json`) |
@@ -50,7 +50,7 @@ Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
| Container | Status |
|-----------|--------|
| `vaultwarden` | Up (healthy), `vaultwarden/server:1.37.1` |
| `vaultwarden` | Up (healthy), `vaultwarden/server:1.37.2` |
| `vw-db` | Up (healthy) — **live** Postgres |
| `vaultwarden-backup` | Up (`pg_dump`) |
| `vaultwarden-pgweb` | Exited (profile `debug`) |
@@ -80,19 +80,33 @@ ansible-playbook playbooks/compose-reconcile.yml --limit vaultwarden \
| Container | Status | Image / notes |
|-----------|--------|---------------|
| `soft-serve` | Up | `ghcr.io/charmbracelet/soft-serve:latest` (`repo.windy.me:2222`) |
| `gitea` | Up | `gitea/gitea@sha256:1c17ecaead42e…` (1.27.3-rootless) — SSH `repo.windy.me:2222`, web `https://repo.windy.me` |
| `gitea-backup` | Up | alpine + sqlite3/rsync sidecar (daily backup 02:00 / prune 03:00, crond) |
### Gitea (replaced Soft Serve 2026-09-18; [Plane VPS-94](https://plane.chans.xyz))
- `/opt/gitea/compose.yml` (+ `Dockerfile.backup`, `scripts/`, `config/app.ini`, `data/`, `secrets/`, `backups/`); 镜像: `compose/gitea/`(参考, 服务器文件为准)
- **rootless 镜像** uid 1000:1000; SQLite `/opt/gitea/data/data/gitea.db`; repos `/opt/gitea/data/data/git/repositories/`; app.ini `/opt/gitea/config/app.ini`(600, 含 SECRET_KEY)
- SSH: 内置 server 容器内 `:2322`(`SSH_LISTEN_PORT` 非特权), Traefik TCP entrypoint `ssh`(`:2222``gitea:2322`, `HostSNI(*)`, `tls=false`) on `vw-net`; clone URL `ssh://git@repo.windy.me:2222/windy/<repo>.git`(owner 段 `windy`)
- **host key 复用 soft-serve**(`SSH_SERVER_HOST_KEYS=/secrets/soft_serve_host_ed25519`, ed25519, 指纹 `SHA256:PdxZRe74…`): 客户端 known_hosts 零变更; 仅公钥认证(密码认证未启用)
- Web: `https://repo.windy.me`(Traefik websecure + letsencrypt); `DISABLE_REGISTRATION=true`, Actions 关闭; 管理员 `windy`(凭据仅存服务器 `/opt/gitea/.admin-credentials`, 勿入库/入 Plane)
- 仓库: 16 个(顶层 11 + `cdia/` 4 + `windyboy/go-caatsm`), 2026-09-18 自 soft-serve `push --mirror` 迁移, 逐仓 `ls-remote` ref 全集 + HEAD symref 两端一致; 可见性仅 `dotfiles-personal` private, 其余 public(与 soft-serve 现状一致)
- 备份: sidecar 每日 02:00 → `backups/gitea_<TS>/{app.ini.tar.gz, gitea.db, repos.tar.gz}`(app.ini 含恢复必需 SECRET_KEY), 03:00 prune 保留 14 份; 已验证手动备份产物 109.9M
- 回滚: `/opt/soft-serve` 未删(compose stop + sidecar 停, 数据与旧备份冻结保留), 回滚 = Traefik `:2222` 指回 `soft-serve:23231` + 客户端 remote 回改旧无 owner 段路径; 观察 24 周后清理(历史: W1N-244~248)
| `traefik` | Up | `traefik:v3.6.2` (`/opt/traefik`, public `:80`/`:443`) |
| `nghttpx-proxy` + `squid-backend` | Up | HTTP forward-proxy stack (`/opt/nghttpx`), network `nghttpx_internal-net`; details TBD |
Directories for `authelia`, `conduit`, `dendrite`, `mastodon`, `rustdesk`, `zitadel`, etc. exist under `/opt` but have no running containers; treat them as dormant, not documented services.
**Disk cleanup 2026-09-18** ([Plane vps VPS-93](https://plane.chans.xyz)): root 71% → **23%** (~33G freed) keeping soft-serve / vaultwarden / traefik (nghttpx kept running per operator choice). Removed: unused Docker images + orphan volumes (incl. `zitadel_data` 801M), dormant `/opt` dirs (dendrite + its disabled `dendrite.service` unit, mastodon, dailysync, keycloak, media-repo, authelia, conduit, npm, manager, fusion, zitadel, rustdesk), rootless podman storage (6.4G stale goauthentik), home dev caches, apt cache, journal 3.8G→162M (+`SystemMaxUse=200M` drop-in, active next boot), truncated container logs (nghttpx 550M / traefik / squid). Follow-up: nghttpx-proxy logs grow ~25M/day (INFO per-connection); root-cause log-level/rotation fix still open (needs container restart approval).
Remaining running services on this host: `gitea`, `vaultwarden` stack, `traefik`, `nghttpx-proxy` + `squid-backend` (undocumented forward proxy, `/opt/nghttpx`). `/opt/soft-serve` kept stopped as rollback (24 weeks, data intact). `/home/windy/authelia` (76M) left in place — outside approved cleanup scope.
## Verified
Last checked: **2026-08-01 18:55 CST** — operational.
Last checked: **2026-09-18** — operational; disk cleanup done (see note above, Plane vps VPS-93). Prior full probe: 2026-08-29.
- `vaultwarden` + `vw-db` healthy; `DATABASE_URL``pg:5432/vaultwarden`
- `https://auth.wsvc.info/` **200**, `/admin` **200**, `/api/config` OK (`disableUserRegistration: true`)
- Identity wrong-password → **400** business error (DB readable, not 500)
- SMTP: container → `mx2:587` OK; STARTTLS cert CN=`mx2.windy.me`; **AUTH OK** with effective `config.json` password (synced with `.env` / `.smtp-credentials`)
- LE cert CN=`auth.wsvc.info`
- PG counts: users=1, ciphers=1327
- SMTP: container → `mx2:587` OK; **AUTH OK** with effective `config.json` password (synced with `.env` / `.smtp-credentials`, fingerprint match)
- PG counts: users=1, ciphers=1360
- Image `vaultwarden/server:1.37.2` (**upgraded 2026-08-29** from 1.37.1; required for Bitwarden clients 2026.8.0+); post-upgrade 404 fixed by Traefik restart, then 200
- vps-health local check **installed 2026-08-29** (`vps-healthcheck.timer` daily 06:15 + `/usr/local/lib/vps-health/run`); `health-report.yml --limit vaultwarden` now passes (**ok**, was failing due to missing check infra + script bugs fixed: trim_blocks render, pgweb debug-profile false positive, SMTP probe moved host-side since image lacks python3)
+30 -20
View File
@@ -8,29 +8,38 @@ diagnosis and procedures that are deliberately interactive or destructive; see
For a live-verified map of the **internal LAN** (gw, gfw, dns, ubnt, APs) and
the software deployed there, see [the LAN overview](../docs/lan-overview.md).
| Host | Role | SSH | IPv4 | Status | Facts |
|------|------|-----|------|--------|-------|
| mx2.windy.me | mailcow (primary MX prio 20) | `ssh -4 windy@mx2.windy.me` | 194.163.160.244 | active | [hosts/mx2.windy.me.md](../hosts/mx2.windy.me.md) |
| us2.wsvc.info | Vaultwarden/Postgres (+ Traefik, Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | 193.9.44.165 | active | [hosts/us2.wsvc.info.md](../hosts/us2.wsvc.info.md) |
| mx.windy.me | mail (secondary MX prio 30) | TBD | see AAAA/A | stub | |
| repo.windy.me | Soft Serve git (on us2) | `ssh -p 2222 windy@repo.windy.me` | 193.9.44.165 | stub | see us2 |
| auth.wsvc.info | Vaultwarden public hostname | — (HTTPS) | → us2 | active | see us2 |
| us1.wsvc.info | PowerDNS secondary (ns2 host) | TBD | 202.91.35.141 | stub | Auth 5.0.5; see hk2 |
| us4.wsvc.info | WireGuard VPN | `ssh -4 windy@us4.wsvc.info` | 185.201.226.122 | active | [hosts/us4.wsvc.info.md](../hosts/us4.wsvc.info.md) |
| hk2.chans.xyz | PowerDNS auth (ns1) | `ssh -4 windy@hk2.chans.xyz` | 154.36.174.161 | active | [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) |
| ns1.wsvc.info | PowerDNS public NS name | — (DNS) | → hk2 `154.36.174.161` | active | see hk2 |
| ns2.wsvc.info | Secondary NS (AXFR/NOTIFY peer) | — (DNS) | → us1 `202.91.35.141` | active | see hk2 |
| pdns.wsvc.info | Poweradmin UI | — (HTTPS) | → hk2 | active | see hk2 |
| pgweb.wsvc.info | PowerDNS Postgres UI | — (HTTPS) | → hk2 | active | see hk2 |
| **synapse.chans.xyz** | Matrix homeserver (ESS: Synapse + MAS + Element) | `ssh -4 windy@synapse.chans.xyz` | `169.58.86.13` | **active** | [hosts/synapse.chans.xyz.md](../hosts/synapse.chans.xyz.md) |
| **gfw.windy.lan** | OpenWrt (ImmortalWrt) LAN gateway / OpenClash (PVE VM 140) | `ssh -4 root@192.168.66.1` | `192.168.66.1` | **active** | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy (PVE VM 120) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | **active** | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | **active** | [hosts/gw.md](../hosts/gw.md) |
| **ubnt** | UniFi Network Controller (PVE VM 160) | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | **active** | [hosts/ubnt.md](../hosts/ubnt.md) |
| **hass.windy.lan** | Home Assistant (HAOS, PVE VM 180, LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | **active** | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
**Ansible 列**`✓` = 该主机在 [`ansible/inventory/hosts.yml`](../ansible/inventory/hosts.yml)
(执行真相),用其 inventory key(见括号注)跑 playbook`—` = 不由 Ansible 管理,
原因是该平台无 ansible 覆盖或仅是公网别名/服务端点。
| Host | Role | SSH | IPv4 | Ansible | Status | Facts |
|------|------|-----|------|---------|--------|-------|
| mx2.windy.me | mailcow (primary MX prio 20) | `ssh -4 windy@mx2.windy.me` | 194.163.160.244 | ✓ (mx2) | active | [hosts/mx2.windy.me.md](../hosts/mx2.windy.me.md) |
| us2.wsvc.info | Vaultwarden/Postgres (+ Traefik, Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | 193.9.44.165 | ✓ (us2) | active | [hosts/us2.wsvc.info.md](../hosts/us2.wsvc.info.md) |
| mx.windy.me | mail (secondary MX prio 30) | TBD | see AAAA/A | — (stub) | stub | — |
| repo.windy.me | Gitea git (on us2) | `ssh -p 2222 windy@repo.windy.me` | 193.9.44.165 | — (service on us2) | stub | see us2 |
| auth.wsvc.info | Vaultwarden public hostname | — (HTTPS) | → us2 | — (alias) | active | see us2 |
| us1.wsvc.info | PowerDNS secondary (ns2 host) | TBD | 202.91.35.141 | — (stub) | stub | Auth 5.0.5; see hk2 |
| us4.wsvc.info | WireGuard VPN | `ssh -4 windy@us4.wsvc.info` | 185.201.226.122 | ✓ (us4) | active | [hosts/us4.wsvc.info.md](../hosts/us4.wsvc.info.md) |
| hk2.chans.xyz | PowerDNS auth (ns1) | `ssh -4 windy@hk2.chans.xyz` | 154.36.174.161 | ✓ (hk2) | active | [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) |
| ns1.wsvc.info | PowerDNS public NS name | — (DNS) | → hk2 `154.36.174.161` | — (alias) | active | see hk2 |
| ns2.wsvc.info | Secondary NS (AXFR/NOTIFY peer) | — (DNS) | → us1 `202.91.35.141` | — (alias) | active | see hk2 |
| pdns.wsvc.info | Poweradmin UI | — (HTTPS) | → hk2 | — (alias) | active | see hk2 |
| pgweb.wsvc.info | PowerDNS Postgres UI | — (HTTPS) | → hk2 | — (alias) | active | see hk2 |
| **synapse.chans.xyz** | Matrix homeserver (ESS: Synapse + MAS + Element) | `ssh -4 windy@synapse.chans.xyz` | `169.58.86.13` | ✓ (matrix_vps) | **active** | [hosts/synapse.chans.xyz.md](../hosts/synapse.chans.xyz.md) |
| **gfw.windy.lan** | OpenWrt (ImmortalWrt) LAN gateway / OpenClash (PVE VM 140) | `ssh -4 root@192.168.66.1` | `192.168.66.1` | — (OpenWrt, no ansible) | **active** | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy (PVE VM 120) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | ✓ (dns_windy_lan) | **active** | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | — (EdgeOS, no ansible) | **active** | [hosts/gw.md](../hosts/gw.md) |
| **ubnt** | UniFi Network Controller (PVE VM 160) | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | ✓ (ubnt) | **active** | [hosts/ubnt.md](../hosts/ubnt.md) |
| **hass.windy.lan** | Home Assistant (HAOS, x88 Pro physical box, LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | — (HAOS, no ansible) | **active** | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
| **pgdb** | TimescaleDB PG18 (Docker) — HA recorder backend (PVE VM, LAN55) | `ssh -4 windy@192.168.55.15` | `192.168.55.15` | — (no ansible) | **active** | [hosts/pgdb.md](../hosts/pgdb.md) |
`status: stub` = known to exist; fill `hosts/<name>.md` when next touched.
**命名映射**ansible inventory key ↔ 本表主机名 —— `matrix_vps``synapse.chans.xyz`
`dns_windy_lan``dns.windy.lan`。inventory key 不随主机名改(防止破坏 `--limit` 用法),
通过 inventory 内的 `display_name` 变量与文档交叉引用。
### Matrix services (synapse.chans.xyz)
| URL | Service | Notes |
@@ -39,4 +48,5 @@ the software deployed there, see [the LAN overview](../docs/lan-overview.md).
| https://synapse.chans.xyz | Synapse API | Client-Server + Federation API |
| https://account.chans.xyz | MAS | Matrix Authentication Service (local passwords) |
| https://admin.chans.xyz | Element Admin | Admin console (MAS admin auth) |
| https://plane.chans.xyz | Plane | Project management (Helm `plane-ce` v1.4.1, ns `plane`) |
| `mrtc.chans.xyz` | MatrixRTC | **Reserved** not deployed |
+42
View File
@@ -0,0 +1,42 @@
# Runbook index
Entry point for all runbooks. Before operational work, read the repo entry
[`AGENTS.md`](../AGENTS.md) and the spec [`RUNBOOKS.md`](../RUNBOOKS.md). New
runbooks start from [`_template.md`](_template.md).
## Route by intent
| Intent | Runbook | Type |
|---|---|---|
| mailcow health check | [mailcow-health.md](mailcow-health.md) | read-only |
| mailcow update | [mailcow-update.md](mailcow-update.md) | change (gated) |
| mailcow SMTP/IMAP client | [mailcow-smtp-client.md](mailcow-smtp-client.md) | reference |
| Vaultwarden health check | [vaultwarden-health.md](vaultwarden-health.md) | read-only |
| Vaultwarden SQLite→PG migrate | [vaultwarden-sqlite-to-postgres.md](vaultwarden-sqlite-to-postgres.md) | change (destructive) |
| PowerDNS health check | [pdns-health.md](pdns-health.md) | read-only |
| RustDesk health check | [rustdesk-health.md](rustdesk-health.md) | read-only |
| Matrix health check | [matrix-health.md](matrix-health.md) | read-only |
| Plane health check | [plane-health.md](plane-health.md) | read-only |
| pgdb health check | [pgdb-health.md](pgdb-health.md) | read-only |
| pgdb DB restore (pg_restore) | [pgdb-restore.md](pgdb-restore.md) | change (procedure) |
| pgdb image/compose update | [pgdb-update.md](pgdb-update.md) | change (gated) |
| AdGuard Home health check | [adguard-home-health.md](adguard-home-health.md) | read-only |
| Host disk cleanup (logs/apt/docker) | [host-disk-cleanup.md](host-disk-cleanup.md) | change (gated) |
| Matter packet capture | [matter-packet-capture.md](matter-packet-capture.md) | read-only |
| Home Assistant maintenance | [home-assistant-maintenance.md](home-assistant-maintenance.md) | change (gated) |
| matrix_e2ee integration update | [matrix-e2ee-update.md](matrix-e2ee-update.md) | change (gated) |
| Routine Ansible operations | [ansible-operations.md](ansible-operations.md) | change (allowlisted) |
| Linear issue → mergeable change | [issue-to-merge.md](issue-to-merge.md) | delivery |
| Failing health/playbook run | [fix-ci.md](fix-ci.md) | change |
| Release a reviewed change to production | [release.md](release.md) | change (gated) |
| Roll back a change | [rollback.md](rollback.md) | change (gated) |
| Controlled network configuration | [network-change.md](network-change.md) | change (gated) |
| Network outage / service recovery | [network-recovery.md](network-recovery.md) | recovery |
## Notes
- `fix-ci.md`, `release.md`, `rollback.md`, `network-change.md`, `network-recovery.md`
are adapted from the upstream guide to this repo's VPS-ops context (execution
layer is Ansible + SSH + Linear, not a software CI/CD pipeline).
- Health runbooks are read-only; they stop (`STOP`) when live state conflicts
with the expected state instead of mutating production.
+119
View File
@@ -0,0 +1,119 @@
# Runbook: <名称>
## Purpose
<说明本 Runbook 要解决的问题及成功结果,1–2 行。>
## Scope
- 适用环境:<production / staging / LAN …>
- 适用对象:<服务、主机、组件或告警类型>
- 不适用情形:<需要改用其他 runbook 或转人工的场景>
## Ownership
- Owner<团队或角色>
- Last reviewed<YYYY-MM-DD>
- Related systems<主机名 / 服务名>
## Preconditions
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
## Inputs
| 输入 | 来源 | 是否必需 | 校验方法 |
|---|---|---:|---|
| <参数> | <来源> | 是/否 | <如何确认有效> |
## Safety
### Non-negotiable rules
- 先只读诊断,后执行变更。
- 不得把删除现有配置作为首次恢复动作。
- 不得猜测或编造缺失参数。
- 不得绕过失败的测试、检查或审批。
- 每次变更后必须完成对应验证。
- 破坏性操作必须获得明确批准。
### Stop conditions
- 实际状态与本文档的前提或预期结果冲突。
- 缺少必要输入、权限、审批或回滚能力。
- 验证失败且本文档没有明确的下一步。
- 影响范围超出 Scope。
### Approval gates
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|---|---|---:|---|
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
## Procedure
### Step 1 — Diagnose
**Action**
<执行只读诊断动作。>
**Expected**
<列出预期输出、状态或证据。>
**Decision**
- 若 <条件 A>,进入 Step 2。
- 若 <条件 B>,进入 Troubleshooting A。
- 若无法判断或状态冲突,`STOP` 并记录证据。
### Step 2 — Change
**Action**
<描述单一、可审计的变更动作。>
**Expected**
<变更后应出现的状态。>
**Verification**
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
**Rollback**
- 触发条件:<什么情况需要回滚>
- 回滚动作:<如何撤销>
- 回滚验证:<如何确认恢复成功>
## Troubleshooting
### Troubleshooting A — <异常名称>
- 证据收集:<日志、指标、命令输出、链接>
- 允许动作:<仅限已验证且低风险的动作>
- 下一步:<回到某步 / 转入另一 runbook / STOP 并升级>
## Final Verification
只有同时满足以下标准,流程才算成功:
- <功能或服务状态>
- <自动化测试或健康检查>
- <监控指标或告警状态>
- <变更记录、PR 或 Issue 已更新>
## Failure Handling
若未能完成:
1. 停止进一步变更。
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
4. 按 <升级渠道> 交接,不继续猜测。
## References
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
+21
View File
@@ -1,5 +1,20 @@
# AdGuard Home health — dns.windy.lan
## Purpose
Read-only health check of the AdGuard Home LAN DNS service.
## Scope
- Applicable: [dns.windy.lan](../hosts/dns.windy.lan.md) (`192.168.66.36`).
- Read-only: does not expose query-log contents or secrets; does not change configuration.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: dns.windy.lan (`/opt/adguardhome`)
This runbook is read-only. It does not expose query-log contents or secrets.
Routine checks run through Ansible on demand:
@@ -58,3 +73,9 @@ a known-bad-signature test; an enabled DO bit alone is not validation.
Private PTR forwarding is intentionally absent because the EdgeRouter does
not currently answer private PTR requests.
## Safety
- Read-only: never change the DNS policy or the `agh-ui-access.service` nftables rule during this check.
- Do not infer a broken DNS policy from an empty `allowed_clients`.
- If live state conflicts with an expected value, `STOP` and report.
+42
View File
@@ -1,8 +1,29 @@
# Runbook: routine operations through Ansible
## Purpose
Routine operations (health, reconcile, maintenance) through the Ansible playbooks.
## Scope
- Applicable: every inventory host, run from `ansible/`.
- Not applicable: arbitrary remote commands — the reconcile playbook is allowlisted and gated.
Run commands from `ansible/`. The inventory forces IPv4 and uses the `windy`
account with sudo. Do a read-only health pass before any reconciliation.
## Safety
- Read-only health pass before any reconciliation.
- Mutating playbooks require explicit confirmation variables; do not bypass them.
- If a reconcile target or service name is not allowlisted, `STOP` — do not invent one.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: Ansible control-plane + all inventory hosts
## Health report (read-only)
```bash
@@ -48,6 +69,27 @@ Do not use this playbook for a Mailcow update, database migration, DNS record
change, or secret rotation. Those operations require their dedicated reviewed
and, where appropriate, interactive procedures.
## Deploy repo-owned Compose (static projects)
Repo source: `compose/<project>/compose.yml` (non-secret; secrets come from the
server-local `.env` via `${VAR}`). Mechanism and per-project status:
[`compose/README.md`](../compose/README.md).
```bash
# Read-only: staged-file diff + allowlist/confirmation asserts, no writes
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden --check --diff
ansible-playbook playbooks/compose-deploy.yml --limit powerdns --check --diff
# Apply: stage repo file → validate `docker compose config -q` against the
# server .env → backup current file (*.bak-<ts>) → promote → `up -d` (gated)
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden \
-e '{"compose_deploy_confirm": true}'
```
The playbook never writes, reads, or transfers the server `.env`. A failed
validation never touches the live compose file. Hosts without an allowlisted
`compose_repo_project` fail the assert — do not invent targets.
## Host-level maintenance
These playbooks cover every inventory host, including the Matrix K3s node:
+76
View File
@@ -0,0 +1,76 @@
# Runbook: fix a failing health/playbook run
> Adapted from the upstream guide's `fix-ci`. This repo has no software CI; the
> equivalent "pipeline" is the Ansible **health report** and the gated playbooks.
> This runbook covers diagnosing and fixing a failed or warning/critical run.
## Purpose
Diagnose and fix a failing Ansible health-report or playbook run without
skipping checks or changing unrelated code.
## Scope
- Applicable: `ansible-playbook playbooks/health-report.yml` and the gated playbooks under `ansible/playbooks/`.
- Not applicable: production changes beyond fixing the run; network/DNS changes → `network-change.md`.
## Safety
- Do not skip or weaken a failing check to make it pass.
- Do not change unrelated hosts or services.
- Prefer read-only diagnosis before mutation; destructive fixes require approval.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: Ansible health report / gated playbooks
## Procedure
### Step 1 — Reproduce and read
**Action** — re-run the failing playbook with `--limit <host>` and capture the task that failed.
```bash
cd ansible
ansible-playbook playbooks/health-report.yml --limit <host> -v
```
**Expected** — a specific failed task, host, and message (warning vs critical).
**Decision** — clear failure → Step 2; ambiguous → `STOP` and collect `-vvv` output + the relevant `latest.json`.
### Step 2 — Diagnose
**Action** — inspect the corresponding service on the host using the matching health runbook (`mailcow-health.md`, `vaultwarden-health.md`, `pdns-health.md`, etc.).
**Expected** — a root cause (container down, cert expired, queue backlog, drift).
**Decision** — root cause found → Step 3; live state conflicts with the runbook's assumptions → `STOP`.
### Step 3 — Fix within scope
**Action** — apply the minimal fix the service runbook prescribes (e.g. `compose-reconcile` for a config drift, or a documented update). Use only allowlisted/gated playbooks.
**Verification** — re-run the health report and confirm it passes.
**Rollback** — revert to the prior config/state and re-run; see `rollback.md` for the general procedure.
## Troubleshooting
### Troubleshooting A — Intermittent/flaky failure
- Evidence: timing, DNS stub flakiness (use `1.1.1.1`/`8.8.8.8` for probes).
- Allowed: re-run once with the documented resolver workaround.
- Next: still failing → `STOP` and escalate.
## Final Verification
- Health report passes for the affected host.
- No checks were skipped or weakened; the fix is committed/documented.
## References
- [`ansible-operations.md`](ansible-operations.md)
- Per-service health runbooks under [`runbooks/`](.)
+229 -19
View File
@@ -1,13 +1,47 @@
# Runbook: Home Assistant maintenance (hass.windy.lan)
Target: [hass.windy.lan](../hosts/hass.windy.lan.md) (HAOS, `machine: green`)
Upstream: HAOS 18.1 / Core 2026.8.x / Supervisor 2026.07.5 (verified 2026-08-13)
Target: [hass.windy.lan](../hosts/hass.windy.lan.md) (physical x88 Pro box, HAOS `machine: green`)
Upstream: HAOS 18.2 / Supervisor 2026.09.0 / Core 2026.9.1 (verified 2026-09-13)
This runbook covers routine Home Assistant maintenance through the **`ha`
supervisor CLI**. All commands are wrapped by a single script
[`scripts/ha-maintenance.sh`](scripts/ha-maintenance.sh); the sections below
document the exact commands it runs, for manual/agent use.
## Purpose
Run routine Home Assistant maintenance on `hass.windy.lan` (health snapshot,
config validation, log inspection, updates, and recovery) through the `ha`
supervisor CLI.
## Scope
Applies to `hass.windy.lan` only (HAOS, `machine: green`). Covers both
read-only checks and gated mutating operations; the "Command families
intentionally NOT scripted" table below lists what is deliberately out of
scope.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: hass.windy.lan (HAOS, `machine: green`)
## Safety
- Prefer read-only checks first; the health snapshot mutates nothing.
- Every mutating mode (update / restart / rebuild / rollback / reboot /
backup / restore / add-on lifecycle) refuses to run without `--yes`.
- `--restore` overwrites the current installation; `--rollback-os`,
`--reboot`, and `--rebuild-core` are disruptive. Run them only from a
planned recovery with the backup verified.
- Never commit `SUPERVISOR_TOKEN` or a long-lived `HA_TOKEN`; read entity
state via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`).
- The `--restart-core` wrapper exits 1 silently on ssh failure — treat an
empty/exit-1 result as failure and confirm with `ha core info`.
- If live state conflicts with a documented expectation, `STOP` and report;
do not improvise command families outside this script.
## Access pattern
`ha` authenticates to the Supervisor with `SUPERVISOR_TOKEN`. Interactive SSH
@@ -26,6 +60,17 @@ The script runs its whole procedure in one remote login (`sudo -n -i bash -s`)
so the banner appears once, then strips it with `awk` up to the
`System is ready! Use browser or app to configure.` line.
**Restart wrapper (verified 2026-08-14, W1N-107):**
`./ha-maintenance.sh --restart-core --yes` exited 1 with no output in <1s
and **did not restart Core**. The wrapper pipes a remote script through
`ssh … 2>/dev/null | awk …`; with `set -uo pipefail`, an ssh failure is
silent and the pipeline returns empty/exit 1 **before any remote command
runs**. That is not a MOTD-strip artifact after a successful restart.
The working restart was
`ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core restart'`
(~131s, `Command completed successfully.`). Treat empty/exit 1 as
failure; confirm with elapsed time and `ha core info`.
## Script usage
```bash
@@ -33,7 +78,7 @@ cd runbooks/scripts
./ha-maintenance.sh # read-only health snapshot
./ha-maintenance.sh --check-config # validate core configuration
./ha-maintenance.sh --logs core 200 # tail core logs
./ha-maintenance.sh --logs core 2500 # tail core logs (use 2500 after a restart)
./ha-maintenance.sh --logs supervisor # tail supervisor logs (default 100)
./ha-maintenance.sh --logs host 50 # tail host journald logs
./ha-maintenance.sh --logs apps:<slug> # tail an add-on log
@@ -55,7 +100,7 @@ cd runbooks/scripts
- `--restore` overwrites the current installation — run only from a planned
recovery, with the backup verified.
## Command reference (verified 2026-08-13)
## Command reference (verified 2026-08-14)
All verified against the live host. MOTD prepends each command's output; strip
with the `awk` pattern above or read the last block.
@@ -65,7 +110,7 @@ with the `awk` pattern above or read the last block.
| Purpose | Command |
|---|---|
| General overview | `ha info` |
| Core version/status | `ha core info` |
| Core version/status | `ha core info` (this CLI build has no `state:` field; success is a normal info dump) |
| Core config validation | `ha core check` |
| Core stats | `ha core stats` |
| Supervisor status | `ha supervisor info` (incl. add-on list) |
@@ -78,7 +123,7 @@ with the `awk` pattern above or read the last block.
| Reload stores/versions | `ha refresh-updates` |
| Job manager | `ha jobs info` |
| Resolution center | `ha resolution info` |
| Core logs | `ha core logs -n 100` (`-f` follow, `-b` boot id) |
| Core logs | `ha core logs -n 100` (`-f` follow, `-b` boot id). Default 100 misses setup; use `-n 2500` after a custom-component restart. `/config/home-assistant.log` may be missing — `ha core logs` is the source of truth. |
| Supervisor logs | `ha supervisor logs -n 100` |
| Host journald logs | `ha host logs -n 100` |
| Add-on logs | `ha apps logs <slug> -n 100` |
@@ -120,14 +165,21 @@ add-on states (any `state: error`?); `resolution info` issues; disk free.
Expect `Command completed successfully.` before a Core restart.
`ha core check` / a YAML reload is **not** enough after copying Python
custom-component files — restart Core.
### 3. Inspect logs
```bash
./ha-maintenance.sh --logs core 200
./ha-maintenance.sh --logs core 2500 # after a Core restart / custom-component copy
./ha-maintenance.sh --logs supervisor
./ha-maintenance.sh --logs apps:core_mosquitto
```
Default `--logs core` (100 lines) is too short to catch coordinator pickle /
setup errors. `/config/home-assistant.log` may be absent while
`ha core logs` still has history.
### 4. Apply updates (mutating)
```bash
@@ -155,6 +207,140 @@ booted. After a bad OS update, `ha os boot-slot other` boots the previous slot.
./ha-maintenance.sh --backup pre-migration --yes # named backup
```
### 7. Install or update a custom component (manual zip)
Home Assistant loads custom integrations from
`/config/custom_components/<domain>/` (on this HAOS host `/config`
`/homeassistant`). Official lookup:
`<config>/custom_components/<domain>` then built-in
`homeassistant/components/<domain>`
([Integration file structure](https://developers.home-assistant.io/docs/creating_integration_file_structure)).
A folder named after the domain, with at least `manifest.json` and
`__init__.py`, is enough. **Restart Core** after copying — `ha core check`
and a YAML reload do not pick up new Python packages.
This host's live trees are **file copies**, not git clones. Do not
`git pull` inside `custom_components/`.
#### Official plugin paths (CSG)
[windyboy/china_southern_power_grid_stat README](https://github.com/windyboy/china_southern_power_grid_stat):
[HACS](https://hacs.xyz/) **or**
[手动下载安装](https://github.com/windyboy/china_southern_power_grid_stat/releases).
This host uses the zip path. **Do not HACS-update this integration here.**
HACS still tracks upstream `CubicPill/china_southern_power_grid_stat`
`v1.2.0` and would overwrite the fork. Releases have no uploaded zip
assets — use GitHub **Source code (zip)** / zipball of the tag.
Worked SSH example (tag, backup, `rsync`, `__pycache__`, restart):
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § Manual
custom-component install.
#### Procedure
1. **Backup the live tree off `custom_components/`.** HA scans every
directory under `custom_components/` whose `manifest.json` `domain`
matches. A `*.bak-*` folder next to the live tree makes Core import
the backup (`No module named '...bak-YYYYMMDD-...'`, W1N-106). CSG
backups: `/homeassistant/.csg-backups/`.
2. **Copy only the inner `custom_components/<domain>/` tree**, not the
repo root and not an extra nested folder.
3. **Wipe `__pycache__` as root.** `rsync --delete` as `hassio` cannot
unlink Core-owned `.pyc` (permission denied, exit 23); stale
`cpython-314` bytecode can keep the old coordinator in memory until
restart. Then restart:
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i rm -rf /homeassistant/custom_components/<domain>/__pycache__ \
/homeassistant/custom_components/<domain>/*/__pycache__ &&
sudo -n -i ha core restart'
```
4. **Wait 12 min**, then `ha core info` (this CLI build has no `state:`
field; success is a normal info dump). Confirm `manifest.json`
`version` matches the tag.
5. **Read enough Core logs.** Default `ha core logs` is too short to
catch setup. Use `-n 2500` (or `--logs core 2500`) and look for
`Setting up <domain>` plus the first coordinator errors.
6. **First poll can time out.** If last-month sensors have numbers but
this-month stay `unknown`/`unavailable`, reload the config entry
(UI: integration → Reload). Supervisor:
```bash
# entry id from .storage/core.config_entries (CSG: 01KGCQDSZCF523A9X6SV3BZ1B9)
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i python3 -c "
import os, urllib.request
req = urllib.request.Request(
\"http://supervisor/core/api/config/config_entries/entry/<ENTRY_ID>/reload\",
method=\"POST\",
headers={\"Authorization\": \"Bearer \" + os.environ[\"SUPERVISOR_TOKEN\"]},
)
print(urllib.request.urlopen(req, timeout=60).status)
"'
```
7. **Do not edit the dashboard or `templates/csg_sensors.yaml` for an
install.** Entity IDs did not change across v1.3.0/v1.3.1/v1.3.2.
(The old `| float(0)` fake-zero follow-up was resolved 2026-08-29 by
W1N-239: template sensors now carry `availability` templates and show
`unavailable` instead of fake zeros when native CSG sensors are down.
Template edits go through that issue, not the install path.)
#### Verify (CSG, after v1.3.2 / W1N-118)
| Check | Expect |
|---|---|
| `manifest.json` `version` | `1.3.2` |
| `ha core logs` after this restart | `Setting up china_southern_power_grid_stat`; **no** `cannot pickle 'mappingproxy'` |
| Config entry | `state: loaded` |
| `sensor.0800041935246530_balance` | numeric (may be `0.0`) |
| `sensor.0800041935246530_this_month_total_usage` | numeric after reload if first poll timed out |
| Native `*_total_cost` / `current_ladder` | may stay `unknown` (CSG marketing calendar SQL error); dashboard uses W1N-114 `csg_*` ladder/cost templates |
`monetary` + `total_increasing` warnings on this-month/year cost sensors
are a remaining plugin issue, not an install failure.
There is no long-lived `HA_TOKEN` in the agent environment. Read entity
states via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`) at
`http://supervisor/core/api/states/<entity_id>`.
#### CSG display refactor 2026-09-04 (VPS-90)
Template/dashboard changes made **after** pricing cross-check (8月账单
198.64 元 vs 模板 198.65 元,≤0.01 元;阶梯常量 0.589/0.639/0.889、
260/600 夏档未动):
- `templates/csg_sensors.yaml` Block B 新增
`sensor.csg_this_month_avg_price`(本月阶梯电费÷本月用电,`元/kWh`);
**csg_* template sensors = 15**
- Panel `power-monitor``lovelace.dashboard_unknown`):环比行改名
「环比上月同期」;glance「本月/上月」去重为单卡「上月」(本月行归
💰核心数据卡);⚡阶梯电价卡加「本月实际均价」行。实体引用 20→21。
- `automations.yaml` +2 提醒:`automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`
10-25/ `automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`4-25
09:00 Matrix 提醒人工切「本月累计」gauge 季节档(max/segments 不可模板化)。
- 金额单位混排(原生 CNY vs 模板 元)**保留**`config/entity_registry/update`
拒绝自定义文本单位(`extra keys not allowed … Got '元'`),已定案接受。
**WS 改面板(2026.8,本机实测,后续沿用)**: core/主机 python 无 ws 库、
core 容器内经 supervisor 代理 WS 被拒(loop prevention)。用
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i sh -c "docker run --rm -i --network host -e SUPERVISOR_TOKEN \
--entrypoint python3 r.hassbus.com/home-assistant/aarch64-hassio-supervisor:2026.08.0 \
- < /tmp/x.py"'
```
`ws://172.30.32.2/core/websocket`aiohttpheader `Authorization: Bearer
$SUPERVISOR_TOKEN`,随后 auth 帧同 token)。命令名 **`lovelace/config`**(读)
+ **`lovelace/config/save`**(写,url_path + 全量 config);`lovelace/config/get`
已不存在(unknown_command)。备份与细节见
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § CSG 面板重构 2026-09-04。
## Command families intentionally NOT scripted
These exist in `ha` but are either rare, dangerous, or better done in the web
@@ -185,21 +371,42 @@ host for exact syntax.
## Docs vs actual CLI discrepancies
The [official HAOS common-tasks docs](https://www.home-assistant.io/common-tasks/os/)
reference `ha backups list` and `ha host update`. **Neither exists in the
installed CLI** (2026-08-13): backups are inspected with `ha backups info
<slug>` (a slug is required) and `host` has no `update` subcommand. Trust the
server CLI (`ha <cmd> --help`) over the docs.
also mention `ha host update`, which **does not exist** in this CLI
(2026-08-14). Docs' `ha backups list` is not a subcommand either:
`ha backups --help` lists freeze/info/new/options/reload/remove/restore/thaw;
extra positional args (`list`, `nonsense`, ...) are ignored and the default
list still prints with exit 0. The list command is plain `ha backups`.
Per-backup: `ha backups info <slug>` (slug required). Trust the server CLI
(`ha <cmd> --help`) over the docs.
## Known issues on hass.windy.lan (observed 2026-08-13)
This CLI's `ha core info` also has no `state:` field (verified 2026-08-14).
Wait for a successful info dump after restart, not a `state: running` line.
From the health snapshot — follow-ups are optional, no action taken:
## Known issues on hass.windy.lan (2026-08-14)
- **2 add-ons in `state: error`**: `core_openthread_border_router`,
`a0d7b954_ssh` (duplicate Advanced SSH & Web Terminal install).
- **`ha resolution info` issues**: `systemd_unit_failed`
(`systemd-vconsole-setup.service`), `no_current_backup`, 2×
`corrupt_repository` (store `d5369777`, `a0d7b954`).
- `host info` reports `disk_life_time: 10` (disk lifetime warning threshold).
2026-08-13 snapshot items were resolved same day (W1N-70/71/72/73/74/75/76):
OTBR and the duplicate SSH add-on uninstalled, resolution-center empty,
full backup `pre-maintenance-20260813` (slug `411a4ba5`). Remaining:
- **Bluetooth hci0 instability (RTL8821CS)**: `bluetooth_auto_recovery`
power-reset times out every ~2 min; kernel `hci0 hardware error`. No BLE
entities exist, so no user impact. HAOS image ships `x88-bt-hci-recovery`
workaround units.
- `host info` reports `disk_life_time: 10` (boot eMMC ~10% life left) —
monitor on each snapshot; plan disk replacement / data-disk migration.
- **Home PPPoE IPv4 to CSG is blackholed** (`curl -4` to
`218.19.148.218:443` times out). `end0` IPv6 works (`curl -6
https://95598.csg.cn` → HTTP 200). Entry `ip_family: ipv4` still
matches the stored option; first post-restart poll can still time out
— reload the config entry rather than reinstalling.
- **WSL HTTP proxy**: LAN `hass.windy.lan:8123` through Mihomo returns
empty `502`. Bypass proxy or add `.windy.lan` to `NO_PROXY` before
debugging UI/API from the workstation
([hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § HTTP proxy
gotcha).
- **No long-lived HA token in the agent environment.** Read entity
states via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`) at
`http://supervisor/core/api/states/...`, not a committed `HA_TOKEN`.
## Pass criteria
@@ -207,4 +414,7 @@ From the health snapshot — follow-ups are optional, no action taken:
- Mutating modes refuse to run without `--yes` (incl. `--restore`, `--app`)
- `--check-config` returns success
- Update / rollback / restore / reboot confirmed only after explicit `--yes`
- Custom-component zip install: live `manifest.json` version matches the
tag; backups not under `custom_components/`; Core restarted; logs show
`Setting up <domain>` without import / pickle errors
- Update the **Verified** line on [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md)
+276
View File
@@ -0,0 +1,276 @@
# Runbook: Host disk cleanup (unbounded container logs / apt cache / docker artifacts)
## Purpose
Reclaim space on a root filesystem that is filling up (≥70% used) on a Docker
Compose host, by fixing unbounded container log growth at the source, clearing
apt/journal caches, and removing unused Docker images/volumes. Success: root
usage drops to a safe band (≤55% used, or per acceptance in the tracking issue)
and log growth stays bounded afterwards.
## Scope
- 适用环境: production single-root-fs hosts running Docker Compose stacks
(first application: `hk2.chans.xyz`; reusable for `mx2.windy.me` / `us2.wsvc.info`
which run the same unbounded-`json.log` pattern).
- 适用对象: root filesystem usage; container stdout/stderr log files
(`/var/lib/docker/containers/*/*-json.log`); `/var/cache/apt`; systemd journal;
unused Docker images / anonymous volumes / build cache.
- 不适用情形: hosts without systemd-journald or without Docker; LAN/HAOS hosts
(use their own runbooks); cases needing disk *growth* (provider resize) rather
than cleanup; anything touching service data volumes or `/opt/*` configs
(STOP and use the service-specific runbook instead).
## Ownership
- Owner: windy (operator) + agent executing per approval
- Last reviewed: 2026-09-02
- Related systems: hk2.chans.xyz (PowerDNS auth / AdGuard Home / Traefik / RustDesk compose stacks)
## Preconditions
- SSH access to the target host with **passwordless sudo** (`sudo -n true` must succeed).
- A recorded `df -h` baseline and `docker system df` baseline.
- **Explicit user approval** for every service touch listed in Approval gates
(recorded in the tracking issue, e.g. Plane `vps` VPS-81).
- No open incident on the target host.
- Container log growth root cause identified in Diagnose before mutating.
## Inputs
| Input | Source | Required | Validation |
|---|---:|---|
| Target host | inventory/hosts.md | yes | SSH login + `uname -r` |
| df/docker baseline | live read-only probe | yes | recorded before first mutation |
| Approved service touches | user confirmation in tracking issue | yes | issue comment states approval |
| Image keep-list (rollback pins) | operator decision in issue | yes | review `docker image ls` before rmi |
| Backup of any config edited | local copy with timestamp | yes | exists before edit |
## Safety
### Non-negotiable rules
- Prefer read-only diagnosis before mutation (never mutate on an unmeasured disk).
- Never use `rm` on a live container log — use `truncate -s 0` (keeps the fd valid).
- Never run `docker image prune -a` when a keep-list is intended — no keep-list
exists; delete explicitly with `docker rmi`.
- Never run `docker volume prune -a` — plain `docker volume prune` (no `-a`)
removes only unused anonymous volumes; named/in-use volumes stay.
- After every mutation, verify the expected state (`df -h`, container status).
- Destructive actions require explicit approval (Approval gates).
### Stop conditions
- Live state conflicts with this runbook's preconditions or expectations (e.g.
root usage differs wildly from baseline, or a container is unhealthy).
- Missing approval, missing backup, or missing rollback ability.
- A verification step fails with no documented next step.
- Any step would touch a volume/container/mount that is not on the approved list.
### Approval gates
| Action | Risk | Explicit approval | Approval record |
|---|---:|---|---|
| `docker restart <chatty container>` | low (sec-level blip of that service only) | yes | tracking issue (VPS-81 T1) |
| `systemctl restart systemd-journald` | low (sec-level, no state loss) | yes | tracking issue (VPS-81 T2) |
| `apt-get clean` | low (re-downloadable) | no | — |
| `journalctl --vacuum-*` / journald drop-in | low | no (restart above is gated) | — |
| `docker rmi` of unused images | medium (rollback pin removed unless kept) | yes (keep-list) | tracking issue (VPS-81 T3) |
| `docker volume prune` | medium (data in anonymous volumes lost) | yes | tracking issue (VPS-81 T4) |
| `docker builder prune` | low | no | — |
## Procedure
### Step 1 — Diagnose
**Action**
Read-only: `df -h`, `df -i`, `sudo du -x -h --max-depth=1 /`, `docker system df`,
and locate oversized container logs:
`sudo ls -la /var/lib/docker/containers/*/*-json.log`. Map a big log to its
container (`docker inspect -f '{{.Name}} {{.LogPath}}' <id>`), then inspect what
it logs (`sudo tail -c 400000 <logpath>`; count `[debug]` lines) and find the
config flag driving it (e.g. AGH `log.verbose` in its YAML; note `log.file: ""`
means the app's own rotation keys are inert and output goes to the container log).
**Expected**
A full accounting of root usage and identification of: (a) any unbounded
container log and its root-cause flag; (b) reclaimable apt cache; (c) journal
size and journald limits; (d) unused images (0 dangling expected) and unused
anonymous volumes.
**Decision**
- If root is ≥70% used or any container log is unbounded → Step 2.
- If root is healthy and logs are bounded → STOP (no change needed; record evidence).
- If state conflicts with expectations (e.g. missing sudo, unexpected mount) → STOP.
### Step 2 — Fix noisy container logging at the source, then truncate
**Action**
1. Back up the app config: `sudo cp <config> <config>.bak-YYYYMMDD-<issue>`.
2. Disable the debug/verbose flag (e.g. `log.verbose: true → false` in the AGH YAML).
3. Apply config with a container restart: `docker restart <container>` (config-level
change; **no recreate** needed and daemon.json rotation would not apply anyway).
4. Truncate the accumulated logs: `sudo truncate -s 0 <json.log>` for the chatty
container(s) (and any other oversized ones, e.g. traefik).
5. Record `df -h` before/after.
**Expected**
`docker logs <container>` no longer shows the `[debug]` flood; the `*-json.log`
stops growing; several GB reclaimed.
**Verification**
- `sudo tail -c 200000 <json.log>` after ≥1 minute → no new debug lines.
- `df -h` improvement recorded.
- Container still `Up (healthy)`.
**Rollback**
- Trigger: log volume unchanged, service degraded, or debug output is actually needed.
- Action: restore the config backup and `docker restart <container>`.
- Verify: original verbose behaviour back; container healthy.
### Step 3 — Clear apt cache and cap journald
**Action**
1. `sudo apt-get clean` (clears only `/var/cache/apt/archives`; `/var/lib/apt/lists`
is not cleared by it and regenerates on `apt update` — optional/low value, skip).
2. `sudo journalctl --vacuum-size=100M`.
3. Write drop-in `/etc/systemd/journald.conf.d/00-disk-<issue>.conf`:
`[Journal]` + `SystemMaxUse=200M`.
4. `sudo systemctl restart systemd-journald` (approved service touch).
5. Record `df -h` before/after.
**Expected**
Archives cleared (~1.4G on hk2), journal ≤100M, future journal capped at 200M.
**Verification**
- `du -sh /var/cache/apt/archives` → ~0.
- `journalctl --disk-usage` → ≤100M.
- `systemctl show systemd-journald -p ...` or restart log confirms new limit;
`journalctl -b` still readable.
**Rollback**
- Trigger: journald fails to start or logs lost unexpectedly.
- Action: remove the drop-in, `sudo systemctl restart systemd-journald`.
- Verify: journald active, prior journal entries still listed.
### Step 4 — Remove unused Docker images (explicit keep-list)
**Action**
1. Enumerate unused images: `docker image ls` cross-checked against the images of
running containers (`docker ps --format '{{.Image}}'`). Re-enumerate at
execution time — the list drifts.
2. Present the exact removal list to the operator; keep the agreed rollback pin(s)
(e.g. `powerdns/pdns-auth-50:5.0.5`) and delete the rest explicitly:
`docker rmi <repo:tag> ...` (per image).
3. Record `df -h` before/after.
**Expected**
Only in-use images + kept pins remain; ~12.5G reclaimed (reclaim is an upper
bound — layers shared with kept images are not freed; measure with `df`, do not
promise the estimate).
**Verification**
- `docker image ls` shows only the expected set.
- `docker system df` images reclaimable ≈ 0 for the removed set.
- All containers still `Up`.
**Rollback**
- Trigger: an image that was actually needed was removed.
- Action: re-pull it from the registry (`docker pull <repo:tag>`); if a kept pin
must change, update the compose pin and `up -d`.
- Verify: image present; affected service healthy.
### Step 5 — Remove unused anonymous volumes and build cache
**Action**
1. Enumerate volumes: `docker volume ls`, and confirm which are referenced by
containers (`docker inspect` Mounts). Expected targets: anonymous volumes with
no container reference.
2. `docker volume prune` (**no `-a`**) — engine ≥ v23 removes only unused
anonymous volumes; in-use volumes (e.g. PG data) are protected by container
references in every version.
3. `docker builder prune -f`.
4. Record `df -h` before/after.
**Expected**
Unused anonymous volumes (~1.2G on hk2) and build cache gone; in-use volumes intact.
**Verification**
- `docker volume ls` shows only in-use volumes.
- Services that own volumes (e.g. postgres) report healthy and data present.
- `df -h` improvement recorded.
**Rollback**
- Trigger: data loss suspected in a removed volume.
- Action: restore from backup if the volume ever contained data; verify against
the pre-prune enumeration (targets must be anonymous + unreferenced before prune).
- Note: this is why target enumeration is recorded before pruning.
## Troubleshooting
### Troubleshooting A — Log still grows after disabling verbose
- Evidence: `sudo tail -c 200000 <json.log>` still shows new lines; app config re-checked.
- Allowed actions: check for a second verbose source (container entrypoint flags,
other apps in the same log); check `docker inspect <c> --format '{{.HostConfig.LogConfig}}'`.
- Next step: back to Step 2 or STOP if a container-level log-opts change (recreate)
would be needed — that is a separate approval.
### Troubleshooting B — `docker rmi` fails (image in use)
- Evidence: `image is being used by stopped container ...`.
- Allowed actions: identify the stopped container (`docker ps -a`); confirm it is
not needed; remove it only with explicit approval.
- Next step: re-run rmi for the remaining images; never force-delete blindly.
### Troubleshooting C — `docker volume prune` would remove more than expected
- Evidence: prune dry-run/listing includes a named or referenced volume.
- Allowed actions: abort; do not add `-a`; re-check references.
- Next step: STOP and report to the operator with the enumeration.
## Final Verification
The flow is successful only when all of the following hold:
- `df -h` root usage is in the agreed band (VPS-81: 76% → ≤55% used; measure, do not assume).
- `docker system df` shows reclaimable ≈ 0 for images/volumes targeted.
- All containers `Up` (health checks pass); public services verified
(`dig @<host-ip> SOA <zone>` for DNS hosts; service URLs reachable).
- Tracking issue updated with before/after `df`, actions, and the one-week
observation checkpoint for log growth.
## Failure Handling
If the flow cannot complete:
1. Stop further mutation.
2. Collect command output, timestamps, and the exact step that failed.
3. Record completed steps, actual results, unmet expectations, and whether a
rollback ran.
4. Hand over per the tracking issue with evidence; do not guess further.
## References
- Plane `vps` issue VPS-81 "hk2: 释放根盘空间" (+ subtasks VPS-82…88) — plan, review findings, approvals.
- [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) — host facts.
- [RUNBOOKS.md](../RUNBOOKS.md) — runbook spec; [runbooks/README.md](README.md) — index.
+93
View File
@@ -0,0 +1,93 @@
# Runbook: issue → mergeable change
## Purpose
Turn an approved Linear `vps` issue into a reviewed, mergeable change in this
repo (docs, runbooks, hosts facts, or Ansible playbooks).
## Scope
- Applicable: repo content under `docs/`, `runbooks/`, `hosts/`, `inventory/`, `ansible/`.
- Not applicable: mutating production state directly — that goes through `release.md` / `ansible-operations.md`.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: Linear MCP (`vps` project), git
## Inputs
| Input | Source | Required | Validation |
|---|---|---:|---|
| Issue identifier | Linear (`vps` project) | Yes | `linear_get_issue <id>` returns a description |
| Current repo state | `git status` / `git log` | Yes | Clean or intended worktree |
## Safety
- Scope is locked to the issue: do not bundle unrelated changes.
- Never commit secrets (see `AGENTS.md` §Safety).
- Verify every change; do not merge a change whose verification was skipped.
## Procedure
### Step 1 — Read the issue
**Action** — `linear_get_issue <id>`, read description and acceptance criteria.
**Expected** — clear scope, action, and verification for the change.
**Decision** — if the issue is ambiguous or lacks verification criteria, `STOP`
and ask for clarification (add a `needs-info` label if applicable). Otherwise go to Step 2.
### Step 2 — Inspect and change
**Action** — read the relevant files, then make the minimal change the issue asks for.
**Expected** — diff is scoped to the issue.
**Decision** — if the change needs production mutation, `STOP` and route to
`release.md`. Otherwise go to Step 3.
### Step 3 — Verify
**Action** — run `scripts/validate-repo.sh` from the repo root (covers secret
scan, inventory cross-check, markdown link check, runbook-spec check, and
Ansible `--syntax-check`); for changes that alter playbook behavior, also run
a read-only `ansible-playbook --check` where possible.
**Verification** — `scripts/validate-repo.sh` exits 0; the concrete checks
must match the change type.
**Decision** — verification passed → Step 4; failed → Troubleshooting A.
### Step 4 — Commit and link
**Action** — commit with a message containing the full issue ID (e.g. `W1N-123: …`); open a PR if the change is substantial; link the issue via `linear_save_comment`.
**Verification** — `git log -1` shows the issue ID; the issue has the commit/PR pointer.
**Rollback** — `git revert <sha>` or `git checkout <branch>` to drop the change; re-verify after.
## Troubleshooting
### Troubleshooting A — Verification failed
- Evidence: command output, failing check.
- Allowed: fix the change within scope; re-run verification.
- Next: still failing → `STOP` and report in the issue.
## Final Verification
- Change matches the issue scope.
- Verification passed and the issue is updated with evidence.
## Failure Handling
If unfinished: stop, collect the failed check output, record completed steps, and
hand back to the issue — do not guess.
## References
- [`docs/agents/issue-tracker.md`](../docs/agents/issue-tracker.md)
- [`RUNBOOKS.md`](../RUNBOOKS.md)
+20
View File
@@ -1,5 +1,20 @@
# Runbook: mailcow health (mx2)
## Purpose
Read-only health check of the mailcow stack on mx2.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
- Read-only: does not change mailcow configuration or service state.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: mx2.windy.me (`/opt/mail`)
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Path: `/opt/mail`
Prefer: the Ansible health report (`ansible/playbooks/health-report.yml --limit mailcow`),
@@ -71,6 +86,11 @@ The sanitized Ansible health profile is `mailcow` (`ansible/playbooks/healthchec
The server-local timer emits a sanitized result at `/var/lib/vps-health/latest.json`.
It does not change Mailcow configuration or service state.
## Safety
- Read-only: never mutate configuration or service state during this check.
- If live state conflicts with an expected value below, `STOP` and report; do not "fix" on the fly.
## Pass criteria
- Compose stack up; watchdog ~100%
+21
View File
@@ -1,5 +1,20 @@
# Runbook: use mailcow SMTP / IMAP (client)
## Purpose
Reference for configuring mail clients against the mailcow SMTP/IMAP endpoints.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md) client submission (587/465) and IMAP/POP (993/995).
- Not applicable: server-side mailcow configuration or administration.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: mx2.windy.me (SMTP/IMAP client endpoints)
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Prerequisite: a mailbox on `windy.me` (password from mailcow UI, not the admin account unless it is that mailbox).
@@ -52,3 +67,9 @@ Do not commit or paste real passwords into this repo.
- Port/TLS mode mismatch (587 vs 465)
- Account active in mailcow; not rate-limited / fail2banned after bad attempts
- Apps that store SMTP in their own config (e.g. Vaultwarden `config.json`) may keep a **stale** password even when `.env` is correct — verify AUTH against the effective config ([vaultwarden-health](vaultwarden-health.md) §5)
## Safety
- Do not commit or paste real passwords into this repo or chat.
- Use submission (587/465) for client sending; never use port 25 as a desktop/app outbound port.
- If live state conflicts with the endpoint values above, `STOP` and report; do not change server-side settings during this reference check.
+28
View File
@@ -1,9 +1,37 @@
# Runbook: mailcow update (mx2)
## Purpose
Update the mailcow stack on mx2 to the latest supported release.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
- Not applicable: config changes beyond the update, DB migration, secret rotation.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: mx2.windy.me (`/opt/mail`)
## Approval gates
| Action | Risk | Explicit approval |
|---|---|---|
| Run `./update.sh` (recreates containers, brief mail interruption) | Medium | Yes — user confirmation required |
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Path: `/opt/mail`
**Confirm with the user before running an update.**
## Safety
- Never run the update without explicit user confirmation.
- Never pass secrets into the chat log; do not commit `mailcow.conf`.
- If a step fails, capture `docker compose ps` and logs and stop before further changes.
- If live state conflicts with this runbook's assumptions (e.g. unexpected `mailcow.conf` values), `STOP` and report.
## Before
1. Run [mailcow-health](mailcow-health.md) (Ansible health report). Record baseline.
+171
View File
@@ -0,0 +1,171 @@
# matrix_e2ee update (hass.windy.lan)
## Purpose
Update the custom **`matrix_e2ee`** integration on `hass.windy.lan` while
preserving a verified rollback point and confirming that Home Assistant loads it.
## Scope
- Applicable: deploying a reviewed `matrix_e2ee` source revision to
`hass.windy.lan`.
- Not applicable: Home Assistant Core upgrades, integration configuration
changes, or recovery without a usable live-tree backup.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-23
- Related systems: hass.windy.lan (HAOS, `machine: green`)
## Approval gates
| Action | Risk | Explicit approval |
|---|---|---|
| Replace the live integration tree and restart Home Assistant Core | Medium | Yes — user confirmation required |
## Safety
- Do not replace the live tree or restart Core without explicit user confirmation.
- If any precondition or verification fails, `STOP` and record evidence before continuing.
## Preconditions
- The source repo at `/home/windy/project/ha-matrix-e2ee` is on the **target
state**: either a release tag (`git tag -l 'v*'`) or a commit whose
`manifest.json` `version` is the target. Note v0.3.0 was deployed from an
**untagged** `main` HEAD (`216cc99`), so the tag check alone is not enough —
confirm the working-tree `custom_components/matrix_e2ee/manifest.json`.
- The working tree matches HEAD: `git status --short` clean (only ignorables)
and `git diff HEAD -- custom_components/` empty. Record
`git rev-parse HEAD` for the docs/Linear record — HEAD can move during a
session, so re-check right before rsync (verified 2026-08-18: HEAD moved
from a `w1n-180` branch merge to `main` mid-deploy).
- The remote host is reachable and `sudo -n -i ha core info` succeeds.
- The workstation HTTP proxy does not interfere — LAN hosts must be reachable
without proxying (unset `http_proxy` / `HTTP_PROXY` if needed).
- If the agent sandbox hits `Bad owner or permissions on /etc/ssh/ssh_config.d/20-systemd-ssh-proxy.conf`, add `-F /dev/null` to the `ssh` / `rsync` commands below.
- Domain is **`matrix_e2ee`** (double-e). Older notes may say `matrix_e2e`;
paths, events, and services all use `matrix_e2ee`.
## Procedure
### 1. Backup the live tree
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i mkdir -p /homeassistant/.matrix-e2ee-backups &&
sudo -n -i cp -a /homeassistant/custom_components/matrix_e2ee \
/homeassistant/.matrix-e2ee-backups/matrix_e2ee.bak-$(date +%Y%m%d)-v<OLD_VERSION>'
```
The backup lives in `/homeassistant/.matrix-e2ee-backups/` — a directory
separated from `custom_components/` to avoid HA scanning it as a custom
component domain.
### 2. Rsync the new source
```bash
rsync -a --delete -e 'ssh -o BatchMode=yes' \
/home/windy/project/ha-matrix-e2ee/custom_components/matrix_e2ee/ \
hassio@hass.windy.lan:/homeassistant/custom_components/matrix_e2ee/
```
The `--delete` cannot remove Core-owned `__pycache__` — that is handled
in the next step. Source `.py` files and `manifest.json` are transferred
correctly even with the `__pycache__` errors, but **rsync exits with code 23
(`some files/attrs were not transferred`)** — that is expected, not a failure.
Confirm the transfer by checking the manifest on the host before restarting.
### 3. Wipe `__pycache__` (as root) and restart Core
Quote the nested `__pycache__` glob — remote login shell is zsh and will
fail with `no matches found` if left unquoted.
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
"sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee/__pycache__ \
'/homeassistant/custom_components/matrix_e2ee/*/__pycache__' &&
sudo -n -i ha core restart"
```
Stale `cpython-314` bytecode in Core-owned `__pycache__` keeps the old
coordinator in memory until restart. Wipe before restart.
Wait for `Command completed successfully.` (typically 12 min).
### 4. Verify the deployment
#### 4a. Confirm manifest version
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i cat /homeassistant/custom_components/matrix_e2ee/manifest.json'
```
Expect `"version": "<NEW_VERSION>"`.
#### 4b. Check Core logs for matrix_e2ee
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i ha core logs -n 2500' | grep -E 'matrix_e2ee|Setting up matrix' | head -20
```
Expect:
- `Setup of domain matrix_e2ee took ...` (older wording `Setting up matrix_e2ee` may appear)
- `matrix_e2ee restored existing device; user=@hass:chans.xyz device=rO1R915ncu`
- No `ERROR` level messages from `custom_components.matrix_e2ee`
- Blocking-call WARNINGs from `_patch_nio_sas_timeout` / nio store I/O are expected
#### 4c. Verify the entry is loaded (optional, via Supervisor API)
No trailing slash on the entries URL (trailing `/` returns 404 on Core 2026.8.1).
```bash
ssh -o BatchMode=yes hassio@hass.windy.lan \
"sudo -n -i python3 - <<'PY'
import os, json, urllib.request
req = urllib.request.Request(
'http://supervisor/core/api/config/config_entries/entry',
headers={'Authorization': 'Bearer ' + os.environ['SUPERVISOR_TOKEN']},
)
entries = json.loads(urllib.request.urlopen(req, timeout=30).read())
for e in entries:
if e['domain'] == 'matrix_e2ee':
print(f\"{e['domain']}: state={e['state']} source={e['source']}\")
PY"
```
Expect `state: loaded`.
### 5. Record the deployment
- Update the `matrix_e2ee` live-tree section in
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md): new version, source
commit (`git rev-parse HEAD`), backup name, and any new feature notes.
- Record the operation in the Linear `vps` project (scope, action,
verification, follow-up); see [docs/agents/issue-tracker.md](../docs/agents/issue-tracker.md).
## Rollback
If Core fails to start after the update:
```bash
# Restore the backup
ssh -o BatchMode=yes hassio@hass.windy.lan \
'sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee &&
sudo -n -i cp -a /homeassistant/.matrix-e2ee-backups/matrix_e2ee.bak-<DATE>-v<OLD_VERSION> \
/homeassistant/custom_components/matrix_e2ee &&
sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee/__pycache__ &&
sudo -n -i ha core restart'
```
If a full HA backup exists (pre-update), restore via `ha backups restore <slug>`.
## References
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — current live version and config
- [docs/home-assistant-matrix.md](../docs/home-assistant-matrix.md) — integration architecture and verification model
- [home-assistant-maintenance.md](home-assistant-maintenance.md) — general HA maintenance procedures
- [ha-matrix-e2ee source](https://github.com/windyboy/ha-matrix-e2ee) — GitHub repo

Some files were not shown because too many files have changed in this diff Show More