Compare commits
79
Commits
d43fc5aa01
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ff1a92110c | ||
|
|
cc3fb99c14 | ||
|
|
6879d79cc6 | ||
|
|
c445c5f512 | ||
|
|
3c83246f24 | ||
|
|
c0c975584a | ||
|
|
0034cec925 | ||
|
|
c9dcde1274 | ||
|
|
037c4ccaa5 | ||
|
|
17bb171578 | ||
|
|
cecf7e6331 | ||
|
|
e957bc2bb1 | ||
|
|
de52cb8b57 | ||
|
|
b15e19bce9 | ||
|
|
bee54a6858 | ||
|
|
d6747028b4 | ||
|
|
f174aa1219 | ||
|
|
5b5f6042e6 | ||
|
|
50136b2ffd | ||
|
|
6707cebc88 | ||
|
|
908ff5412a | ||
|
|
e58283210a | ||
|
|
8c73d1f894 | ||
|
|
bc0a86245d | ||
|
|
13032fd0bb | ||
|
|
d2063e7496 | ||
|
|
2d95f87897 | ||
|
|
b61513c93e | ||
|
|
ab808088b8 | ||
|
|
c7dc4fd25c | ||
|
|
7bc7d3f99b | ||
|
|
fea9a6560f | ||
|
|
a95b626636 | ||
|
|
27fe9c078e | ||
|
|
aaa4ee312e | ||
|
|
6b298491a8 | ||
|
|
32631e2996 | ||
|
|
1426b4ecfe | ||
|
|
1f5e58bf17 | ||
|
|
e5819eeba3 | ||
|
|
079332e082 | ||
|
|
7cedba7f51 | ||
|
|
343c5db415 | ||
|
|
885d977531 | ||
|
|
b0c01b2551 | ||
|
|
047ac03346 | ||
|
|
1ec9246156 | ||
|
|
1936b8f5fe | ||
|
|
eda6536ddb | ||
|
|
1dc880362d | ||
|
|
70aea6cd72 | ||
|
|
8303d78caf | ||
|
|
ebfe7b8488 | ||
|
|
88eaefda33 | ||
|
|
fcb76d3d5a | ||
|
|
6ae835037b | ||
|
|
b5617fd3a9 | ||
|
|
f5842568b9 | ||
|
|
6f8a4918f0 | ||
|
|
5b0f7950e6 | ||
|
|
c0cf82d4af | ||
|
|
95ec2350af | ||
|
|
3de4beb028 | ||
|
|
d54ec71aea | ||
|
|
2ffd9f9f9c | ||
|
|
f255785b72 | ||
|
|
2fd354c2a9 | ||
|
|
82203038f0 | ||
|
|
096e1ce8b6 | ||
|
|
8550053287 | ||
|
|
e501b93d65 | ||
|
|
62b8fbb8b7 | ||
|
|
efa6cf0899 | ||
|
|
086740b16e | ||
|
|
1f6d028ab5 | ||
|
|
035587e3bf | ||
|
|
e7296e664a | ||
|
|
d0d5e5a704 | ||
|
|
baca89be83 |
+13
@@ -15,8 +15,21 @@ id_*
|
||||
.ansible/
|
||||
facts/
|
||||
|
||||
# Local agent-harness / tooling config (not repo content).
|
||||
.agents/
|
||||
.claude/
|
||||
.omp/
|
||||
.opencode/
|
||||
.zcode/
|
||||
.mcp.json
|
||||
WATCHDOG.yml
|
||||
skills-lock.json
|
||||
|
||||
# Editor and operating-system files.
|
||||
.DS_Store
|
||||
.vscode/
|
||||
.idea/
|
||||
*~
|
||||
# Agent working scratch (not repo content).
|
||||
.agent-work/
|
||||
.tmp-*
|
||||
|
||||
@@ -0,0 +1,8 @@
|
||||
repos:
|
||||
- repo: local
|
||||
hooks:
|
||||
- id: validate-repo
|
||||
name: validate repository
|
||||
entry: scripts/validate-repo.sh
|
||||
language: system
|
||||
pass_filenames: false
|
||||
@@ -2,36 +2,58 @@
|
||||
|
||||
This repo is the **agent ops handbook + fact source** for maintaining personal VPS hosts. Prefer verifying live state over assuming docs are complete.
|
||||
|
||||
Also readable as `agent.md` (symlink → this file).
|
||||
|
||||
## How to work
|
||||
|
||||
1. Read [`inventory/hosts.md`](inventory/hosts.md) for the machine list.
|
||||
2. Open the matching [`hosts/<name>.md`](hosts/) for SSH, roles, paths, and quirks.
|
||||
3. For common tasks, follow a runbook under [`runbooks/`](runbooks/).
|
||||
3. For common tasks, follow a runbook under [`runbooks/`](runbooks/). Pick the
|
||||
most specific applicable one from [`runbooks/README.md`](runbooks/README.md);
|
||||
the spec is [`RUNBOOKS.md`](RUNBOOKS.md) and new runbooks start from
|
||||
[`runbooks/_template.md`](runbooks/_template.md).
|
||||
4. Prefer read-only checks first; change only after confirming current state.
|
||||
5. For routine checks and approved service reconciliation, run the matching
|
||||
Ansible playbook from `ansible/`; see [routine Ansible operations](runbooks/ansible-operations.md).
|
||||
6. Default SSH access (`ssh -4 windy@<host>`) is for focused diagnostics,
|
||||
imperative upstream procedures, and incident work. Prefer **IPv4** from this
|
||||
WSL client (AAAA often exists but IPv6 route does not).
|
||||
|
||||
> **Agent sandbox SSH quirk (verified 2026-08-20):** the agent shell runs in
|
||||
> a sandboxed user namespace — system files such as
|
||||
> `/etc/ssh/ssh_config.d/20-systemd-ssh-proxy.conf` appear owned by `nobody`,
|
||||
> so plain `ssh` aborts with `Bad owner or permissions on ...`. Always use
|
||||
> `ssh -F /dev/null` from the agent shell and pass options explicitly
|
||||
> (`~/.ssh/config` is skipped; e.g. `ssh -F /dev/null -p 2222
|
||||
> -i ~/.ssh/id_ed25519 windy@repo.windy.me`). `sudo` never works in the
|
||||
> sandbox (`NoNewPrivs`, no capabilities, `/` read-only). The host itself is
|
||||
> healthy — to inspect or act on the real host from the sandbox use
|
||||
> `/mnt/c/WINDOWS/system32/wsl.exe -u root -- <cmd>` (real root: keep
|
||||
> read-only unless a change is approved).
|
||||
|
||||
7. Record each material VPS operation, incident, configuration change, or
|
||||
verification outcome in the corresponding **Linear `vps` project**. Include
|
||||
scope, action, verification, and remaining follow-up; never put passwords,
|
||||
tokens, private keys, recovery keys, or private room IDs in Linear.
|
||||
verification outcome in the corresponding **Plane `vps` project**
|
||||
(self-hosted `plane.chans.xyz`, Plane MCP `mcp__plane__*`, following the
|
||||
`plane-workflow` skill). **Linear is retired as a record source (2026-09-03)
|
||||
— do not create Linear issues;** existing W1N-* entries are read-only
|
||||
history. Include scope, action, verification, and remaining follow-up; never
|
||||
put passwords, tokens, private keys, recovery keys, or private room IDs in
|
||||
Plane or Linear.
|
||||
|
||||
## Active hosts (quick map)
|
||||
### Runbook execution rules
|
||||
|
||||
| Host | Role | SSH | Facts |
|
||||
|------|------|-----|--------|
|
||||
| **mx2.windy.me** | mailcow (`/opt/mail`, project `cow`) | `ssh -4 windy@mx2.windy.me` | [hosts/mx2.windy.me.md](hosts/mx2.windy.me.md) |
|
||||
| **us2.wsvc.info** | Vaultwarden + Traefik (+ Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | [hosts/us2.wsvc.info.md](hosts/us2.wsvc.info.md) |
|
||||
| **hk2.chans.xyz** | PowerDNS auth ns1 (`/opt/pdns`) | `ssh -4 windy@hk2.chans.xyz` | [hosts/hk2.chans.xyz.md](hosts/hk2.chans.xyz.md) |
|
||||
| **synapse.chans.xyz** | Matrix ESS (Synapse + MAS + Element) on K3s | `ssh -4 windy@synapse.chans.xyz` | [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md) |
|
||||
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy | `ssh -4 windy@192.168.66.36` | [hosts/dns.windy.lan.md](hosts/dns.windy.lan.md) |
|
||||
| **gfw.windy.lan** | OpenWrt LAN gateway / OpenClash | `ssh -4 root@192.168.66.1` | [hosts/gfw.windy.lan.md](hosts/gfw.windy.lan.md) |
|
||||
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | [hosts/gw.md](hosts/gw.md) |
|
||||
| **ubnt** | UniFi Network Controller | `ssh -4 windy@192.168.66.46` | [hosts/ubnt.md](hosts/ubnt.md) |
|
||||
Before operational work: inspect `runbooks/`, select the most specific
|
||||
applicable runbook, follow its steps in order, do not skip verification steps,
|
||||
and respect its STOP and approval conditions. If no runbook applies, diagnose
|
||||
only — do not mutate production state. When live state conflicts with a
|
||||
runbook's assumptions, `STOP` and report; never invent missing parameters or
|
||||
bypass failed checks. The spec is [`RUNBOOKS.md`](RUNBOOKS.md).
|
||||
|
||||
## Active hosts
|
||||
|
||||
The canonical machine list (roles, SSH endpoints, Ansible coverage, status) is
|
||||
[`inventory/hosts.md`](inventory/hosts.md) — the single human-readable source
|
||||
of truth. Per-host facts live in [`hosts/`](hosts/). The Ansible execution
|
||||
inventory is [`ansible/inventory/hosts.yml`](ansible/inventory/hosts.yml). Do
|
||||
not maintain a second copy of the machine table here.
|
||||
|
||||
### Public services
|
||||
|
||||
@@ -41,7 +63,7 @@ Also readable as `agent.md` (symlink → this file).
|
||||
| SMTP `mx2.windy.me:587` (STARTTLS) or `:465` | mx2 | client submission; full email + mailbox password — [runbook](runbooks/mailcow-smtp-client.md) |
|
||||
| IMAP `mx2.windy.me:993` | mx2 | same mailbox credentials |
|
||||
| https://auth.wsvc.info | us2 (`/opt/vaultwarden`) | Vaultwarden (Postgres, **operational**) — client Server URL |
|
||||
| `repo.windy.me:2222` | us2 (`/opt/soft-serve`) | Soft Serve (stub details) |
|
||||
| `repo.windy.me` (git SSH `:2222` / web HTTPS) | us2 (`/opt/gitea`) | Gitea — 1.27.3-rootless pinned, backup sidecar; details in [hosts/us2.wsvc.info.md](hosts/us2.wsvc.info.md) |
|
||||
| DNS `ns1.wsvc.info:53` | hk2 (`/opt/pdns`, Auth **5.0.6**) | PowerDNS auth — zones `windy.me`, `wsvc.info`, `chans.xyz` |
|
||||
| https://pdns.wsvc.info | hk2 (`poweradmin`) | Poweradmin UI |
|
||||
| https://pgweb.wsvc.info | hk2 (`pgweb`) | PowerDNS Postgres browser |
|
||||
@@ -49,6 +71,7 @@ Also readable as `agent.md` (symlink → this file).
|
||||
| https://synapse.chans.xyz | synapse | Synapse Client-Server + Federation API |
|
||||
| https://account.chans.xyz | synapse | Matrix Authentication Service (local passwords) |
|
||||
| https://admin.chans.xyz | synapse | Element Admin console (MAS admin auth) |
|
||||
| https://plane.chans.xyz | synapse (`plane`, Helm `plane-ce` 1.8.0 / v1.4.1) | Plane project management (self-hosted, K3s) |
|
||||
|
||||
### Upstream docs
|
||||
|
||||
@@ -58,6 +81,10 @@ Also readable as `agent.md` (symlink → this file).
|
||||
|
||||
**Matrix (ESS on synapse):** Matrix homeserver running on `synapse.chans.xyz` via the official ESS (Element Server Suite) Helm chart with Synapse + MAS + Element Web + Admin. DNS zone `chans.xyz` managed by hk2 PowerDNS. Before changing config, read [docs/matrix-upstream.md](docs/matrix-upstream.md) and [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md). K3s cluster on this node has hostPort 80/443 for Traefik (no ServiceLB). Health: [matrix-health](runbooks/matrix-health.md).
|
||||
|
||||
**Plane (on synapse):** Self-hosted Plane project management at `plane.chans.xyz`, Helm release `plane-app` (chart `plane-ce-1.8.0`, app `v1.4.1`) in ns `plane` on the same K3s node as Matrix. Config from `/home/windy/plane-k3s/values.yaml`; workload/cert/ingress details in [hosts/synapse.chans.xyz.md](hosts/synapse.chans.xyz.md). Its Postgres/MinIO PVCs are **not** backed up.
|
||||
|
||||
**RustDesk:** Self-hosted RustDesk server on `hk2.chans.xyz` (`/opt/rustdesk`, containers `hbbs`/`hbbr`, image pinned `1.1.14`). The `hbbs -r` relay hostname must resolve to the host's public IP `154.36.174.161` — use `hk2.chans.xyz` (never `hk2.wsvc.info`, which has no DNS record). Health: [rustdesk-health](runbooks/rustdesk-health.md).
|
||||
|
||||
## Runbooks & scripts
|
||||
|
||||
| Task | Path |
|
||||
@@ -71,12 +98,26 @@ Also readable as `agent.md` (symlink → this file).
|
||||
| PowerDNS health (hk2) | [runbooks/pdns-health.md](runbooks/pdns-health.md) |
|
||||
| PowerDNS upstream refs | [docs/pdns-upstream.md](docs/pdns-upstream.md) |
|
||||
| Matrix health | [runbooks/matrix-health.md](runbooks/matrix-health.md) |
|
||||
| Plane health | [runbooks/plane-health.md](runbooks/plane-health.md) |
|
||||
| RustDesk health (hk2) | [runbooks/rustdesk-health.md](runbooks/rustdesk-health.md) |
|
||||
| AdGuard Home health | [runbooks/adguard-home-health.md](runbooks/adguard-home-health.md) |
|
||||
| Host disk cleanup | [runbooks/host-disk-cleanup.md](runbooks/host-disk-cleanup.md) |
|
||||
| Home Assistant maintenance | [runbooks/home-assistant-maintenance.md](runbooks/home-assistant-maintenance.md) + [scripts/ha-maintenance.sh](runbooks/scripts/ha-maintenance.sh) |
|
||||
| matrix_e2ee update (hass.windy.lan) | [runbooks/matrix-e2ee-update.md](runbooks/matrix-e2ee-update.md) |
|
||||
| Matrix upstream refs | [docs/matrix-upstream.md](docs/matrix-upstream.md) |
|
||||
| Hermes Agent Matrix channel | [docs/hermes-matrix.md](docs/hermes-matrix.md) |
|
||||
| UniFi local-service proxy bypass | [docs/unifi-openclash-localhost.md](docs/unifi-openclash-localhost.md) |
|
||||
| UniFi SSO login setting (Ansible) | `cd ansible && ansible-playbook playbooks/unifi-sso.yml --limit unifi` |
|
||||
| Routine Ansible operations | [runbooks/ansible-operations.md](runbooks/ansible-operations.md) |
|
||||
| Routine make commands | `make help` (wraps `ansible-operations.md` read-only + gated flows) |
|
||||
| Issue → mergeable change | [runbooks/issue-to-merge.md](runbooks/issue-to-merge.md) |
|
||||
| Fix failing health/playbook run | [runbooks/fix-ci.md](runbooks/fix-ci.md) |
|
||||
| Release a reviewed change | [runbooks/release.md](runbooks/release.md) |
|
||||
| Roll back a change | [runbooks/rollback.md](runbooks/rollback.md) |
|
||||
| Controlled network change | [runbooks/network-change.md](runbooks/network-change.md) |
|
||||
| Network outage recovery | [runbooks/network-recovery.md](runbooks/network-recovery.md) |
|
||||
|
||||
Full index: [runbooks/README.md](runbooks/README.md). Spec: [RUNBOOKS.md](RUNBOOKS.md).
|
||||
|
||||
Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.yml --limit mailcow`. The local stub resolver is flaky; DNS probes use `1.1.1.1` / `8.8.8.8`.
|
||||
|
||||
@@ -84,12 +125,22 @@ Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.
|
||||
|
||||
### Issue tracker
|
||||
|
||||
Issues are tracked in Linear and created/updated via the Linear MCP (`vps` project). See `docs/agents/issue-tracker.md`.
|
||||
Issues are tracked in **Plane** — self-hosted at `plane.chans.xyz`, project
|
||||
`vps` — and created/updated via the Plane MCP (`mcp__plane__*`), following the
|
||||
`plane-workflow` skill. **Linear is retired as a record source (2026-09-03); do
|
||||
not create Linear issues.** Existing W1N-* entries are read-only history.
|
||||
`docs/agents/issue-tracker.md` documents the retired Linear workflow and is
|
||||
stale; treat this section as authoritative.
|
||||
|
||||
### Triage labels
|
||||
|
||||
Default triage labels: needs-triage, needs-info, ready-for-agent, ready-for-human, wontfix. See `docs/agents/triage-labels.md`.
|
||||
|
||||
### Domain docs
|
||||
|
||||
Domain-documentation conventions, including lazily created `CONTEXT.md` and
|
||||
`docs/adr/` entries when needed, are described in [`docs/agents/domain.md`](docs/agents/domain.md).
|
||||
|
||||
## Safety
|
||||
|
||||
- Never commit secrets: passwords, API keys, private keys, `.env`, `mailcow.conf` DB passwords, Vaultwarden `ADMIN_TOKEN` / `.smtp-credentials`.
|
||||
@@ -121,9 +172,14 @@ Bills, rough notes, and personal clutter stay in the Obsidian vault. This repo h
|
||||
## Layout
|
||||
|
||||
```
|
||||
AGENTS.md / agent.md # this entry (agent.md → AGENTS.md)
|
||||
inventory/hosts.md # machine index
|
||||
AGENTS.md # this entry
|
||||
RUNBOOKS.md # runbook spec (six-field model, naming, review rules)
|
||||
inventory/hosts.md # machine index (human-readable source of truth)
|
||||
ansible/ # playbooks, roles, sanitized control-plane inventory
|
||||
compose/ # repo-owned non-secret Compose sources (+ .env.example)
|
||||
hosts/ # per-host facts
|
||||
runbooks/ # step-by-step ops
|
||||
docs/ # upstream doc indexes / design notes
|
||||
runbooks/ # step-by-step ops (README.md = index, _template.md = template)
|
||||
docs/ # upstream refs / design notes / research records (active + archive/)
|
||||
scripts/validate-repo.sh # repo-wide validation (run before merging)
|
||||
Makefile # routine validate / health / gated ansible wrappers
|
||||
```
|
||||
|
||||
@@ -0,0 +1,183 @@
|
||||
# VPS ops hub — routine validate / health / gated Ansible wrappers.
|
||||
# See runbooks/ansible-operations.md for playbook semantics.
|
||||
|
||||
SHELL := /usr/bin/env bash
|
||||
.SHELLFLAGS := -eu -o pipefail -c
|
||||
|
||||
.DEFAULT_GOAL := help
|
||||
|
||||
REPO_ROOT := $(CURDIR)
|
||||
ANSIBLE_DIR := $(REPO_ROOT)/ansible
|
||||
export ANSIBLE_LOCAL_TEMP := $(REPO_ROOT)/.ansible/tmp
|
||||
export ANSIBLE_HOME := $(REPO_ROOT)/.ansible
|
||||
|
||||
LIMIT ?=
|
||||
EXTRA ?=
|
||||
VERBOSE ?= 0
|
||||
CONFIRM ?= 0
|
||||
TARGETS ?=
|
||||
TRAEFIK ?= 0
|
||||
|
||||
LIMIT_FLAG := $(if $(LIMIT),--limit $(LIMIT),)
|
||||
VERBOSE_FLAG := $(if $(filter 1,$(VERBOSE)),-v,$(if $(filter 2,$(VERBOSE)),-vvv,))
|
||||
|
||||
.PHONY: help validate check deps galaxy syntax ansible-prep \
|
||||
ping inventory audit health health-mailcow health-matrix \
|
||||
maint-preview baseline compose-check \
|
||||
install-healthchecks install-matrix-healthchecks compose-deploy reconcile
|
||||
|
||||
help:
|
||||
@printf '%s\n' \
|
||||
'VPS ops hub — make targets (run from repo root)' \
|
||||
'' \
|
||||
'Variables: LIMIT=<group|host> CONFIRM=1 TARGETS=<svc[,svc]> TRAEFIK=1 VERBOSE=0|1|2 EXTRA=...' \
|
||||
'' \
|
||||
'Local / repo:' \
|
||||
' validate, check scripts/validate-repo.sh (pre-merge gate)' \
|
||||
' deps, galaxy ansible-galaxy collection install' \
|
||||
' syntax ansible-playbook --syntax-check all playbooks' \
|
||||
'' \
|
||||
'Read-only remote (ansible):' \
|
||||
' ping ansible managed -m ping' \
|
||||
' inventory ansible-inventory --graph' \
|
||||
' audit playbooks/audit.yml' \
|
||||
' health [LIMIT=…] playbooks/health-report.yml' \
|
||||
' health-mailcow health --limit mailcow' \
|
||||
' health-matrix health --limit matrix' \
|
||||
' maint-preview playbooks/maintenance-preview.yml' \
|
||||
' baseline playbooks/baseline.yml' \
|
||||
' compose-check compose-deploy --check --diff (requires LIMIT=)' \
|
||||
'' \
|
||||
'Mutating (require CONFIRM=1; host-scoped targets require LIMIT=):' \
|
||||
' install-healthchecks playbooks/healthchecks.yml' \
|
||||
' install-matrix-healthchecks playbooks/matrix-healthchecks.yml' \
|
||||
' compose-deploy playbooks/compose-deploy.yml' \
|
||||
' reconcile playbooks/compose-reconcile.yml (requires TARGETS=)' \
|
||||
'' \
|
||||
'Examples:' \
|
||||
' make validate' \
|
||||
' make health LIMIT=mailcow' \
|
||||
' make compose-check LIMIT=vaultwarden' \
|
||||
' make compose-deploy LIMIT=vaultwarden CONFIRM=1' \
|
||||
' make reconcile LIMIT=powerdns TARGETS=auth CONFIRM=1' \
|
||||
' make reconcile LIMIT=vaultwarden TARGETS=vaultwarden TRAEFIK=1 CONFIRM=1' \
|
||||
'' \
|
||||
'Advanced (not wrapped — use ansible-playbook directly):' \
|
||||
' us4-firewalld, unifi-sso, k3s-server, matrix-stack, wireguard-harden,' \
|
||||
' restic, rustdesk, email-alerts, mailcow update runbook'
|
||||
|
||||
validate check:
|
||||
@bash "$(REPO_ROOT)/scripts/validate-repo.sh"
|
||||
|
||||
deps galaxy: ansible-prep
|
||||
@command -v ansible-galaxy >/dev/null 2>&1 || { echo "ansible-galaxy not found; install Ansible first." >&2; exit 1; }
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-galaxy collection install -r requirements.yml
|
||||
|
||||
syntax: ansible-prep
|
||||
@if ! command -v ansible-playbook >/dev/null 2>&1; then \
|
||||
echo "ansible-playbook not found; syntax check skipped." >&2; \
|
||||
exit 0; \
|
||||
fi
|
||||
@fail=0; \
|
||||
for p in "$(ANSIBLE_DIR)"/playbooks/*.yml; do \
|
||||
if ! (cd "$(ANSIBLE_DIR)" && ansible-playbook --syntax-check "playbooks/$$(basename "$$p")" >/dev/null 2>&1); then \
|
||||
echo "syntax-check failed: $$p" >&2; \
|
||||
fail=1; \
|
||||
fi; \
|
||||
done; \
|
||||
exit $$fail
|
||||
|
||||
ansible-prep:
|
||||
@mkdir -p "$(ANSIBLE_HOME)/tmp" "$(ANSIBLE_HOME)/ssh-control"
|
||||
|
||||
define require_ansible
|
||||
@command -v ansible-playbook >/dev/null 2>&1 || { echo "ansible-playbook not found; install Ansible first." >&2; exit 1; }
|
||||
endef
|
||||
|
||||
define require_limit
|
||||
@if [ -z "$(LIMIT)" ]; then \
|
||||
echo "LIMIT is required (e.g. LIMIT=mailcow, LIMIT=vaultwarden, LIMIT=powerdns)." >&2; \
|
||||
exit 1; \
|
||||
fi
|
||||
endef
|
||||
|
||||
define require_confirm
|
||||
@if [ "$(CONFIRM)" != "1" ]; then \
|
||||
echo "Mutating operation blocked. Re-run with CONFIRM=1" >&2; \
|
||||
exit 1; \
|
||||
fi
|
||||
endef
|
||||
|
||||
define require_targets
|
||||
@if [ -z "$(TARGETS)" ]; then \
|
||||
echo "TARGETS is required (comma-separated service names, e.g. TARGETS=auth or TARGETS=vaultwarden)." >&2; \
|
||||
exit 1; \
|
||||
fi
|
||||
endef
|
||||
|
||||
ping: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible managed -m ping $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
inventory: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-inventory --graph $(EXTRA)
|
||||
|
||||
audit: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/audit.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
health: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
health-mailcow: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml --limit mailcow $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
health-matrix: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/health-report.yml --limit matrix $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
maint-preview: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/maintenance-preview.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
baseline: ansible-prep
|
||||
$(require_ansible)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/baseline.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
compose-check: ansible-prep
|
||||
$(require_ansible)
|
||||
$(require_limit)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-deploy.yml --check --diff --limit $(LIMIT) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
install-healthchecks: ansible-prep
|
||||
$(require_ansible)
|
||||
$(require_confirm)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/healthchecks.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
install-matrix-healthchecks: ansible-prep
|
||||
$(require_ansible)
|
||||
$(require_confirm)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/matrix-healthchecks.yml $(LIMIT_FLAG) $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
compose-deploy: ansible-prep
|
||||
$(require_ansible)
|
||||
$(require_limit)
|
||||
$(require_confirm)
|
||||
@cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-deploy.yml --limit $(LIMIT) \
|
||||
-e '{"compose_deploy_confirm": true}' $(VERBOSE_FLAG) $(EXTRA)
|
||||
|
||||
reconcile: ansible-prep
|
||||
$(require_ansible)
|
||||
$(require_limit)
|
||||
$(require_targets)
|
||||
$(require_confirm)
|
||||
@json=$$(python3 -c 'import json,sys; t=[x.strip() for x in sys.argv[1].split(",") if x.strip()]; \
|
||||
(not t) and sys.exit("TARGETS must contain at least one non-empty service name"); \
|
||||
d={"service_reconcile_confirm": True, "service_reconcile_targets": t}; \
|
||||
(sys.argv[2]=="1") and d.update({"service_reconcile_restart_traefik": True}); \
|
||||
print(json.dumps(d))' "$(TARGETS)" "$(TRAEFIK)"); \
|
||||
cd "$(ANSIBLE_DIR)" && ansible-playbook playbooks/compose-reconcile.yml --limit $(LIMIT) \
|
||||
-e "$$json" $(VERBOSE_FLAG) $(EXTRA)
|
||||
+89
@@ -0,0 +1,89 @@
|
||||
# RUNBOOKS — 仓库级规范
|
||||
|
||||
本文件统一所有 Runbook 的字段、命名、评审与变更规则。上游参考:[docs/archive/agent-runbook-guide.md](docs/archive/agent-runbook-guide.md)。
|
||||
|
||||
## 目录结构
|
||||
|
||||
```text
|
||||
runbooks/
|
||||
├── README.md # 意图 → 文件 路由索引(本目录的入口)
|
||||
├── _template.md # 新建 runbook 的标准模板(复制后填写)
|
||||
├── <intent>.md # 每份 runbook 只描述一种可识别的操作意图
|
||||
└── ...
|
||||
```
|
||||
|
||||
## 最小字段模型
|
||||
|
||||
每份 runbook 必须显式包含以下控制信息,否则盲目执行或错误恢复的风险会升高:
|
||||
|
||||
| 字段 | 作用 | 写作要求 |
|
||||
|---|---|---|
|
||||
| **Action** | 定义当前要执行的动作 | 可观察、可执行的动词;避免“检查一下”“适当调整” |
|
||||
| **Expected** | 描述正常状态或预期输出 | 具体信号、阈值、状态码、测试结果或页面表现 |
|
||||
| **Decision** | 定义分支与下一跳 | “条件 → 下一步”;无法判断时指向 `STOP` |
|
||||
| **Verification** | 确认变更真正生效 | 每个有副作用的步骤后执行,不可跳过 |
|
||||
| **Stop condition** | 规定何时不得继续 | 列出信息缺失、状态冲突、权限不足、验证失败等 |
|
||||
| **Rollback** | 如何恢复到变更前状态 | 触发条件、前提、撤销步骤、回滚后验证 |
|
||||
|
||||
> 只读类 runbook 不产生副作用,可省略 Rollback;但必须保留 Stop condition(状态与预期冲突即 `STOP` 并记录证据)。
|
||||
|
||||
**只读类变体(read-only variant)**:只读 runbook(health 类、参考类)不强制
|
||||
六字段模型,但必须包含以下最小结构,否则不视为达标:
|
||||
|
||||
- `## Purpose`(1–2 行)+ `## Scope`(适用/不适用)
|
||||
- `## Safety` 或等效章节,其中**必须**含显式 Stop condition(状态与预期冲突即
|
||||
`STOP` 并记录证据;不得在执行中自行"顺手修复")
|
||||
- 只读健康类另含可观察的 `## Pass criteria`(或等效的 Expected 信号)
|
||||
- 每份 runbook 顶部/元信息区必须标注 `Last reviewed: <YYYY-MM-DD>`
|
||||
|
||||
## 命名与拆分规则
|
||||
|
||||
- 文件名采用小写连字符,反映**操作意图**而非目标主机,例如 `mailcow-health.md`、`release.md`。
|
||||
- 一份文件只描述一种意图。流程出现明显分叉时拆分为独立文件,不堆叠“万能流程”。
|
||||
- 只读诊断与变更操作应分离:health 类 runbook 保持只读,变更走 `ansible-operations.md`、`release.md`、`rollback.md` 或对应 gated playbook。
|
||||
|
||||
## 章节约定
|
||||
|
||||
- 每份 runbook 顶部含 `## Purpose`(1–2 行)与 `## Scope`(适用/不适用情形)。
|
||||
- 变更型 runbook 必须记录明确的审批门:门控命令式使用 `## Approval gates` 表;
|
||||
流程式在步骤中记录审批动作、证据位置和未批准时的 `STOP`。破坏性/不可逆操作必须获得明确批准。
|
||||
- 语言约定:**runbook 正文统一使用英文**(由 agent 逐字执行,降低二义性);
|
||||
元规范文件(AGENTS.md / RUNBOOKS.md / 模板注释)可保留中文。
|
||||
- 变更型 runbook 的两种形态:
|
||||
- **流程式(Procedure 型)**:使用六字段模型,适用多分支/多步骤变更
|
||||
(现有:`fix-ci.md`、`issue-to-merge.md`、`network-change.md`、
|
||||
`network-recovery.md`、`release.md`、`rollback.md`)。
|
||||
- **门控命令式(gated command reference)**:已稳定、低歧义、可验证的
|
||||
操作以命令集 + 门控呈现(现有:`mailcow-update.md`、
|
||||
`ansible-operations.md`、`home-assistant-maintenance.md`、
|
||||
`matrix-e2ee-update.md`、`vaultwarden-sqlite-to-postgres.md`),必须含 Approval gates 或确认变量
|
||||
要求 + 显式 STOP,不替代流程式形态。新写的变更 runbook 默认用流程式。
|
||||
- 统一在 `## Safety` 或正文中复用以下通用安全规则(更严格要求优先)。
|
||||
|
||||
```markdown
|
||||
## Safety Rules
|
||||
|
||||
- Never delete an existing configuration as the first recovery action.
|
||||
- Prefer read-only diagnosis before mutation.
|
||||
- After every mutation, verify the expected state.
|
||||
- If actual state conflicts with this runbook, STOP.
|
||||
- Do not invent missing parameters.
|
||||
- Do not bypass failed tests.
|
||||
- Destructive actions require explicit approval.
|
||||
```
|
||||
|
||||
## 评审与变更规则
|
||||
|
||||
- 新建/修改 runbook 与代码同仓评审,随系统演进更新。
|
||||
- 每份 runbook 标注 `Last reviewed`;流程执行过程中发现的偏差记入对应的 Linear `vps` 项目 issue。
|
||||
- 破坏性流程(迁移、删除、DNS 变更、网络变更)保持人工审批,不自动下沉。
|
||||
|
||||
## 成熟路径
|
||||
|
||||
1. **人工处理** → 现场处置与复盘,记录证据。
|
||||
2. **Markdown runbook** → 固化步骤与证据要求,Agent 可辅助诊断。
|
||||
3. **Agent + runbook** → 严格按流程执行,受 Stop/Approval 约束。
|
||||
4. **Script / Ansible / Skill** → 把已稳定、低歧义、可验证的操作程序化(本仓库的执行层是 Ansible playbook)。
|
||||
5. **人工审批 + 自动执行** → 审批门控下的自动变更(如 gated playbook + 确认变量)。
|
||||
|
||||
原则:先证据后变更,先小范围后扩大,先验证后结束,不确定则停止。
|
||||
@@ -11,3 +11,8 @@ host_key_checking = True
|
||||
become = True
|
||||
become_method = sudo
|
||||
become_ask_pass = False
|
||||
|
||||
[ssh_connection]
|
||||
# Keep SSH control sockets inside the repo (gitignored .ansible/) so playbook
|
||||
# runs work in sandboxed/CI environments without touching ~/.ansible.
|
||||
ssh_args = -C -o ControlMaster=auto -o ControlPersist=60s -o ControlPath=.ansible/ssh-control/%h-%p-%r
|
||||
|
||||
@@ -15,18 +15,23 @@ all:
|
||||
mx2:
|
||||
ansible_host: mx2.windy.me
|
||||
ansible_host_ipv4: 194.163.160.244
|
||||
display_name: mx2.windy.me
|
||||
service_role: mailcow
|
||||
compose_project_dir: /opt/mail
|
||||
healthcheck_profile: mailcow
|
||||
healthcheck_profiles: [mailcow]
|
||||
service_reconcile_services:
|
||||
all:
|
||||
compose_args: [--force-recreate]
|
||||
us2:
|
||||
ansible_host: us2.wsvc.info
|
||||
ansible_host_ipv4: 193.9.44.165
|
||||
display_name: us2.wsvc.info
|
||||
service_role: vaultwarden
|
||||
compose_project_dir: /opt/vaultwarden
|
||||
healthcheck_profile: vaultwarden
|
||||
compose_repo_project: vaultwarden
|
||||
compose_remote_file: docker-compose.yml
|
||||
healthcheck_profiles: [vaultwarden]
|
||||
restic_backup_profile: vaultwarden
|
||||
service_reconcile_services:
|
||||
vaultwarden:
|
||||
compose_args: [--force-recreate]
|
||||
@@ -34,9 +39,13 @@ all:
|
||||
hk2:
|
||||
ansible_host: hk2.chans.xyz
|
||||
ansible_host_ipv4: 154.36.174.161
|
||||
display_name: hk2.chans.xyz
|
||||
service_role: powerdns
|
||||
compose_project_dir: /opt/pdns
|
||||
healthcheck_profile: pdns
|
||||
compose_repo_project: pdns
|
||||
compose_remote_file: compose.yml
|
||||
healthcheck_profiles: [pdns, rustdesk, hk2aux]
|
||||
restic_backup_profile: pdns
|
||||
service_reconcile_services:
|
||||
auth:
|
||||
compose_args: [--force-recreate]
|
||||
@@ -45,12 +54,17 @@ all:
|
||||
backup:
|
||||
compose_args: [--no-deps, --force-recreate]
|
||||
service_reconcile_traefik_restart_targets: [poweradmin]
|
||||
# RustDesk server (same host, separate compose project)
|
||||
rustdesk_compose_dir: /opt/rustdesk
|
||||
rustdesk_relay: hk2.chans.xyz:21117
|
||||
rustdesk_image: rustdesk/rustdesk-server:1.1.14
|
||||
us4:
|
||||
ansible_host: us4.wsvc.info
|
||||
ansible_host_ipv4: 185.201.226.122
|
||||
display_name: us4.wsvc.info
|
||||
service_role: wireguard
|
||||
compose_project_dir: /opt/wireguard
|
||||
healthcheck_profile: wireguard
|
||||
healthcheck_profiles: [wireguard]
|
||||
wireguard_image: >-
|
||||
lscr.io/linuxserver/wireguard@sha256:ac43e1226878d2611315172d6ea357a95cb326ee73124b91108118efc8666889
|
||||
service_reconcile_services:
|
||||
@@ -59,9 +73,10 @@ all:
|
||||
dns_windy_lan:
|
||||
ansible_host: 192.168.66.36
|
||||
ansible_host_ipv4: 192.168.66.36
|
||||
display_name: dns.windy.lan
|
||||
service_role: adguardhome
|
||||
compose_project_dir: /opt/adguardhome
|
||||
healthcheck_profile: adguardhome
|
||||
healthcheck_profiles: [adguardhome]
|
||||
service_reconcile_services:
|
||||
adguardhome:
|
||||
compose_args: [--no-deps, --force-recreate]
|
||||
@@ -74,6 +89,9 @@ all:
|
||||
powerdns:
|
||||
hosts:
|
||||
hk2:
|
||||
rustdesk:
|
||||
hosts:
|
||||
hk2:
|
||||
wireguard:
|
||||
hosts:
|
||||
us4:
|
||||
@@ -85,6 +103,7 @@ all:
|
||||
ubnt:
|
||||
ansible_host: 192.168.66.46
|
||||
ansible_host_ipv4: 192.168.66.46
|
||||
display_name: ubnt
|
||||
vars:
|
||||
service_role: unifi
|
||||
compose_project_dir: /home/windy/unifi-9
|
||||
@@ -105,6 +124,7 @@ all:
|
||||
matrix_vps:
|
||||
ansible_host: 169.58.86.13
|
||||
ansible_host_ipv4: 169.58.86.13
|
||||
display_name: synapse.chans.xyz
|
||||
service_role: matrix_k3s
|
||||
matrix_server_name: chans.xyz
|
||||
matrix_synapse_host: synapse.chans.xyz
|
||||
|
||||
@@ -64,7 +64,7 @@
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
host: "{{ inventory_hostname }}"
|
||||
profile: "{{ healthcheck_profile }}"
|
||||
profiles: "{{ healthcheck_profiles | default([]) | join(', ') }}"
|
||||
os: "{{ ansible_distribution }} {{ ansible_distribution_version }}"
|
||||
kernel: "{{ ansible_kernel }}"
|
||||
compose_rc: "{{ audit_compose_ps.rc }}"
|
||||
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
# Deploy repo-owned Compose declarations (compose/<project>/compose.yml) to
|
||||
# inventory hosts. Non-secret source; server-local .env provides the values.
|
||||
# Gated: apply requires compose_deploy_confirm=true; --check is a read-only
|
||||
# diff + validation. See runbooks/ansible-operations.md.
|
||||
- name: Deploy repo-owned Compose declarations
|
||||
hosts: docker_hosts
|
||||
become: true
|
||||
gather_facts: false
|
||||
serial: 1
|
||||
roles:
|
||||
- role: compose_deploy
|
||||
tags: [compose, deploy, mutating]
|
||||
@@ -0,0 +1,20 @@
|
||||
---
|
||||
# Deploy/reconcile the self-hosted RustDesk server (hbbs + hbbr) on hk2.
|
||||
#
|
||||
# Safe by default: run with --check for a read-only report, or supply
|
||||
# rustdesk_confirm=true to deploy the compose file and recreate the stack.
|
||||
#
|
||||
# # Read-only report
|
||||
# ansible-playbook playbooks/rustdesk.yml --limit rustdesk --check
|
||||
#
|
||||
# # Apply (deploy compose + recreate hbbs/hbbr)
|
||||
# ansible-playbook playbooks/rustdesk.yml --limit rustdesk \
|
||||
# -e '{"rustdesk_confirm": true}'
|
||||
- name: Deploy and reconcile RustDesk server
|
||||
hosts: rustdesk
|
||||
become: true
|
||||
gather_facts: false
|
||||
serial: 1
|
||||
roles:
|
||||
- role: rustdesk
|
||||
tags: [rustdesk, mutating]
|
||||
@@ -0,0 +1,456 @@
|
||||
---
|
||||
# Narrow reconciliation for the audited us4 public zone. This playbook never
|
||||
# reloads or restarts firewalld and deliberately does not manage Docker rules.
|
||||
- name: Safely remove audited stale firewalld allowances from us4
|
||||
hosts: wireguard
|
||||
become: true
|
||||
gather_facts: false
|
||||
serial: 1
|
||||
any_errors_fatal: true
|
||||
vars:
|
||||
us4_firewalld_confirm: false
|
||||
us4_console_confirm: false
|
||||
us4_firewalld_zone: public
|
||||
us4_firewalld_keep_services:
|
||||
- dhcpv6-client
|
||||
- http
|
||||
- https
|
||||
- smtp
|
||||
- ssh
|
||||
us4_firewalld_stale_services:
|
||||
- imap
|
||||
- imaps
|
||||
- smtp-submission
|
||||
- smtps
|
||||
us4_firewalld_stale_ports:
|
||||
- 24/tcp
|
||||
- 6443/tcp
|
||||
- 8443/tcp
|
||||
us4_firewalld_expected_containers:
|
||||
- nghttpx-proxy
|
||||
- semaphoreui-postgres-1
|
||||
- semaphoreui-semaphore-1
|
||||
- squid-backend
|
||||
- traefik
|
||||
- trlm-server-trilium-1
|
||||
- wireguard
|
||||
us4_firewalld_backup_root: /var/backups/us4-firewall
|
||||
|
||||
tasks:
|
||||
- name: Require the audited host and explicit apply confirmations
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- inventory_hostname == 'us4'
|
||||
- ansible_host == 'us4.wsvc.info'
|
||||
- ansible_host_ipv4 == '185.201.226.122'
|
||||
- ansible_check_mode or (us4_firewalld_confirm | bool)
|
||||
- ansible_check_mode or (us4_console_confirm | bool)
|
||||
fail_msg: >-
|
||||
Apply is allowed only for audited host us4 after the provider console
|
||||
has been tested. Set both us4_firewalld_confirm=true and
|
||||
us4_console_confirm=true. Check mode does not require confirmation.
|
||||
|
||||
- name: Verify the remote host identity
|
||||
ansible.builtin.command:
|
||||
argv: [hostname, -f]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
register: us4_firewalld_hostname
|
||||
|
||||
- name: Reject an unexpected remote host
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- us4_firewalld_hostname.stdout == 'us4.wsvc.info'
|
||||
|
||||
- name: Verify required services are active
|
||||
ansible.builtin.command:
|
||||
argv: [systemctl, is-active, --quiet, "{{ item }}"]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
loop:
|
||||
- atd
|
||||
- firewalld
|
||||
|
||||
- name: Verify firewalld Python bindings used by ansible.posix
|
||||
ansible.builtin.command:
|
||||
argv: [python3, -c, "import dbus, firewall, firewall.client"]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
|
||||
- name: Verify the default firewalld zone
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --get-default-zone]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
register: us4_firewalld_default_zone
|
||||
|
||||
- name: Read runtime public-zone services
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --zone=public, --list-services]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
register: us4_firewalld_runtime_services
|
||||
|
||||
- name: Read permanent public-zone services
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --permanent, --zone=public, --list-services]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
register: us4_firewalld_permanent_services
|
||||
|
||||
- name: Read runtime public-zone ports
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --zone=public, --list-ports]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
register: us4_firewalld_runtime_ports
|
||||
|
||||
- name: Read permanent public-zone ports
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --permanent, --zone=public, --list-ports]
|
||||
check_mode: false
|
||||
changed_when: false
|
||||
register: us4_firewalld_permanent_ports
|
||||
|
||||
- name: Normalize the audited public-zone state
|
||||
ansible.builtin.set_fact:
|
||||
us4_firewalld_pre_services: "{{ us4_firewalld_runtime_services.stdout.split() | sort }}"
|
||||
us4_firewalld_pre_permanent_services: "{{ us4_firewalld_permanent_services.stdout.split() | sort }}"
|
||||
us4_firewalld_pre_ports: "{{ us4_firewalld_runtime_ports.stdout.split() | sort }}"
|
||||
us4_firewalld_pre_permanent_ports: "{{ us4_firewalld_permanent_ports.stdout.split() | sort }}"
|
||||
|
||||
- name: Fail closed on public-zone drift or unknown allowances
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- us4_firewalld_default_zone.stdout == us4_firewalld_zone
|
||||
- us4_firewalld_pre_services == us4_firewalld_pre_permanent_services
|
||||
- us4_firewalld_pre_ports == us4_firewalld_pre_permanent_ports
|
||||
- us4_firewalld_keep_services | difference(us4_firewalld_pre_services) | length == 0
|
||||
- us4_firewalld_pre_services | difference(us4_firewalld_keep_services + us4_firewalld_stale_services) | length == 0
|
||||
- us4_firewalld_pre_ports | difference(us4_firewalld_stale_ports) | length == 0
|
||||
fail_msg: >-
|
||||
The public zone differs from the audited baseline. Stop and review it;
|
||||
this playbook will not infer whether an unknown allowance is required.
|
||||
|
||||
- name: Select only audited stale entries that currently exist
|
||||
ansible.builtin.set_fact:
|
||||
us4_firewalld_cleanup_services: >-
|
||||
{{ us4_firewalld_stale_services | intersect(us4_firewalld_pre_services) | sort }}
|
||||
us4_firewalld_cleanup_ports: >-
|
||||
{{ us4_firewalld_stale_ports | intersect(us4_firewalld_pre_ports) | sort }}
|
||||
|
||||
- name: Report the proposed reconciliation
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
keep_services: "{{ us4_firewalld_keep_services }}"
|
||||
remove_services: "{{ us4_firewalld_cleanup_services }}"
|
||||
remove_ports: "{{ us4_firewalld_cleanup_ports }}"
|
||||
reload_or_restart: false
|
||||
|
||||
- name: Create rollback material when cleanup is required
|
||||
when:
|
||||
- not ansible_check_mode
|
||||
- us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0
|
||||
block:
|
||||
- name: Create the protected firewall backup root
|
||||
ansible.builtin.file:
|
||||
path: "{{ us4_firewalld_backup_root }}"
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0700"
|
||||
|
||||
- name: Create a backup timestamp
|
||||
ansible.builtin.command:
|
||||
argv: [date, +%Y%m%dT%H%M%S%z]
|
||||
changed_when: false
|
||||
register: us4_firewalld_backup_timestamp
|
||||
|
||||
- name: Set the protected backup directory
|
||||
ansible.builtin.set_fact:
|
||||
us4_firewalld_backup_dir: >-
|
||||
{{ us4_firewalld_backup_root }}/{{ us4_firewalld_backup_timestamp.stdout }}
|
||||
us4_firewalld_rollback_command: >-
|
||||
{{ us4_firewalld_backup_root }}/{{ us4_firewalld_backup_timestamp.stdout }}/rollback-phase1.sh
|
||||
|
||||
- name: Create the protected backup directory
|
||||
ansible.builtin.file:
|
||||
path: "{{ us4_firewalld_backup_dir }}"
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0700"
|
||||
|
||||
- name: Back up the complete firewalld configuration
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- tar
|
||||
- --create
|
||||
- --gzip
|
||||
- "--file={{ us4_firewalld_backup_dir }}/firewalld.tgz"
|
||||
- --directory=/etc
|
||||
- firewalld
|
||||
changed_when: true
|
||||
|
||||
- name: Capture the pre-change runtime ruleset
|
||||
ansible.builtin.shell:
|
||||
cmd: >-
|
||||
umask 077 && nft list ruleset >
|
||||
{{ us4_firewalld_backup_dir | quote }}/nft-ruleset.txt
|
||||
executable: /bin/bash
|
||||
changed_when: true
|
||||
|
||||
- name: Install the exact pre-change rollback script
|
||||
ansible.builtin.copy:
|
||||
dest: "{{ us4_firewalld_rollback_command }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0700"
|
||||
content: |
|
||||
#!/bin/sh
|
||||
set -eu
|
||||
exec >>/var/log/us4-firewalld-phase1-rollback.log 2>&1
|
||||
printf '%s rollback start\n' "$(date -Is)"
|
||||
add_service() {
|
||||
service=$1
|
||||
/usr/bin/firewall-cmd --permanent --zone=public \
|
||||
--query-service="$service" >/dev/null 2>&1 ||
|
||||
/usr/bin/firewall-cmd --permanent --zone=public \
|
||||
--add-service="$service"
|
||||
/usr/bin/firewall-cmd --zone=public \
|
||||
--query-service="$service" >/dev/null 2>&1 ||
|
||||
/usr/bin/firewall-cmd --zone=public --add-service="$service"
|
||||
}
|
||||
add_port() {
|
||||
port=$1
|
||||
/usr/bin/firewall-cmd --permanent --zone=public \
|
||||
--query-port="$port" >/dev/null 2>&1 ||
|
||||
/usr/bin/firewall-cmd --permanent --zone=public \
|
||||
--add-port="$port"
|
||||
/usr/bin/firewall-cmd --zone=public \
|
||||
--query-port="$port" >/dev/null 2>&1 ||
|
||||
/usr/bin/firewall-cmd --zone=public --add-port="$port"
|
||||
}
|
||||
{% for service in us4_firewalld_cleanup_services %}
|
||||
add_service {{ service }}
|
||||
{% endfor %}
|
||||
{% for port in us4_firewalld_cleanup_ports %}
|
||||
add_port {{ port }}
|
||||
{% endfor %}
|
||||
/usr/bin/firewall-cmd --check-config
|
||||
printf '%s rollback complete\n' "$(date -Is)"
|
||||
|
||||
- name: Schedule the 15-minute automatic rollback
|
||||
ansible.builtin.shell:
|
||||
cmd: |
|
||||
set -euo pipefail
|
||||
output=$(printf '%s\n' {{ us4_firewalld_rollback_command | quote }} | at now + 15 minutes 2>&1)
|
||||
job_id=$(printf '%s\n' "$output" | sed -n 's/^job \([0-9][0-9]*\).*/\1/p')
|
||||
test -n "$job_id"
|
||||
printf '%s\n' "$job_id"
|
||||
executable: /bin/bash
|
||||
changed_when: true
|
||||
register: us4_firewalld_rollback_job
|
||||
|
||||
- name: Record the automatic rollback job
|
||||
ansible.builtin.set_fact:
|
||||
us4_firewalld_rollback_job_id: "{{ us4_firewalld_rollback_job.stdout }}"
|
||||
us4_firewalld_rollback_cancelled: false
|
||||
|
||||
- name: Persist the rollback job ID beside the backup
|
||||
ansible.builtin.copy:
|
||||
dest: "{{ us4_firewalld_backup_dir }}/phase1-at-job-id"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0600"
|
||||
content: "{{ us4_firewalld_rollback_job_id }}\n"
|
||||
|
||||
- name: Reconcile and verify the audited public zone
|
||||
block:
|
||||
- name: Remove audited stale firewalld services
|
||||
ansible.posix.firewalld:
|
||||
zone: "{{ us4_firewalld_zone }}"
|
||||
service: "{{ item }}"
|
||||
state: disabled
|
||||
permanent: true
|
||||
immediate: true
|
||||
loop: "{{ us4_firewalld_stale_services }}"
|
||||
|
||||
- name: Remove audited stale firewalld ports
|
||||
ansible.posix.firewalld:
|
||||
zone: "{{ us4_firewalld_zone }}"
|
||||
port: "{{ item }}"
|
||||
state: disabled
|
||||
permanent: true
|
||||
immediate: true
|
||||
loop: "{{ us4_firewalld_stale_ports }}"
|
||||
|
||||
- name: Verify the permanent firewalld configuration
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --check-config]
|
||||
when: not ansible_check_mode
|
||||
changed_when: false
|
||||
|
||||
- name: Read reconciled runtime services
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --zone=public, --list-services]
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
register: us4_firewalld_after_runtime_services
|
||||
|
||||
- name: Read reconciled permanent services
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --permanent, --zone=public, --list-services]
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
register: us4_firewalld_after_permanent_services
|
||||
|
||||
- name: Read reconciled runtime ports
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --zone=public, --list-ports]
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
register: us4_firewalld_after_runtime_ports
|
||||
|
||||
- name: Read reconciled permanent ports
|
||||
ansible.builtin.command:
|
||||
argv: [firewall-cmd, --permanent, --zone=public, --list-ports]
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
register: us4_firewalld_after_permanent_ports
|
||||
|
||||
- name: Require the exact audited post-change public zone
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- us4_firewalld_after_runtime_services.stdout.split() | sort == us4_firewalld_keep_services | sort
|
||||
- us4_firewalld_after_permanent_services.stdout.split() | sort == us4_firewalld_keep_services | sort
|
||||
- us4_firewalld_after_runtime_ports.stdout.split() | length == 0
|
||||
- us4_firewalld_after_permanent_ports.stdout.split() | length == 0
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Verify a fresh independent SSH and sudo path
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- ssh
|
||||
- -4
|
||||
- -o
|
||||
- BatchMode=yes
|
||||
- -o
|
||||
- ConnectTimeout=10
|
||||
- -o
|
||||
- ControlMaster=no
|
||||
- -o
|
||||
- ControlPath=none
|
||||
- windy@us4.wsvc.info
|
||||
- sudo -n true
|
||||
delegate_to: localhost
|
||||
become: false
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
vars:
|
||||
ansible_become: false
|
||||
|
||||
- name: Verify public HTTPS routes
|
||||
ansible.builtin.uri:
|
||||
url: "{{ item.url }}"
|
||||
follow_redirects: all
|
||||
status_code: "{{ item.status }}"
|
||||
validate_certs: true
|
||||
use_proxy: false
|
||||
loop:
|
||||
- {url: https://update.wsvc.info/, status: 200}
|
||||
- {url: https://us4-gate.wsvc.info/, status: 401}
|
||||
- {url: https://trlm.wsvc.info/, status: 200}
|
||||
delegate_to: localhost
|
||||
become: false
|
||||
when: not ansible_check_mode
|
||||
vars:
|
||||
ansible_become: false
|
||||
|
||||
- name: Verify the secondary MX TCP listener externally
|
||||
ansible.builtin.wait_for:
|
||||
host: "{{ ansible_host_ipv4 }}"
|
||||
port: 25
|
||||
state: started
|
||||
connect_timeout: 5
|
||||
timeout: 10
|
||||
delegate_to: localhost
|
||||
become: false
|
||||
when: not ansible_check_mode
|
||||
vars:
|
||||
ansible_become: false
|
||||
|
||||
- name: Verify all expected containers are running
|
||||
ansible.builtin.command:
|
||||
argv: [docker, ps, --format, "{{ '{{.Names}}' }}"]
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
register: us4_firewalld_running_containers
|
||||
|
||||
- name: Reject missing application containers
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- us4_firewalld_expected_containers | difference(us4_firewalld_running_containers.stdout_lines) | length == 0
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Verify Fail2ban remains active
|
||||
ansible.builtin.command:
|
||||
argv: [fail2ban-client, status]
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
register: us4_firewalld_fail2ban
|
||||
|
||||
- name: Require all audited Fail2ban jails
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- item in us4_firewalld_fail2ban.stdout
|
||||
loop:
|
||||
- postfix-postscreen
|
||||
- postfix-sasl
|
||||
- recidive
|
||||
- sshd
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Verify the deployed WireGuard health check
|
||||
ansible.builtin.command:
|
||||
argv: [/usr/local/lib/vps-health/run]
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
register: us4_firewalld_wireguard_health
|
||||
|
||||
- name: Cancel automatic rollback only after all checks pass
|
||||
ansible.builtin.command:
|
||||
argv: [at, -r, "{{ us4_firewalld_rollback_job_id }}"]
|
||||
changed_when: true
|
||||
when:
|
||||
- not ansible_check_mode
|
||||
- us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0
|
||||
|
||||
- name: Mark the automatic rollback as cancelled
|
||||
ansible.builtin.set_fact:
|
||||
us4_firewalld_rollback_cancelled: true
|
||||
when:
|
||||
- not ansible_check_mode
|
||||
- us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0
|
||||
|
||||
rescue:
|
||||
- name: Preserve the automatic rollback and stop
|
||||
ansible.builtin.fail:
|
||||
msg: >-
|
||||
A reconciliation or verification task failed. No reload was
|
||||
attempted. If cleanup was required, its automatic rollback remains
|
||||
scheduled; do not remove it manually.
|
||||
|
||||
always:
|
||||
- name: Report backup and rollback disposition
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
backup: >-
|
||||
{{ us4_firewalld_backup_dir |
|
||||
default('not-created-in-check-mode' if ansible_check_mode else 'not-required') }}
|
||||
automatic_rollback: >-
|
||||
{{ 'not-created-in-check-mode' if ansible_check_mode else
|
||||
('cancelled-after-success' if (us4_firewalld_rollback_cancelled | default(false)) else
|
||||
'scheduled-or-executed') if
|
||||
(us4_firewalld_cleanup_services | length > 0 or us4_firewalld_cleanup_ports | length > 0)
|
||||
else 'not-required' }}
|
||||
@@ -0,0 +1,5 @@
|
||||
---
|
||||
collections:
|
||||
# us4-firewalld.yml was source-reviewed and exercised with this version.
|
||||
- name: ansible.posix
|
||||
version: 2.2.2
|
||||
@@ -0,0 +1,8 @@
|
||||
---
|
||||
# Allowlist of compose/ projects this playbook may deploy. A host may only
|
||||
# reference a project listed here (see tasks: "Require a repo compose project").
|
||||
compose_repo_projects:
|
||||
- vaultwarden
|
||||
- pdns
|
||||
- adguardhome
|
||||
- unifi
|
||||
@@ -0,0 +1,84 @@
|
||||
---
|
||||
# Deploy the repo-owned, sanitized Compose declaration to the host.
|
||||
#
|
||||
# Safety model:
|
||||
# - Only hosts with an inventory `compose_repo_project` (allowlisted) are valid.
|
||||
# - The repo file is staged to `<file>.dsh-new` and validated with
|
||||
# `docker compose config --quiet` against the server-local .env BEFORE it
|
||||
# replaces anything. A failed validation never touches the live file.
|
||||
# - The current file is kept as `*.bak-<timestamp>` before promotion.
|
||||
# - Apply mode requires `compose_deploy_confirm=true`; `--check` gives a
|
||||
# read-only diff + validation without writes.
|
||||
# - The playbook never writes, reads, or transfers the server .env.
|
||||
|
||||
- name: Require an allowlisted repo compose project for this host
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- compose_repo_project is defined
|
||||
- compose_repo_project in compose_repo_projects
|
||||
fail_msg: >-
|
||||
No allowlisted compose_repo_project for {{ inventory_hostname }}.
|
||||
Supported: {{ compose_repo_projects | join(', ') }}.
|
||||
|
||||
- name: Require explicit confirmation for apply mode
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- ansible_check_mode or (compose_deploy_confirm | bool)
|
||||
fail_msg: >-
|
||||
This playbook replaces the server compose file and may recreate
|
||||
containers. Run with --check for a read-only diff, or supply
|
||||
compose_deploy_confirm=true to apply.
|
||||
|
||||
- name: Stage the repo compose file next to the live one
|
||||
ansible.builtin.copy:
|
||||
src: "{{ playbook_dir }}/../../compose/{{ compose_repo_project }}/compose.yml"
|
||||
dest: "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
|
||||
mode: "0644"
|
||||
diff: true
|
||||
register: compose_stage
|
||||
|
||||
- name: Validate staged compose against the server .env (read-only)
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- docker
|
||||
- compose
|
||||
- -f
|
||||
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
|
||||
- --project-directory
|
||||
- "{{ compose_project_dir }}"
|
||||
- config
|
||||
- --quiet
|
||||
register: compose_validate
|
||||
changed_when: false
|
||||
failed_when: compose_validate.rc != 0
|
||||
|
||||
- name: Show staged-vs-live difference
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ compose_stage.diff | default('(no change)') }}"
|
||||
when: ansible_check_mode
|
||||
|
||||
- name: Back up the current compose file (apply mode)
|
||||
ansible.builtin.shell:
|
||||
cmd: >-
|
||||
cp -a '{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}'
|
||||
'{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.bak-$(date +%Y%m%d-%H%M%S)'
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Promote the validated compose file (apply mode)
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- mv
|
||||
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}.dsh-new"
|
||||
- "{{ compose_project_dir }}/{{ compose_remote_file | default('compose.yml') }}"
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Apply the compose declaration (apply mode)
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- docker
|
||||
- compose
|
||||
- --project-directory
|
||||
- "{{ compose_project_dir }}"
|
||||
- up
|
||||
- -d
|
||||
when: not ansible_check_mode
|
||||
@@ -7,9 +7,13 @@ healthcheck_timer_on_calendar: '*-*-* 06:15:00'
|
||||
healthcheck_timer_randomized_delay_sec: 15m
|
||||
healthcheck_backup_max_age_hours: 30
|
||||
healthcheck_tls_warn_days: 21
|
||||
healthcheck_profiles:
|
||||
# Map of profile name -> installed script filename. A host selects which
|
||||
# profiles it runs via the `healthcheck_profiles` list (inventory).
|
||||
healthcheck_profile_scripts:
|
||||
mailcow: mailcow.sh
|
||||
vaultwarden: vaultwarden.sh
|
||||
pdns: pdns.sh
|
||||
wireguard: wireguard.sh
|
||||
adguardhome: adguardhome.sh
|
||||
rustdesk: rustdesk.sh
|
||||
hk2aux: hk2aux.sh
|
||||
|
||||
@@ -1,9 +1,10 @@
|
||||
---
|
||||
- name: Validate known health-check profile
|
||||
- name: Validate known health-check profiles
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- healthcheck_profile in healthcheck_profiles
|
||||
fail_msg: "Unsupported healthcheck_profile: {{ healthcheck_profile }}"
|
||||
- item in healthcheck_profile_scripts
|
||||
fail_msg: "Unsupported healthcheck_profile: {{ item }}"
|
||||
loop: "{{ healthcheck_profiles }}"
|
||||
|
||||
- name: Install health-check directories
|
||||
ansible.builtin.file:
|
||||
@@ -28,13 +29,14 @@
|
||||
group: root
|
||||
mode: "0755"
|
||||
|
||||
- name: Install service health-check script
|
||||
- name: Install service health-check scripts
|
||||
ansible.builtin.template:
|
||||
src: "{{ healthcheck_profiles[healthcheck_profile] }}.j2"
|
||||
dest: "{{ healthcheck_install_root }}/{{ healthcheck_profiles[healthcheck_profile] }}"
|
||||
src: "{{ healthcheck_profile_scripts[item] }}.j2"
|
||||
dest: "{{ healthcheck_install_root }}/{{ healthcheck_profile_scripts[item] }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0755"
|
||||
loop: "{{ healthcheck_profiles }}"
|
||||
|
||||
- name: Install health-check dispatcher
|
||||
ansible.builtin.template:
|
||||
|
||||
@@ -7,7 +7,7 @@ set -uo pipefail
|
||||
RESULT_DIR='{{ healthcheck_state_dir }}'
|
||||
LOG_DIR='{{ healthcheck_log_dir }}'
|
||||
HOST_NAME="$(hostname -f 2>/dev/null || hostname)"
|
||||
CHECK_NAME='{{ healthcheck_profile }}'
|
||||
CHECK_NAME="$(basename "$0" .sh)"
|
||||
STATUS=ok
|
||||
EXIT_CODE=0
|
||||
DETAILS=()
|
||||
@@ -31,9 +31,27 @@ compose_ps() {
|
||||
}
|
||||
|
||||
check_compose() {
|
||||
local output
|
||||
output="$(compose_ps)" || { record critical 'compose_ps_failed'; return; }
|
||||
if grep -qiE 'Exited|Restarting|[[:space:]]Dead[[:space:]]' <<<"$output"; then
|
||||
local output services bad
|
||||
# Only flag containers of *active* services (config --services excludes
|
||||
# debug/profile-gated services such as vaultwarden's pgweb, which is
|
||||
# intentionally stopped unless started with --profile debug).
|
||||
services="$(docker compose --project-directory '{{ compose_project_dir }}' config --services 2>/dev/null)" || { record critical 'compose_ps_failed'; return; }
|
||||
output="$(docker compose --project-directory '{{ compose_project_dir }}' ps --all --format json 2>&1)" || { record critical 'compose_ps_failed'; return; }
|
||||
bad="$(printf '%s\n' "$output" | python3 -c '
|
||||
import json, sys
|
||||
services = set(sys.argv[1].split())
|
||||
for line in sys.stdin:
|
||||
line = line.strip()
|
||||
if not line:
|
||||
continue
|
||||
try:
|
||||
c = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
if c.get("Service") in services and c.get("State") in ("exited", "restarting", "dead"):
|
||||
print(c.get("Service"))
|
||||
' "$services")"
|
||||
if [[ -n "$bad" ]]; then
|
||||
record critical 'compose_unhealthy_container'
|
||||
else
|
||||
record ok 'compose_ok'
|
||||
@@ -77,8 +95,11 @@ check_tls_days() {
|
||||
}
|
||||
|
||||
emit_result() {
|
||||
local tmp detail_json
|
||||
tmp="$(mktemp "${RESULT_DIR}/latest.json.XXXXXX")"
|
||||
# Per-check JSON at latest-<check>.json. The dispatcher merges these into
|
||||
# latest.json so multiple profiles on one host do not overwrite each other.
|
||||
local tmp path detail_json
|
||||
path="${RESULT_DIR}/latest-${CHECK_NAME}.json"
|
||||
tmp="$(mktemp "${RESULT_DIR}/.latest-${CHECK_NAME}.XXXXXX")"
|
||||
detail_json="$(printf '%s\n' "${DETAILS[@]:-unknown:no_details}" | python3 -c 'import json,sys; print(json.dumps([line.rstrip() for line in sys.stdin if line.strip()]))')"
|
||||
python3 - "$tmp" "$HOST_NAME" "$CHECK_NAME" "$STATUS" "$EXIT_CODE" "$detail_json" <<'PY'
|
||||
import json, sys
|
||||
@@ -90,6 +111,75 @@ with open(path, 'w', encoding='utf-8') as f:
|
||||
f.write('\n')
|
||||
PY
|
||||
chmod 0640 "$tmp"
|
||||
mv "$tmp" "${RESULT_DIR}/latest.json"
|
||||
mv "$tmp" "$path"
|
||||
cat "$path"
|
||||
}
|
||||
|
||||
aggregate_result() {
|
||||
# Merge the just-run per-check files into latest.json. With a single complete
|
||||
# check this is a verbatim copy, preserving the historical one-object shape.
|
||||
# With several checks it emits one object whose status is the worst of all
|
||||
# checks; each check's own status/details are retained under `checks`. An
|
||||
# expected check with no fresh result file (profile crashed before writing)
|
||||
# is aggregated as `unknown`, so latest.json can never go stale while the
|
||||
# dispatcher reports a failure.
|
||||
case "$#" in
|
||||
0) return 0 ;;
|
||||
1) if [[ -f "${RESULT_DIR}/latest-$1.json" ]]; then
|
||||
cp -f "${RESULT_DIR}/latest-$1.json" "${RESULT_DIR}/latest.json"
|
||||
else
|
||||
python3 - "$RESULT_DIR" "$HOST_NAME" "$@" <<'PY'
|
||||
import json, os, sys
|
||||
rdir, host = sys.argv[1], sys.argv[2]
|
||||
checks = sys.argv[3:]
|
||||
levels = {'ok': 0, 'warning': 1, 'unknown': 2, 'critical': 3}
|
||||
worst, worst_code = 'ok', 0
|
||||
items = []
|
||||
for c in checks:
|
||||
p = os.path.join(rdir, 'latest-%s.json' % c)
|
||||
if os.path.exists(p):
|
||||
d = json.load(open(p))
|
||||
st, code = d['status'], d['exit_code']
|
||||
items.append({'check': d['check'], 'status': st,
|
||||
'exit_code': code, 'details': d['details']})
|
||||
else:
|
||||
st, code = 'unknown', 3
|
||||
items.append({'check': c, 'status': st, 'exit_code': code,
|
||||
'details': ['unknown:check_did_not_complete']})
|
||||
if levels[st] > levels[worst]:
|
||||
worst, worst_code = st, code
|
||||
out = {'schema': 1, 'host': host, 'check': 'aggregate', 'status': worst,
|
||||
'exit_code': worst_code, 'checks': items}
|
||||
open(os.path.join(rdir, 'latest.json'), 'w').write(
|
||||
json.dumps(out, sort_keys=True, separators=(',', ':')) + '\n')
|
||||
PY
|
||||
fi ;;
|
||||
*) python3 - "$RESULT_DIR" "$HOST_NAME" "$@" <<'PY'
|
||||
import json, os, sys
|
||||
rdir, host = sys.argv[1], sys.argv[2]
|
||||
checks = sys.argv[3:]
|
||||
levels = {'ok': 0, 'warning': 1, 'unknown': 2, 'critical': 3}
|
||||
worst, worst_code = 'ok', 0
|
||||
items = []
|
||||
for c in checks:
|
||||
p = os.path.join(rdir, 'latest-%s.json' % c)
|
||||
if os.path.exists(p):
|
||||
d = json.load(open(p))
|
||||
st, code = d['status'], d['exit_code']
|
||||
items.append({'check': d['check'], 'status': st,
|
||||
'exit_code': code, 'details': d['details']})
|
||||
else:
|
||||
st, code = 'unknown', 3
|
||||
items.append({'check': c, 'status': st, 'exit_code': code,
|
||||
'details': ['unknown:check_did_not_complete']})
|
||||
if levels[st] > levels[worst]:
|
||||
worst, worst_code = st, code
|
||||
out = {'schema': 1, 'host': host, 'check': 'aggregate', 'status': worst,
|
||||
'exit_code': worst_code, 'checks': items}
|
||||
open(os.path.join(rdir, 'latest.json'), 'w').write(
|
||||
json.dumps(out, sort_keys=True, separators=(',', ':')) + '\n')
|
||||
PY
|
||||
esac
|
||||
chmod 0640 "${RESULT_DIR}/latest.json"
|
||||
cat "${RESULT_DIR}/latest.json"
|
||||
}
|
||||
|
||||
@@ -1,4 +1,24 @@
|
||||
#!/usr/bin/env bash
|
||||
set -o pipefail
|
||||
'{{ healthcheck_install_root }}/{{ healthcheck_profiles[healthcheck_profile] }}' 2>&1 | tee -a '{{ healthcheck_log_dir }}/healthcheck.log'
|
||||
exit "${PIPESTATUS[0]}"
|
||||
source '{{ healthcheck_install_root }}/health-common.sh'
|
||||
# Run every enabled health-check profile, exit with the worst (max) code, and
|
||||
# merge the per-check results into /var/lib/vps-health/latest.json.
|
||||
rc=0
|
||||
# Drop per-check results from any prior run so a profile that crashes before
|
||||
# reporting cannot leak a stale healthy result into the aggregate.
|
||||
{% for profile in healthcheck_profiles %}
|
||||
rm -f '{{ healthcheck_state_dir }}/latest-{{ healthcheck_profile_scripts[profile] | replace('.sh', '') }}.json'
|
||||
{% endfor %}
|
||||
{% for profile in healthcheck_profiles %}
|
||||
'{{ healthcheck_install_root }}/{{ healthcheck_profile_scripts[profile] }}' 2>&1 | tee -a '{{ healthcheck_log_dir }}/healthcheck.log'
|
||||
this_rc="${PIPESTATUS[0]}"
|
||||
[ "$this_rc" -gt "$rc" ] && rc="$this_rc"
|
||||
{% endfor %}
|
||||
# Collect profile check names line-by-line (robust against Jinja trim_blocks
|
||||
# whitespace control, which would otherwise merge this into one line).
|
||||
aggregate_args=""
|
||||
{% for profile in healthcheck_profiles %}
|
||||
aggregate_args="$aggregate_args {{ healthcheck_profile_scripts[profile] | replace('.sh', '') }}"
|
||||
{% endfor %}
|
||||
aggregate_result $aggregate_args
|
||||
exit "$rc"
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
#!/usr/bin/env bash
|
||||
set -uo pipefail
|
||||
source '{{ healthcheck_install_root }}/health-common.sh'
|
||||
|
||||
require_command docker
|
||||
require_command ss
|
||||
|
||||
# Auxiliary services co-located on hk2.chans.xyz (separate compose projects
|
||||
# under /opt, fronted by Traefik). Verified live 2026-08-12.
|
||||
|
||||
# traefik
|
||||
if docker inspect traefik >/dev/null 2>&1; then
|
||||
[[ "$(docker inspect traefik --format '{{ '{{' }}.State.Running{{ '}}' }}' 2>/dev/null)" == true ]] \
|
||||
&& record ok 'traefik_running' || record critical 'traefik_not_running'
|
||||
else
|
||||
record critical 'traefik_container_missing'
|
||||
fi
|
||||
|
||||
# adguardhome (hk2 variant: DoH 5443, DoT 853)
|
||||
if docker inspect adguardhome >/dev/null 2>&1; then
|
||||
[[ "$(docker inspect adguardhome --format '{{ '{{' }}.State.Running{{ '}}' }}' 2>/dev/null)" == true ]] \
|
||||
&& record ok 'adguard_running' || record critical 'adguard_not_running'
|
||||
else
|
||||
record critical 'adguard_container_missing'
|
||||
fi
|
||||
|
||||
# remark42
|
||||
if docker inspect remark42 >/dev/null 2>&1; then
|
||||
[[ "$(docker inspect remark42 --format '{{ '{{' }}.State.Running{{ '}}' }}' 2>/dev/null)" == true ]] \
|
||||
&& record ok 'remark42_running' || record critical 'remark42_not_running'
|
||||
else
|
||||
record critical 'remark42_container_missing'
|
||||
fi
|
||||
|
||||
# nginx-manager was removed 2026-08-12 (leftover config, never running).
|
||||
# Warn if a container by that name ever reappears.
|
||||
if docker inspect nginx-manager >/dev/null 2>&1; then
|
||||
record warning 'nginx_manager_unexpectedly_running'
|
||||
else
|
||||
record ok 'nginx_manager_not_running'
|
||||
fi
|
||||
|
||||
# Listening ports (Traefik 80/443/8080, AdGuard DoH 5443 / DoT 853).
|
||||
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)80$' \
|
||||
&& record ok 'traefik_http_80' || record critical 'traefik_http_80_missing'
|
||||
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)443$' \
|
||||
&& record ok 'traefik_https_443' || record critical 'traefik_https_443_missing'
|
||||
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)8080$' \
|
||||
&& record ok 'traefik_dashboard_8080' || record warning 'traefik_dashboard_8080_missing'
|
||||
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)5443$' \
|
||||
&& record ok 'adguard_doh_5443' || record critical 'adguard_doh_5443_missing'
|
||||
ss -H -ltn 2>/dev/null | awk '{print $4}' | grep -Eq '(^|:)853$' \
|
||||
&& record ok 'adguard_dot_853' || record critical 'adguard_dot_853_missing'
|
||||
|
||||
emit_result
|
||||
exit "$EXIT_CODE"
|
||||
@@ -0,0 +1,25 @@
|
||||
#!/usr/bin/env bash
|
||||
set -uo pipefail
|
||||
source '{{ healthcheck_install_root }}/health-common.sh'
|
||||
|
||||
require_command docker
|
||||
require_command dig
|
||||
|
||||
# hbbs / hbbr must both be running (separate compose project at /opt/rustdesk).
|
||||
output="$(docker compose --project-directory /opt/rustdesk ps --all 2>&1)"
|
||||
if grep -qiE 'Exited|Restarting|[[:space:]]Dead[[:space:]]' <<<"$output"; then
|
||||
record critical 'rustdesk_unhealthy_container'
|
||||
else
|
||||
record ok 'rustdesk_compose_ok'
|
||||
fi
|
||||
|
||||
# hbbs must advertise the relay hostname that resolves to this host's public IP.
|
||||
cmd="$(docker inspect hbbs --format '{{ '{{' }}json .Config.Cmd{{ '}}' }}' 2>/dev/null)" || record critical 'rustdesk_hbbs_missing'
|
||||
grep -q 'hk2.chans.xyz:21117' <<<"$cmd" || record critical 'rustdesk_relay_misconfigured'
|
||||
|
||||
# The advertised relay hostname must resolve to this host's public IP.
|
||||
resolved="$(dig +short hk2.chans.xyz A 2>/dev/null)"
|
||||
grep -q '154.36.174.161' <<<"$resolved" || record critical 'rustdesk_relay_dns_missing'
|
||||
|
||||
emit_result
|
||||
exit "$EXIT_CODE"
|
||||
@@ -15,11 +15,12 @@ grep -Fq 'vw-db' <<<"$health" || record critical 'postgres_missing'
|
||||
check_https 'https://auth.wsvc.info/' '^200$'
|
||||
check_tls_days auth.wsvc.info 443
|
||||
|
||||
# Read effective config only inside the service and report booleans/fingerprints,
|
||||
# never its SMTP password or other secret fields.
|
||||
smtp_result="$(docker compose --project-directory '{{ compose_project_dir }}' exec -T vaultwarden python3 - <<'PY' 2>&1
|
||||
# Read effective config from the mounted vw-data dir on the host and run the
|
||||
# SMTP AUTH probe from the host (the vaultwarden image has no python3; the
|
||||
# host does). Never print the SMTP password.
|
||||
smtp_result="$(python3 - <<'PY' 2>&1
|
||||
import json, pathlib, smtplib, ssl
|
||||
cfg=json.loads(pathlib.Path('/data/config.json').read_text())
|
||||
cfg=json.loads(pathlib.Path('{{ compose_project_dir }}/vw-data/config.json').read_text())
|
||||
host=cfg.get('smtp_host'); port=int(cfg.get('smtp_port') or 0)
|
||||
user=cfg.get('smtp_username')
|
||||
smtp_secret=cfg.get('smtp_password')
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
---
|
||||
restic_enabled: false
|
||||
# Approved Restic source profile (a key of restic_sources) for the host. Set
|
||||
# per-host in inventory; the role fails if it is not an approved source.
|
||||
restic_backup_profile: ""
|
||||
restic_binary: /usr/bin/restic
|
||||
restic_config_path: /etc/vps-restic/repository.env
|
||||
restic_state_dir: /var/lib/vps-restic
|
||||
|
||||
@@ -10,8 +10,8 @@
|
||||
- name: Validate supported Restic source profile
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- healthcheck_profile in restic_sources
|
||||
fail_msg: "No approved Restic source profile for {{ healthcheck_profile }}."
|
||||
- restic_backup_profile in restic_sources
|
||||
fail_msg: "No approved Restic source profile for {{ restic_backup_profile }}."
|
||||
|
||||
- name: Verify Restic binary exists on target
|
||||
ansible.builtin.stat:
|
||||
|
||||
@@ -3,4 +3,4 @@ set -euo pipefail
|
||||
# Repository and password credentials are host-local in {{ restic_config_path }}.
|
||||
# shellcheck source=/dev/null
|
||||
source '{{ restic_config_path }}'
|
||||
exec '{{ restic_binary }}' backup --tag '{{ healthcheck_profile }}' --tag "$(hostname -s)" {% for source in restic_sources[healthcheck_profile] %}{{ source | quote }} {% endfor %}
|
||||
exec '{{ restic_binary }}' backup --tag '{{ restic_backup_profile }}' --tag "$(hostname -s)" {% for source in restic_sources[restic_backup_profile] %}{{ source | quote }} {% endfor %}
|
||||
|
||||
@@ -2,4 +2,4 @@
|
||||
set -euo pipefail
|
||||
# shellcheck source=/dev/null
|
||||
source '{{ restic_config_path }}'
|
||||
exec '{{ restic_binary }}' forget --prune --keep-daily {{ restic_keep_daily }} --keep-weekly {{ restic_keep_weekly }} --keep-monthly {{ restic_keep_monthly }} --tag '{{ healthcheck_profile }}'
|
||||
exec '{{ restic_binary }}' forget --prune --keep-daily {{ restic_keep_daily }} --keep-weekly {{ restic_keep_weekly }} --keep-monthly {{ restic_keep_monthly }} --tag '{{ restic_backup_profile }}'
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
[Unit]
|
||||
Description=Restic backup for approved {{ healthcheck_profile }} sources
|
||||
Description=Restic backup for approved {{ restic_backup_profile }} sources
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
# RustDesk server deployment (hbbs + hbbr) on hk2.
|
||||
# Safe by default: without rustdesk_confirm=true the role only reports whether
|
||||
# the declared compose file matches live state and refuses to recreate the stack.
|
||||
rustdesk_confirm: false
|
||||
# Compose project directory.
|
||||
rustdesk_compose_dir: /opt/rustdesk
|
||||
# Relay (hbbr) hostname:port advertised to every client via `hbbs -r`.
|
||||
# MUST resolve to this host's public IP (154.36.174.161). The known-bad value
|
||||
# 'hk2.wsvc.info' has no DNS record and must never be used.
|
||||
rustdesk_relay: hk2.chans.xyz:21117
|
||||
# Pinned server image (used for both hbbs and hbbr).
|
||||
rustdesk_image: rustdesk/rustdesk-server:1.1.14
|
||||
@@ -0,0 +1,85 @@
|
||||
---
|
||||
# Deploy/reconcile the self-hosted RustDesk server (hbbs + hbbr).
|
||||
# Idempotent: deploys the declared compose file; only recreates the stack with
|
||||
# explicit confirmation.
|
||||
|
||||
- name: Validate relay address is set and not the known-bad value
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- rustdesk_relay | length > 0
|
||||
- "'hk2.wsvc.info' not in rustdesk_relay"
|
||||
fail_msg: >-
|
||||
rustdesk_relay must be a resolvable relay address. The known-bad
|
||||
'hk2.wsvc.info' has no DNS record and must not be used.
|
||||
|
||||
- name: Ensure compose project directory exists
|
||||
ansible.builtin.file:
|
||||
path: "{{ rustdesk_compose_dir }}"
|
||||
state: directory
|
||||
owner: windy
|
||||
group: root
|
||||
mode: "0755"
|
||||
|
||||
- name: Deploy compose file
|
||||
ansible.builtin.template:
|
||||
src: compose.yml.j2
|
||||
dest: "{{ rustdesk_compose_dir }}/compose.yml"
|
||||
owner: windy
|
||||
group: windy
|
||||
mode: "0644"
|
||||
register: rustdesk_compose_deployed
|
||||
|
||||
- name: Report no change needed
|
||||
ansible.builtin.debug:
|
||||
msg: "compose.yml already matches declared state; no change needed."
|
||||
when: not rustdesk_compose_deployed.changed
|
||||
|
||||
- name: Refuse to recreate without explicit confirmation
|
||||
ansible.builtin.fail:
|
||||
msg: >-
|
||||
compose.yml differs from declared state but rustdesk_confirm is not true.
|
||||
Supply rustdesk_confirm=true to deploy the file and recreate the stack.
|
||||
when:
|
||||
- rustdesk_compose_deployed.changed
|
||||
- not (rustdesk_confirm | bool)
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Apply compose stack
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- docker
|
||||
- compose
|
||||
- --project-directory
|
||||
- "{{ rustdesk_compose_dir }}"
|
||||
- up
|
||||
- -d
|
||||
when:
|
||||
- rustdesk_compose_deployed.changed
|
||||
- rustdesk_confirm | bool
|
||||
changed_when: true
|
||||
register: rustdesk_apply
|
||||
|
||||
- name: Verify hbbs relay command
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- docker
|
||||
- inspect
|
||||
- hbbs
|
||||
- --format
|
||||
- '{{ "{{" }}json .Config.Cmd{{ "}}" }}'
|
||||
register: rustdesk_hbbs_cmd
|
||||
changed_when: false
|
||||
when:
|
||||
- rustdesk_compose_deployed.changed
|
||||
- rustdesk_confirm | bool
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Assert hbbs advertises the declared relay
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- "'{{ rustdesk_relay }}' in rustdesk_hbbs_cmd.stdout"
|
||||
fail_msg: "hbbs is not advertising the declared relay {{ rustdesk_relay }}."
|
||||
when:
|
||||
- rustdesk_compose_deployed.changed
|
||||
- rustdesk_confirm | bool
|
||||
- not ansible_check_mode
|
||||
@@ -0,0 +1,34 @@
|
||||
networks:
|
||||
rustdesk-net:
|
||||
external: false
|
||||
|
||||
services:
|
||||
hbbs:
|
||||
container_name: hbbs
|
||||
ports:
|
||||
- 21115:21115
|
||||
- 21116:21116
|
||||
- 21116:21116/udp
|
||||
- 21118:21118
|
||||
image: {{ rustdesk_image }}
|
||||
command: "hbbs -r {{ rustdesk_relay }}"
|
||||
volumes:
|
||||
- ./hbbs:/root
|
||||
networks:
|
||||
- rustdesk-net
|
||||
depends_on:
|
||||
- hbbr
|
||||
restart: unless-stopped
|
||||
|
||||
hbbr:
|
||||
container_name: hbbr
|
||||
ports:
|
||||
- 21117:21117
|
||||
- 21119:21119
|
||||
image: {{ rustdesk_image }}
|
||||
command: hbbr
|
||||
volumes:
|
||||
- ./hbbr:/root
|
||||
networks:
|
||||
- rustdesk-net
|
||||
restart: unless-stopped
|
||||
@@ -0,0 +1,51 @@
|
||||
# compose/ — repo-owned Compose declarations
|
||||
|
||||
Non-secret Compose sources for the Docker hosts. Secrets are **never** in these
|
||||
files: every secret is a `${VAR}` reference resolved from the **server-local
|
||||
`.env`** (docker compose reads `.env` from the project directory automatically).
|
||||
|
||||
## Source-of-truth matrix
|
||||
|
||||
| Project | Host | Compose source | Mechanism |
|
||||
|---------|------|----------------|-----------|
|
||||
| `vaultwarden` | us2 (`/opt/vaultwarden`) | `compose/vaultwarden/compose.yml` | static file + `compose-deploy.yml` |
|
||||
| `pdns` | hk2 (`/opt/pdns`) | `compose/pdns/compose.yml` | static file + `compose-deploy.yml` |
|
||||
| `pgdb` | pgdb (`/opt/database`, 无 ansible) | `compose/pgdb/compose.yml` | static file(手动部署:scp → `docker compose config -q` → `up -d`;服务器文件名 `docker-compose.yml`) |
|
||||
| `soft-serve` | us2 (`/opt/soft-serve`, 已退役停用) | `compose/soft-serve/compose.yml` (+ `Dockerfile.backup`, `scripts/`) | static file(参考镜像; 2026-09-18 被 gitea 替换 VPS-94, 数据保留作回滚) |
|
||||
| `gitea` | us2 (`/opt/gitea`) | `compose/gitea/compose.yml` (+ `Dockerfile.backup`, `scripts/`) | static file(参考镜像, 未接入 compose-deploy; 服务器文件为准; 2026-09-18 替换 soft-serve, VPS-94) |
|
||||
| `adguardhome` | dns.windy.lan (`/opt/adguardhome`) | — (待从 LAN 提取) | static file (pending) |
|
||||
| `unifi` | ubnt (`/home/windy/unifi-9`) | — (待从 LAN 提取) | static file (pending) |
|
||||
| `wireguard` | us4 (`/opt/wireguard`) | `ansible/templates/wireguard-compose.yml.j2` | role-rendered (inventory vars) |
|
||||
| `rustdesk` | hk2 (`/opt/rustdesk`) | `ansible/roles/rustdesk/templates/compose.yml.j2` | role-rendered (inventory vars) |
|
||||
| `mailcow` | mx2 (`/opt/mail`) | — (mailcow update generator owns it) | excluded by design |
|
||||
|
||||
Mechanism rule: **static** `compose/<project>/compose.yml` for declarations that
|
||||
do not vary per host; **role-rendered j2** for declarations driven by inventory
|
||||
vars (image pins, relay host). One mechanism per project; do not duplicate a
|
||||
project in both.
|
||||
|
||||
## Deploying a static project
|
||||
|
||||
```bash
|
||||
cd ansible
|
||||
|
||||
# Read-only diff + validation against the server .env (no writes)
|
||||
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden --check --diff
|
||||
|
||||
# Apply: stage repo file → validate `docker compose config -q` → backup current
|
||||
# file → promote → `docker compose up -d` (gated)
|
||||
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden \
|
||||
-e '{"compose_deploy_confirm": true}'
|
||||
```
|
||||
|
||||
See [`../runbooks/ansible-operations.md`](../runbooks/ansible-operations.md).
|
||||
|
||||
## Adding a project
|
||||
|
||||
1. Sanitize the live compose so every secret is `${VAR}` from `.env`
|
||||
(prefer `${VAR:?missing VAR}` for required keys).
|
||||
2. Commit `compose/<project>/compose.yml` + `.env.example` (key names only).
|
||||
3. Add `compose_repo_project` (+ `compose_remote_file` if not `compose.yml`) to
|
||||
the host in `ansible/inventory/hosts.yml`, and allowlist the project in
|
||||
`ansible/roles/compose_deploy/defaults/main.yml`.
|
||||
4. Verify with `--check --diff` (zero diff) then a gated apply.
|
||||
@@ -0,0 +1,4 @@
|
||||
# compose/gitea — 秘密一律走服务器本地 .env, 不入库
|
||||
# 迁移期一次性: Gitea 管理员生成的 token (mirror-migrate.sh 读取, 用后撤销)
|
||||
GITEA_MIGRATE_USER=
|
||||
GITEA_MIGRATE_TOKEN=
|
||||
@@ -0,0 +1,3 @@
|
||||
FROM alpine:3.20
|
||||
RUN apk add --no-cache sqlite rsync tzdata
|
||||
WORKDIR /scripts
|
||||
@@ -0,0 +1,59 @@
|
||||
# Gitea on us2 — reference compose (Plane VPS-94, 迁移完成 2026-09-18)
|
||||
# 参考镜像, 服务器 /opt/gitea 文件为准 (同 soft-serve 约定, 未接入 compose-deploy)
|
||||
# rootless 镜像: uid 1000 原生非 root; 数据 /var/lib/gitea (宿主 ./data), 配置 /etc/gitea (宿主 ./config)
|
||||
# SSH: 容器内监听 2322 (非特权, SSH_LISTEN_PORT), 对外 repo.windy.me:2222 经 Traefik TCP entrypoint `ssh`
|
||||
services:
|
||||
gitea:
|
||||
image: gitea/gitea@sha256:1c17ecaead42eb3b5391553d8708103a4beb0e86edf5b9ebc1eb269c318845f2 # 1.27.3-rootless
|
||||
container_name: gitea
|
||||
restart: unless-stopped
|
||||
user: "1000:1000"
|
||||
environment:
|
||||
TZ: Asia/Shanghai
|
||||
volumes:
|
||||
- ./data:/var/lib/gitea
|
||||
- ./config:/etc/gitea
|
||||
- ./secrets:/secrets:ro # 复用的 soft-serve host key (SSH_SERVER_HOST_KEYS)
|
||||
networks:
|
||||
- traefik
|
||||
labels:
|
||||
- traefik.enable=true
|
||||
# Web UI: repo.windy.me (2026-09-18 操作者决定复用现有域名, 免 DNS 变更)
|
||||
- traefik.http.routers.gitea-web.rule=Host(`repo.windy.me`)
|
||||
- traefik.http.routers.gitea-web.entrypoints=websecure
|
||||
- traefik.http.routers.gitea-web.tls.certresolver=letsencrypt
|
||||
- traefik.http.services.gitea-web.loadbalancer.server.port=3000
|
||||
# SSH: 接管 :2222 (entrypoint 已存在, router 动态生效, 无需重启 Traefik)
|
||||
- traefik.tcp.routers.gitea-ssh.entrypoints=ssh
|
||||
- traefik.tcp.routers.gitea-ssh.rule=HostSNI(`*`)
|
||||
- traefik.tcp.routers.gitea-ssh.tls=false
|
||||
- traefik.tcp.services.gitea-ssh.loadbalancer.server.port=2322
|
||||
|
||||
gitea-backup:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile.backup
|
||||
container_name: gitea-backup
|
||||
restart: unless-stopped
|
||||
volumes:
|
||||
- ./data:/data:ro
|
||||
- ./config:/config:ro
|
||||
- ./backups:/backup
|
||||
- ./scripts:/scripts
|
||||
environment:
|
||||
TZ: Asia/Shanghai
|
||||
BACKUP_UID: 1000
|
||||
BACKUP_GID: 1000
|
||||
entrypoint: >
|
||||
/bin/sh -ec "
|
||||
umask 077 &&
|
||||
touch /backup/backup.log &&
|
||||
crontab /scripts/crontab.txt &&
|
||||
echo '[INFO] gitea backup cron installed' &&
|
||||
crond -f -l 8
|
||||
"
|
||||
|
||||
networks:
|
||||
traefik:
|
||||
external: true
|
||||
name: vw-net
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
#!/bin/sh
|
||||
set -eu
|
||||
umask 077
|
||||
D() { date "+%Y-%m-%d %H:%M:%S"; }
|
||||
TS=$(date +%Y%m%d_%H%M%S)
|
||||
OUT="/backup/gitea_${TS}"
|
||||
mkdir -p "$OUT"
|
||||
echo "[$(D)] Starting gitea backup -> $OUT"
|
||||
# rootless 布局: app.ini=/etc/gitea(宿主 ./config), db+repos=/var/lib/gitea/data(宿主 ./data/data)
|
||||
# app.ini 含 SECRET_KEY/INTERNAL_TOKEN — 恢复 2FA/session/mirror 凭据必需
|
||||
tar czf "$OUT/app.ini.tar.gz" -C /config app.ini
|
||||
sqlite3 /data/data/gitea.db ".backup '$OUT/gitea.db'"
|
||||
rsync -a /data/data/git/repositories/ "$OUT/repos/"
|
||||
tar czf "$OUT/repos.tar.gz" -C "$OUT" repos
|
||||
rm -rf "$OUT/repos"
|
||||
chmod 600 "$OUT"/*
|
||||
if [ -n "${BACKUP_UID:-}" ] && [ -n "${BACKUP_GID:-}" ]; then
|
||||
chown -R "$BACKUP_UID:$BACKUP_GID" "$OUT" /backup/backup.log
|
||||
fi
|
||||
echo "[$(D)] Backup OK: $(du -sh "$OUT" | cut -f1)"
|
||||
@@ -0,0 +1,4 @@
|
||||
# Run gitea backup daily at 02:00
|
||||
0 2 * * * /bin/sh /scripts/backup.sh >> /backup/backup.log 2>&1
|
||||
# Prune backups older than 14 days daily at 03:00
|
||||
0 3 * * * /bin/sh /scripts/prune.sh >> /backup/backup.log 2>&1
|
||||
Executable
+41
@@ -0,0 +1,41 @@
|
||||
#!/bin/sh
|
||||
# 一次性迁移辅助 (Plane VPS-94 Phase 2): 在 gitea 容器内执行。
|
||||
# 已于 2026-09-18 执行完成 (16 仓), 留档备查; 复用时按 VPS-94 流程重生成一次性 token。
|
||||
# 用法:
|
||||
# GITEA_MIGRATE_USER=<user> GITEA_MIGRATE_TOKEN=<token> \
|
||||
# docker exec -e GITEA_MIGRATE_USER -e GITEA_MIGRATE_TOKEN gitea \
|
||||
# /scripts/mirror-migrate.sh [public_repo ...]
|
||||
# 每仓: API 建仓 (默认 private, 参数中列出的为 public) -> push --mirror。
|
||||
# default_branch 按源仓 symbolic-ref HEAD 设置, 避免非 main 源仓在 Gitea 显示为空。
|
||||
# 结束后按 VPS-94 Phase 3 逐仓核对 git ls-remote ref 全集。
|
||||
set -eu
|
||||
MUSER="${GITEA_MIGRATE_USER:?need GITEA_MIGRATE_USER}"
|
||||
TOKEN="${GITEA_MIGRATE_TOKEN:?need GITEA_MIGRATE_TOKEN}"
|
||||
SRC="/migration-src"
|
||||
API="http://localhost:3000/api/v1"
|
||||
PUBLIC_REPOS=" $* "
|
||||
|
||||
migrate_one() {
|
||||
dir="$1"
|
||||
git -C "$dir" rev-parse --git-dir >/dev/null 2>&1 || { echo "[SKIP] $dir (not a git repo)"; return 0; }
|
||||
name=$(basename "$dir"); name=${name%.git}
|
||||
def_branch=$(git -C "$dir" symbolic-ref --short HEAD)
|
||||
case "$PUBLIC_REPOS" in *" $name "*) private=false ;; *) private=true ;; esac
|
||||
echo "[MIGRATE] $name (default=$def_branch private=$private)"
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' -X POST "$API/user/repos" \
|
||||
-H "Authorization: token $TOKEN" -H "Content-Type: application/json" \
|
||||
-d "{\"name\":\"$name\",\"private\":$private,\"default_branch\":\"$def_branch\",\"auto_init\":false}")
|
||||
case "$code" in
|
||||
201) : ;;
|
||||
409) echo " [WARN] $name 已存在, 直接补推" ;;
|
||||
*) echo " [FAIL] create HTTP $code"; return 1 ;;
|
||||
esac
|
||||
git -C "$dir" push --mirror "http://$MUSER:$TOKEN@localhost:3000/$MUSER/$name.git"
|
||||
echo " [OK] $name pushed"
|
||||
}
|
||||
|
||||
for dir in "$SRC"/*.git "$SRC"/cdia; do
|
||||
[ -d "$dir" ] || continue
|
||||
migrate_one "$dir"
|
||||
done
|
||||
echo "[DONE] 全部处理完毕; 迁移后记得撤销一次性 token"
|
||||
Executable
+5
@@ -0,0 +1,5 @@
|
||||
#!/bin/sh
|
||||
set -eu
|
||||
D() { date "+%Y-%m-%d %H:%M:%S"; }
|
||||
ls -dt /backup/gitea_* 2>/dev/null | tail -n +15 | xargs -r rm -rf
|
||||
echo "[$(D)] Pruned. Kept $(ls -d /backup/gitea_* 2>/dev/null | wc -l) backups (max 14)"
|
||||
@@ -0,0 +1,39 @@
|
||||
# .env.example — PowerDNS stack (hk2.chans.xyz, /opt/pdns)
|
||||
#
|
||||
# Non-secret key reference ONLY. Real values live in the server-local .env
|
||||
# (never commit them). Compose requires the `:?`-marked keys to be present.
|
||||
|
||||
# Runtime
|
||||
TZ=Asia/Shanghai
|
||||
|
||||
# Postgres superuser (db + backup + pgweb)
|
||||
PGUSER=
|
||||
PGPASSWORD=
|
||||
DB_HOST=db
|
||||
DB_PORT=5432
|
||||
|
||||
# Application database (auth / poweradmin / backup)
|
||||
DB_NAME=pdns
|
||||
DB_USER=pdns
|
||||
DB_PASS=
|
||||
ADMIN_DB=pdnsadmin
|
||||
|
||||
# Backups
|
||||
CRON_SCHEDULE=0 3 * * *
|
||||
RETENTION_DAYS=7
|
||||
MAX_BACKUPS=7
|
||||
DUMP_ROLES=true
|
||||
|
||||
# PowerDNS auth API
|
||||
PDNS_API_KEY=
|
||||
|
||||
# Poweradmin (first-run admin + session)
|
||||
PA_SESSION_KEY=
|
||||
PA_ADMIN_USERNAME=
|
||||
PA_ADMIN_PASSWORD=
|
||||
PA_ADMIN_EMAIL=
|
||||
PA_ADMIN_FULLNAME=
|
||||
|
||||
# pgweb debug profile
|
||||
PGWEB_USER=
|
||||
PGWEB_PASS=
|
||||
@@ -0,0 +1,159 @@
|
||||
networks:
|
||||
frontend:
|
||||
name: traefik
|
||||
external: true
|
||||
|
||||
backend:
|
||||
internal: true
|
||||
|
||||
edge:
|
||||
|
||||
services:
|
||||
db:
|
||||
image: postgres:16
|
||||
container_name: pdns-db
|
||||
environment:
|
||||
POSTGRES_DB: postgres
|
||||
POSTGRES_USER: ${PGUSER:?missing PGUSER}
|
||||
POSTGRES_PASSWORD: ${PGPASSWORD:?missing PGPASSWORD}
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
PGTZ: ${TZ:-Asia/Shanghai}
|
||||
volumes:
|
||||
# Keep the existing mount path to avoid moving the current data directory.
|
||||
- dbdata:/var/lib/postgresql
|
||||
- ./db-init-generated:/docker-entrypoint-initdb.d:ro
|
||||
- ./backup:/backup:ro
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U \"$${POSTGRES_USER}\" -d \"$${POSTGRES_DB}\""]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 10
|
||||
restart: unless-stopped
|
||||
networks: [backend, edge]
|
||||
|
||||
auth:
|
||||
image: powerdns/pdns-auth-50:5.0.6
|
||||
container_name: pdns-auth
|
||||
depends_on:
|
||||
db:
|
||||
condition: service_healthy
|
||||
ports:
|
||||
- "53:53/udp"
|
||||
- "53:53/tcp"
|
||||
- "127.0.0.1:8081:8081"
|
||||
environment:
|
||||
PDNS_API_KEY: ${PDNS_API_KEY:?missing PDNS_API_KEY}
|
||||
DB_NAME: ${DB_NAME:?missing DB_NAME}
|
||||
DB_USER: ${DB_USER:?missing DB_USER}
|
||||
DB_PASS: ${DB_PASS:?missing DB_PASS}
|
||||
TEMPLATE_FILES: secrets
|
||||
volumes:
|
||||
- ./auth/pdns.conf:/etc/powerdns/pdns.conf:ro
|
||||
- ./auth/templates.d:/etc/powerdns/templates.d:ro
|
||||
- ./auth/keys:/var/lib/powerdns
|
||||
- ./auth/import:/import
|
||||
- ./auth/export:/export
|
||||
- ./auth/logs:/var/log/pdns
|
||||
healthcheck:
|
||||
test:
|
||||
[
|
||||
"CMD-SHELL",
|
||||
"python3 -c \"import json, os, urllib.request; req = urllib.request.Request('http://127.0.0.1:8081/api/v1/servers/localhost', headers={'X-API-Key': os.environ['PDNS_API_KEY']}); data = json.load(urllib.request.urlopen(req, timeout=3)); assert data['daemon_type'] == 'authoritative'\""
|
||||
]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 12
|
||||
restart: unless-stopped
|
||||
networks: [backend, edge]
|
||||
|
||||
poweradmin:
|
||||
image: poweradmin/poweradmin:stable
|
||||
container_name: poweradmin
|
||||
depends_on:
|
||||
db:
|
||||
condition: service_healthy
|
||||
auth:
|
||||
condition: service_healthy
|
||||
environment:
|
||||
DB_TYPE: pgsql
|
||||
DB_HOST: ${DB_HOST:-db}
|
||||
DB_PORT: ${DB_PORT:-5432}
|
||||
DB_NAME: ${DB_NAME:?missing DB_NAME}
|
||||
DB_USER: ${DB_USER:?missing DB_USER}
|
||||
DB_PASS: ${DB_PASS:?missing DB_PASS}
|
||||
PA_PDNS_API_URL: http://auth:8081
|
||||
PA_PDNS_API_KEY: ${PDNS_API_KEY:?missing PDNS_API_KEY}
|
||||
PA_DNS_BACKEND: sql
|
||||
PDNS_VERSION: ${PDNS_VERSION:-50}
|
||||
DNS_NS1: ${DNS_NS1:-ns1.wsvc.info}
|
||||
DNS_NS2: ${DNS_NS2:-ns2.wsvc.info}
|
||||
DNS_HOSTMASTER: ${DNS_HOSTMASTER:-hostmaster.wsvc.info}
|
||||
PA_APP_TITLE: ${PA_APP_TITLE:-Poweradmin}
|
||||
PA_TIMEZONE: ${TZ:-Asia/Shanghai}
|
||||
PA_SESSION_KEY: ${PA_SESSION_KEY:?missing PA_SESSION_KEY}
|
||||
PA_CREATE_ADMIN: ${PA_CREATE_ADMIN:-1}
|
||||
PA_ADMIN_USERNAME: ${PA_ADMIN_USERNAME:?missing PA_ADMIN_USERNAME}
|
||||
PA_ADMIN_PASSWORD: ${PA_ADMIN_PASSWORD:?missing PA_ADMIN_PASSWORD}
|
||||
PA_ADMIN_EMAIL: ${PA_ADMIN_EMAIL:?missing PA_ADMIN_EMAIL}
|
||||
PA_ADMIN_FULLNAME: ${PA_ADMIN_FULLNAME:?missing PA_ADMIN_FULLNAME}
|
||||
TRUSTED_PROXIES: private_ranges
|
||||
DEBUG: "false"
|
||||
restart: unless-stopped
|
||||
networks: [backend, frontend]
|
||||
labels:
|
||||
- "traefik.enable=true"
|
||||
- "traefik.docker.network=traefik"
|
||||
- "traefik.http.routers.poweradmin.rule=Host(`pdns.wsvc.info`)"
|
||||
- "traefik.http.routers.poweradmin.entrypoints=websecure"
|
||||
- "traefik.http.routers.poweradmin.tls.certresolver=letsencrypt"
|
||||
- "traefik.http.services.poweradmin.loadbalancer.server.port=80"
|
||||
|
||||
backup:
|
||||
# Use postgres:16 so bash/pg_dump/flock exist without runtime package installs.
|
||||
# backend is internal:true — Alpine apk at start cannot reach mirrors.
|
||||
image: postgres:16
|
||||
container_name: pdns-backup
|
||||
depends_on:
|
||||
db:
|
||||
condition: service_healthy
|
||||
environment:
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
DB_HOST: ${DB_HOST:-db}
|
||||
DB_PORT: ${DB_PORT:-5432}
|
||||
DB_USER: ${PGUSER:?missing PGUSER}
|
||||
DB_PASS: ${PGPASSWORD:?missing PGPASSWORD}
|
||||
DB_NAME: ${DB_NAME:?missing DB_NAME}
|
||||
RETENTION_DAYS: ${RETENTION_DAYS:-7}
|
||||
MAX_BACKUPS: ${MAX_BACKUPS:-7}
|
||||
DUMP_ROLES: ${DUMP_ROLES:-true}
|
||||
CRON_SCHEDULE: ${CRON_SCHEDULE:?missing CRON_SCHEDULE}
|
||||
volumes:
|
||||
- ./backup:/backup
|
||||
- ./scripts:/scripts:ro
|
||||
entrypoint: ["/bin/bash", "/scripts/backup-scheduler.sh"]
|
||||
restart: unless-stopped
|
||||
networks: [backend]
|
||||
|
||||
pgweb:
|
||||
image: sosedoff/pgweb:0.16.2
|
||||
container_name: pdns_pgweb
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
PGWEB_DATABASE_URL: "postgres://${PGUSER:?missing PGUSER}:${PGPASSWORD:?missing PGPASSWORD}@${DB_HOST:-db}:${DB_PORT:-5432}/${DB_NAME:?missing DB_NAME}?sslmode=disable"
|
||||
PGWEB_AUTH_USER: ${PGWEB_USER:?missing PGWEB_USER}
|
||||
PGWEB_AUTH_PASS: ${PGWEB_PASS:?missing PGWEB_PASS}
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
depends_on:
|
||||
db:
|
||||
condition: service_healthy
|
||||
networks: [backend, frontend]
|
||||
labels:
|
||||
- "traefik.enable=true"
|
||||
- "traefik.docker.network=traefik"
|
||||
- "traefik.http.routers.pgweb.rule=Host(`pgweb.wsvc.info`)"
|
||||
- "traefik.http.routers.pgweb.entrypoints=websecure"
|
||||
- "traefik.http.routers.pgweb.tls.certresolver=letsencrypt"
|
||||
- "traefik.http.services.pgweb.loadbalancer.server.port=8081"
|
||||
|
||||
volumes:
|
||||
dbdata: {}
|
||||
@@ -0,0 +1,5 @@
|
||||
# pgdb compose secrets — copy to /opt/database/.env on the host, chmod 600.
|
||||
# NEVER commit the real values. Generate: openssl rand -hex 24
|
||||
POSTGRES_PASSWORD=change-me-strong-hex
|
||||
PGWEB_AUTH_USER=pgweb
|
||||
PGWEB_AUTH_PASS=change-me-strong-hex
|
||||
@@ -0,0 +1,65 @@
|
||||
# pgdb (192.168.55.15) — TimescaleDB + pgweb GUI + nightly backup
|
||||
#
|
||||
# Deploy: copy this file to /opt/database/docker-compose.yml on pgdb,
|
||||
# create /opt/database/.env (chmod 600) from .env.example, plus
|
||||
# /opt/database/pgweb-bookmarks/{hass,scribe}.toml (chmod 600, contains DB password).
|
||||
# Then: docker compose config --quiet && docker compose up -d
|
||||
#
|
||||
# Rollback: previous launch command is kept at /opt/database/run
|
||||
# (container is stateless; data lives on /srv/pgdata).
|
||||
services:
|
||||
timescaledb:
|
||||
image: timescale/timescaledb:latest-pg18
|
||||
container_name: timescaledb
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "192.168.55.15:5432:5432" # bind VM IP only (no IPv6 wildcard)
|
||||
environment:
|
||||
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
|
||||
volumes:
|
||||
- /srv/pgdata:/var/lib/postgresql # data disk (ext4 /dev/sdb1)
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U postgres"]
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 10s
|
||||
|
||||
pgweb:
|
||||
image: sosedoff/pgweb:latest
|
||||
container_name: pgweb
|
||||
restart: unless-stopped
|
||||
# bind/listen/readonly/sessions/bookmarks-only/bookmarks-dir are CLI flags (no env equivalent in v0.17.0)
|
||||
command: ["pgweb", "--bind", "0.0.0.0", "--listen", "8081", "--readonly", "--sessions", "--bookmarks-only", "--bookmarks-dir", "/bookmarks"]
|
||||
ports:
|
||||
- "192.168.55.15:8081:8081" # LAN only + basic auth (see .env)
|
||||
environment:
|
||||
PGWEB_AUTH_USER: ${PGWEB_AUTH_USER}
|
||||
PGWEB_AUTH_PASS: ${PGWEB_AUTH_PASS}
|
||||
PGWEB_BOOKMARKS_DIR: /bookmarks
|
||||
volumes:
|
||||
- ./pgweb-bookmarks:/bookmarks:ro # bookmark .toml files (contain DB password, keep 0600)
|
||||
depends_on:
|
||||
timescaledb:
|
||||
condition: service_healthy
|
||||
|
||||
pg-backup:
|
||||
image: prodrigestivill/postgres-backup-local:latest # latest = postgres 18 base (pg_dump 18.x)
|
||||
container_name: pg-backup
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
POSTGRES_HOST: timescaledb
|
||||
POSTGRES_DB: "hass scribe postgres"
|
||||
POSTGRES_USER: postgres
|
||||
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
|
||||
POSTGRES_EXTRA_OPTS: "-Fc" # custom-format dumps (pg_restore)
|
||||
SCHEDULE: "0 2 * * *" # nightly 02:00 (TZ=Asia/Shanghai -> local 02:00)
|
||||
BACKUP_ON_START: "TRUE" # immediate backup on first start
|
||||
BACKUP_SUFFIX: ".dump"
|
||||
HEALTHCHECK_PORT: "80" # go-cron health endpoint for the image healthcheck
|
||||
TZ: "Asia/Shanghai" # match original host-cron 02:00 local (container default is UTC)
|
||||
volumes:
|
||||
- /opt/database/backups:/backups # POSIX fs required; root disk, separate from data disk
|
||||
depends_on:
|
||||
timescaledb:
|
||||
condition: service_healthy
|
||||
@@ -0,0 +1,24 @@
|
||||
[Unit]
|
||||
Description=Reconcile pgdb compose stack (timescaledb + pgweb + pg-backup) at boot
|
||||
Documentation=file:///opt/database/docker-compose.yml
|
||||
After=network-online.target docker.service
|
||||
Wants=network-online.target
|
||||
Requires=docker.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
RemainAfterExit=yes
|
||||
WorkingDirectory=/opt/database
|
||||
# Idempotent boot-time reconcile. docker's own restore can fail to bind the
|
||||
# published ports (192.168.55.15:5432/8081) when the VM IP is not yet usable
|
||||
# right after boot (EADDRNOTAVAIL, observed 2026-08-30): timescaledb/pgweb
|
||||
# then stay stopped until a manual `docker compose up`. This unit retries
|
||||
# `docker compose up -d` (a no-op when the stack is healthy) until the port
|
||||
# listens, and force-recreates as a last resort to recover a network-detached
|
||||
# container. Data lives on bind mounts (/srv/pgdata, /opt/database/backups),
|
||||
# so recreation is safe.
|
||||
ExecStart=/bin/bash -c 'for i in $(seq 1 12); do docker compose up -d --remove-orphans; sleep 2; if ss -tln | grep -q "192.168.55.15:5432"; then exit 0; fi; sleep 3; done; echo "pgdb-compose: retries exhausted, force-recreating"; docker compose up -d --force-recreate; sleep 10; ss -tln | grep -q "192.168.55.15:5432"'
|
||||
TimeoutStartSec=180
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -0,0 +1,2 @@
|
||||
# Soft Serve initial admin public key (used only on first boot)
|
||||
SOFT_SERVE_INITIAL_ADMIN_KEYS=ssh-ed25519 AAAA... # replace with admin public key
|
||||
@@ -0,0 +1,3 @@
|
||||
FROM alpine:3.20
|
||||
RUN apk add --no-cache sqlite tzdata
|
||||
WORKDIR /scripts
|
||||
@@ -0,0 +1,59 @@
|
||||
services:
|
||||
soft-serve:
|
||||
image: charmcli/soft-serve:v0.12.2
|
||||
container_name: soft-serve
|
||||
restart: unless-stopped
|
||||
# non-root (uid 1000 = windy; 与 backup sidecar BACKUP_UID 一致)
|
||||
user: "1000:1000"
|
||||
|
||||
environment:
|
||||
SOFT_SERVE_DATA_PATH: /var/lib/soft-serve
|
||||
SOFT_SERVE_INITIAL_ADMIN: windy
|
||||
SOFT_SERVE_INITIAL_ADMIN_KEYS: ${SOFT_SERVE_INITIAL_ADMIN_KEYS}
|
||||
|
||||
volumes:
|
||||
- ./data:/var/lib/soft-serve
|
||||
- soft-serve-app:/soft-serve
|
||||
|
||||
networks:
|
||||
- traefik
|
||||
|
||||
labels:
|
||||
- traefik.enable=true
|
||||
|
||||
# SSH over TCP via Traefik (entryPoint ssh -> container port 23231)
|
||||
- traefik.tcp.routers.softserve-ssh.entrypoints=ssh
|
||||
- traefik.tcp.routers.softserve-ssh.rule=HostSNI(`*`)
|
||||
- traefik.tcp.routers.softserve-ssh.tls=false
|
||||
- traefik.tcp.services.softserve-ssh.loadbalancer.server.port=23231
|
||||
|
||||
soft-serve-backup:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile.backup
|
||||
container_name: soft-serve-backup
|
||||
restart: unless-stopped
|
||||
volumes:
|
||||
- ./data:/data:ro
|
||||
- ./backups:/backup
|
||||
- ./scripts:/scripts
|
||||
environment:
|
||||
TZ: Asia/Shanghai
|
||||
BACKUP_UID: 1000
|
||||
BACKUP_GID: 1000
|
||||
entrypoint: >
|
||||
/bin/sh -ec "
|
||||
umask 077 &&
|
||||
touch /backup/backup.log &&
|
||||
crontab /scripts/crontab.txt &&
|
||||
echo '[INFO] soft-serve backup cron installed' &&
|
||||
crond -f -l 8
|
||||
"
|
||||
|
||||
volumes:
|
||||
soft-serve-app:
|
||||
|
||||
networks:
|
||||
traefik:
|
||||
external: true
|
||||
name: vw-net
|
||||
@@ -0,0 +1,15 @@
|
||||
#!/bin/sh
|
||||
set -eu
|
||||
umask 077
|
||||
D() { date "+%Y-%m-%d %H:%M:%S"; }
|
||||
TS=$(date +%Y%m%d_%H%M%S)
|
||||
OUT="/backup/soft-serve_${TS}"
|
||||
mkdir -p "$OUT"
|
||||
echo "[$(D)] Starting soft-serve backup -> $OUT"
|
||||
tar czf "$OUT/repos-config.tar.gz" -C /data repos hooks config.yaml ssh
|
||||
sqlite3 /data/soft-serve.db ".backup '$OUT/soft-serve.db'"
|
||||
chmod 600 "$OUT/repos-config.tar.gz" "$OUT/soft-serve.db"
|
||||
if [ -n "${BACKUP_UID:-}" ] && [ -n "${BACKUP_GID:-}" ]; then
|
||||
chown -R "$BACKUP_UID:$BACKUP_GID" "$OUT" /backup/backup.log
|
||||
fi
|
||||
echo "[$(D)] Backup OK: $(du -sh "$OUT" | cut -f1)"
|
||||
@@ -0,0 +1,4 @@
|
||||
# Run soft-serve backup daily at 02:00
|
||||
0 2 * * * /bin/sh /scripts/backup.sh >> /backup/backup.log 2>&1
|
||||
# Prune backups older than 14 days daily at 03:00
|
||||
0 3 * * * /bin/sh /scripts/prune.sh >> /backup/backup.log 2>&1
|
||||
@@ -0,0 +1,5 @@
|
||||
#!/bin/sh
|
||||
set -eu
|
||||
D() { date "+%Y-%m-%d %H:%M:%S"; }
|
||||
ls -dt /backup/soft-serve_* 2>/dev/null | tail -n +15 | xargs -r rm -rf
|
||||
echo "[$(D)] Pruned. Kept $(ls -d /backup/soft-serve_* 2>/dev/null | wc -l) backups (max 14)"
|
||||
@@ -0,0 +1,38 @@
|
||||
# .env.example — Vaultwarden (us2.wsvc.info, /opt/vaultwarden)
|
||||
#
|
||||
# Non-secret key reference ONLY. Real values live in the server-local .env
|
||||
# (never commit them). Copy the keys below into the server .env if a key is
|
||||
# missing; the compose file requires them via ${VAR} / env_file.
|
||||
|
||||
# Service identity
|
||||
DOMAIN=https://auth.wsvc.info
|
||||
TEMPLATES_FOLDER=
|
||||
|
||||
# Postgres (compose services vaultwarden / backup / pg / pgweb)
|
||||
DB_HOST=pg
|
||||
DB_PORT=5432
|
||||
DB_NAME=vaultwarden
|
||||
DB_USER=vaultwarden
|
||||
DB_PASS=
|
||||
|
||||
# pgweb debug profile
|
||||
PGWEB_USER=
|
||||
PGWEB_PASS=
|
||||
PGWEB_DATABASE_URL=
|
||||
|
||||
# SMTP (mailcow mx2.windy.me:587 starttls)
|
||||
SMTP_HOST=mx2.windy.me
|
||||
SMTP_PORT=587
|
||||
SMTP_SECURITY=starttls
|
||||
SMTP_USERNAME=
|
||||
SMTP_PASSWORD=
|
||||
SMTP_FROM=
|
||||
HELO_NAME=
|
||||
|
||||
# Admin console
|
||||
ADMIN_TOKEN=
|
||||
|
||||
# Runtime
|
||||
UID=1000
|
||||
GID=1000
|
||||
IP_HEADER=X-Forwarded-For
|
||||
@@ -0,0 +1,107 @@
|
||||
services:
|
||||
vaultwarden:
|
||||
image: vaultwarden/server:1.37.2
|
||||
container_name: vaultwarden
|
||||
restart: unless-stopped
|
||||
env_file: ".env"
|
||||
environment:
|
||||
DOMAIN: "https://auth.wsvc.info"
|
||||
DATABASE_URL: "postgresql://${DB_USER}:${DB_PASS}@${DB_HOST}:${DB_PORT}/${DB_NAME}"
|
||||
volumes:
|
||||
- ./vw-data:/data
|
||||
extra_hosts:
|
||||
- "mx2.windy.me:194.163.160.244"
|
||||
networks:
|
||||
- net
|
||||
depends_on:
|
||||
pg:
|
||||
condition: service_healthy
|
||||
labels:
|
||||
- "traefik.enable=true"
|
||||
- "traefik.docker.network=vw-net"
|
||||
|
||||
- "traefik.http.routers.vaultwarden.rule=Host(`auth.wsvc.info`)"
|
||||
- "traefik.http.routers.vaultwarden.entrypoints=websecure"
|
||||
- "traefik.http.routers.vaultwarden.tls=true"
|
||||
- "traefik.http.routers.vaultwarden.tls.certresolver=letsencrypt"
|
||||
|
||||
- "traefik.http.services.vaultwarden.loadbalancer.server.port=80"
|
||||
|
||||
backup:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile.backup
|
||||
container_name: vaultwarden-backup
|
||||
restart: unless-stopped
|
||||
volumes:
|
||||
- ./backups:/backup
|
||||
- ./scripts:/scripts
|
||||
#user: "${UID:-1000}:${GID:-1000}"
|
||||
|
||||
environment:
|
||||
DB_HOST: ${DB_HOST}
|
||||
DB_PORT: ${DB_PORT}
|
||||
DB_USER: ${DB_USER}
|
||||
DB_NAME: ${DB_NAME}
|
||||
DB_PASS: ${DB_PASS}
|
||||
BACKUP_UID: ${UID:-0}
|
||||
BACKUP_GID: ${GID:-0}
|
||||
TZ: Asia/Shanghai
|
||||
entrypoint: >
|
||||
/bin/sh -ec "
|
||||
umask 077 &&
|
||||
printf '%s:%s:*:%s:%s\n' \"$$DB_HOST\" \"$$DB_PORT\" \"$$DB_USER\" \"$$DB_PASS\" > /root/.pgpass &&
|
||||
chmod 600 /root/.pgpass &&
|
||||
touch /backup/backup.log &&
|
||||
crontab /scripts/crontab.txt &&
|
||||
echo '[INFO] Backup cron installed' &&
|
||||
echo '[INFO] Starting crond...' &&
|
||||
crond -f -l 8
|
||||
"
|
||||
networks: [net]
|
||||
|
||||
pg:
|
||||
image: postgres:16
|
||||
container_name: vw-db
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
POSTGRES_DB: ${DB_NAME}
|
||||
POSTGRES_USER: ${DB_USER}
|
||||
POSTGRES_PASSWORD: ${DB_PASS}
|
||||
TZ: Asia/Shanghai
|
||||
PGTZ: Asia/Shanghai
|
||||
volumes:
|
||||
- vwdata:/var/lib/postgresql/data
|
||||
- ./backups:/backup # to import existing dump
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U ${DB_USER} -d ${DB_NAME}"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 10
|
||||
networks: [net]
|
||||
|
||||
pgweb:
|
||||
profiles: ["debug"]
|
||||
image: sosedoff/pgweb:0.16.2
|
||||
container_name: vaultwarden-pgweb
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
# 用 Vaultwarden 的数据库参数拼接连接串
|
||||
#DATABASE_URL: "postgres://${DB_USER}:${DB_PASS}@${DB_HOST}:${DB_PORT}/${DB_NAME}?sslmode=disable"
|
||||
PGWEB_AUTH_USER: ${PGWEB_USER}
|
||||
PGWEB_AUTH_PASS: ${PGWEB_PASS}
|
||||
TZ: Asia/Shanghai
|
||||
#ports:
|
||||
# - "8082:8081" # 本地访问 http://localhost:8082
|
||||
depends_on:
|
||||
pg:
|
||||
condition: service_healthy
|
||||
networks: [net]
|
||||
|
||||
networks:
|
||||
net:
|
||||
name: vw-net
|
||||
external: true
|
||||
|
||||
volumes:
|
||||
vwdata: {}
|
||||
@@ -0,0 +1,51 @@
|
||||
# Domain Docs
|
||||
|
||||
How the engineering skills should consume this repo's domain documentation when exploring the codebase.
|
||||
|
||||
## Before exploring, read these
|
||||
|
||||
- **`CONTEXT.md`** at the repo root, or
|
||||
- **`CONTEXT-MAP.md`** at the repo root if it exists — it points at one `CONTEXT.md` per context. Read each one relevant to the topic.
|
||||
- **`docs/adr/`** — read ADRs that touch the area you're about to work in. In multi-context repos, also check `src/<context>/docs/adr/` for context-scoped decisions.
|
||||
|
||||
If any of these files don't exist, **proceed silently**. Don't flag their absence; don't suggest creating them upfront. The `/domain-modeling` skill (reached via `/grill-with-docs` and `/improve-codebase-architecture`) creates them lazily when terms or decisions actually get resolved.
|
||||
|
||||
## File structure
|
||||
|
||||
Single-context repo (most repos):
|
||||
|
||||
```
|
||||
/
|
||||
├── CONTEXT.md
|
||||
├── docs/adr/
|
||||
│ ├── 0001-event-sourced-orders.md
|
||||
│ └── 0002-postgres-for-write-model.md
|
||||
└── src/
|
||||
```
|
||||
|
||||
Multi-context repo (presence of `CONTEXT-MAP.md` at the root):
|
||||
|
||||
```
|
||||
/
|
||||
├── CONTEXT-MAP.md
|
||||
├── docs/adr/ ← system-wide decisions
|
||||
└── src/
|
||||
├── ordering/
|
||||
│ ├── CONTEXT.md
|
||||
│ └── docs/adr/ ← context-specific decisions
|
||||
└── billing/
|
||||
├── CONTEXT.md
|
||||
└── docs/adr/
|
||||
```
|
||||
|
||||
## Use the glossary's vocabulary
|
||||
|
||||
When your output names a domain concept (in an issue title, a refactor proposal, a hypothesis, a test name), use the term as defined in `CONTEXT.md`. Don't drift to synonyms the glossary explicitly avoids.
|
||||
|
||||
If the concept you need isn't in the glossary yet, that's a signal — either you're inventing language the project doesn't use (reconsider) or there's a real gap (note it for `/domain-modeling`).
|
||||
|
||||
## Flag ADR conflicts
|
||||
|
||||
If your output contradicts an existing ADR, surface it explicitly rather than silently overriding:
|
||||
|
||||
> _Contradicts ADR-0007 (event-sourced orders) — but worth reopening because…_
|
||||
@@ -0,0 +1,17 @@
|
||||
# docs/archive — 归档文档
|
||||
|
||||
归档 = 单次调研、已过期,或与当前运维无行动指向的内容。恢复使用前先确认
|
||||
内容仍与线上状态一致(本仓库原则:先证据后变更,live state 优先)。
|
||||
|
||||
## 归档清单
|
||||
|
||||
| 文件 | 归档日期 | 原位置 | 说明 |
|
||||
|------|---------|--------|------|
|
||||
| `lan-dns-alternatives.md` | 2026-08-17 | `docs/` | DNS 技术选型调研,零引用,无在途决策 |
|
||||
| `agent-runbook-guide.md` | 2026-08-22 | `docs/` | 上游参考存档;仓库落地规范为 `RUNBOOKS.md` |
|
||||
| `lan-core-switch-upgrade-plan.md` | 2026-08-22 | `docs/` | SE5420 历史规划参考;执行以 `docs/lan-se5420-deployment-guide.md` 为准 |
|
||||
| `lan-rb5009-upgrade.md` | 2026-08-22 | `docs/` | 未采购的 ER-X→RB5009 休眠备选方案;其中 PVE 透传与 VLAN10 调研仍可参考 |
|
||||
| `se5420-review-claim-verification-2026-08.md` | 2026-08-22 | `docs/` | 一次性评审复核调研(现场只读复核结论) |
|
||||
|
||||
> 购物类文档(打印机购买指南、交换机选型调研)已按整改计划移入 Obsidian
|
||||
> vault(`~/Documents/vault/my-vault/02_Areas/House/`),不在本目录。
|
||||
@@ -0,0 +1,326 @@
|
||||
# Agent Runbook 实用指南(v1)
|
||||
|
||||
> **定位**:本指南用于把团队的重复性运维、交付与故障处理经验写成可由 Agent 安全执行的流程。它适用于以 Git 仓库为中心的工程协作模式,优先采用 **Markdown + Git 版本控制 + 明确的 Agent 路由规则**,而不是一开始引入复杂的自动化平台。
|
||||
|
||||
> 本文件为上游参考存档。仓库内落地规范见 [`RUNBOOKS.md`](../../RUNBOOKS.md),标准模板见 [`runbooks/_template.md`](../../runbooks/_template.md),索引见 [`runbooks/README.md`](../../runbooks/README.md)。
|
||||
|
||||
## 1. 什么是 Agent Runbook
|
||||
|
||||
Runbook 是预先设计的、可重复执行的操作流程,用于处理部署、告警、故障、配置变更、CI 修复等标准化工作。传统 Runbook 的主要读者是人;**Agent Runbook 则必须把人的隐性判断显式化**,使 Agent 能知道做什么、看到什么才算正常、下一步去哪里、何时停止以及如何撤销。
|
||||
|
||||
Google SRE 强调在事故发生前设计响应流程、系统化排障,并逐步将重复性运维工作自动化。[1] [2] AWS Systems Manager Automation 则把可执行 Runbook 建模为顺序步骤:每个步骤调用一个动作,前一步输出可以传递给后续步骤。[3] 这两种思路共同构成了 Agent Runbook 的实用基础。
|
||||
|
||||
| 层次 | 核心问题 | 应承担的职责 |
|
||||
|---|---|---|
|
||||
| `AGENTS.md` | **何时使用哪份流程?** | 工作路由、通用操作约束、无匹配流程时的默认行为 |
|
||||
| `runbooks/*.md` | **这件事按什么流程做?** | 前置条件、分步操作、决策分支、验证、停止条件与回滚 |
|
||||
| Skill / MCP / Tool | **有哪些可调用能力?** | 具体能力、参数、权限边界和使用说明 |
|
||||
| Shell / GitHub / Linear / SSH 等 | **实际如何执行?** | 对系统、代码库或外部服务执行操作 |
|
||||
|
||||
## 2. 设计目标与适用边界
|
||||
|
||||
Agent Runbook 的目标不是让 Agent 在所有异常下“想办法修好”,而是在一个**已知、受控、可验证、可回退**的边界中提高执行一致性。它应当优先覆盖高频、后果明确、流程稳定的操作,例如 CI 失败定位、Issue 到合并请求、发布前检查、标准部署、回滚及网络变更。
|
||||
|
||||
| 适合纳入 Runbook | 暂不适合直接自动执行 |
|
||||
|---|---|
|
||||
| 明确输入、固定步骤、可观察结果的操作 | 目标或验收标准尚不清楚的探索性任务 |
|
||||
| 可在每次修改后验证状态的变更 | 缺失关键参数、权限或上下文的任务 |
|
||||
| 具有安全回滚路径的发布与配置调整 | 高破坏性、不可逆或影响面未知的操作 |
|
||||
| 可由权限与审批规则约束的运维流程 | 与既有流程事实冲突、无法判断根因的异常场景 |
|
||||
|
||||
> **基本原则**:当实际状态与 Runbook 的假设冲突,Agent 应停止并呈报,而不是补全未知信息、绕过检查或继续试错。
|
||||
|
||||
## 3. Agent Runbook 的最小字段
|
||||
|
||||
与普通人工 Runbook 相比,Agent Runbook 必须显式包含以下六类控制信息。缺少其中任一项,都会增加盲目执行或错误恢复的风险。
|
||||
|
||||
| 字段 | 作用 | 写作要求 |
|
||||
|---|---|---|
|
||||
| **Action** | 定义当前要执行的动作 | 使用可观察、可执行的动词;避免“检查一下”“适当调整”等模糊表述 |
|
||||
| **Expected** | 描述正常状态或预期输出 | 给出具体信号、阈值、状态码、测试结果或页面表现 |
|
||||
| **Decision** | 定义分支与下一跳 | 用“条件 → 下一步”的形式;无法判断时指向 `STOP` |
|
||||
| **Verification** | 确认变更真正生效 | 在每个有副作用的步骤后执行,不能被跳过 |
|
||||
| **Stop condition** | 规定何时不得继续 | 明确列出信息缺失、状态冲突、权限不足、验证失败等条件 |
|
||||
| **Rollback** | 描述如何恢复到变更前状态 | 标明触发条件、前提、撤销步骤及回滚后的验证方式 |
|
||||
|
||||
## 4. 推荐目录与路由机制
|
||||
|
||||
建议把流程与代码一起保存在 Git 仓库中。这样 Runbook 可以评审、版本化、随系统演进更新,也能与相关 Issue、PR 和配置建立可追溯关系。
|
||||
|
||||
```text
|
||||
repo/
|
||||
├── AGENTS.md
|
||||
├── RUNBOOKS.md
|
||||
├── runbooks/
|
||||
│ ├── README.md
|
||||
│ ├── issue-to-merge.md
|
||||
│ ├── fix-ci.md
|
||||
│ ├── release.md
|
||||
│ ├── rollback.md
|
||||
│ ├── network-change.md
|
||||
│ └── network-recovery.md
|
||||
└── ...
|
||||
```
|
||||
|
||||
### `AGENTS.md`:只做路由与通用约束
|
||||
|
||||
`AGENTS.md` 不应重复流程细节。它只需要规定 Agent 在进行操作类工作前,先查找最具体且适用的 Runbook,并严格遵守其中的步骤、验证、停止和审批要求。
|
||||
|
||||
```markdown
|
||||
# Operational Rules
|
||||
|
||||
Before performing operational work:
|
||||
|
||||
1. Inspect `runbooks/`.
|
||||
2. Select the most specific applicable runbook.
|
||||
3. Follow its steps in order.
|
||||
4. Do not skip verification steps.
|
||||
5. Respect STOP and approval conditions.
|
||||
6. If no runbook applies, diagnose only; do not mutate production state.
|
||||
|
||||
## Routing
|
||||
|
||||
- CI failure → `runbooks/fix-ci.md`
|
||||
- GitHub issue implementation → `runbooks/issue-to-merge.md`
|
||||
- Deployment → `runbooks/release.md`
|
||||
- Rollback → `runbooks/rollback.md`
|
||||
- Network configuration → `runbooks/network-change.md`
|
||||
- Network outage → `runbooks/network-recovery.md`
|
||||
```
|
||||
|
||||
### `RUNBOOKS.md`:仓库级规范
|
||||
|
||||
`RUNBOOKS.md` 用于统一所有 Runbook 的字段、命名、评审要求和变更规则。每份 Runbook 只描述一种可识别的操作意图;如果流程已有明显分叉,应拆分为独立文件,而不是堆叠成长篇“万能流程”。
|
||||
|
||||
## 5. 规范模板
|
||||
|
||||
以下模板可直接保存为 `runbooks/_template.md` 使用。
|
||||
|
||||
```markdown
|
||||
# Runbook: <名称>
|
||||
|
||||
## Purpose
|
||||
说明本 Runbook 要解决的问题及成功结果。
|
||||
|
||||
## Scope
|
||||
- 适用环境:<如 development / staging / production>
|
||||
- 适用对象:<服务、仓库、组件或告警类型>
|
||||
- 不适用情形:<需要改用其他 Runbook 或转人工的场景>
|
||||
|
||||
## Ownership
|
||||
- Owner:<团队或角色>
|
||||
- Last reviewed:<YYYY-MM-DD>
|
||||
- Related systems:<系统名称>
|
||||
|
||||
## Preconditions
|
||||
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
|
||||
|
||||
## Inputs
|
||||
| 输入 | 来源 | 是否必需 | 校验方法 |
|
||||
|---|---|---:|---|
|
||||
| <参数> | <来源> | 是/否 | <如何确认有效> |
|
||||
|
||||
## Safety
|
||||
### Non-negotiable rules
|
||||
- 先只读诊断,后执行变更。
|
||||
- 不得把删除现有配置作为首次恢复动作。
|
||||
- 不得猜测或编造缺失参数。
|
||||
- 不得绕过失败的测试、检查或审批。
|
||||
- 每次变更后必须完成对应验证。
|
||||
- 破坏性操作必须获得明确批准。
|
||||
|
||||
### Stop conditions
|
||||
- 实际状态与本文档的前提或预期结果冲突。
|
||||
- 缺少必要输入、权限、审批或回滚能力。
|
||||
- 验证失败且本文档没有明确的下一步。
|
||||
- 影响范围超出 Scope。
|
||||
|
||||
### Approval gates
|
||||
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|
||||
|---|---|---:|---|
|
||||
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Diagnose
|
||||
|
||||
**Action**
|
||||
|
||||
<执行只读诊断动作。>
|
||||
|
||||
**Expected**
|
||||
|
||||
<列出预期输出、状态或证据。>
|
||||
|
||||
**Decision**
|
||||
|
||||
- 若 <条件 A>,进入 Step 2。
|
||||
- 若 <条件 B>,进入 Troubleshooting A。
|
||||
- 若无法判断或状态冲突,`STOP` 并记录证据。
|
||||
|
||||
### Step 2 — Change
|
||||
|
||||
**Action**
|
||||
|
||||
<描述单一、可审计的变更动作。>
|
||||
|
||||
**Expected**
|
||||
|
||||
<变更后应出现的状态。>
|
||||
|
||||
**Verification**
|
||||
|
||||
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
|
||||
|
||||
**Rollback**
|
||||
|
||||
- 触发条件:<什么情况需要回滚>
|
||||
- 回滚动作:<如何撤销>
|
||||
- 回滚验证:<如何确认恢复成功>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Troubleshooting A — <异常名称>
|
||||
|
||||
- 证据收集:<日志、指标、命令输出、链接>
|
||||
- 允许动作:<仅限已验证且低风险的动作>
|
||||
- 下一步:<回到某步 / 转入另一 Runbook / STOP 并升级>
|
||||
|
||||
## Final Verification
|
||||
|
||||
只有同时满足以下标准,流程才算成功:
|
||||
|
||||
- <功能或服务状态>
|
||||
- <自动化测试或健康检查>
|
||||
- <监控指标或告警状态>
|
||||
- <变更记录、PR 或 Issue 已更新>
|
||||
|
||||
## Failure Handling
|
||||
|
||||
若未能完成:
|
||||
|
||||
1. 停止进一步变更。
|
||||
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
|
||||
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
|
||||
4. 按 <升级渠道> 交接,不继续猜测。
|
||||
|
||||
## References
|
||||
|
||||
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
|
||||
```
|
||||
|
||||
## 6. 编写步骤的标准写法
|
||||
|
||||
每个步骤应只承担一个清晰目的,并使用“动作—预期—决策”的闭环表达。如下表所示,前者会导致 Agent 自主扩大操作范围,后者则为其提供安全边界。
|
||||
|
||||
| 不推荐写法 | 推荐写法 |
|
||||
|---|---|
|
||||
| “检查部署是否正常,不正常就修复。” | “读取部署状态与最近一次发布记录。若所有副本 `Ready` 且版本等于目标版本,进入 Final Verification;若副本未就绪,收集事件与日志并进入 Troubleshooting A;若版本不匹配且原因未知,`STOP`。” |
|
||||
| “必要时修改配置。” | “仅当配置差异与变更单 `CHG-123` 完全一致且审批已记录时,应用指定键的值;应用后运行健康检查;失败则按 Rollback 回退。” |
|
||||
| “测试失败时可先跳过。” | “任何必需测试失败均不得继续部署。记录失败测试、日志和提交版本;仅按 Troubleshooting B 处理。” |
|
||||
|
||||
## 7. 通用安全规则
|
||||
|
||||
以下规则适合在每份 Runbook 的 `Safety` 章节中复用。若某流程存在更严格要求,应以更严格要求为准。
|
||||
|
||||
```markdown
|
||||
## Safety Rules
|
||||
|
||||
- Never delete an existing configuration as the first recovery action.
|
||||
- Prefer read-only diagnosis before mutation.
|
||||
- After every mutation, verify the expected state.
|
||||
- If actual state conflicts with this runbook, STOP.
|
||||
- Do not invent missing parameters.
|
||||
- Do not bypass failed tests.
|
||||
- Destructive actions require explicit approval.
|
||||
```
|
||||
|
||||
这些约束体现了一个关键顺序:**先证据,后变更;先小范围,后扩大;先验证,后结束;不确定则停止。** 特别是停止条件必须可操作,例如“权限不足”“缺少变更单”“生产状态与前提不一致”“错误率超过 1%”等,而不应写成“情况复杂时停止”。
|
||||
|
||||
## 8. 运行与审计流程
|
||||
|
||||
Agent 执行 Runbook 时,应按照固定运行模型工作。每一步的输入、动作、输出和下一跳都应可追踪,这与 AWS 自动化 Runbook 的顺序步骤和输出传递思想一致。[3]
|
||||
|
||||
```text
|
||||
输入与前置条件
|
||||
↓
|
||||
只读诊断
|
||||
↓
|
||||
确认预期状态或决策分支
|
||||
↓
|
||||
获取审批(如需要)
|
||||
↓
|
||||
执行最小变更
|
||||
↓
|
||||
立即验证
|
||||
↓
|
||||
成功收尾 / 回滚 / 停止并升级
|
||||
```
|
||||
|
||||
| 阶段 | Agent 必须产出的证据 | 禁止行为 |
|
||||
|---|---|---|
|
||||
| 输入确认 | 参数来源、环境、目标资源、权限与审批状态 | 用猜测值补全必需参数 |
|
||||
| 诊断 | 命令输出、日志、指标或页面状态 | 在未诊断前直接修改生产状态 |
|
||||
| 变更 | 实际执行内容、变更范围、时间 | 将多个无关变更混在一起执行 |
|
||||
| 验证 | 测试、健康检查、监控状态与预期对比 | 以“命令执行成功”代替业务验证 |
|
||||
| 失败处理 | 已做步骤、异常证据、回滚状态和升级对象 | 无限制重试或绕过失败检查 |
|
||||
|
||||
## 9. 从人工操作到自动化的成熟路径
|
||||
|
||||
不建议在流程尚未稳定时先构建复杂 DSL 或全自动编排。应先积累真实案例,把可重复部分固化为 Markdown Runbook,再把已稳定、低歧义、可验证的操作迁移到脚本、CI、Skill 或自动化系统。Google SRE 将能够由机器替代的重复性人工工作视为应逐步消除的 toil。[4]
|
||||
|
||||
| 阶段 | 主要形式 | 人的角色 | 自动化边界 |
|
||||
|---|---|---|---|
|
||||
| 1. 人工处理 | 现场处置与复盘 | 执行、判断、记录 | 不自动化 |
|
||||
| 2. Markdown Runbook | 固化步骤与证据要求 | 审核流程与异常判断 | Agent 可辅助诊断 |
|
||||
| 3. Agent + Runbook | 严格按流程执行 | 审批高风险动作、处理例外 | 受停止条件约束的执行 |
|
||||
| 4. Script / Skill / CI / Automation | 把稳定步骤程序化 | 处理异常和维护自动化 | 自动完成重复性操作 |
|
||||
| 5. 人工审批 + 自动执行 | 常规流程端到端运行 | 决策、审计与治理 | 审批门控下的自动变更 |
|
||||
|
||||
## 10. 上线前检查清单
|
||||
|
||||
在将一份新 Runbook 交给 Agent 使用前,建议由流程所有者按以下清单审核。
|
||||
|
||||
| 检查项 | 合格标准 |
|
||||
|---|---|
|
||||
| 问题边界 | Purpose 与 Scope 清楚描述适用和不适用情形 |
|
||||
| 输入 | 所有必需输入都有来源、格式和校验方法 |
|
||||
| 步骤 | 每一步均有 Action、Expected 与明确的下一跳 |
|
||||
| 变更控制 | 所有修改动作都有 Verification;关键动作有 Rollback |
|
||||
| 安全控制 | Stop conditions、审批门槛和禁止行为已列明 |
|
||||
| 异常处理 | 失败时知道收集什么证据、交给谁,而非继续猜测 |
|
||||
| 可维护性 | 有 Owner、最近复审日期与关联文档;已在版本控制中评审 |
|
||||
| 可演练性 | 已在安全环境或历史案例上走通至少一次 |
|
||||
|
||||
## 11. 建议的首批 Runbook
|
||||
|
||||
首次落地时,应优先选择频率较高、输入相对明确、变更可回退的场景。以下集合通常能覆盖大部分工程协作的基础需求。
|
||||
|
||||
| Runbook | 目的 | 关键安全控制 |
|
||||
|---|---|---|
|
||||
| `issue-to-merge.md` | 从已明确 Issue 到可评审变更 | Scope 锁定、测试门槛、PR 证据 |
|
||||
| `fix-ci.md` | 诊断并修复 CI 失败 | 不跳过测试、不修改无关代码 |
|
||||
| `release.md` | 执行标准发布 | 发布窗口、审批、健康检查、回滚点 |
|
||||
| `rollback.md` | 恢复到已知稳定版本 | 明确触发条件、版本选择、回滚后验证 |
|
||||
| `network-change.md` | 实施受控网络配置变更 | 影响评估、变更单、回退配置 |
|
||||
| `network-recovery.md` | 处理网络异常与服务恢复 | 只读诊断优先、状态冲突即停止 |
|
||||
|
||||
## 12. 结论
|
||||
|
||||
Agent Runbook 的价值不在于把每一项运维工作立即自动化,而在于将团队的工程判断编码为**可路由、可验证、可停止、可回滚**的操作系统。对于多数团队,从仓库中的 `AGENTS.md`、`RUNBOOKS.md` 和一组 Markdown Runbook 起步,已经足够实用。
|
||||
|
||||
当某个流程经过多次执行、输入稳定、异常分支收敛且验证可靠后,再将其下沉为脚本、CI 或其他自动化能力。这样既能逐步降低重复性 toil,也能始终保留人类对高风险和例外情形的决策权。[4]
|
||||
|
||||
## References
|
||||
|
||||
[1]: https://sre.google/sre-book/managing-incidents/ "Google SRE Book — Managing Incidents"
|
||||
[2]: https://sre.google/sre-book/effective-troubleshooting/ "Google SRE Book — Effective Troubleshooting"
|
||||
[3]: https://docs.aws.amazon.com/systems-manager/latest/userguide/automation-documents.html "AWS Systems Manager — Creating your own runbooks"
|
||||
[4]: https://sre.google/sre-book/eliminating-toil/ "Google SRE Book — Eliminating Toil"
|
||||
[5]: https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-automation.html "AWS Systems Manager Automation"
|
||||
[6]: https://docs.aws.amazon.com/systems-manager-automation-runbooks/latest/userguide/automation-runbook-reference.html "AWS Systems Manager Automation Runbook Reference"
|
||||
[7]: https://learn.microsoft.com/en-us/azure/automation/manage-runbooks "Microsoft Learn — Manage runbooks in Azure Automation"
|
||||
|
||||
---
|
||||
|
||||
**来源**:Manus AI《Agent Runbook 实用指南(v1.0)》,本仓库存档为规范参考。
|
||||
+10
-10
@@ -1,13 +1,13 @@
|
||||
# LAN 核心交换机升级计划(保留 ER-X)
|
||||
|
||||
**状态:** SE5420 **已采购**(2026-08-09)。**实施与验证以 [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) 为准**;
|
||||
**状态:** SE5420 **已采购**(2026-08-09)。**实施与验证以 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 为准**;
|
||||
本文为历史规划参考,**不得作为现场执行步骤**;所有实际操作均以部署指南为准。
|
||||
**锁定硬件:** TP-Link **`TL-SE5420`**(16 × 2.5GbE RJ45 + 4 × 10GbE SFP+)。
|
||||
**目标:** SE5420 承接全部 LAN 物理接入与二层转发;ER-X 继续承担公网、NAT、防火墙、
|
||||
LAN66/LAN55 网关与 DHCP。
|
||||
|
||||
**拓扑与流量的详细说明**(职责、逻辑网、流量路径、Wi-Fi 分工、验收边界)见:
|
||||
[lan-erx-se5420-network.md](lan-erx-se5420-network.md)。
|
||||
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)。
|
||||
|
||||
SE5420 官方资料:静态功耗 8 W、最大功耗 32 W;VLAN、LACP、STP/RSTP/MSTP、ACL、
|
||||
CLI/SNMP、配置导入导出与固件下载。无 PoE——AP 使用本地取电 + 普通网线。
|
||||
@@ -65,7 +65,7 @@ VLAN tag。
|
||||
```
|
||||
|
||||
完整端口表、流量路径与 Wi-Fi 分工见
|
||||
[lan-erx-se5420-network.md](lan-erx-se5420-network.md)。
|
||||
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)。
|
||||
|
||||
## VLAN 与端口设计
|
||||
|
||||
@@ -124,8 +124,8 @@ VLAN tag。
|
||||
### 阶段 4:VLAN10 升级专用 Wi-Fi(独立项目)
|
||||
|
||||
不与本次 Done 捆绑。仅 U6;网关 `gfw`;详见
|
||||
[lan-erx-se5420-network.md](lan-erx-se5420-network.md) 第 6.5 / 8 节与
|
||||
[unifi-network.md](unifi-network.md)。
|
||||
[lan-erx-se5420-network.md](../lan-erx-se5420-network.md) 第 6.5 / 8 节与
|
||||
[unifi-network.md](../unifi-network.md)。
|
||||
|
||||
## 性能预期与不变瓶颈
|
||||
|
||||
@@ -148,9 +148,9 @@ VLAN tag。
|
||||
|
||||
## 参考
|
||||
|
||||
- [ER-X + SE5420 网络与拓扑说明](lan-erx-se5420-network.md)
|
||||
- [LAN 概览](lan-overview.md)
|
||||
- [ER-X 配置记录](edgerouter-x-configuration.md)
|
||||
- [UniFi 网络](unifi-network.md)
|
||||
- [`gfw`](../hosts/gfw.windy.lan.md)
|
||||
- [ER-X + SE5420 网络与拓扑说明](../lan-erx-se5420-network.md)
|
||||
- [LAN 概览](../lan-overview.md)
|
||||
- [ER-X 配置记录](../edgerouter-x-configuration.md)
|
||||
- [UniFi 网络](../unifi-network.md)
|
||||
- [`gfw`](../../hosts/gfw.windy.lan.md)
|
||||
- [TL-SE5420 官方规格](https://www.tp-link.com.cn/product_2899.html?v=specification)
|
||||
@@ -0,0 +1,605 @@
|
||||
# Home-LAN DNS alternatives for the windy LAN (research, 2026-08)
|
||||
|
||||
**Status: research only. No configuration was changed.** This page evaluates
|
||||
resolvers/splitters that are genuinely better than — or meaningfully different
|
||||
from — the current "AdGuard Home (AGH) + mosdns" setup on
|
||||
[`dns.windy.lan`](../../hosts/dns.windy.lan.md) (`.36`), for a GFW-constrained
|
||||
China home LAN. Claims are cited to primary sources (official repos, official
|
||||
docs, upstream READMEs); anything not verified is flagged as such.
|
||||
|
||||
> 2026-08-12: facts in this page's scope recap were refreshed by W1N-56 live
|
||||
> verification — mosdns on `.1` is **not idle**, it is clash's
|
||||
> `nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution); the
|
||||
> canonical decision record is
|
||||
> [`lan-dns-architecture.md`](../lan-dns-architecture.md) (final verdict aligned,
|
||||
> Phase 0 kill-test evidence incl. a measured upstream-blackhole degradation
|
||||
> gap).
|
||||
|
||||
Scope recap (from [`lan-overview.md`](../lan-overview.md), verified 2026-08-06):
|
||||
|
||||
- Clients get DNS via EdgeRouter DHCP option 6 → AGH `192.168.66.36:53`.
|
||||
- AGH upstreams: `dns.alidns.com` + `doh.pub` DoH (load-balanced), fallback
|
||||
`https://adg.chans.xyz/dns-query`. **DNSSEC disabled** (known-bad-signature
|
||||
check failed on the selected path). Rewrites: `hass.local` / `hass.windy.lan`.
|
||||
- `gfw` OpenWrt (`.1`) runs OpenClash fake-ip + TPROXY; dnsmasq → clash DNS
|
||||
`127.0.0.1#7874`. `mosdns` on `127.0.0.1:6052` is clash's
|
||||
`nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution: domestic →
|
||||
AGH `.36:53`, foreign → `223.5.5.5`/`119.29.29.29`); it is **not** in the LAN
|
||||
client query path.
|
||||
- No local authoritative PTR source yet; private reverse DNS is a known gap.
|
||||
|
||||
---
|
||||
|
||||
## 1. TL;DR / recommendation
|
||||
|
||||
**The current stack is already 80% of the answer.** AGH is a strong LAN DNS
|
||||
front-end (filtering, rewrites, per-client upstreams, query log, web UI) and its
|
||||
upstream layer — **per-domain upstreams** plus a **per-domain list loaded from a
|
||||
file** (`upstream_dns_file`) — is exactly the mechanism the official docs
|
||||
recommend for accelerating China CDN domains while keeping everything else on a
|
||||
trusted path. [AGH configuration: upstreams](https://adguard-dns.io/kb/adguard-home/configuration/).
|
||||
|
||||
The genuinely worthwhile changes, in order of value:
|
||||
|
||||
1. **Add geo-split inside AGH** via `upstream_dns_file` fed by a converted
|
||||
`accelerated-domains.china.conf` ([felixonmars/dnsmasq-china-list](https://github.com/felixonmars/dnsmasq-china-list)):
|
||||
domestic CDN domains → `dns.alidns.com` / `doh.pub`; everything else →
|
||||
the trusted foreign path (currently `adg.chans.xyz`). This is a documented
|
||||
AGH use case, requires **no new daemon**, and removes the need for mosdns.
|
||||
This is the top recommendation.
|
||||
2. **Re-enable real DNSSEC** by putting validation behind AGH: AGH's
|
||||
`enable_dnssec` only sets the DO bit — it does not validate
|
||||
([AGH config: DNSSEC](https://adguard-dns.io/kb/adguard-home/configuration/)).
|
||||
The two realistic ways are (a) point the foreign/trusted default upstream at
|
||||
a validating resolver ([unbound](https://unbound.docs.nlnetlabs.nl/en/latest/),
|
||||
[blocky](https://0xerr0r.github.io/blocky/latest/configuration/#dnssec-validation))
|
||||
and re-test a known-bad-signature domain; or (b) insert a validating
|
||||
resolver (blocky is the lightest) between AGH and the upstreams.
|
||||
3. **mosdns on `.1` is resolved, not idle** — it is clash's
|
||||
`nameserver`/`default-nameserver` (DIRECT-rule real-IP resolution, verified
|
||||
2026-08-12), so "delete it" is off the table; its role is documented in
|
||||
[`lan-dns-architecture.md`](../lan-dns-architecture.md) §1. If a future change
|
||||
moves this role to an AGH-side companion, keep in mind mosdns's cache strips
|
||||
EDNS0 and it performs no DNSSEC validation
|
||||
([mosdns v5 executable plugins](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)).
|
||||
|
||||
Top-3 alternatives worth pursuing (see §3 for detail):
|
||||
|
||||
| Rank | Option | Why |
|
||||
|------|--------|-----|
|
||||
| 1 | **AGH with China-list geo-split (`upstream_dns_file`)** | Documented AGH pattern; single box; no new service; keeps filtering/rewrites/UI. |
|
||||
| 2 | **Blocky as validating backend behind AGH** | The only "new software" option that adds real in-process DNSSEC validation + conditional per-domain upstreams + ECS in one static binary ([blocky README](https://github.com/0xERR0R/blocky), [config](https://0xerr0r.github.io/blocky/latest/configuration/)). |
|
||||
| 3 | **Unbound as validating recursive resolver** (replaces forwarders for the foreign path, or whole path) | True validation, full recursion (fail-open by nature), private `local-zone`s; heavier ops than AGH's file-driven split. |
|
||||
|
||||
Explicitly **not** recommended as replacements here: smartdns and chinadns-ng
|
||||
(both excellent *splitters*, but neither validates DNSSEC and both lack AGH's
|
||||
filtering/UI/query-log layer, so they add a daemon without closing the DNSSEC
|
||||
gap); mihomo/sing-box DNS as the primary path (couples DNS to the proxy and is
|
||||
fail-closed; keep for proxy-side concerns only); knot-resolver/dnsdist (overkill
|
||||
for a single-operator home LAN).
|
||||
|
||||
---
|
||||
|
||||
## 2. Requirement matrix
|
||||
|
||||
Legend: **●** native/built-in · **◐** possible with config/lists · **○** absent/
|
||||
not applicable. "Geo-split" = route domestic vs foreign names to different
|
||||
upstreams. "Anti-pollution" = a mechanism to avoid/adjudicate poisoned answers
|
||||
(IP-verdict or trusted-upstream routing). "DNSSEC" = performs validation
|
||||
in-process (not just forwards DO).
|
||||
|
||||
| Candidate | Geo-split | Anti-pollution | DNSSEC (validate) | Cache | Private names / rewrites | Ops simplicity | License |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| AGH (current) | ◐ per-domain upstreams + list file | ◐ via trusted foreign upstream | ○ (DO bit only) | ● | ● rewrites, per-client, private-PTR | ● Docker + UI | GPL-3.0 |
|
||||
| mosdns v5 | ● domain/ip list matchers | ◐ forward foreign→trusted | ○ | ● (strips EDNS0) | ● hosts/redirect/reverse_lookup | ◐ single binary, YAML, no UI | GPL-3.0 |
|
||||
| smartdns | ● nameserver groups + domain lists | ● bogus-nxdomain / blacklist-ip / trusted groups | ○ (no option in config ref) | ● serve-expired | ● address / local-domain / lease file | ◐ single binary, optional WebUI plugin | GPL-3.0 |
|
||||
| chinadns-ng | ● chnlist/gfwlist + tag:none IP-test | ● IP verdict via chnroute ipset/nftset | ○ | ● cache/stale/verdict | ◐ hosts / dns-rr-ip | ◐ single static binary, config file | AGPL-3.0 |
|
||||
| dnsmasq-china-list | ◐ (data only) | ◐ (via host resolver) | ◐ via host | ◐ via host | ◐ via host | ◐ feed lists | WTFPL |
|
||||
| unbound | ◐ forward-zones / RPZ / views | ◐ forward-zones + bogus-nxdomain | ● | ● serve-expired | ● local-zone / local-data | ◐ config daemon, no UI | BSD-style (NLnet) |
|
||||
| blocky | ◐ conditional per-domain + client groups | ◐ blocking lists + conditional routing | ● | ● prefetch | ● customDNS / rewrite / hosts | ◐ single binary, YAML, REST (no full web UI) | Apache-2.0 |
|
||||
| Technitium | ◐ conditional-forwarder zones / apps | ◐ blocked lists + forwarding | ● | ● persistent | ● zones, stub, split-horizon | ● .NET + web console | GPL-3.0 |
|
||||
| sing-box | ● DNS rules (geoip/geosite) | ● rule-based servers + (proxy) sniffing | ○ | ● LRU + optimistic | ● hosts / local server | ◐ single binary, JSON | GPLv3-family (metadata "other") |
|
||||
| mihomo | ● nameserver-policy + fallback-filter | ● geoip verdict + geosite | ○ | ● (cache-algorithm) | ● hosts; fake-ip-filter for `.lan` | ◐ single binary, YAML | not cleanly verifiable (repo obfuscated) |
|
||||
| knot-resolver | ◐ policy modules | ◐ policy + RPZ | ● | ● persistent | ◐ hints / local data | ◐ systemd, Lua config | open source (CZ-NIC) |
|
||||
| dnsdist | ◐ Lua rules (custom) | ◐ custom policies | ○ (balancer, not validator) | ○ (no cache of its own) | ○ | ○ power tool | GPL (PowerDNS) |
|
||||
|
||||
Notes:
|
||||
|
||||
- "Geo-split" for AGH/blocky/unbound/Technitium is real but requires feeding a
|
||||
China domain list; chinadns-ng/mihomo additionally offer the **IP-verdict**
|
||||
path for domains not in any list (query both, adopt CN result only if the
|
||||
answer IP is mainland).
|
||||
- mosdns v5's `cache` plugin ignores request EDNS0 and strips response EDNS0
|
||||
([cache plugin](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)) —
|
||||
relevant because AGH in front of it relies on the DO bit for DNSSEC-capable
|
||||
upstreams.
|
||||
- License for mihomo/sing-box/knot-resolver marked conservative: GitHub
|
||||
metadata is "other"/custom or deliberately obfuscated; see §3 caveats.
|
||||
|
||||
---
|
||||
|
||||
## 3. Per-candidate evaluation
|
||||
|
||||
### 3.1 AdGuard Home — advanced upstream routing / built-ins
|
||||
|
||||
What it is: Go DNS proxy + adblock + DHCP, LAN DNS front-end
|
||||
([official](https://adguard-dns.io/kb/adguard-home/overview/)).
|
||||
|
||||
Capabilities relevant here (all from the official configuration page):
|
||||
|
||||
- **Per-domain upstreams** dnsmasq-style: `[/domain/]upstream`, wildcards,
|
||||
`#` = "default upstreams", empty `//` = unqualified names
|
||||
([upstreams for domains](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-for-domains)).
|
||||
- **List from file** `upstream_dns_file` — the docs *explicitly* call out China
|
||||
CDN acceleration via dnsmasq lists, with the `server=/0-100.com/114.114.114.114`
|
||||
→ `[/0-100.com/]114.114.114.114` conversion
|
||||
([loading upstreams from file](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-from-file)).
|
||||
- Upstream modes: `load_balance`, `parallel`, `fastest_addr`; plus `fallback_dns`
|
||||
used only when primary upstreams fail
|
||||
([config file: dns](https://adguard-dns.io/kb/adguard-home/configuration/)).
|
||||
- Per-client upstreams (`clients.persistent[].upstreams`), rewrites
|
||||
(`filtering.rewrites`, incl. wildcard), `local_ptr_upstreams` for private PTR,
|
||||
ECS (`edns_client_subnet` with `use_custom` coarse prefix), optimistic cache
|
||||
([same page](https://adguard-dns.io/kb/adguard-home/configuration/)).
|
||||
- **DNSSEC is DO-bit only**: `enable_dnssec` "defines whether the proxy should
|
||||
set the DO flag in the upstream requests" — validation must happen upstream
|
||||
([same page](https://adguard-dns.io/kb/adguard-home/configuration/)).
|
||||
- DoH/DoT/DoQ/DoH3 serving, `bind_hosts`/ACL guidance
|
||||
([running securely](https://adguard-dns.io/kb/adguard-home/running-securely/)).
|
||||
|
||||
Verdict: **Already installed and capable of the geo-split itself.** The current
|
||||
setup under-uses it: only a load-balanced CN pair + fallback, no per-domain
|
||||
routing and no validating upstream. This is the cheapest "better" state — see §5.
|
||||
|
||||
### 3.2 mosdns v5 — installed, active as clash nameserver (gateway-side)
|
||||
|
||||
What it is: "一个 DNS 转发器" (a DNS forwarder) — plugin-based, sequence-driven
|
||||
([README](https://github.com/IrineSistiana/mosdns), GPL-3.0, ~3.7k★).
|
||||
|
||||
What it does (verified from the v5 wiki and source tree):
|
||||
|
||||
- Servers: `udp_server`, `tcp_server` (TLS→DoT), `quic_server`, `http_server`
|
||||
(DoH); upstreams in `forward` support `udp`, `tcp`, `tls`, `https`, `quic`,
|
||||
HTTP/3, concurrent racing (`concurrent: n` picks the fastest) and socks5
|
||||
([server plugins](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/fu-wu-qi-cha-jian.md),
|
||||
[executable plugins](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)).
|
||||
- Geo-split: v5 data providers are **`domain_set` / `ip_set` (text list files)**
|
||||
plus `qname`/`resp_ip` matchers — verified from the current source tree
|
||||
([plugin/data_provider](https://github.com/IrineSistiana/mosdns/tree/main/plugin/data_provider))
|
||||
— and an `ipset`/`nftset` exec plugin to push answer IPs to kernel sets. The
|
||||
old v4-style `geosite`/`geoip` `.dat` plugins are **not present** in the v5
|
||||
tree; the v5 wiki's own matcher page currently states there are no matcher
|
||||
plugins to document
|
||||
([matcher page](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/pi-pei-qi-cha-jian.md)).
|
||||
Plan on chnlist/gfwlist-style text lists, not `geosite.dat`.
|
||||
- Cache: yes, incl. optional lazy cache and disk dump; **request EDNS0 is
|
||||
ignored and response EDNS0 stripped** by the cache plugin
|
||||
([cache](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)).
|
||||
- Private names: `hosts` (domain-rules style, not OS /etc/hosts syntax),
|
||||
`redirect`, `arbitrary` (zone records), `reverse_lookup` (PTR/HTTP lookup).
|
||||
- Ops: single binary + YAML; `mosdns service install` ships a systemd/launchd
|
||||
helper ([v5 overview](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5.md));
|
||||
Docker image exists. No web UI of its own.
|
||||
|
||||
Verdict: capable splitter/forwarder, but **adds no DNSSEC and no filtering
|
||||
layer**, and its cache interferes with EDNS0/DO handling. As a *back-end* splitter
|
||||
behind AGH it is a legitimate choice only if DNSSEC stays off. Given AGH can do
|
||||
the same per-domain split natively (3.1), mosdns's marginal value here is
|
||||
concurrent upstream racing and ipset/nftset integration — neither is needed at
|
||||
this LAN's scale. Either wire it up properly or remove it.
|
||||
|
||||
### 3.3 smartdns
|
||||
|
||||
What it is: local DNS server that queries multiple upstreams, **speed-tests the
|
||||
answer IPs and returns the fastest**; DoH/DoT/DoQ/DoH3; GPL-3.0, ~11.2k★
|
||||
([README](https://github.com/pymumu/smartdns)).
|
||||
|
||||
Capabilities (from the official config reference and FAQ):
|
||||
|
||||
- Multi upstream + "returns the fastest IP", unlike dnsmasq all-servers
|
||||
([README](https://github.com/pymumu/smartdns)).
|
||||
- Domain groups: `server ... -group <name>` + `nameserver /domain/group` routing,
|
||||
per-`bind` port flags (`-group`, `-no-speed-check`…), client rules/MAC/IP
|
||||
([config options](https://pymumu.github.io/smartdns/configuration/)).
|
||||
- Anti-pollution tooling: `bogus-nxdomain` (return NXDOMAIN for poisoned IPs),
|
||||
`blacklist-ip`, `whitelist-ip`, `ignore-ip`, `ipset`/`nftset` export
|
||||
([same page](https://pymumu.github.io/smartdns/configuration/)).
|
||||
- ECS: global `edns-client-subnet` and per-server `-subnet`
|
||||
([same page](https://pymumu.github.io/smartdns/configuration/)).
|
||||
- Cache: `cache-size`, `serve-expired` (RFC-like stale), `prefetch-domain`,
|
||||
persistent cache file ([same page](https://pymumu.github.io/smartdns/configuration/)).
|
||||
- Private names: `address`, `cname`, `local-domain`, `dnsmasq-lease-file`
|
||||
([same page](https://pymumu.github.io/smartdns/configuration/)).
|
||||
- **DNSSEC: no validation option appears anywhere in the official config
|
||||
reference or FAQ** — its pollution model is blacklist/whitelist + trusted
|
||||
groups + speed selection, not DNSSEC ([config options](https://pymumu.github.io/smartdns/configuration/),
|
||||
[FAQ](https://pymumu.github.io/smartdns/faq/)). Flagged: verify on the version
|
||||
you deploy before relying on it.
|
||||
|
||||
Verdict: the classic China-home "best-IP" resolver; good splitter, no DNSSEC,
|
||||
speed-test model optimizes for latency rather than anti-pollution correctness.
|
||||
Not better than AGH+China-list for this LAN; at most a back-end splitter behind
|
||||
AGH, with the same DNSSEC caveat as mosdns.
|
||||
|
||||
### 3.4 chinadns-ng / chinadns2 / dnsmasq-china-list
|
||||
|
||||
**chinadns-ng** (the requested "china-dns-ng"; actual repo `zfl9/chinadns-ng`,
|
||||
AGPL-3.0, Zig, ~1.4k★) is the maintained rewrite of shadowsocks/ChinaDNS:
|
||||
|
||||
- Two upstream groups (china / trust) + `chnlist.txt` / `gfwlist.txt` domain
|
||||
lists; domains are tagged `chn`/`gfw`/`none`
|
||||
([README](https://github.com/zfl9/chinadns-ng)).
|
||||
- `tag:none` names are queried on **both** upstreams and the china answer is
|
||||
adopted only if its A/AAAA is a mainland IP (tested against a `chnroute`
|
||||
ipset/nftset loaded into the kernel); verdict caching avoids re-testing and
|
||||
leaks ([README: 原理/verdict-cache](https://github.com/zfl9/chinadns-ng)).
|
||||
- Cache with stale + pre-refresh + optional persistence; DoT upstream
|
||||
(wolfssl build); `hosts` + `dns-rr-ip` local records; `nftset` add for
|
||||
chn/gfw IPs; **no DoH by design** and **no DNSSEC** — the author's stated
|
||||
philosophy is "one job, done well"
|
||||
([README](https://github.com/zfl9/chinadns-ng)).
|
||||
- Resource footprint is tiny: ~140 KB baseline, ~2.4 MB with 73k+ chnlist +
|
||||
5.7k gfwlist entries ([README](https://github.com/zfl9/chinadns-ng)).
|
||||
|
||||
**chinadns2** (`zfl9/chinadns2`) is the older C predecessor; effectively
|
||||
superseded by chinadns-ng for new deployments (README not directly fetched —
|
||||
treat as legacy line).
|
||||
|
||||
**dnsmasq-china-list** (felixonmars, ~6.1k★) is data, not a daemon:
|
||||
`accelerated-domains.china.conf`, `bogus-nxdomain.china.conf`,
|
||||
`apple.china.conf`, `google.china.conf`, with generators for **dnsmasq,
|
||||
unbound, bind, dnscrypt-proxy**
|
||||
([README](https://github.com/felixonmars/dnsmasq-china-list), WTFPL per repo).
|
||||
|
||||
Verdict: chinadns-ng is the strongest *pure splitter* for GFW networks (IP
|
||||
verdict beats pure list-based routing for unknown domains), but it cannot
|
||||
validate DNSSEC and brings no filtering UI. As AGH's backend it duplicates what
|
||||
AGH's per-domain upstreams already do; its IP-test mode requires shipping
|
||||
`chnroute` ipset/nftset into the host. dnsmasq-china-list is best used as the
|
||||
**data feed** for the AGH `upstream_dns_file` recommendation in §5.
|
||||
|
||||
### 3.5 unbound
|
||||
|
||||
What it is: validating, recursive, caching resolver from NLnet Labs
|
||||
([docs](https://unbound.docs.nlnetlabs.nl/en/latest/)).
|
||||
|
||||
- **Real DNSSEC validation by default** (trust anchor, chain of trust); the
|
||||
official home-network guide turns it on explicitly
|
||||
([home resolver guide](https://unbound.docs.nlnetlabs.nl/en/latest/use-cases/home-resolver.html)).
|
||||
- Full recursion → does not hard-depend on any upstream or proxy; serve-expired
|
||||
(RFC 8767), aggressive NSEC, DoH/DoT/DoQ serving and TLS upstreams,
|
||||
forward-zone/stub-zone/authority-zone, RPZ filtering, views, ECS module
|
||||
([docs index](https://unbound.docs.nlnetlabs.nl/en/latest/)).
|
||||
- Private names: `local-zone`/`local-data` for `*.windy.lan`-style names
|
||||
([unbound.conf(5)](https://unbound.docs.nlnetlabs.nl/en/latest/manpages/unbound.conf.html)).
|
||||
- No built-in China split: you assemble it with forward-zones fed by
|
||||
dnsmasq-china-list (`make unbound` generator) + `bogus-nxdomain`; no UI, no
|
||||
per-client grouping comparable to AGH.
|
||||
|
||||
Verdict: the gold standard for the **validation** half. Best used as (a) the
|
||||
validating upstream behind AGH for the foreign/trusted path, or (b) a full
|
||||
recursive resolver replacing the forwarders if you accept losing AGH-style
|
||||
filtering/UI on top — keep AGH in front for that. System-package based, heavier
|
||||
to operate than blocky but battle-tested.
|
||||
|
||||
### 3.6 blocky
|
||||
|
||||
What it is: Go DNS proxy + ad-blocker, "fast and lightweight", single static
|
||||
binary, stateless, Apache-2.0, ~6.9k★
|
||||
([README](https://github.com/0xERR0R/blocky)).
|
||||
|
||||
- **In-process DNSSEC validation**: `dnssec.validate` with DO bit, RRSIG
|
||||
verification, chain-of-trust, NSEC/NSEC3, custom trust anchors, SERVFAIL on
|
||||
bogus ([DNSSEC validation docs](https://0xerr0r.github.io/blocky/latest/configuration/#dnssec-validation)).
|
||||
- Upstreams: `parallel_best` (2 random resolvers, fastest answer), `strict`,
|
||||
`random`; per-client/per-subnet upstream **groups**; UDP/TCP/DoT/DoH/DoQ/DoH3;
|
||||
DNS stamps; bootstrap DNS
|
||||
([upstreams](https://0xerr0r.github.io/blocky/latest/configuration/#upstreams-configuration)).
|
||||
- Conditional forwarding + `customDNS` mapping/rewrite (the AGH-rewrite
|
||||
equivalent), hosts files, per-domain upstream routing
|
||||
([custom DNS / conditional](https://0xerr0r.github.io/blocky/latest/configuration/#custom-dns)).
|
||||
- ECS: `ecs.useAsClient` / `ecs.forward`
|
||||
([ECS](https://0xerr0r.github.io/blocky/latest/configuration/#edns-client-subnet-options)).
|
||||
- Cache with min/max TTL + **prefetching**; optional **Redis** cache/state sync
|
||||
between instances; query log to SQLite/Postgres/CSV; Prometheus metrics; REST
|
||||
API ([README](https://github.com/0xERR0R/blocky),
|
||||
[config](https://0xerr0r.github.io/blocky/latest/configuration/)).
|
||||
- No full web admin UI (metrics/REST/logs only) — an ops trade-off vs AGH's UI.
|
||||
|
||||
Verdict: the most attractive *new software* option for this LAN **as a backend
|
||||
behind AGH**: it adds real DNSSEC validation + conditional upstream routing +
|
||||
ECS with a single binary and YAML. It has no China-IP-verdict split built in —
|
||||
feed it the China domain list via `conditional.mapping`/upstream groups, which
|
||||
is fine at this scale. One caveat: no GUI means AGH stays the human-facing
|
||||
front, so AGH→blocky is strictly additive.
|
||||
|
||||
### 3.7 Technitium DNS Server
|
||||
|
||||
What it is: self-hosted authoritative **and** recursive DNS server, .NET,
|
||||
web console, GPL-3.0, ~9.5k★
|
||||
([README](https://github.com/TechnitiumSoftware/DnsServer)).
|
||||
|
||||
- **DNSSEC validation** for recursive resolution, forwarders, and conditional
|
||||
forwarders (RSA/ECDSA/EdDSA, NSEC/NSEC3); can also *serve* signed zones
|
||||
([README](https://github.com/TechnitiumSoftware/DnsServer)).
|
||||
- Conditional forwarder zones + bulk conditional forwarding app; blocked-domain
|
||||
lists with regex support and per-client variants; split-horizon/geolocation
|
||||
via DNS Apps; ECS; QNAME minimization
|
||||
([README](https://github.com/TechnitiumSoftware/DnsServer)).
|
||||
- Serving side: DoH/DoT/DoQ/DoH3 server, built-in DHCP, persistent cache,
|
||||
caching with serve-stale/prefetch, clustering, HTTP/SOCKS5 proxy for DNS
|
||||
(e.g. over Tor) ([README](https://github.com/TechnitiumSoftware/DnsServer)).
|
||||
- Heavier footprint (needs .NET; Docker image available) and a full web console
|
||||
with many features this LAN won't use.
|
||||
|
||||
Verdict: capable and genuinely feature-rich (a real AGH alternative in the
|
||||
"everything in one box" sense — filtering, private zones, validation, DHCP), but
|
||||
it's more moving parts than this LAN needs, and its geo-split still requires
|
||||
manual conditional-forwarder lists. Not chosen over the lighter AGH+backend
|
||||
approach.
|
||||
|
||||
### 3.8 sing-box / mihomo built-in DNS as the split resolver (fake-ip)
|
||||
|
||||
The "third option": let the proxy engine's DNS own resolution, AGH on top.
|
||||
|
||||
**sing-box** DNS object: multiple server types (local, udp, tcp, tls, https,
|
||||
http3, quic, fakeip, hosts, dhcp, mdns…), rule-based server selection by
|
||||
geoip/geosite, LRU cache + optimistic serving, per-query timeout, `client_subnet`
|
||||
(ECS), `reverse_mapping`
|
||||
([sing-box DNS docs](https://sing-box.sagernet.org/configuration/dns/)).
|
||||
|
||||
**mihomo** (Clash.Meta lineage; docs at
|
||||
[wiki.metacubex.one](https://wiki.metacubex.one/en/config/dns/)):
|
||||
`nameserver-policy` (geosite/rule-set/domain keys) routes specific domains to
|
||||
specific resolvers; `fallback` + `fallback-filter` (geoip=CN, geosite=gfw,
|
||||
ipcidr, domain) adjudicate pollution — a CN resolver's answer is adopted only if
|
||||
the IP is mainland, otherwise the overseas fallback's answer is used;
|
||||
`fake-ip`/`redir-host` enhanced mode, `fake-ip-filter` with e.g. `'*.lan'` to
|
||||
keep local names on real-IP; per-DNS-server ECS; cache-algorithm
|
||||
([mihomo DNS config](https://wiki.metacubex.one/en/config/dns/)).
|
||||
|
||||
Assessment for THIS LAN:
|
||||
|
||||
- The **pollution adjudication is strong** (geoip-verdict fallback, geosite
|
||||
lists), and mihomo already runs on the gateway — so "clash DNS as splitter" is
|
||||
tempting.
|
||||
- But the DNS service is **coupled to the proxy**: foreign resolution rides the
|
||||
proxy path, so when OpenClash/subscription is down, fake-ip mapping and
|
||||
foreign lookups break (partial fail-open only if `direct-nameserver`/fallback
|
||||
are carefully set). The LAN requirement says **must not hard-depend on the
|
||||
proxy (fail-open)**.
|
||||
- fake-ip adds an indirection layer for anything in front of it (AGH on top
|
||||
resolves client IPs against fake-ip ranges; leaks/loops need careful rules).
|
||||
- Neither engine **validates DNSSEC** (no RRSIG verification).
|
||||
- sing-box repo license shows "other" in GitHub metadata (not cleanly
|
||||
verifiable); mihomo's repo currently carries **deliberately obfuscated content**
|
||||
("Void Terminal" parody) — treat `wiki.metacubex.one` as the authoritative
|
||||
docs and expect the GitHub surface to change.
|
||||
|
||||
Verdict: keep clash/mihomo DNS exactly where it is (proxy-side, TPROXY/fake-ip),
|
||||
do **not** make it the LAN resolver of record. If you ever want its IP-verdict
|
||||
quality outside the proxy, chinadns-ng gives the same idea with zero proxy
|
||||
dependency.
|
||||
|
||||
### 3.9 knot-resolver / dnsdist — power-resolver options
|
||||
|
||||
**knot-resolver** (CZ-NIC): minimal caching validating resolver, modular/Lua,
|
||||
full DNSSEC validation, forwarding over TLS, query policies, RPZ, views/ACLs,
|
||||
DNS64, persistent cache, serve-stale, even XDP fast-path
|
||||
([docs](https://knot-resolver.readthedocs.io/en/stable/)). As powerful as
|
||||
unbound but with more configuration surface (Lua); overkill for a one-operator
|
||||
home LAN, though it would do the validating-resolver role well.
|
||||
|
||||
**dnsdist** (PowerDNS): "highly DNS-, DoS- and abuse-aware loadbalancer" —
|
||||
routes traffic to backend servers, Lua/YAML config, runtime console, metrics
|
||||
([overview](https://dnsdist.org/)). It is a **balancer, not a validator/cache**
|
||||
— it fronts other resolvers. Overkill; only relevant if you wanted a
|
||||
multi-backend DNS LB, which this LAN does not.
|
||||
|
||||
### 3.10 Emerging / also-considered options
|
||||
|
||||
- **AdGuard Home + dnsmasq-china-list** — covered in §3.1/§5; this is the
|
||||
"emerging best practice" for China CDN splits on AGH and is officially
|
||||
documented.
|
||||
- **pi-hole** — adblock/dashboard equivalent of AGH but no per-domain upstream
|
||||
routing worth choosing it over AGH here (not deeply verified for this write-up;
|
||||
AGH already satisfies the role).
|
||||
- **dnscrypt-proxy** — encrypted forwarder with stamp support; a transport
|
||||
option, not a splitter/validator (not deeply verified for this write-up).
|
||||
- **coredns** — plugin-based; geo-split is DIY via plugins; no DNSSEC
|
||||
validation by default (not deeply verified for this write-up).
|
||||
|
||||
### 3.11 Other popular options (survey supplement, 2026-08-12)
|
||||
|
||||
Follow-up survey of additional popular solutions not covered above, evaluated
|
||||
against this LAN's constraints (fail-open, keep DNS on `.36`, DNSSEC goal).
|
||||
None of these change the §4/§5 recommendation.
|
||||
|
||||
**Encrypted-forwarder micro-tools (AGH downstream options, not replacements):**
|
||||
|
||||
- **dnscrypt-proxy** — the classic OpenWrt encrypted forwarder with
|
||||
China-list support and DNS-stamp routing. No in-process DNSSEC validation and
|
||||
no filtering UI; overlaps with AGH's own DoH upstream layer, so its marginal
|
||||
value here is low.
|
||||
- **dnsproxy** (AdGuardTeam) — lightweight DoH/DoT/DoQ forwarder/server.
|
||||
Functionally a subset of AGH's upstream layer; only useful if forwarding logic
|
||||
is deliberately split out of AGH.
|
||||
- **Stubby** — dnsmasq→stubby→DoT (privacy-community pattern). Pure
|
||||
forwarding, no split/filter/validation; adopting it alone would be a
|
||||
downgrade from AGH.
|
||||
|
||||
**Managed / cloud DNS (zero-ops, not self-hosted):**
|
||||
|
||||
- **NextDNS / ControlD / AdGuard DNS / Cloudflare** — hosted filtering, logs,
|
||||
per-device policies. This LAN already self-hosts AGH + a private
|
||||
`adg.chans.xyz` fallback, so a cloud service would be a downgrade in control
|
||||
(data leaves the LAN). Only realistic use: add one as an extra foreign-path
|
||||
upstream inside AGH's `upstream_dns_file`.
|
||||
|
||||
**Heavier all-in-one resolvers:**
|
||||
|
||||
- **PowerDNS Recursor** — real DNSSEC validation + Lua policy, authoritative
|
||||
and recursive in one. Capable but overlaps unbound; over-provisioned here.
|
||||
- **BIND9** — classic authoritative/recursive; can validate DNSSEC and, more
|
||||
interestingly, serve as a local **authoritative zone** that would close the
|
||||
private-PTR gap. As a LAN resolver it lacks AGH's filtering/UI and is heavier
|
||||
to operate; a small dnsmasq authoritative zone is a lighter way to achieve the
|
||||
PTR goal (still deferred until a local authoritative source exists).
|
||||
- **hickory-dns / trust-dns** (Rust) — emerging recursive resolver, DNSSEC
|
||||
friendly, smaller ecosystem/ops track record than unbound/blocky; not yet
|
||||
worth switching for this LAN.
|
||||
|
||||
**Popular stack patterns (structure, not new software):**
|
||||
|
||||
- **Pi-hole + unbound** — the most common global self-hosted combo
|
||||
(filtering front-end + validating backend). AGH already occupies the
|
||||
Pi-hole role here (and does more), so the equivalent is **AGH + unbound/
|
||||
blocky** — exactly the report's recommendation #2.
|
||||
- **dnsmasq + china-list + smartdns** (classic OpenWrt trio) — routes the
|
||||
China list on the gateway itself. Equivalent to co-locating DNS with the
|
||||
proxy host (`.1`), which violates the fail-open requirement; not recommended
|
||||
for this LAN.
|
||||
|
||||
Verdict: the survey adds no better candidate. dnsproxy/dnscrypt-proxy duplicate
|
||||
AGH's upstream layer, cloud DNS is a control downgrade, and the only genuinely
|
||||
new capability (a local authoritative source for PTR) is better served by a
|
||||
small dnsmasq authoritative zone than by replacing the resolver.
|
||||
|
||||
---
|
||||
|
||||
## 4. Architecture recommendation for this LAN
|
||||
|
||||
### 4.1 Preferred architecture (change is config-only)
|
||||
|
||||
```
|
||||
clients (DHCP option 6 = .36)
|
||||
│ UDP/TCP :53
|
||||
▼
|
||||
AGH .36 (filtering, rewrites, query log, per-client upstreams)
|
||||
│ upstream_dns_file:
|
||||
│ [/cn-domain-list/] dns.alidns.com doh.pub ← CN CDN domains (China list)
|
||||
│ default: https://adg.chans.xyz/dns-query … ← trusted/foreign path
|
||||
└→ validating resolver (unbound OR blocky) for the foreign path (optional phase 2)
|
||||
```
|
||||
|
||||
- Front = AGH stays the single LAN DNS box (filtering/rewrites/UI/query log
|
||||
are its strong suit and are already operating).
|
||||
- Split = AGH per-domain upstreams fed by a converted dnsmasq-china-list; no
|
||||
new daemon. This is the documented AGH pattern
|
||||
([upstreams from file](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-from-file)).
|
||||
- Validation = add a validating resolver behind AGH for the trusted path
|
||||
(blocky simplest; unbound most battle-tested) and re-run the known-bad-signature
|
||||
check that failed before; then flip `enable_dnssec`.
|
||||
|
||||
### 4.2 Why not the alternatives as front-ends
|
||||
|
||||
- **smartdns / chinadns-ng as the LAN resolver**: they are pure splitters —
|
||||
no adblock layer, no query log/UI, no DNSSEC. Replacing AGH with either is a
|
||||
capability downgrade; behind AGH they duplicate AGH's built-in split while
|
||||
adding a daemon and losing validation. Only chinadns-ng's IP-verdict mode is
|
||||
genuinely beyond AGH, and it needs kernel ipset/nftset plumbing.
|
||||
- **mosdns as the AGH backend**: viable splitter, but no validation and its
|
||||
cache strips EDNS0/DO ([cache plugin](https://irine-sistiana.gitbook.io/mosdns-wiki/mosdns-v5/ru-he-pei-zhi-mosdns/ke-zhi-xing-cha-jian.md)),
|
||||
which fights the DNSSEC goal. It is already idle on the box — configure it
|
||||
deliberately or remove it.
|
||||
- **mihomo/sing-box DNS as the resolver of record**: fail-closed + proxy-coupled
|
||||
+ no validation. Keep as proxy-side concern (§3.8).
|
||||
- **knot-resolver / dnsdist / Technitium**: capable but over-provisioned;
|
||||
Technitium is the only one that would *replace* AGH wholesale, and there's no
|
||||
benefit worth the migration here.
|
||||
|
||||
### 4.3 Deployment location
|
||||
|
||||
- **Keep DNS on `dns.windy.lan` (.36)**. It is already the DHCP-advertised
|
||||
resolver; it is a separate VM from the proxy host; DNS therefore stays
|
||||
independent of OpenClash state (fail-open), which is an explicit requirement.
|
||||
- **Do not move it to `gfw` (.1)**: the gateway is where OpenClash injects
|
||||
TPROXY/fake-ip/DNS-hijack rules; co-locating LAN DNS there couples DNS to the
|
||||
proxy and its restart/update lifecycle.
|
||||
- A standalone resolver VM adds nothing: both current VMs already sit on the
|
||||
same PVE hypervisor ([lan-overview.md](../lan-overview.md) §Positioning facts),
|
||||
so a hypervisor outage takes out either placement equally; a second physical
|
||||
host for HA is out of scope for a home LAN.
|
||||
- If you ever run a validating resolver + AGH on `.36`, verify outbound from
|
||||
`.36` to the foreign upstreams is not re-hijacked by OpenClash (loop check
|
||||
already mandated in the [AGH review](../adguard-home-official-review-2026-08.md)).
|
||||
|
||||
### 4.4 Fail-open, DNSSEC, private names — by candidate
|
||||
|
||||
| Concern | How the recommended stack behaves |
|
||||
|---|---|
|
||||
| Fail-open when proxy/subscription down | AGH forwards directly to DoH upstreams; `.36`'s outbound is not forced through the proxy in normal ops (no TUN policy routing on `.36` — [dns host facts](../../hosts/dns.windy.lan.md)). With unbound/blocky behind, foreign resolution recurses/validates directly, independent of OpenClash. Avoid mihomo-DNS-as-resolver, which is proxy-coupled. |
|
||||
| DNSSEC validation | Only unbound, blocky, knot-resolver, Technitium validate in-process. AGH sets DO only; mosdns/smartdns/chinadns-ng/mihomo/sing-box do not. Plan: validate behind AGH, or accept "validating public upstream" (confirm with `dig +dnssec`/known-bad test). |
|
||||
| Private names / rewrites | AGH `rewrites` (already in use for `hass.windy.lan`) + `local_ptr_upstreams` once a local PTR source exists. blocky: `customDNS` mapping/rewrite + hosts. unbound: `local-zone`. All adequate. |
|
||||
| Query log / visibility | AGH is the best at this of everything evaluated (14-day anonymized log already configured). |
|
||||
|
||||
---
|
||||
|
||||
## 5. What would make the current AGH + mosdns setup genuinely better
|
||||
|
||||
Concrete, in increasing effort:
|
||||
|
||||
1. **Implement the China-list geo-split in AGH itself**
|
||||
(`upstream_dns_file` + converted `accelerated-domains.china.conf`, default
|
||||
upstreams = trusted foreign path, `fallback_dns` kept). Official AGH docs
|
||||
describe exactly this pattern
|
||||
([loading upstreams from file](https://adguard-dns.io/kb/adguard-home/configuration/#upstreams-from-file));
|
||||
list source: [dnsmasq-china-list](https://github.com/felixonmars/dnsmasq-china-list).
|
||||
Wire a refresh path (cron/ansible) so the list stays current. Re-test CDN
|
||||
resolution and the DNSSEC known-bad domain after.
|
||||
2. **Put a validating resolver on the trusted path** (unbound or blocky), re-run
|
||||
the known-bad-signature check, then enable AGH DNSSEC. Without this, AGH's
|
||||
`enable_dnssec` is only a DO-flag — the exact reason it is currently off
|
||||
([AGH DNSSEC semantics](https://adguard-dns.io/kb/adguard-home/configuration/),
|
||||
[host facts](../../hosts/dns.windy.lan.md)).
|
||||
3. **Either fully configure mosdns (systemd service, sequence, lists) or remove
|
||||
it.** Leaving an idle `127.0.0.1:6052` listener documented as "not the active
|
||||
path" is drift. If kept, plan around no-EDNS0 cache + no validation; if
|
||||
removed, drop the listener and its config to reduce surface.
|
||||
4. **Close the private-PTR gap**: once a local authoritative source exists (e.g.
|
||||
dnsmasq on `gw`, or a tiny authoritative zone), point AGH
|
||||
`local_ptr_upstreams` at it as the AGH review recommends
|
||||
([AGH review](../adguard-home-official-review-2026-08.md));
|
||||
don't set it before that source exists
|
||||
([dns host facts](../../hosts/dns.windy.lan.md)).
|
||||
5. **Optional: ECS** for CDN geo-accuracy — AGH `edns_client_subnet.use_custom`
|
||||
with a coarse fixed prefix (or blocky `ecs.forward`) if measurements show a
|
||||
benefit; note many CN resolvers ignore ECS
|
||||
([AGH ECS](https://adguard-dns.io/kb/adguard-home/configuration/)).
|
||||
|
||||
If the DNS engineering budget is one afternoon, do #1 + #3. If the goal is
|
||||
"real DNSSEC or nothing", do #1 + #2 + #3. Replacing the stack is only
|
||||
justified if you want to abandon AGH's UI/filtering entirely — nothing evaluated
|
||||
here beats it on that axis for this LAN.
|
||||
|
||||
---
|
||||
|
||||
## Caveats / not verified
|
||||
|
||||
- **Live behavior not tested**: all capability claims are from primary docs
|
||||
reviewed 2026-08-12; DNSSEC behavior of `dns.alidns.com`/`doh.pub`/the
|
||||
`adg.chans.xyz` path and mosdns's actual version on `.36` need on-box
|
||||
`dig +dnssec` verification (per [adguard-home-health](../../runbooks/adguard-home-health.md)).
|
||||
- **smartdns DNSSEC**: the official config reference lists no DNSSEC option;
|
||||
if a newer version added one, it is not reflected here
|
||||
([config options](https://pymumu.github.io/smartdns/configuration/)).
|
||||
- **mosdns geosite/geoip**: v5 source tree (fetched 2026-08-12) contains only
|
||||
`domain_set`/`ip_set` data providers; if a `geosite.dat` plugin exists in a
|
||||
release branch, it is not in `main`
|
||||
([plugin/data_provider](https://github.com/IrineSistiana/mosdns/tree/main/plugin/data_provider)).
|
||||
- **mihomo**: the GitHub repo currently shows deliberately obfuscated metadata
|
||||
(see §3.8); capabilities cited from
|
||||
[wiki.metacubex.one](https://wiki.metacubex.one/en/config/dns/).
|
||||
sing-box/knot-resolver/dnsdist license identifiers via GitHub metadata are
|
||||
"other"/custom — treat the specific SPDX ids with caution.
|
||||
- **chinadns2** README was not retrieved (404 on the raw URL); treated as the
|
||||
legacy predecessor of chinadns-ng and not evaluated in depth.
|
||||
- Obsidian/personal notes were not consulted; this is upstream-docs-only.
|
||||
|
||||
## Related docs
|
||||
|
||||
- [lan-overview.md](../lan-overview.md) — full topology (verified 2026-08-06)
|
||||
- [hosts/dns.windy.lan.md](../../hosts/dns.windy.lan.md) — AGH host facts
|
||||
- [hosts/gfw.windy.lan.md](../../hosts/gfw.windy.lan.md) — OpenClash facts
|
||||
- [adguard-home-official-review-2026-08.md](../adguard-home-official-review-2026-08.md) — prior AGH config review
|
||||
- [runbooks/adguard-home-health.md](../../runbooks/adguard-home-health.md)
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
**状态:** 规划文档(未采购、未接线、未改生产配置)。
|
||||
**重要变更(2026-08-09):** **SE5420 已采购**,网络升级改为「保留 ER-X + SE5420 核心」路径——
|
||||
实施与验证以 [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) 为准。
|
||||
实施与验证以 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 为准。
|
||||
本文保留为「ER-X 网关未来替换为 RB5009」的备选方案;其中 PVE 透传调研与 VLAN10 实现方法仍适用。
|
||||
|
||||
---
|
||||
@@ -389,9 +389,9 @@ logread -e netifd
|
||||
|
||||
## 8. 参考
|
||||
|
||||
- 现网地图:[lan-overview.md](lan-overview.md)
|
||||
- ER-X 现状:[edgerouter-x-configuration.md](edgerouter-x-configuration.md)、[hosts/gw.md](../hosts/gw.md)
|
||||
- UniFi:[unifi-network.md](unifi-network.md)、[hosts/ubnt.md](../hosts/ubnt.md)
|
||||
- `gfw`:[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md)
|
||||
- 作废方案(**不再实施**):[lan-erx-se5420-network.md](lan-erx-se5420-network.md)、[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
|
||||
- 现网地图:[lan-overview.md](../lan-overview.md)
|
||||
- ER-X 现状:[edgerouter-x-configuration.md](../edgerouter-x-configuration.md)、[hosts/gw.md](../../hosts/gw.md)
|
||||
- UniFi:[unifi-network.md](../unifi-network.md)、[hosts/ubnt.md](../../hosts/ubnt.md)
|
||||
- `gfw`:[hosts/gfw.windy.lan.md](../../hosts/gfw.windy.lan.md)
|
||||
- 作废方案(**不再实施**):[lan-erx-se5420-network.md](../lan-erx-se5420-network.md)、[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
|
||||
- MikroTik RB5009 官方:<https://mikrotik.com/product/rb5009ug_s_in>、RouterOS v7 手册
|
||||
@@ -0,0 +1,74 @@
|
||||
# SE5420 实施评审主张核实(2026-08-10)
|
||||
|
||||
> **核对基准(历史快照):** 本文于 2026-08-10 针对 [lan-se5420-deployment-guide.md](../lan-se5420-deployment-guide.md) 的**评审前版本**(`35577d0`)撰写。该指南自 `ffb37a9`("finalize SE5420 deployment guide per review")起已按本评审修订,当前 `origin/main` 章节已重组:旧 §3.3 → §4.3、旧 §6(gfw)→ §11、旧 §7(SSID)→ §12、旧 §9(验收/IPv6)→ §13 + §11.4。文末「当前指南处理情况」列出各主张的现行状态;实施以部署指南现行为准。
|
||||
|
||||
**范围。** 本文核对对 `lan-se5420-deployment-guide.md` 的评审意见。结论分为
|
||||
“已证实”(规范/一手资料直接支持)、“基本证实”(架构推论成立但仍须读取现场配置)和
|
||||
“需现场核实”(不能仅由文档或产品手册断言)。这不是实施变更,也不替代维护窗前的
|
||||
`uci show firewall`、交换机当前 VLAN 表和 PVE bridge 配置检查。
|
||||
|
||||
## 核实结论
|
||||
|
||||
| 评审主张 | 结论 | 依据与限定 |
|
||||
| --- | --- | --- |
|
||||
| `firewall.ubunt_upg.masq=1` 是错误方向,应在实际出站的 `wan` zone 做 IPv4 NAT | **已证实** | OpenWrt 明确规定 masquerade 是**按出站 zone/interface**控制;`masq` 通常在 `wan`。因此,对 `ubunt_upg → wan` 流量把 `masq` 放在源 zone 不是该需求的正确 zone 语义。若 `wan` 已 masq,不应重复开启;也可用 `masq_src` 只限 `192.168.10.0/24`。见 [OpenWrt firewall configuration](https://openwrt.org/docs/guide-user/firewall/firewall_configuration) 和 [fw4 masq 测试](https://lxr.openwrt.org/source/firewall4/tests/02_zones/02_masq)。 |
|
||||
| `ubunt_upg → wan` 允许所有经 gfw `wan` 可路由的目的地,不等于只上互联网 | **已证实** | `forwarding` 的 `src`/`dest` 是 zone-to-zone 单向许可,未按“Internet”语义区分目标 IP;规则可用 `dest_ip` 限制。故若 gfw 的 `wan` 接在 LAN66 且 ER-X 可路由 LAN55,评审所列 LAN66/LAN55 风险成立。最终可达网段仍须以 gfw 路由表、ER-X 路由/防火墙现场检查为准。见 [OpenWrt forwarding/rule 参考](https://openwrt.org/docs/guide-user/firewall/firewall_configuration)。 |
|
||||
| 不应把 `ubunt_upg.forward` 改为 `ACCEPT`;`forward_policy` 不是必要的标准 zone 选项;匿名 `uci add` 不可重复执行 | **基本证实** | OpenWrt zone 的标准项是 `forward`,forwarding 是独立 section,参考页未定义 `forward_policy`。单独的具名 forwarding 足以允许跨 zone 路径,因此保持 zone 内 `forward=REJECT` 是较小权限配置。匿名 section 每执行一次都会新增一节,这是 UCI 的操作语义;实施应先读现场配置并使用具名 section。 |
|
||||
| VLAN10 必须有显式 IPv6 策略,否则可能绕过仅 IPv4 的 NAT/隔离 | **已证实** | fw4 将 `masq`(IPv4)和 `masq6`(IPv6)分开;forwarding 默认 family 是 `any`。仅写 IPv4 DHCP/NAT/地址规则不能表达 VLAN10 的 IPv6 RA、DHCPv6、路由和过滤策略。是否已经存在可用 IPv6 前缀、以及 OpenClash 是否接管 IPv6,必须现场验证。见 [OpenWrt firewall configuration](https://openwrt.org/docs/guide-user/firewall/firewall_configuration)。 |
|
||||
| 管理 SVI + 默认路由与“不开 SVI/静态路由/一切 L3”矛盾 | **已证实** | 指南(历史版)§3.3 同时要求 VLAN66 `192.168.66.253/24` 与默认路由,又要求不开 SVI/静态路由。TL-SE5420 官方称其为三层交换机,支持静态路由、RIP、DHCP server/relay。准确目标应是:仅保留 VLAN66 管理 L3 interface/默认网关,不给 VLAN55/10 建 L3 interface,且禁用不需要的 L3 服务和跨 VLAN routing。见 [TL-SE5420 官方页](https://www.tp-link.com.cn/product_2899.html?v=specification) 与 [官方安装手册](https://service.tp-link.com.cn/download/202310/TL-SE5420%20V1.0%E5%AE%89%E8%A3%85%E6%89%8B%E5%86%8C%201.0.2.pdf)。 |
|
||||
| 必须明确移除 VLAN1 成员,PVID 变更本身不等于 access-port VLAN membership | **已证实** | PVID/native VLAN 只处理进入端口的未标记帧;access/trunk 的允许 VLAN 列表是独立概念。指南(历史版)§3.3 只列 VLAN66/55 member 和 PVID,未写移除 VLAN1 或 ingress filtering。验收应检查 VLAN1 member、VLAN1 管理 IP、端口允许 VLAN 和 tagged-frame ingress policy。见 [Ubiquiti 对 native/tagged/access/trunk 的定义](https://help.ui.com/hc/en-us/articles/26136855808919-Switch-Port-VLAN-Assignment-Trunk-Access-Ports)(术语与 802.1Q 语义)以及 [Linux bridge VLAN 配置示例](https://www.kernel.org/doc/html/v5.19/networking/dsa/b53.html)(显式 `bridge vlan del ... vid 1`)。TL-SE5420 具体 GUI/CLI 行为仍以其固件手册核验。 |
|
||||
| NAS 不应在未先完成双端 LACP 时同时接两口;LACP 不使单 TCP 流自动达到 5G | **基本证实** | 这是标准二层环路/聚合变更控制结论:没有已协商的 LAG 时,两条同 VLAN 并行链路会构成潜在环路;STP 只能作为保护而非实施方法。官方产品页列出 LACP 相关资料,但本次未取得 TL-SE5420/TrueNAS 对端的精确配置与当前 NAS 连接状态,故“必然环路/双 IP”不能在桌面审阅中断言。单连接吞吐受链路散列限制是 802.3ad 的常见实现特性,应以 NAS 与交换机的 hash policy 和 `iperf3` 实测验收。 |
|
||||
| 非 VLAN-aware 的 PVE `vmbr0` 不提供 VLAN10 的端口级隔离 | **已证实** | PVE 将 bridge 描述为虚拟交换机;VLAN-aware mode 才能给 guest NIC 赋 VLAN tag,或显式 trunk。Linux 内核说明:`vlan_filtering=0` 时 bridge 不考虑 VLAN tag,且默认关闭;开启后才按 MAC **和 VLAN tag**转发及进行严格 VID 检查。因此“共享非 VLAN-aware bridge 可让可控 guest 主动消费 VLAN10,不能作为严格隔离边界”成立。不能仅凭该结论断言每个 guest 必定收到每个单播帧:未知单播/广播会泛洪,已学习的单播会按 FDB 转发。见 [PVE 网络配置](https://pve.proxmox.com/wiki/Network_Configuration) 和 [Linux bridge 文档](https://docs.kernel.org/networking/switchdev.html)。 |
|
||||
| AP VLAN10 tagged frame 在一个普通 untagged LAN66 access path 上会“自动去 tag 并泄漏到 LAN66” | **不成立/需改写** | 802.1Q 的 native VLAN 是对**未标记**流量的 VLAN;tagged VLAN 需被显式允许于 trunk。因此通常的正确表述是:若上游不允许 VLAN10 tag,AP 到 VLAN10 网关/DHCP 的路径不存在,SSID 会成为不可用入口。实际设备的端口模式(包括是否错误地配置为 all/trunk、是否接受 tagged ingress)须现场查看,不能泛称必然去标签。见 [Ubiquiti VLAN 端口定义](https://help.ui.com/hc/en-us/articles/26136855808919-Switch-Port-VLAN-Assignment-Trunk-Access-Ports) 和 [Ubiquiti VLAN troubleshooting](https://help.ui.com/hc/en-us/articles/9592924981911-Virtual-Network-VLAN-Troubleshooting)。 |
|
||||
| 仅保留 SSH 会话不是移动 PVE/ER-X 物理上联时的真正带外回滚路径;应全程保持 SE5420 Console | **已证实** | 这是直接的操作依赖判断:TCP SSH 的承载链路被拔除时会断,不能证明回滚可达。TL-SE5420 官方安装手册确认该机有 Type-C Console,且本仓库指南本身也把恢复出厂流程建立在 Console 上。故应在迁移前接通 Console、标注旧/新端口、逐根迁移并用 MAC 表与链路/错误计数验证。见 [官方安装手册](https://service.tp-link.com.cn/download/202310/TL-SE5420%20V1.0%E5%AE%89%E8%A3%85%E6%89%8B%E5%86%8C%201.0.2.pdf)。 |
|
||||
| “所有设备均不得直连 ER-X”不是避免环路的必要条件 | **已证实** | 环路取决于同一 L2 广播域存在多条并行二层路径,不取决于是否还有一个独立终端直接接 ER-X。应禁止的是一个下级交换机/桥接主机同时形成平行路径。此项仍需以 ER-X switch0 VLAN/bridge 现场配置和实际接线图确认。 |
|
||||
| 性能不应承诺全面 2.5G;同 VLAN 才可能在 SE5420 本地交换超 1G,跨 55/66 与 Internet 受 ER-X/宽带限制 | **已证实** | TL-SE5420 的 2.5G 端口仅提高经其本地二层转发的链路上限;跨子网必须由网关路由,Internet 另受 WAN/PPPoE 约束。产品页确认 16×2.5G + 4×10G SFP+,但 ER-X、NAS、PC、AP 的实际协商速率和 NIC/布线能力必须由 `ethtool`/端口状态及 `iperf3` 验证。见 [TL-SE5420 官方规格](https://www.tp-link.com.cn/product_2899.html?v=specification)。 |
|
||||
|
||||
## 已核对的文档内事实
|
||||
|
||||
现行指南的**历史版本**(`35577d0`)确实包含评审指出的关键文字:旧 §3.3 的管理 IP/默认路由与“不开 SVI/静态路由”;旧 §6 的 `ubunt_upg.masq`、`forward=ACCEPT`、`forward_policy`、匿名 forwarding;旧 §7 只验“不可达 LAN66”;旧 §9 对新增 VLAN10 写“IPv6 行为与升级前一致”。因此上述评审不是对未出现内容的假设。
|
||||
|
||||
但这些内容在 `ffb37a9` 起的修订中已被修正或重组:`ubunt_upg.masq` 与 `forward_policy` 已删除,gfw 防火墙改为 §11(§11.3 第 2 步明确“不要给 `ubunt_upg` zone 加 masq”);VLAN10 IPv6 在 §11.4 显式写为“本阶段不提供”;SVI/L3 边界在 §4.3 第 16 步单列“L3 明确边界检查”。本文按历史快照保留评审结论,读者应以现行部署指南为准。
|
||||
|
||||
本仓库的 `hosts/gfw.windy.lan.md` 还记录 gfw 的 `eth0` 在 LAN66、`eth1` 在 LAN55,故 `ubunt_upg → wan` 的隔离结论应在执行前以当前 `ip route`、`uci show firewall`、`nft list ruleset` 复核,而不能从方案文字直接把规则写死。
|
||||
|
||||
## gfw 现场只读复核(2026-08-10)
|
||||
|
||||
已通过 `ssh -4 root@192.168.66.1` 仅读取配置和运行规则,未修改设备。该结果会改变
|
||||
评审中两项“当前状态”的表述:
|
||||
|
||||
| 现场事实 | 对评审的影响 |
|
||||
| --- | --- |
|
||||
| `wan` zone 已有 `masq='1'`;现有配置另有具名 `ubunt_upg_nat`,运行时渲染为 `oifname "eth0"` 且只匹配 `ip saddr 192.168.10.0/24 masquerade`。 | “必须在 wan 开 masq”的**方向原则**正确,但“当前无 masq”不正确。现有显式 SNAT 已在实际出 `eth0` 时执行;计划中再将 `masq` 加到 `ubunt_upg` 仍是多余且方向错误。 |
|
||||
| 当前放行是具名 `ubunt_upg_to_lan`,不是 `ubunt_upg→wan`;其运行链先拒绝 `192.168.66.0/24`,再允许到 `lan`。gfw 的 IPv4 default route 是 `192.168.66.254`。 | 计划新增 `ubunt_upg→wan` 会是与当前设计不同、过宽的改动。现有 LAN66 阻断规则在该链中先匹配;但对经 ER-X 可达的 LAN55/其他内网仍没有显式拒绝,故隔离评审的**剩余风险成立**。应以明确内网前缀 deny + 所需外网 allow 重写,而不是加 WAN forwarding。 |
|
||||
| `ubunt_upg` DHCPv6 和 RA 都是 `disabled`;运行路由表仅有各接口的 IPv6 link-local route,没有 IPv6 default route;全局 IPv6 forwarding 是 `1`。 | 评审“VLAN10 未明确 IPv6 策略”的表述对计划文本仍成立,但“IPv6 可能立即绕过”的事实判断在当前状态**未获证实**:现有 RA/DHCPv6 已关闭且无 IPv6 默认路由。实施文档仍应把这项显式写为“IPv6 不提供”,并在启用前复查。 |
|
||||
| 系统是 ImmortalWrt **25.12.0**,`/usr/bin/apk` 存在(apk-tools 3.0.5)。 | 评审中“ImmortalWrt 21.02.5 应使用 opkg”的版本判断错误/过时;在本机上 `apk add tcpdump` 是可用包管理器。仍应先检查软件包可用性,避免在维护文档中把两种命令并列为未经验证的替代方案。 |
|
||||
|
||||
这些命令输出未含凭据、令牌或私钥,故仅记录了安全相关的摘要;不将完整防火墙快照提交至仓库。
|
||||
|
||||
## 当前指南处理情况(2026-08-13 核对)
|
||||
|
||||
对 `origin/main`(`2fd354c`)逐项核对评审主张:
|
||||
|
||||
| 评审主张 | 现行状态 | 现行位置 |
|
||||
| --- | --- | --- |
|
||||
| `ubunt_upg.masq=1` 方向错误 | ✅ 已修复 | §11.3 第 2 步「不要给 `ubunt_upg` zone 加 masq」 |
|
||||
| `ubunt_upg→wan` 不等于只上互联网 | ⚠️ 原则成立,指南已禁止新增宽泛 forwarding;LAN55/RFC1918 显式 deny 仍为待办 | §11.1b「待补缺口」、§11.3 |
|
||||
| 勿改 `forward=ACCEPT`;无 `forward_policy`;匿名 uci 不可重复 | ✅ 已修复 | §11.3 第 1 步保持 REJECT;指南已无 `forward_policy` |
|
||||
| VLAN10 须显式 IPv6 策略 | ✅ 已修复 | §11.4「本阶段不提供 VLAN10 IPv6」 |
|
||||
| 管理 SVI + 默认路由 vs「不开一切 L3」矛盾 | ✅ 已消解 | §4.3 第 16 步「L3 明确边界检查」 |
|
||||
| 须移除 VLAN1 成员;PVID≠membership | ✅ 指南已加强;现网仍偏离(W1N-54) | §4.3 第 9–12 步;VLAN1 不可删说明 |
|
||||
| NAS 双口未 LACP 前勿并行 | ✅ 已体现 | §7 第 7 步「仅口 8,口 12 断开」 |
|
||||
| 非 VLAN-aware PVE bridge 不能作隔离边界 | ✅ 已体现 | §9.2 要求 VLAN-aware + `bridge-vids` |
|
||||
| AP tagged 帧在 access 口自动去 tag | ✅ 本文已纠正(不成立) | — |
|
||||
| SSH 非真正带外;须 Console | ✅ 已体现 | 开头第 2 条、§4.1、§16 |
|
||||
| 「所有设备不得直连 ER-X」非必要 | ✅ 已体现 | 全程三条第 1 条、§7 |
|
||||
| 不应承诺全面 2.5G | ✅ 已体现 | §8 第 7 条、§14 |
|
||||
|
||||
## 实施前的最低限度现场证据
|
||||
|
||||
1. gfw:保存并审阅 `uci show firewall`、`ip route`、`ip -6 route`、`nft list ruleset`;确认 wan 的 masq 与所有 WAN→内网、VLAN10→内网匹配次序。
|
||||
2. SE5420 Console:导出/截图 VLAN1、55、66 member 和 PVID/ingress-filter 状态;确认唯一管理 L3 interface 和路由/relay/DHCP 状态。
|
||||
3. PVE:记录 `/etc/network/interfaces`、VM NIC VLAN tags 和 `bridge vlan show`,再决定是否把 VLAN-aware 改造另开窗口。
|
||||
4. AP:从实际设备 `info` 或控制器记录确认 Inform URL(本仓库目前记录 `http://192.168.66.46:9080/inform`),并验证 VLAN10 tag 只经 U6/PVE trunk。
|
||||
5. NAS:单网口稳定后,另窗配置并验证两端 LACP,第二根线最后插入;用多流及单流 `iperf3` 分开验收。
|
||||
@@ -0,0 +1,345 @@
|
||||
# Home Assistant × Matrix integration
|
||||
|
||||
Reference for wiring the Home Assistant [Matrix integration](https://www.home-assistant.io/integrations/matrix)
|
||||
to the self-hosted Matrix homeserver at [`synapse.chans.xyz`](../hosts/synapse.chans.xyz.md).
|
||||
Deliberately contains no Matrix passwords, access tokens, or room encryption material.
|
||||
|
||||
> **Status (2026-08-15, W1N-139):** the built-in `matrix` integration has been
|
||||
> **retired** on `hass.windy.lan` and replaced by the custom **`matrix_e2ee`**
|
||||
> integration. The sections below on the built-in integration are kept for
|
||||
> reference only. See [matrix_e2ee](#matrix-e2ee-custom-e2e-integration) for the
|
||||
> active setup and [Device verification (SAS) model](#device-verification-sas-model)
|
||||
> for how device trust works.
|
||||
|
||||
## Purpose
|
||||
|
||||
The integration lets Home Assistant send messages to Matrix rooms and react to
|
||||
messages/reactions in Matrix rooms. "Reacting" is done by firing a
|
||||
`matrix_command` event when one of the configured commands matches; automations
|
||||
then trigger on that event. Sending is done through the `notify.matrix` platform
|
||||
and the `matrix.send_message` / `matrix.react` actions.
|
||||
|
||||
## Environment mapping
|
||||
|
||||
| Integration setting | This deployment |
|
||||
|---|---|
|
||||
| `homeserver` | `https://synapse.chans.xyz` (client-server base URL) |
|
||||
| `username` | full Matrix ID, e.g. `@ha_bot:chans.xyz` |
|
||||
| `password` | MAS local-password account password (see below) |
|
||||
| Room IDs / aliases | full forms with the identity domain, e.g. `!cUrbafjkfsMDVwdRDQ:chans.xyz` or `#room:chans.xyz` |
|
||||
|
||||
- Identity domain is `chans.xyz` (not `synapse.chans.xyz`); user IDs and room
|
||||
aliases carry the `:chans.xyz` suffix.
|
||||
- Authentication on this homeserver is MAS (Matrix Authentication Service) with
|
||||
local-password accounts. The integration logs in with `m.login.password`
|
||||
(username + password), so the bot account must be a local-password account —
|
||||
same as the Hermes account documented in [`hermes-matrix.md`](hermes-matrix.md).
|
||||
If MAS is later switched to OAuth2/OIDC-only (no legacy password login), the
|
||||
integration's password login will stop working; keep that in mind before such
|
||||
a change.
|
||||
- Public registration is disabled. Create/reset the dedicated bot account via
|
||||
MAS / Element Admin.
|
||||
|
||||
## Use a separate bot account (mandatory)
|
||||
|
||||
The docs are explicit: to prevent infinite loops when reacting to commands,
|
||||
the integration **must** use a separate account from any account whose messages
|
||||
it reacts to. Use a dedicated account such as `@ha_bot:chans.xyz`, not a human
|
||||
account.
|
||||
|
||||
## configuration.yaml (example)
|
||||
|
||||
```yaml
|
||||
# The Matrix integration
|
||||
matrix:
|
||||
homeserver: https://synapse.chans.xyz
|
||||
username: "@ha_bot:chans.xyz"
|
||||
password: supersecurepassword
|
||||
rooms:
|
||||
- "#hasstest:chans.xyz"
|
||||
commands:
|
||||
- word: my_command
|
||||
name: my_command
|
||||
```
|
||||
|
||||
After changing `configuration.yaml`, restart Home Assistant to apply the
|
||||
changes. The integration then shows under **Settings → Devices & services**;
|
||||
its entities are on the integration card and the Entities tab.
|
||||
|
||||
### Configuration variables
|
||||
|
||||
| Variable | Meaning |
|
||||
|---|---|
|
||||
| `username` | Full Matrix ID the bot logs in as, e.g. `@ha_bot:chans.xyz`. The `@` has a special YAML meaning, so always quote it. |
|
||||
| `password` | The bot account's password (MAS local password). |
|
||||
| `homeserver` | Full client-server URL of the homeserver. |
|
||||
| `rooms` | Rooms the bot should join and listen in. List **all** rooms commands are to be received in, even if a command scopes itself to fewer rooms. Accepts internal room ID (`!…:chans.xyz`) or alias (`#room:chans.xyz`). |
|
||||
| `commands` | Commands to listen for. Each fires a `matrix_command` event when triggered. |
|
||||
|
||||
### Command types
|
||||
|
||||
| Key | Triggers when |
|
||||
|---|---|
|
||||
| `word` | A message starts with `!<word>`. Arguments after the word are captured as a list in the event's `data`. |
|
||||
| `expression` | A message matches the Python regexp. The regexp group dictionary is captured in the event's `data`. |
|
||||
| `reaction` | A message is reacted to with the given emoji. |
|
||||
| `name` | The command name, exposed as an attribute of the fired event. |
|
||||
|
||||
A command can be scoped to specific rooms with a per-command `rooms` list (the
|
||||
room must still be listed under the top-level `rooms`).
|
||||
|
||||
## Event data
|
||||
|
||||
When a command triggers, a `matrix_command` event fires with:
|
||||
|
||||
- `name` — the command name.
|
||||
- `data` — for `word` commands, a list of arguments (everything after the word,
|
||||
split on spaces); for `expression` commands, the group dictionary of the
|
||||
matching regexp.
|
||||
- `event_id` — the received message's identifier.
|
||||
- `thread_parent` — the root message ID of the thread; equals `event_id` when
|
||||
the message is not inside a thread.
|
||||
|
||||
## Notifications (notify.matrix)
|
||||
|
||||
Deliver notifications from Home Assistant to a Matrix room (direct or group):
|
||||
|
||||
```yaml
|
||||
notify:
|
||||
- name: matrix_notify
|
||||
platform: matrix
|
||||
default_room: "#hasstest:chans.xyz"
|
||||
```
|
||||
|
||||
- The target room must already exist; get its canonical ID from the room
|
||||
settings dialog (`!<randomid>:chans.xyz`) or an alias (`#roomname:chans.xyz`).
|
||||
Quote the room ID/alias in YAML to escape the `!` / `#` characters.
|
||||
- The notifying account may need to be invited to the room, depending on room
|
||||
policy.
|
||||
|
||||
Message formats (`data.format`): `text` (default) and `html`. Images can be
|
||||
attached via `data.images` (list of file paths); files from outside allowed
|
||||
folders require `homeassistant.allowlist_external_dirs` to list the source
|
||||
folder.
|
||||
|
||||
Reply inside a thread by passing the root message ID into `data.thread_id`:
|
||||
|
||||
```yaml
|
||||
action: notify.matrix_notify
|
||||
data:
|
||||
message: "Reply message goes here"
|
||||
data:
|
||||
thread_id: "{{ trigger.event.data.thread_parent }}"
|
||||
```
|
||||
|
||||
## Actions
|
||||
|
||||
- `matrix.react` — send a reaction to a message in a Matrix room
|
||||
(`reaction`, `room`, `message_id`).
|
||||
- `matrix.send_message` — send a message to one or more Matrix rooms.
|
||||
|
||||
## Comprehensive example (adapted)
|
||||
|
||||
```yaml
|
||||
matrix:
|
||||
homeserver: https://synapse.chans.xyz
|
||||
username: "@ha_bot:chans.xyz"
|
||||
password: supersecurepassword
|
||||
rooms:
|
||||
- "#hasstest:chans.xyz"
|
||||
- "#someothertest:chans.xyz"
|
||||
commands:
|
||||
- word: testword
|
||||
name: testword
|
||||
rooms:
|
||||
- "#someothertest:chans.xyz"
|
||||
- expression: "My name is (?P<name>.*)"
|
||||
name: introduction
|
||||
- reaction: 👍
|
||||
name: thumbsup
|
||||
|
||||
notify:
|
||||
- name: matrix_notify
|
||||
platform: matrix
|
||||
default_room: "#hasstest:chans.xyz"
|
||||
|
||||
automation:
|
||||
- alias: "Respond to !testword"
|
||||
triggers:
|
||||
- trigger: event
|
||||
event_type: matrix_command
|
||||
event_data:
|
||||
command: testword
|
||||
actions:
|
||||
- action: notify.matrix_notify
|
||||
data:
|
||||
message: "It looks like you wrote !testword"
|
||||
```
|
||||
|
||||
## matrix_e2ee (custom E2E integration)
|
||||
|
||||
Custom integration [`windyboy/ha-matrix-e2ee`](https://github.com/windyboy/ha-matrix-e2ee),
|
||||
release **v0.3.12** (Matrix activity events + push diagnostics), deployed on
|
||||
`hass.windy.lan` 2026-08-20 (upgraded from v0.3.2, W1N-182/#34 emoji-wait
|
||||
wizard fix; v0.3.9 brought the Connection health binary sensor, SAS/command
|
||||
allowlist split, URL normalization and single-entry enforcement, W1N-156/W1N-190).
|
||||
Runs a dedicated bot with a **persistent E2EE device identity**.
|
||||
|
||||
- Domain `matrix_e2ee`; Config Flow (UI) with YAML import migration, not in HACS. Does **not**
|
||||
override the built-in `matrix` integration.
|
||||
- Dependencies are declared **explicitly** in `manifest.json` to work around Home
|
||||
Assistant's `is_installed` dropping the `[e2e]` extra (W1N-140):
|
||||
`matrix-nio[e2e]==0.26.0` + `vodozemac` + `peewee` + `cachetools` + `atomicwrites`.
|
||||
- **v0.2.0 migration:** YAML `matrix_e2ee:` block was auto-imported into a Config Entry
|
||||
(`source: import`) on first startup, then removed. All settings now managed via
|
||||
**Settings → Devices & Services → Matrix E2EE → Configure**.
|
||||
See [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) for the deployed state.
|
||||
|
||||
### Services & events
|
||||
|
||||
- Services (all admin-only since v0.1.4):
|
||||
- `send_message` (`message`, `room_id`)
|
||||
- `start_verification` (`user_id`, `device_id`)
|
||||
- `confirm_verification` (`transaction_id`)
|
||||
- `cancel_verification` (`transaction_id`)
|
||||
- `reauthenticate` (`password`) — soft-logout only
|
||||
- `get_fingerprint` (no fields; returns bot's own `ed25519`/`curve25519` keys; added v0.1.3)
|
||||
- `verify_device_by_fingerprint` (`user_id`, `device_id`, `ed25519`; added v0.1.3,
|
||||
renamed from `verify_device` in v0.1.4; requires exact `ed25519` match)
|
||||
- Events:
|
||||
- `matrix_e2ee_command` (`room_id`, `sender`, `command`, `args` only —
|
||||
never the raw body)
|
||||
- `matrix_e2ee_error` (codes, no secrets)
|
||||
- `matrix_e2ee_verification` (`stage`, `transaction_id`, `user_id`, `device_id`,
|
||||
optional `emojis`, optional `expires_at`; `expires_at` added v0.1.3)
|
||||
- `matrix_e2ee_fingerprint` (`user_id`, `device_id`, `ed25519`, `curve25519` —
|
||||
public keys only; added v0.1.3)
|
||||
- `matrix_e2ee_message_received` (`room_id`, `sender`, `event_id`; added v0.3.12
|
||||
activity events)
|
||||
- `matrix_e2ee_verification_done` (`transaction_id`, `user_id`, `device_id`;
|
||||
added v0.3.12)
|
||||
- v0.3.12 also adds an `event.` platform entity (`Bot activity`,
|
||||
`event_types: ["message", "command", "verification_done"]`) and a diagnostic
|
||||
Connection binary sensor (`binary_sensor.*_connection`, CONNECTIVITY class).
|
||||
- `notify.matrix_e2ee` is **not implemented** (upstream deferred) — notifications
|
||||
must call `matrix_e2ee.send_message` (message + room_id).
|
||||
- Commands fire Home Assistant events only; the integration never calls
|
||||
`domain.service` itself. Map commands in automations.
|
||||
- Encrypted rooms fail-closed on unverified devices.
|
||||
- Since v0.1.4: `start_verification`, `confirm_verification`, `cancel_verification`,
|
||||
`verify_device_by_fingerprint`, and `reauthenticate` are enforced as HA admin-only
|
||||
via `async_register_admin_service`; non-admin users cannot call them.
|
||||
|
||||
### Storage & recovery
|
||||
|
||||
- `.storage/matrix_e2ee_session.json` (`user_id`, `device_id`, `access_token`,
|
||||
`pickle_key`) and `.storage/matrix_e2ee_store/` (Olm/Megolm, device trust,
|
||||
sync token). Both stay on the HA persistent volume and are in HA backups.
|
||||
- Soft logout → `matrix_e2ee.reauthenticate` (keeps `device_id` + crypto store;
|
||||
rejected outside soft-logout state since v0.1.3).
|
||||
- Hard logout / store loss → delete session + store, restart with password, re-SAS
|
||||
(a **new device**; old history not decryptable).
|
||||
|
||||
## Device verification (SAS + fingerprint) model
|
||||
|
||||
Researched 2026-08-15 (W1N-139 stage-6 pre-study), updated for v0.1.3/v0.1.4.
|
||||
Sources: matrix.org
|
||||
[cross-signing guide](https://matrix.org/docs/guides/implementing-more-advanced-e-2-ee-features-such-as-cross-signing/),
|
||||
matrix-nio [examples](https://matrix-nio.readthedocs.io/en/latest/examples.html),
|
||||
[element-android#6832](https://github.com/vector-im/element-android/issues/6832),
|
||||
Element [device-verification](https://element.io/features/device-verification).
|
||||
|
||||
`matrix_e2ee` supports three verification paths (the wizard — v0.3.0
|
||||
bot-initiated, reworked in v0.3.1/v0.3.2 to wait for a peer-initiated inbound
|
||||
SAS from the user's Matrix client with emoji comparison — automates the SAS
|
||||
flow):
|
||||
|
||||
### 1. SAS (mutual, manual confirmation since v0.1.4)
|
||||
|
||||
- SAS is device-to-device: exchange ephemeral keys → derive emojis → **a human on
|
||||
each side compares and confirms** (`m.key.verification.mac`).
|
||||
- Matrix distinguishes two cases (spec uses *should*, not *must*):
|
||||
- **same user, two devices** → to-device messages (SAS);
|
||||
- **two different users** → **in-room (DM) messages**, verifying the *user*
|
||||
(cross-signing master key), not a specific device.
|
||||
- Cross-signing: each user has master / self-signing / user-signing keys. A device
|
||||
looks "verified" to another user via the chain
|
||||
`my master → my user-signing → their master → their self-signing → their device`.
|
||||
- Element's "Verify" button only starts **in-DM user verification**; it has no
|
||||
"verify a specific device of another user via to-device" flow (matrix.org
|
||||
recommends hiding per-device verification for other users).
|
||||
- `matrix_e2ee` implements **raw to-device device SAS** (`start_verification`/
|
||||
`confirm_verification`), **no cross-signing / in-room**. This is a non-standard
|
||||
cross-user path: works with matrix-nio + Element Web/Desktop (reported in
|
||||
element-android#6832), **not** on Element Android/X.
|
||||
- **v0.1.3**: inbound SAS auto-complete was added; SAS events include `expires_at`.
|
||||
- **v0.1.4 (breaking)**: auto-confirm was removed. **Every** device — including
|
||||
another device of the bot's own account — requires explicit `confirm_verification`
|
||||
after emoji comparison. Only the bot's own account or users in `allowed_users`
|
||||
may initiate SAS (`verification_peer_denied` otherwise).
|
||||
- **v0.2.1**: storage I/O moved off the event loop (`asyncio.to_thread`,
|
||||
W1N-167); own-keys query on startup so inbound SAS can build a session (W1N-166).
|
||||
- **v0.2.2** (not deployed): intermediate version.
|
||||
- **v0.2.3**: sync loop runs as a background task (fixes bootstrap setup timeout,
|
||||
W1N-168); SAS double-send of key and MAC fixed (W1N-169).
|
||||
- **v0.2.6**: `_log_verification_state()` tracks SAS state transitions with
|
||||
`async_write_ha_state` for diagnosis (W1N-174);
|
||||
`_bridge_verification_request()` handles inbound
|
||||
`m.key.verification.request` → `m.key.verification.ready` since nio lacks a
|
||||
`request` framework (W1N-173).
|
||||
- **v0.2.5**: bridge `m.key.verification.request` → `ready` (nio lacks
|
||||
request framework, W1N-173).
|
||||
- **v0.2.4**: `_patch_nio_sas_timeout()` works around nio 0.26.0
|
||||
`_last_event_time` bug (SAS timed out at 60s regardless of activity — now uses
|
||||
`_max_age` 5 min); `_repair_dropped_start()` recovers SAS `start` events nio
|
||||
dropped when the peer device was unknown (W1N-170/W1N-172);
|
||||
`VERIFICATION_TIMEOUT_SECONDS` 600→240 (fires before nio's `_max_age`).
|
||||
- **v0.2.11**: `receive_mac_event` no longer overrides canceled state (W1N-179/#31).
|
||||
- **v0.3.12**: Matrix activity events (`matrix_e2ee_message_received`,
|
||||
`matrix_e2ee_verification_done`) + `event.` Bot activity entity + Connection
|
||||
diagnostic binary sensor.
|
||||
- **v0.3.9**: SAS driver gate split from the command allowlist — new
|
||||
`verification_peer_users` option (W1N-156/#41); SAS/sync logs demoted
|
||||
warning→info/debug (W1N-188/#38); Connection health binary sensor
|
||||
(W1N-185/#40); URL normalization + single-entry enforcement (W1N-190/#42).
|
||||
- **v0.3.8**: `m.key.verification.done` handshake completion for
|
||||
request-based SAS (W1N-183/#35).
|
||||
- **v0.3.2**: wizard waits for the inbound SAS to show emojis before moving
|
||||
to the compare step (`_wait_for_inbound` requires `latest_sas_snapshot()` to
|
||||
return `emojis`) — W1N-182/#34.
|
||||
- **v0.3.1**: verification wizard now waits for a peer-initiated inbound SAS
|
||||
(options flow no longer starts verification from the bot; `latest_sas_snapshot()`
|
||||
skips verified/canceled transactions) — GitHub #33.
|
||||
- **v0.3.0**: bot-initiated device verification wizard (W1N-180/#32).
|
||||
- Inbound SAS is gated to `allowed_users` (v0.1.3); **since v0.3.9 (W1N-156)
|
||||
the gate is the separate `verification_peer_users` allowlist**, which is
|
||||
unset on hass.windy.lan — only the bot's own account may drive SAS until
|
||||
`@zhiqiang:chans.xyz` is added there.
|
||||
|
||||
### 2. One-sided fingerprint (added v0.1.3, hardened v0.1.4)
|
||||
|
||||
- Call `matrix_e2ee.get_fingerprint` to get the bot's own `ed25519` device key
|
||||
(read it from the `matrix_e2ee_fingerprint` event).
|
||||
- In Element, open the bot user's sessions and use "Manually verify by text".
|
||||
Compare the session key with the fingerprint.
|
||||
- To trust another device from the bot's side, call
|
||||
`matrix_e2ee.verify_device_by_fingerprint` with the peer's `user_id`, `device_id`,
|
||||
and `ed25519` key. The match is exact (since v0.1.4's rename from `verify_device`).
|
||||
Feed the **peer** key, not the bot's own key.
|
||||
- This trusts from one side only; the peer still trusts the bot independently.
|
||||
- Both `get_fingerprint` and `verify_device_by_fingerprint` are HA admin-only.
|
||||
|
||||
### Consequence
|
||||
|
||||
Whether `@zhiqiang`'s device can be verified depends on which Element client
|
||||
they use. Open options recorded in W1N-139 (A: Web SAS test; B: upstream
|
||||
in-room/cross-signing; C: unencrypted-room downgrade).
|
||||
|
||||
## References
|
||||
|
||||
- Home Assistant Matrix integration: <https://www.home-assistant.io/integrations/matrix>
|
||||
- Matrix host facts: [`hosts/synapse.chans.xyz.md`](../hosts/synapse.chans.xyz.md)
|
||||
- Matrix deployment and upstream index: [`matrix-upstream.md`](matrix-upstream.md)
|
||||
- Hermes Agent Matrix channel (MAS local-password + access-token pattern): [`hermes-matrix.md`](hermes-matrix.md)
|
||||
- HA host facts: [`hosts/hass.windy.lan.md`](../hosts/hass.windy.lan.md)
|
||||
- HA maintenance runbook: [`runbooks/home-assistant-maintenance.md`](../runbooks/home-assistant-maintenance.md)
|
||||
+188
-60
@@ -1,39 +1,103 @@
|
||||
# 内网 DNS 架构调研与优化建议
|
||||
|
||||
> 状态:2026-08-12 调研,Linear **W1N-56**。基于网络工程师视角,方案待实施评审。
|
||||
> 2026-08-12 现场核查修正:gfw 上 mosdns 已并入 OpenClash DNS 链(作为 clash 的
|
||||
> `nameserver`,DIRECT 规则真实 IP 解析用),**不再闲置**;推荐方案(AGH 前端 +
|
||||
> `.36` 伴生 mosdns 后端)仍待评审落地。
|
||||
> 相关:`docs/lan-overview.md`、`hosts/dns.windy.lan.md`、`hosts/gfw.windy.lan.md`。
|
||||
> 状态:2026-08-12 调研定稿,Linear **W1N-56**。**Phase 0 验证(2026-08-12)全部完成,最终裁决(用户定稿)已对齐**;本期未改动任何生产 DNS 路径。
|
||||
> 2026-08-12 现场核查修正:gfw 上 mosdns **不是闲置**——它是 OpenClash clash 的
|
||||
> `nameserver`/`default-nameserver`(DIRECT 规则真实 IP 解析),处于活动链路,裁决第 5 条
|
||||
> 的"删除闲置 mosdns"前提不成立,处置改为"正式纳管并文档化"。
|
||||
> 相关:`docs/lan-overview.md`、`hosts/dns.windy.lan.md`、`hosts/gfw.windy.lan.md`、W1N-40。
|
||||
> 2026-08-13 (W1N-62):gfw mosdns 国外分支已从"明文国内公网 DNS"改为**加密 DoH**
|
||||
> (自建 `https://adg.chans.xyz/dns-query`,hk2),新增 `foreign_upstream`/`foreign_fallback`
|
||||
> (primary=DoH, secondary=明文国内 DNS, threshold 1000ms),`bootstrap` 用现有国内公网 IP
|
||||
> (防自举循环)。实测:mosdns 平面国外域名 A/AAAA 恢复(google AAAA `2607:f8b0…`)、
|
||||
> `dup.baidustatic.com`→`0.0.0.0`(AGH 拦截保留)、clash 7874 fake-ip 平面不变。
|
||||
> **重要修正(定稿)**:DoH 流量实测为 **gfw→hk2 直连,未经 clash 代理**——nft output 链
|
||||
> (mangle mark/tcp redirect)计数为 0、`/proc/net/tcp` 存在到 hk2:443 的 established
|
||||
> 连接,路由自身 TCP 输出当前并未被 OpenClash 重定向,故"经代理访问加密 DNS"的假设不成立。
|
||||
> **最终决策:接受直连,不强行走代理**——`foreign_upstream` 以自建解析器
|
||||
> `adg.chans.xyz`(hk2)为主力,自有 VPS 直连即可达、无被墙/污染问题、DoH/TLS 已加密、应答干净,
|
||||
> 走代理毫无增益反而把 DNS 平面耦合进 clash;实测 kill clash 期间国外查询 0.01s 正常应答,
|
||||
> watchdog 自动拉起,直连使 DNS 平面独立于代理(优于过代理)。
|
||||
> **多上游冗余(同日)**:`concurrent: 3`,新增 `https://dns.quad9.net/dns-query` 与
|
||||
> `https://dns.cloudflare.com/dns-query`(2026-08-13 本网络实测可达;`dns.quad101.net`
|
||||
> TLS 握手失败已排除)——hk2 故障时仍由干净的国外 DoH 应答,最后才退化明文国内兜底。
|
||||
> 验证:google AAAA 由 hk2 的 `2607:f8b0…` 变为新上游的 `2404:6800…`,国内外/拦截/代理平面无回归。
|
||||
> 备份:`config.yaml.bak-foreign-doh-20260813-103746` / `config.yaml.bak-foreign-doh-20260813-103813` /
|
||||
> `config.yaml.bak-multi-doh-20260813-105421`。
|
||||
|
||||
## 1. 现状(实测)
|
||||
## 1. 现状(实测 2026-08-12)
|
||||
|
||||
| 角色 | 部署 | 职责 | 是否在活动路径 |
|
||||
|------|------|------|----------------|
|
||||
| **AdGuard Home** | `192.168.66.36`(PVE VM 120,Docker host 网络) | EdgeRouter DHCP 通告给 LAN55/66 客户端的 DNS;广告/过滤、查询统计、Web 面板 | ✅ **是** |
|
||||
| **mosdns** | `192.168.66.1`(gfw OpenWrt)监听 `127.0.0.1:6052` | OpenClash custom DNS 的 `nameserver`(DIRECT 规则真实 IP 分流:国内→AGH `.36`,国外→国内公网 DNS `223.5.5.5`/`119.29.29.29`) | ✅ 网关侧(clash 消费,不面向客户端) |
|
||||
| **OpenClash / clash(meta)** | `192.168.66.1`(gfw) | gateway 自身/被劫持流量的 fake-ip + 代理,DNS 走 dnsmasq→clash `#7874` | 仅网关侧 |
|
||||
| **AdGuard Home** | `192.168.66.36`(PVE VM 120,Docker host 网络) | EdgeRouter DHCP 通告给 LAN55/66 客户端的唯一 DNS;广告/过滤、查询统计、Web 面板 | ✅ **是(LAN 客户端唯一入口)** |
|
||||
| **mosdns** | `192.168.66.1`(gfw OpenWrt)监听 `127.0.0.1:6052`(仅本机) | clash 的 `nameserver`/`default-nameserver`:DIRECT 规则真实 IP 分流(国内 → AGH `.36:53`,国外 → `223.5.5.5`/`119.29.29.29`) | ✅ 网关侧(clash 消费,不面向客户端) |
|
||||
| **OpenClash / clash(meta)** | `192.168.66.1`(gfw) | gateway 自身/被劫持流量的 fake-ip + 代理;DNS 走 dnsmasq→clash `#7874` | 仅网关侧与 VLAN10 |
|
||||
|
||||
**关键事实:** LAN 客户端 DNS 直连 `66.36`,**不经过** gfw(EdgeRouter `service dns forwarding` cache 512,通告 `.36`)。所以 gfw/clash 的 fake-ip 分流对"直连 AGH 的客户端"不起作用;gfw 上 mosdns 作为 clash 的 `nameserver` 供 DIRECT 规则连接的真实 IP 解析用(国内→AGH,国外→国内公网 DNS),不面向 LAN 客户端。
|
||||
**关键事实(全部实测):**
|
||||
- LAN 客户端 DNS 直连 `66.36`,**不经过** gfw(EdgeRouter `service dns forwarding` cache 512,通告 `.36`)。
|
||||
- AGH 上游:DoH `dns.alidns.com`(→`223.5.5.5`/`223.6.6.6`)+ `doh.pub`(→`120.53.53.53`/`1.12.12.12`),
|
||||
`upstream_mode: load_balance`、`fastest_timeout: 1s`、`upstream_timeout: 10s`;bootstrap
|
||||
`223.5.5.5`/`223.6.6.6`(公网 IP,无自举循环);兜底 DoH `adg.chans.xyz`(→`hk2.chans.xyz`→`154.36.174.161`)。
|
||||
DNS 监听 UDP/TCP **v4 only(`0.0.0.0:53`)**,无 v6 监听;`enable_dnssec: false`;cache 4 MB;ratelimit 20。
|
||||
- AGH 宿主出网:默认路由 `via 192.168.66.254`(EdgeRouter),**直连,不经 gfw**;DoH 实测可达
|
||||
(223.5.5.5:443 → HTTP 400/0.05s,120.53.53.53 → 502/0.06s,adg.chans.xyz 冷连接 ~3.5s)。
|
||||
- gfw 自身出网:OpenClash `openclash_mangle_output` 对非本地区域/国内 IP 流量统一
|
||||
`mark 0x162 → tproxy 127.0.0.1:7895`(**gfw 自身流量默认走代理**,含 clash 的
|
||||
nameserver-policy DoH);mosdns 的上游(AGH `.36` 本地区域、`223.5.5.5` 国内 IP)均被 bypass,保持直连。
|
||||
- gfw DNS 链:`server=127.0.0.1#7874`(dnsmasq)→ clash:`nameserver: [127.0.0.1:6052]`(mosdns)、
|
||||
`default-nameserver: [127.0.0.1:6052]`、`nameserver-policy` 国外域名 → DoH `https://1.1.1.1/dns-query`、
|
||||
`enhanced-mode: fake-ip`(198.18.0.1/16)、`ipv6: false`;nft 有 UDP/53 hijack → dnsmasq。
|
||||
- **mosdns 配置缺陷(2026-08-12 发现并修复)**:`main` sequence 的国内分支
|
||||
(`matches: qname $domestic_domains → exec: $domestic_upstream`)之后**缺少
|
||||
`matches: has_resp → accept` 守卫**。mosdns v5 的 `sequence` 在 `forward` 成功后不会停止,
|
||||
只有 `accept`/`reject`/`return` 或错误会终止——因此命中 `geosite_cn` 的查询会被转发**两次**
|
||||
(AGH 与 223.5.5.5/119.29.29.29),最终应答来自最后一个 forward(国内公网 DNS),**AGH 的
|
||||
拦截/rewrite 对 DIRECT 国内域名静默失效**。实测证据:`dup.baidustatic.com`(在 `geosite_cn`
|
||||
且在 AGH 拦截表)经 mosdns 返回真实 IP `183.60.227.49` 而非 `0.0.0.0`。已修复(备份
|
||||
`/etc/mosdns/config.yaml.bak-20260812`),修复后同一域名返回 `0.0.0.0`,taobao/google 解析
|
||||
与 clash 链均无回归。**§6 的二期示例同款缺陷已一并修正。**
|
||||
同日追加加固:`domestic_fallback`(fallback 插件:`primary: domestic_upstream`(AGH)、
|
||||
`secondary: default_upstream`(223.5.5.5/119.29.29.29)、`threshold: 500ms`)使 AGH 宕机时
|
||||
DIRECT 国内真实 IP 查询回退国内公网 DNS,不再直接报错;实测:AGH 停止时缓存未命中查询由
|
||||
fallback 应答(NXDOMAIN/真实 IP),AGH 恢复后主路径即时应答且拦截(`0.0.0.0`)恢复。
|
||||
备份:`/etc/mosdns/config.yaml.bak-fallback-20260812`。
|
||||
- AGH rewrites(实测):`hass.windy.lan`/`hass.local` → `192.168.55.11`;`dns.windy.lan` → `.36`;
|
||||
`ubnt.windy.lan` → `.46`;`gfw.windy.lan` → `.1`;`nas.windy.local` → `.32`。
|
||||
- 拦截:仅启用 **AdGuard DNS filter**(filter_1);实测 `doubleclick.net`/`googleadservices.com` → `0.0.0.0`。
|
||||
|
||||
**现状缺口:**
|
||||
1. AGH 上游是**固定 DoH**(alidns/doh.pub,兜底 adg.chans.xyz),**没有"国内/国外分流"能力** → 国外域名解析易受 DNS 污染/时延差,也无法为不同 region 选最优上游。
|
||||
2. 国内/国外分流逻辑(geo)与代理分流逻辑(clash fake-ip)混在网关上,职责不清。
|
||||
**现状缺口(实测确认):**
|
||||
1. AGH 上游是固定 DoH,**无"国内/国外分流"能力**;国外域名解析质量依赖唯一兜底路径。
|
||||
2. **兜底失效**(kill-test 证实):主上游黑洞时,兜底 `adg.chans.xyz` 在客户端 15s 窗口内不生效
|
||||
(`upstream_timeout: 10s` + TCP 重试行为),缓存未命中查询无有界降级——见 §8。
|
||||
3. 国外域名 AAAA 经国内路径全部置空(见 §8),v6 解析缺位。
|
||||
|
||||
## 2. 两个候选方案评估
|
||||
## 2. 最终裁决对齐(用户定稿 2026-08-12)
|
||||
|
||||
| # | 裁决 | 本issue处理 |
|
||||
|---|------|------------|
|
||||
| 1 | **保留 AGH `.36` 为唯一 LAN DNS 入口**(现有方案增强版),不改 EdgeRouter DHCP 通告 | ✅ 现状保持;本期零改动 |
|
||||
| 2 | **否决"AGH 全局转发到 Clash fake-IP"**——DNS 平面必须与流量转发平面一致 | ✅ 分层方案(§4)明确 AGH 上游为**真实 IP** 解析路径,不与 fake-ip 混用 |
|
||||
| 3 | **AGH → mosdns 仅为二期可选项**(经实测确有需求后启用) | ✅ §4 为二期方案;§8 kill-test 已给出"实测需求"证据(降级缺口) |
|
||||
| 4 | 代理 VLAN10 将来用独立 OpenClash DNS 平面(fake-ip + TPROXY),不污染普通 LAN | ✅ 现状即此(dnsmasq→clash,非面向 LAN 客户端);文档记录 |
|
||||
| 5 | **删除或明确禁用** `.1` 上未使用的 mosdns | ⚠️ 前提修正:mosdns 是 clash 的 nameserver,处于活动链路(§1)。处置改为**正式纳管并文档化**(本文件 + `hosts/gfw.windy.lan.md`),不删除 |
|
||||
| 6 | 先完成验证再改动生产路径 | ✅ 本期完成全部 Phase 0 验证(§8),**未改任何生产 DNS 路径** |
|
||||
| 7 | 建立 `home.arpa` 内部域(替代 `.local`) | ⏳ 后续任务:当前命名空间为 `.lan`(AGH rewrites + EdgeRouter DHCP domain),`hass.local` 兼容保留至迁移完成;home.arpa 需联动 AGH rewrites、DHCP domain、客户端,另行排期 |
|
||||
| 8 | 高可用时增加第二个等价 AGH(独立物理故障域) | ⏳ 备用方案,记录不实施 |
|
||||
|
||||
## 3. 两个候选方案评估
|
||||
|
||||
### 方案 A:AGH 单独作为统一入口(现状演进)
|
||||
- 优点:单解析点、面板/拦截/日志集中、维护简单。
|
||||
- 缺点:AGH 对 geo 分流 + 防污染支持弱(官方定位是"过滤/家长控制",见 adguard README "Encrypted DNS upstream... requires additional software")。固定 DoH 上游无法按域名 region 选路。→ **不足以解决防污染/分流问题。**
|
||||
- 缺点:AGH 对 geo 分流 + 防污染支持弱(官方定位是"过滤/家长控制")。固定 DoH 上游无法按域名
|
||||
region 选路;且 §8 kill-test 显示主上游全挂时缓存未命中查询无有界降级。→ **不足以解决防污染/分流/降级问题。**
|
||||
|
||||
### 方案 B:mosdns 作为智能上游分流器
|
||||
mosdns(v5)用 `sequence` 编排:`geosite/geoip` 匹配器 → 国内域名转发国内 DoH、国外域名转发加密 DoH(防污染),可加 `cache`、`reject`(屏蔽)。
|
||||
- 优点:真正解决"国内快 / 国外不被污染"的分流;性能高(百万域名表也不卡)。
|
||||
- 缺点:纯转发器,无 Web 面板、无每客户端统计、拦截要靠域名表(不如 AGH 体验)。→ 单独当入口会退回原始体验。
|
||||
mosdns(v5)用 `sequence` 编排:`geosite/geoip` 匹配器 → 国内域名转发国内 DoH、国外域名转发加密 DoH(防污染),可加 `cache`、`reject`。
|
||||
- 优点:真正解决"国内快 / 国外不被污染"的分流;性能高。
|
||||
- 缺点:纯转发器,无 Web 面板、无每客户端统计、拦截靠域名表。→ 单独当入口会退回原始体验。
|
||||
|
||||
**结论:两个方案是互补的,不是二选一。** 单用 A 无法分流防污染,单用 B 失去 AGH 的管理体验。
|
||||
**结论:两个方案互补,不是二选一。** 单用 A 无法分流防污染且降级无界;单用 B 失去 AGH 管理体验。
|
||||
|
||||
## 3. 推荐:分层架构(AGH 前端 + mosdns 后端)
|
||||
## 4. 二期可选项:分层架构(AGH 前端 + mosdns 后端,经实测需求后启用)
|
||||
|
||||
```
|
||||
局域网客户端(DHCP DNS = 192.168.66.36)
|
||||
@@ -44,33 +108,31 @@ AdGuard Home (66.36) ── 前端:广告/过滤、拦截表、每客户端
|
||||
▼
|
||||
mosdns(66.36 伴生容器) ── 后端:智能分流 + 防污染
|
||||
│ - geosite:cn → 国内 DoH/UDP(aliDNS / 腾讯 DNSPod)
|
||||
│ - 其他 → 加密 DoH(Cloudflare/Google/自建 adg.chans.xyz)
|
||||
│ - 其他 → 加密 DoH(自建 adg.chans.xyz 等)
|
||||
▼
|
||||
上游 DoH
|
||||
```
|
||||
|
||||
职责分离,每个工具只做自己最擅长的事:
|
||||
- **AGH = 策略/拦截/可观测**(拦截表、每客户端日志、面板)。AGH 原生干不了"按域名选路",所以不做分流。
|
||||
- **mosdns = 智能转发**(geo 分流 + 加密防污染)。不用它当入口,所以保持 AGH 的 UX。
|
||||
- **OpenClash(gfw)= 代理选路**(fake-ip + 规则决定"哪些流量走代理")。与"DNS 解析选上游"是**两个独立决策**,分开放在不同工具最干净——DNS 解析在 66.36 做,代理路由在网关做,互不耦合。
|
||||
职责分离:
|
||||
- **AGH = 策略/拦截/可观测**(拦截表、每客户端日志、面板)。
|
||||
- **mosdns = 智能转发**(geo 分流 + 加密防污染 + 内置 cache,可显著缩短降级窗口)。
|
||||
- **OpenClash(gfw)= 代理选路**(fake-ip + 规则)。与 DNS 解析选上游是两个独立决策,分开放最干净。
|
||||
|
||||
### 推荐部署位置:mosdns 与 AGH 同机(66.36),而非 gf(.1)
|
||||
**部署位置:mosdns 与 AGH 同机(66.36 伴生容器),而非 gfw(.1)**:单点即 AGH 所在;不受网关重启/
|
||||
OpenClash churn 影响;可纳入现有 compose/ansible 管理;不占用 OpenWrt 资源。放 gfw 会与 clash 的
|
||||
DNS 处理互相干扰、耦合,且网关重启即断全 LAN DNS。**不推荐放 .1。**
|
||||
|
||||
| 位置 | 评价 |
|
||||
|------|------|
|
||||
| **66.36 伴生容器(推荐)** | 单点即 AGH 所在;AGH→mosdns 走本机/近端一跳;不受网关重启/OpenClash churn 影响;可纳入现有 ansible compose 管理;不占用 OpenWrt 资源 |
|
||||
| gf(.1) | 虽近网络边缘,但该网关已有 clash fake-ip + 多种劫持规则,再叠 mosdns 会与 clash 的 DNS 处理互相干扰、耦合;且网关重启即断 DNS(影响整个 LAN)。**不推荐** |
|
||||
> 启用条件(kill-test 实测需求,§8):主上游全挂时,当前 AGH 单入口对缓存未命中查询无有界降级。
|
||||
> 分层方案(或下调 `upstream_timeout` + 改 failover 模式)可修;启用与否由用户在二期决定。
|
||||
|
||||
> 注意:若把 mosdns 放 gfw,必须先理清与 OpenClash `dnsmasq→clash #7874` + `nft fw4 DNS-hijack` 的先后/覆盖关系,否则会出现"部分设备解析走了 clash、部分走了 mosdns"的混乱。放 66.36 则完全避开这个冲突。
|
||||
## 5. 更优替代方案(一并考虑)
|
||||
|
||||
## 4. 更优替代方案(也一并考虑)
|
||||
1. **分层(推荐二期,见 §4)**:AGH(66.36)→ mosdns(66.36 伴生)→ 上游。体验最好、职责最清。
|
||||
2. **纯 mosdns + 前端面板**:损失拦截/统计管理体验。**不推荐**用于替换。
|
||||
3. **AGH 只挂一个带分流的上游(第三方 DoH 聚合)**:失去可控性且不可信。不推荐做主路径。
|
||||
4. **全部交给 OpenClash fake-ip,关闭 AGH**:让"代理网关"成为全 LAN DNS 单点;且 AGH 拦截/日志也没了。**不推荐。**
|
||||
|
||||
1. **分层(推荐,见上)**:AGH(66.36)→ mosdns(66.36 伴生)→ 上游。体验最好、职责最清。
|
||||
2. **纯 mosdns + 前端面板**:不用 AGH,用 mosdns + 其他统计面板。→ 会明显损失拦截/统计管理体验,除非你讨厌 AGH 的 Docker 部署。**不推荐**用于替换。
|
||||
3. **AGH 只挂一个带分流的上游(第三方 DoH 聚合)**:例如接一个已做分流的公共 DoH。→ 失去可控性,且不可信。不推荐做主路径。
|
||||
4. **全部交给 OpenClash fake-ip,关闭 AGH**:把 LAN 客户端 DNS 指到 gfw。→ 让"代理网关"成为全 LAN DNS 单点,网关重启/代理抖动整个内网断网;且 AGH 的拦截/日志也没了。**不推荐。**
|
||||
|
||||
## 5. mosdns 配置要点(mosdns v5,预留实施)
|
||||
## 6. mosdns 配置要点(mosdns v5,二期实施预留)
|
||||
|
||||
核心是 `sequence` + 上游拆分 + 缓存 + 屏蔽:
|
||||
|
||||
@@ -80,39 +142,105 @@ plugins:
|
||||
type: sequence
|
||||
args:
|
||||
- exec: cache 1024 # 缓存加速
|
||||
- matches: has_resp
|
||||
exec: accept
|
||||
# 国内分流:命中 geosite:cn → 国内 DoH
|
||||
- matches: [ qname &geosite:cn ]
|
||||
exec: forward https://dns.alidns.com/dns-query
|
||||
# 广告域名可选屏蔽(或交给 AGH 前置拦截,二选一)
|
||||
# - matches: [ qname &./blocklist.txt ]
|
||||
# exec: reject 3
|
||||
# 广告域名屏蔽交 AGH 前置,不重复维护
|
||||
# 其余(国外)→ 加密 DoH 防污染
|
||||
- exec: forward https://1.1.1.1/dns-query
|
||||
# 备选国外上游/兜底
|
||||
- matches: [ has_resp ]
|
||||
exec: accept
|
||||
- exec: forward_addr https://208.67.222.222:443/dns-query
|
||||
- exec: forward https://adg.chans.xyz/dns-query
|
||||
- type: udp_server
|
||||
args: { entry: main, listen: "127.0.0.1:5353" }
|
||||
- type: tcp_server
|
||||
args: { entry: main, listen: "127.0.0.1:5353" }
|
||||
```
|
||||
|
||||
> **注意**:`sequence` 中每个 `forward` 分支之后必须跟 `matches: has_resp → accept`
|
||||
> (或改用 `goto`/`jump` + `return` 结构),否则查询会继续执行后续规则被二次转发,
|
||||
> 最终应答来自最后一个 forward——gfw 上 mosdns 的同类缺陷(2026-08-12)已实测并修复(见 §1)。
|
||||
|
||||
要点:
|
||||
- 上游可加 `upstream` 的 `concurrent > 1` 与 `addr` 做多/故障切换。
|
||||
- `geosite:cn` / `geoip:cn` 数据插件自动从 repo 更新;国内用 aliDNS/腾讯,国外用 DoH(Cloudflare/Google/自建 adg.chans.xyz)。
|
||||
- 屏蔽交由 AGH 前置(推荐),不要 AGH 和 mosdns 都自己维护一套拦截表(重复)。
|
||||
- 上游可加 `upstream` 的 `concurrent > 1` 与多地址故障切换;mosdns 自带 cache,能保证上游故障时
|
||||
缓存命中仍即时应答(对应 §8 认定的降级缺口)。
|
||||
- `geosite:cn` / `geoip:cn` 数据自动更新;国内 aliDNS/腾讯,国外可用自建 `adg.chans.xyz`(实测
|
||||
唯一能返回国外 AAAA 的路径,§8)。
|
||||
- 屏蔽交 AGH 前置,AGH 与 mosdns 不各自维护拦截表。
|
||||
|
||||
## 6. 迁移 / 实施顺序(待评审)
|
||||
## 7. 迁移 / 实施顺序(二期,待用户确认启用)
|
||||
|
||||
1. 在 66.36 起 mosdns 伴生容器(`/opt/mosdns` + compose,固定 digest,纳入 ansible)。
|
||||
2. AGH「上游 DNS 服务器」改为指向 mosdns(`http://127.0.0.1:5353/dns-query` 或 `127.0.0.1:5353`)。AGH 的 `bootstrap` 仍用公网 IP(避免 AGH → mosdns → AGH 死循环)。
|
||||
3. 验证:国内域名(如 `taobao.com`)、国外域名(如 `google.com`)、被拦截域名、每客户端日志。
|
||||
4. 确认后,`disable`/移除 gfw 上闲置的 mosdns(6052)以免混淆。
|
||||
5. 回归:EdgeRouter 通告不变(仍 `.36`),因此 LAN 客户端无感;重启 AGH/mosdns 单点验证。
|
||||
1. 在 66.36 起 mosdns 伴生容器(`/opt/mosdns` + compose,**按 digest 固定镜像**,纳入 ansible)。
|
||||
2. AGH「上游 DNS 服务器」改为指向 mosdns(`127.0.0.1:5353`,bootstrap 仍用公网 IP,避免
|
||||
AGH → mosdns → AGH 死循环);只保留**一条语义一致的上游路径**,保持可回滚(备份 yaml + `--check-config`)。
|
||||
3. 验证:国内域名、国外域名、被拦截域名、每客户端日志、AAA A 解析(§8 基线)。
|
||||
4. 回归:EdgeRouter 通告不变(仍 `.36`),LAN 客户端无感;重启 AGH/mosdns 单点验证(§8 kill-test 模板)。
|
||||
5. 上线后重跑 §8 kill-test,确认降级窗口有界。
|
||||
|
||||
## 7. 风险与备注
|
||||
- mosdns 仅监听 `127.0.0.1`(不对外),由 AGH 消费;避免 LAN 直连 mosdns 造成两套入口。
|
||||
- AGH 上游指向本机 mosdns 时,务必配 bootstrap 公网 IP,否则自举死循环。
|
||||
- 本方案不改 EdgeRouter DHCP/通告,不改 gfw OpenClash 代理规则,只动 66.36 上的 DNS 链路,风险可控。
|
||||
- 与 W1N-40「审查并修正 AdGuard Home」联动:该 issue 侧重 AGH 本身,本 issue 侧重整体 DNS 分层。
|
||||
## 8. Phase 0 验证证据(2026-08-12 全部实测)
|
||||
|
||||
### 8.1 基线与功能
|
||||
|
||||
| 项 | 结果 |
|
||||
|----|------|
|
||||
| rewrites:`hass.windy.lan` / `hass.local` | → `192.168.55.11` ✅(兼容保留) |
|
||||
| rewrites:`dns.windy.lan` / `gfw.windy.lan` / `ubnt.windy.lan` / `nas.windy.local` | → `.36` / `.1` / `.46` / `.32` ✅ |
|
||||
| 国内解析 `taobao.com`(经 AGH) | 真实 CN IP(59.82.x 等)✅ |
|
||||
| 国外解析 `google.com` / `github.com`(经 AGH) | 真实 IP(142.250.x / 20.205.x),**无 fake-IP 泄漏** ✅ |
|
||||
| 广告拦截 `doubleclick.net` / `googleadservices.com` | → `0.0.0.0` ✅ |
|
||||
| clash 7874 `google.com` | `198.18.1.101`(fake-ip,仅网关/VLAN10 平面)✅ |
|
||||
| clash 7874 `taobao.com` | 真实 IP(经 mosdns→AGH)✅ |
|
||||
| dnsmasq :53(.1)`google.com` / `taobao.com` | fake-ip / 真实 IP ✅ |
|
||||
| mosdns 6052 直连 `taobao.com` / `google.com` | 真实 IP(59.82.x / 142.250.73.78)✅ |
|
||||
| DNSSEC:`dnssec-failed.org`(经 AGH) | 返回正常应答 `96.99.227.255`(非 SERVFAIL)→ 当前路径不校验,W1N-40 结论复现,**维持关闭** |
|
||||
|
||||
### 8.2 出口路径与 fake-IP 泄漏
|
||||
|
||||
- AGH 宿主默认路由 `via 192.168.66.254`(EdgeRouter),**直连出网,不经 gfw**;DoH 端点实测可达
|
||||
(见 §1)。LAN 客户端经 AGH 的解析结果全部为真实 IP,无 `198.18/16` 泄漏。
|
||||
- gfw 自身流量默认进代理(`openclash_mangle_output` mark 0x162 → tproxy :7895),
|
||||
clash nameserver-policy 的 `1.1.1.1` DoH 实测 35ms 可达(走代理链路,不依赖直连)。
|
||||
- mosdns 上游(AGH `.36`、`223.5.5.5`)命中本地/国内 bypass 规则,保持直连——设计意图达成。
|
||||
|
||||
### 8.3 IPv6 / RDNSS / AAAA
|
||||
|
||||
- LAN 有 IPv6 SLAAC(EdgeRouter dhcpv6-pd /60 → eth0 host-address + switch0,**仅 `service slaac`**,
|
||||
**无 RDNSS/dns-server 通告**);`.36` 有全局 v6 地址 + RA 默认路由。
|
||||
- **RDNSS 未通告** → v6 客户端无 v6 DNS,回退 v4 DNS(`.36`);AGH 仅监听 `0.0.0.0:53`(v4 only),无 v6 DNS 服务。
|
||||
- **AAAA 解析实测**:`baidu.com` 公网本就无 AAAA(dns.google NOERROR/0,权威 NS 而已);
|
||||
`taobao.com` AAAA 经国内路径正常(`2408:4001:f10::6f` 等);**国外域名(`google.com`)经
|
||||
`223.5.5.5` UDP、alidns DoH、`8.8.8.8` UDP 全部返回空**,而 dns.google 与 `adg.chans.xyz`
|
||||
DoH 均能返回 `2404:6800:4005:81a::200e` → **国内路径对国外域 AAAA 置空;`adg.chans.xyz`
|
||||
兜底是当前唯一能返回国外 AAAA 的路径**。clash `ipv6: false` 亦不返回 AAAA。
|
||||
|
||||
### 8.4 自举(bootstrap)循环
|
||||
|
||||
- AGH `bootstrap_dns: [223.5.5.5, 223.6.6.6]`(公网 IP,非 AGH 自身)→ 无自举循环;AGH 解析
|
||||
DoH 主机名不经过自身。二期方案要求 AGH→mosdns 时 bootstrap 仍用公网 IP(§7)。
|
||||
|
||||
### 8.5 Kill-test 矩阵(2026-08-12,全部实测)
|
||||
|
||||
| # | 场景 | 结果 |
|
||||
|---|------|------|
|
||||
| 1 | 重启 `.1` mosdns(init.d) | ✅ 直连 6052 与 clash 链恢复 |
|
||||
| 2 | 重启 `.1` dnsmasq | ✅ 真实 IP 与 fake-ip 双路径恢复 |
|
||||
| 3 | kill `.1` clash 核心 | ✅ LAN DNS(AGH)不受影响;gfw dnsmasq→clash **有界 3s 失败**(无卡死);OpenClash watchdog ~15s 自动拉起 |
|
||||
| 4 | stop/start `.1` OpenClash | ✅ DNS 平面独立于代理;clash 与 tproxy 规则恢复 |
|
||||
| 5 | 重启 `.36` AGH 容器 | ✅ 全量恢复:rewrites/拦截/国内外解析/DNSSEC 行为不变 |
|
||||
| 6 | 黑洞 alidns DoH(223.5.5.5/223.6.6.6:443) | ✅ ~0.5s 内经 `doh.pub` 应答(load_balance 生效) |
|
||||
| 7 | 黑洞全部主上游,兜底存活(adg.chans.xyz) | ⚠️ **客户端 15s 内无应答**——兜底未在窗口内生效 |
|
||||
| 8 | 黑洞全部上游(含兜底),缓存未命中 | ⚠️ **25s 内无应答、无 SERVFAIL**——解析器对缓存未命中查询"卡死" |
|
||||
| 9 | 黑洞全部上游,缓存命中 | ✅ 瞬时 NOERROR(cache 兜底) |
|
||||
|
||||
**结论(Phase 0 门禁):** 国内解析在代理停止/上游单点故障/组件重启下均维持可用;
|
||||
但**"主上游全挂"时缓存未命中查询无有界降级**——`upstream_timeout: 10s` 与 TCP 重试行为使
|
||||
兜底 `adg.chans.xyz` 在实践中无法在客户端期望窗口内生效。这是 §4 二期分层方案(或下调
|
||||
`upstream_timeout` + failover 模式)的**实测需求依据**;按最终裁决,本期不改生产路径。
|
||||
|
||||
## 9. 风险与备注
|
||||
|
||||
- mosdns 仅监听 `127.0.0.1`(不对外),由 clash 消费;二期若启用,保持同样的边界,避免 LAN 出现两套入口。
|
||||
- AGH 上游指向本机 mosdns 时务必配公网 bootstrap,否则自举死循环。
|
||||
- 本方案不改 EdgeRouter DHCP/通告、不改 gfw OpenClash 代理规则,只动 66.36 上的 DNS 链路,风险可控。
|
||||
- 已知降级缺口(§8.5 #7/#8):主上游全挂时缓存未命中查询无有界降级;启用二期前,LAN 客户端会感知
|
||||
超时(约 10s+)。缓解:AGH cache 已覆盖高频域;根治需二期。
|
||||
- 与 W1N-40「审查并修正 AdGuard Home」联动:该 issue 侧重 AGH 本身,本 issue 侧重整体 DNS 分层。
|
||||
@@ -307,7 +307,7 @@ VLAN10 / 升级 SSID / 客人 SSID **失败或未做,不否决**本次核心
|
||||
|
||||
## 13. 参考
|
||||
|
||||
- 实施阶段与清单:[lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md)
|
||||
- 实施阶段与清单:[lan-core-switch-upgrade-plan.md](archive/lan-core-switch-upgrade-plan.md)
|
||||
- 现网地图:[lan-overview.md](lan-overview.md)
|
||||
- ER-X:[edgerouter-x-configuration.md](edgerouter-x-configuration.md)、[hosts/gw.md](../hosts/gw.md)
|
||||
- UniFi / VLAN10 前置:[unifi-network.md](unifi-network.md)
|
||||
|
||||
+76
-11
@@ -13,6 +13,11 @@ from each section below.
|
||||
> **Verified live on 2026-08-06** by read-only SSH from the WSL client. No
|
||||
> changes were made. `gfw.windy.lan` root SSH was re-verified the same day after
|
||||
> the key was installed; its facts below are from the fresh probe.
|
||||
>
|
||||
> **IPv6 re-verified 2026-08-20** (read-only): UniFi controller `Default`
|
||||
> network IPv6 enabled (SLAAC/RA), both APs hold global SLAAC addresses, and
|
||||
> `zhiqiangf` key-only AP SSH re-confirmed. See
|
||||
> [unifi-network.md](unifi-network.md).
|
||||
|
||||
---
|
||||
|
||||
@@ -35,10 +40,14 @@ from each section below.
|
||||
│ ubnt — UniFi Network Controller (192.168.66.46)
|
||||
```
|
||||
|
||||
> **SE5420 purchased (2026-08-09):** TP-Link `TL-SE5420` acquired; deployment plan is
|
||||
> **SE5420 live (2026-08-22):** TP-Link `TL-SE5420` (purchased 2026-08-09) is
|
||||
> online — management `192.168.66.253` reachable, web UI on :80/:443; LAN55
|
||||
> 上联为 ER-X `switch0` **单口**(`eth1` up、`eth2`/`eth3` down,2026-08-22
|
||||
> 只读核实)→ `switch0` 不再是 LAN55 全量抓包点(同段有线单播在 SE5420 本地
|
||||
> 交换),全量点只能靠 SE5420 port mirroring。迁移状态见部署计划
|
||||
> [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md). Design/planning refs:
|
||||
> [lan-erx-se5420-network.md](lan-erx-se5420-network.md),
|
||||
> [lan-core-switch-upgrade-plan.md](lan-core-switch-upgrade-plan.md).
|
||||
> [lan-core-switch-upgrade-plan.md](archive/lan-core-switch-upgrade-plan.md).
|
||||
|
||||
---
|
||||
|
||||
@@ -47,19 +56,20 @@ from each section below.
|
||||
| Host | Role | SSH | IPv4 | Facts |
|
||||
|------|------|-----|------|-------|
|
||||
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | [hosts/gw.md](../hosts/gw.md) |
|
||||
| **PVE** | Proxmox host (`.66.26`/vmbr0 · `.55.26`/vmbr1) — hosts gfw/dns/ubnt/haos VMs | `ssh -4 root@192.168.66.26` | `192.168.66.26` | — |
|
||||
| **PVE** | Proxmox host (`.66.26`/vmbr0 · `.55.26`/vmbr1) — hosts gfw/dns/ubnt VMs | `ssh -4 root@192.168.66.26` | `192.168.66.26` | — |
|
||||
| **gfw.windy.lan** | OpenWrt LAN gateway / OpenClash — **PVE VM 140** | `ssh -4 root@192.168.66.1` | `192.168.66.1` | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
|
||||
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy — **PVE VM 120** (`pihole`) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
|
||||
| **ubnt** | UniFi Network Controller — **PVE VM 160** | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | [hosts/ubnt.md](../hosts/ubnt.md) |
|
||||
| **haos** | Home Assistant (HAOS) — **PVE VM 180** (LAN55) | — | `192.168.55.11` | — |
|
||||
| **hass.windy.lan** | Home Assistant (HAOS) — **x88 Pro physical box** (LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
|
||||
| **pgdb** | TimescaleDB PG18 (Docker) — HA recorder 后端 — **PVE VM** (LAN55) | `ssh -4 windy@192.168.55.15` | `192.168.55.15` | [hosts/pgdb.md](../hosts/pgdb.md) |
|
||||
| **NAS/FreeNAS** | NAS; `transmission` jail runs here (`.51`) | — | — | — |
|
||||
| **U6 Lite** | UniFi AP (LAN66) | `ssh -4 zhiqiangf@192.168.66.6` | `192.168.66.6` | [docs/unifi-network.md](../docs/unifi-network.md) |
|
||||
| **UAP-AC-Lite** | UniFi AP (LAN55) | `ssh -4 zhiqiangf@192.168.55.5` | `192.168.55.5` | [docs/unifi-network.md](../docs/unifi-network.md) |
|
||||
|
||||
> **Positioning facts (verified 2026-08-09):** `dns`/`ubnt`/`gfw`/`haos` are all VMs on PVE
|
||||
> (no separate physical hosts); `transmission` is a FreeNAS/NAS jail. Only gw, PVE,
|
||||
> NAS, U6, UAP-AC-Lite, and wired PCs/NAS are physical SE5420 ports. See
|
||||
> [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) §1.
|
||||
> **Positioning facts:** `dns`/`ubnt`/`gfw`/`pgdb` are VMs on PVE; `haos` is a **physical x88 Pro
|
||||
> box** (HAOS bare-metal, `machine: green`), not a PVE VM (corrected 2026-08-15).
|
||||
> `transmission` is a FreeNAS/NAS jail. Physical SE5420 ports: gw, PVE, haos, NAS,
|
||||
> U6, UAP-AC-Lite, and wired PCs. See [lan-se5420-deployment-guide.md](lan-se5420-deployment-guide.md) §1.
|
||||
|
||||
---
|
||||
|
||||
@@ -76,7 +86,10 @@ from each section below.
|
||||
| Port-forwards | `hass`→192.168.55.11:8123 · `transmission`→192.168.66.51:51413 · `ssh`→192.168.66.36:22 (orig 5822) · `openvpn`→192.168.66.32:1194 · WAN iface pppoe0 |
|
||||
| Management | SSH TCP 22 · EdgeOS GUI HTTP 80 / HTTPS 443 |
|
||||
|
||||
**Static DHCP mappings (LAN66):** `OnePlus-12`=.37, `gfw`=.1, `hp-nas`=.32, `pihole`=.36, `pve`=.26, `transmission`=.51, `ubnt-6`=.6, `ubnt-app`=.46, `windy-pc`=.99. LAN55: `Aqara-Hub-M3-10CB`=.248.
|
||||
**Static DHCP mappings (LAN66):** `OnePlus-12`=.37, `gfw`=.1, `hp-nas`=.32, `pihole`=.36, `pve`=.26, `transmission`=.51, `ubnt-6`=.6, `ubnt-app`=.46, `windy-pc`=.99. LAN55: `Aqara-Hub-M3-10CB`=.248, `SmartThings-Station`=.48, `espressif`=.47,
|
||||
`hass`=.11, `hass-wifi`=.250, `ihost`=.12, `midea_ac_0418`=.10,
|
||||
`midea_e3_0198`=.42, `roborock-wm-a141`=.43, `samsung-hub`=.251,
|
||||
`matter`=.41 (added 2026-08-20).
|
||||
|
||||
> **Note:** `LAN_IN`/`LAN_OUT` are defined but not applied to an interface, so LAN55
|
||||
> and LAN66 are bidirectionally reachable by default. Do not rely on those rules as
|
||||
@@ -93,7 +106,7 @@ from each section below.
|
||||
| SSH | `ssh -4 root@192.168.66.1` (key-only, verified 2026-08-06) |
|
||||
| OpenClash | `/etc/openclash/clash` (clash_meta core) + config `/etc/openclash/pass-cat.yaml` |
|
||||
| Mode | **fake-ip + TPROXY transparent proxy** (`operation_mode=fake-ip`, `en_mode=fake-ip`, `proxy_mode=rule`) |
|
||||
| DNS | dnsmasq → clash DNS `127.0.0.1#7874`; `mosdns` also listens on `127.0.0.1:6052` (not the active path) |
|
||||
| DNS | dnsmasq → clash DNS `127.0.0.1#7874`; clash `nameserver` = mosdns `127.0.0.1:6052` (DIRECT 规则真实 IP 解析,非客户端路径) |
|
||||
| nft | `table inet fw4` with OpenClash TPROXY/redirect + DNS-hijack rules; residual `table inet passwall` (0 packets, unused) |
|
||||
|
||||
**OpenClash listeners:** HTTP `7890` · SOCKS `7891` · Redirect `7892` · Mixed `7893` · TPROXY `7895` · DNS `7874` · dashboard `9090`. `8443` is **not** an OpenClash listener (only in its TLS-sniffing port list).
|
||||
@@ -145,6 +158,19 @@ See [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md).
|
||||
|
||||
---
|
||||
|
||||
## hass.windy.lan — Home Assistant (HAOS)
|
||||
|
||||
| Item | Value |
|
||||
|------|-------|
|
||||
| IPv4 | `192.168.55.11` (LAN55) |
|
||||
| DNS | `hass.windy.lan` (AdGuard rewrite; legacy `hass.local` alias) |
|
||||
| SSH | `ssh hassio@hass.windy.lan` (key-only, verified 2026-08-13) |
|
||||
| Web UI | `http://hass.windy.lan:8123` |
|
||||
| WAN | gw port-forward `hass` → `192.168.55.11:8123` |
|
||||
| Platform | HAOS on physical x88 Pro box; kernel `6.1.115-haos` (aarch64), `machine: green` |
|
||||
|
||||
---
|
||||
|
||||
## Managed access points
|
||||
|
||||
| Name | Model | Mgmt IP | Firmware | Network | Inform |
|
||||
@@ -155,6 +181,43 @@ See [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md).
|
||||
Both reported **Connected** to `http://192.168.66.46:9080/inform` on 2026-08-06.
|
||||
AP SSH account is `zhiqiangf` (key-only, verified). See [docs/unifi-network.md](../docs/unifi-network.md).
|
||||
|
||||
**IPv6 (verified 2026-08-20):** both APs hold global SLAAC IPv6 addresses on
|
||||
`br0` — U6 Lite `240e:3bd:235:1fb1::/64` (LAN66), UAP-AC-Lite
|
||||
`240e:3bd:235:1fb2::/64` (LAN55) — with RA default routes via `gw`; the
|
||||
controller's `Default` network has IPv6 enabled (SLAAC). Prefixes are dynamic
|
||||
(PPPoE PD), so they rotate on redial. Details:
|
||||
[docs/unifi-network.md](../docs/unifi-network.md).
|
||||
|
||||
**SSID cleanup (2026-08-21, W1N-207):** the SmartThings Element/vWire provisioning
|
||||
SSIDs (`element-8a0d5133c9438f12`, `vwire-8b2d67469e455785`, `vport-F09FC22004E9`)
|
||||
were removed/disabled in the controller (`element_adopt` setting off, element wlanconf
|
||||
deleted, connectivity `x_mesh_essid`/`x_mesh_psk` cleared, device `x_vwirekey` removed,
|
||||
`vwire_enabled`/`mesh_sta_vap_enabled=false`) and cleared from both APs; all
|
||||
vwire/vport/element flags on the remaining SSIDs are now `disabled`.
|
||||
|
||||
**Stable ULA on gw: not feasible (2026-08-21, W1N-207):** EdgeOS v3.0.1
|
||||
`interfaces switch switch0` rejects a static `ipv6 address`, and an explicit
|
||||
`router-advert` node *replaces* the DHCPv6-PD-slaac RA (drops the delegated GUA
|
||||
prefix from radvd → LAN55 loses IPv6 egress after RA expiry). Attempted and rolled
|
||||
back cleanly (no `save`; gw config unchanged). Consequence: after a PD rotation,
|
||||
restart HA's matter-server (see [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md))
|
||||
to clear stale IPv6 mDNS caches.
|
||||
|
||||
**LAN55 RA environment (observed 2026-08-21):** besides `gw`, the SmartThings
|
||||
Station (.48) and Aqara M3 (.248) act as Thread border routers and advertise ULA
|
||||
prefixes (`fd00:5a7:6415:1::/64`, `fd97:d580:16fe:1::/64`); several LAN55 hosts
|
||||
(HA, PVE, UAP-AC-Lite) have IPv6 forwarding enabled and mark themselves as
|
||||
routers in NDP. This is normal Thread-BDR behaviour and was not the Matter
|
||||
failure cause.
|
||||
|
||||
**Matter 灯泡(2026-08-21 实测,W1N-207):** 两盏 ESP32-C2 Matter 灯泡
|
||||
(VP `0x4891/0x4100`,OUI `34:98:7a`)——工作盏 MAC `34:98:7a:25:a1:f0`;故障盏
|
||||
MAC `34:98:7a:27:7f:08`(hostname `matter`,动态 .145)。故障盏已在 Aqara fabric
|
||||
`4DF2B1455D19402D` 内、宣告 `CM=0`(不在配对模式)且缺 GUA → 找回需**恢复出厂**
|
||||
后扫它自己的二维码。DHCP 保留 `matter`(.45 → MAC `34:98:7a:27:10:bc`)与故障盏
|
||||
MAC 不符,保留从未租出(待修,见 [hosts/gw.md](../hosts/gw.md))。完整排障知识:
|
||||
[docs/matter-pairing-troubleshoot.md](matter-pairing-troubleshoot.md)。
|
||||
|
||||
---
|
||||
|
||||
## Quick orientation (who runs what)
|
||||
@@ -165,6 +228,7 @@ AP SSH account is `zhiqiangf` (key-only, verified). See [docs/unifi-network.md](
|
||||
| Transparent/explicit proxy (OpenClash) | gfw.windy.lan | `ssh -4 root@192.168.66.1` |
|
||||
| LAN DNS (AdGuard Home) + Mihomo proxy | dns.windy.lan | `ssh -4 windy@192.168.66.36` |
|
||||
| UniFi controller + dockge | ubnt | `ssh -4 windy@192.168.66.46` |
|
||||
| Home Assistant | hass.windy.lan | `ssh hassio@hass.windy.lan` · UI `:8123` |
|
||||
| Wi-Fi APs | U6 Lite / UAP-AC-Lite | via controller |
|
||||
|
||||
---
|
||||
@@ -175,9 +239,10 @@ AP SSH account is `zhiqiangf` (key-only, verified). See [docs/unifi-network.md](
|
||||
- [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) — OpenClash listeners
|
||||
- [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) — AdGuard Home + Mihomo detail
|
||||
- [hosts/ubnt.md](../hosts/ubnt.md) — UniFi controller + proxy contract
|
||||
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — Home Assistant (HAOS) SSH + LAN access
|
||||
- [docs/unifi-network.md](../docs/unifi-network.md) — APs, inform endpoint, recovery
|
||||
- [docs/unifi-third-party-vlan10-dhcp.md](unifi-third-party-vlan10-dhcp.md) — VLAN Wi-Fi feasibility and DHCP boundary
|
||||
- [docs/unifi-openwrt-vlan10-implementation-examples.md](unifi-openwrt-vlan10-implementation-examples.md) — supported topology and examples
|
||||
- [docs/edgerouter-x-configuration.md](../docs/edgerouter-x-configuration.md) — effective gw config
|
||||
- [docs/unifi-openclash-localhost.md](../docs/unifi-openclash-localhost.md) — proxy bypass
|
||||
- [runbooks/adguard-home-health.md](../runbooks/adguard-home-health.md) — AGH health
|
||||
- [runbooks/adguard-home-health.md](../runbooks/adguard-home-health.md) — AGH health
|
||||
@@ -106,7 +106,7 @@
|
||||
|
||||
- 第一步不建 VLAN10、不向 ER-X 送任何 tag、口 4/6(PVE、U6)不做 trunk。
|
||||
- NAS 只接口 8,口 12 断开(LACP 是独立维护窗)。
|
||||
- 不占口的 VM:dns(.36=VM120)、ubnt(.46=VM160)、gfw(.1=VM140)、haos(.55.11=VM180);transmission(.51) 是 NAS jail。
|
||||
- 不占口的 VM:dns(.36=VM120)、ubnt(.46=VM160)、gfw(.1=VM140);haos(.55.11) 是物理 x88 Pro 盒子(非 VM);transmission(.51) 是 NAS jail。
|
||||
|
||||
## 3. 开箱与固件升级
|
||||
|
||||
@@ -207,7 +207,7 @@
|
||||
4. **验证:**
|
||||
- `ip -br addr`:`vmbr1` = `192.168.55.26/24`;
|
||||
- `ping -c3 192.168.55.254` → 通。
|
||||
5. 逐台验证 VM(顺序:gfw → dns → ubnt → haos):
|
||||
5. 逐台验证(顺序:gfw → dns → ubnt → haos;前三个是 VM,haos 是物理盒子):
|
||||
```bash
|
||||
ssh -4 root@192.168.66.26 'qm list'
|
||||
```
|
||||
@@ -215,7 +215,7 @@
|
||||
- dns:`ping -c3 192.168.66.36` → 通;
|
||||
- ubnt:`ping -c3 192.168.66.46` → 通;
|
||||
- haos:`ping -c3 192.168.55.11` → 通(注意是 55 网段)。
|
||||
6. 每个 VM 再验业务:gfw 的 OpenClash 面板/DNS 正常、dns 的 AdGuard UI 能开、ubnt 控制器 Connected、haos 界面能开。不以"宿主开机"代替。
|
||||
6. 每台再验业务:gfw 的 OpenClash 面板/DNS 正常、dns 的 AdGuard UI 能开、ubnt 控制器 Connected、haos 界面能开。不以"宿主开机"代替。
|
||||
|
||||
## 7. 迁移 AP 与接入设备
|
||||
|
||||
@@ -462,7 +462,7 @@ ssh -4 root@192.168.66.1 'uci show network; uci show firewall; uci show dhcp; ip
|
||||
|
||||
- 每次实质变更后在 Linear `vps` 项目记录 scope / action / verification / 遗留 follow-up。
|
||||
- 本仓库不记录 SE5420 口令、ER-X 配置快照(含 PPPoE/口令)、gfw 凭据。
|
||||
- 实施前先读 `se5420-review-claim-verification-2026-08.md` 的现场只读复核结论。
|
||||
- 实施前先读 `archive/se5420-review-claim-verification-2026-08.md` 的现场只读复核结论。
|
||||
|
||||
## 16. 回滚
|
||||
|
||||
@@ -486,10 +486,10 @@ ssh -4 root@192.168.66.1 'uci show network; uci show firewall; uci show dhcp; ip
|
||||
## 参考
|
||||
|
||||
- 设计说明:[lan-erx-se5420-network.md](lan-erx-se5420-network.md)
|
||||
- 评审核实:[se5420-review-claim-verification-2026-08.md](se5420-review-claim-verification-2026-08.md)
|
||||
- 评审核实:[se5420-review-claim-verification-2026-08.md](archive/se5420-review-claim-verification-2026-08.md)
|
||||
- 现网地图:[lan-overview.md](lan-overview.md)
|
||||
- 官方安装手册(Markdown 版):[se5420-official-manuals/tl-se5420-install-manual.md](se5420-official-manuals/tl-se5420-install-manual.md)
|
||||
- 官方 PDF:<https://service.tp-link.com.cn/download/202310/TL-SE5420%20V1.0安装手册%201.0.2.pdf>
|
||||
- 规格 / 固件:<https://www.tp-link.com.cn/product_2899.html?v=specification> · <https://www.tp-link.com.cn/product_2899.html?v=download>
|
||||
- Omada VLAN 指南:<https://support.omadanetworks.com/en/document/12981/> · <https://support.omadanetworks.com/en/document/13135/>
|
||||
- ER-X:[edgerouter-x-configuration.md](edgerouter-x-configuration.md);gfw:[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md);PVE VLAN10:[lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09](lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09)
|
||||
- ER-X:[edgerouter-x-configuration.md](edgerouter-x-configuration.md);gfw:[hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md);PVE VLAN10:[lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09](archive/lan-rb5009-upgrade.md#阶段-5-附pve-上-vlan10-透传实现-调研-2026-08-09)
|
||||
|
||||
@@ -108,3 +108,4 @@ Steps:
|
||||
- Matrix Authentication Service: <https://github.com/element-hq/matrix-authentication-service>
|
||||
- Matrix spec: <https://spec.matrix.org/>
|
||||
- Federation tester: <https://federationtester.matrix.org/>
|
||||
- Home Assistant Matrix integration: [home-assistant-matrix.md](home-assistant-matrix.md)
|
||||
|
||||
@@ -0,0 +1,215 @@
|
||||
# Matter 配网排障手册
|
||||
|
||||
> 基于 Matter 1.5.1 Core Spec §4.3.1 与本环境(EdgeRouter X + UniFi AP + Aqara M3 +
|
||||
> Home Assistant)2026-08-21 实测整理。配套 Linear W1N-207。
|
||||
|
||||
## 1. Matter 配网协议要点(发现即一切)
|
||||
|
||||
- **发现走 mDNS(DNS-SD)**,UDP **5353**,组播 `224.0.0.251` / `ff02::fb`。
|
||||
**不经过单播 DNS(如 AdGuard .36)、不需要反向 DNS、不需要 DHCPv6**(SLAAC 即满足 Matter
|
||||
的 IPv6 要求)。
|
||||
- 服务类型:
|
||||
- `_matterc._udp` — 可配网设备(Commissionable),**配对模式才有效**
|
||||
- `_matter._tcp` — 已配设备(Operational),TXT 里含 fabric 信息
|
||||
- 子类型(配对方按此过滤):
|
||||
- `_L<全12位 discriminator>`(如 `_L3266`)— 按二维码里的完整 discriminator 精确匹配
|
||||
- `_S<高4位>`(如 `_S12`)
|
||||
- `_V<vendorId>`、`_T<deviceType>`(可选)
|
||||
- `_CM`(仅真正处于配对模式时发布)
|
||||
- TXT 关键键:`D=`(discriminator,规范 **SHALL** 必填)、`VP=`(vendor+product)、
|
||||
**`CM=`**、`RI=`(rotating id)、`PH=`/`PI=`(配对提示)。
|
||||
- 配对端口:**TCP 5540**(PASE/CASE)。部分生态(Aqara M3)为 Thread 中继节点用 **5552**。
|
||||
- 实例名:64 位随机 hex;**进入配对模式时更换**(可用作"是否重新进过配对"的信号)。
|
||||
- 规范参考:[Matter 1.5.1 Core Spec §4.3.1](https://csa-iot.org/wp-content/uploads/2026/03/23-27349-010_Matter-1.5.1-Core-Specification.pdf)、
|
||||
[Google Home: Commissionable and Operational Discovery](https://developers.home.google.com/matter/primer/commissionable-and-operational-discovery)、
|
||||
[Matter Handbook: Discovery](https://handbook.buildwithmatter.com/how-it-works/discovery/)、
|
||||
[connectedhomeip: IP commissioning](https://pigweed.googlesource.com/third_party/github/project-chip/connectedhomeip/+show/59edd2ff8506b1e3dabb7040d716f0e75a2312d1/docs/guides/ip_commissioning.md)。
|
||||
|
||||
## 2. 关键判据:CM=0 = 不在配对模式
|
||||
|
||||
规范 §4.3.1.2 / §4.3.1.7:
|
||||
|
||||
- 设备可以长期宣告 `_matterc`(**Extended Discovery**),但 **`CM=0` 表示"当前不接受配网"**。
|
||||
- **已在 fabric 里的设备**(宣告里同时有 `_matter._tcp` + `_I<fabric>._sub` 运营记录)重配时
|
||||
通常报 `CM=0` —— 它已配好,不是新设备。
|
||||
- **配对方不能把已配设备当新设备加** → 重加/找回必须先**恢复出厂**(清 fabric,重启后以
|
||||
`CM=1` 全新配对模式宣告),再用**它自己的二维码**添加。
|
||||
- 常见误判:抓包看到 `_matterc` 宣告就以为"在配对模式"——**必须看 `CM=`**。
|
||||
|
||||
## 3. 本环境实测事实(2026-08-21,W1N-207)
|
||||
|
||||
| 事实 | 状态 |
|
||||
|---|---|
|
||||
| LAN55 IPv6/mDNS 链路 | ✅ 全正常(RA→交换机→AP→客户端;mDNS 双向通;igmp snooping off、mdns on、无客户端隔离、无组播增强、PMF off、WPA2、仅 2.4G) |
|
||||
| Matter 不依赖单播 DNS/.36、反向 DNS、DHCPv6 | ✅ 已排除(.36 健康且不在路径上) |
|
||||
| HA matter-server 曾宣告两代前的旧 GUA | ✅ 已修复(重启 `core_matter_server`;宣告恢复当前前缀) |
|
||||
| ISP PD /60 随重拨轮换 → Matter IPv6 缓存反复失效 | ⚠️ 环境性根因;对策 = 重拨后重启 matter-server + 重启 M3 |
|
||||
| EdgeOS 上静态 ULA 不可行 | ✅ 已尝试并回滚(switch0 不支持静态 `ipv6 address`;显式 router-advert 会替换 PD-slaac RA) |
|
||||
| 在用的两盏 ESP32-C2 Matter 灯泡(VP `0x4891/0x4100`;2026-08-23 复核) | 工作盏 MAC 已变为 `fc:e8:c0:25:a1:f0`(`.146`,hostname `espressif`;原 `34:98:7a:25:a1:f0` 全网消失,疑固件更新后换 MAC——末 3 字节相同);新盏 `34:98:7a:27:10:bc`(`.148`,hostname `matter`)。两盏各宣告 **3 个 fabric** 运营实例:Aqara `4DF2B1455D19402D`、`2F6E56020E1996E7`、HA `DCE86145C137AF0E`(见 §8) |
|
||||
| 故障盏 `34:98:7a:27:7f:08`(曾 .145,Aqara fabric,`CM=0` 缺 GUA) | 2026-08-23 复核:无租约、ARP incomplete、AP 无日志 = **已离网**(退役/退换) |
|
||||
| **失败模式 C(2026-08-23 实测,两盏同时)**:mDNS 活、5540 死 | 灯泡 ping 通(v4/v6)、DHCP 正常续租、mDNS 应答并宣告 `_matter._tcp`(SRV :5540、TXT `T=1`、当前前缀 GUA),但 **TCP 5540 在 IPv4 与 IPv6(fe80+GUA)均 RST 拒绝** → 配对方无法建立 CASE,App 显示离线;hass matter-server 侧无任何 established :5540 会话(详见 §8) |
|
||||
| ISP PD 前缀再次轮换(2026-08-23 → `240e:3bd:238:4812::/64`;08-22 为 `235:1fb2`) | hass 与 `.148` 均持当前前缀 GUA;hass 残留 `.146` 旧前缀 GUA 的 **FAILED** 邻居项(旧地址缓存仍被某端尝试) |
|
||||
| DHCP 保留 `matter`(.45 → MAC `…10:bc`) | ⚠️ 保留仍未生效:新灯泡(`…10:bc`)实际拿到动态 `.148` 而非保留的 `.45`(待修,见 hosts/gw.md) |
|
||||
| **新灯泡(2026-08-22 添加成功)**:MAC `34:98:7a:27:10:bc`(=DHCP 保留目标 MAC),hostname `matter`,IP `.148`,VP `4891/4100`,D=`3377` | ✅ 已入 **Aqara fabric `4DF2B1455D19402D`**;**经 BLE 配网**(Aqara Home App)——线上**无 TCP 5540** 属正常(BLE 会话对 AP/hass 抓包不可见) |
|
||||
| ESP32-C2 灯泡 firmware 挂死模式(2026-08-22 实测) | 入网后宣告 `_matterc`(CM=1、D=3377)约 **3 秒后网络栈完全静默**:STA 收发计数冻结、不掉线不重启、配对方(手机/M3 `_L3377` 查询)无应答 → 加不上。**对策=断电 10 秒重启**重新进配网模式(实例名更换:`3F4E2C66F2DA85CD`→`E5BA8E28E4DE23A0`),随即 App 添加即成功 |
|
||||
| 遗留 SSID(element/vwire/vport) | ✅ 已清理 |
|
||||
|
||||
## 4. 抓包方法(BusyBox 兼容)
|
||||
|
||||
> 完整指令集(实时 / 落盘轮转 / 定向抓取 / Wireshark 解密)见
|
||||
> [runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md)。
|
||||
> 下面是最常用的两条。
|
||||
|
||||
**视角必须在 LAN55**。**HA matter-server 作配对方时推荐直接在 hass `end0` 抓**——配对方
|
||||
必然参与配对流程的每一条通讯(mDNS 本段组播 + 自己的 TCP 5540 全程),覆盖最全;AP `br0`
|
||||
能看到全部 mDNS 组播 + 无线客户端单播,但**看不到有线↔有线单播**(如 Thread 设备经有线 M3
|
||||
配对时 HA↔M3 的 5540 在 AP 侧不可见)。66 网段电脑看不到 55 的组播。BusyBox 注意点仅适用
|
||||
AP(**不要用 `--line-buffered`**;引号外层双引号、内层单引号);hass 是 HAOS 全量 tcpdump。
|
||||
|
||||
完整抓取(跑配对时保持窗口开着,`Ctrl+C` 结束):
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
hass 侧(HA matter-server 作配对方,推荐;非交互 ssh 需显式 `sudo -n -i`):
|
||||
|
||||
```bash
|
||||
ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
精简过滤(只看 Matter 信号):
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -vvv -tt 'udp port 5353 or tcp port 5540 or tcp port 5552' | grep -E '_matterc|_matter|_L[0-9]+|_S[0-9]+|_CM|_V[0-9]+|_T[0-9]+|\.5540|\.5552'"
|
||||
```
|
||||
|
||||
存 pcap 供 Wireshark:把上面 `-w /tmp/matter.pcap` 追加到 tcpdump 参数(去掉 `-vvv`),
|
||||
`scp zhiqiangf@192.168.55.5:/tmp/matter.pcap .` 拉回本地分析。
|
||||
|
||||
> **落盘务必轮转**:AP `/tmp` 只有约 60MB。用
|
||||
> `-C 5 -W 12 -w /tmp/matter.pcap`(每 5MB 轮转、最多 12 个文件)防止写满,
|
||||
> 详见 runbook Step 3(落盘轮转)。
|
||||
> **Matter 载荷是加密的**:mDNS(5353)明文可读;5540 上的 Matter 报文要看明文
|
||||
> 需要 Wireshark matter-dissector + 会话密钥,详见 runbook Step 5(解密)。
|
||||
|
||||
### 阶段对照表
|
||||
|
||||
| 阶段 | 应该看到 | 对应问题 |
|
||||
|---|---|---|
|
||||
| 发现(设备侧) | `_matterc._udp` + `_L3266._sub` + `_S12._sub` + TXT `D=3266 CM=1` + SRV `:5540` + AAAA | **无宣告**=设备没入网/没进配对模式;**`CM=0`**=不在配对模式(已配设备);**无 `_L3266`**=固件子类型缺失 |
|
||||
| 发现(配对方侧) | M3/手机查询 `_L3266._sub._matterc._udp` | 查询有、无应答 = 码/discriminator 不匹配或设备不在线 |
|
||||
| 配对握手 | 到设备 IP **TCP 5540 SYN/SYN-ACK** 双向 | **SYN 无 ACK**=设备不可达/防火墙;**完全无 5540**=发现阶段没完成 |
|
||||
| 配完后 | 设备宣告 `_matter._tcp` + `_I<fabric>._sub` | 出现 = 已入网成功 |
|
||||
| BLE 配网(手机 App 直连设备 BLE,如 Aqara Home) | 线上**无 TCP 5540**(BLE 会话对 AP/hass 抓包不可见);设备入网后仍先 mDNS 宣告 `_matterc` | 成功判据=最终宣告 `_matter._tcp` + `_I<fabric>._sub`;无 5540 **不代表**失败 |
|
||||
|
||||
## 5. 排障决策树(按顺序)
|
||||
|
||||
1. 抓包看**有没有 `_matterc` 宣告**:没有 → 设备不通电 / 没连上 Wi-Fi / 没进配对模式
|
||||
(先解决"设备在线",网络侧已反复验证正常)。
|
||||
2. 有宣告但 **`CM=0`** → 设备已配 / 不在配对模式 → **恢复出厂**后重试(用它自己的二维码)。
|
||||
3. 有宣告 `CM=1` 但**无 `_L<disc>` 子类型** → 固件 mDNS 缺陷 → 升固件或换通用发现配对方。
|
||||
4. `CM=1` + 子类型齐全但**无 TCP 5540** → 配对方没匹配上(查码/discriminator)或设备不可达。
|
||||
5. 有 5540 但配对中断 → 查 `CM` 源(码是否正确)、设备电源、fabric 状态(是否需先清)。
|
||||
|
||||
## 6. 相关文档
|
||||
|
||||
- [runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md) — Matter 抓包指令集(实时/落盘轮转/定向/解密)
|
||||
- [docs/lan-overview.md](lan-overview.md) — LAN 拓扑、SSID 清理、ULA 不可行
|
||||
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — matter-server 重拨运维规范
|
||||
- [docs/unifi-network.md](unifi-network.md) — UniFi 网络/IPv6/SSID 记录
|
||||
- [hosts/gw.md](../hosts/gw.md) — DHCP 保留 `matter` MAC 错位(待修)
|
||||
|
||||
## 7. 2026-08-22 实测记录:添加新 ESP32-C2 Matter 灯泡(成功 + 失败路径全记录)
|
||||
|
||||
> 场景:手机 App(Aqara Home)添加一盏**新的** ESP32-C2 Matter 灯泡
|
||||
> `34:98:7a:27:10:bc`(hostname `matter`,最终 IP `.148`)。中途换了灯泡并断电重启,
|
||||
> 共经历 **2 种失败模式** 和 **1 条成功路径**,全部抓包实证。
|
||||
>
|
||||
> 抓包点:UAP-AC-Lite `192.168.55.5` `br0`(轮转 `udp 5353 or tcp 5540 or tcp 5552`
|
||||
> + 定向全量 `ether host 34:98:7a:27:10:bc`)+ hostapd/stahtd 日志 + gw DHCP/ARP 交叉验证。
|
||||
> 本地用 tshark 4.7.2 分析。
|
||||
|
||||
### 时间线(CST,2026-08-22)
|
||||
|
||||
| 时间 | 事件 | 判据 / 说明 |
|
||||
|---|---|---|
|
||||
| 10:50:28 | 启动轮转抓包 | — |
|
||||
| 10:51:12–19 | 手机 `.143`(OnePlus,连 wifi0ap0=`ubnt-windy-2`)重新关联;查询 `_matter._tcp` | 运营查询(浏览已配设备),**不是**配网(配网应查 `_matterc._udp`) |
|
||||
| ~10:56 | 用户报「配置 wifi 后挂起,不能加入 wifi」 | 首次失败 |
|
||||
| 11:00:15 | M3 `.248` 查询 `_L3266._sub._matterc` | 无应答(是另一台设备的 discriminator,无关) |
|
||||
| 11:01:24–42 | **失败模式 A**:故障盏 `…7f:08` 尝试关联 `wifi0ap1`(`ubnt-haas`):发 1 次 open-auth 帧(algorithm 0)→ AP 回 `status_code=0` → **客户端不再发 assoc 请求** → 18s 后 `auth_failures=1` + disassociated | auth 阶段卡死(client 侧);非密码错——密码错会先 assoc 再 4-way 失败 |
|
||||
| 11:04:01–05 | **新灯泡 `…10:bc` 关联 `wifi0ap1` 成功**,WPA2 4-way 完成,DHCP 拿 `.148`(tracker `soft failure`: ip_delta 3.76s,avg_rssi -68);随即宣告 `_matterc`:实例 `3F4E2C66F2DA85CD`,TXT `VP=4891+4100 D=3377 CM=1`,SRV :5540,**有 GUA** | 发现阶段判据全过 |
|
||||
| 11:04:05 之后 | **失败模式 B**:灯泡网络栈完全静默——STA 收发计数冻结(rx=89/tx=5 持续 12s+ 不变)、**不掉线不重启** | firmware 挂死 |
|
||||
| 11:04:59–11:07:27 | 手机查 `_matterc` ×4、`.60` 解析实例、M3 查 `_L3377._sub._matterc` ×5(discriminator 3377 正是新盏)——**全部无应答**;TCP 5540/5552 全程 0 | 配对方找不到设备 → App 报「加不上」 |
|
||||
| 11:09:35–48 | **断电 10 秒重启**:灯泡重新关联 `wifi0ap1` ×2 | 对策生效 |
|
||||
| 11:10:26 | DHCP 重新拿 `.148`;STA 计数恢复持续增长(活跃) | — |
|
||||
| 11:10–11:15 | 重新宣告 `_matterc`(**新实例 `E5BA8E28E4DE23A0`**——入配对模式实例名更换,符合规范);App 走 **BLE 配网** | 线上无 TCP 5540(BLE 对 AP 不可见,属正常) |
|
||||
| 11:15:23 | 灯泡宣告 **`_matter._tcp`**:`4DF2B1455D19402D-02EF2FF12DFAF10E`(Aqara fabric)+ SRV :5540 + GUA + A | ✅ **添加成功**(已入 Aqara fabric `4DF2B1455D19402D`) |
|
||||
|
||||
### 结论与经验
|
||||
|
||||
1. **同族灯泡(VP 4891/4100,OUI 34:98:7a)存在两种不同失败模式**:
|
||||
- 故障盏 `…7f:08`:auth 阶段卡死(auth 帧后不发 assoc);此前(08-21 21:27)成功关联后伴随
|
||||
`ip_failures=1`(拿不到 IP)+ 缺 GUA —— 属更深层故障,需恢复出厂,本次未处理,仍离线。
|
||||
- 新盏 `…10:bc`:入网 + 宣告 `_matterc`(CM=1)成功后约 3 秒固件挂死(全静默)。
|
||||
**断电 10 秒重启即恢复**,是最简单有效的对策。
|
||||
2. **AP 抓包看不到 BLE 配网**:Aqara Home App 对 WiFi Matter 设备走 BLE 配网时,线上只有
|
||||
mDNS/DHCP,**无 TCP 5540 不代表失败**;成功判据 = 设备最终宣告 `_matter._tcp` + `_I<fabric>._sub`。
|
||||
3. **发现判据回顾**:`_matterc` + TXT(`CM=1`、`D=`、`VP=`)+ SRV :5540 + AAAA(GUA) + A 全齐才算
|
||||
设备真的在配对模式;配对方按 `_L<disc>._sub._matterc` 精确匹配 discriminator(本例 D=3377)。
|
||||
4. 新盏 RSSI -68、DHCP 3.76s,射频偏弱,可能加剧 firmware 不稳定(待观察)。
|
||||
5. DHCP 保留 `matter`(.45→`…10:bc`)**仍未生效**:新盏实际拿动态 `.148`(待修,见 hosts/gw.md)。
|
||||
6. 识别「配对方在找但设备不答」的快速方法:抓包里配对方持续查 `_matterc`/`_L<disc>` 而目标 MAC
|
||||
零应答 + STA 收发计数冻结 = 设备侧挂死;此时**先断电重启设备**,不要怀疑网络/AP。
|
||||
|
||||
### 后续:新盏 11:24 起离线循环(同一盏 `…10:bc`,2026-08-22)
|
||||
|
||||
配网成功后约 10 分钟(11:15–11:24 可控制),灯泡进入**持续性故障循环**:
|
||||
|
||||
| 时间 | 事件 | 模式 |
|
||||
|---|---|---|
|
||||
| 11:24:57 | `EVENT_STA_LEAVE`(真掉线) | 掉线 |
|
||||
| 11:25:12 | 重连 `auth_failures=2` | **auth 卡死**(同故障盏 `…7f:08` 11:01 的模式) |
|
||||
| 11:25:23–25 | 重连成功,WPA2 完成,重新拿 `.148` | — |
|
||||
| 11:25:38 | `soft failure`,ip_delta 2.65s,**avg_rssi -73**(-68→-73 持续变差) | 射频偏弱 |
|
||||
| 11:25 之后 | STA 计数冻结(rx=119/tx=64 不动);M3 持续查询其运营实例 `4DF2B1455D19402D-02EF2FF12DFAF10E._matter._tcp` **无应答** → App 显示「离线」 | 静默挂死 |
|
||||
|
||||
**结论**:三盏 ESP32-C2 灯泡中两盏(`…7f:08`、`…10:bc`)故障,表现覆盖 auth 卡死 / 静默挂死 /
|
||||
随机掉线三种形态;工作盏 `…25:a1:f0` 正常。网络侧(AP、M3、DHCP、mDNS)均验证正常。
|
||||
**疑似根因(按可能性)**:① ESP32-C2 Matter 灯泡 firmware 缺陷(同批次)② 射频偏弱
|
||||
(RSSI -73,天线/距离/遮挡)加剧不稳定 ③ 供电不稳(brownout 造成 Wi-Fi 栈崩溃重启)。
|
||||
**待办**:移近 AP 或改善供电后观察;App 内查固件更新;仍复发则考虑退换。
|
||||
|
||||
## 8. 2026-08-23 状态核查:两盏「半在线」——mDNS 宣告正常但 TCP 5540 无监听(失败模式 C)
|
||||
|
||||
> 全程**只读**核查(gw DHCP/ARP、AP hostapd 日志、hass matter-server 状态 + mDNS 抓包、
|
||||
> 对灯泡 v4/v6 的 TCP 5540 探测,09:0x CST)。结论:**网络侧全部健康;两盏灯泡网络栈活着、
|
||||
> mDNS 运营宣告正常,但 Matter 会话端点(TCP 5540)无监听**——配对方无法建立 CASE,
|
||||
> App 内应显示离线/不可达。
|
||||
|
||||
| 对象 | 状态(2026-08-23) |
|
||||
|---|---|
|
||||
| 新盏 `34:98:7a:27:10:bc`(`.148`,hostname `matter`) | DHCP 04:40 续租;gw ARP 完整;ping 通(93–122ms,ESP32 省电时延);08-22 16:40 起稳定关联 `wifi0ap1`,关联时 `avg_rssi -70`。mDNS 宣告 3 实例:`4DF2B1455D19402D-02EF079EEB480D07`(**新 node ID——08-22 之后被重新配网过**)、`2F6E56020E1996E7-137147AF27BE4EB6`、`DCE86145C137AF0E-0000000000000011`(HA fabric);host 记录 A `.148` + fe80 + **当前前缀** GUA `240e:3bd:238:4812:*`。支持单播 legacy mDNS 查询(`dig -p 5353 @.148 _matter._tcp.local PTR` 可用) |
|
||||
| 工作盏(MAC 已变)`fc:e8:c0:25:a1:f0`(`.146`,hostname `espressif`) | DHCP 07:17 续租;ping 通 v4/v6(v6 fe80 38–61ms)。mDNS 宣告 3 实例:`4DF2B1455D19402D-02EF4CA3F856B615`、`2F6E56020E1996E7-EE8F2E4F1A77BF05`、`DCE86145C137AF0E-000000000000000B`。原 MAC `34:98:7a:25:a1:f0` 全网消失(无租约/ARP/AP 日志)而新 MAC 末 3 字节相同 → 疑固件更新后改 MAC。**拒绝单播 5353**(ICMP port unreachable),只应答组播查询——同族固件行为差异。hass 残留其旧前缀 GUA `240e:3bd:235:1fb2:fee8:c0ff:fe25:a1f0` 的 **FAILED** 邻居项 |
|
||||
| 故障盏 `34:98:7a:27:7f:08`(曾 `.145`) | 无租约、ARP incomplete、AP 日志零事件 = 已离网 |
|
||||
| **TCP 5540 探测(两盏)** | IPv4(LAN66 与 hass 本段)、IPv6(fe80%end0 + 当前 GUA)全部 **RST(Connection refused)** —— SRV 宣告 :5540 且 TXT `T=1`,但实际无监听 |
|
||||
| hass matter-server | `started`,v9.0.4,无更新;宣告自身运营实例 `DCE86145C137AF0E-…1B669`(v4+v6,当前 GUA);**无任何 established :5540 会话**;core/add-on 日志无 matter 错误 |
|
||||
| 其他 Matter 控制器 | Aqara M3 `.248` 在线(有线 0.8ms),宣告含自身 fabric 节点 `4DF2B1455D19402D-11E158E46D24A000`;SmartThings `.48` 在线并周期查询 `_matter._tcp.local`;手机(当前前缀 GUA)也在浏览。LAN55 共见 **5 个 fabric**:`4DF2B1455D19402D`(M3)、`DCE86145C137AF0E`(HA)、`2F6E56020E1996E7`、`03BCFAEDD6153944`、`6A6FF80C2DB84DEE` |
|
||||
|
||||
**判定**:失败模式 C = TCP/IP 栈与 mDNS 守护进程活着(主动 RST、DHCP 续租、ping 通),
|
||||
但 Matter 应用层监听不存在。与模式 A(auth 卡死)、模式 B(全静默挂死)同族不同形态;
|
||||
**两盏同时处于同一状态**更指向共同诱因(固件缺陷,或 PD 轮换等共同事件后未恢复)。
|
||||
**对策(推荐,未执行)**:逐盏断电 10 秒重启(模式 B 的已验证对策),重启后复测
|
||||
TCP 5540 恢复监听即可确认。
|
||||
|
||||
**核查方法备忘**(只读,可复用):
|
||||
|
||||
- gw:`show dhcp leases` / `show arp`(经 `/opt/vyatta/bin/vyatta-op-cmd-wrapper`)。
|
||||
- AP:`grep -i <mac> /var/log/messages`(hostapd 关联事件 + stahtd RSSI/soft failure)。
|
||||
- hass:`sudo -n -i ha apps info core_matter_server`;`ip -6 neigh show dev end0`
|
||||
(看灯泡 fe80/旧新前缀 GUA 与 FAILED 项);被动抓包
|
||||
`sudo -n -i timeout 65 tcpdump -ni end0 -s 0 -tt 'udp port 5353'`——配对方周期查询
|
||||
会自然引出灯泡宣告,无需主动发包。
|
||||
- 5540 探测:hass 上 python3 对 v4 / fe80%end0 / GUA 各 connect 一次;RST=无监听,
|
||||
超时=不可达(两者含义不同)。
|
||||
@@ -0,0 +1,45 @@
|
||||
# Plane CE 加固草稿(docs/plane-hardening/)
|
||||
|
||||
> **状态:草稿,未应用、未提交。** 对应追踪:Plane vps 项目条目(2026-09-03,**记录源**;Linear W1N-277 已取消,Linear 自 2026-09-03 起不再作为记录源)。
|
||||
> 线上实例:`plane.chans.xyz`(synapse K3s,ns `plane`,release `plane-app` = chart `plane-ce-1.8.0` / app `v1.4.1`)。
|
||||
> 依据:2026-09-03 只读核查(13 条审查意见中 11 条属实、#3 基本属实、#9 指标归属错误)+ 上游 chart 模板逐条核对。
|
||||
|
||||
## 文件
|
||||
|
||||
| 文件 | 内容 |
|
||||
|------|------|
|
||||
| `values.hardened.yaml` | 可选硬化 values(external secrets 引用、requireExplicitSecrets、minio pin、上传限额对齐);含 HTTP→HTTPS `extraObjects` 示例 |
|
||||
| `secrets.yaml.example` | 6 组外部 Secret 结构占位(只含 key 名,真实值仅存宿主机) |
|
||||
| `backup/plane-backup.yaml` | **PostgreSQL 备份 CronJob**(pg_dump `-Fc`,hostPath `/var/backups/plane`;MinIO 已按实际用量剔除) |
|
||||
| `backup/README.md` | 备份方案说明(排程/容量/保留/还原/阻塞) |
|
||||
|
||||
## 应用顺序(每步先 diff 后执行,全部需用户逐项确认)
|
||||
|
||||
### 现在就值得做:DB 备份(P0,见 backup/)
|
||||
`plane-backup.yaml` 部署 + 手动触发验证一次即可;88 MB 库每日快照几乎零成本。
|
||||
|
||||
### 可选(顺手做一次,不是必须)
|
||||
- **Phase A 密钥外部化**(零行为变化、无停机,约 15 分钟):按 `secrets.yaml.example`
|
||||
在宿主机建 6 个 Secret(值先复制当前集群),用 `values.hardened.yaml` 跑
|
||||
`helm diff upgrade` → `helm upgrade`;验证后删除 chart 生成的旧 Secret。
|
||||
价值:默认密钥不再落在 chart 公开常量上,作为保险。
|
||||
- **MCP API Key 轮换**:若审查对话出过你的环境,Plane 后台重生成 + 更新
|
||||
`/home/windy/plane-k3s/mcp/mcp.env`(0600)+ 重启 Cursor MCP。
|
||||
- **/god-mode IP 白名单**:若在意管理后台被公网爆破。chart 1.8.0 的 IngressRoute
|
||||
不支持给单条路由追加 middleware → 需 post-renderer 或 upgrade 后 `kubectl patch`
|
||||
(升级会覆盖,需固化);源 IP 清单待提供。
|
||||
|
||||
### 明确暂缓/跳过(个人单节点,等出现症状再处理)
|
||||
- SECRET_KEY 等轮换(Phase B):等真要配 SMTP/OAuth 前再做(避免旧密文不可解)。
|
||||
- NetworkPolicy、有状态组件 resources limits(chart 无 values 开关,需 post-render/patch)、
|
||||
HTTP→HTTPS(草稿已给 `extraObjects` 示例)、metrics-server/Sentry。
|
||||
|
||||
## 关键限制(chart 1.8.0 模板已核对)
|
||||
- `external_secrets.*_existingSecret` 设置后,对应 Secret **必须**包含模板所需全部 key
|
||||
(缺失不自动补),见 `secrets.yaml.example` 注释。
|
||||
- `app_keys_existingSecret` 的 envFrom 在所有 workload 上**最后注入**(后置生效),
|
||||
保证 app/live 共享密钥一致——不要在其后再放同名 key 的 Secret。
|
||||
- `DATABASE_URL`/`AMQP_URL`/`REDIS_URL` 是 chart 生成的派生 URL,内嵌明文密码;
|
||||
外部化后轮换 DB/队列密码时必须同步更新 `plane-app-env`。
|
||||
- minio 的 `MINIO_ROOT_*` 与 `AWS_*` 同源于一个 Secret;升级时 bucket Job 会重跑
|
||||
(需 admin 权限凭据)——换 svcacct 前先确认权限覆盖该 Job。
|
||||
@@ -0,0 +1,48 @@
|
||||
# Plane CE 备份方案(DB-only)— DRAFT (2026-09-03), 未应用
|
||||
|
||||
> 关联:`plane-backup.yaml`(CronJob);追踪:Plane vps 项目条目(记录源,2026-09-03 起不用 Linear)。
|
||||
> 现状(实测):pg 全库 **88 MB**(310 issues / 1 user);MinIO uploads **264 KB**(几乎空)。
|
||||
|
||||
## 范围决策(2026-09-03,实际角度)
|
||||
|
||||
- **做:PostgreSQL 逻辑备份** —— 覆盖现实故障(误删、升级失败、磁盘坏、重装),成本≈0。
|
||||
- **不做:MinIO/附件备份** —— 桶仅 264 KB,个人实例附件可接受丢失;不为它付日常维护。
|
||||
日后附件明显变多再按原完整版思路加 `mc mirror`(历史版本见本目录 git 历史/Plane 条目评论)。
|
||||
- 异机同步暂不启用(见下"局限/阻塞")。
|
||||
|
||||
## 方案
|
||||
|
||||
集群内 CronJob(ns `plane`,每天 **01:30 UTC = 03:30 本地**,控制器按 UTC 跑):
|
||||
|
||||
1. 单容器 `postgres:15.7-alpine`:`pg_dump -Fc`(自定义压缩格式)打 `plane` 库
|
||||
→ `/var/backups/plane/pg/plane-<UTC时间戳>.dump`(hostPath `DirectoryOrCreate`)
|
||||
2. 保留 7 天(`find -mtime +7 -delete`),成功/失败历史各留 3/2
|
||||
3. 凭据:现 chart Secret `plane-app-pgdb-secrets`(Phase A 外部化后改 `plane-pgdb-credentials`)
|
||||
|
||||
## 容量
|
||||
|
||||
- 库 88 MB → `-Fc` 快照约 10–40 MB/天 × 7 天 ≈ **<300 MB**,对 83 G 可用盘可忽略。
|
||||
|
||||
## 还原(未演练;应用前先做一次隔离测试)
|
||||
|
||||
```bash
|
||||
# 目标 PG15 实例(临时起一个 postgres:15.7-alpine 容器或另一台机):
|
||||
# 先建空库: createdb plane (user=plane)
|
||||
pg_restore -h <target> -U plane -d plane --clean --if-exists /var/backups/plane/pg/plane-<TS>.dump
|
||||
# 还原后确认 310 issues 量级一致;附件为空属预期(未备份 MinIO)
|
||||
```
|
||||
|
||||
## 验收(应用前逐项过)
|
||||
|
||||
- [ ] CronJob 建立后手动触发一次:`kubectl -n plane create job --from=cronjob/plane-backup plane-backup-manual-1`,Job `Completed`
|
||||
- [ ] `/var/backups/plane/pg/plane-*.dump` 可被 `pg_restore -l` 列出
|
||||
- [ ] 备份 Job 只依赖 pgdb 服务,不依赖 Plane 应用 Pod(应用故障期间也能出备份)
|
||||
- [ ] 保留清理 dry-run(`find ... -print`)正确;`df -h /` 前后对比记录
|
||||
|
||||
## 局限 / 阻塞
|
||||
|
||||
- **本地方案不是离机备份**:单节点磁盘/整机故障即丢。如日后要离机,纳入
|
||||
[Restic 异机 repository 决策与存取隔离](https://plane.chans.xyz/space/projects/56874283-7e1d-43a8-afa4-631cf1c4ad5b/issues/7825d564-ae15-446b-bced-be26b648346b/)
|
||||
(与 Matrix 备份同一决策);恢复演练纪律见
|
||||
[服务级 restore runbook 与隔离复元演练](https://plane.chans.xyz/space/projects/56874283-7e1d-43a8-afa4-631cf1c4ad5b/issues/a9bea3ba-c958-4a74-b2f1-6bbb653f21d3/)。
|
||||
- 提醒:同一节点 **Matrix 数据价值远高于 Plane 且同样无备份** —— 若投入备份精力,顺序上 Matrix 优先。
|
||||
@@ -0,0 +1,63 @@
|
||||
# Plane CE PostgreSQL backup CronJob — DRAFT (2026-09-03), NOT applied.
|
||||
# ns: plane (synapse K3s single node). Output: hostPath /var/backups/plane (root disk, auto-created).
|
||||
#
|
||||
# Scope decision (2026-09-03, practical): DB-only. MinIO dropped — uploads bucket
|
||||
# measured at 264 KB / 444 KB total; attachments are acceptable loss for this
|
||||
# personal 1-user instance (310 issues / 88 MB DB). Revisit only if usage grows.
|
||||
#
|
||||
# Credentials: read from the CURRENT chart-generated Secret (works today). After the
|
||||
# optional external-secrets migration (docs/plane-hardening/README.md Phase A) switch
|
||||
# the secretKeyRef name to plane-pgdb-credentials.
|
||||
#
|
||||
# Apply:
|
||||
# ssh windy@synapse.chans.xyz 'sudo k3s kubectl apply -n plane -f -' < plane-backup.yaml
|
||||
# Manual run + verify:
|
||||
# sudo k3s kubectl -n plane create job --from=cronjob/plane-backup plane-backup-manual-1
|
||||
# sudo k3s kubectl -n plane get cronjob,job,pods | grep plane-backup
|
||||
# sudo ls -lh /var/backups/plane/pg
|
||||
# Restore steps + tuning: see backup/README.md
|
||||
|
||||
apiVersion: batch/v1
|
||||
kind: CronJob
|
||||
metadata:
|
||||
name: plane-backup
|
||||
namespace: plane
|
||||
spec:
|
||||
# 01:30 UTC daily = 03:30 local (CEST). CronJob controller runs in UTC.
|
||||
schedule: "30 1 * * *"
|
||||
concurrencyPolicy: Forbid
|
||||
successfulJobsHistoryLimit: 3
|
||||
failedJobsHistoryLimit: 2
|
||||
jobTemplate:
|
||||
spec:
|
||||
backoffLimit: 2
|
||||
template:
|
||||
spec:
|
||||
restartPolicy: OnFailure
|
||||
volumes:
|
||||
- name: backup
|
||||
hostPath:
|
||||
path: /var/backups/plane
|
||||
type: DirectoryOrCreate
|
||||
containers:
|
||||
- name: pg-dump
|
||||
image: postgres:15.7-alpine
|
||||
env:
|
||||
- name: PGPASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: plane-app-pgdb-secrets # -> plane-pgdb-credentials after Phase A
|
||||
key: POSTGRES_PASSWORD
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -euo pipefail
|
||||
TS=$(date -u +%Y%m%dT%H%M%SZ)
|
||||
mkdir -p /backup/pg
|
||||
pg_dump -h plane-app-pgdb.plane.svc.cluster.local -U plane -d plane \
|
||||
-Fc -f "/backup/pg/plane-${TS}.dump"
|
||||
find /backup/pg -type f -name 'plane-*.dump' -mtime +7 -delete
|
||||
echo "pg_dump done: /backup/pg/plane-${TS}.dump ($(du -h /backup/pg/plane-${TS}.dump | cut -f1))"
|
||||
volumeMounts:
|
||||
- name: backup
|
||||
mountPath: /backup
|
||||
@@ -0,0 +1,91 @@
|
||||
# External Secret structure for Plane CE hardening — EXAMPLE ONLY.
|
||||
# No real values here; this file is safe to commit. Real values live only on the
|
||||
# host (/home/windy/plane-k3s, 0600/0700) and in the cluster.
|
||||
#
|
||||
# Phase A — create each Secret with the CURRENT cluster values first (zero change):
|
||||
# # current source Secrets (chart-generated):
|
||||
# kubectl -n plane get secret plane-app-app-secrets -o jsonpath='{.data.SECRET_KEY}' | base64 -d
|
||||
# kubectl -n plane get secret plane-app-live-secrets -o jsonpath='{.data.REDIS_URL}' | base64 -d
|
||||
# kubectl -n plane get secret plane-app-pgdb-secrets -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d
|
||||
# kubectl -n plane get secret plane-app-rabbitmq-secrets -o jsonpath='{.data.RABBITMQ_DEFAULT_PASS}' | base64 -d
|
||||
# kubectl -n plane get secret plane-app-doc-store-secrets -o jsonpath='{.data}' | base64 -d
|
||||
#
|
||||
# e.g. kubectl -n plane create secret generic plane-app-keys \
|
||||
# --from-literal=SECRET_KEY="$(<copy from above>)" \
|
||||
# --from-literal=LIVE_SERVER_SECRET_KEY="$(<copy from above>)"
|
||||
#
|
||||
# All keys below are REQUIRED by chart templates/plane-ce-1.8.0 (verified 2026-09-03):
|
||||
# missing keys are NOT auto-filled once an existingSecret is referenced.
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: plane-app-keys # external_secrets.app_keys_existingSecret
|
||||
namespace: plane
|
||||
type: Opaque
|
||||
stringData:
|
||||
SECRET_KEY: "" # current: copy from plane-app-app-secrets; rotate only in Phase B
|
||||
LIVE_SERVER_SECRET_KEY: "" # current: same value as above / plane-app-live-secrets
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: plane-app-env # external_secrets.app_env_existingSecret
|
||||
namespace: plane
|
||||
type: Opaque
|
||||
stringData:
|
||||
REDIS_URL: "" # redis://plane-app-redis.plane.svc.cluster.local:6379/
|
||||
DATABASE_URL: "" # postgresql://plane:plane@plane-app-pgdb.plane.svc.cluster.local/plane
|
||||
AMQP_URL: "" # amqp://plane:plane@plane-app-rabbitmq.plane.svc.cluster.local/
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: plane-live-env # external_secrets.live_env_existingSecret
|
||||
namespace: plane
|
||||
type: Opaque
|
||||
stringData:
|
||||
REDIS_URL: "" # redis://plane-app-redis.plane.svc.cluster.local:6379/
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: plane-pgdb-credentials # external_secrets.pgdb_existingSecret
|
||||
namespace: plane
|
||||
type: Opaque
|
||||
stringData:
|
||||
POSTGRES_PASSWORD: "" # Phase A: keep current ('plane'); Phase B: ALTER USER first, then sync
|
||||
POSTGRES_DB: "plane"
|
||||
POSTGRES_USER: "plane"
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: plane-rabbitmq-credentials # external_secrets.rabbitmq_existingSecret
|
||||
namespace: plane
|
||||
type: Opaque
|
||||
stringData:
|
||||
RABBITMQ_DEFAULT_USER: "plane"
|
||||
RABBITMQ_DEFAULT_PASS: "" # Phase A: keep current; Phase B: rabbitmqctl change_password first
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: plane-minio-credentials # external_secrets.doc_store_existingSecret
|
||||
namespace: plane
|
||||
type: Opaque
|
||||
stringData:
|
||||
FILE_SIZE_LIMIT: "20971520" # must match env.doc_upload_size_limit
|
||||
AWS_S3_BUCKET_NAME: "uploads"
|
||||
USE_MINIO: "1"
|
||||
MINIO_ROOT_USER: "admin"
|
||||
MINIO_ROOT_PASSWORD: "" # root creds take effect on first init only
|
||||
AWS_ACCESS_KEY_ID: "admin"
|
||||
AWS_SECRET_ACCESS_KEY: "" # == MINIO_ROOT_PASSWORD while minio.local_setup
|
||||
AWS_S3_ENDPOINT_URL: "http://plane-app-minio:9000"
|
||||
@@ -0,0 +1,109 @@
|
||||
# Plane CE hardened values — DRAFT (2026-09-03), NOT applied.
|
||||
# Target file on host: /home/windy/plane-k3s/values.yaml (synapse.chans.xyz)
|
||||
# Reference release: plane-app, chart plane-ce-1.8.0 (values.yaml L1-362 + templates verified 2026-09-03).
|
||||
# No secrets in this file. Secret *values* live only in k8s Secrets (see secrets.yaml.example).
|
||||
#
|
||||
# Two phases:
|
||||
# Phase A: externalize secrets (reference names below) with CURRENT values copied -> zero change.
|
||||
# Phase B: rotate credentials one by one (see README.md). SECRET_KEY rotation is cheap only while
|
||||
# SMTP/OAuth are unconfigured (no encrypted config rows yet).
|
||||
|
||||
planeVersion: v1.4.1
|
||||
|
||||
ingress:
|
||||
enabled: true
|
||||
appHost: plane.chans.xyz
|
||||
ingressClass: traefik
|
||||
traefik:
|
||||
# 20 MiB (chart default). Keep aligned with env.doc_upload_size_limit below.
|
||||
maxRequestBodyBytes: 20971520
|
||||
|
||||
ssl:
|
||||
createIssuer: true
|
||||
issuer: http # HTTP-01; ssl_token_existingSecret not needed
|
||||
email: admin@chans.xyz
|
||||
generateCerts: true
|
||||
|
||||
postgres:
|
||||
storageClass: local-path
|
||||
volumeSize: 5Gi
|
||||
# NOTE: chart 1.8.0 exposes NO resources knob for the bundled datastores
|
||||
# (stateful templates render no resources block). Add limits via
|
||||
# --post-renderer/kustomize or `kubectl -n plane patch sts ...` re-applied on
|
||||
# every upgrade (P2 task; see README.md).
|
||||
|
||||
redis:
|
||||
storageClass: local-path
|
||||
# image: valkey/valkey:7.2.11-alpine # already pinned by chart default; uncomment to make explicit
|
||||
|
||||
minio:
|
||||
# P2: pin. Digest of the currently running :latest (2026-09-03, pod plane-app-minio-wl-0).
|
||||
image: minio/minio@sha256:14cea493d9a34af32f524e538b8346cf79f3321eff8e708c1e2960462bd8936e
|
||||
# image_mc: minio/mc@sha256:... # optional: pin one-shot bucket-init client the same way
|
||||
storageClass: local-path
|
||||
volumeSize: 5Gi
|
||||
|
||||
rabbitmq:
|
||||
storageClass: local-path
|
||||
|
||||
env:
|
||||
# Fail the render instead of ever falling back to the chart's PUBLIC constants
|
||||
# (values.yaml L340-341 in chart 1.8.0). Requires external_secrets below.
|
||||
requireExplicitSecrets: true
|
||||
|
||||
# SECRET_KEY / LIVE_SERVER_SECRET_KEY are deliberately OMITTED here.
|
||||
# They live in k8s Secret `plane-app-keys` (referenced below). With
|
||||
# requireExplicitSecrets=true and app_keys_existingSecret set, the chart renders
|
||||
# neither key itself and app+live workloads both envFrom `plane-app-keys` LAST
|
||||
# (later envFrom wins), which keeps the shared signing key consistent.
|
||||
|
||||
pgdb_name: plane
|
||||
docstore_bucket: uploads
|
||||
# Align app-side upload cap with the Traefik body limit (was 5242880/5MiB).
|
||||
# Keep both at 20MiB, or lower both together.
|
||||
doc_upload_size_limit: "20971520"
|
||||
|
||||
external_secrets:
|
||||
# Shared signing keys (used by app + live). REQUIRED keys: SECRET_KEY, LIVE_SERVER_SECRET_KEY.
|
||||
app_keys_existingSecret: plane-app-keys
|
||||
# REQUIRED keys: REDIS_URL, DATABASE_URL, AMQP_URL (chart-derived URLs; update on DB/queue rotation).
|
||||
app_env_existingSecret: plane-app-env
|
||||
# REQUIRED keys: REDIS_URL.
|
||||
live_env_existingSecret: plane-live-env
|
||||
# REQUIRED keys: POSTGRES_PASSWORD, POSTGRES_DB, POSTGRES_USER.
|
||||
pgdb_existingSecret: plane-pgdb-credentials
|
||||
# REQUIRED keys: RABBITMQ_DEFAULT_USER, RABBITMQ_DEFAULT_PASS.
|
||||
rabbitmq_existingSecret: plane-rabbitmq-credentials
|
||||
# REQUIRED keys: FILE_SIZE_LIMIT, AWS_S3_BUCKET_NAME, USE_MINIO, MINIO_ROOT_USER,
|
||||
# MINIO_ROOT_PASSWORD, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_S3_ENDPOINT_URL.
|
||||
doc_store_existingSecret: plane-minio-credentials
|
||||
# ssl_token_existingSecret: '' # DNS-01 only (cloudflare/digitalocean); unused with HTTP-01
|
||||
|
||||
# Optional, P2: HTTP -> HTTPS 301. The chart's own IngressRoute binds only
|
||||
# 'websecure' (http:// currently 404s). extraObjects is rendered verbatim (toYaml).
|
||||
# Uncomment and `helm upgrade` once reviewed:
|
||||
# extraObjects:
|
||||
# - apiVersion: traefik.io/v1alpha1
|
||||
# kind: Middleware
|
||||
# metadata:
|
||||
# name: plane-https-redirect
|
||||
# namespace: plane
|
||||
# spec:
|
||||
# redirectScheme:
|
||||
# scheme: https
|
||||
# permanent: true
|
||||
# - apiVersion: traefik.io/v1alpha1
|
||||
# kind: IngressRoute
|
||||
# metadata:
|
||||
# name: plane-http-to-https
|
||||
# namespace: plane
|
||||
# spec:
|
||||
# entryPoints: [web]
|
||||
# routes:
|
||||
# - match: Host(`plane.chans.xyz`)
|
||||
# kind: Rule
|
||||
# middlewares:
|
||||
# - name: plane-https-redirect
|
||||
# services:
|
||||
# - name: plane-app-web
|
||||
# port: 3000
|
||||
@@ -1,158 +0,0 @@
|
||||
# 希力威视 SR-S25G3218F 调查(2026-08-08)
|
||||
|
||||
**结论:** 若需求是大量 2.5G 终端、少量 10G 光上联,`SR-S25G3218F` 的端口密度
|
||||
更合适;厂商已公开该型号的固件页,但仍缺少完整规格书、管理手册与兼容矩阵。若需求是 8 条全部可协商
|
||||
1/2.5/5/10G 的铜缆链路,且希望有可查的 L3 能力和固件入口,兮克
|
||||
`SKS8300-8T` 是资料更完整、风险更低的选择;它的代价是主动风扇、外置 12 V 电源、
|
||||
无 SFP+ 光口,且仍不应把消费级/SMB 设备当作安全边界或唯一核心。两者都应在
|
||||
到货可退换期内完成实机验收。
|
||||
|
||||
本页为采购前资料调查,不代表已接入本地网络;检索日期为 2026-08-08。
|
||||
|
||||
## 已能核实的事项
|
||||
|
||||
| 项目 | 结论与证据强度 |
|
||||
|---|---|
|
||||
| 型号/端口 | 京东的希力威视商品标题称该 SKU 为 `SR-S25G3218F`,有 16 个 2.5G 电口和 2 个万兆光口,并宣传 VLAN、端口隔离与 LACP。该店铺被厂商官网列为可购买的「京东旗舰店」,因此可作为销售规格,非技术手册。[京东商品页](https://item.jd.com/100165071727.html);[厂商购买渠道说明](https://en.sirivision.com/contactus/) |
|
||||
| 厂商身份 | 厂商官网为 Shenzhen/Guangdong Sirivision Communication;英文官网说明其自 2016 年起提供接入、汇聚和核心交换机方案。[厂商首页](https://en.sirivision.com/) |
|
||||
| 公开的二手厂家资料 | 同一制造商名义的 Alibaba 出口页将精确型号写成 `16*2.5G+2*10G`、`120Gbps`,并列出 QoS、VLAN、SNMP、L3 与 stackable。这是制造商发布在平台上的销售资料,**不是**官网数据表;其中后五项不能据此视为已验收的功能承诺。[制造商平台页](https://www.alibaba.com/pla/SR-S25G3218F-QoS-Managed-SFP-Switch-1625G210G_1601494946214.html) |
|
||||
| 固件入口 | 厂商已发布此精确型号的[固件页](https://www.sirivision.com/sr-s25g3218f%E5%9B%BA%E4%BB%B6/)。公开变更记录提到“光口自适应”和“增加 DAC 配置”;这证明厂商维护过该路径,**不**代表任意 SFP+/DAC/铜模块均兼容。 |
|
||||
| 本机可计算的带宽 | 端口线速相加为单向 60 Gb/s(16 × 2.5 + 2 × 10);若厂商所谓 `120Gbps` 是全双工交换容量,则数学上吻合。它**不**证明缓冲、PPS、表项规模或实际无阻塞性能。 |
|
||||
|
||||
## 网管/L2/L3 能力边界
|
||||
|
||||
京东标题足以支持把 VLAN、端口隔离、LACP 作为「卖家声称提供」的功能;不得由此推导出
|
||||
ACL、IPv4/IPv6 静态路由、SVI 数量、DHCP relay、OSPF/RIP、VRRP、IGMP、ERPS、
|
||||
802.1X、RADIUS/TACACS+、SSH/HTTPS 管理、SNMP 版本、日志/审计、配置备份或固件
|
||||
安全维护一定存在。
|
||||
|
||||
尤其要注意:厂商官网把真正列出的 2.5G L3 产品标为
|
||||
`SR-S25G3412F (8 × 2.5G + 4 × 10G SFP+)`;其 2.5G 类目只显示 7 个型号,
|
||||
不含 `SR-S25G3218F`。官网也把 L2+、Web Smart、L3 分成不同产品类别。这个目录
|
||||
差异**不是**证明 3218F 没有 L3,而是说明「三层」无法通过官网的精确型号文档确认。
|
||||
[2.5G 产品目录](https://en.sirivision.com/product-category/products/2-5g-switches/);
|
||||
[官网的 10G L3 目录](https://en.sirivision.com/product-category/products/10g-switches/10g-layer3-managed-switches/);
|
||||
[官网的 L2+ 分类示例](https://en.sirivision.com/product-category/products/gigabit-switches/gigabit-layer2-managed-switches/)。
|
||||
|
||||
采购前请向京东/厂商索取**与机身 SKU、硬件 revision 和固件版本对应**的 PDF
|
||||
数据表、管理手册和 release notes,并要求书面回答至少以下问题:
|
||||
|
||||
1. L3 是只有 VLAN Interface/IPv4 静态路由,还是另有 IPv6、ACL、动态路由、DHCP relay
|
||||
等;每项的最大 VLAN、MAC、ARP、路由、ACL、LAG 数量分别是多少?
|
||||
2. LACP 是否符合 802.3ad、一个 LAG 最多多少成员、能否跨两台设备(若销售页的
|
||||
`stackable` 属实,堆叠的线缆/模块、最大成员、控制面和软件版本为何)?
|
||||
3. 管理面是否支持 HTTPS/SSH、禁用 HTTP/Telnet、独立管理 VLAN、SNMPv3、syslog、NTP、
|
||||
配置导出/回滚和已签名或可校验的固件;默认凭据首次登录是否强制修改?
|
||||
|
||||
## 供电、散热和光口:当前不能确认
|
||||
|
||||
针对该精确 SKU,厂商官网目录与公开搜索未找到说明书/数据表,所以以下均为**待确认,
|
||||
不能猜测**:
|
||||
|
||||
- 是否为内置 AC 电源、额定输入范围/最大功耗、是否带电源开关和接地端子;是否完全
|
||||
不提供 PoE(本型号名和京东标题均未写 PoE,但这不足以替代规格书)。
|
||||
- 风扇数量、常态/满载噪声、风向、环境温湿度、机架深度与安装耳;不要将「金属壳」
|
||||
或产品照片等同于无风扇/静音。
|
||||
- 两个槽是否均为 **10G SFP+**,是否可协商 1G SFP;支持的 SR/LR/BiDi 波长距离、
|
||||
DAC/AOC 长度、第三方模块/EERPOM 兼容策略、10GBASE-T SFP+ 模块的功耗/温度限制,
|
||||
以及是否支持 GPON/XPON ONU「猫棒」。
|
||||
|
||||
厂商确实单列「SFP Optical Modules」产品分类,但这不构成 3218F 的兼容清单。
|
||||
[厂商产品导航](https://en.sirivision.com/)。购买光模块/直连线时,应要求厂商按这台
|
||||
设备的硬件/固件 revision 出具兼容型号清单;没有书面清单时,先在可退换期实测两端的
|
||||
链路、重启恢复、热插拔与长时间满载错误计数。
|
||||
|
||||
## 风险与建议验收
|
||||
|
||||
- **文档/生命周期风险(中到高):** 精确型号不在厂商当前官网 2.5G 目录,虽有固件下载页,
|
||||
但未公开完整型号手册、明确 release notes 或兼容矩阵。官网的售后条款也要求按具体产品查询保修期,配件(含光纤头)
|
||||
的保修条款与主机不同;不要把平台页的「3 年」当作中国零售 SKU 的已确认保修。
|
||||
[厂商售后条款](https://en.sirivision.com/after-sale-protection/)
|
||||
- **功能表述风险(高):** 页面将 L2 特性和「三层网管」并列;在命令/网页菜单、
|
||||
手册和测试证明之前,将其当作 L2 VLAN/LACP 设备部署,跨 VLAN 路由仍由现有网关承担。
|
||||
- **双 10G 上联约束(中):** 两个 SFP+ 可作双上联或一个二成员 LAG,但 LAG 增加的是
|
||||
多流量总吞吐,单一 TCP/UDP 流通常仍受一条 10G 链路限制;上级设备也必须匹配 LACP
|
||||
配置。
|
||||
- **管理面风险(中到高):** 家用/低价网管设备常见明文管理、弱默认口令或不透明的固件
|
||||
更新周期;采购后先置于受限管理 VLAN,改口令、升级已验证固件,且不将管理界面暴露
|
||||
到 WAN/访客网。
|
||||
|
||||
最低验收应包括:逐口协商 100M/1G/2.5G、两只不同厂家 SFP+/DAC(仅在卖家承诺支持的
|
||||
范围内)、VLAN trunk/access/PVID、STP/环路保护、LACP 故障切换、端口隔离、满载
|
||||
双向 iperf3 与错误计数、冷启动后的配置保留,以及管理面的 HTTPS/SSH/SNMPv3/配置备份。
|
||||
如无法提供与型号匹配的正式资料或其中任一关键项失败,应在退换期内退货,并选择公开
|
||||
数据表、固件与兼容矩阵更完整的型号。
|
||||
|
||||
## 备选:兮克 SKS8300-8T 对比
|
||||
|
||||
### 已核实的厂商规格
|
||||
|
||||
兮克官网的精确型号页明确将 `SKS8300-8T` 定位为三层管理型 10G 全电口交换机,并列出:
|
||||
|
||||
- 8 × 1/2.5/5/10GBASE-T RJ45;160 Gb/s 交换容量、119.05 Mpps、12 Mbit 缓存、
|
||||
16K MAC、12 KB 巨帧、512 MB DRAM、32 MB Flash,尺寸 207 × 136 × 35 mm;
|
||||
- QoS、ACL、IP+MAC+端口绑定、流分类/优先级标记、多端口镜像、静态/灵活 QinQ、
|
||||
sFlow,以及「基于策略的 IPv4/IPv6 单播路由」。
|
||||
|
||||
这些是厂商能力声明,并非对每一种路由协议或表项上限的承诺;但相对 3218F 的仅有
|
||||
销售标题,它给出了精确型号、转发性能和 L3 范围。[兮克 SKS8300-8T
|
||||
产品页](https://seekswan.com/user/custom-pages/SKS8300-8T.html)
|
||||
|
||||
独立的 OpenWrt 设备资料将其识别为 Realtek RTL9303、512 MB RAM,记录了原厂固件
|
||||
下载入口和串口/TFTP 恢复路径;其硬件数据页列为 12 V / 4 A。这支持「可恢复、可替换
|
||||
系统」的可操作性,但**不是**兮克对原厂功能的支持承诺。
|
||||
[OpenWrt 设备页](https://openwrt.org/toh/xikestor/sks8300-8t);
|
||||
[OpenWrt 硬件数据](https://openwrt.org/toh/hwdata/xikestor/xikestor_sks8300-8t)。
|
||||
|
||||
### 能力、物理与运维比较
|
||||
|
||||
| 维度 | 希力威视 SR-S25G3218F | 兮克 SKS8300-8T |
|
||||
|---|---|---|
|
||||
| 接口/典型用途 | 16 × 2.5G 电口 + 2 × 10G SFP+(销售规格);适合很多 2.5G 终端/NAS,以 10G 光或 DAC 上联。 | 8 × 1/2.5/5/10GBASE-T;适合 10G 铜缆设备、2.5/5G 多速率 NAS/主机。没有 SFP+,光纤上联必须经媒体转换或选另一型号。 |
|
||||
| 可确认的三层范围 | 仅销售/平台资料称 L3;没有精确型号官方手册,不能确认静态路由以外的功能。 | 官网明确写策略型 IPv4/IPv6 单播路由、ACL/QoS/sFlow/QinQ;动态路由、VRRP、IPv6 ACL/SNMP/认证等仍须按当前固件手册确认。 |
|
||||
| 冗余/二层 | 卖家声称 VLAN、端口隔离、LACP;STP/环网的实现与规格未知。 | 官网声明 L3 和多项转发特性,但未在产品页给出 STP/LACP/ERPS 的精确限制;购买前仍索取手册。 |
|
||||
| 散热/噪声 | 无可核实的精确型号风扇、噪声、功耗或风向数据。 | 独立手册镜像和产品图均称智能温控风扇,但厂商产品页未给 dBA;应按「有风扇、可能听得见」规划,不能承诺静音。 |
|
||||
| 供电 | 未找到精确型号官方输入/功耗资料。 | OpenWrt 硬件数据记录 12 V / 4 A;确认随附电源适配器的插头、余量和地区认证。官方产品页未给满载功耗。 |
|
||||
| 固件/恢复 | 有精确型号官方固件页;公开记录包含光口自适应与 DAC 配置改动,但未找到完整 release notes、恢复步骤或兼容矩阵。 | 厂商产品页提供「相关下载」区,OpenWrt 还记录原厂固件入口、RJ45 串口和 U-Boot/TFTP 恢复;原厂镜像是否签名、漏洞修复 SLA、配置回退仍未知。 |
|
||||
|
||||
关于 8T 的风扇、满载功耗(常见转述为 ≤36 W)、温度范围、芯片型号等,本次未找到
|
||||
相应的**厂商原始数据表**;不将第三方手册转录当作已核实规格。若噪声、UPS 容量或
|
||||
机柜散热是购买约束,请先让卖家提供产品铭牌照片、适配器铭牌照片、额定/实测功耗和
|
||||
dBA 测试条件。
|
||||
|
||||
### 选择与验收建议
|
||||
|
||||
- 选 **3218F**:必须有 ≥12 个 2.5G 接入端、10G 光/DAC 上联、且 L3 留给现有路由器。
|
||||
下单前先取得精确型号手册和 SFP+/DAC 兼容承诺;否则端口数量优势不足以抵消资料风险。
|
||||
- 选 **8T**:最多 8 个设备但需要多速率 10G RJ45、明确的 IPv4/IPv6 静态/策略路由和
|
||||
以后自行维护/恢复的余地。不要把其 160 Gb/s 标称交换容量误解为 8 端口同时 10G
|
||||
全双工的性能保证——该标称与端口总线速数学相等,但仍须以实测和厂商 PPS/缓冲说明为准。
|
||||
- 两台都不应单独承担防火墙、访客/IoT 安全隔离或 WAN 暴露;VLAN 的跨网段策略和公网
|
||||
边界留在受支持的网关/防火墙上。先为管理面创建专用 VLAN,仅从管理主机访问,禁用
|
||||
未使用的远程管理协议,备份配置和原厂固件后再接入生产网络。
|
||||
|
||||
## 低功耗核心备选(8 × 2.5G + 2 × SFP+)
|
||||
|
||||
如果核心只需接最多 8 台铜缆终端、上联/连接 NAS 使用 DAC 或光纤 10G,优先考虑没有
|
||||
PoE 的以下两款。它们都满足 VLAN trunk、LACP 和至少两个 10G SFP+ 的需求;不要为
|
||||
AP 选 PoE 版来承担核心,因为 PoE 预算、风扇和待机损耗都会明显增加。
|
||||
|
||||
| 型号 | 端口与管理能力(厂商声明) | 厂商功耗 / 噪声资料 | 对当前 LAN 的判断 |
|
||||
|---|---|---|---|
|
||||
| **TP-Link Omada SG3210X-M2** | 8 × 100M/1G/2.5G RJ45、2 × 10G SFP+,并有 RJ45 和 Micro-USB console。厂商规格列出 802.1Q VLAN、STP/RSTP/MSTP、静态 LAG 和 802.3ad LACP(最多 8 个聚合组、每组最多 8 端口);L3 是 32 个 IPv4/IPv6 接口、48 条静态路由。 | **无风扇**;100–240 V AC 内置电源。`UN 1.20` 数据表:待机最高 **6.0 W**(220 V/50 Hz、25 °C),最高 **15.3 W**(220 V)或 **15.0 W**(110 V)。 | **首选低功耗方案。** 足以做 LAN66 核心、给 PVE/gfw 与 U6 Lite 做 VLAN 10 trunk,并以 SFP+ DAC/光口连接 10G NAS/主机;它不提供 5G/10G RJ45,10G 铜缆需外置转换或 SFP+ 10GBASE-T 模块。 |
|
||||
| **MikroTik CRS310-8G+2S+IN** | 8 × 2.5G RJ45、2 × 10G SFP+;SFP+ 笼支持 1G/2.5G/10G。RouterOS v7(也可选 SwOS)支持 VLAN、链路聚合与 ACL。 | 18–57 V DC 外置供电;官方给出“无附件”最高 **21 W**、总体最高 **34 W**,且机内 **1 个风扇**。厂商没有在该页给出 dBA。 | 可用且软件/文档/恢复路径成熟,但不是本题的静音低功耗优先项:官方最大功耗显著高于 TP-Link,且有风扇。适合明确偏好 RouterOS/SwOS 与其可维护性时选。 |
|
||||
|
||||
功耗数字是各厂商的**上限/待机测试条件**,不是你实际墙插读数;SFP+ 光模块、DAC/AOC,尤其
|
||||
10GBASE-T SFP+ 模块,会另增功耗和热量。对于本网络,用被动 DAC 或短距光模块连接 10G
|
||||
设备,通常比全 RJ45 10G 核心更容易保持低温、低噪。
|
||||
|
||||
`SG3210X-M2` 的上表数据对应 TP-Link 的 `UN 1.20` 数据表;不同地区/硬件版本的包装、
|
||||
认证和功耗标注可能不同,购买中国零售版本前应让卖家确认**准确硬件版本、保修渠道和固件地区**。
|
||||
本次未找到 TP-Link 中国官网的该精确型号页,因此不能把海外官方页面当作大陆现货/售后承诺。
|
||||
MikroTik 同样应通过其官方零售商查询渠道确认本地库存和保修。两台购买前还应确认所选
|
||||
SFP+/DAC 的兼容清单。
|
||||
|
||||
来源:[TP-Link 产品规格](https://www.tp-link.com/uk/business-networking/omada-switch-access-pro/sg3210x-m2/);
|
||||
[TP-Link `UN 1.20` 数据表](https://static.tp-link.com/upload/product-overview/2025/202512/20251224/SG3210X-M2%28UN%29%201.20_datasheet.pdf);
|
||||
[MikroTik 产品页](https://mikrotik.com/product/crs310_8g_2s_in);
|
||||
[MikroTik 用户手册](https://help.mikrotik.com/docs/spaces/UM/pages/214630429/CRS310-8G%2B2S%2BIN)。
|
||||
+64
-3
@@ -68,6 +68,66 @@ db.device.find(
|
||||
).pretty()
|
||||
```
|
||||
|
||||
## IPv6 status (verified 2026-08-20)
|
||||
|
||||
IPv6 is **enabled and live** on the main Wi-Fi networks. Read-only
|
||||
verification, no changes made.
|
||||
|
||||
**Controller (`networkconf` in the `ace` DB):** the `Default` LAN network has
|
||||
`ipv6_enabled: true`, `ipv6_client_address_assignment: slaac`,
|
||||
`ipv6_ra_enabled: true`, `ipv6_ra_priority: high`, and
|
||||
`dhcpdv6_allow_slaac: true`. `ipv6_interface_type: "none"` is expected: the
|
||||
network's gateway is the third-party EdgeRouter (`gw`), so the controller does
|
||||
not manage WAN-side IPv6 — RA/SLAAC is served by the router.
|
||||
|
||||
All active SSIDs map to the `Default` network: `ubnt-windy` (5G),
|
||||
`ubnt-windy-2` (2.4G), `ubnt-haas` (2.4G) — clients on them receive SLAAC IPv6.
|
||||
|
||||
Exception: the dormant `ubnt-upg` VLAN 10 network (and its `ubnt-upg` SSID) has
|
||||
no IPv6 configuration (default off). See
|
||||
[Dedicated Wi-Fi through a third-party gateway](#dedicated-wi-fi-through-a-third-party-gateway).
|
||||
|
||||
**APs (live):** both managed APs hold global SLAAC addresses on `br0` with a
|
||||
default route learned via RA from `gw`:
|
||||
|
||||
| AP | Global IPv6 on `br0` (at check time) | Default route |
|
||||
|---|---|---|
|
||||
| U6 Lite (`192.168.66.6`) | `240e:3bd:235:1fb1:...`/64 | `default via fe80::... dev br0 proto ra` |
|
||||
| UAP-AC-Lite (`192.168.55.5`) | `240e:3bd:235:1fb2:...`/64 | `default via fe80::... dev br0 proto ra` |
|
||||
|
||||
The delegated prefixes are dynamic ISP allocations (PPPoE PD `/60`) and rotate
|
||||
on redial; only the structure is stable.
|
||||
|
||||
**Gateway (`gw`):** the IPv6 routing table shows connected `/64`s on `eth0`
|
||||
(LAN66) and `switch0` (LAN55) plus `::/0` via `pppoe0`.
|
||||
|
||||
Re-verify:
|
||||
|
||||
```bash
|
||||
ssh -4 -o BatchMode=yes zhiqiangf@192.168.66.6 'ip -6 addr show br0; ip -6 route show'
|
||||
ssh -4 -o BatchMode=yes zhiqiangf@192.168.55.5 'ip -6 addr show br0; ip -6 route show'
|
||||
```
|
||||
|
||||
> **2026-08-21 (W1N-207):** SmartThings Element/vWire provisioning SSIDs
|
||||
> (`element-8a0d5133c9438f12`, `vwire-8b2d67469e455785`, `vport-F09FC22004E9`) were
|
||||
> removed (element_adopt setting disabled + element wlanconf deleted + device vwire
|
||||
> fields cleared) and confirmed off on both APs (normal SSIDs unchanged: `ubnt-windy`,
|
||||
> `ubnt-windy-2`, `ubnt-haas`, `ubnt-upg`). Root cause of Matter onboarding failure that
|
||||
> day: HA's matter-server advertised a stale IPv6 GUA (two prefix generations old) in
|
||||
> mDNS; fixed by restarting the add-on — see [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md).
|
||||
>
|
||||
> **SSID ↔ subnet split (Matter-relevant):** `ubnt-windy` (5G) is served only by the
|
||||
> U6 Lite on LAN66; `ubnt-haas` / `ubnt-windy-2` (2.4G) only by the UAP-AC-Lite on
|
||||
> LAN55. mDNS is link-local multicast and does **not** cross the routed 55/66
|
||||
> boundary (no mDNS reflector). Matter commissioning therefore requires phone and
|
||||
> device on the **same subnet (LAN55)**; a phone on 5G (LAN66) cannot discover a
|
||||
> LAN55 Matter device.
|
||||
>
|
||||
> **Cleanup side-effects (left as-is, harmless):** after the direct-DB cleanup,
|
||||
> `db.device.cfgversion` holds placeholder values (`0000000000000000` /
|
||||
> `1111111111111111`) and UAP-AC-Lite has `mesh_sta_vap_enabled=false`; the
|
||||
> controller has not reverted them and no functional impact was observed.
|
||||
|
||||
## Dedicated Wi-Fi through a third-party gateway
|
||||
|
||||
### Architecture boundary discovered on 2026-08-08
|
||||
@@ -151,9 +211,10 @@ Use `ssh zhiqiangf@AP_IP` for the adopted-device account. Do not query or copy
|
||||
the controller's `mgmt` database setting into logs or documentation: it can
|
||||
contain the managed SSH password.
|
||||
|
||||
On 2026-08-06, key-only IPv4 SSH was verified for both managed APs using the
|
||||
`zhiqiangf` account. Verify future access without permitting password or
|
||||
keyboard-interactive fallback:
|
||||
Key-only IPv4 SSH was verified for both managed APs using the `zhiqiangf`
|
||||
account on 2026-08-06 and re-verified 2026-08-20 (BatchMode with password and
|
||||
keyboard-interactive disabled; both APs still log in key-only). Verify future
|
||||
access without permitting password or keyboard-interactive fallback:
|
||||
|
||||
```bash
|
||||
ssh -4 -o BatchMode=yes -o PasswordAuthentication=no \
|
||||
|
||||
@@ -48,9 +48,11 @@ space before increasing retention. DNSSEC is disabled because the selected
|
||||
upstream path did not pass the known-bad-signature validation check; do not
|
||||
enable it without re-testing validated upstreams.
|
||||
|
||||
The compatible names `hass.windy.lan` and legacy `hass.local` currently point
|
||||
to the same Home Assistant address. Migrate clients to `hass.windy.lan`; keep
|
||||
the legacy rewrite until its planned retirement.
|
||||
`hass.windy.lan` points to Home Assistant via this rewrite. The legacy
|
||||
`hass.local` rewrite was removed on 2026-08-14; `hass.local` now resolves only
|
||||
via HAOS mDNS/LLMNR (`hostname: hass`), not via AdGuard Home. The
|
||||
`nas.windy.local` rewrite was likewise removed on 2026-08-14; `.local` names
|
||||
are now left to mDNS only. Remaining rewrites all use `.windy.lan`.
|
||||
|
||||
## Mihomo and routing boundary
|
||||
|
||||
|
||||
@@ -30,6 +30,44 @@ OpenClash runs `/etc/openclash/clash` (clash_meta core) with configuration
|
||||
(`/etc/mosdns/config.yaml`): domestic domains → AGH `.36:53`, foreign →
|
||||
`223.5.5.5`/`119.29.29.29` (Chinese public DNS). mosdns is **not** in the
|
||||
client query path — LAN/VLAN10 clients receive fake-ip from clash :7874.
|
||||
> 2026-08-12: fixed missing `has_resp → accept` guard after the domestic
|
||||
> branch in `/etc/mosdns/config.yaml` (domestic queries were double-forwarded,
|
||||
> final answer came from CN public DNS, bypassing AGH blocking/rewrites;
|
||||
> verified via `dup.baidustatic.com` before/after); added `domestic_fallback`
|
||||
> (fallback plugin: primary=AGH, secondary=CN public DNS, 500ms) so domestic
|
||||
> DIRECT lookups survive an AGH outage. Backups:
|
||||
> `config.yaml.bak-20260812` / `config.yaml.bak-fallback-20260812`. See
|
||||
> [docs/lan-dns-architecture.md](../docs/lan-dns-architecture.md) §1.
|
||||
> 2026-08-13 (W1N-62): foreign branch now uses encrypted DoH
|
||||
> `https://adg.chans.xyz/dns-query` (self-hosted, hk2) via new
|
||||
> `foreign_upstream` / `foreign_fallback` plugins; non-CN queries → DoH,
|
||||
> falls back to CN public DNS after 1000ms. `bootstrap` = existing CN public
|
||||
> DNS IPs (no self-loop). Live-verified: google/youtube real IP + AAAA
|
||||
> restored (2607:f8b0…), `dup.baidustatic.com` → `0.0.0.0` (AGH intercept
|
||||
> kept), clash 7874 fake-ip plane unchanged. **Final decision (2026-08-13):
|
||||
> DoH goes DIRECT to hk2, not via clash proxy** — `foreign_upstream` points
|
||||
> only at the self-hosted resolver `adg.chans.xyz` (hk2), which is directly
|
||||
> reachable and already encrypted (DoH/TLS) with clean answers, so forcing
|
||||
> the proxy adds nothing and would couple the DNS plane to clash (nft output
|
||||
> chains also show OpenClash does not currently redirect router-own TCP).
|
||||
> Kill-test: foreign queries answered during clash outage, watchdog
|
||||
> auto-restarted. Backups: `config.yaml.bak-foreign-doh-20260813-103746` /
|
||||
> `config.yaml.bak-foreign-doh-20260813-103813`.
|
||||
> 2026-08-13 (W1N-62): added redundancy to `foreign_upstream` —
|
||||
> `concurrent: 3`, upstreams = `adg.chans.xyz` (hk2) + `dns.quad9.net` +
|
||||
> `dns.cloudflare.com` (both direct-reachable from CN, live-tested 2026-08-13;
|
||||
> `dns.quad101.net` excluded — TLS handshake fails). Verified: google.com
|
||||
> AAAA now `2404:6800…` (new upstream answering, was `2607:f8b0…` via hk2),
|
||||
> taobao/intercept/clash-fake-ip all unchanged. Backup:
|
||||
> `config.yaml.bak-multi-doh-20260813-105421`.
|
||||
> 2026-09-01: removed `dns.quad9.net` from `foreign_upstream` — recurring
|
||||
> `WARN foreign_upstream … unexpected EOF` bursts (481 log entries) against
|
||||
> Quad9 DoH; endpoint answers on probe but gets intermittently
|
||||
> connection-reset from this network (same failure class as the excluded
|
||||
> `dns.quad101.net`). Remaining upstreams `adg.chans.xyz` (hk2) +
|
||||
> `dns.cloudflare.com` both verified live; google.com A + youtube.com AAAA
|
||||
> resolve through mosdns :6052 after restart. Backup:
|
||||
> `config.yaml.bak-quad9-remove-20260901-201801`.
|
||||
- nft: OpenClash injects TPROXY/redirect + DNS-hijack rules into
|
||||
`table inet fw4`; a residual `table inet passwall` exists with 0 packets (unused)
|
||||
|
||||
|
||||
+61
@@ -34,6 +34,18 @@ new SSH host key out of band before accepting it.
|
||||
IPv6 prefix delegation assigns SLAAC-capable `/64` networks to both LANs.
|
||||
`eth4` applies the WAN IPv4 and IPv6 firewall policies.
|
||||
|
||||
**SE5420 single-uplink topology (verified 2026-08-22):** the TP-Link `TL-SE5420`
|
||||
core switch is deployed — management `192.168.66.253` (TP-Link OUI `f8:c9:03`,
|
||||
web UI on :80/:443). The LAN55 uplink into `switch0` is a **single member
|
||||
port**: `eth1` link up, `eth2`/`eth3` down. All LAN55 wired devices (hass
|
||||
`.11`, Aqara M3 `.248`, SmartThings `.48`, UAP-AC-Lite `.5`) are reached via
|
||||
`switch0` behind that one uplink, so same-segment wired↔wired unicast is
|
||||
switched locally on the SE5420 and never reaches the ER-X. The switch FDB is
|
||||
hardware-offloaded and not readable from the ER-X (`brctl showmacs switch0` →
|
||||
"Operation not supported"; `show mac-address-table` / `show ethernet-switch`
|
||||
are not available on this EdgeOS build) — port link state (`show interfaces
|
||||
ethernet`) plus ARP are the reliable topology checks.
|
||||
|
||||
Detailed effective configuration, including firewall binding and WAN exposure,
|
||||
is recorded in [the EdgeRouter X configuration record](../docs/edgerouter-x-configuration.md).
|
||||
|
||||
@@ -77,6 +89,35 @@ relevant interface/direction to take effect. Use the operational `show
|
||||
firewall` output—not merely the configured rule definitions—to determine the
|
||||
effective policy.
|
||||
|
||||
## PPPoE redial
|
||||
|
||||
To force the `pppoe0` session to reconnect (e.g. to obtain a fresh WAN IP), use
|
||||
the operational `disconnect` / `connect` commands — **not** `renew dhcp
|
||||
interface`, which applies only to DHCP interfaces:
|
||||
|
||||
```bash
|
||||
ssh -4 zhiqiang@192.168.66.254
|
||||
/opt/vyatta/bin/vyatta-op-cmd-wrapper disconnect interface pppoe0
|
||||
/opt/vyatta/bin/vyatta-op-cmd-wrapper connect interface pppoe0
|
||||
```
|
||||
|
||||
`disconnect` tears down the PPP session; `connect` re-dials immediately. A
|
||||
short pause between them (a few seconds, or minutes for cautious ISPs) lets the
|
||||
old session finish teardown before redialing. This briefly drops the whole WAN
|
||||
uplink and may change the public IPv4 and delegated IPv6 `/60`; in-flight
|
||||
sessions and port-forwarded services are interrupted until the new session is
|
||||
up.
|
||||
|
||||
The `zhiqiang` account logs into `vbash`, not the EdgeOS CLI, so operational
|
||||
commands must be invoked through `/opt/vyatta/bin/vyatta-op-cmd-wrapper` and
|
||||
depend on its passwordless `sudo`. The `ubnt` account lands directly in the
|
||||
operational CLI, where the same commands are entered without the wrapper.
|
||||
`show`/`configure` are interactive-only aliases (from
|
||||
`/etc/bash_completion.d/vyatta-{op,cfg}`, loaded via `~/.bashrc`), so a
|
||||
non-interactive `ssh ubnt@… 'show …'` also fails — from a script use the op
|
||||
wrapper above, or `_vyatta_op_run` after sourcing `vyatta-op` with
|
||||
`vyatta_op_templates=/opt/vyatta/share/vyatta-op/templates`.
|
||||
|
||||
## Maintenance notes
|
||||
|
||||
- EdgeOS writes persistent changes through its configuration tree: enter
|
||||
@@ -100,3 +141,23 @@ from `192.168.55.254` reached the UniFi controller at `192.168.66.46` with
|
||||
3/3 ICMP replies. This supports the AP Inform path to
|
||||
`192.168.66.46:9080`; the controller listener and an online LAN55 AP provide
|
||||
the corresponding application-level evidence. No firewall changes were made.
|
||||
|
||||
IPv6 was re-verified by read-only SSH on 2026-08-20 during the UniFi AP/AC
|
||||
check: the IPv6 routing table shows connected `/64`s on `eth0` (LAN66) and
|
||||
`switch0` (LAN55) plus `::/0` via `pppoe0`; both UniFi APs obtained SLAAC
|
||||
addresses from the router's RAs. No configuration changes were made.
|
||||
|
||||
**DHCP 保留 `matter` 失效(2026-08-21 发现,2026-08-23 复核仍未生效,W1N-207):**
|
||||
静态映射 `matter` → .45 / MAC `34:98:7a:27:10:bc`,但该灯泡一直以**动态租约**拿
|
||||
`.148`(hostname `matter`;2026-08-23 09:02 时租约当日 04:40 已续租)。保留 .45 从未
|
||||
被租出。2026-08-23 复核补充:另一盏工作灯泡的 MAC 已变为 `fc:e8:c0:25:a1:f0`
|
||||
(动态 `.146`,hostname `espressif`),原「把 MAC 改为 `34:98:7a:27:7f:08`」的修正
|
||||
建议已过时(该灯泡已离网)。处置:删除该保留,或按现用 MAC(`.148` 的
|
||||
`34:98:7a:27:10:bc` / `.146` 的 `fc:e8:c0:25:a1:f0`)重建,**未执行**。
|
||||
|
||||
**SE5420 部署 + switch0 单上联(2026-08-22 只读核实):** `switch0` 成员口
|
||||
`eth1` link up、`eth2`/`eth3` down(单上联);SE5420 管理面 `192.168.66.253`
|
||||
在线(TP-Link OUI `f8:c9:03`,:80/:443);ARP 显示 LAN55 主机(hass `.11`、
|
||||
M3 `.248`、SmartThings `.48`、UAP-AC-Lite `.5`)全部经 switch0 可达。含义:
|
||||
`switch0` 不再是 LAN55 的全量抓包点(同段有线单播在 SE5420 本地交换),详见
|
||||
[runbooks/matter-packet-capture.md](../runbooks/matter-packet-capture.md)。
|
||||
|
||||
@@ -0,0 +1,751 @@
|
||||
[hosts/hass.windy.lan.md#8DF6]
|
||||
# hass.windy.lan — Home Assistant (HAOS)
|
||||
|
||||
## Role and access
|
||||
|
||||
| Item | Value |
|
||||
|---|---|
|
||||
| Role | Home Assistant automation hub |
|
||||
| IPv4 | `192.168.55.11` (LAN55) |
|
||||
| DNS | `hass.windy.lan` (AdGuard rewrite on `dns.windy.lan`; legacy `hass.local` alias) |
|
||||
| SSH | `ssh hassio@hass.windy.lan` |
|
||||
| **Host** | **x88 Pro physical box** (HAOS bare-metal, `machine: green`; verified 2026-08-18) |
|
||||
| Platform | Home Assistant OS; kernel `6.1.115-haos` (aarch64) |
|
||||
| Web UI | `http://hass.windy.lan:8123` (LAN); WAN port-forward `hass` on gw → `:8123` |
|
||||
|
||||
Use `hassio` for routine SSH inspection. Key-only login was verified on
|
||||
2026-08-13 from the WSL client (`BatchMode=yes`).
|
||||
|
||||
The `ha` supervisor CLI (`/usr/bin/ha`) authenticates with `SUPERVISOR_TOKEN`.
|
||||
Interactive login works because `~hassio/.zprofile` runs `exec sudo -i`, which
|
||||
loads a root environment carrying the supervisor API token. Non-interactive
|
||||
`ssh hassio 'command'` does not source `.zprofile` and fails with
|
||||
`unauthorized: missing or invalid API token`. Run `ha` non-interactively via:
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core info'
|
||||
```
|
||||
|
||||
Verified 2026-08-13 that `sudo -n -i ha core info` works from the WSL client.
|
||||
Never copy the supervisor token into this repository.
|
||||
|
||||
The current SSH ED25519 host-key fingerprint is
|
||||
`SHA256:DMcMOgDzFsFTon1fndXowEP7jlyOK3/AX3PVK8BATvk` (verified 2026-08-13).
|
||||
Verify a changed key out of band before accepting it.
|
||||
|
||||
Do not store Home Assistant long-lived tokens, integration credentials, or
|
||||
recovery codes in this repository.
|
||||
|
||||
## Network
|
||||
|
||||
| Interface | Address / role |
|
||||
|---|---|
|
||||
| `end0` | IPv4 static `192.168.55.11/24` (gw `.254`, DNS `192.168.66.36`); IPv6 SLAAC `auto` with GUA on the current PD-derived /64 (`240e:3bd:235:1fb2:*` at 2026-08-22; rotates on PPPoE redial); primary LAN55 NIC (interface name verified live 2026-08-22 — `end1` does not exist) |
|
||||
| `wlan0` | Supervisor **disabled** (verified 2026-08-14, W1N-104); IPv6 remains off on this RTL8821CS radio |
|
||||
| `wg0` | `10.13.13.2/32`; WireGuard (add-on / integration tunnel) |
|
||||
| `hassio` / `docker0` | internal HAOS Docker bridges (`172.30.32.0/23`, `172.30.232.0/23`) |
|
||||
|
||||
LAN55 clients reach the HTTP API on `dns.windy.lan:80` for the AdGuard Home
|
||||
integration; see [hosts/dns.windy.lan.md](dns.windy.lan.md).
|
||||
|
||||
## API access
|
||||
|
||||
Home Assistant exposes a REST API at `http://hass.windy.lan:8123/api/` (same
|
||||
as `http://192.168.55.11:8123/api/`). Authenticate with a **long-lived access
|
||||
token** created under **Profile → Security → Long-lived access tokens**.
|
||||
|
||||
```bash
|
||||
HA_URL="http://hass.windy.lan:8123"
|
||||
HA_TOKEN="<long-lived-access-token>"
|
||||
|
||||
# Health check — expect {"message":"API running."} and HTTP:200
|
||||
curl -sS -w "\nHTTP:%{http_code}\n" \
|
||||
-H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/"
|
||||
|
||||
# Read one entity state
|
||||
curl -sS -H "Authorization: Bearer $HA_TOKEN" \
|
||||
"$HA_URL/api/states/sensor.csg_30d_max"
|
||||
|
||||
# List entities / recent errors
|
||||
curl -sS -H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/states"
|
||||
curl -sS -H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/error_log"
|
||||
```
|
||||
|
||||
- `401` → token invalid or expired; create a new one.
|
||||
- `404` on `/api/states/<id>` → entity does not exist.
|
||||
- The token is a secret: never commit it here; keep it in the shell
|
||||
environment or a secrets file outside the repo.
|
||||
|
||||
### HTTP proxy gotcha (verified 2026-08-13)
|
||||
|
||||
The WSL client had `http_proxy` set to Mihomo (`192.168.66.99:7890`). LAN
|
||||
hostnames sent **through that proxy** returned empty `502`, even though DNS
|
||||
resolved and the HA UI was up. Direct `192.168.55.11:8123` worked, and
|
||||
`hass.windy.lan:8123` worked only after clearing the HTTP proxy.
|
||||
|
||||
Before debugging a "502" on a LAN URL, check `env | grep -i proxy` and bypass
|
||||
the proxy:
|
||||
|
||||
```bash
|
||||
unset http_proxy HTTP_PROXY all_proxy ALL_PROXY
|
||||
curl -sS -w "\nHTTP:%{http_code}\n" \
|
||||
-H "Authorization: Bearer $HA_TOKEN" "$HA_URL/api/"
|
||||
```
|
||||
|
||||
For a persistent fix, add `.windy.lan` (leading dot) and the LAN ranges to
|
||||
`NO_PROXY`, or add `*.windy.lan` to the proxy's own bypass/skip-proxy list.
|
||||
See `~/.config/zsh/env/local/environment.env` for the client-side setting.
|
||||
|
||||
## Safe verification
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan 'hostname; ip -4 addr show end0'
|
||||
```
|
||||
|
||||
From a LAN client, confirm DNS and UI reachability:
|
||||
|
||||
```bash
|
||||
getent hosts hass.windy.lan
|
||||
# expect 192.168.55.11
|
||||
```
|
||||
|
||||
## Local patches (custom components)
|
||||
|
||||
|
||||
### Manual custom-component install (this host)
|
||||
|
||||
Home Assistant loads custom integrations from
|
||||
`<config>/custom_components/<domain>/` (HAOS: `/config` ≡ `/homeassistant`).
|
||||
A folder named after the integration domain, containing at least
|
||||
`manifest.json` and `__init__.py`, is enough; Core must be restarted after
|
||||
copying files. Official HA lookup order:
|
||||
`<config>/custom_components/<domain>` then built-in
|
||||
`homeassistant/components/<domain>`.
|
||||
See [Integration file structure](https://developers.home-assistant.io/docs/creating_integration_file_structure).
|
||||
|
||||
This host **does not git-clone** custom components. The live tree is a file
|
||||
copy. Do not `git pull` on HA.
|
||||
|
||||
**Official plugin path** (from
|
||||
[windyboy/china_southern_power_grid_stat README](https://github.com/windyboy/china_southern_power_grid_stat)):
|
||||
HACS **or** [手动下载安装](https://github.com/windyboy/china_southern_power_grid_stat/releases).
|
||||
This host uses the latter. Releases here have no uploaded zip assets; use
|
||||
GitHub's **Source code (zip)** / zipball of the tag.
|
||||
|
||||
**UI (Samba / File editor / Studio Code Server):**
|
||||
|
||||
1. Download Source code (zip) from the GitHub Release.
|
||||
2. Extract. Copy only the inner
|
||||
`custom_components/china_southern_power_grid_stat/` tree — not the repo
|
||||
root, not a nested extra folder.
|
||||
3. Place it at `/config/custom_components/china_southern_power_grid_stat/`.
|
||||
4. Restart Core (**Settings → System → Restart**).
|
||||
5. First install only: **Settings → Devices & services → Add integration**.
|
||||
|
||||
**SSH from the workstation** (verified 2026-08-14, W1N-107). Replace `v1.3.1`
|
||||
with the tag being installed:
|
||||
|
||||
```bash
|
||||
TAG=v1.3.1
|
||||
STAGE=/tmp/csg-${TAG}-deploy
|
||||
mkdir -p "$STAGE"
|
||||
gh api "repos/windyboy/china_southern_power_grid_stat/zipball/${TAG}" \
|
||||
> "$STAGE/src.zip"
|
||||
unzip -q "$STAGE/src.zip" -d "$STAGE"
|
||||
SRC=$(find "$STAGE" -type d -path '*/custom_components/china_southern_power_grid_stat' | head -1)
|
||||
# expect .../custom_components/china_southern_power_grid_stat
|
||||
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i mkdir -p /homeassistant/.csg-backups &&
|
||||
sudo -n -i cp -a /homeassistant/custom_components/china_southern_power_grid_stat \
|
||||
/homeassistant/.csg-backups/china_southern_power_grid_stat.bak-$(date +%Y%m%d)-manual'
|
||||
|
||||
rsync -a --delete \
|
||||
-e 'ssh -o BatchMode=yes' \
|
||||
"$SRC/" \
|
||||
hassio@hass.windy.lan:/homeassistant/custom_components/china_southern_power_grid_stat/
|
||||
|
||||
# --delete cannot remove Core-owned __pycache__; wipe as root, then restart
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i rm -rf /homeassistant/custom_components/china_southern_power_grid_stat/__pycache__ \
|
||||
/homeassistant/custom_components/china_southern_power_grid_stat/*/__pycache__ &&
|
||||
sudo -n -i ha core restart'
|
||||
```
|
||||
|
||||
Wait until Core is up (`ha core info` returns, typically 1–2 min; this CLI
|
||||
build does not print a `state:` field).
|
||||
Then:
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i cat /homeassistant/custom_components/china_southern_power_grid_stat/manifest.json'
|
||||
# version must match the tag
|
||||
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i ha core logs -n 2500' | grep -E 'china_southern_power_grid_stat|cannot pickle' || true
|
||||
```
|
||||
|
||||
**Host constraints (do not skip):**
|
||||
|
||||
- Backups **must** live in `/homeassistant/.csg-backups/`. A `*.bak-*`
|
||||
directory next to the live folder is scanned as the same domain and Core
|
||||
fails with `No module named '...bak-YYYYMMDD-...'`.
|
||||
- Do not install this fork via HACS on this host. HACS still tracks
|
||||
`CubicPill/china_southern_power_grid_stat` `v1.2.0`; a HACS update would
|
||||
overwrite the live copy.
|
||||
- First poll after restart can time out to CSG over IPv4; if this-month
|
||||
sensors stay `unknown` while last-month filled, reload the config entry
|
||||
(UI: integration → Reload, or supervisor
|
||||
`POST /core/api/config/config_entries/entry/<id>/reload`).
|
||||
- `runbooks/scripts/ha-maintenance.sh --restart-core --yes` can print
|
||||
nothing and exit 1 in under a second **without restarting Core**. The
|
||||
wrapper's ssh line discards stderr (`2>/dev/null`); with `pipefail`,
|
||||
an ssh failure yields empty stdout + exit 1 before any remote command
|
||||
runs. Do not treat that as a completed restart. Confirm with elapsed
|
||||
time (~2 min for a real restart) and `ha core info`. Prefer
|
||||
`ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core restart'`.
|
||||
|
||||
Full command family: [runbooks/home-assistant-maintenance.md](../runbooks/home-assistant-maintenance.md).
|
||||
|
||||
### `china_southern_power_grid_stat` live tree
|
||||
|
||||
**v1.3.2** (`934f58c`, verified 2026-08-15, W1N-118): manual zipball of
|
||||
GitHub release
|
||||
[v1.3.2](https://github.com/windyboy/china_southern_power_grid_stat/releases/tag/v1.3.2)
|
||||
copied to `/config/custom_components/china_southern_power_grid_stat`.
|
||||
Earlier trees: v1.3.1/`55a293fc` (W1N-107), v1.3.0/`69f13c90` (W1N-106),
|
||||
`a433e8c` (W1N-105), `de01914` (W1N-103), `eb8b174` (W1N-102). Backups:
|
||||
`/homeassistant/.csg-backups/` (w1n102/104/105/106/107/118).
|
||||
|
||||
v1.3.0 crashed the coordinator on first refresh
|
||||
(`TypeError: cannot pickle 'mappingproxy' object` in
|
||||
`copy.deepcopy(self._config)` under Python 3.14 / HA 2026.8.1). v1.3.1
|
||||
wraps those `deepcopy` calls with `dict(...)`. Post-restart 22:13 CST:
|
||||
entry `loaded`, no pickle traceback. Native this-month sensors filled after
|
||||
reloading entry `01KGCQDSZCF523A9X6SV3BZ1B9` (`ip_family: ipv4`). Native
|
||||
cost/ladder sensors can stay `unknown` because CSG
|
||||
`get_month_daily_cost_detail` returns a marketing-system SQL error; the
|
||||
dashboard uses template ladder/cost entities instead. Do not change
|
||||
`templates/csg_sensors.yaml` or the 电力监控 dashboard for an install.
|
||||
|
||||
**`templates/csg_sensors.yaml` hardened 2026-08-29 (W1N-239):** added
|
||||
`availability` templates to all 12 `csg_*` sensors (numeric sensors can't
|
||||
render `unknown`/`unavailable` in `state`; availability suppresses
|
||||
rendering instead — native CSG down ⇒ derived sensors show `unavailable`,
|
||||
no more fake zeros / "一档" / `0%`). `csg_yesterday_kwh` now falls back to
|
||||
`last_month_by_day`'s last entry when `this_month_by_day` is empty (month
|
||||
start); ladder constants (`t1/t2/p1/p2/p3`) deduped into per-block
|
||||
`variables:` (Block B + Block D); `csg_mom_change` parses `date`
|
||||
defensively. Backup:
|
||||
`/homeassistant/.csg-backups/csg_sensors.yaml.bak-20260829-w1n239`.
|
||||
**Verified:** `ha core check` OK; Core restart required (trigger-based
|
||||
template blocks don't settle on `template.reload` — W1N-114 precedent);
|
||||
post-restart all 12 entities numeric & consistent (302.47 kWh→180.28 元,
|
||||
324.03 kWh→194.06 元, mom_change -3.6%, yesterday 7.66 kWh/2026-08-28),
|
||||
no template errors in Core logs.
|
||||
|
||||
**`csg_sensors.yaml` off-by-one fixed 2026-08-29 (W1N-241):** CSG data
|
||||
lags 1 day (`sum(this_month_by_day)` == `this_month_total_usage`, data
|
||||
stops at yesterday), but templates used `now().day` as "days elapsed" →
|
||||
`csg_predicted_usage` underestimated ~1 daily avg (~3%) and
|
||||
`csg_mom_change` compared this-month 28 days vs last-month 29 days
|
||||
(-3.6% vs true -0.3%). Both now derive the day number from
|
||||
`this_month_by_day[-1].date` (fallback `now().day` when empty). Added
|
||||
`sensor.csg_this_month_daily_avg` (month-to-date avg, 302.47/28=10.8) and
|
||||
`sensor.csg_prediction_progress` (usage/predicted %, 90.3) in Block C
|
||||
(trigger adds `csg_predicted_usage`). Backup:
|
||||
`/homeassistant/.csg-backups/csg_sensors.yaml.bak-20260829-w1n241`.
|
||||
**Verified (8/29):** predicted 324.03→334.81, mom_change -3.6→-0.3,
|
||||
daily_avg 10.8, progress 90.3, predicted_cost 194.06→200.94 (334.81 kWh
|
||||
ladder), ladder cost 180.28 unchanged, `ha core check` OK after restart,
|
||||
no template errors; 14 csg_* entities total.
|
||||
|
||||
**电力监控面板(`lovelace.dashboard_unknown` / view `power-monitor`)
|
||||
updated 2026-08-29 (W1N-240 + W1N-242):** 「本月累计」gauge 对齐夏季阶梯:
|
||||
`max:650`、segments `0/260/600`(绿/橙/红 = 一/二/三档;冬季 11-01 需切
|
||||
`max:450`、`0/200/400` — **seasonal switch point**,见下文)。「📊 统计
|
||||
数据」卡新增本年/去年 4 行(原生传感器,口径标注「电费(账单)」、本年
|
||||
「(至今)」)+ 本月日均/预测进度 2 行(`csg_this_month_daily_avg` /
|
||||
`csg_prediction_progress`,W1N-242);面板共引用 **20** 个实体。改前备份:
|
||||
`/homeassistant/.lovelace-backups/dashboard-unknown-power-monitor-20260829-204845.json`
|
||||
(W1N-240)、`-20260829-210708.json`(W1N-242)
|
||||
(改法:WS `lovelace/config/save`,参数 `url_path: dashboard-unknown` +
|
||||
`config`;勿直改 `.storage/`)。验证:WS 读回 18→20 实体 diff ✓、gauge
|
||||
配置一致 ✓、URL `http://hass.windy.lan:8123/dashboard-unknown/power-monitor`。
|
||||
|
||||
**`csg_sensors.yaml` W1N-242:** `csg_predicted_usage` /
|
||||
`csg_mom_change` / `csg_this_month_daily_avg` 三处取 `days[-1]` 前补
|
||||
`sort(attribute='date')`(与 `csg_yesterday_kwh` 一致,防上游乱序取错
|
||||
数据日)。备份 `csg_sensors.yaml.bak-20260829-w1n242`。验证:Core
|
||||
restart 后回归值不变(334.81 / -0.3 / 10.8 / 90.3 / 200.94 / 180.28)。
|
||||
|
||||
**CSG 面板重构 2026-09-04(VPS-90,先核对计价后展示层改动):** 核对
|
||||
`power-monitor` 计价与 8 月账单一致(198.65 vs 账单 198.64,差 ≤0.01 元,
|
||||
因模板用公众圆整价 0.589/0.639/0.889、账单用 6 位精确价),不改阶梯常量。
|
||||
改动:① `csg_sensors.yaml` Block B 新增
|
||||
`sensor.csg_this_month_avg_price`(本月阶梯电费÷本月用电,`元/kWh`,
|
||||
availability 照 W1N-239 惯例;**csg_* 实体 14→15**);② 面板改名「环比上月」
|
||||
→「环比上月同期」;glance「本月/上月」grid 去重为单卡「上月」(本月用电/电费
|
||||
行归 💰核心数据卡);⚡阶梯电价卡加「本月实际均价」行(当前档位/当前电价/
|
||||
本月实际均价/档位剩余;面板唯一实体引用 20→21);③ `automations.yaml` 加
|
||||
2 条提醒:`automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`(10-25 09:00)
|
||||
与 `automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`(4-25 09:00)经
|
||||
`matrix_e2ee.send_message` 提醒切 gauge。④ 金额单位混排(原生 CNY vs 模板
|
||||
元)**维持**:`config/entity_registry/update` 拒绝自定义文本单位
|
||||
(`extra keys not allowed … Got '元'`),用户确认接受。备份:
|
||||
`.lovelace-backups/dashboard-unknown-power-monitor-20260904-204757-pre-refactor.json`、
|
||||
`.csg-backups/csg_sensors.yaml.bak-20260904-204757-pre-refactor`(及
|
||||
`-205301-pre-avgprice`)、`.automations-backups/automations.yaml.bak-*`。
|
||||
**WS 改法(2026.8,本机实测)**:core/主机 python 无 ws 库、core 容器内经
|
||||
supervisor 代理 WS 被拒(loop prevention),用
|
||||
`docker run --rm --network host -e SUPERVISOR_TOKEN`(supervisor 镜像
|
||||
`aarch64-hassio-supervisor:2026.08.0`)连 `ws://172.30.32.2/core/websocket`;
|
||||
命令名 `lovelace/config`(读)+ `lovelace/config/save`(写),
|
||||
`lovelace/config/get` 已不存在(unknown_command)。验证:新实体
|
||||
0.589 元/kWh、15 个 csg_* 数值齐全、回归值不变(14.09/198.65/331.22/
|
||||
304.99/181.89)、automations on、`ha core check` OK、日志无 template 错误。
|
||||
|
||||
**CSG 长期归档(W1N-243, 2026-08-29):** scribe 库新增 `csg_history`
|
||||
表(逐日 usage/cost/ladder/balance + 逐月累计;2026-07-01 起回填,永久),
|
||||
由 TimescaleDB 每日任务 **1008** `csg_daily_snapshot()`(22:30
|
||||
Asia/Shanghai;**TS job 非 pg_cron**)upsert 维护。日费用在原生
|
||||
`latest_day_cost` 缺失时回退 = 昨日用电 × 当前档费率(模板
|
||||
`csg_current_ladder_tariff` 0.639);月费用回退模板
|
||||
`csg_this_month_ladder_cost`。**语义**:day 行 usage/cost 为该日值,
|
||||
ladder/balance 为 22:30 快照值。详见 [hosts/pgdb.md](../hosts/pgdb.md)。
|
||||
|
||||
> **Seasonal gauge switch (W1N-240 已知事项):** 每年 **11-01** 把
|
||||
> `power-monitor` 视图「本月累计」gauge 切到冬季 `max:450` /
|
||||
> `0/200/400`,**5-01** 切回夏季 `max:650` / `0/260/600`(与模板
|
||||
> `now().month` 季节逻辑对齐;模板常量在 Block B/D `variables`)。
|
||||
> **提醒 automation(2026-09-04 起,VPS-90):**
|
||||
> `automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`(10-25)与
|
||||
> `automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`(4-25)09:00 经
|
||||
> `matrix_e2ee.send_message` 发操作步骤提醒;gauge 的 max/segments 无法
|
||||
> 模板化,仍需人工改卡配置。
|
||||
|
||||
Home PPPoE IPv4 to CSG is still blackholed (`curl -4` to `218.19.148.218:443`
|
||||
times out). `end0` IPv6 is enabled (`ipv6.method: auto`); from HA,
|
||||
`curl -6 https://95598.csg.cn` returns HTTP 200 via `240e:f9:8060::1:16`.
|
||||
|
||||
**`tianqi` weather recorder patch (verified 2026-08-13, W1N-75):**
|
||||
`/config/custom_components/tianqi/weather.py` has a local patch adding
|
||||
`_unrecorded_attributes = frozenset({"hourly_temperature", "hourly_skycon",
|
||||
"hourly_cloudrate", "hourly_precipitation"})` to the `WeatherEntity` class.
|
||||
Without it, weather.guangzhou's state attributes (~19 KB, dominated by the 4
|
||||
hourly_* arrays of up to 48 entries) exceed the recorder 16384-byte limit, so
|
||||
the recorder drops **all** attributes for the entity and logs
|
||||
`Recorder.db_schema: State attributes for weather.guangzhou exceed maximum
|
||||
size of 16384 bytes`. The patch excludes only the 4 arrays from recording
|
||||
(live state unchanged; other attributes still stored; ~6.3 KB payload). Backup
|
||||
at `weather.py.bak-w1n75`. **Re-apply after any `tianqi` component update.**
|
||||
The `_unrecorded_attributes` mechanism exists in Core 2026.8.1
|
||||
(`Entity.__init_subclass__` → `state_info["unrecorded_attributes"]`, consumed
|
||||
by recorder `shared_attrs_bytes_from_event`).
|
||||
|
||||
|
||||
### `matrix_e2ee` live tree (E2E Matrix bot, verified 2026-08-20)
|
||||
|
||||
**v0.3.12** (tag `v0.3.12`; feat — Matrix activity events
|
||||
`matrix_e2ee_message_received` / `matrix_e2ee_verification_done` + push
|
||||
diagnostics; v0.3.9 added Connection health binary sensor, SAS/command
|
||||
allowlist split, URL normalization, single-entry enforcement):
|
||||
source copy from `/home/windy/project/ha-matrix-e2ee` `ea421ed` (tag
|
||||
`v0.3.12`) deployed 2026-08-20 via SSH rsync from workstation (upgraded
|
||||
from v0.3.2, backup `matrix_e2ee.bak-20260820-v0.3.2`).
|
||||
Custom **`matrix_e2ee`** integration — **Config Flow** (UI). See
|
||||
[docs/home-assistant-matrix.md](../docs/home-assistant-matrix.md).
|
||||
**Update runbook:** [runbooks/matrix-e2ee-update.md](../runbooks/matrix-e2ee-update.md).
|
||||
|
||||
Earlier: v0.3.2 (tag `v0.3.2`, W1N-182/#34: wizard waits for inbound SAS
|
||||
emojis) deployed 2026-08-18 from `d35c484` (backup
|
||||
`matrix_e2ee.bak-20260818-v0.3.1`); v0.3.1 (GitHub #33: peer-initiated
|
||||
verification wizard fix) deployed 2026-08-18 from `d22e935` (backup
|
||||
`matrix_e2ee.bak-20260818-v0.3.0`); v0.3.0 (W1N-180/#32: bot-initiated
|
||||
verification wizard; W1N-179/#31 `receive_mac_event` cancel-state fix)
|
||||
deployed 2026-08-18 from `216cc99` (backup
|
||||
`matrix_e2ee.bak-20260818-v0.2.10`).
|
||||
|
||||
- Bot `@hass:chans.xyz` reused (E2EE device `rO1R915ncu`). Config Entry
|
||||
`01M04D7C1M4T2GX5VPG7NVQ7GV` (`source: import`, `state: loaded`). All
|
||||
settings via **Settings → Devices & Services → Matrix E2EE → Configure**.
|
||||
- Config Entry options: `allowed_rooms` `["!gidvAzpDzwtzfEDrqu:chans.xyz", "!boxfylDSzOvrWkcsyY:chans.xyz"]`,
|
||||
`allowed_users` `["@zhiqiang:chans.xyz"]`, `command_prefix` `"!"`.
|
||||
**`verification_peer_users` not set** (v0.3.9+ SAS allowlist split from
|
||||
`allowed_users`, W1N-156): defaults to empty → only the bot's own account
|
||||
may drive SAS; `@zhiqiang` is denied until the option is added via
|
||||
Settings → Devices & Services → Matrix E2EE → Configure.
|
||||
- Storage: `/config/.storage/matrix_e2ee_session.json` +
|
||||
`/config/.storage/matrix_e2ee_store/`. Backups:
|
||||
`/homeassistant/.matrix-e2ee-backups/` (incl. `matrix_e2ee.bak-20260820-v0.3.2`,
|
||||
`matrix_e2ee.bak-20260818-v0.3.1`,
|
||||
`matrix_e2ee.bak-20260818-v0.3.0`,
|
||||
`matrix_e2ee.bak-20260818-v0.2.10`,
|
||||
`matrix_e2ee.bak-20260816-v0.2.9`, `matrix_e2ee.bak-20260816-v0.2.8`);
|
||||
full HA backup slugs `3d9d36db` (pre-v0.1.4) + `9f223f35` (pre-v0.2.0).
|
||||
- v0.3.12: Matrix activity events + push diagnostics
|
||||
(`matrix_e2ee_message_received` / `matrix_e2ee_verification_done`).
|
||||
v0.3.9: Connection health binary sensor (W1N-185/#40), config-entry
|
||||
diagnostics (W1N-184/#39), SAS/command allowlist split
|
||||
`verification_peer_users` (W1N-156/#41), SAS/sync logs demoted
|
||||
warning→info/debug (W1N-188/#38), URL normalization + single-entry
|
||||
enforcement (W1N-190/#42).
|
||||
v0.3.8: `m.key.verification.done` handshake for request-based SAS
|
||||
(W1N-183/#35).
|
||||
v0.3.2: wizard waits for inbound SAS emojis before the compare step
|
||||
(W1N-182/#34).
|
||||
v0.3.1: verification wizard waits for a peer-initiated inbound SAS instead
|
||||
of the bot starting SAS (GitHub #33).
|
||||
v0.3.0: bot-initiated device verification wizard (W1N-180/#32).
|
||||
v0.2.11: `receive_mac_event` no longer overrides canceled state (W1N-179/#31).
|
||||
- v0.2.9: restore SAS emoji rendering after vodozemac migration (W1N-175/#29).
|
||||
v0.2.8: SAS commitment unpadded base64 for Element interop (W1N-174/#28).
|
||||
v0.2.7: SAS cancel code/reason logging. v0.2.6: verification state logging +
|
||||
request→ready bridge. v0.2.4: `_patch_nio_sas_timeout()` +
|
||||
`_repair_dropped_start()`; `VERIFICATION_TIMEOUT_SECONDS` 600→240.
|
||||
- Automation `1761188403590`「Matrix 聊天关卫生间灯」: trigger
|
||||
`matrix_e2ee_command` (command `关卫生间灯`), actions `light.turn_off` +
|
||||
`matrix_e2ee.send_message` (room `!gidvAzpDzwtzfEDrqu`).
|
||||
- **SAS not yet completed:** every device requires explicit `confirm_verification`.
|
||||
Encrypted-room commands stay fail-closed until `@zhiqiang`'s device is verified.
|
||||
Since v0.3.9 the SAS driver gate uses `verification_peer_users` (empty on
|
||||
this host) instead of `allowed_users` — add `@zhiqiang:chans.xyz` there
|
||||
before retrying the wizard. Three paths available: SAS manual confirm,
|
||||
fingerprint, or the device verification wizard (v0.3.0 bot-initiated,
|
||||
reworked in v0.3.1/v0.3.2 to wait for a peer-initiated inbound SAS from
|
||||
Element with emoji comparison), see
|
||||
[docs/home-assistant-matrix.md § Device verification](../docs/home-assistant-matrix.md).
|
||||
### Scribe long-term history (3.8.0 setup 2026-08-29; 4.4.0 verified 2026-09-13)
|
||||
|
||||
- **Scribe 4.4.0** (`/homeassistant/custom_components/scribe/`, HACS repo
|
||||
`jonathan-gtd/scribe`, = latest stable 2026-09-12; upgraded 2026-09-13 together
|
||||
with Core 2026.9.1 / HAOS 18.2), configured from
|
||||
`/homeassistant/scribe.yaml` — W1N-238 moved the block out of
|
||||
`configuration.yaml` on 2026-08-29 (main config now carries
|
||||
`scribe: !include scribe.yaml`; content moved verbatim; backup
|
||||
`configuration.yaml.bak-20260829-201724-w1n238`). Config entry
|
||||
`01KC2VFJWEQ3XDHY6TQKHPDVRB`, `source: import` — UI "Configure → Advanced"
|
||||
edits are overridden by the YAML on restart; treat YAML as authoritative.
|
||||
- TimescaleDB at `192.168.55.15:5432/scribe` (DB user `hass`; host in inventory,
|
||||
see [hosts/pgdb.md](../hosts/pgdb.md)). Database re-initialized 2026-08-29 14:06 CST
|
||||
(user-handled; earlier `relation "entities" does not exist` errors resolved).
|
||||
Health: `binary_sensor.scribe_database_connection`.
|
||||
- 2026-08-29 config applied (backup `/homeassistant/configuration.yaml.bak-20260829-scribe`):
|
||||
- `record_events: true` with `include_events` whitelist: `automation_triggered`,
|
||||
`matrix_e2ee_command`, `matrix_e2ee_message_received`,
|
||||
`matrix_e2ee_verification_done`, `script_started`, `tag_scanned`,
|
||||
`mobile_app_notification_action`, `homeassistant_start`, `homeassistant_stop`.
|
||||
- State noise trimmed: `exclude_domains` update/button; glob
|
||||
`sensor.zigbee2mqtt_bridge_*`; 4 hassio cpu/mem-percent entities.
|
||||
- Global `exclude_attributes` drops tianqi `hourly_*` arrays (~19 KB/state —
|
||||
the recorder-side `_unrecorded_attributes` patch does not apply to Scribe).
|
||||
- `enable_stats_io` + `enable_stats_size` on → 14 `sensor.scribe_*` stats
|
||||
entities (`scribe_states_written`, `scribe_events_written`, rates, sizes).
|
||||
- Verified post-restart 14:23 CST: writer started, `scribe_events_written=1`
|
||||
(homeassistant_start), states ~110/min, buffer 3, no scribe log errors.
|
||||
- **4.x upgrade核对 2026-09-13(只读 + 一处配置变更)**:live `manifest.json` =
|
||||
4.4.0。两个 4.0 breaking change 在本机都不需要动作——数据库是 3.x 结构
|
||||
(`states_raw` PK `(metadata_id, time)` 在,4.2.0 的启动态去重因此可用),
|
||||
TimescaleDB 2.29.2 已装。4.1.0 修了 `db_url` 优先级,YAML 里的
|
||||
`!secret scribe_url` 现在是权威。`scribe.yaml` 现有键在 4.4.0 全部仍然合法
|
||||
(未知键被忽略,`extra=vol.ALLOW_EXTRA`)。**配置优先级 YAML > entry
|
||||
`options` > entry `data` > 默认值**,而 `_resolve_settings` 读的是
|
||||
`hass.data[DOMAIN]["yaml_config"]`(只有 `async_setup` 会写),所以
|
||||
**YAML 改动必须重启 Core,reload config entry 不重读 YAML。**
|
||||
- **`stats_io_interval: 300`(2026-09-13 添加**,备份
|
||||
`/homeassistant/scribe.yaml.bak-20260913-191558`)。4.4.0 不再让 HA 每 30s
|
||||
轮询 I/O 统计传感器,改由集成自己每 60s 发布,间隔成为配置项。scribe 自己的
|
||||
传感器此前是本机自写历史的主要来源(变更前 24h:11 019 / 87 461 行状态 =
|
||||
12.6%),60s → 300s 把这部分降约 5 倍(每个 I/O 传感器约 1440 → 288 行/天)。
|
||||
验证:`ha core check` OK;重启 88s;`ScribeWriter started successfully`;
|
||||
无 scribe error/warning;scribe Repairs 问题 0 条;传感器发布间隔实测正好
|
||||
300s(11:19:26 → 11:24:26 UTC)。
|
||||
- **Retention 现在可用但刻意不设**:`retention_states` / `retention_events`
|
||||
(4.0.0)按间隔丢 chunk,留空 = 永久保留,符合本机定位(Scribe 是永久归档,
|
||||
recorder 保留 365 天)。注意 retention 是**绕过** entry `data` 副本读取的
|
||||
(`from_entry_data=False`),所以删掉 YAML 行即撤销策略。`db_schema`、
|
||||
`enable_rollups`、`scribe.purge` 同样未用:图表走 `sensor_minute` +
|
||||
`timescale_database_reader`(见 [hosts/pgdb.md](pgdb.md)),不吃 scribe 自己的
|
||||
视图,配置里也没有任何 `scribe.query` 调用。`flush_interval` 仍是 entry
|
||||
`data` 钉住的 5s——上游下一个版本把默认改成 30s,但 entry 值优先,要采用只能
|
||||
在 YAML 显式写 `flush_interval: 30`。
|
||||
- Recorder stays external-Postgres with `purge_keep_days: 365` (W1N-243,
|
||||
2026-08-29, raised from 30 — ~300 MB/yr, 1% of the 30G pgdb disk) for
|
||||
native UI per-change history; Scribe is the permanent archive. Long-term
|
||||
statistics stay permanent (not purged by `purge_keep_days`). Note:
|
||||
extending retention does **not** recover pre-2026-08-29 raw history
|
||||
(already purged); only `csg_history` day/month values cover that period.
|
||||
|
||||
### Config layout: scribe.yaml + templates/ merge (W1N-238, verified 2026-08-29)
|
||||
|
||||
- `configuration.yaml` line 29: `scribe: !include scribe.yaml`; line 9:
|
||||
`template: !include_dir_merge_list templates`. No `packages/`.
|
||||
- `scribe.yaml` (config root): the Scribe block, content identical to the
|
||||
former inline one; import semantics unchanged.
|
||||
- `templates/`: `csg_sensors.yaml` (12 template sensors, top-level **list**)
|
||||
+ `quick_sensors.yaml` (scaffold for Quick-derived `quick_*` sensors, empty
|
||||
list with convention header). **`!include_dir_merge_list` merges per-file
|
||||
lists; non-list files are silently skipped** — every file in `templates/`
|
||||
must be a top-level list (`- sensor:` blocks). Directory include only picks
|
||||
up `*.yaml`, so the `.bak` / `.pre-*` backups in the dir are ignored. After
|
||||
adding sensors, verify template-platform entity count = 12 + N (entity
|
||||
registry `platform: template`).
|
||||
- Convention (per review + W1N-233): pure sums/averages stay min_max helpers
|
||||
(e.g. `sensor.dang_qian_zong_gong_lu`); only template-logic derivations
|
||||
(ladder pricing, cross-entity conditions) go into `quick_sensors.yaml`.
|
||||
- Post-change verification 20:19 CST: `ha core check` ok, 92 s restart
|
||||
(2026.8.3), `binary_sensor.scribe_database_connection` on,
|
||||
`scribe_states_written` 18581→19426 growing, template entities still 12,
|
||||
csg sensors numeric, no scribe/template log errors.
|
||||
|
||||
### Timescale Plotly card + database reader (verified 2026-08-29)
|
||||
|
||||
Chart stack over the Scribe TimescaleDB archive. Upstream pair (no HACS;
|
||||
manual copies): reader `remmob/timescale_database_reader` **v1.1.0** (main
|
||||
`bb8776a`) + card `remmob/timescale-plotly-card` **2.2.0** (main `217961d`).
|
||||
|
||||
- **Reader integration**: `/homeassistant/custom_components/timescale_database_reader/`.
|
||||
Config entry `01M165P77QT1FQEAVPNZHDT82W` ("Scribe", `source: user`): connects
|
||||
`hass@192.168.55.15:5432/scribe` (credentials = `secrets.yaml` `scribe_url`),
|
||||
`table: sensor_minute`. Exposes no entities/services — it serves WS command
|
||||
`timescale/query` (window ≤ 365 d, ≤ 50 000 rows, `downsample` bucket seconds).
|
||||
Benign startup warning `Error executing test query: column "time" does not
|
||||
exist`: the self-test SQL assumes the LTSS column name; the scribe table uses
|
||||
`minute` — real queries work (verified: 70 rows for a live power sensor).
|
||||
- **Card**: `/homeassistant/www/community/timescale-plotly-card/timescale-plotly-card.js`
|
||||
(root-owned, same convention as HACS dirs). Lovelace resource (storage)
|
||||
id `2e360d17b5aa4ce59c2fd13c43b51215` →
|
||||
`/hacsfiles/timescale-plotly-card/timescale-plotly-card.js`, type `module`.
|
||||
Card config matches the entry by `database: scribe` (name from the reader
|
||||
entry). Updates: replace the file, resource URL unchanged — browsers need a
|
||||
hard refresh or a bumped `?v=` query on the resource URL.
|
||||
- **pgdb side** (`sensor_minute_aggregate` cagg + `sensor_minute` hypertable +
|
||||
every-minute refresh job): see [hosts/pgdb.md](pgdb.md) § Databases.
|
||||
- **Agent-side HA WebSocket without a long-lived token** (verified 2026-08-29):
|
||||
connect `ws://supervisor/core/websocket` with header
|
||||
`Authorization: Bearer $SUPERVISOR_TOKEN`, then send
|
||||
`{"type":"auth","access_token":"$SUPERVISOR_TOKEN"}` — the Supervisor proxy
|
||||
swaps it for a core token (works as the internal Supervisor admin user). Note
|
||||
`lovelace/resources/create` in HA 2026.8 takes `res_type` (NOT
|
||||
`resource_type`).
|
||||
- Scribe stores numeric sensor values in `states_raw.value` with `state` NULL,
|
||||
so `sensor_minute.state` shows `'0'` for numeric sensors; the card plots
|
||||
`avg_state` (from `value`) — expected, not a bug.
|
||||
- **Quick 仪表盘(`dashboard-quick`)图表套件**(2026-08-29 创建,经 WS
|
||||
`lovelace/config/save` 写入;W1N-230 修复 + W1N-231 round-2 改进):
|
||||
5 张 timescale 卡——大功率电器/常驻负载功率(按量级拆图,避免尖峰压扁
|
||||
<70 W 基线)、按插座用电量(`energy_mode` + cumulative/diff,数据质量前提
|
||||
见 pgdb 的 refresh 过程补丁)、室内外温湿度(温度左轴/湿度右轴,4 位置同色
|
||||
配对)、人体感应活动状态(3 个 `motion_state`,banded `state_map`
|
||||
none/small/medium/large → 0-11,per-entity `line_color` 红/蓝/绿)。
|
||||
空调实体引用为 `kong_diao_*`(`kong_tiao` 是笔误,W1N-230 修复;`grep -c
|
||||
kong_tiao` 应为 0)。灯区:2×2 嵌套 grid(`grid_options: {columns: "full"}`,
|
||||
内层 `columns: 2`)+ 4 卡统一 `mushroom-light-card`(显式 name、
|
||||
`use_light_color: false`、内联亮度/色温控制),heading icon
|
||||
`mdi:lightbulb-group`。heading badges:环境 4 温度(迷你/mini数显/数显/广州)、
|
||||
大功率电器 空调/电脑当前功率、常驻负载 总功率
|
||||
(`sensor.dang_qian_zong_gong_lu`,min_max **sum** helper,`round_digits: 0`,
|
||||
任一源掉线 fail-closed → unknown)。常驻负载图卡级 `fill: 'tozeroy'` +
|
||||
冰箱/主网络 per-entity `fill_color`(线色 20% 透明)+ 其余 5 条 `fill: false`
|
||||
(per-entity fill 逐系列退出,卡 JS `seriesConfig.fill !== false`)。
|
||||
布局:视图 `type: sections` + `max_columns: 4`;灯/用电/环境/人体感应
|
||||
`column_span: 4`,功率两图拆两个 `column_span: 2` 分区**并排**(等高 280px,
|
||||
桌面并排、手机回落堆叠;去卡内 title 省半宽图垂直空间)。
|
||||
**分区/卡片是两套尺寸键,不可混用**:分区宽 = `column_span`
|
||||
(`hui-sections-view.ts` 缺省按 1 列渲染,绝不省略);卡片宽 =
|
||||
`grid_options: {columns: <n|"full">}`(`hui-card.ts` 只读 `config.grid_options`,
|
||||
写在卡片上的 `column_span` 被静默忽略;缺省 12 列,分区内格 = 12 × 分区
|
||||
span,故 span-4 分区里缺省卡片只有 1/4 宽)。
|
||||
修改前备份:`/homeassistant/.lovelace-backups/dashboard-quick-*.json`
|
||||
(W1N-230 修复: `20260829-190256`;round-2 改进: `20260829-194040`)。
|
||||
|
||||
- **Quick 时间范围扩容 (2026-09-13, VPS-92)**: 用户反馈「48 小时不够」。
|
||||
各 timescale 卡可选档上调——大功率电器/常驻负载 `…,24h` → `+3d,7d`;
|
||||
环境 `6h,12h,24h,48h` → `+7d,14d,30d`;人体感应 `…,24h` → `+3d,7d`;
|
||||
用电量(按插座) `energy_time_ranges` `today,week,month,custom` → `+3mo`。
|
||||
**默认档未改**(6h / 6h / today / 24h / 12h)。卡片 JS 只接受
|
||||
`<n>m|<n>h|<n>d`(`parseDurationToMs` 正则 `/^(\d+)(m|h|d)$/`,
|
||||
仅 m/h/d,无 w)与命名档 `today|week|month|3mo|6mo|year|years|custom`;
|
||||
`energy_mode` 卡必须用后者。**数据下界注意**:scribe `sensor_minute`
|
||||
目前最早只到 **2026-08-29**,所以 >15d 的档(14d 边缘、30d 明显)前半段
|
||||
会是空白,等归档继续累积才好看。备份
|
||||
`.lovelace-backups/dashboard-quick-20260913-190912-pre-timerange.json`。
|
||||
|
||||
### 地图仪表盘:CARTO keyed tiles via `custom:map-card` (verified 2026-08-30, W1N-261)
|
||||
|
||||
- **背景:** CARTO 自 2026-08-26 起对无 key 栅格瓦片打 "API KEY REQUIRED"
|
||||
水印,内置地图卡/zone 编辑器全部受影响。Core 2026.8.3 的 `MapCardConfig`
|
||||
**没有任何瓦片配置项**(frontend 20260729.7 源码核对:
|
||||
`setup-leaflet-map.ts` 硬编码 CARTO voyager URL)。上游修复是 2026.9.0b1
|
||||
起改用 OSMF 矢量瓦片(frontend PR #53816),stable 预计 2026-09-02 前后。
|
||||
- **变更:** 「地图」仪表盘(url_path `map`,storage)唯一 map 卡替换为
|
||||
`custom:map-card`([nathan-gs/ha-map-card](https://github.com/nathan-gs/ha-map-card)
|
||||
**v1.16.0**,手动安装非 HACS):`tile_layer_url` =
|
||||
`https://{s}.basemaps.cartocdn.com/rastertiles/voyager/{z}/{x}/{y}.png?key=<CARTO_KEY>`
|
||||
(配 `tile_layer_options: {subdomains: abcd, maxZoom: 20}` + OSM/CARTO
|
||||
attribution)。实体不变:2 person + 4 zone(zone 用 `display: icon` +
|
||||
`circle: auto`,circle 读实体 `radius` 属性画半径圈)。
|
||||
- **CARTO key 是 secret**: 只存在于服务端 lovelace 存储(dashboard `map`
|
||||
的卡片配置)和用户本人处;勿写入本仓库或 Linear。
|
||||
- **文件/资源:** `/homeassistant/www/community/ha-map-card/map-card.js`
|
||||
(root:root 644,678554 B,sha256
|
||||
`f30dfb606e858d2216d5198d8cf758ce956d127006ebd7d66d4329153a247ec2`);
|
||||
Lovelace resource(storage)id `9d2b50b52c60420d89ebd041f722cf60` →
|
||||
`/hacsfiles/ha-map-card/map-card.js`,type module(WS
|
||||
`lovelace/resources/create`,2026.8 参数名 `res_type`)。升级 = 手动替换
|
||||
该文件(不在 HACS 管理下,浏览器需强刷)。
|
||||
- **备份:** `/homeassistant/.lovelace-backups/dashboard-map-map-20260830-133714.json`
|
||||
(还原 = 把备份里的 `views[0].cards[0]` 写回后再 WS `lovelace/config/save`
|
||||
url_path `map`)。
|
||||
- **验证 8/30:** 同瓦片无 key=水印 / 带 key=干净(256×256 PNG 视觉对比);
|
||||
resource HTTP 200 text/javascript;WS 读回卡片配置(type/entities/key/
|
||||
attribution/options)全部符合;HA 主机 `curl -4` 带 key 瓦片 200。
|
||||
- **Follow-up:** Core 升 2026.9.0 stable 后内置地图/zone 编辑器自动切
|
||||
OSMF 矢量瓦片;届时可保留 custom 卡(继续 keyed CARTO)或用备份还原
|
||||
内置卡。zone 编辑器等其余内置地图的水印在 2026.9 前无解。
|
||||
|
||||
## Known issues
|
||||
|
||||
**Bluetooth hci0 instability — RTL8821CS (verified 2026-08-13, W1N-74):**
|
||||
The local Bluetooth controller hci0 is an **RTL8821CS** combo chip on the
|
||||
x88 Pro board. Kernel logs show recurring `hci0: hardware error 0x00`,
|
||||
`Opcode 0x200c tx timeout` (HCI_LE_Set_Scan_Parameters), `Unable to disable
|
||||
scanning: -110`, `Peer device has reset` — the chip hardware-stalls during
|
||||
active scanning. HA's `bluetooth_auto_recovery` power-cycle then times out
|
||||
after 5 s and retries every ~2 min:
|
||||
`bluetooth_auto_recovery.recover: Could not reset the power state of the
|
||||
Bluetooth adapter hci0 ... due to timeout after 5 seconds`. The HAOS image
|
||||
already ships custom systemd units to cope (`x88-bt-hci-recovery.service` and
|
||||
a "Patch HA Bluetooth scanner mode for x88 RTL8821CS" service, visible in host
|
||||
journal). **No user impact:** there are **no BLE entities** in HA
|
||||
(xiaomi_ble / bthome / led_ble / bluetooth / esphome domains are all empty;
|
||||
platforms merely load from stray advertisements). Real IoT devices are Zigbee
|
||||
(via Zigbee2MQTT) or WiFi/MQTT/cloud. An ESPHome Bluetooth-proxy ESP32
|
||||
(`/config/esphome/bluetooth.yaml`, bluetooth_proxy: active, WiFi `ubnt-haas`)
|
||||
is configured but currently offline (ESPHome add-on stopped, port 6053
|
||||
unreachable) and produced no entities. Follow-up (optional): disable the
|
||||
local adapter and rely on the ESPHome proxy, or stop the bluetooth
|
||||
integration entirely.
|
||||
|
||||
**eMMC disk lifetime 10% (verified 2026-08-13, W1N-76):** `ha host info`
|
||||
reports `disk_life_time: 10` — the boot eMMC (`/dev/mmcblk2`, CJTD4R
|
||||
`0xacacc064`, 64 GB) has ~10% life left. `disk_free: 40.2/56.4 GB`. Full
|
||||
backup `pre-maintenance-20260813` (slug `411a4ba5`, 144.26 MB) taken
|
||||
2026-08-13 covers current config; monitor `disk_life_time` on each health
|
||||
snapshot and plan a disk replacement / data-disk migration before the eMMC
|
||||
fails.
|
||||
|
||||
## Matter Server (verified 2026-08-21)
|
||||
|
||||
- Add-on `core_matter_server` (`homeassistant/aarch64-addon-matter-server`) runs the Matter
|
||||
commissioner on this host (host networking; add-on container `app_core_matter_server`).
|
||||
- **After the ISP PD prefix rotates (PPPoE redial), the add-on can cache a stale IPv6 GUA
|
||||
in its mDNS advertisement** — clients trying that dead address make Matter
|
||||
commissioning/connection fail. Fix: restart the add-on so it re-enumerates addresses:
|
||||
`ssh hassio@hass.windy.lan 'sudo -n -i ha apps restart core_matter_server'`
|
||||
(`ha addons restart ...` also works; "addons" is deprecated in favor of "apps").
|
||||
- Verified 2026-08-21 (W1N-207): stale `240e:3bd:234:2f22:*` AAAA in mDNS removed by
|
||||
restart; advertisement now carries only current GUA `240e:3bd:235:1fb2:*` + link-local;
|
||||
CASE sessions with Aqara M3 / SmartThings hubs resumed over IPv6 link-local.
|
||||
|
||||
> **Open items (2026-08-21, W1N-207):** a phone on LAN55 was querying five known
|
||||
> `_matter._tcp` instances of which only HA answered — the other Matter nodes are
|
||||
> offline / not announcing (device-side; user to confirm power/Wi-Fi). HA's IPv6
|
||||
> default route via NetworkManager was observed missing once (curl -6 intermittent,
|
||||
> while ping6 and `curl -6 --noproxy` work) — not the Matter root cause; re-check
|
||||
> on the next health snapshot.
|
||||
|
||||
Verified 2026-08-23 (read-only, W1N-207): add-on `started`, version `9.0.4`, no
|
||||
update pending; current GUA `240e:3bd:238:4812:*` (PD rotated again since 08-22)
|
||||
advertised correctly over v4+v6. Both ESP32-C2 bulbs now announce `_matter._tcp`
|
||||
(multi-fabric, including this host's fabric `DCE86145C137AF0E`) — but they
|
||||
**refuse TCP 5540 on IPv4 and IPv6**, so matter-server holds **zero established
|
||||
:5540 sessions** (device-side failure mode C; no errors logged — see
|
||||
[docs/matter-pairing-troubleshoot.md §8](../docs/matter-pairing-troubleshoot.md)).
|
||||
|
||||
## 马桶换气电源(Matter 插座,半计量)+ 电量估算 (2026-09-13)
|
||||
|
||||
**设备**:Matter `Smart Plug`(SIXWGH,`model_id 3596`,hw 1.0 / sw 1.3.0),node 18
|
||||
(0x12),`device_id 5ef1850953466d6e7a9c6b901fbebe1c`,config entry
|
||||
`01JF51VQ48PGJGXX3RNAG6MVAA`,区域**卫生间** (`wei_sheng_jian`),label `power`;
|
||||
2026-09-13 17:58 CST 配对。实体:
|
||||
`switch.wei_sheng_jian_ma_tong_huan_qi_dian_yuan`(插座)、
|
||||
`sensor.…_dian_yuan`(电源 W)、`sensor.…_dian_ya`(电压 V)、
|
||||
`sensor.…_you_gong_dian_liu`(有功电流 A)、`sensor.…_dian_li`(电力 kWh,
|
||||
**永久 unknown**)。
|
||||
|
||||
**根因(实测 Matter 属性,node 18)**:电量簇 0x0091 `FeatureMap = 13`
|
||||
(IMPE|CUME|PERE,即**声明**支持导入/累计/周期电量),但
|
||||
`CumulativeEnergyImported (0x0001)` 恒为 `null`,`PeriodicEnergyImported
|
||||
(0x0003)` 带载也恒为 `{Energy: 0}`;`CumulativeEnergyExported (0x0002)`
|
||||
不存在(EXPE 未声明,自洽)。HA 只用 `CumulativeEnergyImported` 建能量实体
|
||||
(`components/matter/sensor.py:1083`,`allow_none_value=True`)→ 该实体
|
||||
**永远不会出数**。**功率计量本身正常**:0x0090 `FeatureMap = 2` (ALTC),
|
||||
Voltage / ActiveCurrent / ActivePower 都随负载变化(实测 220.3 V / 118 mA /
|
||||
24.7 W,HA `电源` 0.0→24.9 W 有历史)。厂商 `update` 实体报无新固件。
|
||||
|
||||
**处理(方案 A:功率积分补电量)**:
|
||||
|
||||
- 新建 **Integration (Riemann sum) 辅助元素**:config entry
|
||||
`01M2D53T188FW8WEC547ENHSVH`(domain `integration`,state `loaded`),
|
||||
source `sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_dian_yuan`,
|
||||
`method: trapezoidal`、`unit_prefix: k`、`unit_time: h`、`round: 3`、
|
||||
`max_sub_interval: 60s`。
|
||||
- 实体 `sensor.wei_sheng_jian_ma_tong_huan_qi_dian_yuan_energy`(创建时 HA
|
||||
自动生成 `…_dian_yuan_ma_tong_huan_qi_dian_yuan_dian_liang`,随后立即
|
||||
`config/entity_registry/update` 改名为 `<插座>_energy` 以对齐约定;
|
||||
该实体新建、无引用,改名安全),friendly name「马桶换气电源 电力」,
|
||||
unit kWh、`device_class: energy`、**`state_class: total`**——能源仪表盘
|
||||
允许 `TOTAL` 与 `TOTAL_INCREASING`(`components/energy/validate.py:279`)。
|
||||
- **能源仪表盘** (`/energy`):grid 源 `[8]` 由 `…_dian_li` 改为 `…_energy`,
|
||||
其余 8 条插座源未动。注意这 9 条「插座」全部以 `type: grid` 注册,被当作
|
||||
全屋用电代理;`switch` 卡片所在的 Grid 卡片此前第 9 行是空的,即本次修复点。
|
||||
- **Quick 仪表盘**:「用电量(按插座)」图第 11 项由 `…_dian_li` 改为
|
||||
`…_energy`;新增 `column_span: 2` 的「开关」区块(heading + tile
|
||||
`switch.…` + `toggle` feature + 功率徽标)→ 视图 6→7 分区。
|
||||
|
||||
**口径警告**:`…_energy` 是**估算值**(Riemann 积分,只在 HA 运行期间累计、
|
||||
非账单级),与另外 8 个原生计量插座的累计电量口径不同;功率传感器更新
|
||||
间隔约 5–10 s(实测 24.9/24.8/25.0 W 抖动),加 `max_sub_interval: 60s`
|
||||
保证静默时也继续累计。
|
||||
|
||||
**Agent 侧建辅助元素的方法(2026-09-13 实测)**:HA 的 config flow 走
|
||||
**REST**(WS 只有 `config_entries/flow/progress|subscribe`,没有 start)。
|
||||
经 supervisor 代理即可,无需 HA 长连接/长寿命 token:
|
||||
|
||||
```bash
|
||||
# SUPERVISOR_TOKEN 由 sudo -n -i 提供
|
||||
curl -s -X POST -H "Authorization: Bearer $SUPERVISOR_TOKEN" \
|
||||
-H "Content-Type: application/json" -d '{"handler":"integration"}' \
|
||||
http://supervisor/core/api/config/config_entries/flow # → {flow_id, step_id:"user", data_schema}
|
||||
curl -s -X POST -H "Authorization: Bearer $SUPERVISOR_TOKEN" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"name":"…","source":"sensor.x","method":"trapezoidal","round":3,
|
||||
"unit_prefix":"k","unit_time":"h","max_sub_interval":{"minutes":1}}' \
|
||||
http://supervisor/core/api/config/config_entries/flow/<flow_id> # → create_entry
|
||||
```
|
||||
|
||||
`auth/long_lived_access_token` 在 supervisor 代理身份下**失败**
|
||||
(`unknown_error`),故无法用长寿命 token 开浏览器会话;`DurationSelector`
|
||||
的值是 `{"minutes":1}` 形式(`cv.time_period`)。
|
||||
|
||||
**备份/回滚**:`.lovelace-backups/dashboard-quick-20260913-181251-pre-ma-tong-plug.json`
|
||||
(改动前原件)、`…-20260913-183210-pre-repoint.json`(改名/换源前);
|
||||
`.ha-backups/energy-20260913-183135-pre-ma-tong-repoint.json`(能源 prefs)。
|
||||
回滚 = 把能源 prefs 的源 [8] 指回 `…_dian_li` + 还原 Quick 面板 JSON;
|
||||
如需彻底放弃估算电量 = 删除 config entry `01M2D53T188FW8WEC547ENHSVH`。
|
||||
|
||||
**验证 (2026-09-13 18:3x)**:`…_energy` 0.002→0.003 kWh 且随 24.6 W 负载
|
||||
增长(换气扇关掉后回落 0.0 W,累计值保留);`recorder/list_statistic_ids`
|
||||
已含该实体;Quick 面板 WS 读回 7 分区、用电量图 11 项指向新实体、旧
|
||||
`_dian_li` 引用 0 处;能源 prefs 读回 9 源、第 9 条为新实体。
|
||||
**`energy/validate` 已全绿**(9 源 0 issue):创建后 ~5 min 内曾报
|
||||
`statistics_not_defined`(recorder 的统计任务周期是 5 min,`statistics_meta`
|
||||
行由该任务建立),18:39 复核时已自动消失——建辅助元素后**不要**把这条
|
||||
瞬时告警当作失败。
|
||||
|
||||
## Related docs
|
||||
|
||||
- [runbooks/home-assistant-maintenance.md](../runbooks/home-assistant-maintenance.md) — `ha` CLI maintenance runbook + [script](../runbooks/scripts/ha-maintenance.sh); custom-component zip install is §7
|
||||
- [docs/lan-overview.md](../docs/lan-overview.md) — LAN map and gw port-forward
|
||||
- [hosts/dns.windy.lan.md](dns.windy.lan.md) — `hass.windy.lan` / `hass.local` rewrites
|
||||
+73
-3
@@ -56,7 +56,8 @@ See full shape in [docs/pdns-upstream.md](../docs/pdns-upstream.md). Live secret
|
||||
| `poweradmin` | poweradmin | Up (healthy) | `poweradmin/poweradmin:stable` |
|
||||
| `pdns_pgweb` | pgweb | Up | `sosedoff/pgweb:0.16.2` |
|
||||
| `pdns-backup` | backup | Up | `postgres:16` (scheduler) |
|
||||
| `powerdns-admin` | *(orphan)* | Exited | legacy PDA UI — not in active compose |
|
||||
|
||||
> Legacy PDA UI container `powerdns-admin` (orphan, Exited) was removed 2026-08-12 (W1N-59).
|
||||
|
||||
### Network model
|
||||
|
||||
@@ -100,9 +101,60 @@ See full shape in [docs/pdns-upstream.md](../docs/pdns-upstream.md). Live secret
|
||||
|
||||
**Quirk:** `backend` is internal — backup must not use Alpine + runtime `apk`/`crond`. Uses `postgres:16` + `backup-scheduler.sh` (fixed 2026-08-01).
|
||||
|
||||
## Other software on this host (stubs)
|
||||
## RustDesk Server
|
||||
|
||||
`/opt/traefik`, `/opt/adguard`, `/opt/remark42`, `/opt/rustdesk`, `/opt/nginx-manager`, …
|
||||
**Status: operational** (hbbs + hbbr Up; image pinned `1.1.14`; relay address fixed 2026-08-12, W1N-59).
|
||||
|
||||
| Item | Value |
|
||||
|------|--------|
|
||||
| Install path | `/opt/rustdesk` |
|
||||
| Compose | `/opt/rustdesk/compose.yml` |
|
||||
| Containers | `hbbs` (rendezvous), `hbbr` (relay) |
|
||||
| Image | `rustdesk/rustdesk-server:1.1.14` (pinned) |
|
||||
| Relay (hbbr) | `hk2.chans.xyz:21117` — advertised to clients via `hbbs -r` |
|
||||
| Rendezvous (hbbs) | `21115/tcp` (NAT test), `21116/tcp+udp`, `21118/tcp` (ws) |
|
||||
| Relay (hbbr) | `21117/tcp`, `21119/tcp` (ws) |
|
||||
| Public IP | `154.36.174.161` |
|
||||
| Health | [runbooks/rustdesk-health.md](../runbooks/rustdesk-health.md) |
|
||||
|
||||
**Note:** the relay hostname in `hbbs -r` must resolve to this host's public IP
|
||||
(`154.36.174.161`). `hk2.chans.xyz` resolves correctly; the previously used
|
||||
`hk2.wsvc.info` had **no DNS record** and broke relay connectivity for clients
|
||||
(fixed 2026-08-12, W1N-59).
|
||||
|
||||
## Other software on this host (confirmed 2026-08-12)
|
||||
|
||||
Verified live via `docker ps` / port scan. Each runs as a separate compose
|
||||
project under `/opt/<name>` and is fronted by Traefik where noted.
|
||||
|
||||
| Service | Path | Container(s) | Image | Ports / notes |
|
||||
|---------|------|--------------|-------|---------------|
|
||||
| Traefik | `/opt/traefik` | `traefik` | `traefik:v3.6.2` | `80`, `443` (TLS entry), `8080` (dashboard) |
|
||||
| AdGuard Home | `/opt/adguard` | `adguardhome` | `adguard/adguardhome:latest` | DoH `5443`, DoT `853` (bridge; no LAN `:53`) |
|
||||
| Remark42 | `/opt/remark42` | `remark42` | `ghcr.io/umputun/remark42:latest` | no host ports; via Traefik (in-container `8080`) |
|
||||
|
||||
### Traefik dashboard auth
|
||||
|
||||
| Item | Value |
|
||||
|------|-------|
|
||||
| Dashboard URL | `https://npm.chans.xyz` (Traefik `api@internal` router), also host `:8080` |
|
||||
| Auth | HTTP Basic via Traefik `basicauth` middleware (label `dashboard-auth`) |
|
||||
| User | `windy` — stored as a **bcrypt** hash (plaintext never stored) |
|
||||
| Hash generator | `/opt/traefik/generate-dashboard-auth.sh` (bcrypt; auto `$`→`$$` compose escaping) |
|
||||
| Config | `/opt/traefik/compose.yml` (label `traefik.http.middlewares.dashboard-auth.basicauth.users`) |
|
||||
|
||||
**Password rotated 2026-08-12** from apr1/MD5 to bcrypt via the generator script; the
|
||||
plaintext lives only in the operator's password manager, never in this repo.
|
||||
To rotate again: `cd /opt/traefik && ./generate-dashboard-auth.sh windy`, paste the
|
||||
printed label into `compose.yml`, then `docker compose up -d --force-recreate traefik`.
|
||||
|
||||
`/opt/nginx-manager` was a leftover (compose + `data/` + `letsencrypt/`, no running
|
||||
container) and was **removed 2026-08-12**; pre-deletion backup:
|
||||
`/opt/backups/nginx-manager-20260812.tar.gz`.
|
||||
|
||||
Health coverage: these auxiliary services are checked by the `hk2aux`
|
||||
health-check profile (`ansible/roles/healthcheck`). Run:
|
||||
`cd ansible && ansible-playbook playbooks/health-report.yml --limit powerdns`.
|
||||
|
||||
## Ops / runbooks
|
||||
|
||||
@@ -133,6 +185,23 @@ dig @202.91.35.141 SOA wsvc.info +short
|
||||
|
||||
On-server docs: `/opt/pdns/README.md`, `CHANGELOG.md`.
|
||||
|
||||
## Disk / logging (VPS-81, 2026-09-02)
|
||||
|
||||
Root disk cleanup performed (runbook: [host-disk-cleanup](../runbooks/host-disk-cleanup.md)):
|
||||
|
||||
- Root `/` (20G vda1): 76% used → **38% used** (15G → 7.1G; free 4.7G → 12G).
|
||||
- **AGH log flood root cause fixed**: `/opt/adguard/conf/AdGuardHome.yaml`
|
||||
`log.verbose: true → false` (backup `AdGuardHome.yaml.bak-20260902-vps81`).
|
||||
Verbose debug was streaming to stderr → container `json.log` (~120MB/day);
|
||||
`log.file: ""` makes AGH's own rotation keys inert. Restart only (no recreate).
|
||||
- Journald capped: `/etc/systemd/journald.conf.d/00-vps81.conf`
|
||||
`SystemMaxUse=200M`; journal vacuumed to ~96M.
|
||||
- Docker: engine **29.7.2**; 14 unused images removed (kept `pdns-auth-50:5.0.5`
|
||||
rollback pin); 12 orphan anonymous volumes + build cache pruned. In-use
|
||||
volumes intact (`pdns_dbdata`, `b594d738…` PG data, `e855d078…` backup).
|
||||
- Follow-up: re-check AGH `json.log` growth **2026-09-09** (one-week checkpoint);
|
||||
global docker log rotation only if still needed.
|
||||
|
||||
## Verified
|
||||
|
||||
Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded.
|
||||
@@ -142,3 +211,4 @@ Last checked: **2026-08-01 21:40 CST** — operational; docs audit recorded.
|
||||
- `only-notify=` + `also-notify=202.91.35.141`; MASTER `domains.master` cleared
|
||||
- https://pdns.wsvc.info → **302**; https://pgweb.wsvc.info → **401**
|
||||
- Hardening backlog: API/DB credential rotation + TSIG rotate (see upstream doc)
|
||||
- 2026-09-02 (VPS-81): post-cleanup verified — 10 containers Up (adguardhome healthy), DNS SOA/NS + web endpoints OK; see Disk/logging section above.
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
# pgdb — TimescaleDB (PG18, Docker)
|
||||
|
||||
## Role and access
|
||||
|
||||
| Item | Value |
|
||||
|---|---|
|
||||
| Role | TimescaleDB PostgreSQL 18 (Docker) — Home Assistant recorder 后端(`hass`/`scribe` 库) |
|
||||
| IPv4 | `192.168.55.15` (LAN55) |
|
||||
| DNS | (none) |
|
||||
| SSH | `ssh -4 windy@192.168.55.15`(key auth 已验证可用 2026-08-29;agent 沙箱用 `ssh -F /dev/null -o BatchMode=yes`;password auth 亦可) |
|
||||
| Host | PVE 管理的 QEMU VM(i440FX,**VMID 100**),Debian 13 (trixie),内核 6.12.105;宿主机 **pve2 `192.168.55.25`**(Proxmox 9.2.2,SSH `root@192.168.55.25`,`onboot: 1`,QEMU guest agent 已装;2026-08-31 补记) |
|
||||
| Resources | 3 GB RAM(08-30 13:58 由 2G 上调、删除 balloon/ksm/shares 后重启生效)/ 30 GB disk(26 G 空闲) |
|
||||
| Docker | 29.7.2;容器 `timescaledb` = `timescale/timescaledb:latest-pg18`(PG **18.6** + TimescaleDB **2.29.2**,Apache-2.0 版) |
|
||||
| Ports | `192.168.55.15:5432`(PG,IPv4 only);`192.168.55.15:8081`(pgweb GUI,basic auth) |
|
||||
|
||||
## Databases
|
||||
|
||||
| DB | Owner | Size | 用途 |
|
||||
|---|---|---|---|
|
||||
| `hass` | hass | ~406 MB(2026-09-13) | HA recorder(states/events/statistics),客户端 HAOS `192.168.55.11` |
|
||||
| `scribe` | postgres | ~2.6 GB(2026-09-13) | HA scribe 集成(entities/areas/devices 注册表同步 + `states_raw`/`events` hypertable + `csg_history` 长期归档表);体积由 `sensor_minute` 图表管道主导(2.26 GB),见 Known issues |
|
||||
| `postgres` | postgres | ~9 MB | 默认库 |
|
||||
|
||||
## Ops notes
|
||||
|
||||
- **Docker compose 管理**(2026-08-29 改造):`/opt/database/docker-compose.yml`(源码在仓库 `compose/pgdb/`)+ `/opt/database/.env`(0600,密钥)+ `/opt/database/pgweb-bookmarks/`(0600,bookmark 含 DB 密码)。三个服务:
|
||||
| 服务 | 镜像 | 端口 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `timescaledb` | `timescale/timescaledb:latest-pg18` | `192.168.55.15:5432`(IPv4 only) | PG 18.6 + TS 2.29.2;healthcheck pg_isready;`restart: unless-stopped` |
|
||||
| `pgweb` | `sosedoff/pgweb:latest`(v0.17.0) | `192.168.55.15:8081` | Web GUI:http://192.168.55.15:8081;basic auth(用户名/密码见 .env `PGWEB_AUTH_USER/PASS`);`--readonly --sessions --bookmarks-only --bookmarks-dir /bookmarks`(v0.17.0 不读 PGWEB_BOOKMARKS_DIR env,必须用 flag);bookmarks = hass/scribe |
|
||||
| `pg-backup` | `prodrigestivill/postgres-backup-local:latest`(=PG18 客户端) | — | 每日 02:00(`TZ=Asia/Shanghai`,本地时区)`pg_dump -Fc` 三库 → `/opt/database/backups/{daily,weekly,monthly}`;保留 7 天/4 周/6 月;`BACKUP_ON_START` |
|
||||
- **数据盘**:`/dev/sdb1`(32G ext4,label `pgdata`)挂载 `/srv/pgdata`,fstab 按 `UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86`(defaults,noatime)持久化(2026-08-29 迁移)。容器 bind mount `/srv/pgdata:/var/lib/postgresql`。
|
||||
- 容器内 postgres 用户 uid/gid = **70**(Debian 系,非 999);迁移数据后需 `chown -R 70:70`。
|
||||
- **密码**:postgres 超级用户已换强密码(hex,存 `/opt/database/.env` 0600,2026-08-29)。HA 用 `hass` 角色不受影响。
|
||||
- **备份**:由 `pg-backup` 容器接管(2026-08-29),宿主机 cron 与 `/opt/database/pg-backup.sh` 已退役。恢复用 `pg_restore`(custom format)——2026-08-29 已实测还原 hass 库 dump(states 10014 行)成功。
|
||||
- **认证**:外部连接 scram-sha-256(密码必填,改密码有效);容器内 loopback 为 trust(官方镜像默认)。
|
||||
- **回滚**:旧启动命令保留在 `/opt/database/run`(容器无状态,数据在 /srv/pgdata);旧匿名卷 `9375195843b950f4e04c34872409ca095e1136520dd019a8e86e2794be06c236`(根盘 ~82M)保留作兜底,确认稳定后可 `docker volume rm`。
|
||||
- **开机自愈**(2026-08-30):新增 systemd oneshot `pgdb-compose.service`(enabled,源码在仓库 `compose/pgdb/pgdb-compose.service`):`After=network-online.target docker.service`,开机后幂等执行 `docker compose up -d`,重试直到 `192.168.55.15:5432` 监听,重试耗尽 `--force-recreate` 兜底(数据在 bind mount,无损)。原因:2026-08-30 开机竞态——docker 恢复容器时 VM IP 尚未可绑(EADDRNOTAVAIL),timescaledb/pgweb 启动失败且 docker 不重试。手动重跑:`sudo systemctl restart pgdb-compose.service`。
|
||||
- 本机无防火墙(ufw/nft/iptables 均未装)——待办:如要彻底隔离可加 ufw 白名单 192.168.55.11。
|
||||
- `/opt/database/backups/` 根下残留 `*-2026-08-29_1359.dump`(compose 化之前旧备份机制产物)与 `backup.log`——健康检查只看 `daily/`,残留可清理。
|
||||
- **Runbooks**:[pgdb-health](../runbooks/pgdb-health.md)(只读健康检查)、[pgdb-restore](../runbooks/pgdb-restore.md)(pg_restore 还原)、[pgdb-update](../runbooks/pgdb-update.md)(镜像/compose 升级)。
|
||||
- **CSG 长期归档(2026-08-29, W1N-243)**:`csg_history` 表(`period date / kind('day'|'month') / usage_kwh / cost / ladder / balance / updated_at`,PK(period,kind),`GRANT SELECT TO hass`)保存南方电网有价值数据:day = 逐日(昨日用电/费用/阶梯/余额,2026-07-01 起),month = 当月累计(用电/费用,2025-01 起)。由 TimescaleDB 每日任务 **1008** `csg_daily_snapshot()`(22:30 Asia/Shanghai;**TS job 非 pg_cron**,本库未装 pg_cron)upsert 维护:取「最新有值行」防瞬态 unknown 竞态;日费用缺原生 `latest_day_cost` 时回退 = 昨日用电 × 当前档费率(模板 `csg_current_ladder_tariff` 0.639);月费用回退模板 `csg_this_month_ladder_cost`。验证:day 08-28 = 7.66 / 4.89474 / 二档 / 0,month 08 = 302.47 / 180.28。回填来源:集成 attributes `history_data`(59 天)+ `by_month`(19 月)——08-29 前唯一残存历史。回滚:`DROP TABLE csg_history` + `SELECT delete_job(1008)`。
|
||||
|
||||
## Known issues
|
||||
|
||||
- 2026-09-13:**`sensor_minute` 体积构成与压缩窗口(只读诊断,暂不处理)**。`scribe` 库 2.6 GB = `sensor_minute` **2.26 GB**(850 万行 / 16 天,约 56–59 万行/天 = 331 实体 × 1440 分钟 LOCF)+ `states_raw` 290 MB + `events` 1.5 MB;`hass` 库另 406 MB。2.26 GB 中 1.50 GB 是 chunk `[09-03,09-10]`、0.76 GB 是 `[09-10,09-17]`,**都还没到压缩窗口**——TimescaleDB 的 `compress_after` 按 **chunk 结束时间**判断(09-10 结束 + 7 天 = **09-17** 才合格),所以「7 天 chunk + 7 天 compress_after」的设计下限就是盘上常驻近 14 天原始数据;已压缩的 `[08-27,09-03]` 从 1.04 GB → **1.5 MB**(LOCF 重复度极高,~700:1)。任务 1005 健康(30 成功 / 0 失败,最近 09-13 04:18 跑过但无合格 chunk);1002/1003/1006/1007 亦全 Success。稳态估算 ≈ 2 个未压缩 chunk(3–4.5 GB)+ 已压缩归档(约 1.5 MB/周 ≈ 80 MB/年)≈ **4–5 GB 平台期**;pgdata 卷 32 G 当前用 3.2 G,可用 27 G,**无需处理**。复查点 **2026-09-17 之后**:`_hyper_4_6_chunk` 应转为 `compressed=true` 且库体积回落;若仍为 false 才需动手(手动 `compress_chunk()` 或调小 `compress_after`)。可选调优:chunk 间隔 7 天 → 1 天 + `compress_after` → 2 天,把常驻未压缩量压到 <1 GB(`set_chunk_time_interval` 只对新 chunk 生效,旧 chunk 不重切)。诊断命令:`select chunk_name, is_compressed from timescaledb_information.chunks where hypertable_name='sensor_minute';` + `pg_database_size('scribe')`。**注意:这是 pgdb 侧对象,HA/scribe 的 `retention_states` 管不到它;HA 侧唯一杠杆是少记/少画(等于砍图)。**
|
||||
- 2026-08-29:HA 侧 HACS 集成 `custom_components.scribe`(YAML `scribe: db_url:`,连 `scribe` 库)建表被拒(`permission denied for schema public`,hass 无 CREATE 权限),之后持续报 `relation "entities" does not exist`。**已解决**:① `GRANT CREATE ON SCHEMA public TO hass;`(scribe 库)② 重启 HA Core 触发重跑建表。重启后自动创建 `entities`(1591 行)/`users`/`areas`/`devices`/`integrations`/`states_raw` 表并启用 TimescaleDB 时间序列能力。报错已停止(最后一条 06:06 UTC),`states_raw` 持续写入。2026-08-29 复查:scribe 现有**两个** hypertable——`states_raw`(segmentby `metadata_id`、orderby `time`)与 `events`(segmentby `event_type`、orderby `time`),均 1 维 `time`;压缩已配置(`timescaledb_information.compression_settings` 可见对应行;2.29.x 该视图无 `compression_enabled` 列)。
|
||||
- 2026-08-29:**timescale reader 图表对象**(配套 hass 的 `timescale_database_reader` 集成 + `timescale-plotly-card`,上游 SQL `remmob/timescale_database_reader` `SQL/scribe/01+02` @ `bb8776a`,以 postgres 执行):`sensor_minute_aggregate` 连续聚合(1 分钟桶,last(state)/last(value),实时聚合开启)+ `sensor_minute_aggregate_entity` 视图(join `entities`)+ `sensor_minute` hypertable(`minute`/`entity_id`/`state`/`value`,LOCF 前向填充)。任务:1005 `sensor_minute` 压缩(7 天)、1006 `sensor_minute` 保留(10 年)、1007 `every_minute_refresh` 每分钟增量刷新(含 5 分钟回溯窗口修正)。授权:`GRANT SELECT ON sensor_minute_aggregate, sensor_minute_aggregate_entity, sensor_minute, entities TO hass`。种子 19529 行(331 实体,自首个数据点起)。**刻意跳过**了上游脚本对 `states_raw` 的 3 个月保留 + 压缩策略语句——与"`states_raw` 永久归档"定位冲突,如需磁盘回收属用户决策(scribe 自己的压缩任务 1000/1001 未动)。
|
||||
- 2026-08-29:**`sensor_minute_refresh` 本地补丁(类比 tianqi 补丁,重跑上游 02 SQL 后需重打)**:值 CASE 的 `ELSE 0` → `ELSE NULL`。原因:scribe 对 unavailable 分钟 value 为 NULL,上游刷新过程兜底写 0;对差分模式的用电图,0→计数器回升会把插座的**生命周期累计值**(最高 1588 kWh)算进掉线那一小时。同日一次性清理既有脏 0:头部占位行 DELETE 505 行(各实体首次非零分钟之前的 value=0);`sensor.%_energy` 与温湿度实体的 value=0 → NULL(10+16 行,物理上不可能的真 0,图表渲染为断点)。功率实体的中途 0 是真实待机读数,保留。
|
||||
- `hass` 库的 recorder 表仍为普通表(无 hypertable);`scribe` 集成负责时间序列历史(`states_raw` + `events` hypertable)。
|
||||
|
||||
## Verification history
|
||||
|
||||
- 2026-08-31:**13:58 重启根因确认,非停电**(W1N-263):pve2(`192.168.55.25`)任务日志显示 08-30 **13:58:00 `root@pam` 在 PVE Web UI 修改 VM 100 配置**(`-delete allow-ksm,balloon,shares -memory 3072`),**13:58:06 点 Reboot**(`qmreboot` → 客机 13:58:08 干净 ACPI 关机 → 13:58:13 自动重启)。宿主机全程在线(08-30 09:00 开机至今连续运行 1d12h+),`.66.26` PVE 及各 VM 均无重启——排除停电。HA recorder 在窗口(13:58:46–47)报 2 次 `Connection refused`,DB 恢复后自动重连,**无数据丢失**(`hass.states`/`scribe.states_raw` 13:55–14:02 逐分钟无缺口,recorder 内存队列吸收回写)。13:58:47 三容器已起,13:58:56 自愈单元 `pgdb-compose.service` 执行成功——本次自愈按设计工作。同日下午 12:54–12:55 另有一次**客机内自重启**(无 PVE 任务,工作站 SSH 会话相邻)。08-29 22:19→08-30 09:00 宿主机停机 10h41m 为**干净关机**(systemd 有序关闭,非停电)。
|
||||
- 2026-08-30:**开机竞态故障 + 修复**(W1N-260):09:01 开机后 docker 恢复容器时绑定 `192.168.55.15:5432/8081` 失败(EADDRNOTAVAIL)→ timescaledb/pgweb 停摆至 12:16,pg-backup 开机备份失败(解析不到 timescaledb)→ unhealthy。12:22 `docker compose up -d --force-recreate` 修复(三容器回 `database_default`、端口发布、今日备份、pgweb 恢复);用户重启 HA Core 后写入管道恢复。12:43 新增开机自愈 unit `pgdb-compose.service`(enabled,已实测幂等 reconcile)。pgdb-health 8 项全绿。
|
||||
- 2026-08-29:首次检查(只读)+ 修复 scribe 权限 + 安装夜间备份。见 Linear vps 项目登记。
|
||||
- 2026-08-29:**compose 改造完成**(W1N-227,用户已验收):裸 `docker run` → `/opt/database/docker-compose.yml` 三服务(timescaledb + pgweb + pg-backup);superuser 换强密码;端口收紧 IPv4;备份容器化(TZ=Asia/Shanghai,cron 02:00 本地);`pg_restore` 还原实测通过;pgweb UI 用户确认可查 hass/scribe 数据。源码在仓库 `compose/pgdb/`。
|
||||
- 2026-08-29:**运维 runbook 落地**(W1N-228,已验收):新增 `runbooks/pgdb-health.md`(只读,8 项诊断全绿)、`pgdb-restore.md`(流程式,temp-DB 安全还原 + 审批门)、`pgdb-update.md`(门控命令式,回滚=/opt/database/run + 旧卷);README 索引与 validate-repo.sh 分类同步更新;runbook 命令已对活主机逐条实测(含 `pg_restore -l` 校验当日 dump)。同日修正:SSH key auth 可用(facts 原记"密钥未安装"已过时);scribe 新增 `events` hypertable。
|
||||
- 2026-08-29:**CSG 长期归档 + recorder 365d**(W1N-243):建 `csg_history` 表 + attributes 回填(逐日 59 + 逐月 19)+ 每日任务 1008(函数 v2:最新有值行读取、日费用阶梯回退);hass `purge_keep_days` 30→365(备份 `configuration.yaml.bak-20260829-purge365`)。见 Linear vps W1N-243。
|
||||
@@ -24,6 +24,7 @@ ssh -4 windy@synapse.chans.xyz
|
||||
| DB | ESS embedded PostgreSQL 17 (PVC 20Gi, local-path) |
|
||||
| Cache | ESS embedded Redis (PVC 2Gi) |
|
||||
| Chart | `oci://ghcr.io/element-hq/ess-helm/matrix-stack`, version `26.7.2` |
|
||||
| Plane | Helm `plane-ce-1.8.0` (app `v1.4.1`), namespace `plane` — self-hosted Plane project management |
|
||||
|
||||
### Matrix service endpoints
|
||||
|
||||
@@ -51,8 +52,50 @@ All other ports internal only (no K3s API, no database, no Redis exposed).
|
||||
|
||||
- `ess` — all ESS workloads (Synapse, MAS, Element, Postgres, Redis, HAProxy)
|
||||
- `matrix-system` — cluster base resources (ResourceQuota, LimitRange, mrtc-placeholder)
|
||||
- `plane` — Plane project management (Helm release `plane-app`)
|
||||
- `cert-manager` — cert-manager
|
||||
|
||||
## Plane (project management)
|
||||
|
||||
Self-hosted [Plane](https://github.com/makeplane/plane) on the same K3s node, deployed via the official `plane-ce` Helm chart.
|
||||
|
||||
| Item | Detail |
|
||||
|------|--------|
|
||||
| Release | `plane-app` (ns `plane`), chart `plane-ce-1.8.0`, app `v1.4.1`, revision 1 |
|
||||
| URL | https://plane.chans.xyz |
|
||||
| Install date | 2026-09-01 |
|
||||
| Values source | `/home/windy/plane-k3s/values.yaml` (plain file, not a git repo) |
|
||||
| Images | `artifacts.plane.so/makeplane/*` (`plane-frontend`, `plane-backend`, `plane-admin`, `plane-live`), pullPolicy `Always` |
|
||||
| Ingress | Traefik `IngressRoute` `plane-app-ingress` — `/`→web, `/api` `/auth`→api, `/spaces`→space, `/god-mode`→admin, `/live`→live, `/uploads`→minio; `maxRequestBodyBytes` 20Mi |
|
||||
| TLS | Own namespace `Issuer` `plane-app-cert-issuer` (HTTP-01, LE prod, `admin@chans.xyz`); cert `plane-app-ssl-cert` (CN `plane.chans.xyz`) |
|
||||
| DB | Bundled Postgres `15.7-alpine` (PVC 5Gi, local-path) |
|
||||
| Cache/queue | Bundled Redis (PVC 100Mi), RabbitMQ `3.13.6-management-alpine` (PVC 100Mi) |
|
||||
| Storage | Bundled MinIO (`minio/minio:latest`, root user `admin`, PVC 5Gi) — S3 for uploads/docs |
|
||||
| Resources | Every workload: cpu 50m/500m, mem 50Mi/1000Mi, replicas 1 |
|
||||
| SMTP | Not configured (no `smtp` values) — Plane invites/password resets won't email yet |
|
||||
|
||||
Workloads (all 1/1 Running): 7 Deployments (`plane-app-{admin,api,beat-worker,live,space,web,worker}-wl`) + 4 StatefulSets (`plane-app-{minio,pgdb,rabbitmq,redis}-wl`); init Jobs `api-migrate-1` / `minio-bucket-1` Completed. All PVCs Bound on `local-path` (root disk).
|
||||
|
||||
### Plane configuration notes
|
||||
|
||||
- **`planeVersion: v1.4.1`** pinned in values.yaml; chart tracks Plane's own tags.
|
||||
- **Secrets**: Helm-generated Opaque secrets (`plane-app-app-secrets`, `-doc-store-secrets`, `-pgdb-secrets`, `-rabbitmq-secrets`, `-live-secrets`); `requireExplicitSecrets: false`. Values live in `$SECRET_KEY`, `DATABASE_URL`, `AMQP_URL`, `REDIS_URL` etc.
|
||||
- **Sentry / CORS**: `sentry_dsn` and `cors_allowed_origins` empty (defaults fine for single-host).
|
||||
- **MinIO is `latest` tag** — pin a version for reproducibility.
|
||||
- **Backup**: NOT covered by `/var/backups/matrix` (which is paused anyway) — Plane Postgres/MinIO PVCs have no backup tier yet.
|
||||
|
||||
### Plane verification
|
||||
|
||||
```bash
|
||||
# Release + workloads
|
||||
sudo helm list -A
|
||||
sudo k3s kubectl -n plane get deploy,sts,pods -o wide
|
||||
# Cert + ingress
|
||||
sudo k3s kubectl -n plane get certificate,ingressroute
|
||||
# Endpoint
|
||||
curl -4 -s -o /dev/null -w '%{http_code}\n' https://plane.chans.xyz/
|
||||
```
|
||||
|
||||
## Local backup
|
||||
|
||||
| Item | Detail |
|
||||
@@ -62,7 +105,7 @@ All other ports internal only (no K3s API, no database, no Redis exposed).
|
||||
| Retention | 7 days |
|
||||
| Disk warning | 80% (healthcheck), 90% (backup stops) |
|
||||
| Content | Planned: PostgreSQL `synapse` + `mas` logical dumps, media store archive, `/etc/matrix-bootstrap` |
|
||||
| Status | **Not operational** — no current Matrix backup or recovery tier |
|
||||
| Status | **Not operational** — no current Matrix backup or recovery tier. **Plane data (its own Postgres + MinIO PVCs in ns `plane`) is also not covered by any backup.** |
|
||||
|
||||
## Health checks
|
||||
|
||||
@@ -101,5 +144,6 @@ diagnosis and imperative recovery work.
|
||||
- MatrixRTC / Element Call / LiveKit / Coturn not deployed (`mrtc.chans.xyz` reserved only)
|
||||
- SMTP email not yet configured (requires manual secret bootstrap followed by a
|
||||
reviewed Ansible stack deployment)
|
||||
- Plane `minio` image uses `latest` tag (pin a version)
|
||||
- No off-site Restic backup
|
||||
- Single-node K3s (no HA for control plane)
|
||||
|
||||
@@ -53,6 +53,11 @@ of `8080`. During adoption or recovery, use the documented `:9080/inform` URL;
|
||||
an AP left on `:8080` can remain reachable by ping and SSH while showing
|
||||
offline in the controller.
|
||||
|
||||
IPv6 is enabled on the controller's `Default` network (`ipv6_enabled: true`,
|
||||
client assignment SLAAC; RA is served by `gw`, so `ipv6_interface_type` is
|
||||
`none`); both managed APs hold global SLAAC addresses — verified 2026-08-20.
|
||||
See [docs/unifi-network.md](../docs/unifi-network.md).
|
||||
|
||||
## Safe reconciliation and verification
|
||||
|
||||
```bash
|
||||
|
||||
+26
-12
@@ -5,12 +5,12 @@
|
||||
| Role | Multi-service VPS (Vaultwarden, Traefik, Soft Serve, …) |
|
||||
| SSH | `ssh -4 windy@us2.wsvc.info` (prefer IPv4 from WSL) |
|
||||
| IPv4 | `193.9.44.165` |
|
||||
| Also DNS | `auth.wsvc.info` → this host; `repo.windy.me` → this host (Soft Serve) |
|
||||
| Also DNS | `auth.wsvc.info` → this host; `repo.windy.me` → this host (Gitea) |
|
||||
| Public HTTPS | Traefik on `:80` / `:443` (`/opt/traefik`) |
|
||||
|
||||
## Vaultwarden (Bitwarden-compatible)
|
||||
|
||||
**Status: operational** (Postgres live, HTTPS 200, healthy containers, SMTP AUTH OK — last probe 2026-08-01 18:55 CST).
|
||||
**Status: operational** (Postgres live, HTTPS 200, healthy containers, SMTP AUTH OK — last probe 2026-08-29).
|
||||
|
||||
Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
|
||||
|
||||
@@ -21,9 +21,9 @@ Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
|
||||
| Env file | `/opt/vaultwarden/.env` |
|
||||
| Admin overrides | `/opt/vaultwarden/vw-data/config.json` (**wins over env**) |
|
||||
| Public URL / `DOMAIN` | `https://auth.wsvc.info` |
|
||||
| Image | `vaultwarden/server:1.37.1` (pinned) |
|
||||
| Image | `vaultwarden/server:1.37.2` (pinned) |
|
||||
| Live DB | **Postgres 16** (`vw-db` / service `pg`) via compose `DATABASE_URL` |
|
||||
| Data (probe) | users=1, ciphers=1327 |
|
||||
| Data (probe) | users=1, ciphers=1360 |
|
||||
| Cold SQLite | `backups/sqlite-cold/db.sqlite3.pre-pg-20260801` (not used live) |
|
||||
| Pre-migrate backup | `backups/pre-pg-migrate-20260801_161204/` |
|
||||
| Data dir | `./vw-data` → `/data` (attachments, rsa keys, `config.json`) |
|
||||
@@ -50,7 +50,7 @@ Upstream docs: [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)
|
||||
|
||||
| Container | Status |
|
||||
|-----------|--------|
|
||||
| `vaultwarden` | Up (healthy), `vaultwarden/server:1.37.1` |
|
||||
| `vaultwarden` | Up (healthy), `vaultwarden/server:1.37.2` |
|
||||
| `vw-db` | Up (healthy) — **live** Postgres |
|
||||
| `vaultwarden-backup` | Up (`pg_dump`) |
|
||||
| `vaultwarden-pgweb` | Exited (profile `debug`) |
|
||||
@@ -80,19 +80,33 @@ ansible-playbook playbooks/compose-reconcile.yml --limit vaultwarden \
|
||||
|
||||
| Container | Status | Image / notes |
|
||||
|-----------|--------|---------------|
|
||||
| `soft-serve` | Up | `ghcr.io/charmbracelet/soft-serve:latest` (`repo.windy.me:2222`) |
|
||||
| `gitea` | Up | `gitea/gitea@sha256:1c17ecaead42e…` (1.27.3-rootless) — SSH `repo.windy.me:2222`, web `https://repo.windy.me` |
|
||||
| `gitea-backup` | Up | alpine + sqlite3/rsync sidecar (daily backup 02:00 / prune 03:00, crond) |
|
||||
|
||||
### Gitea (replaced Soft Serve 2026-09-18; [Plane VPS-94](https://plane.chans.xyz))
|
||||
|
||||
- `/opt/gitea/compose.yml` (+ `Dockerfile.backup`, `scripts/`, `config/app.ini`, `data/`, `secrets/`, `backups/`); 镜像: `compose/gitea/`(参考, 服务器文件为准)
|
||||
- **rootless 镜像** uid 1000:1000; SQLite `/opt/gitea/data/data/gitea.db`; repos `/opt/gitea/data/data/git/repositories/`; app.ini `/opt/gitea/config/app.ini`(600, 含 SECRET_KEY)
|
||||
- SSH: 内置 server 容器内 `:2322`(`SSH_LISTEN_PORT` 非特权), Traefik TCP entrypoint `ssh`(`:2222` → `gitea:2322`, `HostSNI(*)`, `tls=false`) on `vw-net`; clone URL `ssh://git@repo.windy.me:2222/windy/<repo>.git`(owner 段 `windy`)
|
||||
- **host key 复用 soft-serve**(`SSH_SERVER_HOST_KEYS=/secrets/soft_serve_host_ed25519`, ed25519, 指纹 `SHA256:PdxZRe74…`): 客户端 known_hosts 零变更; 仅公钥认证(密码认证未启用)
|
||||
- Web: `https://repo.windy.me`(Traefik websecure + letsencrypt); `DISABLE_REGISTRATION=true`, Actions 关闭; 管理员 `windy`(凭据仅存服务器 `/opt/gitea/.admin-credentials`, 勿入库/入 Plane)
|
||||
- 仓库: 16 个(顶层 11 + `cdia/` 4 + `windyboy/go-caatsm`), 2026-09-18 自 soft-serve `push --mirror` 迁移, 逐仓 `ls-remote` ref 全集 + HEAD symref 两端一致; 可见性仅 `dotfiles-personal` private, 其余 public(与 soft-serve 现状一致)
|
||||
- 备份: sidecar 每日 02:00 → `backups/gitea_<TS>/{app.ini.tar.gz, gitea.db, repos.tar.gz}`(app.ini 含恢复必需 SECRET_KEY), 03:00 prune 保留 14 份; 已验证手动备份产物 109.9M
|
||||
- 回滚: `/opt/soft-serve` 未删(compose stop + sidecar 停, 数据与旧备份冻结保留), 回滚 = Traefik `:2222` 指回 `soft-serve:23231` + 客户端 remote 回改旧无 owner 段路径; 观察 2–4 周后清理(历史: W1N-244~248)
|
||||
| `traefik` | Up | `traefik:v3.6.2` (`/opt/traefik`, public `:80`/`:443`) |
|
||||
| `nghttpx-proxy` + `squid-backend` | Up | HTTP forward-proxy stack (`/opt/nghttpx`), network `nghttpx_internal-net`; details TBD |
|
||||
|
||||
Directories for `authelia`, `conduit`, `dendrite`, `mastodon`, `rustdesk`, `zitadel`, etc. exist under `/opt` but have no running containers; treat them as dormant, not documented services.
|
||||
**Disk cleanup 2026-09-18** ([Plane vps VPS-93](https://plane.chans.xyz)): root 71% → **23%** (~33G freed) keeping soft-serve / vaultwarden / traefik (nghttpx kept running per operator choice). Removed: unused Docker images + orphan volumes (incl. `zitadel_data` 801M), dormant `/opt` dirs (dendrite + its disabled `dendrite.service` unit, mastodon, dailysync, keycloak, media-repo, authelia, conduit, npm, manager, fusion, zitadel, rustdesk), rootless podman storage (6.4G stale goauthentik), home dev caches, apt cache, journal 3.8G→162M (+`SystemMaxUse=200M` drop-in, active next boot), truncated container logs (nghttpx 550M / traefik / squid). Follow-up: nghttpx-proxy logs grow ~25M/day (INFO per-connection); root-cause log-level/rotation fix still open (needs container restart approval).
|
||||
|
||||
Remaining running services on this host: `gitea`, `vaultwarden` stack, `traefik`, `nghttpx-proxy` + `squid-backend` (undocumented forward proxy, `/opt/nghttpx`). `/opt/soft-serve` kept stopped as rollback (2–4 weeks, data intact). `/home/windy/authelia` (76M) left in place — outside approved cleanup scope.
|
||||
|
||||
## Verified
|
||||
|
||||
Last checked: **2026-08-01 18:55 CST** — operational.
|
||||
Last checked: **2026-09-18** — operational; disk cleanup done (see note above, Plane vps VPS-93). Prior full probe: 2026-08-29.
|
||||
|
||||
- `vaultwarden` + `vw-db` healthy; `DATABASE_URL` → `pg:5432/vaultwarden`
|
||||
- `https://auth.wsvc.info/` **200**, `/admin` **200**, `/api/config` OK (`disableUserRegistration: true`)
|
||||
- Identity wrong-password → **400** business error (DB readable, not 500)
|
||||
- SMTP: container → `mx2:587` OK; STARTTLS cert CN=`mx2.windy.me`; **AUTH OK** with effective `config.json` password (synced with `.env` / `.smtp-credentials`)
|
||||
- LE cert CN=`auth.wsvc.info`
|
||||
- PG counts: users=1, ciphers=1327
|
||||
- SMTP: container → `mx2:587` OK; **AUTH OK** with effective `config.json` password (synced with `.env` / `.smtp-credentials`, fingerprint match)
|
||||
- PG counts: users=1, ciphers=1360
|
||||
- Image `vaultwarden/server:1.37.2` (**upgraded 2026-08-29** from 1.37.1; required for Bitwarden clients 2026.8.0+); post-upgrade 404 fixed by Traefik restart, then 200
|
||||
- vps-health local check **installed 2026-08-29** (`vps-healthcheck.timer` daily 06:15 + `/usr/local/lib/vps-health/run`); `health-report.yml --limit vaultwarden` now passes (**ok**, was failing due to missing check infra + script bugs fixed: trim_blocks render, pgweb debug-profile false positive, SMTP probe moved host-side since image lacks python3)
|
||||
|
||||
+133
-1
@@ -9,10 +9,74 @@
|
||||
| Compose file | `/opt/wireguard/compose.yml` |
|
||||
| Container | `wireguard` |
|
||||
| Image policy | Immutable digest, updated only in an approved maintenance window |
|
||||
| Public port | UDP `51820` on IPv4 and IPv6 |
|
||||
| Public endpoint | `us4.wsvc.info:51820/udp`; DNS publishes only A `185.201.226.122` (no native AAAA) |
|
||||
| Tunnel subnet | `10.13.13.0/24` |
|
||||
| Routing policy | IPv4-only full tunnel (`ALLOWEDIPS=0.0.0.0/0`); IPv6 traffic is not guaranteed to use the VPN |
|
||||
|
||||
Upstream image documentation:
|
||||
[LinuxServer.io WireGuard](https://docs.linuxserver.io/images/docker-wireguard/).
|
||||
|
||||
## Deployment configuration
|
||||
|
||||
The repository-owned, non-secret Compose declaration is rendered from
|
||||
`ansible/templates/wireguard-compose.yml.j2`. The live declaration was verified
|
||||
on 2026-08-12 with these core settings:
|
||||
|
||||
| Setting | Live value / intent |
|
||||
|---------|---------------------|
|
||||
| Image | `lscr.io/linuxserver/wireguard@sha256:ac43e1226878d2611315172d6ea357a95cb326ee73124b91108118efc8666889` |
|
||||
| Image version | `1.0.20260223-r0-ls119` (build 2026-07-30) |
|
||||
| Required capability | `NET_ADMIN` only; host kernel already supplies WireGuard/iptables, so `SYS_MODULE` and `/lib/modules` are not granted |
|
||||
| Filesystem | Read-only container root; executable tmpfs at `/run`; writable bind mount `/opt/wireguard/config:/config` |
|
||||
| Restart | `unless-stopped` |
|
||||
| Server mode | Named peers `ha`, `phone`, `mbp`; runtime and configured peer counts both `3` |
|
||||
| Client DNS | `1.1.1.1` |
|
||||
| Tunnel routing | IPv4 full tunnel, `0.0.0.0/0`; no client IPv6 tunnel |
|
||||
| Runtime interface | `wg0`, server address `10.13.13.1/32`, listen port `51820` |
|
||||
| Forwarding/NAT | IPv4 forwarding enabled in the container namespace; `wg0` forwarding allowed and egress masqueraded on `eth+`; IPv6 forwarding disabled |
|
||||
|
||||
Docker binds UDP `51820` on both host socket families, but the public hostname
|
||||
has no AAAA record. Clients using `us4.wsvc.info` therefore reach the server over
|
||||
IPv4.
|
||||
|
||||
## Other host services and firewall (2026-08-12)
|
||||
|
||||
This host also carries the `windy.me` secondary MX and several web applications;
|
||||
do not build its firewall allowlist from the WireGuard role alone.
|
||||
|
||||
| Port | Owner / purpose | Effective public state |
|
||||
|------|-----------------|------------------------|
|
||||
| TCP `22` | SSH management | Open |
|
||||
| TCP `25` | Postfix, `mx.windy.me` (MX priority 30) | Open; retain until the secondary-MX role is explicitly retired |
|
||||
| TCP `80`, `443` | Traefik for `update.wsvc.info`, `us4-gate.wsvc.info`, and `trlm.wsvc.info` | Open |
|
||||
| TCP `3000` | Semaphore UI direct Docker publish | Open; redundant with the Traefik route and should be removed or bound to loopback |
|
||||
| TCP `8080` | Traefik direct Docker publish | Open; redundant with the authenticated dashboard route and should be removed or bound to loopback |
|
||||
| UDP `51820` | WireGuard | Required public endpoint |
|
||||
| TCP `9443` | Host nghttpx-to-Squid proxy | Listening but blocked by the current firewall |
|
||||
| UDP `123` | ntpsec | Listening but blocked by the current firewall |
|
||||
|
||||
PostgreSQL (`5433`/`5434`/`5435`), MariaDB (`3306`), and the host Squid TCP
|
||||
listener (`3128`) are loopback-only. Squid also owns wildcard UDP sockets, which
|
||||
are not allowed by the current public zone.
|
||||
|
||||
UFW is not installed. Firewalld `2.3.1` is active with nftables. On 2026-08-12,
|
||||
the reviewed `ansible/playbooks/us4-firewalld.yml` reconciliation removed the
|
||||
stale `imap`, `imaps`, `smtp-submission`, and `smtps` services plus TCP `24`,
|
||||
`6443`, and `8443` without reloading or restarting firewalld. Runtime and
|
||||
permanent public-zone state now match exactly: services `dhcpv6-client`, `http`,
|
||||
`https`, `smtp`, and `ssh`, with no explicit ports.
|
||||
|
||||
Docker-published ports are accepted through Docker's DNAT/FORWARD chains, so
|
||||
the public-zone cleanup does not close `3000` or `8080`. Their Compose bindings
|
||||
remain a separate, staged follow-up after the required observation window.
|
||||
Firewalld logged Docker chain/policy conflicts during the 2026-08-10 boots;
|
||||
treat any firewall reload or service restart as a maintenance-window operation
|
||||
and reverify Docker routing. Tracking: Linear `W1N-60`.
|
||||
|
||||
`mx.windy.me` also publishes AAAA `2602:f9f3:0:2::878`, while the host currently
|
||||
has no global IPv6 address or IPv6 default route. Treat that as a separate
|
||||
secondary-MX reachability issue.
|
||||
|
||||
## Safety
|
||||
|
||||
- Private keys, preshared keys, peer configuration files, and QR codes remain
|
||||
@@ -21,6 +85,17 @@
|
||||
- Local rollback archives are stored in `/opt/wireguard/backups` (directory
|
||||
mode `0700`, archives mode `0600`). They contain private keys, are not an
|
||||
off-host disaster-recovery backup, and must never leave the server.
|
||||
- Live private keys, preshared keys, generated peer configs, QR images, and
|
||||
`wg0.conf` are mode `0600`. Template-only `peer.conf` and `server.conf` files
|
||||
are mode `0644` and do not contain generated key material.
|
||||
- `/opt/wireguard/config` is mode `0755`, but its sensitive files are `0600`.
|
||||
The current files are owned by the image's numeric UID/GID rather than the
|
||||
declared `PUID=1000` / `PGID=1000`; the root-run WireGuard processes can use
|
||||
them, but reconcile ownership only after a protected backup and maintenance
|
||||
review.
|
||||
- `LOG_CONFS` is currently unset and the inspected container log contained no
|
||||
QR-code/config banners. Do not enable config logging; generated QR images are
|
||||
credentials.
|
||||
- Do not delete, move, or regenerate `/opt/wireguard/config` during
|
||||
maintenance.
|
||||
- Before a container recreation, validate `docker compose config` and retain a
|
||||
@@ -35,6 +110,23 @@ cd ansible
|
||||
ansible-playbook playbooks/health-report.yml --limit wireguard
|
||||
```
|
||||
|
||||
Preview the narrow, fail-closed public-zone reconciliation:
|
||||
|
||||
```bash
|
||||
ansible-galaxy collection install -r requirements.yml
|
||||
ansible-playbook playbooks/us4-firewalld.yml --limit us4 --check --diff
|
||||
```
|
||||
|
||||
Apply it only after testing the provider console and keeping an independent SSH
|
||||
rollback session open. The playbook creates a protected server-local backup and
|
||||
a 15-minute automatic rollback before changing rules; it cancels that rollback
|
||||
only after SSH, HTTPS, SMTP, Docker, Fail2ban, and WireGuard checks pass:
|
||||
|
||||
```bash
|
||||
ansible-playbook playbooks/us4-firewalld.yml --limit us4 \
|
||||
-e '{"us4_firewalld_confirm": true, "us4_console_confirm": true}'
|
||||
```
|
||||
|
||||
The image update and recreate procedure is deliberately separate and requires
|
||||
an immutable image digest in the server-side Compose file plus an explicit
|
||||
maintenance-window confirmation:
|
||||
@@ -59,3 +151,43 @@ ansible-playbook playbooks/wireguard-harden.yml --limit wireguard \
|
||||
- Validate a known client can handshake and sends IPv4 traffic through the VPN.
|
||||
- Do not treat inactive mobile peers as a failure solely because their latest
|
||||
handshake is old.
|
||||
|
||||
## Live audit snapshot (2026-08-12)
|
||||
|
||||
The WireGuard service itself is healthy and its installation is broadly
|
||||
reasonable:
|
||||
|
||||
- The sanitized Ansible health report returned `status=ok`; Compose is valid,
|
||||
the container is running with zero restarts, `wg0` exists, and UDP `51820` is
|
||||
listening.
|
||||
- One of three peers had a current handshake during the audit. Two peers had
|
||||
not handshaken since the current container/interface start; confirm those
|
||||
clients only if they are expected to be active.
|
||||
- The image is immutable-digest pinned, key-bearing files are protected, the
|
||||
container root is read-only, and the container has `NET_ADMIN` without the
|
||||
broader `SYS_MODULE` capability.
|
||||
- Debian `13.6`, kernel `6.12.101+deb13-amd64`, Docker Engine `29.7.2`, and
|
||||
Docker Compose `v5.4.0` were observed. No Debian package updates or reboot
|
||||
requirement were pending.
|
||||
|
||||
Open host-level follow-up (do not conflate these with a WireGuard outage):
|
||||
|
||||
1. **Disk capacity:** `/` was 90% used with about 3.4 GiB free. Docker reported
|
||||
about 2.48 GB of reclaimable images and the system journal used about 1.9
|
||||
GB, but do not prune or vacuum without reviewing retention and rollback
|
||||
needs first.
|
||||
2. **Docker exposure:** the firewalld public-zone cleanup is complete, but
|
||||
Docker still publishes `3000` and `8080` outside the ordinary host INPUT
|
||||
path. Remove those redundant Compose bindings in separate maintenance units
|
||||
after the observation window, and confirm provider firewall rules first.
|
||||
3. **Image maintenance:** the upstream `latest` amd64 image had advanced to
|
||||
`1.0.20260223-r0-ls120` (build 2026-08-06). Review and pin its immutable
|
||||
digest in a maintenance window rather than updating unattended.
|
||||
4. **Host hygiene:** `apache2.service`, `certbot.service`, and
|
||||
`postgresql@9.6-main.service` were in a failed state while unrelated Docker
|
||||
workloads remained active. Establish ownership and remove or repair stale
|
||||
units separately.
|
||||
5. **Resource/log limits:** the WireGuard container has no memory, CPU, or PID
|
||||
limit and uses Docker's `json-file` log driver without a per-container
|
||||
rotation setting. Current log size was small, but limits/rotation should be
|
||||
considered during a reviewed Compose update.
|
||||
|
||||
+30
-19
@@ -8,28 +8,38 @@ diagnosis and procedures that are deliberately interactive or destructive; see
|
||||
For a live-verified map of the **internal LAN** (gw, gfw, dns, ubnt, APs) and
|
||||
the software deployed there, see [the LAN overview](../docs/lan-overview.md).
|
||||
|
||||
| Host | Role | SSH | IPv4 | Status | Facts |
|
||||
|------|------|-----|------|--------|-------|
|
||||
| mx2.windy.me | mailcow (primary MX prio 20) | `ssh -4 windy@mx2.windy.me` | 194.163.160.244 | active | [hosts/mx2.windy.me.md](../hosts/mx2.windy.me.md) |
|
||||
| us2.wsvc.info | Vaultwarden/Postgres (+ Traefik, Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | 193.9.44.165 | active | [hosts/us2.wsvc.info.md](../hosts/us2.wsvc.info.md) |
|
||||
| mx.windy.me | mail (secondary MX prio 30) | TBD | see AAAA/A | stub | — |
|
||||
| repo.windy.me | Soft Serve git (on us2) | `ssh -p 2222 windy@repo.windy.me` | 193.9.44.165 | stub | see us2 |
|
||||
| auth.wsvc.info | Vaultwarden public hostname | — (HTTPS) | → us2 | active | see us2 |
|
||||
| us1.wsvc.info | PowerDNS secondary (ns2 host) | TBD | 202.91.35.141 | stub | Auth 5.0.5; see hk2 |
|
||||
| us4.wsvc.info | WireGuard VPN | `ssh -4 windy@us4.wsvc.info` | 185.201.226.122 | active | [hosts/us4.wsvc.info.md](../hosts/us4.wsvc.info.md) |
|
||||
| hk2.chans.xyz | PowerDNS auth (ns1) | `ssh -4 windy@hk2.chans.xyz` | 154.36.174.161 | active | [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) |
|
||||
| ns1.wsvc.info | PowerDNS public NS name | — (DNS) | → hk2 `154.36.174.161` | active | see hk2 |
|
||||
| ns2.wsvc.info | Secondary NS (AXFR/NOTIFY peer) | — (DNS) | → us1 `202.91.35.141` | active | see hk2 |
|
||||
| pdns.wsvc.info | Poweradmin UI | — (HTTPS) | → hk2 | active | see hk2 |
|
||||
| pgweb.wsvc.info | PowerDNS Postgres UI | — (HTTPS) | → hk2 | active | see hk2 |
|
||||
| **synapse.chans.xyz** | Matrix homeserver (ESS: Synapse + MAS + Element) | `ssh -4 windy@synapse.chans.xyz` | `169.58.86.13` | **active** | [hosts/synapse.chans.xyz.md](../hosts/synapse.chans.xyz.md) |
|
||||
| **gfw.windy.lan** | OpenWrt (ImmortalWrt) LAN gateway / OpenClash (PVE VM 140) | `ssh -4 root@192.168.66.1` | `192.168.66.1` | **active** | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
|
||||
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy (PVE VM 120) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | **active** | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
|
||||
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | **active** | [hosts/gw.md](../hosts/gw.md) |
|
||||
| **ubnt** | UniFi Network Controller (PVE VM 160) | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | **active** | [hosts/ubnt.md](../hosts/ubnt.md) |
|
||||
**Ansible 列**:`✓` = 该主机在 [`ansible/inventory/hosts.yml`](../ansible/inventory/hosts.yml)
|
||||
(执行真相),用其 inventory key(见括号注)跑 playbook;`—` = 不由 Ansible 管理,
|
||||
原因是该平台无 ansible 覆盖或仅是公网别名/服务端点。
|
||||
|
||||
| Host | Role | SSH | IPv4 | Ansible | Status | Facts |
|
||||
|------|------|-----|------|---------|--------|-------|
|
||||
| mx2.windy.me | mailcow (primary MX prio 20) | `ssh -4 windy@mx2.windy.me` | 194.163.160.244 | ✓ (mx2) | active | [hosts/mx2.windy.me.md](../hosts/mx2.windy.me.md) |
|
||||
| us2.wsvc.info | Vaultwarden/Postgres (+ Traefik, Soft Serve, …) | `ssh -4 windy@us2.wsvc.info` | 193.9.44.165 | ✓ (us2) | active | [hosts/us2.wsvc.info.md](../hosts/us2.wsvc.info.md) |
|
||||
| mx.windy.me | mail (secondary MX prio 30) | TBD | see AAAA/A | — (stub) | stub | — |
|
||||
| repo.windy.me | Gitea git (on us2) | `ssh -p 2222 windy@repo.windy.me` | 193.9.44.165 | — (service on us2) | stub | see us2 |
|
||||
| auth.wsvc.info | Vaultwarden public hostname | — (HTTPS) | → us2 | — (alias) | active | see us2 |
|
||||
| us1.wsvc.info | PowerDNS secondary (ns2 host) | TBD | 202.91.35.141 | — (stub) | stub | Auth 5.0.5; see hk2 |
|
||||
| us4.wsvc.info | WireGuard VPN | `ssh -4 windy@us4.wsvc.info` | 185.201.226.122 | ✓ (us4) | active | [hosts/us4.wsvc.info.md](../hosts/us4.wsvc.info.md) |
|
||||
| hk2.chans.xyz | PowerDNS auth (ns1) | `ssh -4 windy@hk2.chans.xyz` | 154.36.174.161 | ✓ (hk2) | active | [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) |
|
||||
| ns1.wsvc.info | PowerDNS public NS name | — (DNS) | → hk2 `154.36.174.161` | — (alias) | active | see hk2 |
|
||||
| ns2.wsvc.info | Secondary NS (AXFR/NOTIFY peer) | — (DNS) | → us1 `202.91.35.141` | — (alias) | active | see hk2 |
|
||||
| pdns.wsvc.info | Poweradmin UI | — (HTTPS) | → hk2 | — (alias) | active | see hk2 |
|
||||
| pgweb.wsvc.info | PowerDNS Postgres UI | — (HTTPS) | → hk2 | — (alias) | active | see hk2 |
|
||||
| **synapse.chans.xyz** | Matrix homeserver (ESS: Synapse + MAS + Element) | `ssh -4 windy@synapse.chans.xyz` | `169.58.86.13` | ✓ (matrix_vps) | **active** | [hosts/synapse.chans.xyz.md](../hosts/synapse.chans.xyz.md) |
|
||||
| **gfw.windy.lan** | OpenWrt (ImmortalWrt) LAN gateway / OpenClash (PVE VM 140) | `ssh -4 root@192.168.66.1` | `192.168.66.1` | — (OpenWrt, no ansible) | **active** | [hosts/gfw.windy.lan.md](../hosts/gfw.windy.lan.md) |
|
||||
| **dns.windy.lan** | AdGuard Home LAN DNS + Mihomo explicit proxy (PVE VM 120) | `ssh -4 windy@192.168.66.36` | `192.168.66.36` | ✓ (dns_windy_lan) | **active** | [hosts/dns.windy.lan.md](../hosts/dns.windy.lan.md) |
|
||||
| **gw** | EdgeRouter X primary LAN gateway | `ssh -4 zhiqiang@192.168.66.254` | `192.168.66.254` | — (EdgeOS, no ansible) | **active** | [hosts/gw.md](../hosts/gw.md) |
|
||||
| **ubnt** | UniFi Network Controller (PVE VM 160) | `ssh -4 windy@192.168.66.46` | `192.168.66.46` | ✓ (ubnt) | **active** | [hosts/ubnt.md](../hosts/ubnt.md) |
|
||||
| **hass.windy.lan** | Home Assistant (HAOS, x88 Pro physical box, LAN55) | `ssh hassio@hass.windy.lan` | `192.168.55.11` | — (HAOS, no ansible) | **active** | [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) |
|
||||
| **pgdb** | TimescaleDB PG18 (Docker) — HA recorder backend (PVE VM, LAN55) | `ssh -4 windy@192.168.55.15` | `192.168.55.15` | — (no ansible) | **active** | [hosts/pgdb.md](../hosts/pgdb.md) |
|
||||
|
||||
`status: stub` = known to exist; fill `hosts/<name>.md` when next touched.
|
||||
|
||||
**命名映射**:ansible inventory key ↔ 本表主机名 —— `matrix_vps` ↔ `synapse.chans.xyz`、
|
||||
`dns_windy_lan` ↔ `dns.windy.lan`。inventory key 不随主机名改(防止破坏 `--limit` 用法),
|
||||
通过 inventory 内的 `display_name` 变量与文档交叉引用。
|
||||
|
||||
### Matrix services (synapse.chans.xyz)
|
||||
|
||||
| URL | Service | Notes |
|
||||
@@ -38,4 +48,5 @@ the software deployed there, see [the LAN overview](../docs/lan-overview.md).
|
||||
| https://synapse.chans.xyz | Synapse API | Client-Server + Federation API |
|
||||
| https://account.chans.xyz | MAS | Matrix Authentication Service (local passwords) |
|
||||
| https://admin.chans.xyz | Element Admin | Admin console (MAS admin auth) |
|
||||
| https://plane.chans.xyz | Plane | Project management (Helm `plane-ce` v1.4.1, ns `plane`) |
|
||||
| `mrtc.chans.xyz` | MatrixRTC | **Reserved** – not deployed |
|
||||
|
||||
@@ -0,0 +1,42 @@
|
||||
# Runbook index
|
||||
|
||||
Entry point for all runbooks. Before operational work, read the repo entry
|
||||
[`AGENTS.md`](../AGENTS.md) and the spec [`RUNBOOKS.md`](../RUNBOOKS.md). New
|
||||
runbooks start from [`_template.md`](_template.md).
|
||||
|
||||
## Route by intent
|
||||
|
||||
| Intent | Runbook | Type |
|
||||
|---|---|---|
|
||||
| mailcow health check | [mailcow-health.md](mailcow-health.md) | read-only |
|
||||
| mailcow update | [mailcow-update.md](mailcow-update.md) | change (gated) |
|
||||
| mailcow SMTP/IMAP client | [mailcow-smtp-client.md](mailcow-smtp-client.md) | reference |
|
||||
| Vaultwarden health check | [vaultwarden-health.md](vaultwarden-health.md) | read-only |
|
||||
| Vaultwarden SQLite→PG migrate | [vaultwarden-sqlite-to-postgres.md](vaultwarden-sqlite-to-postgres.md) | change (destructive) |
|
||||
| PowerDNS health check | [pdns-health.md](pdns-health.md) | read-only |
|
||||
| RustDesk health check | [rustdesk-health.md](rustdesk-health.md) | read-only |
|
||||
| Matrix health check | [matrix-health.md](matrix-health.md) | read-only |
|
||||
| Plane health check | [plane-health.md](plane-health.md) | read-only |
|
||||
| pgdb health check | [pgdb-health.md](pgdb-health.md) | read-only |
|
||||
| pgdb DB restore (pg_restore) | [pgdb-restore.md](pgdb-restore.md) | change (procedure) |
|
||||
| pgdb image/compose update | [pgdb-update.md](pgdb-update.md) | change (gated) |
|
||||
| AdGuard Home health check | [adguard-home-health.md](adguard-home-health.md) | read-only |
|
||||
| Host disk cleanup (logs/apt/docker) | [host-disk-cleanup.md](host-disk-cleanup.md) | change (gated) |
|
||||
| Matter packet capture | [matter-packet-capture.md](matter-packet-capture.md) | read-only |
|
||||
| Home Assistant maintenance | [home-assistant-maintenance.md](home-assistant-maintenance.md) | change (gated) |
|
||||
| matrix_e2ee integration update | [matrix-e2ee-update.md](matrix-e2ee-update.md) | change (gated) |
|
||||
| Routine Ansible operations | [ansible-operations.md](ansible-operations.md) | change (allowlisted) |
|
||||
| Linear issue → mergeable change | [issue-to-merge.md](issue-to-merge.md) | delivery |
|
||||
| Failing health/playbook run | [fix-ci.md](fix-ci.md) | change |
|
||||
| Release a reviewed change to production | [release.md](release.md) | change (gated) |
|
||||
| Roll back a change | [rollback.md](rollback.md) | change (gated) |
|
||||
| Controlled network configuration | [network-change.md](network-change.md) | change (gated) |
|
||||
| Network outage / service recovery | [network-recovery.md](network-recovery.md) | recovery |
|
||||
|
||||
## Notes
|
||||
|
||||
- `fix-ci.md`, `release.md`, `rollback.md`, `network-change.md`, `network-recovery.md`
|
||||
are adapted from the upstream guide to this repo's VPS-ops context (execution
|
||||
layer is Ansible + SSH + Linear, not a software CI/CD pipeline).
|
||||
- Health runbooks are read-only; they stop (`STOP`) when live state conflicts
|
||||
with the expected state instead of mutating production.
|
||||
@@ -0,0 +1,119 @@
|
||||
# Runbook: <名称>
|
||||
|
||||
## Purpose
|
||||
|
||||
<说明本 Runbook 要解决的问题及成功结果,1–2 行。>
|
||||
|
||||
## Scope
|
||||
|
||||
- 适用环境:<production / staging / LAN …>
|
||||
- 适用对象:<服务、主机、组件或告警类型>
|
||||
- 不适用情形:<需要改用其他 runbook 或转人工的场景>
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner:<团队或角色>
|
||||
- Last reviewed:<YYYY-MM-DD>
|
||||
- Related systems:<主机名 / 服务名>
|
||||
|
||||
## Preconditions
|
||||
|
||||
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
|
||||
|
||||
## Inputs
|
||||
|
||||
| 输入 | 来源 | 是否必需 | 校验方法 |
|
||||
|---|---|---:|---|
|
||||
| <参数> | <来源> | 是/否 | <如何确认有效> |
|
||||
|
||||
## Safety
|
||||
|
||||
### Non-negotiable rules
|
||||
|
||||
- 先只读诊断,后执行变更。
|
||||
- 不得把删除现有配置作为首次恢复动作。
|
||||
- 不得猜测或编造缺失参数。
|
||||
- 不得绕过失败的测试、检查或审批。
|
||||
- 每次变更后必须完成对应验证。
|
||||
- 破坏性操作必须获得明确批准。
|
||||
|
||||
### Stop conditions
|
||||
|
||||
- 实际状态与本文档的前提或预期结果冲突。
|
||||
- 缺少必要输入、权限、审批或回滚能力。
|
||||
- 验证失败且本文档没有明确的下一步。
|
||||
- 影响范围超出 Scope。
|
||||
|
||||
### Approval gates
|
||||
|
||||
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|
||||
|---|---|---:|---|
|
||||
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Diagnose
|
||||
|
||||
**Action**
|
||||
|
||||
<执行只读诊断动作。>
|
||||
|
||||
**Expected**
|
||||
|
||||
<列出预期输出、状态或证据。>
|
||||
|
||||
**Decision**
|
||||
|
||||
- 若 <条件 A>,进入 Step 2。
|
||||
- 若 <条件 B>,进入 Troubleshooting A。
|
||||
- 若无法判断或状态冲突,`STOP` 并记录证据。
|
||||
|
||||
### Step 2 — Change
|
||||
|
||||
**Action**
|
||||
|
||||
<描述单一、可审计的变更动作。>
|
||||
|
||||
**Expected**
|
||||
|
||||
<变更后应出现的状态。>
|
||||
|
||||
**Verification**
|
||||
|
||||
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
|
||||
|
||||
**Rollback**
|
||||
|
||||
- 触发条件:<什么情况需要回滚>
|
||||
- 回滚动作:<如何撤销>
|
||||
- 回滚验证:<如何确认恢复成功>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Troubleshooting A — <异常名称>
|
||||
|
||||
- 证据收集:<日志、指标、命令输出、链接>
|
||||
- 允许动作:<仅限已验证且低风险的动作>
|
||||
- 下一步:<回到某步 / 转入另一 runbook / STOP 并升级>
|
||||
|
||||
## Final Verification
|
||||
|
||||
只有同时满足以下标准,流程才算成功:
|
||||
|
||||
- <功能或服务状态>
|
||||
- <自动化测试或健康检查>
|
||||
- <监控指标或告警状态>
|
||||
- <变更记录、PR 或 Issue 已更新>
|
||||
|
||||
## Failure Handling
|
||||
|
||||
若未能完成:
|
||||
|
||||
1. 停止进一步变更。
|
||||
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
|
||||
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
|
||||
4. 按 <升级渠道> 交接,不继续猜测。
|
||||
|
||||
## References
|
||||
|
||||
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
|
||||
@@ -1,5 +1,20 @@
|
||||
# AdGuard Home health — dns.windy.lan
|
||||
|
||||
## Purpose
|
||||
|
||||
Read-only health check of the AdGuard Home LAN DNS service.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: [dns.windy.lan](../hosts/dns.windy.lan.md) (`192.168.66.36`).
|
||||
- Read-only: does not expose query-log contents or secrets; does not change configuration.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: dns.windy.lan (`/opt/adguardhome`)
|
||||
|
||||
This runbook is read-only. It does not expose query-log contents or secrets.
|
||||
|
||||
Routine checks run through Ansible on demand:
|
||||
@@ -58,3 +73,9 @@ a known-bad-signature test; an enabled DO bit alone is not validation.
|
||||
|
||||
Private PTR forwarding is intentionally absent because the EdgeRouter does
|
||||
not currently answer private PTR requests.
|
||||
|
||||
## Safety
|
||||
|
||||
- Read-only: never change the DNS policy or the `agh-ui-access.service` nftables rule during this check.
|
||||
- Do not infer a broken DNS policy from an empty `allowed_clients`.
|
||||
- If live state conflicts with an expected value, `STOP` and report.
|
||||
|
||||
@@ -1,8 +1,29 @@
|
||||
# Runbook: routine operations through Ansible
|
||||
|
||||
## Purpose
|
||||
|
||||
Routine operations (health, reconcile, maintenance) through the Ansible playbooks.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: every inventory host, run from `ansible/`.
|
||||
- Not applicable: arbitrary remote commands — the reconcile playbook is allowlisted and gated.
|
||||
|
||||
Run commands from `ansible/`. The inventory forces IPv4 and uses the `windy`
|
||||
account with sudo. Do a read-only health pass before any reconciliation.
|
||||
|
||||
## Safety
|
||||
|
||||
- Read-only health pass before any reconciliation.
|
||||
- Mutating playbooks require explicit confirmation variables; do not bypass them.
|
||||
- If a reconcile target or service name is not allowlisted, `STOP` — do not invent one.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: Ansible control-plane + all inventory hosts
|
||||
|
||||
## Health report (read-only)
|
||||
|
||||
```bash
|
||||
@@ -48,6 +69,27 @@ Do not use this playbook for a Mailcow update, database migration, DNS record
|
||||
change, or secret rotation. Those operations require their dedicated reviewed
|
||||
and, where appropriate, interactive procedures.
|
||||
|
||||
## Deploy repo-owned Compose (static projects)
|
||||
|
||||
Repo source: `compose/<project>/compose.yml` (non-secret; secrets come from the
|
||||
server-local `.env` via `${VAR}`). Mechanism and per-project status:
|
||||
[`compose/README.md`](../compose/README.md).
|
||||
|
||||
```bash
|
||||
# Read-only: staged-file diff + allowlist/confirmation asserts, no writes
|
||||
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden --check --diff
|
||||
ansible-playbook playbooks/compose-deploy.yml --limit powerdns --check --diff
|
||||
|
||||
# Apply: stage repo file → validate `docker compose config -q` against the
|
||||
# server .env → backup current file (*.bak-<ts>) → promote → `up -d` (gated)
|
||||
ansible-playbook playbooks/compose-deploy.yml --limit vaultwarden \
|
||||
-e '{"compose_deploy_confirm": true}'
|
||||
```
|
||||
|
||||
The playbook never writes, reads, or transfers the server `.env`. A failed
|
||||
validation never touches the live compose file. Hosts without an allowlisted
|
||||
`compose_repo_project` fail the assert — do not invent targets.
|
||||
|
||||
## Host-level maintenance
|
||||
|
||||
These playbooks cover every inventory host, including the Matrix K3s node:
|
||||
@@ -60,6 +102,30 @@ ansible-playbook playbooks/maintenance-preview.yml
|
||||
ansible-playbook playbooks/baseline.yml
|
||||
```
|
||||
|
||||
## us4 firewalld reconciliation
|
||||
|
||||
The us4 playbook owns only the audited `public` zone allowlist. It fails closed
|
||||
on unknown services or ports, never reloads/restarts firewalld, and does not
|
||||
manage Docker-published ports.
|
||||
|
||||
```bash
|
||||
cd ansible
|
||||
ansible-galaxy collection install -r requirements.yml
|
||||
|
||||
# Read-only preview
|
||||
ansible-playbook playbooks/us4-firewalld.yml --limit us4 --check --diff
|
||||
|
||||
# Apply only after testing the provider console and retaining an independent
|
||||
# SSH rollback session.
|
||||
ansible-playbook playbooks/us4-firewalld.yml --limit us4 \
|
||||
-e '{"us4_firewalld_confirm": true, "us4_console_confirm": true}'
|
||||
```
|
||||
|
||||
Apply creates a protected server-local backup and schedules a 15-minute
|
||||
automatic rollback before changing rules. The rollback is cancelled only after
|
||||
the playbook verifies fresh SSH/sudo access, public HTTPS routes, SMTP, Docker,
|
||||
Fail2ban, and WireGuard. Do not bypass either confirmation variable.
|
||||
|
||||
## UniFi SSO login setting (mutating)
|
||||
|
||||
Reconciles `super_sdn.sso_login_enabled` on the UniFi controller (host `ubnt`,
|
||||
|
||||
@@ -0,0 +1,76 @@
|
||||
# Runbook: fix a failing health/playbook run
|
||||
|
||||
> Adapted from the upstream guide's `fix-ci`. This repo has no software CI; the
|
||||
> equivalent "pipeline" is the Ansible **health report** and the gated playbooks.
|
||||
> This runbook covers diagnosing and fixing a failed or warning/critical run.
|
||||
|
||||
## Purpose
|
||||
|
||||
Diagnose and fix a failing Ansible health-report or playbook run without
|
||||
skipping checks or changing unrelated code.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: `ansible-playbook playbooks/health-report.yml` and the gated playbooks under `ansible/playbooks/`.
|
||||
- Not applicable: production changes beyond fixing the run; network/DNS changes → `network-change.md`.
|
||||
|
||||
## Safety
|
||||
|
||||
- Do not skip or weaken a failing check to make it pass.
|
||||
- Do not change unrelated hosts or services.
|
||||
- Prefer read-only diagnosis before mutation; destructive fixes require approval.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: Ansible health report / gated playbooks
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Reproduce and read
|
||||
|
||||
**Action** — re-run the failing playbook with `--limit <host>` and capture the task that failed.
|
||||
|
||||
```bash
|
||||
cd ansible
|
||||
ansible-playbook playbooks/health-report.yml --limit <host> -v
|
||||
```
|
||||
|
||||
**Expected** — a specific failed task, host, and message (warning vs critical).
|
||||
|
||||
**Decision** — clear failure → Step 2; ambiguous → `STOP` and collect `-vvv` output + the relevant `latest.json`.
|
||||
|
||||
### Step 2 — Diagnose
|
||||
|
||||
**Action** — inspect the corresponding service on the host using the matching health runbook (`mailcow-health.md`, `vaultwarden-health.md`, `pdns-health.md`, etc.).
|
||||
|
||||
**Expected** — a root cause (container down, cert expired, queue backlog, drift).
|
||||
|
||||
**Decision** — root cause found → Step 3; live state conflicts with the runbook's assumptions → `STOP`.
|
||||
|
||||
### Step 3 — Fix within scope
|
||||
|
||||
**Action** — apply the minimal fix the service runbook prescribes (e.g. `compose-reconcile` for a config drift, or a documented update). Use only allowlisted/gated playbooks.
|
||||
|
||||
**Verification** — re-run the health report and confirm it passes.
|
||||
|
||||
**Rollback** — revert to the prior config/state and re-run; see `rollback.md` for the general procedure.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Troubleshooting A — Intermittent/flaky failure
|
||||
|
||||
- Evidence: timing, DNS stub flakiness (use `1.1.1.1`/`8.8.8.8` for probes).
|
||||
- Allowed: re-run once with the documented resolver workaround.
|
||||
- Next: still failing → `STOP` and escalate.
|
||||
|
||||
## Final Verification
|
||||
|
||||
- Health report passes for the affected host.
|
||||
- No checks were skipped or weakened; the fix is committed/documented.
|
||||
|
||||
## References
|
||||
|
||||
- [`ansible-operations.md`](ansible-operations.md)
|
||||
- Per-service health runbooks under [`runbooks/`](.)
|
||||
@@ -0,0 +1,420 @@
|
||||
# Runbook: Home Assistant maintenance (hass.windy.lan)
|
||||
|
||||
Target: [hass.windy.lan](../hosts/hass.windy.lan.md) (physical x88 Pro box, HAOS `machine: green`)
|
||||
Upstream: HAOS 18.2 / Supervisor 2026.09.0 / Core 2026.9.1 (verified 2026-09-13)
|
||||
|
||||
This runbook covers routine Home Assistant maintenance through the **`ha`
|
||||
supervisor CLI**. All commands are wrapped by a single script
|
||||
[`scripts/ha-maintenance.sh`](scripts/ha-maintenance.sh); the sections below
|
||||
document the exact commands it runs, for manual/agent use.
|
||||
|
||||
## Purpose
|
||||
|
||||
Run routine Home Assistant maintenance on `hass.windy.lan` (health snapshot,
|
||||
config validation, log inspection, updates, and recovery) through the `ha`
|
||||
supervisor CLI.
|
||||
|
||||
## Scope
|
||||
|
||||
Applies to `hass.windy.lan` only (HAOS, `machine: green`). Covers both
|
||||
read-only checks and gated mutating operations; the "Command families
|
||||
intentionally NOT scripted" table below lists what is deliberately out of
|
||||
scope.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: hass.windy.lan (HAOS, `machine: green`)
|
||||
|
||||
## Safety
|
||||
|
||||
- Prefer read-only checks first; the health snapshot mutates nothing.
|
||||
- Every mutating mode (update / restart / rebuild / rollback / reboot /
|
||||
backup / restore / add-on lifecycle) refuses to run without `--yes`.
|
||||
- `--restore` overwrites the current installation; `--rollback-os`,
|
||||
`--reboot`, and `--rebuild-core` are disruptive. Run them only from a
|
||||
planned recovery with the backup verified.
|
||||
- Never commit `SUPERVISOR_TOKEN` or a long-lived `HA_TOKEN`; read entity
|
||||
state via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`).
|
||||
- The `--restart-core` wrapper exits 1 silently on ssh failure — treat an
|
||||
empty/exit-1 result as failure and confirm with `ha core info`.
|
||||
- If live state conflicts with a documented expectation, `STOP` and report;
|
||||
do not improvise command families outside this script.
|
||||
|
||||
## Access pattern
|
||||
|
||||
`ha` authenticates to the Supervisor with `SUPERVISOR_TOKEN`. Interactive SSH
|
||||
login works because `~hassio/.zprofile` runs `exec sudo -i`; the root login
|
||||
environment carries the token. Non-interactive use must be:
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha <cmd>'
|
||||
```
|
||||
|
||||
Running `ha` as `hassio` directly (or `sudo -n` without `-i`) returns
|
||||
`unauthorized: missing or invalid API token`.
|
||||
|
||||
**MOTD:** every `sudo -n -i` login prints the SSH & Web Terminal MOTD banner.
|
||||
The script runs its whole procedure in one remote login (`sudo -n -i bash -s`)
|
||||
so the banner appears once, then strips it with `awk` up to the
|
||||
`System is ready! Use browser or app to configure.` line.
|
||||
|
||||
**Restart wrapper (verified 2026-08-14, W1N-107):**
|
||||
`./ha-maintenance.sh --restart-core --yes` exited 1 with no output in <1s
|
||||
and **did not restart Core**. The wrapper pipes a remote script through
|
||||
`ssh … 2>/dev/null | awk …`; with `set -uo pipefail`, an ssh failure is
|
||||
silent and the pipeline returns empty/exit 1 **before any remote command
|
||||
runs**. That is not a MOTD-strip artifact after a successful restart.
|
||||
The working restart was
|
||||
`ssh -o BatchMode=yes hassio@hass.windy.lan 'sudo -n -i ha core restart'`
|
||||
(~131s, `Command completed successfully.`). Treat empty/exit 1 as
|
||||
failure; confirm with elapsed time and `ha core info`.
|
||||
|
||||
## Script usage
|
||||
|
||||
```bash
|
||||
cd runbooks/scripts
|
||||
|
||||
./ha-maintenance.sh # read-only health snapshot
|
||||
./ha-maintenance.sh --check-config # validate core configuration
|
||||
./ha-maintenance.sh --logs core 2500 # tail core logs (use 2500 after a restart)
|
||||
./ha-maintenance.sh --logs supervisor # tail supervisor logs (default 100)
|
||||
./ha-maintenance.sh --logs host 50 # tail host journald logs
|
||||
./ha-maintenance.sh --logs apps:<slug> # tail an add-on log
|
||||
|
||||
# Mutating — refuse to run without --yes:
|
||||
./ha-maintenance.sh --update --yes # refresh + update core(--backup)/supervisor/os
|
||||
./ha-maintenance.sh --restart-core --yes # restart Core
|
||||
./ha-maintenance.sh --restart-core --safe-mode --yes # restart Core in safe mode
|
||||
./ha-maintenance.sh --rebuild-core --yes # rebuild Core image (after options change)
|
||||
./ha-maintenance.sh --rollback-os --yes # boot previous OS slot (A/B rollback)
|
||||
./ha-maintenance.sh --reboot --yes # reboot the HAOS host
|
||||
./ha-maintenance.sh --backup [NAME] --yes # full backup (optionally named)
|
||||
./ha-maintenance.sh --restore <slug> --yes # restore a backup (DESTRUCTIVE)
|
||||
./ha-maintenance.sh --app restart core_mosquitto --yes # add-on lifecycle
|
||||
```
|
||||
|
||||
- `--app` action is one of `start|stop|restart|update`; needs an add-on slug.
|
||||
- `HA_HOST` / `HA_SSH_USER` override the defaults (`hass.windy.lan` / `hassio`).
|
||||
- `--restore` overwrites the current installation — run only from a planned
|
||||
recovery, with the backup verified.
|
||||
|
||||
## Command reference (verified 2026-08-14)
|
||||
|
||||
All verified against the live host. MOTD prepends each command's output; strip
|
||||
with the `awk` pattern above or read the last block.
|
||||
|
||||
### Routine / read-only
|
||||
|
||||
| Purpose | Command |
|
||||
|---|---|
|
||||
| General overview | `ha info` |
|
||||
| Core version/status | `ha core info` (this CLI build has no `state:` field; success is a normal info dump) |
|
||||
| Core config validation | `ha core check` |
|
||||
| Core stats | `ha core stats` |
|
||||
| Supervisor status | `ha supervisor info` (incl. add-on list) |
|
||||
| Supervisor stats | `ha supervisor stats` |
|
||||
| OS status | `ha os info` (boot slots A/B) |
|
||||
| Host status | `ha host info` (disk free/total, kernel) |
|
||||
| Network | `ha network info` (`supervisor_internet`) |
|
||||
| Hardware | `ha hardware info` |
|
||||
| Pending updates | `ha available-updates` |
|
||||
| Reload stores/versions | `ha refresh-updates` |
|
||||
| Job manager | `ha jobs info` |
|
||||
| Resolution center | `ha resolution info` |
|
||||
| Core logs | `ha core logs -n 100` (`-f` follow, `-b` boot id). Default 100 misses setup; use `-n 2500` after a custom-component restart. `/config/home-assistant.log` may be missing — `ha core logs` is the source of truth. |
|
||||
| Supervisor logs | `ha supervisor logs -n 100` |
|
||||
| Host journald logs | `ha host logs -n 100` |
|
||||
| Add-on logs | `ha apps logs <slug> -n 100` |
|
||||
| Add-on list | `ha supervisor info` → `addons:` (started/stopped/error) |
|
||||
| Security integrity | `ha security integrity` |
|
||||
|
||||
### Mutating (require --yes)
|
||||
|
||||
| Purpose | Command |
|
||||
|---|---|
|
||||
| Update core (with partial backup) | `ha core update --backup` |
|
||||
| Update supervisor | `ha supervisor update` |
|
||||
| Update OS | `ha os update` |
|
||||
| Update add-on | `ha apps update <slug>` |
|
||||
| Restart core | `ha core restart` / `ha core restart --safe-mode` |
|
||||
| Rebuild core | `ha core rebuild` |
|
||||
| OS rollback | `ha os boot-slot other` |
|
||||
| Reboot host | `ha host reboot` |
|
||||
| Full backup | `ha backups new [--name NAME]` |
|
||||
| Restore backup | `ha backups restore <slug>` |
|
||||
| Add-on start/stop/restart | `ha apps start\|stop\|restart <slug>` |
|
||||
|
||||
## Procedure
|
||||
|
||||
### 1. Health snapshot (read-only)
|
||||
|
||||
```bash
|
||||
./ha-maintenance.sh
|
||||
```
|
||||
|
||||
Review: supervisor `healthy: true`/`supported: true`; core/OS `update_available`;
|
||||
add-on states (any `state: error`?); `resolution info` issues; disk free.
|
||||
|
||||
### 2. Validate config after any `configuration.yaml` change
|
||||
|
||||
```bash
|
||||
./ha-maintenance.sh --check-config
|
||||
```
|
||||
|
||||
Expect `Command completed successfully.` before a Core restart.
|
||||
|
||||
`ha core check` / a YAML reload is **not** enough after copying Python
|
||||
custom-component files — restart Core.
|
||||
|
||||
### 3. Inspect logs
|
||||
|
||||
```bash
|
||||
./ha-maintenance.sh --logs core 2500 # after a Core restart / custom-component copy
|
||||
./ha-maintenance.sh --logs supervisor
|
||||
./ha-maintenance.sh --logs apps:core_mosquitto
|
||||
```
|
||||
|
||||
Default `--logs core` (100 lines) is too short to catch coordinator pickle /
|
||||
setup errors. `/config/home-assistant.log` may be absent while
|
||||
`ha core logs` still has history.
|
||||
|
||||
### 4. Apply updates (mutating)
|
||||
|
||||
```bash
|
||||
./ha-maintenance.sh --update --yes
|
||||
```
|
||||
|
||||
Runs `refresh-updates` → `core update --backup` (partial backup first) →
|
||||
`supervisor update` → `os update`, then re-prints pending updates. Prefer the
|
||||
web UI (**Settings → System → Updates**) for a human-supervised pass.
|
||||
|
||||
### 5. Recovery operations (mutating, only when needed)
|
||||
|
||||
```bash
|
||||
./ha-maintenance.sh --restart-core --safe-mode --yes # start Core without custom integrations
|
||||
./ha-maintenance.sh --rollback-os --yes # OS update broke boot? go back one slot
|
||||
./ha-maintenance.sh --restore <slug> --yes # full restore; overwrites current install
|
||||
```
|
||||
|
||||
OS update policy: HAOS uses two boot slots (A/B); `ha os info` shows which slot
|
||||
booted. After a bad OS update, `ha os boot-slot other` boots the previous slot.
|
||||
|
||||
### 6. Backup before major changes
|
||||
|
||||
```bash
|
||||
./ha-maintenance.sh --backup pre-migration --yes # named backup
|
||||
```
|
||||
|
||||
### 7. Install or update a custom component (manual zip)
|
||||
|
||||
Home Assistant loads custom integrations from
|
||||
`/config/custom_components/<domain>/` (on this HAOS host `/config` ≡
|
||||
`/homeassistant`). Official lookup:
|
||||
`<config>/custom_components/<domain>` then built-in
|
||||
`homeassistant/components/<domain>`
|
||||
([Integration file structure](https://developers.home-assistant.io/docs/creating_integration_file_structure)).
|
||||
A folder named after the domain, with at least `manifest.json` and
|
||||
`__init__.py`, is enough. **Restart Core** after copying — `ha core check`
|
||||
and a YAML reload do not pick up new Python packages.
|
||||
|
||||
This host's live trees are **file copies**, not git clones. Do not
|
||||
`git pull` inside `custom_components/`.
|
||||
|
||||
#### Official plugin paths (CSG)
|
||||
|
||||
[windyboy/china_southern_power_grid_stat README](https://github.com/windyboy/china_southern_power_grid_stat):
|
||||
[HACS](https://hacs.xyz/) **or**
|
||||
[手动下载安装](https://github.com/windyboy/china_southern_power_grid_stat/releases).
|
||||
|
||||
This host uses the zip path. **Do not HACS-update this integration here.**
|
||||
HACS still tracks upstream `CubicPill/china_southern_power_grid_stat`
|
||||
`v1.2.0` and would overwrite the fork. Releases have no uploaded zip
|
||||
assets — use GitHub **Source code (zip)** / zipball of the tag.
|
||||
|
||||
Worked SSH example (tag, backup, `rsync`, `__pycache__`, restart):
|
||||
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § Manual
|
||||
custom-component install.
|
||||
|
||||
#### Procedure
|
||||
|
||||
1. **Backup the live tree off `custom_components/`.** HA scans every
|
||||
directory under `custom_components/` whose `manifest.json` `domain`
|
||||
matches. A `*.bak-*` folder next to the live tree makes Core import
|
||||
the backup (`No module named '...bak-YYYYMMDD-...'`, W1N-106). CSG
|
||||
backups: `/homeassistant/.csg-backups/`.
|
||||
2. **Copy only the inner `custom_components/<domain>/` tree**, not the
|
||||
repo root and not an extra nested folder.
|
||||
3. **Wipe `__pycache__` as root.** `rsync --delete` as `hassio` cannot
|
||||
unlink Core-owned `.pyc` (permission denied, exit 23); stale
|
||||
`cpython-314` bytecode can keep the old coordinator in memory until
|
||||
restart. Then restart:
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i rm -rf /homeassistant/custom_components/<domain>/__pycache__ \
|
||||
/homeassistant/custom_components/<domain>/*/__pycache__ &&
|
||||
sudo -n -i ha core restart'
|
||||
```
|
||||
|
||||
4. **Wait 1–2 min**, then `ha core info` (this CLI build has no `state:`
|
||||
field; success is a normal info dump). Confirm `manifest.json`
|
||||
`version` matches the tag.
|
||||
5. **Read enough Core logs.** Default `ha core logs` is too short to
|
||||
catch setup. Use `-n 2500` (or `--logs core 2500`) and look for
|
||||
`Setting up <domain>` plus the first coordinator errors.
|
||||
6. **First poll can time out.** If last-month sensors have numbers but
|
||||
this-month stay `unknown`/`unavailable`, reload the config entry
|
||||
(UI: integration → Reload). Supervisor:
|
||||
|
||||
```bash
|
||||
# entry id from .storage/core.config_entries (CSG: 01KGCQDSZCF523A9X6SV3BZ1B9)
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i python3 -c "
|
||||
import os, urllib.request
|
||||
req = urllib.request.Request(
|
||||
\"http://supervisor/core/api/config/config_entries/entry/<ENTRY_ID>/reload\",
|
||||
method=\"POST\",
|
||||
headers={\"Authorization\": \"Bearer \" + os.environ[\"SUPERVISOR_TOKEN\"]},
|
||||
)
|
||||
print(urllib.request.urlopen(req, timeout=60).status)
|
||||
"'
|
||||
```
|
||||
|
||||
7. **Do not edit the dashboard or `templates/csg_sensors.yaml` for an
|
||||
install.** Entity IDs did not change across v1.3.0/v1.3.1/v1.3.2.
|
||||
(The old `| float(0)` fake-zero follow-up was resolved 2026-08-29 by
|
||||
W1N-239: template sensors now carry `availability` templates and show
|
||||
`unavailable` instead of fake zeros when native CSG sensors are down.
|
||||
Template edits go through that issue, not the install path.)
|
||||
|
||||
#### Verify (CSG, after v1.3.2 / W1N-118)
|
||||
|
||||
| Check | Expect |
|
||||
|---|---|
|
||||
| `manifest.json` `version` | `1.3.2` |
|
||||
| `ha core logs` after this restart | `Setting up china_southern_power_grid_stat`; **no** `cannot pickle 'mappingproxy'` |
|
||||
| Config entry | `state: loaded` |
|
||||
| `sensor.0800041935246530_balance` | numeric (may be `0.0`) |
|
||||
| `sensor.0800041935246530_this_month_total_usage` | numeric after reload if first poll timed out |
|
||||
| Native `*_total_cost` / `current_ladder` | may stay `unknown` (CSG marketing calendar SQL error); dashboard uses W1N-114 `csg_*` ladder/cost templates |
|
||||
|
||||
`monetary` + `total_increasing` warnings on this-month/year cost sensors
|
||||
are a remaining plugin issue, not an install failure.
|
||||
|
||||
There is no long-lived `HA_TOKEN` in the agent environment. Read entity
|
||||
states via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`) at
|
||||
`http://supervisor/core/api/states/<entity_id>`.
|
||||
|
||||
#### CSG display refactor 2026-09-04 (VPS-90)
|
||||
|
||||
Template/dashboard changes made **after** pricing cross-check (8月账单
|
||||
198.64 元 vs 模板 198.65 元,≤0.01 元;阶梯常量 0.589/0.639/0.889、
|
||||
260/600 夏档未动):
|
||||
|
||||
- `templates/csg_sensors.yaml` Block B 新增
|
||||
`sensor.csg_this_month_avg_price`(本月阶梯电费÷本月用电,`元/kWh`);
|
||||
**csg_* template sensors = 15**。
|
||||
- Panel `power-monitor`(`lovelace.dashboard_unknown`):环比行改名
|
||||
「环比上月同期」;glance「本月/上月」去重为单卡「上月」(本月行归
|
||||
💰核心数据卡);⚡阶梯电价卡加「本月实际均价」行。实体引用 20→21。
|
||||
- `automations.yaml` +2 提醒:`automation.csg_mian_ban_qie_dong_ji_dang_ti_xing`
|
||||
(10-25)/ `automation.csg_mian_ban_qie_xia_ji_dang_ti_xing`(4-25)
|
||||
09:00 Matrix 提醒人工切「本月累计」gauge 季节档(max/segments 不可模板化)。
|
||||
- 金额单位混排(原生 CNY vs 模板 元)**保留**:`config/entity_registry/update`
|
||||
拒绝自定义文本单位(`extra keys not allowed … Got '元'`),已定案接受。
|
||||
|
||||
**WS 改面板(2026.8,本机实测,后续沿用)**: core/主机 python 无 ws 库、
|
||||
core 容器内经 supervisor 代理 WS 被拒(loop prevention)。用
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i sh -c "docker run --rm -i --network host -e SUPERVISOR_TOKEN \
|
||||
--entrypoint python3 r.hassbus.com/home-assistant/aarch64-hassio-supervisor:2026.08.0 \
|
||||
- < /tmp/x.py"'
|
||||
```
|
||||
|
||||
连 `ws://172.30.32.2/core/websocket`(aiohttp,header `Authorization: Bearer
|
||||
$SUPERVISOR_TOKEN`,随后 auth 帧同 token)。命令名 **`lovelace/config`**(读)
|
||||
+ **`lovelace/config/save`**(写,url_path + 全量 config);`lovelace/config/get`
|
||||
已不存在(unknown_command)。备份与细节见
|
||||
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § CSG 面板重构 2026-09-04。
|
||||
|
||||
## Command families intentionally NOT scripted
|
||||
|
||||
These exist in `ha` but are either rare, dangerous, or better done in the web
|
||||
UI; documented here so nothing is a surprise. Use `ha <family> --help` on the
|
||||
host for exact syntax.
|
||||
|
||||
| Family | Notes |
|
||||
|---|---|
|
||||
| `ha audio` | Audio device management; peripheral. |
|
||||
| `ha authentication` | `auth list/reset/cache`; user password ops — do in web UI. `auth list` is local-terminal only. |
|
||||
| `ha cli` | Internal CLI backend info/update; self-maintained. |
|
||||
| `ha dns` | Internal DNS server; only relevant if Supervisor DNS add-on in use. |
|
||||
| `ha docker` | Host Docker backend info/options/registries; HAOS-managed. |
|
||||
| `ha mounts` | Network storage (NFS/CIFS) mounts — configure in **Settings → System → Storage**. |
|
||||
| `ha multicast` / `ha observer` | Internal services; self-maintained. |
|
||||
| `ha network scan/update/vlan` | WiFi AP scan & interface config — prefer web UI networking. |
|
||||
| `ha host disks/options/shutdown/reload` | Disk ops / host options; `shutdown` is equivalent to `--reboot` but off. |
|
||||
| `ha os datadisk list/move/wipe` | Data-disk migration; `wipe` is **local-terminal only** and erases all data. |
|
||||
| `ha os import` | Import config from USB stick. |
|
||||
| `ha os boards` / `os config` | Board / OS settings. |
|
||||
| `ha core options` / `supervisor options` | Core/OS config options (e.g. `--duplicate-log-file`); changes need `ha core rebuild` + restart. |
|
||||
| `ha backups freeze/thaw/remove/options` | Freeze/thaw for external backup tools; removal is destructive. |
|
||||
| `ha jobs options/reset` | Job-manager tuning. |
|
||||
| `ha resolution check/healthcheck/issue/suggestion` | Resolution center management; `healthcheck` runs fixups. |
|
||||
| `ha store add/delete/repair` | Repository management — add repos in web UI app store. |
|
||||
| `ha security info/options` | Security backend options. |
|
||||
|
||||
## Docs vs actual CLI discrepancies
|
||||
|
||||
The [official HAOS common-tasks docs](https://www.home-assistant.io/common-tasks/os/)
|
||||
also mention `ha host update`, which **does not exist** in this CLI
|
||||
(2026-08-14). Docs' `ha backups list` is not a subcommand either:
|
||||
`ha backups --help` lists freeze/info/new/options/reload/remove/restore/thaw;
|
||||
extra positional args (`list`, `nonsense`, ...) are ignored and the default
|
||||
list still prints with exit 0. The list command is plain `ha backups`.
|
||||
Per-backup: `ha backups info <slug>` (slug required). Trust the server CLI
|
||||
(`ha <cmd> --help`) over the docs.
|
||||
|
||||
This CLI's `ha core info` also has no `state:` field (verified 2026-08-14).
|
||||
Wait for a successful info dump after restart, not a `state: running` line.
|
||||
|
||||
## Known issues on hass.windy.lan (2026-08-14)
|
||||
|
||||
2026-08-13 snapshot items were resolved same day (W1N-70/71/72/73/74/75/76):
|
||||
OTBR and the duplicate SSH add-on uninstalled, resolution-center empty,
|
||||
full backup `pre-maintenance-20260813` (slug `411a4ba5`). Remaining:
|
||||
|
||||
- **Bluetooth hci0 instability (RTL8821CS)**: `bluetooth_auto_recovery`
|
||||
power-reset times out every ~2 min; kernel `hci0 hardware error`. No BLE
|
||||
entities exist, so no user impact. HAOS image ships `x88-bt-hci-recovery`
|
||||
workaround units.
|
||||
- `host info` reports `disk_life_time: 10` (boot eMMC ~10% life left) —
|
||||
monitor on each snapshot; plan disk replacement / data-disk migration.
|
||||
- **Home PPPoE IPv4 to CSG is blackholed** (`curl -4` to
|
||||
`218.19.148.218:443` times out). `end0` IPv6 works (`curl -6
|
||||
https://95598.csg.cn` → HTTP 200). Entry `ip_family: ipv4` still
|
||||
matches the stored option; first post-restart poll can still time out
|
||||
— reload the config entry rather than reinstalling.
|
||||
- **WSL HTTP proxy**: LAN `hass.windy.lan:8123` through Mihomo returns
|
||||
empty `502`. Bypass proxy or add `.windy.lan` to `NO_PROXY` before
|
||||
debugging UI/API from the workstation
|
||||
([hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) § HTTP proxy
|
||||
gotcha).
|
||||
- **No long-lived HA token in the agent environment.** Read entity
|
||||
states via Supervisor (`SUPERVISOR_TOKEN` after `sudo -n -i`) at
|
||||
`http://supervisor/core/api/states/...`, not a committed `HA_TOKEN`.
|
||||
|
||||
## Pass criteria
|
||||
|
||||
- Health snapshot completes; supervisor `healthy`/`supported: true`
|
||||
- Mutating modes refuse to run without `--yes` (incl. `--restore`, `--app`)
|
||||
- `--check-config` returns success
|
||||
- Update / rollback / restore / reboot confirmed only after explicit `--yes`
|
||||
- Custom-component zip install: live `manifest.json` version matches the
|
||||
tag; backups not under `custom_components/`; Core restarted; logs show
|
||||
`Setting up <domain>` without import / pickle errors
|
||||
- Update the **Verified** line on [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md)
|
||||
@@ -0,0 +1,276 @@
|
||||
# Runbook: Host disk cleanup (unbounded container logs / apt cache / docker artifacts)
|
||||
|
||||
## Purpose
|
||||
|
||||
Reclaim space on a root filesystem that is filling up (≥70% used) on a Docker
|
||||
Compose host, by fixing unbounded container log growth at the source, clearing
|
||||
apt/journal caches, and removing unused Docker images/volumes. Success: root
|
||||
usage drops to a safe band (≤55% used, or per acceptance in the tracking issue)
|
||||
and log growth stays bounded afterwards.
|
||||
|
||||
## Scope
|
||||
|
||||
- 适用环境: production single-root-fs hosts running Docker Compose stacks
|
||||
(first application: `hk2.chans.xyz`; reusable for `mx2.windy.me` / `us2.wsvc.info`
|
||||
which run the same unbounded-`json.log` pattern).
|
||||
- 适用对象: root filesystem usage; container stdout/stderr log files
|
||||
(`/var/lib/docker/containers/*/*-json.log`); `/var/cache/apt`; systemd journal;
|
||||
unused Docker images / anonymous volumes / build cache.
|
||||
- 不适用情形: hosts without systemd-journald or without Docker; LAN/HAOS hosts
|
||||
(use their own runbooks); cases needing disk *growth* (provider resize) rather
|
||||
than cleanup; anything touching service data volumes or `/opt/*` configs
|
||||
(STOP and use the service-specific runbook instead).
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: windy (operator) + agent executing per approval
|
||||
- Last reviewed: 2026-09-02
|
||||
- Related systems: hk2.chans.xyz (PowerDNS auth / AdGuard Home / Traefik / RustDesk compose stacks)
|
||||
|
||||
## Preconditions
|
||||
|
||||
- SSH access to the target host with **passwordless sudo** (`sudo -n true` must succeed).
|
||||
- A recorded `df -h` baseline and `docker system df` baseline.
|
||||
- **Explicit user approval** for every service touch listed in Approval gates
|
||||
(recorded in the tracking issue, e.g. Plane `vps` VPS-81).
|
||||
- No open incident on the target host.
|
||||
- Container log growth root cause identified in Diagnose before mutating.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Input | Source | Required | Validation |
|
||||
|---|---:|---|
|
||||
| Target host | inventory/hosts.md | yes | SSH login + `uname -r` |
|
||||
| df/docker baseline | live read-only probe | yes | recorded before first mutation |
|
||||
| Approved service touches | user confirmation in tracking issue | yes | issue comment states approval |
|
||||
| Image keep-list (rollback pins) | operator decision in issue | yes | review `docker image ls` before rmi |
|
||||
| Backup of any config edited | local copy with timestamp | yes | exists before edit |
|
||||
|
||||
## Safety
|
||||
|
||||
### Non-negotiable rules
|
||||
|
||||
- Prefer read-only diagnosis before mutation (never mutate on an unmeasured disk).
|
||||
- Never use `rm` on a live container log — use `truncate -s 0` (keeps the fd valid).
|
||||
- Never run `docker image prune -a` when a keep-list is intended — no keep-list
|
||||
exists; delete explicitly with `docker rmi`.
|
||||
- Never run `docker volume prune -a` — plain `docker volume prune` (no `-a`)
|
||||
removes only unused anonymous volumes; named/in-use volumes stay.
|
||||
- After every mutation, verify the expected state (`df -h`, container status).
|
||||
- Destructive actions require explicit approval (Approval gates).
|
||||
|
||||
### Stop conditions
|
||||
|
||||
- Live state conflicts with this runbook's preconditions or expectations (e.g.
|
||||
root usage differs wildly from baseline, or a container is unhealthy).
|
||||
- Missing approval, missing backup, or missing rollback ability.
|
||||
- A verification step fails with no documented next step.
|
||||
- Any step would touch a volume/container/mount that is not on the approved list.
|
||||
|
||||
### Approval gates
|
||||
|
||||
| Action | Risk | Explicit approval | Approval record |
|
||||
|---|---:|---|---|
|
||||
| `docker restart <chatty container>` | low (sec-level blip of that service only) | yes | tracking issue (VPS-81 T1) |
|
||||
| `systemctl restart systemd-journald` | low (sec-level, no state loss) | yes | tracking issue (VPS-81 T2) |
|
||||
| `apt-get clean` | low (re-downloadable) | no | — |
|
||||
| `journalctl --vacuum-*` / journald drop-in | low | no (restart above is gated) | — |
|
||||
| `docker rmi` of unused images | medium (rollback pin removed unless kept) | yes (keep-list) | tracking issue (VPS-81 T3) |
|
||||
| `docker volume prune` | medium (data in anonymous volumes lost) | yes | tracking issue (VPS-81 T4) |
|
||||
| `docker builder prune` | low | no | — |
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Diagnose
|
||||
|
||||
**Action**
|
||||
|
||||
Read-only: `df -h`, `df -i`, `sudo du -x -h --max-depth=1 /`, `docker system df`,
|
||||
and locate oversized container logs:
|
||||
`sudo ls -la /var/lib/docker/containers/*/*-json.log`. Map a big log to its
|
||||
container (`docker inspect -f '{{.Name}} {{.LogPath}}' <id>`), then inspect what
|
||||
it logs (`sudo tail -c 400000 <logpath>`; count `[debug]` lines) and find the
|
||||
config flag driving it (e.g. AGH `log.verbose` in its YAML; note `log.file: ""`
|
||||
means the app's own rotation keys are inert and output goes to the container log).
|
||||
|
||||
**Expected**
|
||||
|
||||
A full accounting of root usage and identification of: (a) any unbounded
|
||||
container log and its root-cause flag; (b) reclaimable apt cache; (c) journal
|
||||
size and journald limits; (d) unused images (0 dangling expected) and unused
|
||||
anonymous volumes.
|
||||
|
||||
**Decision**
|
||||
|
||||
- If root is ≥70% used or any container log is unbounded → Step 2.
|
||||
- If root is healthy and logs are bounded → STOP (no change needed; record evidence).
|
||||
- If state conflicts with expectations (e.g. missing sudo, unexpected mount) → STOP.
|
||||
|
||||
### Step 2 — Fix noisy container logging at the source, then truncate
|
||||
|
||||
**Action**
|
||||
|
||||
1. Back up the app config: `sudo cp <config> <config>.bak-YYYYMMDD-<issue>`.
|
||||
2. Disable the debug/verbose flag (e.g. `log.verbose: true → false` in the AGH YAML).
|
||||
3. Apply config with a container restart: `docker restart <container>` (config-level
|
||||
change; **no recreate** needed and daemon.json rotation would not apply anyway).
|
||||
4. Truncate the accumulated logs: `sudo truncate -s 0 <json.log>` for the chatty
|
||||
container(s) (and any other oversized ones, e.g. traefik).
|
||||
5. Record `df -h` before/after.
|
||||
|
||||
**Expected**
|
||||
|
||||
`docker logs <container>` no longer shows the `[debug]` flood; the `*-json.log`
|
||||
stops growing; several GB reclaimed.
|
||||
|
||||
**Verification**
|
||||
|
||||
- `sudo tail -c 200000 <json.log>` after ≥1 minute → no new debug lines.
|
||||
- `df -h` improvement recorded.
|
||||
- Container still `Up (healthy)`.
|
||||
|
||||
**Rollback**
|
||||
|
||||
- Trigger: log volume unchanged, service degraded, or debug output is actually needed.
|
||||
- Action: restore the config backup and `docker restart <container>`.
|
||||
- Verify: original verbose behaviour back; container healthy.
|
||||
|
||||
### Step 3 — Clear apt cache and cap journald
|
||||
|
||||
**Action**
|
||||
|
||||
1. `sudo apt-get clean` (clears only `/var/cache/apt/archives`; `/var/lib/apt/lists`
|
||||
is not cleared by it and regenerates on `apt update` — optional/low value, skip).
|
||||
2. `sudo journalctl --vacuum-size=100M`.
|
||||
3. Write drop-in `/etc/systemd/journald.conf.d/00-disk-<issue>.conf`:
|
||||
`[Journal]` + `SystemMaxUse=200M`.
|
||||
4. `sudo systemctl restart systemd-journald` (approved service touch).
|
||||
5. Record `df -h` before/after.
|
||||
|
||||
**Expected**
|
||||
|
||||
Archives cleared (~1.4G on hk2), journal ≤100M, future journal capped at 200M.
|
||||
|
||||
**Verification**
|
||||
|
||||
- `du -sh /var/cache/apt/archives` → ~0.
|
||||
- `journalctl --disk-usage` → ≤100M.
|
||||
- `systemctl show systemd-journald -p ...` or restart log confirms new limit;
|
||||
`journalctl -b` still readable.
|
||||
|
||||
**Rollback**
|
||||
|
||||
- Trigger: journald fails to start or logs lost unexpectedly.
|
||||
- Action: remove the drop-in, `sudo systemctl restart systemd-journald`.
|
||||
- Verify: journald active, prior journal entries still listed.
|
||||
|
||||
### Step 4 — Remove unused Docker images (explicit keep-list)
|
||||
|
||||
**Action**
|
||||
|
||||
1. Enumerate unused images: `docker image ls` cross-checked against the images of
|
||||
running containers (`docker ps --format '{{.Image}}'`). Re-enumerate at
|
||||
execution time — the list drifts.
|
||||
2. Present the exact removal list to the operator; keep the agreed rollback pin(s)
|
||||
(e.g. `powerdns/pdns-auth-50:5.0.5`) and delete the rest explicitly:
|
||||
`docker rmi <repo:tag> ...` (per image).
|
||||
3. Record `df -h` before/after.
|
||||
|
||||
**Expected**
|
||||
|
||||
Only in-use images + kept pins remain; ~1–2.5G reclaimed (reclaim is an upper
|
||||
bound — layers shared with kept images are not freed; measure with `df`, do not
|
||||
promise the estimate).
|
||||
|
||||
**Verification**
|
||||
|
||||
- `docker image ls` shows only the expected set.
|
||||
- `docker system df` images reclaimable ≈ 0 for the removed set.
|
||||
- All containers still `Up`.
|
||||
|
||||
**Rollback**
|
||||
|
||||
- Trigger: an image that was actually needed was removed.
|
||||
- Action: re-pull it from the registry (`docker pull <repo:tag>`); if a kept pin
|
||||
must change, update the compose pin and `up -d`.
|
||||
- Verify: image present; affected service healthy.
|
||||
|
||||
### Step 5 — Remove unused anonymous volumes and build cache
|
||||
|
||||
**Action**
|
||||
|
||||
1. Enumerate volumes: `docker volume ls`, and confirm which are referenced by
|
||||
containers (`docker inspect` Mounts). Expected targets: anonymous volumes with
|
||||
no container reference.
|
||||
2. `docker volume prune` (**no `-a`**) — engine ≥ v23 removes only unused
|
||||
anonymous volumes; in-use volumes (e.g. PG data) are protected by container
|
||||
references in every version.
|
||||
3. `docker builder prune -f`.
|
||||
4. Record `df -h` before/after.
|
||||
|
||||
**Expected**
|
||||
|
||||
Unused anonymous volumes (~1.2G on hk2) and build cache gone; in-use volumes intact.
|
||||
|
||||
**Verification**
|
||||
|
||||
- `docker volume ls` shows only in-use volumes.
|
||||
- Services that own volumes (e.g. postgres) report healthy and data present.
|
||||
- `df -h` improvement recorded.
|
||||
|
||||
**Rollback**
|
||||
|
||||
- Trigger: data loss suspected in a removed volume.
|
||||
- Action: restore from backup if the volume ever contained data; verify against
|
||||
the pre-prune enumeration (targets must be anonymous + unreferenced before prune).
|
||||
- Note: this is why target enumeration is recorded before pruning.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Troubleshooting A — Log still grows after disabling verbose
|
||||
|
||||
- Evidence: `sudo tail -c 200000 <json.log>` still shows new lines; app config re-checked.
|
||||
- Allowed actions: check for a second verbose source (container entrypoint flags,
|
||||
other apps in the same log); check `docker inspect <c> --format '{{.HostConfig.LogConfig}}'`.
|
||||
- Next step: back to Step 2 or STOP if a container-level log-opts change (recreate)
|
||||
would be needed — that is a separate approval.
|
||||
|
||||
### Troubleshooting B — `docker rmi` fails (image in use)
|
||||
|
||||
- Evidence: `image is being used by stopped container ...`.
|
||||
- Allowed actions: identify the stopped container (`docker ps -a`); confirm it is
|
||||
not needed; remove it only with explicit approval.
|
||||
- Next step: re-run rmi for the remaining images; never force-delete blindly.
|
||||
|
||||
### Troubleshooting C — `docker volume prune` would remove more than expected
|
||||
|
||||
- Evidence: prune dry-run/listing includes a named or referenced volume.
|
||||
- Allowed actions: abort; do not add `-a`; re-check references.
|
||||
- Next step: STOP and report to the operator with the enumeration.
|
||||
|
||||
## Final Verification
|
||||
|
||||
The flow is successful only when all of the following hold:
|
||||
|
||||
- `df -h` root usage is in the agreed band (VPS-81: 76% → ≤55% used; measure, do not assume).
|
||||
- `docker system df` shows reclaimable ≈ 0 for images/volumes targeted.
|
||||
- All containers `Up` (health checks pass); public services verified
|
||||
(`dig @<host-ip> SOA <zone>` for DNS hosts; service URLs reachable).
|
||||
- Tracking issue updated with before/after `df`, actions, and the one-week
|
||||
observation checkpoint for log growth.
|
||||
|
||||
## Failure Handling
|
||||
|
||||
If the flow cannot complete:
|
||||
|
||||
1. Stop further mutation.
|
||||
2. Collect command output, timestamps, and the exact step that failed.
|
||||
3. Record completed steps, actual results, unmet expectations, and whether a
|
||||
rollback ran.
|
||||
4. Hand over per the tracking issue with evidence; do not guess further.
|
||||
|
||||
## References
|
||||
|
||||
- Plane `vps` issue VPS-81 "hk2: 释放根盘空间" (+ subtasks VPS-82…88) — plan, review findings, approvals.
|
||||
- [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) — host facts.
|
||||
- [RUNBOOKS.md](../RUNBOOKS.md) — runbook spec; [runbooks/README.md](README.md) — index.
|
||||
@@ -0,0 +1,93 @@
|
||||
# Runbook: issue → mergeable change
|
||||
|
||||
## Purpose
|
||||
|
||||
Turn an approved Linear `vps` issue into a reviewed, mergeable change in this
|
||||
repo (docs, runbooks, hosts facts, or Ansible playbooks).
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: repo content under `docs/`, `runbooks/`, `hosts/`, `inventory/`, `ansible/`.
|
||||
- Not applicable: mutating production state directly — that goes through `release.md` / `ansible-operations.md`.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: Linear MCP (`vps` project), git
|
||||
|
||||
## Inputs
|
||||
|
||||
| Input | Source | Required | Validation |
|
||||
|---|---|---:|---|
|
||||
| Issue identifier | Linear (`vps` project) | Yes | `linear_get_issue <id>` returns a description |
|
||||
| Current repo state | `git status` / `git log` | Yes | Clean or intended worktree |
|
||||
|
||||
## Safety
|
||||
|
||||
- Scope is locked to the issue: do not bundle unrelated changes.
|
||||
- Never commit secrets (see `AGENTS.md` §Safety).
|
||||
- Verify every change; do not merge a change whose verification was skipped.
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Read the issue
|
||||
|
||||
**Action** — `linear_get_issue <id>`, read description and acceptance criteria.
|
||||
|
||||
**Expected** — clear scope, action, and verification for the change.
|
||||
|
||||
**Decision** — if the issue is ambiguous or lacks verification criteria, `STOP`
|
||||
and ask for clarification (add a `needs-info` label if applicable). Otherwise go to Step 2.
|
||||
|
||||
### Step 2 — Inspect and change
|
||||
|
||||
**Action** — read the relevant files, then make the minimal change the issue asks for.
|
||||
|
||||
**Expected** — diff is scoped to the issue.
|
||||
|
||||
**Decision** — if the change needs production mutation, `STOP` and route to
|
||||
`release.md`. Otherwise go to Step 3.
|
||||
|
||||
### Step 3 — Verify
|
||||
|
||||
**Action** — run `scripts/validate-repo.sh` from the repo root (covers secret
|
||||
scan, inventory cross-check, markdown link check, runbook-spec check, and
|
||||
Ansible `--syntax-check`); for changes that alter playbook behavior, also run
|
||||
a read-only `ansible-playbook --check` where possible.
|
||||
|
||||
**Verification** — `scripts/validate-repo.sh` exits 0; the concrete checks
|
||||
must match the change type.
|
||||
|
||||
**Decision** — verification passed → Step 4; failed → Troubleshooting A.
|
||||
|
||||
### Step 4 — Commit and link
|
||||
|
||||
**Action** — commit with a message containing the full issue ID (e.g. `W1N-123: …`); open a PR if the change is substantial; link the issue via `linear_save_comment`.
|
||||
|
||||
**Verification** — `git log -1` shows the issue ID; the issue has the commit/PR pointer.
|
||||
|
||||
**Rollback** — `git revert <sha>` or `git checkout <branch>` to drop the change; re-verify after.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Troubleshooting A — Verification failed
|
||||
|
||||
- Evidence: command output, failing check.
|
||||
- Allowed: fix the change within scope; re-run verification.
|
||||
- Next: still failing → `STOP` and report in the issue.
|
||||
|
||||
## Final Verification
|
||||
|
||||
- Change matches the issue scope.
|
||||
- Verification passed and the issue is updated with evidence.
|
||||
|
||||
## Failure Handling
|
||||
|
||||
If unfinished: stop, collect the failed check output, record completed steps, and
|
||||
hand back to the issue — do not guess.
|
||||
|
||||
## References
|
||||
|
||||
- [`docs/agents/issue-tracker.md`](../docs/agents/issue-tracker.md)
|
||||
- [`RUNBOOKS.md`](../RUNBOOKS.md)
|
||||
@@ -1,5 +1,20 @@
|
||||
# Runbook: mailcow health (mx2)
|
||||
|
||||
## Purpose
|
||||
|
||||
Read-only health check of the mailcow stack on mx2.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
|
||||
- Read-only: does not change mailcow configuration or service state.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: mx2.windy.me (`/opt/mail`)
|
||||
|
||||
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
|
||||
Path: `/opt/mail`
|
||||
Prefer: the Ansible health report (`ansible/playbooks/health-report.yml --limit mailcow`),
|
||||
@@ -71,6 +86,11 @@ The sanitized Ansible health profile is `mailcow` (`ansible/playbooks/healthchec
|
||||
The server-local timer emits a sanitized result at `/var/lib/vps-health/latest.json`.
|
||||
It does not change Mailcow configuration or service state.
|
||||
|
||||
## Safety
|
||||
|
||||
- Read-only: never mutate configuration or service state during this check.
|
||||
- If live state conflicts with an expected value below, `STOP` and report; do not "fix" on the fly.
|
||||
|
||||
## Pass criteria
|
||||
|
||||
- Compose stack up; watchdog ~100%
|
||||
|
||||
@@ -1,5 +1,20 @@
|
||||
# Runbook: use mailcow SMTP / IMAP (client)
|
||||
|
||||
## Purpose
|
||||
|
||||
Reference for configuring mail clients against the mailcow SMTP/IMAP endpoints.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md) client submission (587/465) and IMAP/POP (993/995).
|
||||
- Not applicable: server-side mailcow configuration or administration.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: mx2.windy.me (SMTP/IMAP client endpoints)
|
||||
|
||||
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
|
||||
Prerequisite: a mailbox on `windy.me` (password from mailcow UI, not the admin account unless it is that mailbox).
|
||||
|
||||
@@ -52,3 +67,9 @@ Do not commit or paste real passwords into this repo.
|
||||
- Port/TLS mode mismatch (587 vs 465)
|
||||
- Account active in mailcow; not rate-limited / fail2banned after bad attempts
|
||||
- Apps that store SMTP in their own config (e.g. Vaultwarden `config.json`) may keep a **stale** password even when `.env` is correct — verify AUTH against the effective config ([vaultwarden-health](vaultwarden-health.md) §5)
|
||||
|
||||
## Safety
|
||||
|
||||
- Do not commit or paste real passwords into this repo or chat.
|
||||
- Use submission (587/465) for client sending; never use port 25 as a desktop/app outbound port.
|
||||
- If live state conflicts with the endpoint values above, `STOP` and report; do not change server-side settings during this reference check.
|
||||
|
||||
@@ -1,9 +1,37 @@
|
||||
# Runbook: mailcow update (mx2)
|
||||
|
||||
## Purpose
|
||||
|
||||
Update the mailcow stack on mx2 to the latest supported release.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
|
||||
- Not applicable: config changes beyond the update, DB migration, secret rotation.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: mx2.windy.me (`/opt/mail`)
|
||||
|
||||
## Approval gates
|
||||
|
||||
| Action | Risk | Explicit approval |
|
||||
|---|---|---|
|
||||
| Run `./update.sh` (recreates containers, brief mail interruption) | Medium | Yes — user confirmation required |
|
||||
|
||||
Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
|
||||
Path: `/opt/mail`
|
||||
**Confirm with the user before running an update.**
|
||||
|
||||
## Safety
|
||||
|
||||
- Never run the update without explicit user confirmation.
|
||||
- Never pass secrets into the chat log; do not commit `mailcow.conf`.
|
||||
- If a step fails, capture `docker compose ps` and logs and stop before further changes.
|
||||
- If live state conflicts with this runbook's assumptions (e.g. unexpected `mailcow.conf` values), `STOP` and report.
|
||||
|
||||
## Before
|
||||
|
||||
1. Run [mailcow-health](mailcow-health.md) (Ansible health report). Record baseline.
|
||||
|
||||
@@ -0,0 +1,171 @@
|
||||
# matrix_e2ee update (hass.windy.lan)
|
||||
|
||||
## Purpose
|
||||
|
||||
Update the custom **`matrix_e2ee`** integration on `hass.windy.lan` while
|
||||
preserving a verified rollback point and confirming that Home Assistant loads it.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: deploying a reviewed `matrix_e2ee` source revision to
|
||||
`hass.windy.lan`.
|
||||
- Not applicable: Home Assistant Core upgrades, integration configuration
|
||||
changes, or recovery without a usable live-tree backup.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-23
|
||||
- Related systems: hass.windy.lan (HAOS, `machine: green`)
|
||||
|
||||
## Approval gates
|
||||
|
||||
| Action | Risk | Explicit approval |
|
||||
|---|---|---|
|
||||
| Replace the live integration tree and restart Home Assistant Core | Medium | Yes — user confirmation required |
|
||||
|
||||
## Safety
|
||||
|
||||
- Do not replace the live tree or restart Core without explicit user confirmation.
|
||||
- If any precondition or verification fails, `STOP` and record evidence before continuing.
|
||||
|
||||
## Preconditions
|
||||
|
||||
- The source repo at `/home/windy/project/ha-matrix-e2ee` is on the **target
|
||||
state**: either a release tag (`git tag -l 'v*'`) or a commit whose
|
||||
`manifest.json` `version` is the target. Note v0.3.0 was deployed from an
|
||||
**untagged** `main` HEAD (`216cc99`), so the tag check alone is not enough —
|
||||
confirm the working-tree `custom_components/matrix_e2ee/manifest.json`.
|
||||
- The working tree matches HEAD: `git status --short` clean (only ignorables)
|
||||
and `git diff HEAD -- custom_components/` empty. Record
|
||||
`git rev-parse HEAD` for the docs/Linear record — HEAD can move during a
|
||||
session, so re-check right before rsync (verified 2026-08-18: HEAD moved
|
||||
from a `w1n-180` branch merge to `main` mid-deploy).
|
||||
- The remote host is reachable and `sudo -n -i ha core info` succeeds.
|
||||
- The workstation HTTP proxy does not interfere — LAN hosts must be reachable
|
||||
without proxying (unset `http_proxy` / `HTTP_PROXY` if needed).
|
||||
- If the agent sandbox hits `Bad owner or permissions on /etc/ssh/ssh_config.d/20-systemd-ssh-proxy.conf`, add `-F /dev/null` to the `ssh` / `rsync` commands below.
|
||||
- Domain is **`matrix_e2ee`** (double-e). Older notes may say `matrix_e2e`;
|
||||
paths, events, and services all use `matrix_e2ee`.
|
||||
|
||||
## Procedure
|
||||
|
||||
### 1. Backup the live tree
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i mkdir -p /homeassistant/.matrix-e2ee-backups &&
|
||||
sudo -n -i cp -a /homeassistant/custom_components/matrix_e2ee \
|
||||
/homeassistant/.matrix-e2ee-backups/matrix_e2ee.bak-$(date +%Y%m%d)-v<OLD_VERSION>'
|
||||
```
|
||||
|
||||
The backup lives in `/homeassistant/.matrix-e2ee-backups/` — a directory
|
||||
separated from `custom_components/` to avoid HA scanning it as a custom
|
||||
component domain.
|
||||
|
||||
### 2. Rsync the new source
|
||||
|
||||
```bash
|
||||
rsync -a --delete -e 'ssh -o BatchMode=yes' \
|
||||
/home/windy/project/ha-matrix-e2ee/custom_components/matrix_e2ee/ \
|
||||
hassio@hass.windy.lan:/homeassistant/custom_components/matrix_e2ee/
|
||||
```
|
||||
|
||||
The `--delete` cannot remove Core-owned `__pycache__` — that is handled
|
||||
in the next step. Source `.py` files and `manifest.json` are transferred
|
||||
correctly even with the `__pycache__` errors, but **rsync exits with code 23
|
||||
(`some files/attrs were not transferred`)** — that is expected, not a failure.
|
||||
Confirm the transfer by checking the manifest on the host before restarting.
|
||||
|
||||
### 3. Wipe `__pycache__` (as root) and restart Core
|
||||
|
||||
Quote the nested `__pycache__` glob — remote login shell is zsh and will
|
||||
fail with `no matches found` if left unquoted.
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
"sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee/__pycache__ \
|
||||
'/homeassistant/custom_components/matrix_e2ee/*/__pycache__' &&
|
||||
sudo -n -i ha core restart"
|
||||
```
|
||||
|
||||
Stale `cpython-314` bytecode in Core-owned `__pycache__` keeps the old
|
||||
coordinator in memory until restart. Wipe before restart.
|
||||
|
||||
Wait for `Command completed successfully.` (typically 1–2 min).
|
||||
|
||||
### 4. Verify the deployment
|
||||
|
||||
#### 4a. Confirm manifest version
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i cat /homeassistant/custom_components/matrix_e2ee/manifest.json'
|
||||
```
|
||||
|
||||
Expect `"version": "<NEW_VERSION>"`.
|
||||
|
||||
#### 4b. Check Core logs for matrix_e2ee
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i ha core logs -n 2500' | grep -E 'matrix_e2ee|Setting up matrix' | head -20
|
||||
```
|
||||
|
||||
Expect:
|
||||
- `Setup of domain matrix_e2ee took ...` (older wording `Setting up matrix_e2ee` may appear)
|
||||
- `matrix_e2ee restored existing device; user=@hass:chans.xyz device=rO1R915ncu`
|
||||
- No `ERROR` level messages from `custom_components.matrix_e2ee`
|
||||
- Blocking-call WARNINGs from `_patch_nio_sas_timeout` / nio store I/O are expected
|
||||
|
||||
#### 4c. Verify the entry is loaded (optional, via Supervisor API)
|
||||
|
||||
No trailing slash on the entries URL (trailing `/` returns 404 on Core 2026.8.1).
|
||||
|
||||
```bash
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
"sudo -n -i python3 - <<'PY'
|
||||
import os, json, urllib.request
|
||||
req = urllib.request.Request(
|
||||
'http://supervisor/core/api/config/config_entries/entry',
|
||||
headers={'Authorization': 'Bearer ' + os.environ['SUPERVISOR_TOKEN']},
|
||||
)
|
||||
entries = json.loads(urllib.request.urlopen(req, timeout=30).read())
|
||||
for e in entries:
|
||||
if e['domain'] == 'matrix_e2ee':
|
||||
print(f\"{e['domain']}: state={e['state']} source={e['source']}\")
|
||||
PY"
|
||||
```
|
||||
|
||||
Expect `state: loaded`.
|
||||
|
||||
### 5. Record the deployment
|
||||
|
||||
- Update the `matrix_e2ee` live-tree section in
|
||||
[hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md): new version, source
|
||||
commit (`git rev-parse HEAD`), backup name, and any new feature notes.
|
||||
- Record the operation in the Linear `vps` project (scope, action,
|
||||
verification, follow-up); see [docs/agents/issue-tracker.md](../docs/agents/issue-tracker.md).
|
||||
|
||||
## Rollback
|
||||
|
||||
If Core fails to start after the update:
|
||||
|
||||
```bash
|
||||
# Restore the backup
|
||||
ssh -o BatchMode=yes hassio@hass.windy.lan \
|
||||
'sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee &&
|
||||
sudo -n -i cp -a /homeassistant/.matrix-e2ee-backups/matrix_e2ee.bak-<DATE>-v<OLD_VERSION> \
|
||||
/homeassistant/custom_components/matrix_e2ee &&
|
||||
sudo -n -i rm -rf /homeassistant/custom_components/matrix_e2ee/__pycache__ &&
|
||||
sudo -n -i ha core restart'
|
||||
```
|
||||
|
||||
If a full HA backup exists (pre-update), restore via `ha backups restore <slug>`.
|
||||
|
||||
## References
|
||||
|
||||
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — current live version and config
|
||||
- [docs/home-assistant-matrix.md](../docs/home-assistant-matrix.md) — integration architecture and verification model
|
||||
- [home-assistant-maintenance.md](home-assistant-maintenance.md) — general HA maintenance procedures
|
||||
- [ha-matrix-e2ee source](https://github.com/windyboy/ha-matrix-e2ee) — GitHub repo
|
||||
@@ -1,5 +1,20 @@
|
||||
# Matrix Health Check
|
||||
|
||||
## Purpose
|
||||
|
||||
Read-only health check of the Matrix homeserver (ESS on K3s).
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: [synapse.chans.xyz](../hosts/synapse.chans.xyz.md), namespace `ess`.
|
||||
- Read-only: does not change pods, ingress, certificates, or configuration.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: synapse.chans.xyz (ESS chart `26.7.2`, K3s)
|
||||
|
||||
Monitor the Matrix homeserver running on `synapse.chans.xyz` (ESS chart `26.7.2`, K3s node).
|
||||
|
||||
Prefer `cd ansible && ansible-playbook playbooks/health-report.yml --limit matrix`
|
||||
@@ -90,3 +105,9 @@ Backup automation is currently paused. `/var/backups/matrix/` is retained for a
|
||||
| Well-known returns 404/redirect | Root `chans.xyz` ingress missing or misconfigured |
|
||||
| 502 Bad Gateway | Synapse pod restarting or DB down |
|
||||
| SMTP emails not sent | MAS SMTP config incomplete; TCP reachable but AUTH failing — see `runbooks/vaultwarden-health.md` |
|
||||
|
||||
## Safety
|
||||
|
||||
- Read-only: never mutate pods, ingress, certificates, or configuration during this check.
|
||||
- Backup automation is paused; do not treat `/var/backups/matrix/` as a recovery source.
|
||||
- If live state conflicts with an expected value, `STOP` and report.
|
||||
|
||||
@@ -0,0 +1,312 @@
|
||||
# Matter packet capture (read-only)
|
||||
|
||||
## Purpose
|
||||
|
||||
Capture Matter-related traffic on LAN55 (mDNS discovery + PASE/CASE commissioning +
|
||||
operational traffic) to determine whether a device is on the network, is in
|
||||
commissioning mode, and whether the commissioning handshake completes. Capture
|
||||
is read-only and changes no device or network state.
|
||||
|
||||
## Scope
|
||||
|
||||
- Environment: LAN55 (`hass.windy.lan`, Aqara M3, ESP32-C2 Matter bulbs,
|
||||
phone / HA matter-server all on the 55 subnet).
|
||||
- Subject: Matter over Wi-Fi and Thread relay nodes. The Thread 802.15.4 air
|
||||
side itself is not capturable — only IPv6 forwarding by a Thread relay such
|
||||
as the M3 is visible.
|
||||
- Not applicable: BLE commissioning, Thread 802.15.4 frames, cross-subnet
|
||||
multicast (66-subnet hosts cannot see the 55 subnet's mDNS — link-local
|
||||
multicast does not cross the routed 55/66 boundary, there is no reflector).
|
||||
- Read-only: no AP/device/network config is modified; state returns to normal
|
||||
when tcpdump exits.
|
||||
|
||||
### Capture-point selection
|
||||
|
||||
Matter commissioning is a two-party conversation and the commissioner
|
||||
participates in every message of it, so capturing on the commissioner host
|
||||
equals capturing the whole flow.
|
||||
|
||||
| Capture point | Sees | Blind spot | Notes |
|
||||
|---|---|---|---|
|
||||
| **hass `end0` — commissioner side (recommended)** | The full HA-driven commissioning conversation: all mDNS queries/announcements (segment multicast) + the complete TCP 5540 PASE/CASE session | Phone-as-commissioner flows (the phone's session to the device does not pass through hass) | `core_matter_server` uses **host networking**, so tcpdump on `end0` sees the add-on's traffic directly; `/` is overlay with ~42 GB free — no 60 MB tmpfs rotation needed |
|
||||
| **UAP-AC-Lite `br0` (192.168.55.5)** | All mDNS multicast (flooded; igmp snooping off) + all wireless-client unicast + unicast to/from the AP | Wired↔wired unicast — e.g. HA↔M3 TCP 5540 while a Thread device commissions via the M3 (wired, observed) — is switched locally and never traverses the AP | AP `/tmp` is a ~60 MB tmpfs → rotating capture is **mandatory** |
|
||||
|
||||
For the common "add device" case with HA matter-server as the commissioner,
|
||||
capture on hass `end0`. Use the AP `br0` point for wireless-device or
|
||||
phone-driven flows (a wireless client's unicast to/from its AP is only visible
|
||||
there).
|
||||
|
||||
A third point, `gw` `switch0`, is **verified as a limited capture point**
|
||||
(cross-subnet/gateway/mDNS flows only — not a full mirror of LAN55) — see
|
||||
[Capture point: gw switch0](#capture-point-gw-switch0).
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-22
|
||||
- Related systems: UAP-AC-Lite AP `192.168.55.5` (br0), `core_matter_server` on
|
||||
`hass.windy.lan` (`end0`), Aqara M3, ESP32-C2 Matter bulbs
|
||||
|
||||
## Preconditions
|
||||
|
||||
- SSH to the capture point:
|
||||
- AP: `ssh zhiqiangf@192.168.55.5` (key-only, `BatchMode=yes` verified).
|
||||
- hass: `ssh hassio@hass.windy.lan`. Non-interactive SSH does **not** source
|
||||
`.zprofile`, so run tcpdump as `sudo -n -i tcpdump …` (verified 2026-08-22).
|
||||
- tcpdump available:
|
||||
- AP: full 4.9.2 / libpcap 1.8.1 (verified 2026-08-22).
|
||||
- hass: `/usr/bin/tcpdump` via `sudo -n -i` (verified 2026-08-22).
|
||||
- Trigger source ready: put the Matter device into commissioning mode, or have
|
||||
HA/phone perform discovery/commissioning — otherwise no relevant packets.
|
||||
- AP `/tmp` is a ~60 MB tmpfs (61.3 M total, 60.4 M free): rotating capture
|
||||
(`-C`/`-W`) is mandatory on the AP. hass `/` is overlay — rotation optional
|
||||
but keep the habit for long captures.
|
||||
|
||||
## Safety
|
||||
|
||||
### Non-negotiable rules
|
||||
|
||||
- Read-only diagnosis: no installs, config changes, or service restarts on the
|
||||
AP, hass, devices, or network.
|
||||
- pcap files are limited to `/tmp`; pull them off and delete them afterwards
|
||||
(mandatory on the AP; same hygiene on hass).
|
||||
- Never write captured content (including any plaintext key material) into this
|
||||
repository or Linear.
|
||||
|
||||
### Stop conditions
|
||||
|
||||
- Capture point unreachable (ssh fails) → `STOP`, fix the network first.
|
||||
- tcpdump reports "Permission denied" or cannot listen → `STOP` (admin needed;
|
||||
on hass verify `sudo -n -i` works).
|
||||
- Filter expression syntax error → `STOP`, use only expressions verified in
|
||||
this document.
|
||||
- AP `/tmp` nearly full (rotation file count × single-file size ≈ 60 MB) →
|
||||
`STOP` and clean old pcaps.
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Choose the capture point
|
||||
|
||||
**Action**
|
||||
|
||||
- HA matter-server is the commissioner (the "add device" case) → hass `end0`.
|
||||
- Wireless device or phone-driven flow → AP `br0`.
|
||||
|
||||
**Expected**
|
||||
|
||||
- The chosen point is reachable and tcpdump starts listening.
|
||||
|
||||
**Decision**
|
||||
|
||||
- Capture point chosen and reachable → Step 2.
|
||||
- Neither applies or the choice is unclear → `STOP` and record why.
|
||||
|
||||
### Step 2 — Realtime observation (quick confirm traffic appears)
|
||||
|
||||
**Action**
|
||||
|
||||
AP:
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
hass (commissioner side):
|
||||
|
||||
```bash
|
||||
ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 -s 0 -tt 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
Keep the window open, trigger the device behavior (enter commissioning mode /
|
||||
start commissioning / send a command), `Ctrl+C` to stop.
|
||||
|
||||
**Expected**
|
||||
|
||||
- `_matterc._udp` / `_matter._tcp` mDNS announcements (UDP 5353, multicast
|
||||
`224.0.0.251` / `ff02::fb`).
|
||||
- During commissioning: TCP **5540** (PASE/CASE) SYN/SYN-ACK between the device
|
||||
IP and HA/M3.
|
||||
- If the target device's MAC is known, add `and ether host <mac>` to keep only
|
||||
that device (see variants).
|
||||
- `5552` is not a standard Matter port; it is an observed port for the Aqara M3
|
||||
Thread-relay node (see `docs/matter-pairing-troubleshoot.md`).
|
||||
|
||||
**Decision**
|
||||
|
||||
- Expected packets present → Step 3 to save evidence, or judge directly against
|
||||
the stage table (`docs/matter-pairing-troubleshoot.md` §4).
|
||||
- No packets at all → `STOP`: fix device online / commissioning-mode first; the
|
||||
network side is repeatedly verified healthy (see troubleshooting doc).
|
||||
- mDNS present but no 5540 → see troubleshooting doc decision tree, item 4
|
||||
(§5).
|
||||
|
||||
### Step 3 — Rotating capture + pull to WSL
|
||||
|
||||
**Action** (`-C 5` = rotate every 5 MB, `-W 12` = max 12 files, ≈ 60 MB ≤ AP tmpfs)
|
||||
|
||||
AP:
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -C 5 -W 12 -w /tmp/matter.pcap 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
hass (rotation optional — overlay disk):
|
||||
|
||||
```bash
|
||||
ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 -s 0 -C 5 -W 12 -w /tmp/matter.pcap 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
Trigger the traffic, then `Ctrl+C`. Files are `/tmp/matter.pcap`,
|
||||
`/tmp/matter.pcap1`, …
|
||||
|
||||
**Expected**
|
||||
|
||||
- tcpdump prints capture statistics (`N packets captured`).
|
||||
- `ls -la /tmp/matter.pcap*` shows the files; total stays < 60 MB on the AP.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "ls -la /tmp/matter.pcap*"
|
||||
# or
|
||||
ssh hassio@hass.windy.lan "ls -la /tmp/matter.pcap*"
|
||||
```
|
||||
|
||||
**Pull to WSL for analysis and clean up afterwards**
|
||||
|
||||
```bash
|
||||
scp zhiqiangf@192.168.55.5:/tmp/matter.pcap* .
|
||||
# or
|
||||
scp hassio@hass.windy.lan:/tmp/matter.pcap* .
|
||||
|
||||
# clean up on the capture point
|
||||
ssh zhiqiangf@192.168.55.5 "rm -f /tmp/matter.pcap*"
|
||||
ssh hassio@hass.windy.lan "sudo -n -i rm -f /tmp/matter.pcap*"
|
||||
```
|
||||
|
||||
### Step 4 — Wireshark analysis (optional)
|
||||
|
||||
**Action**
|
||||
|
||||
Open the pcap in Wireshark. mDNS (UDP 5353) is plaintext and directly
|
||||
readable; Matter payloads on TCP/UDP 5540 show only the handshake by default —
|
||||
plaintext needs the dissector plus session keys (Step 5).
|
||||
|
||||
**Expected**
|
||||
|
||||
- `mDNS` filter shows all discovery records; `tcp.port==5540` shows the
|
||||
commissioning handshake.
|
||||
|
||||
### Step 5 — Decrypt Matter plaintext (optional, needs session keys)
|
||||
|
||||
Matter payloads are encrypted (AES-CCM); mDNS plaintext contains no keys. To
|
||||
decrypt, one of:
|
||||
|
||||
1. **Capture-side key leak with a chip tool (most common)**: the commissioner
|
||||
(HA matter-server / chip-tool) prints or exports session keys during
|
||||
commissioning; enter them in Wireshark → Preferences → Protocols → Matter.
|
||||
See [matter-dissector README](https://github.com/project-chip/matter-dissector#security-features).
|
||||
2. **well-known CASE keys**: both sides compiled with
|
||||
`MATTER_CONFIG_SECURITY_TEST_MODE` / `CASEUseKnownECDHKey`; not enabled in
|
||||
this environment (ESP32-C2 + HA official matter-server).
|
||||
|
||||
**Expected**
|
||||
|
||||
- Matter dissector expands protocol headers, IM commands, and cluster content.
|
||||
|
||||
**Stop condition (decryption)**: with no session keys or test keys obtainable,
|
||||
do not fabricate keys to force a decrypt — plaintext mDNS + TCP handshake
|
||||
still resolves most troubleshooting; for plaintext payloads, upgrade to
|
||||
exporting keys on the commissioner side, then return to this runbook.
|
||||
|
||||
## Targeted capture variants
|
||||
|
||||
### One device only (known MAC)
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -tt 'ether host 34:98:7a:27:7f:08 and (udp port 5353 or tcp port 5540 or tcp port 5552)'"
|
||||
```
|
||||
|
||||
MACs from `docs/matter-pairing-troubleshoot.md` §3 (working bulb
|
||||
`34:98:7a:25:a1:f0`, broken bulb `34:98:7a:27:7f:08`).
|
||||
|
||||
### mDNS announcements only (no 5540 noise)
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -tt 'udp port 5353'"
|
||||
```
|
||||
|
||||
### Rotating capture with timestamped filename (multiple runs)
|
||||
|
||||
```bash
|
||||
ssh zhiqiangf@192.168.55.5 "tcpdump -ni br0 -s 0 -C 5 -W 12 -w /tmp/matter-\$(date +%H%M%S).pcap 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
> The nested `$(date)` must be escaped as `\$(...)` inside the double-quoted ssh
|
||||
> command so the remote shell expands it.
|
||||
|
||||
For the hass point, prefix the same commands with
|
||||
`ssh hassio@hass.windy.lan "sudo -n -i tcpdump -ni end0 …"`.
|
||||
|
||||
## Capture point: gw switch0
|
||||
|
||||
**Status: verified 2026-08-22 — limited capture point; NOT a full mirror of
|
||||
LAN55.**
|
||||
|
||||
- `gw` `switch0` (`eth1`–`eth3`, `192.168.55.254/24`) is LAN55's L2 aggregation
|
||||
only while devices plug directly into the ER-X. EdgeOS ships tcpdump;
|
||||
`tcpdump -ni switch0` follows Linux bridge semantics.
|
||||
- **Live topology (verified 2026-08-22): the SE5420 core switch is deployed**
|
||||
(management `192.168.66.253` up — TP-Link OUI `f8:c9:03`, web UI on
|
||||
:80/:443) and the ER-X uplink is a **single switch0 member port**: `eth1`
|
||||
link up, `eth2`/`eth3` down. All LAN55 wired devices (hass `.11`, Aqara M3
|
||||
`.248`, SmartThings `.48`, UAP-AC-Lite `.5`) are reached via `switch0`
|
||||
behind that one uplink. Same-segment wired↔wired unicast switches locally on
|
||||
the SE5420 and never reaches `switch0`.
|
||||
- **What `switch0` still sees:** cross-subnet (66↔55) unicast, traffic to/from
|
||||
the gateway itself (DHCP, DNS forwarding, port-forwards), and LAN55 mDNS
|
||||
multicast (flooded up the uplink). Use it only for those flows; for a full
|
||||
commissioning conversation use the hass `end0` or AP `br0` point instead.
|
||||
- **Full mirror:** only via SE5420 port mirroring (the switch cannot run
|
||||
tcpdump). Not configured; out of scope here.
|
||||
- **Verification commands (EdgeOS v3.0.1 build 5862409):**
|
||||
- Interactive: `ssh ubnt@192.168.66.254` (or `zhiqiang`), then
|
||||
`show interfaces ethernet` — port link states are the decisive check
|
||||
(`eth1` up + `eth2`/`eth3` down = single uplink). `configure` (config
|
||||
mode) also accepts `show ...`.
|
||||
- Non-interactive (agent/script): `show`/`configure` are interactive-only
|
||||
aliases on this build; use the op wrapper:
|
||||
```bash
|
||||
ssh ubnt@192.168.66.254 '/opt/vyatta/bin/vyatta-op-cmd-wrapper show interfaces ethernet'
|
||||
```
|
||||
- `show ethernet-switch port all` and `show mac-address-table` are NOT
|
||||
available on this build; the switch FDB is hardware-offloaded
|
||||
(`brctl showmacs switch0` → "Operation not supported"). Port link state
|
||||
+ ARP (`show arp`) are the reliable checks.
|
||||
- SE5420 liveness: `ping 192.168.66.253` and `:80/:443`.
|
||||
- Sample capture at this point (cross-segment/gateway/mDNS flows only;
|
||||
tcpdump needs root — `zhiqiang` has passwordless sudo):
|
||||
```bash
|
||||
ssh zhiqiang@192.168.66.254 "sudo -n tcpdump -ni switch0 -s 0 'udp port 5353 or tcp port 5540 or tcp port 5552'"
|
||||
```
|
||||
|
||||
## Pass criteria
|
||||
|
||||
- Realtime capture consistently shows the target device's mDNS announcements
|
||||
(`_matterc` / `_matter._tcp`) on the chosen point.
|
||||
- Commissioning shows the TCP 5540 handshake (SYN/SYN-ACK/ACK); on the hass
|
||||
`end0` point this includes wired Thread-relay commissioning (HA↔M3), which
|
||||
the AP point cannot see.
|
||||
- Saved pcap opens in Wireshark and filters by `mDNS` / `tcp.port==5540`.
|
||||
|
||||
## References
|
||||
|
||||
- [docs/matter-pairing-troubleshoot.md](../docs/matter-pairing-troubleshoot.md) —
|
||||
troubleshooting decision tree, stage table, device MAC/fabric facts
|
||||
- [docs/unifi-network.md](../docs/unifi-network.md) — UniFi network/IPv6/SSID records
|
||||
- [hosts/hass.windy.lan.md](../hosts/hass.windy.lan.md) — matter-server host
|
||||
networking + `sudo -n -i` non-interactive note
|
||||
- [hosts/gw.md](../hosts/gw.md) — DHCP `matter` reservation MAC mismatch (pending, W1N-207)
|
||||
- [matter-dissector](https://github.com/project-chip/matter-dissector) —
|
||||
Wireshark Matter dissector (incl. decryption)
|
||||
- [Silabs: Using Wireshark to Capture Network Traffic in Matter](https://docs.silabs.com/matter/2.9.1/matter-references/matter-wireshark)
|
||||
@@ -0,0 +1,74 @@
|
||||
# Runbook: controlled network change
|
||||
|
||||
## Purpose
|
||||
|
||||
Apply a controlled network configuration change (DNS records, firewall, LAN
|
||||
gateway, VLAN) with impact assessment, approval, and a rollback path.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: PowerDNS zone records, `us4` firewalld allowlist, LAN gateway/VLAN/DNS changes, WireGuard.
|
||||
- Not applicable: SSH access-policy changes (see `AGENTS.md` §SSH access safety — mandatory lockout-risk procedure).
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: PowerDNS / us4 firewalld / LAN gateway / WireGuard
|
||||
|
||||
## Preconditions
|
||||
|
||||
- A change record (Linear `vps` issue) describes the change, its reason, and rollback.
|
||||
- Read-only impact assessment done (current config captured, blast radius known).
|
||||
|
||||
## Safety
|
||||
|
||||
- Never change DNS or network config without a change record and approval.
|
||||
- Capture the current config first; never delete existing config as the first action.
|
||||
- For DNS: record the current record values and TTL before editing.
|
||||
- For firewall: retain an independent SSH rollback session before applying (see `ansible-operations.md` §us4).
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Assess and capture
|
||||
|
||||
**Action** — capture the current state (e.g. `dig` for DNS, `--check --diff` for firewall, `show` for gateway).
|
||||
|
||||
**Expected** — a baseline of current config and an identified blast radius.
|
||||
|
||||
**Decision** — change fully specified with rollback → Step 2; missing → `STOP`.
|
||||
|
||||
### Step 2 — Approve
|
||||
|
||||
**Action** — confirm approval is recorded in the issue/change record.
|
||||
|
||||
**Decision** — approved → Step 3; not approved → `STOP`.
|
||||
|
||||
### Step 3 — Change
|
||||
|
||||
**Action** — apply the single change (edit the record, run the gated playbook, or change gateway config) and only that change.
|
||||
|
||||
**Expected** — the new value/state is in effect.
|
||||
|
||||
**Verification** — re-query/verify the new state and confirm dependent services still pass health.
|
||||
|
||||
**Rollback** — restore the captured prior config and re-verify.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Troubleshooting A — Change broke dependent service
|
||||
|
||||
- Evidence: health report / endpoint failure.
|
||||
- Allowed: roll back to the captured prior config.
|
||||
- Next: verify; if still broken, escalate.
|
||||
|
||||
## Final Verification
|
||||
|
||||
- New state verified; dependent services healthy.
|
||||
- Change and outcome recorded in the issue.
|
||||
|
||||
## References
|
||||
|
||||
- [`ansible-operations.md`](ansible-operations.md)
|
||||
- [`rollback.md`](rollback.md)
|
||||
- [`network-recovery.md`](network-recovery.md)
|
||||
@@ -0,0 +1,67 @@
|
||||
# Runbook: network outage / service recovery
|
||||
|
||||
## Purpose
|
||||
|
||||
Recover from a network outage or service failure, starting from read-only
|
||||
diagnosis and mutating only when the root cause is confirmed.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: unreachable VPS services, LAN gateway/DNS failures, DNS resolution failures.
|
||||
- Not applicable: planned changes (→ `network-change.md`), SSH access recovery (→ `AGENTS.md` §SSH access safety).
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: VPS services / LAN gateway / DNS
|
||||
|
||||
## Safety
|
||||
|
||||
- Read-only diagnosis first; do not mutate while the root cause is unknown.
|
||||
- If live state conflicts with a runbook's assumptions, `STOP` and report.
|
||||
- Keep the current verified management session open as the recovery path.
|
||||
|
||||
## Procedure
|
||||
|
||||
### Step 1 — Diagnose (read-only)
|
||||
|
||||
**Action** — gather evidence without changing anything:
|
||||
|
||||
```bash
|
||||
# From laptop, pin DNS to a public resolver if the stub is flaky
|
||||
dig @1.1.1.1 +short <host> A
|
||||
curl -4 -sS -I --max-time 10 https://<host>/
|
||||
# From a reachable host, inspect the service
|
||||
ssh -4 windy@<host> 'docker compose ps -a; df -h /; tail -n 50 /var/lib/vps-health/latest.json'
|
||||
```
|
||||
|
||||
**Expected** — a clear picture: is it DNS, connectivity, host, or service?
|
||||
|
||||
**Decision** — root cause localized → Step 2; ambiguous or conflicting → `STOP` and escalate (provider console if host is unreachable).
|
||||
|
||||
### Step 2 — Confirm and route
|
||||
|
||||
**Action** — match the failure to the owning runbook (`mailcow-health.md`, `pdns-health.md`, `matrix-health.md`, etc.) or `network-change.md` for a config fix.
|
||||
|
||||
**Expected** — an applicable runbook with a recovery action.
|
||||
|
||||
**Decision** — applicable → follow it; none → `STOP` (diagnose only, do not mutate).
|
||||
|
||||
### Step 3 — Recover (gated)
|
||||
|
||||
**Action** — apply only the runbook's documented recovery, with approval.
|
||||
|
||||
**Verification** — re-run the health report / endpoint check and confirm recovery.
|
||||
|
||||
**Rollback** — if recovery makes it worse, revert per `rollback.md`.
|
||||
|
||||
## Final Verification
|
||||
|
||||
- Service reachable and health report green.
|
||||
- Incident and recovery recorded in the Linear `vps` issue.
|
||||
|
||||
## References
|
||||
|
||||
- [`network-change.md`](network-change.md)
|
||||
- Per-service health runbooks under [`runbooks/`](.)
|
||||
+22
-1
@@ -1,5 +1,20 @@
|
||||
# PowerDNS health (hk2)
|
||||
|
||||
## Purpose
|
||||
|
||||
Read-only health check of the `/opt/pdns` PowerDNS stack.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: [hk2.chans.xyz](../hosts/hk2.chans.xyz.md), `/opt/pdns`.
|
||||
- Read-only: does not change PowerDNS, DNS records, or secrets.
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-17
|
||||
- Related systems: hk2.chans.xyz (`/opt/pdns`)
|
||||
|
||||
Read-only checks for the `/opt/pdns` stack on **hk2.chans.xyz** (`ns1.wsvc.info`).
|
||||
|
||||
Facts: [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) · Upstream: [docs/pdns-upstream.md](../docs/pdns-upstream.md)
|
||||
@@ -18,7 +33,7 @@ Use these only after the Ansible health report needs investigation.
|
||||
ssh -4 windy@hk2.chans.xyz 'cd /opt/pdns && docker compose ps -a'
|
||||
```
|
||||
|
||||
Expect `auth`, `db`, `poweradmin` healthy; `backup` Up; `pgweb` Up. Ignore stopped orphan `powerdns-admin` unless cleaning orphans.
|
||||
Expect `auth`, `db`, `poweradmin` healthy; `backup` Up; `pgweb` Up. The legacy PDA orphan `powerdns-admin` was removed 2026-08-12 (W1N-59).
|
||||
|
||||
### Version / security poll
|
||||
|
||||
@@ -87,6 +102,12 @@ Expect: `primary=yes`, `also-notify=202.91.35.141`, `only-notify=` empty, `gpgsq
|
||||
The sanitized Ansible health profile is `pdns` (`ansible/playbooks/healthchecks.yml`). It runs locally through `vps-healthcheck.timer`, writes a sanitized JSON result to `/var/lib/vps-health/latest.json`, and uses the API key only inside the PowerDNS container. It does not modify PowerDNS, DNS records, or secrets.
|
||||
|
||||
|
||||
## Safety
|
||||
|
||||
- Read-only: never mutate PowerDNS configuration or DNS records during this check.
|
||||
- Do not paste the API key into chat/logs.
|
||||
- If live state conflicts with an expected value, `STOP` and report.
|
||||
|
||||
## After config changes
|
||||
|
||||
- `auth/pdns.conf`, `auth/templates.d/secrets.j2`, or auth-related `.env` → use the Ansible Compose reconcile playbook with target `auth`
|
||||
|
||||
@@ -0,0 +1,181 @@
|
||||
# Runbook: pgdb health (TimescaleDB + pgweb + pg-backup)
|
||||
|
||||
## Purpose
|
||||
|
||||
Read-only health check of the pgdb TimescaleDB compose stack (PG18 + pgweb GUI + nightly custom-format backups). Confirms the stack is serving Home Assistant (hass/scribe) and that backups are current and restorable.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applicable: [pgdb](../hosts/pgdb.md) (`192.168.55.15`), `/opt/database` compose stack.
|
||||
- Read-only: never mutates containers, databases, backups, or secrets.
|
||||
- Not applicable: restoring data (use [pgdb-restore](pgdb-restore.md)), upgrading images (use [pgdb-update](pgdb-update.md)), HA-side changes (see `hosts/hass.windy.lan.md`).
|
||||
|
||||
## Ownership
|
||||
|
||||
- Owner: personal ops (Windy)
|
||||
- Last reviewed: 2026-08-29
|
||||
- Related systems: pgdb (`/opt/database`, TimescaleDB 18.6 / TS 2.29.2), HA `192.168.55.11` (hass/scribe clients)
|
||||
|
||||
## Access
|
||||
|
||||
SSH to pgdb (key auth works from the WSL agent shell as of 2026-08-29):
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15
|
||||
```
|
||||
|
||||
From the agent sandbox, always use `-F /dev/null` (system ssh config is unreadable there) and prefer IPv4. The compose project lives at `/opt/database` — prefix every `docker compose` call with `cd /opt/database`. Never print `.env` values or bookmarks (they contain DB passwords); compare or use them only inside commands that output non-secret signals (status codes, counts, names).
|
||||
|
||||
## Safety
|
||||
|
||||
### Non-negotiable rules
|
||||
|
||||
- Read-only diagnosis only; never "fix while checking".
|
||||
- Never print passwords or secrets — redact/consume them inside commands.
|
||||
- If live state conflicts with an expected value, `STOP` and record evidence; do not invent parameters or bypass a failed check.
|
||||
- Restore/update work belongs to the change runbooks, not this one.
|
||||
|
||||
### Stop conditions
|
||||
|
||||
- Any container `Exited`, `Restarting`, or not `healthy` where expected.
|
||||
- Expected database/table/hypertable missing or a key query errors.
|
||||
- Write-activity sample does not increase (scribe `states_raw` static).
|
||||
- Latest daily backup older than today, not custom format, or `pg_restore -l` fails.
|
||||
- Disk usage near full on `/srv/pgdata` or `/`.
|
||||
- New `ERROR`/`FATAL` lines in the timescaledb log or backup failures in the pg-backup log.
|
||||
|
||||
## Pass criteria
|
||||
|
||||
- `docker compose ps -a`: `timescaledb` + `pg-backup` **Up (healthy)**, `pgweb` **Up**; ports bound to `192.168.55.15:5432` and `:8081`.
|
||||
- PG 18.x; databases `hass`, `scribe`, `postgres` present; HA (`192.168.55.11`) connected as `hass` to both `hass` and `scribe`.
|
||||
- `hass.states` and `scribe.states_raw` row counts grow between two samples (scribe writes continuously).
|
||||
- TimescaleDB extension 2.29.x; scribe hypertables `states_raw` + `events` (1-dim, `time`); compression configured (segmentby/orderby rows in `timescaledb_information.compression_settings`); `entities` table exists.
|
||||
- pgweb: no credentials → HTTP 401; with credentials → HTTP 200; `/api/bookmarks` → `["hass","scribe"]`.
|
||||
- `daily/*-latest.dump` symlinks point to today's dumps; `file -L` reports `PostgreSQL custom database dump`.
|
||||
- `/srv/pgdata` (`/dev/sdb1`, 32G) and `/` not near full; fstab mounts `/srv/pgdata` by `UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86` with `defaults,noatime`.
|
||||
- timescaledb log: no new `ERROR`/`FATAL`; pg-backup log: recent successful backup.
|
||||
|
||||
## Checks
|
||||
|
||||
### 1. Containers
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose ps -a --format "table {{.Name}}\t{{.Status}}\t{{.Ports}}"'
|
||||
```
|
||||
|
||||
**Expected**
|
||||
|
||||
- `timescaledb` Up (healthy), `pg-backup` Up (healthy), `pgweb` Up.
|
||||
- Ports: `192.168.55.15:5432->5432/tcp` (timescaledb), `192.168.55.15:8081->8081/tcp` (pgweb).
|
||||
|
||||
**Stop** if any container is `Exited`/`Restarting`/`unhealthy`, or a port binding changed.
|
||||
|
||||
### 2. PG core and HA clients
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select version();" | head -1'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select datname from pg_database where datistemplate=false order by 1;"'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select datname, usename, client_addr from pg_stat_activity where client_addr is not null group by 1,2,3 order by 1;"'
|
||||
```
|
||||
|
||||
**Expected**
|
||||
|
||||
- `PostgreSQL 18.x` (verified: 18.6).
|
||||
- Databases: `hass`, `postgres`, `scribe`.
|
||||
- HA sessions: `hass|hass|192.168.55.11` and `scribe|hass|192.168.55.11` (the HAOS recorder/scribe clients from `192.168.55.11`).
|
||||
|
||||
**Stop** if a database is missing, the version is not 18.x, or HA has no live sessions (scribe connectivity is part of the HA pipeline).
|
||||
|
||||
### 3. Write activity
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database
|
||||
A=$(docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select count(*) from states_raw;")
|
||||
sleep 30
|
||||
B=$(docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select count(*) from states_raw;")
|
||||
echo "states_raw $A -> $B"'
|
||||
```
|
||||
|
||||
Also sample `hass.states` once (recorder table, bulk-writes on HA restart): `docker compose exec -T timescaledb psql -U postgres -d hass -Atc "select count(*) from states;"`.
|
||||
|
||||
**Expected** — `states_raw` increases between samples (verified: 2665 → 2693 in 30 s). `states` count is sane (thousands).
|
||||
|
||||
**Stop** if `states_raw` is static across samples while HA is up — writes have stalled.
|
||||
|
||||
### 4. TimescaleDB (extension, hypertables, compression)
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -Atc "select extversion from pg_extension where extname='"'"'timescaledb'"'"';"'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select hypertable_name, num_dimensions from timescaledb_information.hypertables order by 1;"'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select hypertable_name, attname, segmentby_column_index, orderby_column_index from timescaledb_information.compression_settings order by 1,3,4;"'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T timescaledb psql -U postgres -d scribe -Atc "select to_regclass('"'"'public.entities'"'"');"'
|
||||
```
|
||||
|
||||
**Expected**
|
||||
|
||||
- Extension version `2.29.x` (verified: 2.29.2).
|
||||
- Scribe hypertables: `states_raw` and `events`, both `1` dimension.
|
||||
- Compression configured for `states_raw` (segmentby `metadata_id` idx 1, orderby `time` idx 1) and `events` (segmentby `event_type`, orderby `time`). Note: TimescaleDB 2.29.x has **no** `compression_enabled` column in this view — row presence is the enabled signal.
|
||||
- `entities` resolves (scribe registry table).
|
||||
|
||||
**Stop** if the extension version differs from the pinned 2.29.x line, a hypertable is missing, compression rows vanish, or `entities` is absent (scribe schema broke — see Known issues in [hosts/pgdb.md](../hosts/pgdb.md)).
|
||||
|
||||
### 5. pgweb GUI
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database
|
||||
curl -s -o /dev/null -w "no-auth:%{http_code}\n" --max-time 8 http://192.168.55.15:8081/
|
||||
U=$(grep -E "^PGWEB_AUTH_USER=" .env | cut -d= -f2-); P=$(grep -E "^PGWEB_AUTH_PASS=" .env | cut -d= -f2-)
|
||||
curl -s -o /dev/null -w "auth:%{http_code}\n" --max-time 8 -u "$U:$P" http://192.168.55.15:8081/
|
||||
echo -n "bookmarks:"; curl -s --max-time 8 -u "$U:$P" http://192.168.55.15:8081/api/bookmarks; echo'
|
||||
```
|
||||
|
||||
Use `192.168.55.15:8081` (pgweb binds the VM IP only — loopback is not bound). Credentials are read from `.env` on the host and never printed.
|
||||
|
||||
**Expected** — `no-auth:401`, `auth:200`, `bookmarks:["hass","scribe"]`.
|
||||
|
||||
**Stop** if pgweb is unreachable, unauthenticated access is not 401, or bookmarks diverge from `["hass","scribe"]`.
|
||||
|
||||
### 6. Backups
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'ls -l --time-style=long-iso /opt/database/backups/daily/'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'file -L /opt/database/backups/daily/hass-latest.dump'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose exec -T pg-backup pg_restore -l /backups/daily/hass-latest.dump | head -4'
|
||||
```
|
||||
|
||||
**Expected**
|
||||
|
||||
- `daily/*-latest.dump` symlinks point to **today's** `*-YYYYMMDD.dump` (nightly 02:00 local `Asia/Shanghai`; a fresh container also fires `BACKUP_ON_START`).
|
||||
- `file -L` reports `PostgreSQL custom database dump` (pg_restore format; verified v1.16-0).
|
||||
- `pg_restore -l` from the **pg-backup** container lists the archive TOC without error (timescaledb does not mount `/backups`).
|
||||
|
||||
**Stop** if the latest dump is not from today, is not custom format, or `pg_restore -l` fails. Stray non-`daily/` dumps at the `/opt/database/backups/` root are pre-compose leftovers — ignore for health, flag for cleanup.
|
||||
|
||||
### 7. Disk and mount
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'df -h /srv/pgdata /'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'grep -E "srv/pgdata" /etc/fstab'
|
||||
```
|
||||
|
||||
**Expected**
|
||||
|
||||
- `/srv/pgdata` = `/dev/sdb1` 32G (verified: 88M used / 30G avail) and `/` with comfortable headroom.
|
||||
- fstab: `UUID=c9e12e79-1f66-404c-ab7f-b8809be81d86 /srv/pgdata ext4 defaults,noatime 0 2`.
|
||||
|
||||
**Stop** if either filesystem is near full (define threshold before acting) or the fstab entry is missing/changed.
|
||||
|
||||
### 8. Logs
|
||||
|
||||
```bash
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose logs --since 24h timescaledb 2>&1 | grep -E "ERROR|FATAL" | tail -10'
|
||||
ssh -F /dev/null -o BatchMode=yes windy@192.168.55.15 'cd /opt/database && docker compose logs --since 24h pg-backup 2>&1 | tail -5'
|
||||
```
|
||||
|
||||
**Expected**
|
||||
|
||||
- timescaledb: no new `ERROR`/`FATAL`. Known benign history: an old `relation "hass.states" does not exist` from a wrong-schema probe and `compression_enabled` column errors from an outdated query — neither recurs with the commands above.
|
||||
- pg-backup: recent successful run (`SQL backup created successfully` for each database, no restore/cleanup errors).
|
||||
|
||||
**Stop** if repeated `ERROR`/`FATAL` appear or a backup run failed.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user