docs(runbooks): add runbook spec, template, index and 6 first-batch runbooks; light-enhance existing 10

- RUNBOOKS.md: repo-level spec (six-field model, naming, safety, maturity path)
- runbooks/_template.md + README.md: standard template and 16-entry routing index
- new: issue-to-merge, fix-ci, release, rollback, network-change, network-recovery
- light-enhance 10 existing runbooks with Purpose/Scope/Safety headers
- AGENTS.md: point step 3 at index/spec, add runbook execution rules
- docs/agent-runbook-guide.md: archive of Manus AI guide
This commit is contained in:
windyboy
2026-08-17 15:59:46 +08:00
parent 047ac03346
commit b0c01b2551
21 changed files with 1140 additions and 2 deletions
+23 -2
View File
@@ -8,7 +8,10 @@ Also readable as `agent.md` (symlink → this file).
1. Read [`inventory/hosts.md`](inventory/hosts.md) for the machine list. 1. Read [`inventory/hosts.md`](inventory/hosts.md) for the machine list.
2. Open the matching [`hosts/<name>.md`](hosts/) for SSH, roles, paths, and quirks. 2. Open the matching [`hosts/<name>.md`](hosts/) for SSH, roles, paths, and quirks.
3. For common tasks, follow a runbook under [`runbooks/`](runbooks/). 3. For common tasks, follow a runbook under [`runbooks/`](runbooks/). Pick the
most specific applicable one from [`runbooks/README.md`](runbooks/README.md);
the spec is [`RUNBOOKS.md`](RUNBOOKS.md) and new runbooks start from
[`runbooks/_template.md`](runbooks/_template.md).
4. Prefer read-only checks first; change only after confirming current state. 4. Prefer read-only checks first; change only after confirming current state.
5. For routine checks and approved service reconciliation, run the matching 5. For routine checks and approved service reconciliation, run the matching
Ansible playbook from `ansible/`; see [routine Ansible operations](runbooks/ansible-operations.md). Ansible playbook from `ansible/`; see [routine Ansible operations](runbooks/ansible-operations.md).
@@ -20,6 +23,15 @@ Also readable as `agent.md` (symlink → this file).
scope, action, verification, and remaining follow-up; never put passwords, scope, action, verification, and remaining follow-up; never put passwords,
tokens, private keys, recovery keys, or private room IDs in Linear. tokens, private keys, recovery keys, or private room IDs in Linear.
### Runbook execution rules
Before operational work: inspect `runbooks/`, select the most specific
applicable runbook, follow its steps in order, do not skip verification steps,
and respect its STOP and approval conditions. If no runbook applies, diagnose
only — do not mutate production state. When live state conflicts with a
runbook's assumptions, `STOP` and report; never invent missing parameters or
bypass failed checks. The spec is [`RUNBOOKS.md`](RUNBOOKS.md).
## Active hosts (quick map) ## Active hosts (quick map)
| Host | Role | SSH | Facts | | Host | Role | SSH | Facts |
@@ -82,6 +94,14 @@ Also readable as `agent.md` (symlink → this file).
| UniFi local-service proxy bypass | [docs/unifi-openclash-localhost.md](docs/unifi-openclash-localhost.md) | | UniFi local-service proxy bypass | [docs/unifi-openclash-localhost.md](docs/unifi-openclash-localhost.md) |
| UniFi SSO login setting (Ansible) | `cd ansible && ansible-playbook playbooks/unifi-sso.yml --limit unifi` | | UniFi SSO login setting (Ansible) | `cd ansible && ansible-playbook playbooks/unifi-sso.yml --limit unifi` |
| Routine Ansible operations | [runbooks/ansible-operations.md](runbooks/ansible-operations.md) | | Routine Ansible operations | [runbooks/ansible-operations.md](runbooks/ansible-operations.md) |
| Issue → mergeable change | [runbooks/issue-to-merge.md](runbooks/issue-to-merge.md) |
| Fix failing health/playbook run | [runbooks/fix-ci.md](runbooks/fix-ci.md) |
| Release a reviewed change | [runbooks/release.md](runbooks/release.md) |
| Roll back a change | [runbooks/rollback.md](runbooks/rollback.md) |
| Controlled network change | [runbooks/network-change.md](runbooks/network-change.md) |
| Network outage recovery | [runbooks/network-recovery.md](runbooks/network-recovery.md) |
Full index: [runbooks/README.md](runbooks/README.md). Spec: [RUNBOOKS.md](RUNBOOKS.md).
Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.yml --limit mailcow`. The local stub resolver is flaky; DNS probes use `1.1.1.1` / `8.8.8.8`. Routine mailcow health: `cd ansible && ansible-playbook playbooks/health-report.yml --limit mailcow`. The local stub resolver is flaky; DNS probes use `1.1.1.1` / `8.8.8.8`.
@@ -131,8 +151,9 @@ Bills, rough notes, and personal clutter stay in the Obsidian vault. This repo h
``` ```
AGENTS.md / agent.md # this entry (agent.md → AGENTS.md) AGENTS.md / agent.md # this entry (agent.md → AGENTS.md)
RUNBOOKS.md # runbook spec (six-field model, naming, review rules)
inventory/hosts.md # machine index inventory/hosts.md # machine index
hosts/ # per-host facts hosts/ # per-host facts
runbooks/ # step-by-step ops runbooks/ # step-by-step ops (README.md = index, _template.md = template)
docs/ # upstream doc indexes / design notes docs/ # upstream doc indexes / design notes
``` ```
+68
View File
@@ -0,0 +1,68 @@
# RUNBOOKS — 仓库级规范
本文件统一所有 Runbook 的字段、命名、评审与变更规则。上游参考:[docs/agent-runbook-guide.md](docs/agent-runbook-guide.md)。
## 目录结构
```text
runbooks/
├── README.md # 意图 → 文件 路由索引(本目录的入口)
├── _template.md # 新建 runbook 的标准模板(复制后填写)
├── <intent>.md # 每份 runbook 只描述一种可识别的操作意图
└── ...
```
## 最小字段模型
每份 runbook 必须显式包含以下控制信息,否则盲目执行或错误恢复的风险会升高:
| 字段 | 作用 | 写作要求 |
|---|---|---|
| **Action** | 定义当前要执行的动作 | 可观察、可执行的动词;避免“检查一下”“适当调整” |
| **Expected** | 描述正常状态或预期输出 | 具体信号、阈值、状态码、测试结果或页面表现 |
| **Decision** | 定义分支与下一跳 | “条件 → 下一步”;无法判断时指向 `STOP` |
| **Verification** | 确认变更真正生效 | 每个有副作用的步骤后执行,不可跳过 |
| **Stop condition** | 规定何时不得继续 | 列出信息缺失、状态冲突、权限不足、验证失败等 |
| **Rollback** | 如何恢复到变更前状态 | 触发条件、前提、撤销步骤、回滚后验证 |
> 只读类 runbook 不产生副作用,可省略 Rollback;但必须保留 Stop condition(状态与预期冲突即 `STOP` 并记录证据)。
## 命名与拆分规则
- 文件名采用小写连字符,反映**操作意图**而非目标主机,例如 `mailcow-health.md``release.md`
- 一份文件只描述一种意图。流程出现明显分叉时拆分为独立文件,不堆叠“万能流程”。
- 只读诊断与变更操作应分离:health 类 runbook 保持只读,变更走 `ansible-operations.md``release.md``rollback.md` 或对应 gated playbook。
## 章节约定
- 每份 runbook 顶部含 `## Purpose`12 行)与 `## Scope`(适用/不适用情形)。
- 变更型 runbook 额外含 `## Approval gates` 表;破坏性/不可逆操作必须获得明确批准。
- 统一在 `## Safety` 或正文中复用以下通用安全规则(更严格要求优先)。
```markdown
## Safety Rules
- Never delete an existing configuration as the first recovery action.
- Prefer read-only diagnosis before mutation.
- After every mutation, verify the expected state.
- If actual state conflicts with this runbook, STOP.
- Do not invent missing parameters.
- Do not bypass failed tests.
- Destructive actions require explicit approval.
```
## 评审与变更规则
- 新建/修改 runbook 与代码同仓评审,随系统演进更新。
- 每份 runbook 标注 `Last reviewed`;流程执行过程中发现的偏差记入对应的 Linear `vps` 项目 issue。
- 破坏性流程(迁移、删除、DNS 变更、网络变更)保持人工审批,不自动下沉。
## 成熟路径
1. **人工处理** → 现场处置与复盘,记录证据。
2. **Markdown runbook** → 固化步骤与证据要求,Agent 可辅助诊断。
3. **Agent + runbook** → 严格按流程执行,受 Stop/Approval 约束。
4. **Script / Ansible / Skill** → 把已稳定、低歧义、可验证的操作程序化(本仓库的执行层是 Ansible playbook)。
5. **人工审批 + 自动执行** → 审批门控下的自动变更(如 gated playbook + 确认变量)。
原则:先证据后变更,先小范围后扩大,先验证后结束,不确定则停止。
+326
View File
@@ -0,0 +1,326 @@
# Agent Runbook 实用指南(v1
> **定位**:本指南用于把团队的重复性运维、交付与故障处理经验写成可由 Agent 安全执行的流程。它适用于以 Git 仓库为中心的工程协作模式,优先采用 **Markdown + Git 版本控制 + 明确的 Agent 路由规则**,而不是一开始引入复杂的自动化平台。
> 本文件为上游参考存档。仓库内落地规范见 [`RUNBOOKS.md`](../RUNBOOKS.md),标准模板见 [`runbooks/_template.md`](../runbooks/_template.md),索引见 [`runbooks/README.md`](../runbooks/README.md)。
## 1. 什么是 Agent Runbook
Runbook 是预先设计的、可重复执行的操作流程,用于处理部署、告警、故障、配置变更、CI 修复等标准化工作。传统 Runbook 的主要读者是人;**Agent Runbook 则必须把人的隐性判断显式化**,使 Agent 能知道做什么、看到什么才算正常、下一步去哪里、何时停止以及如何撤销。
Google SRE 强调在事故发生前设计响应流程、系统化排障,并逐步将重复性运维工作自动化。[1] [2] AWS Systems Manager Automation 则把可执行 Runbook 建模为顺序步骤:每个步骤调用一个动作,前一步输出可以传递给后续步骤。[3] 这两种思路共同构成了 Agent Runbook 的实用基础。
| 层次 | 核心问题 | 应承担的职责 |
|---|---|---|
| `AGENTS.md` | **何时使用哪份流程?** | 工作路由、通用操作约束、无匹配流程时的默认行为 |
| `runbooks/*.md` | **这件事按什么流程做?** | 前置条件、分步操作、决策分支、验证、停止条件与回滚 |
| Skill / MCP / Tool | **有哪些可调用能力?** | 具体能力、参数、权限边界和使用说明 |
| Shell / GitHub / Linear / SSH 等 | **实际如何执行?** | 对系统、代码库或外部服务执行操作 |
## 2. 设计目标与适用边界
Agent Runbook 的目标不是让 Agent 在所有异常下“想办法修好”,而是在一个**已知、受控、可验证、可回退**的边界中提高执行一致性。它应当优先覆盖高频、后果明确、流程稳定的操作,例如 CI 失败定位、Issue 到合并请求、发布前检查、标准部署、回滚及网络变更。
| 适合纳入 Runbook | 暂不适合直接自动执行 |
|---|---|
| 明确输入、固定步骤、可观察结果的操作 | 目标或验收标准尚不清楚的探索性任务 |
| 可在每次修改后验证状态的变更 | 缺失关键参数、权限或上下文的任务 |
| 具有安全回滚路径的发布与配置调整 | 高破坏性、不可逆或影响面未知的操作 |
| 可由权限与审批规则约束的运维流程 | 与既有流程事实冲突、无法判断根因的异常场景 |
> **基本原则**:当实际状态与 Runbook 的假设冲突,Agent 应停止并呈报,而不是补全未知信息、绕过检查或继续试错。
## 3. Agent Runbook 的最小字段
与普通人工 Runbook 相比,Agent Runbook 必须显式包含以下六类控制信息。缺少其中任一项,都会增加盲目执行或错误恢复的风险。
| 字段 | 作用 | 写作要求 |
|---|---|---|
| **Action** | 定义当前要执行的动作 | 使用可观察、可执行的动词;避免“检查一下”“适当调整”等模糊表述 |
| **Expected** | 描述正常状态或预期输出 | 给出具体信号、阈值、状态码、测试结果或页面表现 |
| **Decision** | 定义分支与下一跳 | 用“条件 → 下一步”的形式;无法判断时指向 `STOP` |
| **Verification** | 确认变更真正生效 | 在每个有副作用的步骤后执行,不能被跳过 |
| **Stop condition** | 规定何时不得继续 | 明确列出信息缺失、状态冲突、权限不足、验证失败等条件 |
| **Rollback** | 描述如何恢复到变更前状态 | 标明触发条件、前提、撤销步骤及回滚后的验证方式 |
## 4. 推荐目录与路由机制
建议把流程与代码一起保存在 Git 仓库中。这样 Runbook 可以评审、版本化、随系统演进更新,也能与相关 Issue、PR 和配置建立可追溯关系。
```text
repo/
├── AGENTS.md
├── RUNBOOKS.md
├── runbooks/
│ ├── README.md
│ ├── issue-to-merge.md
│ ├── fix-ci.md
│ ├── release.md
│ ├── rollback.md
│ ├── network-change.md
│ └── network-recovery.md
└── ...
```
### `AGENTS.md`:只做路由与通用约束
`AGENTS.md` 不应重复流程细节。它只需要规定 Agent 在进行操作类工作前,先查找最具体且适用的 Runbook,并严格遵守其中的步骤、验证、停止和审批要求。
```markdown
# Operational Rules
Before performing operational work:
1. Inspect `runbooks/`.
2. Select the most specific applicable runbook.
3. Follow its steps in order.
4. Do not skip verification steps.
5. Respect STOP and approval conditions.
6. If no runbook applies, diagnose only; do not mutate production state.
## Routing
- CI failure → `runbooks/fix-ci.md`
- GitHub issue implementation → `runbooks/issue-to-merge.md`
- Deployment → `runbooks/release.md`
- Rollback → `runbooks/rollback.md`
- Network configuration → `runbooks/network-change.md`
- Network outage → `runbooks/network-recovery.md`
```
### `RUNBOOKS.md`:仓库级规范
`RUNBOOKS.md` 用于统一所有 Runbook 的字段、命名、评审要求和变更规则。每份 Runbook 只描述一种可识别的操作意图;如果流程已有明显分叉,应拆分为独立文件,而不是堆叠成长篇“万能流程”。
## 5. 规范模板
以下模板可直接保存为 `runbooks/_template.md` 使用。
```markdown
# Runbook: <名称>
## Purpose
说明本 Runbook 要解决的问题及成功结果。
## Scope
- 适用环境:<如 development / staging / production>
- 适用对象:<服务、仓库、组件或告警类型>
- 不适用情形:<需要改用其他 Runbook 或转人工的场景>
## Ownership
- Owner<团队或角色>
- Last reviewed<YYYY-MM-DD>
- Related systems<系统名称>
## Preconditions
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
## Inputs
| 输入 | 来源 | 是否必需 | 校验方法 |
|---|---|---:|---|
| <参数> | <来源> | 是/否 | <如何确认有效> |
## Safety
### Non-negotiable rules
- 先只读诊断,后执行变更。
- 不得把删除现有配置作为首次恢复动作。
- 不得猜测或编造缺失参数。
- 不得绕过失败的测试、检查或审批。
- 每次变更后必须完成对应验证。
- 破坏性操作必须获得明确批准。
### Stop conditions
- 实际状态与本文档的前提或预期结果冲突。
- 缺少必要输入、权限、审批或回滚能力。
- 验证失败且本文档没有明确的下一步。
- 影响范围超出 Scope。
### Approval gates
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|---|---|---:|---|
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
## Procedure
### Step 1 — Diagnose
**Action**
<执行只读诊断动作。>
**Expected**
<列出预期输出、状态或证据。>
**Decision**
- 若 <条件 A>,进入 Step 2。
- 若 <条件 B>,进入 Troubleshooting A。
- 若无法判断或状态冲突,`STOP` 并记录证据。
### Step 2 — Change
**Action**
<描述单一、可审计的变更动作。>
**Expected**
<变更后应出现的状态。>
**Verification**
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
**Rollback**
- 触发条件:<什么情况需要回滚>
- 回滚动作:<如何撤销>
- 回滚验证:<如何确认恢复成功>
## Troubleshooting
### Troubleshooting A — <异常名称>
- 证据收集:<日志、指标、命令输出、链接>
- 允许动作:<仅限已验证且低风险的动作>
- 下一步:<回到某步 / 转入另一 Runbook / STOP 并升级>
## Final Verification
只有同时满足以下标准,流程才算成功:
- <功能或服务状态>
- <自动化测试或健康检查>
- <监控指标或告警状态>
- <变更记录、PR 或 Issue 已更新>
## Failure Handling
若未能完成:
1. 停止进一步变更。
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
4. 按 <升级渠道> 交接,不继续猜测。
## References
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
```
## 6. 编写步骤的标准写法
每个步骤应只承担一个清晰目的,并使用“动作—预期—决策”的闭环表达。如下表所示,前者会导致 Agent 自主扩大操作范围,后者则为其提供安全边界。
| 不推荐写法 | 推荐写法 |
|---|---|
| “检查部署是否正常,不正常就修复。” | “读取部署状态与最近一次发布记录。若所有副本 `Ready` 且版本等于目标版本,进入 Final Verification;若副本未就绪,收集事件与日志并进入 Troubleshooting A;若版本不匹配且原因未知,`STOP`。” |
| “必要时修改配置。” | “仅当配置差异与变更单 `CHG-123` 完全一致且审批已记录时,应用指定键的值;应用后运行健康检查;失败则按 Rollback 回退。” |
| “测试失败时可先跳过。” | “任何必需测试失败均不得继续部署。记录失败测试、日志和提交版本;仅按 Troubleshooting B 处理。” |
## 7. 通用安全规则
以下规则适合在每份 Runbook 的 `Safety` 章节中复用。若某流程存在更严格要求,应以更严格要求为准。
```markdown
## Safety Rules
- Never delete an existing configuration as the first recovery action.
- Prefer read-only diagnosis before mutation.
- After every mutation, verify the expected state.
- If actual state conflicts with this runbook, STOP.
- Do not invent missing parameters.
- Do not bypass failed tests.
- Destructive actions require explicit approval.
```
这些约束体现了一个关键顺序:**先证据,后变更;先小范围,后扩大;先验证,后结束;不确定则停止。** 特别是停止条件必须可操作,例如“权限不足”“缺少变更单”“生产状态与前提不一致”“错误率超过 1%”等,而不应写成“情况复杂时停止”。
## 8. 运行与审计流程
Agent 执行 Runbook 时,应按照固定运行模型工作。每一步的输入、动作、输出和下一跳都应可追踪,这与 AWS 自动化 Runbook 的顺序步骤和输出传递思想一致。[3]
```text
输入与前置条件
只读诊断
确认预期状态或决策分支
获取审批(如需要)
执行最小变更
立即验证
成功收尾 / 回滚 / 停止并升级
```
| 阶段 | Agent 必须产出的证据 | 禁止行为 |
|---|---|---|
| 输入确认 | 参数来源、环境、目标资源、权限与审批状态 | 用猜测值补全必需参数 |
| 诊断 | 命令输出、日志、指标或页面状态 | 在未诊断前直接修改生产状态 |
| 变更 | 实际执行内容、变更范围、时间 | 将多个无关变更混在一起执行 |
| 验证 | 测试、健康检查、监控状态与预期对比 | 以“命令执行成功”代替业务验证 |
| 失败处理 | 已做步骤、异常证据、回滚状态和升级对象 | 无限制重试或绕过失败检查 |
## 9. 从人工操作到自动化的成熟路径
不建议在流程尚未稳定时先构建复杂 DSL 或全自动编排。应先积累真实案例,把可重复部分固化为 Markdown Runbook,再把已稳定、低歧义、可验证的操作迁移到脚本、CI、Skill 或自动化系统。Google SRE 将能够由机器替代的重复性人工工作视为应逐步消除的 toil。[4]
| 阶段 | 主要形式 | 人的角色 | 自动化边界 |
|---|---|---|---|
| 1. 人工处理 | 现场处置与复盘 | 执行、判断、记录 | 不自动化 |
| 2. Markdown Runbook | 固化步骤与证据要求 | 审核流程与异常判断 | Agent 可辅助诊断 |
| 3. Agent + Runbook | 严格按流程执行 | 审批高风险动作、处理例外 | 受停止条件约束的执行 |
| 4. Script / Skill / CI / Automation | 把稳定步骤程序化 | 处理异常和维护自动化 | 自动完成重复性操作 |
| 5. 人工审批 + 自动执行 | 常规流程端到端运行 | 决策、审计与治理 | 审批门控下的自动变更 |
## 10. 上线前检查清单
在将一份新 Runbook 交给 Agent 使用前,建议由流程所有者按以下清单审核。
| 检查项 | 合格标准 |
|---|---|
| 问题边界 | Purpose 与 Scope 清楚描述适用和不适用情形 |
| 输入 | 所有必需输入都有来源、格式和校验方法 |
| 步骤 | 每一步均有 Action、Expected 与明确的下一跳 |
| 变更控制 | 所有修改动作都有 Verification;关键动作有 Rollback |
| 安全控制 | Stop conditions、审批门槛和禁止行为已列明 |
| 异常处理 | 失败时知道收集什么证据、交给谁,而非继续猜测 |
| 可维护性 | 有 Owner、最近复审日期与关联文档;已在版本控制中评审 |
| 可演练性 | 已在安全环境或历史案例上走通至少一次 |
## 11. 建议的首批 Runbook
首次落地时,应优先选择频率较高、输入相对明确、变更可回退的场景。以下集合通常能覆盖大部分工程协作的基础需求。
| Runbook | 目的 | 关键安全控制 |
|---|---|---|
| `issue-to-merge.md` | 从已明确 Issue 到可评审变更 | Scope 锁定、测试门槛、PR 证据 |
| `fix-ci.md` | 诊断并修复 CI 失败 | 不跳过测试、不修改无关代码 |
| `release.md` | 执行标准发布 | 发布窗口、审批、健康检查、回滚点 |
| `rollback.md` | 恢复到已知稳定版本 | 明确触发条件、版本选择、回滚后验证 |
| `network-change.md` | 实施受控网络配置变更 | 影响评估、变更单、回退配置 |
| `network-recovery.md` | 处理网络异常与服务恢复 | 只读诊断优先、状态冲突即停止 |
## 12. 结论
Agent Runbook 的价值不在于把每一项运维工作立即自动化,而在于将团队的工程判断编码为**可路由、可验证、可停止、可回滚**的操作系统。对于多数团队,从仓库中的 `AGENTS.md``RUNBOOKS.md` 和一组 Markdown Runbook 起步,已经足够实用。
当某个流程经过多次执行、输入稳定、异常分支收敛且验证可靠后,再将其下沉为脚本、CI 或其他自动化能力。这样既能逐步降低重复性 toil,也能始终保留人类对高风险和例外情形的决策权。[4]
## References
[1]: https://sre.google/sre-book/managing-incidents/ "Google SRE Book — Managing Incidents"
[2]: https://sre.google/sre-book/effective-troubleshooting/ "Google SRE Book — Effective Troubleshooting"
[3]: https://docs.aws.amazon.com/systems-manager/latest/userguide/automation-documents.html "AWS Systems Manager — Creating your own runbooks"
[4]: https://sre.google/sre-book/eliminating-toil/ "Google SRE Book — Eliminating Toil"
[5]: https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-automation.html "AWS Systems Manager Automation"
[6]: https://docs.aws.amazon.com/systems-manager-automation-runbooks/latest/userguide/automation-runbook-reference.html "AWS Systems Manager Automation Runbook Reference"
[7]: https://learn.microsoft.com/en-us/azure/automation/manage-runbooks "Microsoft Learn — Manage runbooks in Azure Automation"
---
**来源**Manus AI《Agent Runbook 实用指南(v1.0)》,本仓库存档为规范参考。
+34
View File
@@ -0,0 +1,34 @@
# Runbook index
Entry point for all runbooks. Before operational work, read the repo entry
[`AGENTS.md`](../AGENTS.md) and the spec [`RUNBOOKS.md`](../RUNBOOKS.md). New
runbooks start from [`_template.md`](_template.md).
## Route by intent
| Intent | Runbook | Type |
|---|---|---|
| mailcow health check | [mailcow-health.md](mailcow-health.md) | read-only |
| mailcow update | [mailcow-update.md](mailcow-update.md) | change (gated) |
| mailcow SMTP/IMAP client | [mailcow-smtp-client.md](mailcow-smtp-client.md) | reference |
| Vaultwarden health check | [vaultwarden-health.md](vaultwarden-health.md) | read-only |
| Vaultwarden SQLite→PG migrate | [vaultwarden-sqlite-to-postgres.md](vaultwarden-sqlite-to-postgres.md) | change (destructive) |
| PowerDNS health check | [pdns-health.md](pdns-health.md) | read-only |
| RustDesk health check | [rustdesk-health.md](rustdesk-health.md) | read-only |
| Matrix health check | [matrix-health.md](matrix-health.md) | read-only |
| AdGuard Home health check | [adguard-home-health.md](adguard-home-health.md) | read-only |
| Routine Ansible operations | [ansible-operations.md](ansible-operations.md) | change (allowlisted) |
| Linear issue → mergeable change | [issue-to-merge.md](issue-to-merge.md) | delivery |
| Failing health/playbook run | [fix-ci.md](fix-ci.md) | change |
| Release a reviewed change to production | [release.md](release.md) | change (gated) |
| Roll back a change | [rollback.md](rollback.md) | change (gated) |
| Controlled network configuration | [network-change.md](network-change.md) | change (gated) |
| Network outage / service recovery | [network-recovery.md](network-recovery.md) | recovery |
## Notes
- `fix-ci.md`, `release.md`, `rollback.md`, `network-change.md`, `network-recovery.md`
are adapted from the upstream guide to this repo's VPS-ops context (execution
layer is Ansible + SSH + Linear, not a software CI/CD pipeline).
- Health runbooks are read-only; they stop (`STOP`) when live state conflicts
with the expected state instead of mutating production.
+119
View File
@@ -0,0 +1,119 @@
# Runbook: <名称>
## Purpose
<说明本 Runbook 要解决的问题及成功结果,1–2 行。>
## Scope
- 适用环境:<production / staging / LAN …>
- 适用对象:<服务、主机、组件或告警类型>
- 不适用情形:<需要改用其他 runbook 或转人工的场景>
## Ownership
- Owner<团队或角色>
- Last reviewed<YYYY-MM-DD>
- Related systems<主机名 / 服务名>
## Preconditions
- <执行前必须满足的权限、备份、窗口、健康状态或已知信息>
## Inputs
| 输入 | 来源 | 是否必需 | 校验方法 |
|---|---|---:|---|
| <参数> | <来源> | 是/否 | <如何确认有效> |
## Safety
### Non-negotiable rules
- 先只读诊断,后执行变更。
- 不得把删除现有配置作为首次恢复动作。
- 不得猜测或编造缺失参数。
- 不得绕过失败的测试、检查或审批。
- 每次变更后必须完成对应验证。
- 破坏性操作必须获得明确批准。
### Stop conditions
- 实际状态与本文档的前提或预期结果冲突。
- 缺少必要输入、权限、审批或回滚能力。
- 验证失败且本文档没有明确的下一步。
- 影响范围超出 Scope。
### Approval gates
| 动作 | 风险级别 | 是否需要明确批准 | 批准记录位置 |
|---|---|---:|---|
| <动作> | 低/中/高 | 是/否 | <Issue / PR / 变更单> |
## Procedure
### Step 1 — Diagnose
**Action**
<执行只读诊断动作。>
**Expected**
<列出预期输出、状态或证据。>
**Decision**
- 若 <条件 A>,进入 Step 2。
- 若 <条件 B>,进入 Troubleshooting A。
- 若无法判断或状态冲突,`STOP` 并记录证据。
### Step 2 — Change
**Action**
<描述单一、可审计的变更动作。>
**Expected**
<变更后应出现的状态。>
**Verification**
<给出可重复执行的验证命令、测试、监控指标或检查清单。>
**Rollback**
- 触发条件:<什么情况需要回滚>
- 回滚动作:<如何撤销>
- 回滚验证:<如何确认恢复成功>
## Troubleshooting
### Troubleshooting A — <异常名称>
- 证据收集:<日志、指标、命令输出、链接>
- 允许动作:<仅限已验证且低风险的动作>
- 下一步:<回到某步 / 转入另一 runbook / STOP 并升级>
## Final Verification
只有同时满足以下标准,流程才算成功:
- <功能或服务状态>
- <自动化测试或健康检查>
- <监控指标或告警状态>
- <变更记录、PR 或 Issue 已更新>
## Failure Handling
若未能完成:
1. 停止进一步变更。
2. 收集 <命令输出、时间范围、请求 ID、日志链接、截图或复现步骤>。
3. 记录已完成步骤、实际结果、未满足的预期和是否执行过回滚。
4. 按 <升级渠道> 交接,不继续猜测。
## References
- <关联 Issue、PR、架构文档、仪表盘、配置仓库或外部文档>
+15
View File
@@ -1,5 +1,14 @@
# AdGuard Home health — dns.windy.lan # AdGuard Home health — dns.windy.lan
## Purpose
Read-only health check of the AdGuard Home LAN DNS service.
## Scope
- Applicable: [dns.windy.lan](../hosts/dns.windy.lan.md) (`192.168.66.36`).
- Read-only: does not expose query-log contents or secrets; does not change configuration.
This runbook is read-only. It does not expose query-log contents or secrets. This runbook is read-only. It does not expose query-log contents or secrets.
Routine checks run through Ansible on demand: Routine checks run through Ansible on demand:
@@ -58,3 +67,9 @@ a known-bad-signature test; an enabled DO bit alone is not validation.
Private PTR forwarding is intentionally absent because the EdgeRouter does Private PTR forwarding is intentionally absent because the EdgeRouter does
not currently answer private PTR requests. not currently answer private PTR requests.
## Safety
- Read-only: never change the DNS policy or the `agh-ui-access.service` nftables rule during this check.
- Do not infer a broken DNS policy from an empty `allowed_clients`.
- If live state conflicts with an expected value, `STOP` and report.
+15
View File
@@ -1,8 +1,23 @@
# Runbook: routine operations through Ansible # Runbook: routine operations through Ansible
## Purpose
Routine operations (health, reconcile, maintenance) through the Ansible playbooks.
## Scope
- Applicable: every inventory host, run from `ansible/`.
- Not applicable: arbitrary remote commands — the reconcile playbook is allowlisted and gated.
Run commands from `ansible/`. The inventory forces IPv4 and uses the `windy` Run commands from `ansible/`. The inventory forces IPv4 and uses the `windy`
account with sudo. Do a read-only health pass before any reconciliation. account with sudo. Do a read-only health pass before any reconciliation.
## Safety
- Read-only health pass before any reconciliation.
- Mutating playbooks require explicit confirmation variables; do not bypass them.
- If a reconcile target or service name is not allowlisted, `STOP` — do not invent one.
## Health report (read-only) ## Health report (read-only)
```bash ```bash
+70
View File
@@ -0,0 +1,70 @@
# Runbook: fix a failing health/playbook run
> Adapted from the upstream guide's `fix-ci`. This repo has no software CI; the
> equivalent "pipeline" is the Ansible **health report** and the gated playbooks.
> This runbook covers diagnosing and fixing a failed or warning/critical run.
## Purpose
Diagnose and fix a failing Ansible health-report or playbook run without
skipping checks or changing unrelated code.
## Scope
- Applicable: `ansible-playbook playbooks/health-report.yml` and the gated playbooks under `ansible/playbooks/`.
- Not applicable: production changes beyond fixing the run; network/DNS changes → `network-change.md`.
## Safety
- Do not skip or weaken a failing check to make it pass.
- Do not change unrelated hosts or services.
- Prefer read-only diagnosis before mutation; destructive fixes require approval.
## Procedure
### Step 1 — Reproduce and read
**Action** — re-run the failing playbook with `--limit <host>` and capture the task that failed.
```bash
cd ansible
ansible-playbook playbooks/health-report.yml --limit <host> -v
```
**Expected** — a specific failed task, host, and message (warning vs critical).
**Decision** — clear failure → Step 2; ambiguous → `STOP` and collect `-vvv` output + the relevant `latest.json`.
### Step 2 — Diagnose
**Action** — inspect the corresponding service on the host using the matching health runbook (`mailcow-health.md`, `vaultwarden-health.md`, `pdns-health.md`, etc.).
**Expected** — a root cause (container down, cert expired, queue backlog, drift).
**Decision** — root cause found → Step 3; live state conflicts with the runbook's assumptions → `STOP`.
### Step 3 — Fix within scope
**Action** — apply the minimal fix the service runbook prescribes (e.g. `compose-reconcile` for a config drift, or a documented update). Use only allowlisted/gated playbooks.
**Verification** — re-run the health report and confirm it passes.
**Rollback** — revert to the prior config/state and re-run; see `rollback.md` for the general procedure.
## Troubleshooting
### Troubleshooting A — Intermittent/flaky failure
- Evidence: timing, DNS stub flakiness (use `1.1.1.1`/`8.8.8.8` for probes).
- Allowed: re-run once with the documented resolver workaround.
- Next: still failing → `STOP` and escalate.
## Final Verification
- Health report passes for the affected host.
- No checks were skipped or weakened; the fix is committed/documented.
## References
- [`ansible-operations.md`](ansible-operations.md)
- Per-service health runbooks under [`runbooks/`](.)
+89
View File
@@ -0,0 +1,89 @@
# Runbook: issue → mergeable change
## Purpose
Turn an approved Linear `vps` issue into a reviewed, mergeable change in this
repo (docs, runbooks, hosts facts, or Ansible playbooks).
## Scope
- Applicable: repo content under `docs/`, `runbooks/`, `hosts/`, `inventory/`, `ansible/`.
- Not applicable: mutating production state directly — that goes through `release.md` / `ansible-operations.md`.
## Ownership
- Owner: personal ops (Windy)
- Last reviewed: 2026-08-17
- Related systems: Linear MCP (`vps` project), git
## Inputs
| Input | Source | Required | Validation |
|---|---|---:|---|
| Issue identifier | Linear (`vps` project) | Yes | `linear_get_issue <id>` returns a description |
| Current repo state | `git status` / `git log` | Yes | Clean or intended worktree |
## Safety
- Scope is locked to the issue: do not bundle unrelated changes.
- Never commit secrets (see `AGENTS.md` §Safety).
- Verify every change; do not merge a change whose verification was skipped.
## Procedure
### Step 1 — Read the issue
**Action**`linear_get_issue <id>`, read description and acceptance criteria.
**Expected** — clear scope, action, and verification for the change.
**Decision** — if the issue is ambiguous or lacks verification criteria, `STOP`
and ask for clarification (add a `needs-info` label if applicable). Otherwise go to Step 2.
### Step 2 — Inspect and change
**Action** — read the relevant files, then make the minimal change the issue asks for.
**Expected** — diff is scoped to the issue.
**Decision** — if the change needs production mutation, `STOP` and route to
`release.md`. Otherwise go to Step 3.
### Step 3 — Verify
**Action** — run the applicable verification (link check for docs; `ansible-playbook --check` / `--syntax-check` for playbooks).
**Verification** — see `RUNBOOKS.md` §校验; the concrete checks must match the change type.
**Decision** — verification passed → Step 4; failed → Troubleshooting A.
### Step 4 — Commit and link
**Action** — commit with a message containing the full issue ID (e.g. `W1N-123: …`); open a PR if the change is substantial; link the issue via `linear_save_comment`.
**Verification**`git log -1` shows the issue ID; the issue has the commit/PR pointer.
**Rollback**`git revert <sha>` or `git checkout <branch>` to drop the change; re-verify after.
## Troubleshooting
### Troubleshooting A — Verification failed
- Evidence: command output, failing check.
- Allowed: fix the change within scope; re-run verification.
- Next: still failing → `STOP` and report in the issue.
## Final Verification
- Change matches the issue scope.
- Verification passed and the issue is updated with evidence.
## Failure Handling
If unfinished: stop, collect the failed check output, record completed steps, and
hand back to the issue — do not guess.
## References
- [`docs/agents/issue-tracker.md`](../docs/agents/issue-tracker.md)
- [`RUNBOOKS.md`](../RUNBOOKS.md)
+14
View File
@@ -1,5 +1,14 @@
# Runbook: mailcow health (mx2) # Runbook: mailcow health (mx2)
## Purpose
Read-only health check of the mailcow stack on mx2.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
- Read-only: does not change mailcow configuration or service state.
Target: [mx2.windy.me](../hosts/mx2.windy.me.md) Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Path: `/opt/mail` Path: `/opt/mail`
Prefer: the Ansible health report (`ansible/playbooks/health-report.yml --limit mailcow`), Prefer: the Ansible health report (`ansible/playbooks/health-report.yml --limit mailcow`),
@@ -71,6 +80,11 @@ The sanitized Ansible health profile is `mailcow` (`ansible/playbooks/healthchec
The server-local timer emits a sanitized result at `/var/lib/vps-health/latest.json`. The server-local timer emits a sanitized result at `/var/lib/vps-health/latest.json`.
It does not change Mailcow configuration or service state. It does not change Mailcow configuration or service state.
## Safety
- Read-only: never mutate configuration or service state during this check.
- If live state conflicts with an expected value below, `STOP` and report; do not "fix" on the fly.
## Pass criteria ## Pass criteria
- Compose stack up; watchdog ~100% - Compose stack up; watchdog ~100%
+14
View File
@@ -1,5 +1,14 @@
# Runbook: use mailcow SMTP / IMAP (client) # Runbook: use mailcow SMTP / IMAP (client)
## Purpose
Reference for configuring mail clients against the mailcow SMTP/IMAP endpoints.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md) client submission (587/465) and IMAP/POP (993/995).
- Not applicable: server-side mailcow configuration or administration.
Target: [mx2.windy.me](../hosts/mx2.windy.me.md) Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Prerequisite: a mailbox on `windy.me` (password from mailcow UI, not the admin account unless it is that mailbox). Prerequisite: a mailbox on `windy.me` (password from mailcow UI, not the admin account unless it is that mailbox).
@@ -52,3 +61,8 @@ Do not commit or paste real passwords into this repo.
- Port/TLS mode mismatch (587 vs 465) - Port/TLS mode mismatch (587 vs 465)
- Account active in mailcow; not rate-limited / fail2banned after bad attempts - Account active in mailcow; not rate-limited / fail2banned after bad attempts
- Apps that store SMTP in their own config (e.g. Vaultwarden `config.json`) may keep a **stale** password even when `.env` is correct — verify AUTH against the effective config ([vaultwarden-health](vaultwarden-health.md) §5) - Apps that store SMTP in their own config (e.g. Vaultwarden `config.json`) may keep a **stale** password even when `.env` is correct — verify AUTH against the effective config ([vaultwarden-health](vaultwarden-health.md) §5)
## Safety
- Do not commit or paste real passwords into this repo or chat.
- Use submission (587/465) for client sending; never use port 25 as a desktop/app outbound port.
+21
View File
@@ -1,9 +1,30 @@
# Runbook: mailcow update (mx2) # Runbook: mailcow update (mx2)
## Purpose
Update the mailcow stack on mx2 to the latest supported release.
## Scope
- Applicable: [mx2.windy.me](../hosts/mx2.windy.me.md), `/opt/mail`.
- Not applicable: config changes beyond the update, DB migration, secret rotation.
## Approval gates
| Action | Risk | Explicit approval |
|---|---|---|
| Run `./update.sh` (recreates containers, brief mail interruption) | Medium | Yes — user confirmation required |
Target: [mx2.windy.me](../hosts/mx2.windy.me.md) Target: [mx2.windy.me](../hosts/mx2.windy.me.md)
Path: `/opt/mail` Path: `/opt/mail`
**Confirm with the user before running an update.** **Confirm with the user before running an update.**
## Safety
- Never run the update without explicit user confirmation.
- Never pass secrets into the chat log; do not commit `mailcow.conf`.
- If a step fails, capture `docker compose ps` and logs and stop before further changes.
## Before ## Before
1. Run [mailcow-health](mailcow-health.md) (Ansible health report). Record baseline. 1. Run [mailcow-health](mailcow-health.md) (Ansible health report). Record baseline.
+15
View File
@@ -1,5 +1,14 @@
# Matrix Health Check # Matrix Health Check
## Purpose
Read-only health check of the Matrix homeserver (ESS on K3s).
## Scope
- Applicable: [synapse.chans.xyz](../hosts/synapse.chans.xyz.md), namespace `ess`.
- Read-only: does not change pods, ingress, certificates, or configuration.
Monitor the Matrix homeserver running on `synapse.chans.xyz` (ESS chart `26.7.2`, K3s node). Monitor the Matrix homeserver running on `synapse.chans.xyz` (ESS chart `26.7.2`, K3s node).
Prefer `cd ansible && ansible-playbook playbooks/health-report.yml --limit matrix` Prefer `cd ansible && ansible-playbook playbooks/health-report.yml --limit matrix`
@@ -90,3 +99,9 @@ Backup automation is currently paused. `/var/backups/matrix/` is retained for a
| Well-known returns 404/redirect | Root `chans.xyz` ingress missing or misconfigured | | Well-known returns 404/redirect | Root `chans.xyz` ingress missing or misconfigured |
| 502 Bad Gateway | Synapse pod restarting or DB down | | 502 Bad Gateway | Synapse pod restarting or DB down |
| SMTP emails not sent | MAS SMTP config incomplete; TCP reachable but AUTH failing — see `runbooks/vaultwarden-health.md` | | SMTP emails not sent | MAS SMTP config incomplete; TCP reachable but AUTH failing — see `runbooks/vaultwarden-health.md` |
## Safety
- Read-only: never mutate pods, ingress, certificates, or configuration during this check.
- Backup automation is paused; do not treat `/var/backups/matrix/` as a recovery source.
- If live state conflicts with an expected value, `STOP` and report.
+68
View File
@@ -0,0 +1,68 @@
# Runbook: controlled network change
## Purpose
Apply a controlled network configuration change (DNS records, firewall, LAN
gateway, VLAN) with impact assessment, approval, and a rollback path.
## Scope
- Applicable: PowerDNS zone records, `us4` firewalld allowlist, LAN gateway/VLAN/DNS changes, WireGuard.
- Not applicable: SSH access-policy changes (see `AGENTS.md` §SSH access safety — mandatory lockout-risk procedure).
## Preconditions
- A change record (Linear `vps` issue) describes the change, its reason, and rollback.
- Read-only impact assessment done (current config captured, blast radius known).
## Safety
- Never change DNS or network config without a change record and approval.
- Capture the current config first; never delete existing config as the first action.
- For DNS: record the current record values and TTL before editing.
- For firewall: retain an independent SSH rollback session before applying (see `ansible-operations.md` §us4).
## Procedure
### Step 1 — Assess and capture
**Action** — capture the current state (e.g. `dig` for DNS, `--check --diff` for firewall, `show` for gateway).
**Expected** — a baseline of current config and an identified blast radius.
**Decision** — change fully specified with rollback → Step 2; missing → `STOP`.
### Step 2 — Approve
**Action** — confirm approval is recorded in the issue/change record.
**Decision** — approved → Step 3; not approved → `STOP`.
### Step 3 — Change
**Action** — apply the single change (edit the record, run the gated playbook, or change gateway config) and only that change.
**Expected** — the new value/state is in effect.
**Verification** — re-query/verify the new state and confirm dependent services still pass health.
**Rollback** — restore the captured prior config and re-verify.
## Troubleshooting
### Troubleshooting A — Change broke dependent service
- Evidence: health report / endpoint failure.
- Allowed: roll back to the captured prior config.
- Next: verify; if still broken, escalate.
## Final Verification
- New state verified; dependent services healthy.
- Change and outcome recorded in the issue.
## References
- [`ansible-operations.md`](ansible-operations.md)
- [`rollback.md`](rollback.md)
- [`network-recovery.md`](network-recovery.md)
+61
View File
@@ -0,0 +1,61 @@
# Runbook: network outage / service recovery
## Purpose
Recover from a network outage or service failure, starting from read-only
diagnosis and mutating only when the root cause is confirmed.
## Scope
- Applicable: unreachable VPS services, LAN gateway/DNS failures, DNS resolution failures.
- Not applicable: planned changes (→ `network-change.md`), SSH access recovery (→ `AGENTS.md` §SSH access safety).
## Safety
- Read-only diagnosis first; do not mutate while the root cause is unknown.
- If live state conflicts with a runbook's assumptions, `STOP` and report.
- Keep the current verified management session open as the recovery path.
## Procedure
### Step 1 — Diagnose (read-only)
**Action** — gather evidence without changing anything:
```bash
# From laptop, pin DNS to a public resolver if the stub is flaky
dig @1.1.1.1 +short <host> A
curl -4 -sS -I --max-time 10 https://<host>/
# From a reachable host, inspect the service
ssh -4 windy@<host> 'docker compose ps -a; df -h /; tail -n 50 /var/lib/vps-health/latest.json'
```
**Expected** — a clear picture: is it DNS, connectivity, host, or service?
**Decision** — root cause localized → Step 2; ambiguous or conflicting → `STOP` and escalate (provider console if host is unreachable).
### Step 2 — Confirm and route
**Action** — match the failure to the owning runbook (`mailcow-health.md`, `pdns-health.md`, `matrix-health.md`, etc.) or `network-change.md` for a config fix.
**Expected** — an applicable runbook with a recovery action.
**Decision** — applicable → follow it; none → `STOP` (diagnose only, do not mutate).
### Step 3 — Recover (gated)
**Action** — apply only the runbook's documented recovery, with approval.
**Verification** — re-run the health report / endpoint check and confirm recovery.
**Rollback** — if recovery makes it worse, revert per `rollback.md`.
## Final Verification
- Service reachable and health report green.
- Incident and recovery recorded in the Linear `vps` issue.
## References
- [`network-change.md`](network-change.md)
- Per-service health runbooks under [`runbooks/`](.)
+15
View File
@@ -1,5 +1,14 @@
# PowerDNS health (hk2) # PowerDNS health (hk2)
## Purpose
Read-only health check of the `/opt/pdns` PowerDNS stack.
## Scope
- Applicable: [hk2.chans.xyz](../hosts/hk2.chans.xyz.md), `/opt/pdns`.
- Read-only: does not change PowerDNS, DNS records, or secrets.
Read-only checks for the `/opt/pdns` stack on **hk2.chans.xyz** (`ns1.wsvc.info`). Read-only checks for the `/opt/pdns` stack on **hk2.chans.xyz** (`ns1.wsvc.info`).
Facts: [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) · Upstream: [docs/pdns-upstream.md](../docs/pdns-upstream.md) Facts: [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) · Upstream: [docs/pdns-upstream.md](../docs/pdns-upstream.md)
@@ -87,6 +96,12 @@ Expect: `primary=yes`, `also-notify=202.91.35.141`, `only-notify=` empty, `gpgsq
The sanitized Ansible health profile is `pdns` (`ansible/playbooks/healthchecks.yml`). It runs locally through `vps-healthcheck.timer`, writes a sanitized JSON result to `/var/lib/vps-health/latest.json`, and uses the API key only inside the PowerDNS container. It does not modify PowerDNS, DNS records, or secrets. The sanitized Ansible health profile is `pdns` (`ansible/playbooks/healthchecks.yml`). It runs locally through `vps-healthcheck.timer`, writes a sanitized JSON result to `/var/lib/vps-health/latest.json`, and uses the API key only inside the PowerDNS container. It does not modify PowerDNS, DNS records, or secrets.
## Safety
- Read-only: never mutate PowerDNS configuration or DNS records during this check.
- Do not paste the API key into chat/logs.
- If live state conflicts with an expected value, `STOP` and report.
## After config changes ## After config changes
- `auth/pdns.conf`, `auth/templates.d/secrets.j2`, or auth-related `.env` → use the Ansible Compose reconcile playbook with target `auth` - `auth/pdns.conf`, `auth/templates.d/secrets.j2`, or auth-related `.env` → use the Ansible Compose reconcile playbook with target `auth`
+73
View File
@@ -0,0 +1,73 @@
# Runbook: release a reviewed change to production
## Purpose
Apply a reviewed configuration change to production hosts in a controlled,
verifiable, rollback-safe way.
## Scope
- Applicable: reviewed changes to Compose services, DNS, firewall, or Ansible-managed config on production hosts.
- Not applicable: destructive migrations (`vaultwarden-sqlite-to-postgres.md`), SSH access-policy changes (`AGENTS.md` §SSH access safety), or network changes needing a change ticket (`network-change.md`).
## Preconditions
- The change is reviewed and its intent matches a Linear issue / change record.
- Read-only health pass (or `maintenance-preview`) has been run and is green.
- A rollback point exists (backup, prior config, or known-good git/playbook state).
## Safety
- Never release without a verification and a rollback point.
- Never bypass gated confirmation variables.
- One host at a time; re-run health after each.
## Procedure
### Step 1 — Pre-flight
**Action** — run the read-only health report for the target host(s).
```bash
cd ansible && ansible-playbook playbooks/health-report.yml --limit <host>
```
**Expected** — green baseline.
**Decision** — green → Step 2; warning/critical → fix first (`fix-ci.md`), `STOP` the release.
### Step 2 — Approve
**Action** — confirm the change is approved in the issue/change record; confirm the window.
**Decision** — approval recorded → Step 3; missing → `STOP`.
### Step 3 — Apply
**Action** — run the gated playbook for the change, with its confirmation variable.
```bash
# example: reviewed Compose reconcile
cd ansible && ansible-playbook playbooks/compose-reconcile.yml --limit <host> \
-e '{"service_reconcile_confirm": true, "service_reconcile_targets": ["<svc>"]}'
```
**Expected** — playbook completes; service reports the new state.
**Verification** — re-run the health report and confirm green; confirm the specific change took effect (version, config, endpoint).
**Rollback** — on verification failure, revert the change and re-verify (see `rollback.md`).
## Final Verification
- Health report green post-release.
- The change is visible and the issue/change record is updated.
## Failure Handling
On failure: stop further changes, collect the playbook output and health report, and route to `rollback.md` if the change was applied.
## References
- [`ansible-operations.md`](ansible-operations.md)
- [`rollback.md`](rollback.md)
+56
View File
@@ -0,0 +1,56 @@
# Runbook: roll back a change
## Purpose
Restore a known-good state after a change failed verification or caused a
regression.
## Scope
- Applicable: Compose services, DNS records, firewall rules, and Ansible-managed config that have a rollback point.
- Not applicable: destructive migrations without a backup (route to the owning runbook, e.g. `vaultwarden-sqlite-to-postgres.md` §Rollback).
## Preconditions
- A rollback point exists: backup, prior config, or known-good commit/playbook state.
- The failing change and its symptoms are recorded.
## Safety
- Confirm the rollback target before acting; do not roll back unrelated changes.
- Verify after rollback; a rollback that is not verified is not complete.
## Procedure
### Step 1 — Confirm the target
**Action** — record the change to revert (playbook run, config edit, DNS change) and its intended prior state.
**Expected** — the exact prior state is known and reachable.
**Decision** — prior state known → Step 2; unknown → `STOP` and escalate.
### Step 2 — Revert
**Action** — apply the documented reversal for the change type:
- Compose config: re-apply the prior config and reconcile (or `git revert` the change then reconcile).
- DNS: restore the previous record value on PowerDNS.
- Firewall: re-apply the prior allowlist via the gated playbook.
- General repo change: `git revert <sha>`.
**Expected** — the prior configuration is in place.
**Verification** — re-run the relevant health report / endpoint check and confirm green.
**Rollback** — if the rollback itself fails, `STOP`; keep the original management session open and escalate.
## Final Verification
- Health report green after rollback.
- The incident and rollback are recorded in the Linear `vps` issue.
## References
- [`release.md`](release.md)
- [`ansible-operations.md`](ansible-operations.md)
+14
View File
@@ -1,5 +1,14 @@
# RustDesk server health (hk2) # RustDesk server health (hk2)
## Purpose
Read-only health check of the `/opt/rustdesk` server stack.
## Scope
- Applicable: [hk2.chans.xyz](../hosts/hk2.chans.xyz.md), `/opt/rustdesk`.
- Read-only: does not change RustDesk configuration or service state.
Read-only checks for the `/opt/rustdesk` stack on **hk2.chans.xyz**. Read-only checks for the `/opt/rustdesk` stack on **hk2.chans.xyz**.
Facts: [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md) Facts: [hosts/hk2.chans.xyz.md](../hosts/hk2.chans.xyz.md)
@@ -66,6 +75,11 @@ The sanitized Ansible health profile is `rustdesk` (deployed via
locally through `vps-healthcheck.timer`, writes a sanitized JSON result to locally through `vps-healthcheck.timer`, writes a sanitized JSON result to
`/var/lib/vps-health/latest.json`, and never exposes secrets. `/var/lib/vps-health/latest.json`, and never exposes secrets.
## Safety
- Read-only: never mutate configuration or service state during this check.
- If live state conflicts with an expected value, `STOP` and report.
## After config changes ## After config changes
Reviewed changes to `/opt/rustdesk/compose.yml` (relay address, image pin, port Reviewed changes to `/opt/rustdesk/compose.yml` (relay address, image pin, port
+15
View File
@@ -1,5 +1,14 @@
# Runbook: Vaultwarden health (us2) # Runbook: Vaultwarden health (us2)
## Purpose
Read-only health check of the Vaultwarden stack (Postgres backend) on us2.
## Scope
- Applicable: [us2.wsvc.info](../hosts/us2.wsvc.info.md), `/opt/vaultwarden`.
- Read-only: does not change Vaultwarden, Traefik, SMTP, or secrets.
Target: [us2.wsvc.info](../hosts/us2.wsvc.info.md) Target: [us2.wsvc.info](../hosts/us2.wsvc.info.md)
Path: `/opt/vaultwarden` Path: `/opt/vaultwarden`
URL: https://auth.wsvc.info/ URL: https://auth.wsvc.info/
@@ -96,6 +105,12 @@ Expect `AUTH_OK`. `535` usually means stale password in `config.json` (see step
The sanitized Ansible health profile is `vaultwarden` (`ansible/playbooks/healthchecks.yml`). It runs locally through `vps-healthcheck.timer`, writes a sanitized JSON result to `/var/lib/vps-health/latest.json`, and uses only server-side credentials for the SMTP AUTH probe. It does not modify Vaultwarden, Traefik, SMTP, or secrets. The sanitized Ansible health profile is `vaultwarden` (`ansible/playbooks/healthchecks.yml`). It runs locally through `vps-healthcheck.timer`, writes a sanitized JSON result to `/var/lib/vps-health/latest.json`, and uses only server-side credentials for the SMTP AUTH probe. It does not modify Vaultwarden, Traefik, SMTP, or secrets.
## Safety
- Read-only: never mutate configuration, containers, or secrets during this check.
- Never print passwords — compare lengths/hashes only.
- If live state conflicts with an expected value below, `STOP` and report.
## Pass criteria ## Pass criteria
- Healthy compose; `DATABASE_URL` points at `pg` - Healthy compose; `DATABASE_URL` points at `pg`
@@ -1,5 +1,20 @@
# Runbook: Vaultwarden SQLite → Postgres (us2) # Runbook: Vaultwarden SQLite → Postgres (us2)
## Purpose
Replay checklist for the exceptional SQLite→Postgres migration on us2.
## Scope
- Applicable: [us2.wsvc.info](../hosts/us2.wsvc.info.md), `/opt/vaultwarden`.
- Not applicable: routine reconcile — this is destructive and intentionally not automated.
## Safety
- Destructive: requires explicit approval; confirm backups exist before starting.
- Do not commit `.env` or password-bearing files to git.
- If any step diverges from this checklist, `STOP` and report.
Target: [us2.wsvc.info](../hosts/us2.wsvc.info.md) Target: [us2.wsvc.info](../hosts/us2.wsvc.info.md)
Upstream: [Using the PostgreSQL Backend](https://github.com/dani-garcia/vaultwarden/wiki/Using-the-PostgreSQL-Backend) · [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md) Upstream: [Using the PostgreSQL Backend](https://github.com/dani-garcia/vaultwarden/wiki/Using-the-PostgreSQL-Backend) · [docs/vaultwarden-upstream.md](../docs/vaultwarden-upstream.md)