Files
esh-pfi-infrastructure/docs/README.md
T
vh ad2b0e97aa docs/runbooks/nh3-prune-ritual: write up the quarterly NH3 prune ceremony
New runbook captures the three-phase process:

  Phase 1 — Drop --append-only via DSM Container Manager web UI
  Phase 2 — sudo resticprofile forget --prune --verbose on each of
            nh3-docker, nh3-dev, irv-ml1 (interactive sudo per host)
  Phase 3 — Restore --append-only via DSM

Why each phase looks the way it does, what to expect (largely no-op
runs for the first 6 months while no snapshots have aged out of the
keep window), how to verify each phase non-destructively (curl 401
on the rest-server root proves the container's up + serving), what
to do if Phase 2 fails with `repository is configured as append-only`
(skipped Phase 1 / DSM didn't apply), and the path to future
automation (find docker bin path on DSM, NOPASSWD-lock syncuser to
the specific recreate command).

Includes a "last run history" table seeded with today's first
post-pipeline run (no-op, irv-ml1 only had 3 snapshots due to the
04-25→27 CUDA stall).

Cross-referenced from docs/README.md (runbook tree), docs/
orientation.md (where-to-look table), and STATUS.md item 9 (which
now points at the runbook + records the next-round date 2026-07-27).
2026-04-27 20:54:08 -07:00

54 lines
2.2 KiB
Markdown

# docs/
Navigation map for the documentation tree. New session? Read
[`orientation.md`](orientation.md) first — it's the narrative overview
of the fleet, backup architecture, governing principles, and gotchas,
and it points at everything else.
## Tree
```
docs/
├── orientation.md # start here — fleet overview + where-to-look guide
├── runbooks/ # ops runbooks (recovery, deployment phases)
│ ├── disaster-recovery.md
│ ├── nh3-prune-ritual.md
│ └── pbs-deployment.md
└── pfi/ # PFI-specific reference (services, models, VMs)
├── docker-stack.md
├── model-list.md
├── proxmox-vms.md
├── recommended-model-settings.md
├── vm-102-matrix-appservice.md
└── vm-102-matrix-synapse.md
```
## What goes where
- **`runbooks/`** — step-by-step ops procedures. Anything you'd reach
for during an incident or while standing up new infrastructure.
Examples: disaster recovery (blast-radius tiers + restoration steps),
PBS deployment (9-phase rollout). New runbook → new file here.
- **`pfi/`** — PFI-specific reference material that's too narrow for the
top-level CLAUDE.md but doesn't change incident response. AI model
inventory, recommended inference settings, Matrix bridge config,
Proxmox VM map. New stable reference → new file here.
- **Top-level (`docs/orientation.md`, `docs/README.md`)** — narrative
guides about the workspace itself, not about specific infra.
## Cross-references
- Fleet topology + servers table: top-level [`CLAUDE.md`](../CLAUDE.md).
- Open work + recent milestones: top-level [`STATUS.md`](../STATUS.md).
- Durable cross-session facts: `~/.claude/projects/-home-lkraven-development-eshpfi-management/memory/`.
## Conventions
- Markdown, GitHub-flavored. CommonMark renders fine in most viewers.
- File names are lowercase-kebab-case, descriptive. No dates in
filenames — git history covers that.
- One topic per file. If a file grows past ~500 lines, look for a
natural split before adding more.
- No checked-in binaries or checksums. Build/release artifacts belong
in a build pipeline or `tools/`, not `docs/`.