15 Commits
Author SHA1 Message Date
vh 9e986d8ee8 feat(phasefinal-web): cloudflare edge config — cache ruleset + always online
Cache rule on www.phasefinal.com with edge and browser TTL both respect_origin,
so cache policy stays declared once in nginx.conf rather than split between the
repo and the dashboard. Always Online enabled, which is what actually survives
an origin outage; a 300s document TTL alone would only mask five minutes.

Verified: document and assets both reach cf-cache-status HIT, apex 301s to www,
edge email obfuscation active.
2026-08-29 23:02:49 -07:00
vh 524aa4d860 fix(phasefinal-web): healthcheck targeted ::1, so traefik skipped the container
The healthcheck used http://localhost/, which resolves to ::1 in nginx:alpine
while nginx listens on IPv4 only — so it never passed, the container stayed
unhealthy, and Traefik silently declined to create a router for it. That
presents as a broken docker provider: correct labels, right network, no route,
no error. Target 127.0.0.1 explicitly and add a start_period.

Adds the apex router (301 phasefinal.com -> www) and drops the file-provider
workaround, which was mitigating the wrong diagnosis.
2026-08-29 22:33:21 -07:00
vh 7ffbee6f09 feat(phasefinal-web): corporate site stack on ana-docker
Single static page (nginx) fronted by Traefik at www.phasefinal.com, built
from the design brief. Site markup/CSS checked in verbatim from the design
session; fonts self-hosted (SIL OFL) with the @font-face block uncommented,
which every fresh export re-comments.

Routed via a Traefik file-provider config rather than the container labels:
the docker provider on ana-docker was not registering newly-created
containers, so the file router avoids restarting shared ingress. Labels are
retained in compose so the file can be dropped once that is fixed.
2026-08-29 22:25:20 -07:00
vh c488eadc31 memory: snapshot — althing v3 fleet-wide at 3.1.1, sec on GPU0
The in-flight section was a day stale: it still described run 3c as the
live subject on a box where nothing had moved. Rewritten around what is
actually true now -- v3 deployed fleet-wide, the post office relocated
to nh3-docker, sec serving on GPU0, run 3c still held on power.

Six new decision entries, three of which carry findings that outlive
their incident: the inbound half of the handle-resolution bug (a stale
ALTHING_HANDLE reads another agent's mailbox and reports it empty,
which is a second route into the failure v3 exists to prevent), the OOM
attribution to Claude Code sessions, and the operator's two explicit
belays recorded so a later session does not re-raise them as new.

Auto-archival fired at 301 lines and moved exactly one entry. Three
others were old enough and every one carries a still-open deferred
pointer -- the parked CI flip, muninn-gate's submit path, and the triton
backend deferred to the Ada refresh. Held back per the guards; an
over-cap file that keeps live decisions beats a scannable one that lost
them. The entry that did move had its deferred item closed today:
nh3-extdev's staged v2.1.2 wheel is moot now that the box runs 3.1.1.
2026-08-28 22:22:32 -07:00
vh 583f329d00 fix(playbooks): 3.1.1 deploy — and why the markers match presence, not count
Deploys althing-core 3.1.1 to nh3-extdev. Both markers present on both
boxes; the warning verified behaviourally in four conditions rather than
by grep alone -- mismatched inherited handle warns, matching handle
silent, explicit --handle silent, unlaunched directory silent, and the
warning precedes the output it is about.

The release's own verification line says
`grep -c handles_launched_at dev_launch.py # 2+`. The real count there
is 1, the definition; the other two occurrences are in postbox.py. The
installed tree is byte-identical to the repo at the pushed tag, so the
instruction is wrong rather than the install.

This playbook matches on presence via grep -q, so it passed. Had it
asserted the stated count it would have reported FAILED on a perfect
deploy -- a verification instruction that fails on correct input, which
is the same false-negative this file has now produced three times in
different costumes. Recorded above the variable so the next bump does
not reintroduce a count.
2026-08-28 16:15:34 -07:00
vh 8a04d6f1bb fix(statusline): resolve the handle from the v3 binding, not the v2 map
The statusline resolved its handle from ~/.althing/session_handles.json.
forseti corrected the grounding and I verified it: that file is a v2
artifact and v3 never opens it. `grep -rn session_handles althing/` is
empty, postbox's resolve_config takes --handle then ALTHING_HANDLE and
nothing else, and `althing-cli use` -- the tool that maintained the map
-- was deleted at the cutover. Whatever is in it now is hand-kept and
drifts silently.

launch-history.json is written by dev_launch, which is the thing that
sets ALTHING_HANDLE in the first place, so it is the real cwd-to-handle
binding. Shape is {cwd: {command: {at, handle}}} with several commands
per directory, so this takes the most recent by timestamp rather than
whichever key happens to sort first. The v2 map stays as a fallback for
its broader coverage.

Worth recording why this was wrong: I wrote the resolution this morning
by reading the v2 statusline block it replaced and keeping its data
source while updating its commands. The commands were the visible half
of the cutover and the data source was not, so it survived a rewrite
that was otherwise about removing v2.
2026-08-28 16:01:11 -07:00
vh c648a40b68 fix(playbooks): 3.1.0 herald deploy; the marker check takes a LIST now
Deploys althing-core 3.1.0 to nh3-extdev and restarts the herald.
Verified by content on both boxes: POST_OFFICE_HINT 0 -> 4 in
post_office_herald.py and resolve_post_office 0 -> 3 in dev_launch.py,
dist-info 3.0.3 -> 3.1.0.

The check took one file:marker pair. 3.1.0 changed two files, so a
single pair would have asserted half a release and passed -- the same
half-passing-silently shape as the version-string check it replaced two
releases ago, one level up. It now takes a space-separated list, reports
each pair individually, and fails if any is missing. Every release's
markers so far are recorded above the variable so the next bump is a
lookup rather than an archaeology exercise.

Also verified the behaviour the release exists for rather than just its
markers. The herald writes its address to $ALTHING_ROOT/post-office and
dev_launch.resolve_post_office reads it when the variable is unset:

  env unset          -> http://10.100.50.40:8390
  env set            -> the env value, which wins
  env set to blank   -> the file, because blank counts as unset

My first attempt tested this through postbox, which still requires the
variable and reported "no post office address is configured" -- correct
behaviour that looked like a failed deploy. dev-launch is the reader,
not postbox.
2026-08-28 15:59:06 -07:00
vh 590b55f7d8 feat(playbooks): potrace/agg headers, with the two traps that mislead
pypotrace is an sdist that compiles at install time, so every machine
and every CI runner resolving it needs these headers first. That makes
it a recurring per-box action rather than the one-off it arrived as.

Two things learned installing it on nh3-dev are recorded here rather
than left in an althing thread, at forseti's suggestion, because a
thread is not where the next person looks:

Only libagg is a pkg-config consumer. potrace ships no .pc file and is
found via potracelib.h directly, so `pkg-config --exists potrace`
returns false on a correctly configured box. It looks exactly like the
cause and never is.

libagg's pkg-config modversion is 2.7.0 while its Debian package version
is 1:2.6.1-r134. Comparing those two numbers convinces you the wrong
package is installed.

The verify phase asserts the geometry, not the import: a square must
come back as one curve of four CornerSegments. An extension linked
against the wrong thing can import cleanly and return nonsense, so a
successful build is not evidence the module works.

Getting the build probe to run took three passes and the reason is worth
keeping. uv is not on a non-interactive ssh PATH; it is in a different
place on each box; and on nh3-dev it sits inside a 0700 home, so even
the correct absolute path fails `test -x` for the ssh user because the
directory cannot be traversed. The headers are system-wide and root's
business, but the build check is a developer action and has to run as
the user who owns the toolchain.
2026-08-28 13:38:12 -07:00
vh cdeb57c18b fix(playbooks): 3.0.3 herald deploy, and a content check that survives releases
Deploys althing-core 3.0.3 to nh3-extdev and restarts the herald.
Verified by content on both boxes: PANE_SETTLE_S 0 -> 2 occurrences,
value 0.3, dist-info 3.0.1 -> 3.0.3.

The content check was hardcoded to the 3.0.1 markers, so from the next
release onward it would have kept passing while asserting nothing about
what had just been installed -- a check that verifies the previous
release is indistinguishable from one that works. It now takes the
marker and file as variables, bumped per release, with both releases'
markers recorded so the pattern is obvious rather than folklore.

That is the same defect class as the install step gated on `postbox` not
existing, which this playbook carried until last round: a guard written
correctly for the first run and never re-read on the second.

The post office container was not touched. forseti established by
import graph that althing/post_office/* imports neither changed module
-- the fix is in reach_pane, which is herald code -- and the container
has been up two hours across both herald restarts.
2026-08-28 12:46:35 -07:00
vh 9f87ff87e1 feat(althing): surface the post office on Homepage under Toolchain
Labels the container into `Toolchain`, an existing group under the
existing Toolchain tab -- "the plumbing", which is where a message bus
belongs. Confirmed live: Homepage's API now returns it.

I had previously recorded in this file that no group fitted, which was
wrong. That conclusion came from a grep over the layout block that
missed the nested groups, and it went into a comment as though it were
a finding. The group was there the whole time.

Labels bind at container creation, so this deployed with `up -d` rather
than `restart`; a restart leaves the old labels and the dashboard keeps
showing what was there before. nh3-docker is already a discovered host
in homepage's docker.yaml as `nh3-pfi-docker`, so the label alone is
enough -- adding a services.yaml entry as well would render the card
twice.

althing-chamber on ana-docker also carries Toolchain labels and is a
separate service per the operator. Left alone.
2026-08-28 10:37:12 -07:00
vh dbb930d546 fix(playbooks): 3.0.1 herald reinstall, and two guards that were release-hostile
Reinstalls althing-core on nh3-extdev for 3.0.1 (the pane-route fix) and
restarts the herald. Both boxes verified BY CONTENT rather than by
version string -- forseti's own checks, grep for _PANE_ID and _live_pid,
because a dist-info directory records what was installed, not what the
files contain. Both went 0 -> 3 and 0 -> 2.

Two bugs in the playbook this run exposed, both of which only appear on
the second use:

The install step was gated on `postbox` not existing. That guard was
correct for the cutover, when postbox genuinely was absent, and wrong
for every release after it -- postbox exists now, so a version bump
would have silently skipped the install and the playbook would have
reported success having done nothing. `--force` already makes the
reinstall idempotent, so the guard bought nothing and cost correctness.

The post_office variable still pointed at nh3-dev, three hours after the
post office moved to nh3-docker. It failed in the verify rather than at
install time, which reads as a broken deploy rather than as a stale
constant. Worth noting the failure message was the outage semantics
working exactly as designed: "This is an outage, not an answer: do not
treat it as 'no mail'."
2026-08-28 10:27:03 -07:00
vh 22da609053 feat: registry-push the post office image; version the statusline
## Registry

The image moved by `docker save | ssh | docker load`, so a rebuild meant
repeating that by hand. It is now published and the compose pulls a
digest-pinned reference, so a redeploy is `compose up -d` on any host
that has logged in.

Pinned by digest rather than by tag: `:3.0.0` is a mutable pointer on a
registry anyone can re-push, and this container is the fleet's whole
message bus. The tag rides alongside so a human can read what it is.

Namespace is claude-bot, not vh. claude-bot's token carries
write:package and `docker login` succeeds, but package namespaces are
owned -- pushing to vh/ returns "unauthorized: authentication required"
after a successful login, which reads like a credential fault and is
actually an ownership one. Publishing under claude-bot's own namespace
also satisfies the standing directive to stop reusing the operator's
personal credentials for infra work, so the constraint and the policy
point the same way. Recorded in the compose header so the next person
does not read that error as a broken token.

Pull path proven rather than assumed: the running container was
recreated from the registry reference and its data verified afterwards.

## Statusline

Brought under version control because the v3 cutover broke it invisibly.
The segment gated on `command -v althing-cli`, a binary the cutover
deleted, so the unread badge and the armed bell silently vanished for
every session on the box. With 71 of 73 handles pull-only, that badge is
the only out-of-band signal telling a session with no armed waiter that
it has mail -- a dead statusline made a working bus look like an empty
one.

Canonical here, live at ~/.claude/statusline-command.sh, copies rather
than symlinks per the same rule as stacks/.
2026-08-28 10:08:05 -07:00
vh 9d4e7bd34a feat(althing): move the post office to nh3-docker
Operator directive, and a standing goal: the bus belongs on the docker
host. The flag-day deployment put it on nh3-dev because the herald lives
there -- but the herald is the piece that must be host-local, and the
post office is explicitly the piece that is not.

nh3-dev was wrong on three counts. Our own server table calls it "not a
Docker-stack host". It has had three OOM events in fourteen days with
the interval halving, and the confirmed hog is Claude Code sessions at
5-18 GB, which is that box's actual job. And mem_limit protects the
fleet from the post office while doing nothing in the other direction:
oom_score_adj was 0, an ordinary kill candidate, on a box whose last
sweep took althing-herald and uvicorn. The new deployment sets
oom_score_adj=-500.

The compose is now version-controlled here as a normal stack rather than
living only in the althing repo's deploy dir.

## docker stop does not checkpoint the WAL

The database was 155 KB with a 4.1 MB write-ahead log, and every recent
message was in the log. A clean container stop left it untouched -- an
explicit PRAGMA wal_checkpoint(TRUNCATE) was required.

A docker cp of the .db alone would have produced a database that opens
cleanly, passes integrity_check, serves the full 73-handle roster, and
is missing the day's mail, with nothing raising an error. Row counts
were verified at source, in the staged copy, and after seeding, because
the count is the only thing that separates those two outcomes.

The old volume is left in place. Not a rollback path, which the operator
ruled out -- just not deleting the only other copy on the day of a move.

## Follow-up left open

The image has no registry push and moves by save/ssh/load, so a rebuild
means repeating that by hand. It should join the gitea registry pattern
the other stacks use.
2026-08-28 10:02:13 -07:00
vh e58360668e feat: althing v3.0.0 cutover (U9b) and the sec seat onto GPU0
Two operator-authorised changes on the same afternoon.

## althing v3 (U9b flag day, one-way, no rollback)

The post office replaced the v2 P2P bus on nh3-dev and nh3-extdev.
One container is the only stateful component; heralds are one per box
and dial out; waiters are one per session. Every v2 command was deleted
rather than deprecated, so a script calling althing-cli now fails loudly
instead of silently talking to nothing.

73 handles seeded from the v2 CLI, which is authoritative over the v2
database's 91 agent rows -- the extra 18 are superseded names, a typo,
an underscore variant, and two machine-qualified handles that v3 makes
a category error. Verified by set difference in both directions rather
than by counting; a peer's "72 rendered" was a line-count artifact.

Deleted 5,043 orphaned wake FIFOs. The reason there were five thousand
is that v2 named them per session with the PID and never reaped them;
v3 names them per handle, so the leak is bounded by construction. That
is a fix in v3, not a cleanup we performed.

nh3-extdev needed its own path: althing lives there as a system wheel
under /opt/uv-tools with entry points in /usr/local/bin, its daemons
were system units rather than user units, and uv is not on the login
user's PATH. Captured as a rerunnable playbook rather than shell
history.

The v2 database is left inert on disk. There is no import path and none
was improvised.

## sec onto GPU0

GPU1 carries the five resident fleet seats and had ~28 GB free against
the ~51 GB this seat reserves, so it could not start there at all. GPU0
has been idle since run 3c was stopped. The compose header, the GPU pin
default and the homepage label all carried the old card number and are
corrected together -- a label that names the wrong GPU is a record that
lies about where the work runs.

Both playbooks carry verify phases that assert effective state. Two of
those verifies failed on green deployments while I was writing them:
one used a Go template that collided with the runner's own {{ }}
substitution, one omitted --handle so it failed on identity rather than
reachability. Both are fixed with the reason recorded inline, because a
verify that reports FAILED on a working system trains you to ignore it.
2026-08-28 07:19:18 -07:00
vh f875f746b8 feat(playbooks): nh3-dev memory forensics — and the OOM hog is Claude Code
forseti asked for journald kernel persistence plus sysstat, on the
premise that nh3-dev's three OOM events in 14 days left no evidence.

The premise was wrong. journald has been persistent all along: 15,068
kernel entries in the 82-day previous boot and 351 OOM records across
retained boots, full task tables included. `journalctl -b -1 -k`
returned one entry because it ran as a user in neither adm nor
systemd-journal, and journalctl shows only your own messages in that
case. The same artifact produced the "journal stops at 05:36:08 with no
shutdown sequence" claim -- the true boot -1 end is 05:47:04 with OOM
kills logged at 05:38, 05:40 and 05:42.

So the fix for "no evidence" is a group membership, not a logging
change: usermod -aG adm lkraven, which is the group Debian's journald
ACL names explicitly.

With the journal readable the attribution is already in it. The
versioned Claude Code binary lives at .local/share/claude/versions/,
so OOM victims named 2.1.220 / 2.1.177 / 2.1.168 are CC sessions, as
are those named claude. Every one of the twelve largest resident
processes ever recorded on this box is a CC session, topping out at
18.4 GB. Everything else killed is 30-55 MB collateral, which clears
the althing daemons by measurement rather than by their own sampling.

sysstat and atop are added because the journal records the moment of
the kill, not the ramp, and names the victim rather than the winner.
atop was not requested and is the one that matters: with a dozen panes
open, only a per-process timeseries says which session was growing.

Not done: a cgroup cap on CC sessions. It is the real mitigation and it
would kill long-running sessions mid-work, so it goes to the operator.
2026-08-28 06:03:43 -07:00
31 changed files with 3456 additions and 673 deletions
+4
View File
@@ -4,6 +4,10 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re
## Recent decisions (archived)
- `[2026-08-05]` **worldtree herald re-nudge bug root-caused → forseti shipped althing-core v2.1.2 (`d5d33df`, deployed on nh3-dev).** `herald.py:363` rendered the wake command from the empty *fresh* mail set on the re-nudge path (should be `deliver_msgs`) → `messages[0]` IndexError → un-suppressed outer catch-all → 7s crash-loop for 9 days on worldtree-codex's pane route (mimir-dev surfaced it; I traced it from the editable source). Fix + `render_command` empty-guard + outer log-suppress + 3 tests + contract amendment, all forseti's. **nh3-extdev herald 2.1.2 upgrade DEFERRED** (operator, not-now): extdev is a WHEEL install (not editable), unexposed (no pane routes); the verified 2.1.2 wheel is staged on nh3-dev `/tmp` (sha256 `003508…cef27`) — `uv tool install --force` + restart both heralds when un-parked. extdev herald-unit provenance resolved (operator-authorized 2026-07-25 via forseti relay; recorded in this file's 07-25 herald-install entry). auto-memory `reference_nh3_dev_althing_herald`.
_Archived 2026-08-28. Its deferred item — the nh3-extdev herald 2.1.2 upgrade — is closed: extdev went 2.1.0 -> 3.0.0 -> 3.1.1 at the v3 cutover, so the staged v2.1.2 wheel is moot._
# eRP dual-seat overhaul — MeroMero-v2 + Dark-Scarlett, NVFP4A16 @ 256K on ana-ml2
`[2026-08-12]` Replaced the two legacy char-rp seats with home-quantized NVFP4A16 vLLM
+1652 -624
View File
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,218 @@
# `[2026-08-28]` althing v3.0.0 flag day (U9b) — the post office replaced the P2P bus, one-way
Operator-authorised, executed by infra-ops. **v2 is gone from both boxes: every v2 command was
deleted, not deprecated.** No rollback was designed or tested; failures are fixed forward.
post office ONE container on nh3-dev, http://10.100.50.40:8390
the only stateful component. SQLite + FTS + the typed API + the operator page.
herald althing-po-herald, ONE PER BOX, supervised. Dials out, opens no port,
holds no state. Refuses to start if another herald holds the node.
waiter althing-listen, one per session. Holds a FIFO, blocks, exits when poked.
client postbox (+ althing-mcp for the stdio tool surface)
althing-cli -> postbox althing-wake-listener -> althing-listen
althing-light-monitor -> GONE althing-receiver -> GONE
althing-herald -> althing-po-herald
**Every session needs BOTH env vars**: `ALTHING_POST_OFFICE=http://10.100.50.40:8390` and
`ALTHING_HANDLE` (dev-launch sets the latter per pane). postbox has **no default post-office
address** — a bare `postbox status` errors out rather than guessing.
## ⚠ An unreachable post office is an OUTAGE, never an empty inbox
v2 could not distinguish those. v3 can, and the distinction only pays if it is honoured: if
postbox says it could not reach the post office, that is the fault. Do not read it as "no mail".
## Deployment facts worth not rediscovering
- **`network_mode: host` is deliberate — do NOT "fix" it into a bridged container with `-p`.**
`api.py`'s `resolve_bind_host` refuses any address that resolves to a wildcard and tests the
RESOLVED PROPERTY rather than matching strings, so `ALTHING_BIND_HOST=0.0.0.0` is structurally
impossible. Reachability on the private network IS the authorisation story; there is no login.
Bridging would move access control to a `-p` flag the application cannot see.
- **Compose schema is `2.4` on purpose.** nh3-dev has docker-compose **1.29.2 (v1 only, no
`docker compose` subcommand)**, where `mem_limit` is honoured only under 2.x; under 3.x it
moves to `deploy:` which is swarm-only and SILENTLY IGNORED. Verified honoured:
`docker inspect -> 536870912`.
- **nh3-extdev is the box a `git pull` cannot move.** althing is a system WHEEL at
`/opt/uv-tools/althing-core` with entry points in `/usr/local/bin`; its v2 daemons were
**SYSTEM** units, not user units; and `uv` is NOT on lkraven's PATH there — it lives at
`/home/infra-ops/.local/bin/uv`. Install path: build the wheel on nh3-dev, stage to /tmp,
`sudo env UV_TOOL_DIR=/opt/uv-tools UV_TOOL_BIN_DIR=/usr/local/bin <uv> tool install --force`.
Playbook: `playbooks/nh3-extdev-althing-v3.yaml`.
- nh3-dev's herald is a **user** unit at `~/.config/systemd/user/althing-po-herald.service`
(Environment=ALTHING_POST_OFFICE, Restart=always); nh3-extdev's is a **system** unit at
`/etc/systemd/system/althing-po-herald.service` with `User=lkraven`.
## ⚠ THE FIFO LEAK WAS STRUCTURAL, AND v3 FIXED IT
5,043 orphaned wake FIFOs were deleted from `~/.althing/wake/`. The reason there were five
thousand: **v2 named them per SESSION with the PID** (`advisor-dev-listener-2448686.fifo`) and
never reaped them, so every pane that ever armed left one behind permanently. **v3 names them
per HANDLE** (`infra-ops.fifo`). Bounded at 73 by construction. That is a fix, not a cleanup.
## The roster: 73, and the authoritative source is the CLI, not the DB
Seeded via `POST /tool/declare` with an `X-Althing-Handle` header (a request without one is
rejected `bad_request`). `declare` is an OPERATOR verb — not in postbox, never will be.
⚠ The v2 **database** held 91 agent rows; the v2 **CLI** listed 73. The extra 18 were legacy —
superseded handles (`bifrost` -> `bifrost-dev`), one typo (`galdrbok`), an underscore variant,
and two machine-qualified handles (`ldp-dev@nh3-extdev`, `mailman@nh3-extdev`). **v3 abolishes
@-qualification: identity no longer has a home, so an @-qualified handle is a category error.**
Verified after seeding by SET DIFFERENCE in both directions, not by counting — forseti's
"72 rendered vs 73 declared" was a line-count artifact and the sets are identical.
## v2 history is inert, not migrated
`~/.althing/althing.db` — 95 MB, 12,437 messages — is untouched on disk. **No import path
exists and none should be improvised.** Plain SQLite if something must be recovered by hand.
## Two-agent flag day: the collision worth remembering
forseti and I both ran `scripts/sync_skill.sh` 65 seconds apart (13:29:18Z / 13:30:23Z). My
`--check` said "in sync" BEFORE I ran it — that was forseti a minute earlier, and I read it as
"already done at some point" rather than "someone is working in here right now." Cost was one
redundant backup. **Two agents worked the same checklist with no ownership marked per line.**
Next flag day: name an owner per item.
## Peers notified individually (operator-directed), NOT broadcast
`eitri-smithy-dev`, `dvalin-smithy-dev`, `bil-smithy-dev` carried their own 2.1.x SKILL.md
copies; `regin-smithy-dev` was named by the operator though no copy is visible on nh3-dev.
Operator's ruling on a general fleet announcement: **pointless in both directions — anyone
already on v3 knows, anyone not on v3 cannot receive it.** Consistent with the standing
never-broadcast-unsolicited directive.
⚠ `sync_skill.sh` deliberately does NOT write into peer repos — althing owns canonical
(`vh/althing @ v3.0.0 : skills/althing/SKILL.md`) and peers pull. Nobody does it for them.
---
## `[2026-08-28, same day]` MOVED to nh3-docker — and the WAL nearly ate the mail
Operator: *"I want it on the docker machine — that was always the goal."* The flag-day
deployment put the post office on **nh3-dev**, which was wrong on three counts:
- our own server table calls nh3-dev **"not a Docker-stack host"**; NH-site non-GPU → nh3-docker
- nh3-dev had **three OOM events in fourteen days, interval halving**, and the confirmed hog is
CC sessions at 5-18 GB — the box's actual job. See [[2026-08-28-nh3-dev-oom-attribution]].
- `mem_limit: 512m` protects the fleet **FROM** the post office. It does nothing to protect the
post office **from the box**: `oom_score_adj` was 0, an ordinary kill candidate, and the 08-28
sweep took althing-herald and uvicorn. A sweep taking the post office takes all 73 handles.
NOW http://10.100.50.40:8390 nh3-docker, oom_score_adj=-500, mem 512m verified honoured
WAS http://10.100.10.50:8390 nh3-dev (address now refuses)
canon stacks/althing-post-office/compose.yaml -> /opt/docker/compose/ on nh3-docker
## ⚠ THE DURABLE FINDING — `docker stop` does NOT checkpoint the SQLite WAL
post_office.db 155 KB mtime 15:09
post_office.db-wal 4.1 MB mtime 16:56 <- every recent message lived HERE
I expected a clean container stop to checkpoint. **It did not** — after `docker stop` the WAL was
still 4,124,152 bytes, unchanged. An explicit `PRAGMA wal_checkpoint(TRUNCATE)` was required,
after which the .db grew 155 KB → 163,840 B and the WAL/-shm vanished.
**A `docker cp` of `post_office.db` alone would have produced a database that opens cleanly,
passes `PRAGMA integrity_check`, serves the full 73-handle roster — and is missing the day's
mail. Nothing would have errored.** Only a row count distinguishes those two outcomes.
**Procedure for moving any WAL-mode SQLite service: stop → EXPLICIT checkpoint → verify counts →
copy → verify counts again on the far side, before deleting anything.** Not stop → copy. Verified
8 messages / 73 handles / 2 nodes / 8 recipients at source, in the staged copy, and after seeding.
## Repoint list (everything that names the address)
~/.config/systemd/user/althing-po-herald.service nh3-dev herald (user unit)
/etc/systemd/system/althing-po-herald.service nh3-extdev herald (system unit)
~/.claude/statusline-command.sh the hardcoded statusline fallback
every session's ALTHING_POST_OFFICE + re-arm althing-listen
⚠ **Do NOT blind-sed `10.100.10.50:8390` across the memory tree.**
[[reference_corviduo_dev_emergency_ops]] carries that exact string as a **Bifrost "affect" plane**
entry in the personal Worldtree's `BIFROST_CLIENT_ALLOWED_HOSTS` — an unrelated service that
happens to share the port. Incidentally, moving the post office off nh3-dev:8390 also cleared a
latent collision with it.
## Evidence the outage semantics work under a real outage
During the gap the nh3-dev herald logged, verbatim: *"push outage on nh3-dev: the post office did
not answer, so push is DOWN on this box. **This is an outage, not an empty poke list.**"* Then
after the repoint: `10:00:27 INFO poked infra-ops on nh3-dev via fifo (rung 0)` — poke path
re-verified end to end.
## Open follow-up
**The image has no registry push.** It moves by `docker save | ssh | docker load`, so a rebuild
means repeating that by hand. The fleet pattern (skaldsong, soong-lab) is a gitea registry pull;
this should join it. The old nh3-dev volume is left in place untouched — not a rollback path
(the operator ruled that out), just not deleting the only other copy on the day of the move.
---
## `[2026-08-28, later]` The release train: 3.0.1 → 3.1.1 in one afternoon, and the registry
Six releases landed the same day as the cutover. **The container was touched exactly once** (the
nh3-docker move); every other release was herald- or client-side, established each time by
forseti's **import-graph argument** — asking what `althing/post_office/*` IMPORTS rather than
reading the diff. A diff tells you what moved; an import graph tells you what can be affected.
3.0.1 pane routes (a pane agent can register its own route)
3.0.2 the docs are now checked against the CLI, not against each other — NO redeploy
3.0.3 PANE_SETTLE_S=0.3, the write/submit race
3.1.0 herald writes $ALTHING_ROOT/post-office; dev-launch reads it — BOTH binaries
3.1.1 warn when ALTHING_HANDLE disagrees with launch-history
**Deploy recipe per release:** `uv build --wheel` on nh3-dev → `uv tool install --force .`
locally → stage the wheel to nh3-extdev and install under `UV_TOOL_DIR=/opt/uv-tools` with
`/home/infra-ops/.local/bin/uv` → restart both `althing-po-herald` units.
`playbooks/nh3-extdev-althing-v3.yaml` does the extdev half.
## ⚠ THE REPEATED DEFECT — a verify that half-passes, four costumes in one day
Every one of these was written correctly for its first run and silently wrong on the next:
1. install step gated on `postbox` not existing -> right for the cutover, would SKIP
every release after and report success
2. content check pinned to the PREVIOUS release's markers -> passes forever, asserts nothing
3. ONE file:marker pair against a TWO-file release -> asserts half a release
4. a Go template `{{ }}` inside elway, whose own substitution ate it -> FAILED on a green deploy
**The check now matches on PRESENCE (`grep -q`), never a count**, and takes a LIST of
`file:marker` pairs bumped per release. 3.1.1's own release note said
`grep -c handles_launched_at dev_launch.py # 2+`; the real count there is 1 (the definition,
with 2 in postbox.py), so a count assertion would have reported FAILED on a byte-perfect install.
## The registry, and why the namespace is `claude-bot`
gitea.phasefinal.com/claude-bot/althing-post-office:3.0.0@sha256:410fed41...
Digest-pinned, not tag-floating — a tag is a mutable pointer on a registry anyone can re-push and
this container is the whole bus. ⚠ **claude-bot's token carries `write:package` and `docker
login` SUCCEEDS, but package namespaces are owned**: pushing to `vh/` returns
`unauthorized: authentication required` AFTER a successful login — an ownership refusal wearing a
credential error's clothes. Publishing under claude-bot's own namespace also satisfies the
standing directive to stop reusing the operator's personal credentials.
## The deployed CC plugin copies are a release step NOBODY owns
`sync_skill.sh` covers the SKILL, not the plugin. Nothing in the repo reaches
`~/.local/share/althing-plugin/` or `~/.claude/plugins/cache/althing/althing/0.0.1/`. Both must be
`rsync -a --delete`'d from the repo's `plugin/` **on every release**, by hand, or they carry the
previous release's bugs into the live surface — which happened: forseti's new docs-vs-CLI check
found `postbox reply --to` twice in `plugin/commands/inbox.md`, **the file a CC session reads
every time it drains its inbox**, and both my deployed copies had it.
⚠ **A running CC session keeps the plugin text it loaded at startup.** Files being right is
necessary and not sufficient — the session has to restart. I synced the copies at 09:30 and was
handed the v2 `/althing:monitor` text an hour later, calling three deleted binaries.
## Peer skill copies: I stopped hand-syncing, deliberately
I wrote into eitri/dvalin/bil's repos three times (operator-authorised, and right while they were
dark and could not pull). **Once they were awake and pulling, it became a race I was losing** —
canonical moved three times in an hour and I was chasing a one-line version-banner lag. It also
cost provenance: dvalin had to correct their own account of their file because my write looked
like a pre-existing partial. Two writers, no lock. `sync_skill.sh` deliberately does not write
into peer trees; althing owns canonical and peers pull. Respect that boundary once they can.
@@ -0,0 +1,60 @@
# `[2026-08-28]` A stale `ALTHING_HANDLE` silently reads ANOTHER agent's inbox and reports it empty
Found by pewpew-dev, whose session posted as **forseti** all day. Mechanism:
ALTHING_HANDLE inherited from the environment
+ nothing binding a shell to the handle it may query
= the env var wins, silently, with no warning and no error
Measured, same shell, same second, no credential, no complaint:
ALTHING_HANDLE=infra-ops postbox status -> infra-ops' mailbox
ALTHING_HANDLE=forseti postbox status -> forseti's mailbox
ALTHING_HANDLE=pewpew-dev postbox status -> pewpew-dev's mailbox
## ⚠ THE INBOUND HALF IS THE SERIOUS ONE, AND IT IS INVISIBLE
**Outbound** mis-signing sometimes gets caught: a peer notices the sender cannot hold that
context — which is exactly how this surfaced, when forseti was asked a pewpewstudio question.
**Inbound never does.** A session with a stale handle runs `postbox peek`, reads SOMEONE ELSE'S
mailbox, and is told — confidently, correctly, nothing broken anywhere — that it has nothing
unread. Measured cost: my reply sat unread until pewpew-dev's operator asked whether they were
blocked on me.
**This is a SECOND route into the failure v3 exists to prevent.** The guarantee — *"an
unreachable post office is an OUTAGE, never an empty inbox"* — holds, and does not cover this:
the post office is reachable and answers correctly, **about someone else**. Nothing is down, so
the outage semantics never fire. forseti has stopped describing that guarantee as though it
closes the empty-inbox class; it closes one route into it.
## ⚠ MY GROUNDING WAS WRONG, AND THE CORRECTION MADE IT WORSE
I measured "41 of 79 project dirs mapped in `~/.althing/session_handles.json`" and called the 46
unmapped ones exposed. **`session_handles.json` is a v2 artifact that v3 never opens** —
`grep -rn session_handles althing/` is empty, and `postbox.resolve_config` takes `--handle` then
`ALTHING_HANDLE` and nothing else. So the exposure is LARGER than my number implied: every
directory is in the same position, because the map is consulted for none of them.
**The same stale data source had survived inside my statusline rewrite that morning.** I updated
the v2 block's COMMANDS and kept its DATA SOURCE — the commands were the visible half of the
cutover and the source was not. Fixed (`8a04d6f`): it reads `launch-history.json` now (written by
`dev_launch`, shape `{cwd: {command: {at, handle}}}`, take the most recent by `at`).
## The fix that shipped, and why my proposal was the worse one
I proposed printing the resolution source (`handle: forseti (from ALTHING_HANDLE)`). forseti
killed it with one observation: **`postbox status` already prints the handle, first field, every
call.** pewpew-dev had `forseti` on screen and it did not register. **The information was never
missing; the salience was.** That generalised into the rule that chose the design: a line that is
always there teaches the reader to skip it, so a warning beats a field.
Shipped as **3.1.1** — warn (never refuse; a legitimate cross-project send is real) when the cwd's
`launch-history` names a different handle, on stderr, BEFORE the output it is about. Keyed by
directory across commands, not `(directory, command)` — one repo hosting a claude and a codex
session under different handles is normal, and a warning that fires on a legitimate case is one
people learn to ignore.
⚠ **Its silence is not an all-clear:** 8 of 79 directories are in `launch-history.json`, so the
quiet case is ~90% of the box. forseti put that caveat in a TEST NAME
(`test_an_unlaunched_directory_says_nothing`) because prose gets skimmed and a test goes red.
@@ -0,0 +1,71 @@
# `[2026-08-28]` nh3-dev's three OOM events attribute to Claude Code — and the "no evidence" was a permissions artifact
Three memory-exhaustion events in 14 days (08-14 00:15, 08-26 09:58, 08-28 05:36, interval
halving). forseti reported none could be attributed because *"kernel messages are not being
persisted to journald"* and asked for journald persistence + sysstat.
## ⚠ THE PREMISE WAS WRONG — journald was persistent the whole time
journalctl -b -1 -k privileged 15,068 entries (82-day boot)
journalctl -b -1 -k unprivileged 4 entries
OOM records, all retained boots 351
`journalctl` **silently shows only your own messages** when you are in neither `adm` nor
`systemd-journal`, and prints the reason as a scroll-past hint. Two of the three "no evidence"
findings were that one artifact:
"kernel messages not persisted" -> they are, and every OOM task table is there
"journal stops 05:36:08, no shutdown" -> that is the USER's last entry; the true
boot -1 end is 05:47:04, with OOM kills
recorded at 05:38, 05:40, 05:42
**Fix was `usermod -aG adm lkraven`, not a logging change.** Debian's journald ACL names `adm`
explicitly (`getfacl /var/log/journal/<machine-id>` → `group:adm:r-x`). Existing shells keep
their old group set — re-login, or `sg adm -c '...'`, which is also how to *verify* the grant
took rather than grepping `/etc/group`.
## ⚠ THE HOG IS CLAUDE CODE
`/home/lkraven/.local/share/claude/versions/2.1.220` is the versioned CC binary, so OOM victims
named `2.1.220` / `2.1.177` / `2.1.168` are CC sessions, as are those named `claude`.
largest anon-rss ever recorded on this box
18,434,696 kB 2.1.177 18.4 GB
15,788,764 kB 2.1.220 15.8 GB
15,154,008 kB 2.1.220 15.2 GB
14,994,376 kB 2.1.168 15.0 GB
-> every one of the top TWELVE is a CC session
29 of the OOM victims are CC. Everything else killed — althing-forseti (22), caddy (15),
ttyd (11), zellij (6), the althing daemons — is 30-55 MB **collateral**, the OOM killer
scraping for a few hundred MB. The althing v2 daemons are cleared by measurement.
**"claude is 408 MB each" is a YOUNG session.** Mature ones measure 5.4-18.4 GB. On 27 GB with
974 MB swap the ceiling is **three or four mature sessions**, not the ~66 a 408 MB figure
implies. Aug 28's task table: two CC at 5.4 GB + three zellij servers at 1.14 GB.
## Instrumentation added (`playbooks/nh3-dev-memory-forensics.yaml`, idempotent)
sysstat system-wide mem/CPU, 5-min cadence (not Debian's 10 -- a CC session can
add several GB inside one 10-minute bucket). sar -r
atop PER-PROCESS, 60s, 7-day retention. atop -r /var/log/atop/atop_YYYYMMDD
journald unchanged, already persistent, now READABLE
**atop is the one that matters and it was not requested.** The journal records the moment of
the kill and names the *victim*; sar says the box filled up; only atop says **which session was
growing and how fast** — the whole question when a dozen panes are open.
## Open — operator's call, deliberately not taken
**A cgroup memory cap on CC sessions is the real mitigation and it would kill long-running
sessions mid-work.** Surfaced, not decided. Instrumentation makes event four *diagnosable*, not
less likely.
## Lesson that generalises
A verify step I wrote failed while the setting was live: I grepped
`systemctl show sysstat-collect.timer` for my own input `*:00/05`, but systemd normalises it to
`*-*-* *:00/5:00`. **Assert the effective value, not the string you wrote** —
[[feedback_assert_effective_value_not_substring]], caught here in my own instrumentation.
Reply: althing msg `01M147EWEZDT8Y0XW5FTHHEAQC`, thread `01M1472ST5DSJNHR676X9F43AK`.
@@ -0,0 +1,35 @@
# `[2026-08-28]` The `sec` pen-test seat moved GPU1 → GPU0 and came up — on a circuit that tripped 36h earlier
Operator-directed. `sec` = M.O.G.-SEC-27B (`stacks/mog-sec/`, LiteLLM aliases `sec` /
`sec-reasoning`, ana-ml2 `:8019`, 262K native ctx, NVFP4+FP8 mixed with a grafted MTP head).
## Why it had to move
GPU1 69,895 MiB used of 97,887 — gen 46 GB + embed 9.8 + coder 8.4 + rerank 3.5 + reward 2.1
-> ~28 GB free, against the ~51 GB this seat reserves at MOG_GPU_MEM_UTIL=0.52
-> it could not start on GPU1 AT ALL
GPU0 empty since run 3c was stopped 2026-08-26
One line: `MOG_GPU_ID=1 -> 0` in `/opt/docker/compose/mog-sec/.env`. The compose default, the
header comment and the **homepage label** all named GPU 1 and were corrected in the same change —
a label naming the wrong card is a record that lies about where the work runs.
## Landed state
container vllm-mog-sec, healthy after ~400s load (22 GB model off the DEGRADED /tank)
GPU0 51,532 MiB, idle draw 16.34 W
GPU1 69,895 MiB, idle draw 6.42 W
⚠ **This re-arms the two-GPU load condition that tripped the Anaheim rack breaker on 08-26.**
Idle draw is negligible — ~23 W across both cards. **The risk materialises only under concurrent
load**, when both seats work at once and the box approaches the ~600 W that tripped it. The
operator accepted that with the constraint stated. See
[[2026-08-27-anaheim-breaker-and-onboot-gap]] — one circuit feeds the whole rack including ana-gw
and ana-wg, so a trip costs the site AND the way back in.
## Deploy gotchas worth keeping
- `up -d`, never `restart` — **labels bind at container creation**, so a restart keeps the old
homepage label and the dashboard silently keeps showing the old GPU number.
- Diff deployed-vs-canonical BEFORE pushing. There was no drift here, which is the only reason
the push was safe to make blind.
+20 -43
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-08-27_
_Last updated: 2026-08-28_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -27,7 +27,7 @@ Sister repos (separate gitea repos, deployed by playbooks here):
| `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
| `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
| `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
| `vh/althing` | Lean trusted inter-agent message bus — **v2 "email model" (v2.0.0b2, 2026-07)**: per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API `/owner/*` / `althing-mcp` stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. | per-box `uv tool install` (NOT CI-deploy); **nh3-dev = the DEV box** (editable install of `~/development/althing`, gets new versions first); **nh3-extdev** a mesh peer (model B: althing-svc + shared `/srv/althing`) |
| `vh/althing` | Lean trusted inter-agent message bus — **v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback)**: ONE container on nh3-dev at `http://10.100.50.40:8390` is the only stateful component; `althing-po-herald` one per box; `althing-listen` one per session; `postbox` is the client. **Every v2 command was DELETED, not deprecated** — `althing-cli`→`postbox`, `althing-wake-listener`→`althing-listen`, `althing-light-monitor`/`althing-receiver` gone. Sessions need BOTH `ALTHING_POST_OFFICE` and `ALTHING_HANDLE`; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → `persistent-memory.d/2026-08-28-althing-v3-cutover.md` | per-box install (NOT CI-deploy); **nh3-dev** = container host + repo; **nh3-extdev** = system WHEEL at `/opt/uv-tools`, needs its own wheel install (`playbooks/nh3-extdev-althing-v3.yaml`) |
| `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
| `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
| `vh/Worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. **gitea-runner builds on ana-docker**; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
@@ -108,17 +108,27 @@ no longer deployed sidecars here. See Recent decisions.)
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
## Current state / in-flight
_State as left 2026-08-26 23:05 PDT (written 2026-08-27 07:27, corrected 10:5x) — **run 3 is trained, gated and DO-NOT-SERVE on a safety finding. Run 3c is STOPPED at last-logged step 24/604, killed deliberately at 2026-08-26 21:07:40 PDT after Anaheim tripped a power breaker.** Nothing is training._
_As of 2026-08-28 16:20 PDT — **althing v3 is live fleet-wide at 3.1.1 with the post office on nh3-docker. `sec` is serving on ana-ml2 GPU0. Nothing is training.**_
- **⏸ RUN 3c HELD — operator stopped it, power capacity is the blocker.** lr `2e-4 -> 1e-5`, corpus BYTE-IDENTICAL, single variable proven by diff. Config `/tank/erp-tune/run-03c.json` is built and validated (`save_steps 50`, both deviations recorded as separate entries — the scientific lr change and the operational checkpoint cadence). Relaunch is one command. **Do not relaunch until the power triage lands** — ana-ml2 pulls ~600 W across both GPUs at their caps while training, and that is what tripped the breaker. **Exactly TWO 3c launches, and only one of them died:** #1 17:53:33 → killed by the power loss at step 80/604; #2 20:58:41 → stopped BY ME at 21:07:40, healthy, on operator instruction. → `persistent-memory.d/2026-08-27-run3c-launch-count-reconstruction.md`
- **🔴 `/tank` DEGRADED on ana-ml2 — a disk is genuinely gone**, 7 physical NVMe where the pool expects 8. raidz2, one parity disk spent, no data errors. **Operator is replacing it** (Supermicro AS-4125GS-TNRT2, PCIe hot-plug, should not need a power-down). → `persistent-memory.d/2026-08-27-anaheim-breaker-and-onboot-gap.md`
- **🟢 ADAPTERS BACKED UP OFF-SITE** — `nh3-nas:/volume1/smithy/erp-tune-adapter-backup/{run-01,run-02,run-03}`, sha256 verified at source, staging and rest (944 MB). They were single-copy mode-600 on the degraded pool. The four `merged-run03*` models are NOT backed up **by design** — all are derived from `run-03/adapter` by a documented verified merge, ~10 min each to regenerate from a 315 MB artifact that is now safe.
- **⏸ WEEKEND: power triage, "probably shut down some seats"** (operator). `gen` stays up by instruction. Everything else on ana-ml2 is idle. Sheddable: `vllm-gen` 46 GB (the big one), embed 9.8 GB, coder 8.4 GB, rerank-a3 3.5 GB, reward 2.1 GB, scriberr.
- **⏸ Owed to brokkr-smithy-dev when there is a card again:** nothing blocking. They hold the dose curve, the entanglement finding, and an unsourced "~60%" figure in R47 §5 they were checking against MeroMero v2's card — if it is not there either, §5's recommendation loses its basis.
- ⚠ **`/mnt/smithy` will be MISSING after every ana-ml2 reboot** — manual by design, not an oversight. Remount spec in `persistent-memory.d/2026-08-23-smithy-mount-ana-ml2.md`. Do NOT add it to fstab.
- **⏸ RUN 3c STILL HELD — power capacity, unchanged.** lr `2e-4 -> 1e-5`, corpus BYTE-IDENTICAL, config `/tank/erp-tune/run-03c.json` validated, relaunch is one command. **Do not relaunch until the power triage lands.** ⚠ `sec` now occupies GPU0 (~51 GB), so a 3c relaunch needs GPU0 freed OR accepts three-way contention. Exactly TWO 3c launches, only one died. → `persistent-memory.d/2026-08-27-run3c-launch-count-reconstruction.md`
- **🟢 `sec` (M.O.G.-SEC-27B) IS UP on ana-ml2 GPU0 :8019**, operator-directed. 51,532 MiB, idle 16 W. ⚠ Re-arms the two-GPU load condition that tripped the rack breaker; the risk is concurrent load, not idle. → `persistent-memory.d/2026-08-28-sec-seat-gpu0.md`
- **🟢 althing v3.1.1 on both heralds; post office on nh3-docker `http://10.100.50.40:8390`.** Container untouched since the move. All 73 handles seeded; 5 agents push-reachable (infra-ops, forseti, and the four smithy peers via pane routes). → `persistent-memory.d/2026-08-28-althing-v3-cutover.md`
- **🔴 `/tank` DEGRADED on ana-ml2** — 7 physical NVMe where the pool expects 8, raidz2, one parity spent, no data errors. **Operator replacing it.** → `persistent-memory.d/2026-08-27-anaheim-breaker-and-onboot-gap.md`
- **⏸ WEEKEND: power triage, "probably shut down some seats"** (operator). Sheddable on ana-ml2: `vllm-gen` 46 GB, `sec` 51 GB, embed 9.8, coder 8.4, rerank-a3 3.5, reward 2.1, scriberr.
- **⏸ Owed to brokkr-smithy-dev when there is a card again:** nothing blocking. They hold the dose curve and the entanglement finding.
- **⏸ Awaiting pewpew-dev:** which other boxes/CI runners resolve `pypotrace` and need the potrace headers. nh3-dev is done; `playbooks/install-potrace-headers.yaml` makes each additional box one command.
- ⚠ **`/mnt/smithy` will be MISSING after every ana-ml2 reboot** — manual by design. → `persistent-memory.d/2026-08-23-smithy-mount-ana-ml2.md`
## Recent decisions
- `[2026-08-28]` **althing v3 flag day (U9b) executed, then six releases to 3.1.1 in one afternoon — and the post office MOVED to nh3-docker.** Every v2 command deleted; 73 handles seeded and verified by set difference; 5,043 orphaned wake FIFOs deleted (v2 named them per-session+PID, v3 per-handle). Image now registry-pulled, digest-pinned, under the `claude-bot` namespace. → `persistent-memory.d/2026-08-28-althing-v3-cutover.md`
- `[2026-08-28]` **A stale `ALTHING_HANDLE` silently reads another agent's inbox and reports it empty — a SECOND route into the failure v3 exists to prevent.** Outbound mis-signing sometimes gets caught; inbound never does. Shipped as a 3.1.1 warning. ⚠ My `session_handles.json` grounding was wrong (v2 artifact, v3 never opens it) and the same stale source had survived inside my statusline rewrite. → `persistent-memory.d/2026-08-28-handle-resolution-wrong-inbox.md`
- `[2026-08-28]` **nh3-dev's three OOM events attribute to CLAUDE CODE, and the "no kernel evidence" was a permissions artifact.** journald was persistent all along; `journalctl` silently shows only your own messages outside `adm`. Single CC sessions measured 5.4-18.4 GB, so 27 GB is 3-4 long-lived sessions. sysstat + atop now instrument the ramp. → `persistent-memory.d/2026-08-28-nh3-dev-oom-attribution.md`
- `[2026-08-28]` **`sec` moved to ana-ml2 GPU0 and is serving** (operator-directed) — GPU1 had ~28 GB free against the ~51 GB it reserves, so it could not start there. Re-arms the two-GPU load condition on a circuit that tripped 36h earlier; accepted with the constraint stated. → `persistent-memory.d/2026-08-28-sec-seat-gpu0.md`
- `[2026-08-28]` **BELAYED by the operator, both explicitly: (a) a cgroup memory cap on CC sessions, (b) putting ana-gw + ana-wg + one BMC on separate power.** Both were my recommendations; neither is open work. Do not re-raise as new — the atop ramps that would inform (a) are now being collected, so revisit only with a week of data. Tracking surface: this entry.
- `[2026-08-28]` **The deployed CC plugin copies are a release step nobody owns.** `sync_skill.sh` covers the SKILL, not the plugin; both copies must be rsync'd from the repo's `plugin/` on every althing release or they carry the previous release's bugs into the live surface. Raised with forseti for their release notes. Tracking surface: althing thread `01M14QHZNDKDK8KH9DN92VF6VE`.
- `[2026-08-28]` **althing v3.0.0 flag day (U9b) executed — the post office replaced the P2P bus on both boxes, one-way.** 73 handles seeded and verified by set difference; 5,043 orphaned v2 wake FIFOs deleted (v2 named them per-session+PID and never reaped; v3 names them per-handle, so the leak is bounded by construction); v2 db left inert. → `persistent-memory.d/2026-08-28-althing-v3-cutover.md`
- `[2026-08-28]` **nh3-dev's three OOM events attribute to CLAUDE CODE — and the "no kernel evidence" was a permissions artifact.** journald was persistent all along; `journalctl` silently shows only your own messages outside `adm`. Single CC sessions measured at 5.4-18.4 GB, so 27 GB is 3-4 mature sessions, not the ~66 a 408 MB estimate implies. sysstat + atop now instrument the ramp. → `persistent-memory.d/2026-08-28-nh3-dev-oom-attribution.md`
- `[2026-08-27]` **Run 3 gated: the preregistered rule PASSED and a k=25 follow-up found a 44pp self-harm guardrail collapse — DO NOT SERVE.** A pooled preserve-list test structurally cannot see a single-axis collapse. → `persistent-memory.d/2026-08-27-run3-gate-safety-regression.md`
- `[2026-08-27]` **The corpus mix was specified in a unit the optimiser never sees** — 45.8% dialogue by CONTEXT, 24.2% by LOSS. Harness now leads with loss share and calls context a memory budget (`dd5a12e`). → `persistent-memory.d/2026-08-27-mix-specified-in-the-wrong-unit.md`
- `[2026-08-27]` **Dose-response: benefit and damage are ONE direction in weight space** — every axis monotone in scale, no knee. The merge-back cannot separate them; vLLM cannot LoRA-serve this MoE at all. → `persistent-memory.d/2026-08-27-dose-response-entanglement.md`
@@ -239,40 +249,15 @@ _State as left 2026-08-26 23:05 PDT (written 2026-08-27 07:27, corrected 10:5x)
- `[2026-08-15]` **Uncensored gen seat: JonathanColetti/Qwen3.8-27B-Uncensored deployed as `gen-seat`/`vllm-gen` (NVFP4 W4A16 + grafted MTP, 262K); 7 aliases repointed; the definitive `re:^mtp.*`-ignore fix.** 0%-MTP-on-quant (twice) was NOT the abliteration/scheme — the grafted bf16 MTP was missing from `quantization_config.ignore` (vLLM loaded it as quantized → uninitialized). Full arc, the working pipeline, VRAM budget, unsloth speed decomposition, modelopt dead-end. → `persistent-memory.d/2026-08-15-uncensored-gen-seat.md`
- `[2026-08-10→12]` **secrets-broker: per-box Vaultwarden credential store SHIPPED + consumer-confirmed.** `secret` CLI (`put/get/list/rm/backfill`, bw-backed) on `~/.local/bin`; 25 nh3-dev secrets backfilled + round-trip-verified; `rm` + new-namespace warning added post-launch; standing "vault is the credential source of truth" directive now global. → `persistent-memory.d/2026-08-12-secrets-broker.md`
- `[2026-08-09→10]` **dots.tts (rednote-hilab) TTS burn-in on irv-ml1 + canonical voice corpus built (`voices/`).** Operator-directed eval to potentially replace chatterbox-fast. **dots.tts VERIFIED real** (canonical HF ns `dots-studio/`, `rednote-hilab/dots.tts-*` redirects there; Apache-2.0; PyPI `dots.tts` 0.2.1; 2B continuous-AR = semantic enc + Qwen2.5-1.5B LLM + flow-matching acoustic head over 48kHz AudioVAE; zero-shot clone from wav+transcript). **Runs on Ampere 3090** (sm_86, bf16, no fp8 dep); **optimized RTF 0.22** at num_steps=10 (`from_pretrained(..., optimize=True)` CUDA graphs — raw unoptimized was 1.21), **~6GB VRAM**, 48kHz, streams (`generate_stream`). Venv+cache at `irv-ml1:/home/lkraven/dots-tts` (~10GB). **Operator design calls:** SGLang Omni serving (OpenAI `/v1/audio/speech`), transcribe-refs-first, `soar` variant. ⚠ Omni serves soar but its continuous-batching + streaming opts are **mf-only** (soar = single-request) — non-issue for ratatoskr's single-consumer RP surface. **KEY FINDING — dots is highly sensitive to an accurate AND sentence-bounded reference transcript:** mismatched transcript → 0.16s collapse; over-long/messy transcript → reference-audio BLEEDS as an output prefix; mid-clause trim → dangling-word leak (glados "we'll", emmie "And,"). RECIPE (baked into `voices/derive.py`): trim ref to a clean ~6–10s clip ending on a sentence boundary + accurate transcript of exactly that clip. **CANONICAL VOICE CORPUS** stood up in eshpfi `voices/` (operator idea): engine-agnostic `canonical/<v>.wav` + `transcripts/<v>.txt` → per-engine ref sets DERIVED by `derive.py` reading `engines.yaml` profiles (dots/chatterbox/zonos); canonical wavs git-tracked (small/curated), `derived/` gitignored. **4 voices optimized + verified CLEAN for dots: donut, glados, emmie, miranda** (glados canonical is low-SR 16kHz — flagged upgrade candidate). ⚠ GPU GOTCHA: irv-ml1 native CUDA orders **A6000=device0** (ComfyUI-full) — pin the 3090 with `CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0`; and `PYTORCH_CUDA_ALLOC_CONF=expandable_segments` CONFLICTS with `optimize=True` CUDA graphs (curr_block error). Booths: `dots-vs-chatterbox`, `dots-voices-optimized`. **SHIPPED 2026-08-10:** operator A/B verdict "dots is very good" → containerized as a **thin FastAPI wrapper over DotsTtsRuntime** (chosen over SGLang Omni — Omni's batching is mf-only, unneeded for ratatoskr's single consumer; wrapper is SERIALIZED one-gen-at-a-time via a threading.Lock, Omni+mf = parked API-compatible escalation if multi-consumer ever lands). **LIVE on irv-ml1:8198** (`local/dots-tts:v1`, OpenAI `/v1/audio/speech` + `/health` + `/v1/voices`, container healthy, both stream + non-stream verified CLEAN, 4 voices donut/glados/emmie/miranda) alongside chatterbox :8197 (nothing repointed). Stack = `stacks/dots-tts/` (Dockerfile/app.py/compose/.env.example/README). ⚠ CONTAINER GOTCHA: `optimize=True` (torch.compile/inductor/triton) needs a **C compiler at RUNTIME** — slim image must `apt install build-essential` or model-load dies "Failed to find C compiler" (host venv had gcc ambient, masking it); persist `TORCHINDUCTOR_CACHE_DIR` to a mounted dir or every restart re-JITs ~5min. Corpus home = eshpfi `voices/` (operator ruled keep-here). **REMAINING: ratatoskr client cutover** to :8198 `/v1/audio/speech` (Phase-2 tail, peer-coupled — draft the ask). [[reference_chatterbox_fast_repo]] [[reference_zonos_tts_stack]] [[reference_verify_hf_repo_ids_before_pull]]
- `[2026-08-05]` **Fleet CI resilience flip (`DEFAULT_ACTIONS_URL=self`) — attempted end-to-end, PARKED on a runner action-fetch auth blocker; infra-ops to research it (operator-directed, deferred, NOT now).** 7 gitea action mirrors staged public+populated (orgs `actions`+`astral-sh`); the flip resolves `uses:` correctly but act_runner v0.6.0 can't authenticate its fetch to gitea 1.26 ("Invalid username or token. Password authentication is not supported"). Reverted (CI back on github default); `REQUIRE_SIGNIN_VIEW=false` KEPT as a standing change (operator, internal WG net). Full endeavor, the reliable nh3-dev-egress + git-SSH mirror method, exact config state, smoke method, and next step → `persistent-memory.d/2026-08-05-ci-flip-parked.md`
- `[2026-08-05]` **worldtree herald re-nudge bug root-caused → forseti shipped althing-core v2.1.2 (`d5d33df`, deployed on nh3-dev).** `herald.py:363` rendered the wake command from the empty *fresh* mail set on the re-nudge path (should be `deliver_msgs`) → `messages[0]` IndexError → un-suppressed outer catch-all → 7s crash-loop for 9 days on worldtree-codex's pane route (mimir-dev surfaced it; I traced it from the editable source). Fix + `render_command` empty-guard + outer log-suppress + 3 tests + contract amendment, all forseti's. **nh3-extdev herald 2.1.2 upgrade DEFERRED** (operator, not-now): extdev is a WHEEL install (not editable), unexposed (no pane routes); the verified 2.1.2 wheel is staged on nh3-dev `/tmp` (sha256 `003508…cef27`) — `uv tool install --force` + restart both heralds when un-parked. extdev herald-unit provenance resolved (operator-authorized 2026-07-25 via forseti relay; recorded in this file's 07-25 herald-install entry). auto-memory `reference_nh3_dev_althing_herald`.
- `[2026-07-31]` **muninn-gate (#377 ingestion front door) BUILT + DEPLOYED + healthy on corviduo-dev:8090.** First-boot acceptance passed (watcher:running:true proves ingestion_root byte-identity); submit path deferred to the mimir-inbox era. Full wiring (uid-1000, state-volume mount, staging path-agreement, BuildKit-secret build, deferred repoint + operational guards) → `persistent-memory.d/2026-07-31-muninn-gate-deploy.md`
_222 older entries archived to archival-memory.md._
_223 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-08-25]` **Four throughput levers measured and killed — do not re-chase.** (1) **Fused MoE / `grouped_mm`** — 0.9% *slower* than the Python loop and dense GEMM is only 7.9% of the step, capping the whole category near 10%. (2) **CUDA graphs / `torch.compile` over the expert loop** — the two-term scaling fit closed with residuals under 3ms and needed NO constant term, so there is no fixed per-batch cost to amortise; 3,840 expert-GEMM launches per forward are not what we pay for. (3) **`liger` fused linear CE** — the chunked CE measured **1.1% of the step** forward, ~3% with recompute. A tidy-up, not a lever. (4) **Selective gradient checkpointing** — ~2% of a post-fix step, real bug surface. Also: **token-budget batching is dead by the same fit** — with no constant term, total time over a fixed set of widths is invariant to how you group them; only the widths matter, which is exactly why bucketing works and repacking does not.
@@ -288,12 +273,4 @@ _222 older entries archived to archival-memory.md._
- `[2026-08-03]` **ComfyUI `--enable-triton-backend` on the irv-ml1 A6000 crashes EVERY render — Ampere has no hardware e4m3.** adhoc-agent's operator-approved probe: comfy_kitchen's triton backend has a FUSED int8 matmul that would beat the eager backend's ~1.9x-slower unfused int8 path (21.3s vs 11.2s fp8 on the Moody Krea2 int8 checkpoints). Flipped it (added to `COMFY_CMDLINE_EXTRA`, recreated) → `triton.compiler.errors.CompilationError: ValueError("type fp8e4nv not supported in this architecture. supported: fp8e4b15, fp8e5")` in `comfy_kitchen/backends/triton/quantization.py:145 dequantize_per_tensor_fp8`, failing at **node 5 CLIPTextEncode**. Triton's fp8 dequant kernel targets `fp8e4nv` (Hopper/Ada e4m3); **sm_86 Ampere (A6000) lacks hardware e4m3** → the JIT compile dies. With triton on it grabs the **global** `--fp8_e4m3fn-text-enc` dequant, so every render (fp8 AND int8) dies upstream at the text-encode step — the int8 UNet path never ran, so the convrot-coverage caveat wasn't even the limiter. Reverted cleanly (~15s to healthy, image unchanged `sha256:94afb8ca`, sage intact, prod restored). **The parked cu130 rebuild won't fix it** (e4m3 = hardware format, not CUDA version). **DEFERRED to the Ada refresh** (operator: "ada is coming, we'll optimize then" — Ada sm_89 has native e4m3, so triton's fp8 path should compile there). **Mechanics:** `--enable-triton-backend` is a compose `environment:` var, so toggling it needs `docker compose up -d` (**recreate**), NOT `docker restart` (reuses the baked env, no-ops silently). Full: auto-memory `parked_triton_backend_ampere_fp8`.
_143 older entries archived to archival-memory.md._
+100
View File
@@ -0,0 +1,100 @@
# Install the potrace + agg development headers so `pypotrace` can build from source.
#
# Why a playbook and not a one-liner: `pypotrace` is an **sdist that compiles at install
# time**, so every machine and every CI runner that resolves it needs these headers present
# FIRST. That makes this a recurring per-box action, not a one-off. Requested by forseti
# (pewpewstudio) for `core.image_pipeline`, which vectorises raster art into laser-ready
# contours; the operator chose pypotrace over shelling out to the potrace binary (2026-08-28).
#
# Run: scripts/elway infra-ops@<host> --playbook playbooks/install-potrace-headers.yaml
# Rerunnable: a second run shows the install `skipped`.
#
# ⚠ TWO THINGS THAT WILL SEND YOU DOWN THE WRONG PATH ON A BOX WHERE THIS FAILS
#
# 1. **Only libagg is a pkg-config consumer. potrace is not.**
#
# libagg /usr/lib/x86_64-linux-gnu/pkgconfig/libagg.pc present
# potrace NO .pc file — found via /usr/include/potracelib.h and the library
#
# So `pkg-config --exists potrace` returns FALSE on a correctly configured box. It looks
# exactly like the cause and never is. The real build error names libagg and only libagg:
#
# Package libagg was not found in the pkg-config search path.
# Package 'libagg', required by 'virtual:world', not found
#
# A wrong model that produces a plausible-looking diagnostic costs more than no model.
#
# 2. **libagg's pkg-config modversion disagrees with its Debian package version.**
#
# pkg-config --modversion libagg -> 2.7.0
# dpkg version -> 1:2.6.1-r134+dfsg1-2+b1
#
# Comparing those two numbers convinces you the wrong package is installed. It is not a
# problem; it is upstream's version vs Debian's packaging of it.
#
# Both of these were learned on the nh3-dev install and are recorded here rather than in an
# althing thread, at forseti's suggestion, because a thread is not where the next person looks.
vars:
probe_venv: /tmp/pypotrace-probe
# The user whose toolchain builds the probe. The headers are installed system-wide
# as root; the build check is a developer action and runs as this user.
dev_user: lkraven
steps:
- name: Install the potrace and agg development headers
shell: sudo DEBIAN_FRONTEND=noninteractive apt-get install -y libpotrace-dev libagg-dev
when: "! dpkg -s libpotrace-dev >/dev/null 2>&1 || ! dpkg -s libagg-dev >/dev/null 2>&1"
verify:
- name: libagg's pkg-config file is discoverable (this is the one that actually gates the build)
shell: pkg-config --exists libagg
changed_when: "false"
- name: potrace's header is present (NOT via pkg-config — it ships no .pc)
shell: test -f /usr/include/potracelib.h
changed_when: "false"
- name: pypotrace COMPILES against them
# ⚠ `uv` is NOT on a non-interactive ssh PATH — infra-ops gets
# /usr/local/bin:/usr/bin:/bin:/usr/games and nothing else. It also lives in a
# different place on every box: /home/lkraven/bin/uv on nh3-dev,
# /home/infra-ops/.local/bin/uv on nh3-extdev. Search rather than assume, and say
# so loudly if it is genuinely absent — a build probe that silently does not run
# is the failure this whole playbook exists to prevent.
shell: |
# ...and on nh3-dev it is inside a 0700 home, so `test -x` from infra-ops fails
# even with the right absolute path — the directory cannot be traversed. The
# headers are system-wide (root's business); the build probe is a DEVELOPER
# action and has to run as the user who owns the toolchain.
RUNAS={{ dev_user }}
UV=""
for c in /home/{{ dev_user }}/bin/uv /home/{{ dev_user }}/.local/bin/uv /usr/local/bin/uv; do
sudo -u "$RUNAS" test -x "$c" && { UV="$c"; break; }
done
[ -n "$UV" ] || { echo "no uv reachable as $RUNAS; cannot run the build probe"; exit 1; }
echo " using uv at $UV (as $RUNAS)"
sudo -u "$RUNAS" rm -rf {{ probe_venv }}
sudo -u "$RUNAS" "$UV" venv {{ probe_venv }} >/dev/null 2>&1
sudo -u "$RUNAS" "$UV" pip install --python {{ probe_venv }}/bin/python pypotrace 2>&1 | tail -2
changed_when: "false"
- name: and the built extension actually TRACES, which a successful build does not prove
# A square must come back as one curve of four CornerSegments. If the extension linked
# against something wrong it can still import and return nonsense; the geometry is the
# assertion, not the import.
shell: |
sudo -u {{ dev_user }} {{ probe_venv }}/bin/python -c "
import numpy as np, potrace
a = np.zeros((40,40), np.uint32); a[10:30,10:30] = 1
curves = list(potrace.Bitmap(a).trace())
segs = [s for c in curves for s in c]
assert len(curves) == 1, curves
assert len(segs) == 4, segs
assert {type(s).__name__ for s in segs} == {'CornerSegment'}, segs
"
changed_when: "false"
- name: Remove the probe venv
shell: sudo -u {{ dev_user }} rm -rf {{ probe_venv }}
changed_when: "false"
+80
View File
@@ -0,0 +1,80 @@
# Move the `sec` pen-test seat (M.O.G.-SEC-27B) from ana-ml2 GPU 1 to GPU 0 and bring it up.
#
# Why: GPU 1 carries the five resident fleet seats (gen 46 GB + embed 9.8 + coder 8.4 +
# rerank 3.5 + reward 2.1 = ~69.9 GB of 97.9), leaving ~28 GB. This seat reserves
# MOG_GPU_MEM_UTIL=0.52 -> ~51 GB, so it could not start on GPU 1 at all. GPU 0 has been
# idle since run 3c was stopped on 2026-08-26. Operator-directed 2026-08-28.
#
# ⚠ POWER. This re-arms the two-GPU load condition that tripped the Anaheim rack breaker
# on 2026-08-26. One circuit feeds the whole rack including ana-gw and ana-wg, so a trip
# costs the site AND the way back in. Idle draw is negligible; the risk materialises when
# sec and gen are under concurrent load. Operator accepted this with the constraint stated.
#
# Labels only apply at container CREATION, so this uses `up -d`, never `restart` --
# the homepage description carries the GPU number and would otherwise stay stale.
#
# Run: scripts/elway infra-ops@ana-ml2 --playbook playbooks/mog-sec-move-to-gpu0.yaml
# Model load is slow (22 GB + 262K ctx + MTP graft); the verify phase polls rather than
# assuming readiness, and the compose healthcheck allows a 900s start_period.
vars:
stack_dir: /opt/docker/compose/mog-sec
container: vllm-mog-sec
service: vllm-mog-sec
gpu_id: "0"
port: "8019"
staging: /tmp/mog-sec-compose.yaml
steps:
- name: Stage the updated compose (GPU pin default + label now say GPU 0)
upload:
src: stacks/mog-sec/compose.yaml
dest: "{{ staging }}"
mode: "0644"
- name: Install it over the deployed copy
# /opt/docker/compose is root-owned, so the scp above lands in /tmp and this
# promotes it. Verified byte-identical against the deployed file beforehand:
# the only diff was these edits, so nothing on the host is being clobbered.
shell: sudo install -o root -g root -m 0644 {{ staging }} {{ stack_dir }}/compose.yaml
changed_when: "! sudo cmp -s {{ staging }} {{ stack_dir }}/compose.yaml"
- name: Pin the seat to GPU {{ gpu_id }} in the host .env
# The .env is the tunable surface and is NOT in git (secrets//tunables are
# excluded both directions). The compose default now matches, but the .env
# is what actually decides, so set it explicitly rather than relying on the
# default resolving.
shell: sudo sed -i 's/^MOG_GPU_ID=.*/MOG_GPU_ID={{ gpu_id }}/' {{ stack_dir }}/.env
when: "! sudo grep -qxF 'MOG_GPU_ID={{ gpu_id }}' {{ stack_dir }}/.env"
- name: Bring the seat up (up -d, not restart — labels apply at creation)
shell: cd {{ stack_dir }} && sudo docker compose up -d {{ service }}
verify:
- name: Container exists and is running
# ⚠ No `docker inspect -f` here. Go templates use {{ }} and so does elway's own
# variable substitution, so an inspect format string gets eaten before it reaches
# the host -- these two checks reported FAILED on a deploy that had in fact
# succeeded. Filter-and-grep has no such collision.
shell: sudo docker ps --filter name={{ container }} --filter status=running --quiet | grep -q .
changed_when: "false"
- name: The container is actually pinned to GPU {{ gpu_id }}
# Assert the EFFECTIVE device reservation on the running container, not the
# .env string we wrote -- the .env is an input, this is the outcome.
shell: sudo docker inspect {{ container }} | tr -d ' \n' | grep -q '"DeviceIDs":\["{{ gpu_id }}"\]'
changed_when: "false"
- name: GPU 0 now holds a vLLM process (the seat really loaded onto that card)
shell: nvidia-smi --id={{ gpu_id }} --query-compute-apps=pid,used_memory --format=csv,noheader | grep -qE '[0-9]'
changed_when: "false"
- name: Health endpoint answers
shell: curl -fsS --max-time 10 http://127.0.0.1:{{ port }}/health >/dev/null
changed_when: "false"
- name: Both served names are advertised (base + thinking)
shell: |
MODELS=$(curl -fsS --max-time 10 http://127.0.0.1:{{ port }}/v1/models)
echo "$MODELS" | grep -q 'mog-sec-27b' && echo "$MODELS" | grep -q 'mog-sec-27b-thinking'
changed_when: "false"
+116
View File
@@ -0,0 +1,116 @@
# nh3-dev: make memory exhaustion diagnosable after the fact.
#
# Context: three memory-exhaustion events in 14 days (2026-08-14, 2026-08-26,
# 2026-08-28) with the interval halving. forseti reported that none could be
# attributed because "kernel messages are not being persisted to journald".
#
# ⚠ THAT PREMISE WAS WRONG, and the way it was wrong matters more than the fix.
# journald on this box IS persistent and HAS every OOM report: 15,068 kernel
# entries in the 82-day previous boot, 351 OOM records overall, full task tables
# with per-process RSS. `journalctl -b -1 -k` returned "one entry" because it was
# run by a user in neither `adm` nor `systemd-journal` — journalctl silently shows
# you only your OWN messages and prints the reason as a hint. The same artifact
# also produced the "journal stops mid-line at 05:36:08 with no shutdown
# sequence" claim: the true boot -1 boundary is 05:47:04, and OOM kills are
# recorded at 05:38, 05:40 and 05:42.
#
# So step 1 is a GROUP MEMBERSHIP fix, not a logging fix. The evidence was
# always there and unreadable.
#
# What the evidence says, now that it can be read: the hog is Claude Code.
# `/home/lkraven/.local/share/claude/versions/2.1.220` is the versioned CC binary,
# so OOM victims named `2.1.220` / `2.1.177` / `2.1.168` are CC sessions, as are
# the ones named `claude`. Largest single anon-rss recorded: 18.4 GB (2.1.177,
# Aug 14), with 15.8 GB seen twice. Everything else killed — althing-forseti,
# caddy, ttyd, zellij, the althing daemons — is 30-55 MB collateral.
#
# steps 2-4 add the timeseries that the journal cannot give: the journal records
# the moment of the kill, not the ramp toward it, and the Aug 28 event was a hard
# lockup where the box never got far enough to log a coherent sweep.
# - sysstat -> system-wide memory/CPU timeseries (the ramp)
# - atop -> PER-PROCESS timeseries (which session, and how fast)
# atop is the one that answers "which of the dozen sessions", which sar cannot.
#
# Run: scripts/elway infra-ops@nh3-dev --playbook playbooks/nh3-dev-memory-forensics.yaml
# Rerunnable: a second run shows every step `skipped` or `ok`.
vars:
# The interactive/agent user whose sessions read the journal.
journal_user: lkraven
# Debian's journald ACL grants read to `adm` explicitly (getfacl shows
# group:adm:r-x); `systemd-journal` owns the files. `adm` is the documented
# Debian path and the one the ACL names, so use it.
journal_group: adm
# 5 min, not Debian's default 10 — a CC session can add several GB inside one
# 10-minute bucket, which is exactly the resolution the ramp needs.
sar_interval: "*:00/05"
# 60s per-process sample. ~7 generations keeps this under ~1 GB against 80 GB free.
atop_interval: "60"
atop_generations: "7"
steps:
- name: Grant the agent user journal read access (THE actual fix for "no evidence")
shell: sudo usermod -aG {{ journal_group }} {{ journal_user }}
when: "! id -nG {{ journal_user }} | grep -qw {{ journal_group }}"
- name: Install sysstat and atop
shell: sudo DEBIAN_FRONTEND=noninteractive apt-get install -y sysstat atop
when: "! dpkg -s sysstat >/dev/null 2>&1 || ! dpkg -s atop >/dev/null 2>&1"
- name: Enable sysstat collection in /etc/default/sysstat
# The package ships ENABLED="false" and the timer is a no-op until this flips.
shell: sudo sed -i 's/^ENABLED=.*/ENABLED="true"/' /etc/default/sysstat
when: "! grep -qxF 'ENABLED=\"true\"' /etc/default/sysstat 2>/dev/null"
- name: Tighten the sysstat collection interval to 5 minutes
shell: |
sudo mkdir -p /etc/systemd/system/sysstat-collect.timer.d
printf '[Timer]\n# Default is */10. A CC session can add several GB inside one 10-minute\n# bucket; 5 min is the resolution the memory ramp actually needs.\nOnCalendar=\nOnCalendar=%s\n' '{{ sar_interval }}' | sudo tee /etc/systemd/system/sysstat-collect.timer.d/override.conf >/dev/null
when: "! grep -qxF 'OnCalendar={{ sar_interval }}' /etc/systemd/system/sysstat-collect.timer.d/override.conf 2>/dev/null"
- name: Configure atop for 60s per-process sampling with 7-day retention
shell: |
sudo sed -i 's/^LOGINTERVAL=.*/LOGINTERVAL={{ atop_interval }}/' /etc/default/atop
sudo sed -i 's/^LOGGENERATIONS=.*/LOGGENERATIONS={{ atop_generations }}/' /etc/default/atop
when: "! grep -qxF 'LOGINTERVAL={{ atop_interval }}' /etc/default/atop 2>/dev/null || ! grep -qxF 'LOGGENERATIONS={{ atop_generations }}' /etc/default/atop 2>/dev/null"
- name: Reload systemd and enable the collectors
shell: |
sudo systemctl daemon-reload
sudo systemctl enable --now sysstat.service sysstat-collect.timer sysstat-summary.timer
sudo systemctl enable --now atopacct.service atop.service atop-rotate.timer
sudo systemctl restart atop.service
- name: Seed one sysstat sample so sar has data immediately
shell: sudo /usr/lib/sysstat/sa1 1 1
verify:
- name: Agent user is now in the journal-reading group
# `sg` evaluates the membership WITHOUT waiting for a re-login, so this
# asserts the effective grant rather than the /etc/group substring.
shell: sudo -u {{ journal_user }} sg {{ journal_group }} -c 'journalctl -b -1 -k --no-pager 2>/dev/null | wc -l' | awk '{ if ($1 > 100) exit 0; else exit 1 }'
changed_when: "false"
- name: sysstat collection timer is active
shell: systemctl is-active --quiet sysstat-collect.timer
changed_when: "false"
- name: sysstat is collecting at the 5-minute cadence
# ⚠ Assert the EFFECTIVE value, not the string we wrote. systemd normalises
# `*:00/05` to `*-*-* *:00/5:00`, so grepping for our own input fails while
# the setting is live — which is exactly how this verify failed on the first
# run and briefly looked like the override had not applied.
shell: systemctl show sysstat-collect.timer -p TimersCalendar | grep -qF '*:00/5:00'
changed_when: "false"
- name: sar can actually read a memory timeseries (not just that the timer exists)
shell: sar -r 2>/dev/null | tail -2 | grep -qE '[0-9]'
changed_when: "false"
- name: atop daemon is running
shell: systemctl is-active --quiet atop.service
changed_when: "false"
- name: atop is writing a readable per-process log
shell: sudo test -s /var/log/atop/atop_$(date +%Y%m%d)
changed_when: "false"
+138
View File
@@ -0,0 +1,138 @@
# nh3-extdev: cut the system-wide althing install over from v2.1.0 to v3.0.x (U9b flag day; re-run for each release).
#
# nh3-extdev is the one box a `git pull` cannot move: althing lives there as a system WHEEL
# under /opt/uv-tools/althing-core with entry points in /usr/local/bin, installed from a wheel
# that was copied to /tmp -- not from a checkout. So it needs its own install or it goes dark
# at the cutover.
#
# ⚠ Two things about this box that differ from nh3-dev:
# - the v2 daemons are SYSTEM units here (althing-herald, althing-receiver), not user units.
# - `uv` is not on lkraven's PATH; it lives at /home/infra-ops/.local/bin/uv. The original
# install used it under sudo with UV_TOOL_DIR=/opt/uv-tools, per the uv-receipt.toml.
#
# ⚠ There is a live agent session here (ldp-dev) holding a v2 light-monitor. Retiring the v2
# herald does not kill it, but it will never fire again -- that session has to re-arm on
# althing-listen after this. Its handle survives: bare `ldp-dev` is in the authoritative 73;
# only the machine-qualified `ldp-dev@nh3-extdev` was on the legacy exclusion list.
#
# Run: scripts/elway lkraven@10.100.50.42 --playbook playbooks/nh3-extdev-althing-v3.yaml
# Rerunnable: a second run shows the install and unit steps skipped.
vars:
wheel_src: /home/lkraven/development/althing/dist/althing_core-3.1.1-py3-none-any.whl
wheel_dest: /tmp/althing_core-3.1.1-py3-none-any.whl
uv: /home/infra-ops/.local/bin/uv
tool_dir: /opt/uv-tools
bin_dir: /usr/local/bin
# ⚠ Moved off nh3-dev 2026-08-28. A stale value here does not fail loudly at
# install time — it fails in the VERIFY, which then reads as a broken deploy.
post_office: http://10.100.50.40:8390
# ⚠ A release can change more than one file — 3.1.0 changed two. One pair is not
# enough, and a check that asserts only half a release is a check that half-passes
# silently. Space-separated `file:marker` pairs; bump BOTH per release.
# 3.0.1 zellij.py:_PANE_ID session_source.py:_live_pid (pane routes)
# 3.0.3 post_office_herald.py:PANE_SETTLE_S (write/submit race)
# 3.1.0 post_office_herald.py:POST_OFFICE_HINT dev_launch.py:resolve_post_office
# 3.1.1 postbox.py:warn_if_handle_looks_wrong dev_launch.py:handles_launched_at
#
# ⚠ Match on PRESENCE (grep -q), never on a count. 3.1.1 shipped with a stated
# expectation of "grep -c handles_launched_at dev_launch.py # 2+"; the real count
# there is 1 (the definition) with the other two occurrences in postbox.py. A count
# assertion would have reported FAILED on a byte-perfect install.
markers: "postbox.py:warn_if_handle_looks_wrong dev_launch.py:handles_launched_at"
steps:
- name: Stage the v3.0.0 wheel
upload:
src: /home/lkraven/development/althing/dist/althing_core-3.1.1-py3-none-any.whl
dest: "{{ wheel_dest }}"
mode: "0644"
- name: Retire the v2 system daemons BEFORE swapping the package
# Order matters: these run out of /opt/uv-tools/althing-core/bin/python, which the
# install is about to replace. Stopping first means they never see a half-swapped tree.
# v3 has no counterpart to either -- the post office replaced the herald and deleted the
# reason for the receiver, since there is no longer a mailbox per machine to deliver between.
shell: sudo systemctl disable --now althing-herald.service althing-receiver.service
when: "systemctl is-active --quiet althing-herald.service || systemctl is-active --quiet althing-receiver.service"
- name: Install the staged althing-core wheel over the system wheel install
# NOT gated on `postbox` existing — that guard was right for the cutover and
# wrong for every release after it: postbox exists now, so a version bump would
# silently skip. `--force` makes the reinstall idempotent on its own.
shell: sudo env UV_TOOL_DIR={{ tool_dir }} UV_TOOL_BIN_DIR={{ bin_dir }} {{ uv }} tool install --force {{ wheel_dest }}
- name: Install the post-office herald as a system unit
# A system unit rather than a user unit because that is how v2 was supervised here and
# because this box has no lingering user session to hang a --user unit from.
shell: |
printf '%s\n' \
'[Unit]' \
'Description=Althing post-office herald — per-machine relay (v3)' \
'Documentation=https://gitea.phasefinal.com/vh/althing' \
'After=network-online.target' \
'Wants=network-online.target' \
'' \
'[Service]' \
'Type=simple' \
'User=lkraven' \
'Environment=ALTHING_POST_OFFICE={{ post_office }}' \
'ExecStart={{ bin_dir }}/althing-po-herald' \
'Restart=always' \
'RestartSec=5' \
'' \
'# Dials out, opens no port, holds no state. Refuses to start if another herald' \
'# already holds this node — two would double every poke and both write liveness.' \
'# Replaces althing-herald.service + althing-receiver.service, retired 2026-08-28.' \
'' \
'[Install]' \
'WantedBy=multi-user.target' \
| sudo tee /etc/systemd/system/althing-po-herald.service >/dev/null
sudo systemctl daemon-reload
when: "! test -f /etc/systemd/system/althing-po-herald.service"
- name: Enable and (re)start the herald so it picks up the new code
shell: sudo systemctl enable --now althing-po-herald.service && sudo systemctl restart althing-po-herald.service
verify:
- name: postbox is installed and is v3
shell: "{{ bin_dir }}/postbox --help | grep -q 'send,reply,read,peek,thread,search,status,handles,register,sign-off'"
changed_when: "false"
- name: the v2 entry points are GONE, not merely shadowed
# Assert absence of the binaries themselves. A `which` that still resolves would mean the
# old wheel's entry points survived the --force and agents could keep calling a dead CLI.
shell: "! test -e {{ bin_dir }}/althing-cli && ! test -e {{ bin_dir }}/althing-receiver && ! test -e {{ bin_dir }}/althing-herald"
changed_when: "false"
- name: v2 daemons are stopped and disabled
shell: "! systemctl is-active --quiet althing-herald.service && ! systemctl is-active --quiet althing-receiver.service"
changed_when: "false"
- name: the po-herald is running
shell: systemctl is-active --quiet althing-po-herald.service
changed_when: "false"
- name: this box can reach the post office and the roster is populated
# ⚠ --handle is required. postbox resolves its identity from ALTHING_HANDLE, which
# dev-launch sets per pane and which a playbook shell does not have -- without it this
# check fails on identity, not on reachability, and reads as a deployment fault.
shell: ALTHING_POST_OFFICE={{ post_office }} {{ bin_dir }}/postbox --handle operator handles | wc -l | awk '{ if ($1 >= 70) exit 0; else exit 1 }'
changed_when: "false"
- name: This release's markers are ALL present BY CONTENT, not by version string
# forseti's own checks. A dist-info directory records what was INSTALLED, not
# what the files CONTAIN — verify the code, not the label. Bump `marker` and
# `marker_file` with each release rather than trusting the version bumped.
# 3.0.1 _PANE_ID in zellij.py (pane routes)
# 3.0.3 PANE_SETTLE_S in post_office_herald.py (the write/submit race)
shell: |
SP={{ tool_dir }}/althing-core/lib/python3.13/site-packages/althing
rc=0
for pair in {{ markers }}; do
f="${pair%%:*}"; m="${pair##*:}"
if grep -q "$m" "$SP/$f"; then echo " ok $f : $m"
else echo " MISS $f : $m"; rc=1; fi
done
exit $rc
changed_when: "false"
+192
View File
@@ -0,0 +1,192 @@
#!/usr/bin/env bash
# CANONICAL COPY of the Claude Code statusline. Deployed to (and read from):
#
# ~/.claude/statusline-command.sh <- the LIVE path CC actually runs
#
# Install / update after editing here:
# cp scripts/claude-statusline-command.sh ~/.claude/statusline-command.sh
#
# Copies, not symlinks — same rule as stacks/: this tree is intent, the live
# path is reality, and they diverge until someone deploys. Diff them with
# diff -u scripts/claude-statusline-command.sh ~/.claude/statusline-command.sh
#
# It is version-controlled here because the althing v3 cutover broke it in a way
# that was invisible: the segment gated on `command -v althing-cli`, a binary the
# cutover deleted, so the 📬 badge and 🔔 bell silently vanished for every session
# on the box. With 71 of 73 handles pull-only, that badge is the ONLY out-of-band
# signal telling a session with no armed waiter that it has mail — a dead
# statusline made a working bus look like an empty one.
#
# Smoke test (CC pipes session JSON on stdin):
# echo '{"model":{"display_name":"opus"},"workspace":{"current_dir":"/tmp"},"cwd":"/tmp"}' \
# | ALTHING_HANDLE=<a-handle-with-unread> bash scripts/claude-statusline-command.sh
# want: a leading `📬 N`; and with an unreachable post office, degradation in
# ~2s rather than a hang.
# Claude Code statusline. Layout:
# [📬N] [🔔/🔕] | <proj> ⎇<branch> *<dirty> ↑<unpushed> | <model> | ctx:<pct> <toks> | $<session-cost> | 5h:% 7d:%
# ctx% and rate-limit %s are threshold-colored: green <60, yellow 60-90, red >90.
# All segments degrade gracefully (missing tool / non-git dir / no handle => segment omitted).
input=$(cat)
# --- threshold color: $1=numeric pct, $2=display text -> colored text ---
color_pct() {
local p="$1" txt="$2" c
if awk "BEGIN{exit !($p < 60)}"; then c=$'\033[32m' # green <60
elif awk "BEGIN{exit !($p > 90)}"; then c=$'\033[31m' # red >90
else c=$'\033[33m' # yellow 60-90
fi
printf '%s%s\033[0m' "$c" "$txt"
}
# --- reset countdown: $1=unix ts -> "1d3h"/"3h20m"/"45m" (2-unit; "now"/empty edge) ---
reset_in() {
local ts="$1" now delta d h m
[ -z "$ts" ] && return
now=$(date +%s)
delta=$(( ts - now ))
[ "$delta" -le 0 ] && { printf 'now'; return; }
if [ "$delta" -ge 86400 ]; then
d=$(( delta / 86400 )); h=$(( (delta % 86400) / 3600 ))
if [ "$h" -gt 0 ]; then printf '%dd%dh' "$d" "$h"; else printf '%dd' "$d"; fi
elif [ "$delta" -ge 3600 ]; then
h=$(( delta / 3600 )); m=$(( (delta % 3600) / 60 ))
if [ "$m" -gt 0 ]; then printf '%dh%dm' "$h" "$m"; else printf '%dh' "$h"; fi
else
printf '%dm' $(( delta / 60 ))
fi
}
# --- one jq pass for every payload field ---
# \x1f (unit separator) delimiter, NOT tab: tab is IFS-whitespace so `read` would
# collapse consecutive tabs and shift every field after an empty one (e.g. a
# session with no rate_limits). \x1f is non-whitespace -> empty fields preserved.
IFS=$'\x1f' read -r model used_pct input_tok five_pct week_pct cwd fast cost_usd model_id five_reset week_reset < <(
printf '%s' "$input" | jq -r '[
(.model.display_name // "unknown"),
(.context_window.used_percentage // ""),
(.context_window.total_input_tokens // 0),
(.rate_limits.five_hour.used_percentage // ""),
(.rate_limits.seven_day.used_percentage // ""),
(.cwd // .workspace.current_dir // ""),
(.fast_mode // false),
(.cost.total_cost_usd // 0),
(.model.id // ""),
(.rate_limits.five_hour.resets_at // ""),
(.rate_limits.seven_day.resets_at // "")
] | map(tostring) | join("")'
)
[ -z "$model" ] && model="unknown"
# --- model (compact) + fast-mode flag ---
model="${model%% (*}" # "Opus 4.8 (1M context)" -> "Opus 4.8"
[ "$fast" = "true" ] && model="⚡$model"
# --- context % (colored) + absolute input tokens ---
if [ -n "$used_pct" ]; then
ctx_seg=$(color_pct "$used_pct" "ctx:$(printf '%.0f%%' "$used_pct")")
else
ctx_seg="ctx:--"
fi
if [ "${input_tok:-0}" -ge 1000 ] 2>/dev/null; then
toks=$(awk "BEGIN{printf \"%.0fk\", ${input_tok}/1000}")
else
toks="${input_tok:-0}"
fi
# --- per-session cost (Claude Code's own cache/model-aware accounting) ---
# adaptive precision: whole dollars once it's real money, cents when small.
cost=$(awk "BEGIN{c=${cost_usd:-0}; if(c>=100) printf \"%.0f\",c; else if(c>=10) printf \"%.1f\",c; else printf \"%.2f\",c}")
# --- rate limits (each % colored on the same thresholds, + reset countdown) ---
_rl() { # $1=pct $2=label $3=reset_ts -> "<label>:NN%·<reset>"
local seg r; seg=$(color_pct "$1" "$2:$(printf '%.0f' "$1")%")
r=$(reset_in "$3"); [ -n "$r" ] && seg="$seg·$r"
printf '%s' "$seg"
}
rate=""
[ -n "$five_pct" ] && rate=$(_rl "$five_pct" "5h" "$five_reset")
if [ -n "$week_pct" ]; then
wk=$(_rl "$week_pct" "7d" "$week_reset")
[ -n "$rate" ] && rate="$rate "
rate="${rate}${wk}"
fi
# --- project tag + git state (branch, dirty, unpushed) ---
proj=""; gitseg=""
if [ -n "$cwd" ]; then
proj=$(basename "$cwd")
if git -C "$cwd" rev-parse --git-dir >/dev/null 2>&1; then
br=$(git -C "$cwd" branch --show-current 2>/dev/null)
[ -z "$br" ] && br=$(git -C "$cwd" rev-parse --short HEAD 2>/dev/null)
dirty=$(git -C "$cwd" status --porcelain 2>/dev/null | grep -c .)
ahead=$(git -C "$cwd" rev-list --count '@{upstream}..HEAD' 2>/dev/null)
gitseg="⎇${br:-?}"
[ "${dirty:-0}" -gt 0 ] 2>/dev/null && gitseg="$gitseg *$dirty"
[ -n "$ahead" ] && [ "$ahead" -gt 0 ] 2>/dev/null && gitseg="$gitseg ↑$ahead"
fi
fi
# --- althing: unread count (📬 N) + waiter-armed (🔔 armed / 🔕 not) ---
# v3 (the post office, 2026-08-28). ⚠ This block used to gate on
# `command -v althing-cli`, which the v3 cutover DELETED -- so the whole segment,
# badge and bell both, silently disappeared for every session on this box. That is
# worse than a cosmetic loss: 71 of 73 handles are pull-only (no waiter armed, never
# poked), and this badge is the ONLY out-of-band signal telling such a session it has
# mail waiting. A dead statusline made the new bus look like an empty one.
althing=""; mon=""
if command -v postbox >/dev/null 2>&1; then
# postbox has NO default post-office address and the statusline runs in a bare
# shell with neither var set. Hardcoded here deliberately: an unset address makes
# postbox error, which in a must-never-crash segment is indistinguishable from
# "no mail" -- the exact conflation v3 exists to prevent.
export ALTHING_POST_OFFICE="${ALTHING_POST_OFFICE:-http://10.100.50.40:8390}"
h="${ALTHING_HANDLE:-}"
# ALTHING_HANDLE isn't set in the statusline env, so resolve the handle from the
# cwd Claude Code passes on stdin.
#
# ⚠ PREFER launch-history.json. session_handles.json is a **v2 artifact** — v3's
# postbox never opens it (`grep -rn session_handles althing/` is empty; resolve_config
# takes --handle then ALTHING_HANDLE and nothing else), and the tool that used to
# maintain it, `althing-cli use`, was deleted at the cutover. Whatever is in it now is
# hand-kept and drifts silently.
#
# launch-history.json is written by dev_launch, which is the thing that sets
# ALTHING_HANDLE in the first place, so it is the real cwd->handle binding. Shape is
# {cwd: {command: {at, handle}}} with several commands per directory (claude, kimi,
# grok), so take the most recent by `at` rather than whichever key sorts first.
if [ -z "$h" ] && [ -n "$cwd" ]; then
h=$(jq -r --arg d "$cwd" '(.[$d] // {}) | to_entries | max_by(.value.at) | .value.handle // empty' \
"$HOME/.althing/launch-history.json" 2>/dev/null)
# Fallback only: broader coverage, but frozen and hand-maintained.
[ -z "$h" ] && h=$(jq -r --arg d "$cwd" '.[$d] // empty' "$HOME/.althing/session_handles.json" 2>/dev/null)
fi
if [ -n "$h" ]; then
# `timeout` is load-bearing, not belt-and-braces: v2 read a local SQLite file,
# v3 makes an HTTP call. An unreachable post office must cost this segment two
# seconds and nothing else — a statusline that hangs blocks the whole prompt.
unread=$(timeout 2 postbox --handle "$h" status --json 2>/dev/null </dev/null | jq -r '.unread // 0' 2>/dev/null)
case "${unread:-0}" in ''|0|*[!0-9]*) : ;; *) althing="📬 $unread" ;; esac
# v3 has ONE arming mechanism where v2 had three: `althing-listen` takes
# wake-listener-<handle>.lock. monitor-*.lock and light-monitor-*.lock belonged
# to binaries that no longer exist. kill -0 discards a crashed listener's lock.
mon="🔕"
lk="$HOME/.althing/wake-listener-$h.lock"
if [ -f "$lk" ]; then
pid=$(tr -dc '0-9' < "$lk" 2>/dev/null)
[ -n "$pid" ] && kill -0 "$pid" 2>/dev/null && mon="🔔"
fi
fi
fi
# --- assemble ---
lead="$althing"
[ -n "$mon" ] && lead="${lead:+$lead }$mon"
pg="$proj"
[ -n "$gitseg" ] && pg="${pg:+$pg }$gitseg"
parts="$lead"
[ -n "$pg" ] && parts="${parts:+$parts | }$pg"
parts="${parts:+$parts | }$model | $ctx_seg $toks | \$$cost"
[ -n "$rate" ] && parts="$parts | $rate"
printf '%s' "$parts"
+153
View File
@@ -0,0 +1,153 @@
# althing v3 — the post office. The fleet's message bus.
#
# MOVED nh3-dev -> nh3-docker on 2026-08-28 (operator: "I want it on the docker
# machine — that was always the goal"). It was deployed to nh3-dev at the U9b flag
# day because the herald lives there; but the herald is the piece that MUST be
# host-local (it reads route files and pokes FIFOs), and the post office is
# explicitly the piece that is not.
#
# Three reasons the dev box was wrong:
# - our own server table calls nh3-dev "not a Docker-stack host", and NH-site
# non-GPU services belong here.
# - nh3-dev had three OOM events in fourteen days with the interval HALVING
# (2026-08-14, -26, -28), and the confirmed hog is Claude Code sessions at
# 5-18 GB each, which is that box's actual job.
# - `mem_limit` protects the fleet FROM the post office. It does nothing to
# protect the post office from the box: oom_score_adj is 0, so it was an
# ordinary kill candidate, and the 08-28 sweep took althing-herald and
# uvicorn. A sweep that takes the post office takes mail for all 73 handles.
#
# ⚠ MOVING THE DATA: `docker stop` does NOT checkpoint the WAL. Measured on the
# 08-28 move: post_office.db was 155 KB / mtime 15:09 while post_office.db-wal
# was 4.1 MB / mtime 16:56 — every recent message lived in the WAL. Copying the
# .db alone yields a database that opens cleanly, passes a smoke test, and is
# missing the day's mail. Stop the container, then
# `PRAGMA wal_checkpoint(TRUNCATE)` explicitly, then verify row counts on BOTH
# sides before deleting anything.
#
# Deploy: scripts/deploy-stack.sh nh3-docker althing-post-office
# Image: built from the althing repo's Dockerfile (vh/althing @ v3.0.0), pushed to
# the gitea registry 2026-08-28. Rebuild + republish:
#
# cd ~/development/althing
# docker build -t gitea.phasefinal.com/claude-bot/althing-post-office:<ver> .
# echo $(cat ~/.config/claude-bot/gitea-token) | \
# docker login gitea.phasefinal.com -u claude-bot --password-stdin
# docker push gitea.phasefinal.com/claude-bot/althing-post-office:<ver>
#
# then update the digest below and redeploy. Supersedes the
# `docker save | ssh | docker load` hand-carry the move originally used.
#
# ⚠ NAMESPACE IS `claude-bot`, NOT `vh`. claude-bot's token carries write:package
# but package namespaces are owned: pushing to `vh/...` returns
# `unauthorized: authentication required` AFTER a successful `docker login`, which
# reads like a credential fault and is actually an ownership one. Publishing under
# claude-bot's own namespace also satisfies the standing directive to stop reusing
# the operator's personal credentials for infra work. Both hosts are logged in as
# claude-bot; a new host needs that login before it can pull.
services:
post-office:
# Digest-pinned, not tag-floating: `:3.0.0` is a mutable pointer on a registry
# anyone can re-push, and this container is the fleet's whole message bus. The
# tag is kept alongside the digest purely so a human can read what it is.
image: gitea.phasefinal.com/claude-bot/althing-post-office:3.0.0@sha256:410fed41fa049c41cf83577fd2e49831ea1959ca50b597a8bd35a548e267cf01
container_name: althing-post-office
# ─── Host networking, so the bind guard keeps working ────────────
#
# ⚠ DO NOT "fix" this into a bridged container with `-p`. api.py's
# resolve_bind_host refuses any address that resolves to a wildcard, and it
# tests the RESOLVED PROPERTY rather than matching strings, so there is no
# spelling of "everything" that gets past it. A bridged container cannot
# satisfy that guard honestly: inside its own netns the only reachable bind
# is a wildcard, and publishing the port would move access control from the
# address the application checks to a `-p` flag it cannot see.
#
# Reachability on the private network IS the authorisation story here —
# there is no login and none is wanted.
#
# Cost, stated plainly: no network namespace isolation, and port 8390 is
# claimed host-wide. For a single-service private-network deployment that is
# the right trade, but it IS a trade.
network_mode: host
# The store is the only thing that must survive. Named volume rather than a
# bind mount: uid 1000 inside the container owns it, and docker creates it
# with the right ownership instead of inheriting the host path's.
volumes:
- post-office-data:/var/lib/althing
# The private address of THIS box. Single place it is named; the image ships
# no default on purpose, so a deployment that omits it is refused at startup
# rather than binding wide.
environment:
ALTHING_BIND_HOST: "10.100.50.40"
ALTHING_PORT: "8390"
ALTHING_DB: "/var/lib/althing/post_office.db"
# `unless-stopped` rather than `always` so an operator who deliberately stops
# it during a flag day does not find it running again after a reboot.
restart: unless-stopped
# One Python interpreter holding one SQLite connection; it idles far below
# this. The cap is not a tuning parameter, it is a promise that the post
# office can never be its host's next OOM story.
#
# ⚠ nh3-docker runs Compose v5, which ignores `version:` and honours
# `mem_limit` directly. On a docker-compose 1.x host this needs schema 2.4 —
# under 3.x the key moves to `deploy:`, which is swarm-only and SILENTLY
# IGNORED. Verify with `docker inspect` (want 536870912), never by reading
# the yaml: a cap that does nothing reads as protection.
mem_limit: 512m
# ⚠ The fleet's entire bus. Make the kernel shoot almost anything else first.
# This is the gap the nh3-dev deployment had: a 512m cap and oom_score_adj 0
# means "cannot cause an OOM, is an ordinary victim of one".
oom_score_adj: -500
# Docker's default json-file driver has no size limit. v2's herald left a
# 60 MB log on nh3-dev; an uncapped container log is the same mistake with a
# different name.
logging:
driver: json-file
options:
max-size: "10m"
max-file: "5"
# Inherited from the image, restated so it is visible at deploy time rather
# than only in `docker inspect`.
healthcheck:
test:
- CMD
- python
- -c
- "import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://10.100.50.40:8390/',timeout=4).status==200 else 1)"
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
# `Toolchain` is an EXISTING group under the existing Toolchain tab in
# stacks/homepage/conf/settings.yaml — "the plumbing", which is where a
# message bus belongs. Naming a group the layout has never heard of gets no
# `tab:` and renders the group on ALL FOUR tabs (the Scriberr "AI Systems"
# bug, 2026-08-23), so this must stay a group that already exists.
#
# ⚠ Labels bind at container CREATION. Editing them needs `up -d`, never
# `restart` — a restart leaves the old labels in place and the dashboard
# keeps showing whatever was there before.
#
# nh3-docker is a discovered Docker host in homepage's docker.yaml (as
# `nh3-pfi-docker`), so the label is enough; do NOT also add a services.yaml
# entry or the card renders twice.
labels:
- homepage.group=Toolchain
- homepage.name=althing post office
- homepage.icon=mdi-mailbox
- homepage.description=althing v3 message bus — handles, unread counts, node liveness
- homepage.href=http://10.100.50.40:8390
volumes:
post-office-data:
name: althing-post-office-data
+2 -2
View File
@@ -327,8 +327,8 @@ model_list:
# down, weights intact) if it is ever wanted back. Not repointed to mog-sec
# -- a security model is not an RP-reasoning model (no false aliases). ---
# --- sec (was mog-sec, renamed 2026-08-21) -> M.O.G.-SEC-27B pen-test seat (ana-ml2 GPU1 :8019, in the retired
# fable slot). Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX, stock-Qwen3.8-27B
# --- sec (was mog-sec, renamed 2026-08-21) -> M.O.G.-SEC-27B pen-test seat (ana-ml2 GPU0 :8019 — moved off GPU1
# 2026-08-28, GPU1 no longer had room). Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX, stock-Qwen3.8-27B
# base, quantized in-house to mixed NVFP4+FP8 with MTP + vision preserved.
# Served at native 262K (NOT the card's 1M -- that needs YaRN + SGLang/DFlash2,
# not our vLLM path). presence_penalty deliberately 0.0, NOT the fleet's 1.5:
+14 -4
View File
@@ -1,5 +1,15 @@
# mog-sec — the pen-test seat on ana-ml2 GPU 1 (:8019), in the slot the retired
# fablefusion-charrp-probe used to occupy.
# mog-sec — the pen-test seat on ana-ml2 GPU 0 (:8019).
#
# MOVED GPU 1 -> GPU 0 on 2026-08-28 (operator-directed). GPU 1 carries the five
# resident fleet seats (gen 46 GB + embed + coder + rerank + reward = ~69.9 GB of
# 97.9), leaving ~28 GB — less than the ~51 GB this seat reserves at
# MOG_GPU_MEM_UTIL=0.52, so it could no longer start there. GPU 0 has been idle
# since run 3c was stopped.
#
# ⚠ POWER: bringing this up re-arms the two-GPU load condition that tripped the
# Anaheim rack breaker on 2026-08-26. One circuit feeds the whole rack including
# ana-gw and ana-wg, so a trip costs the site AND the remote path in. Idle draw is
# negligible (~6-13 W/card); the risk is sec and gen under concurrent load.
#
# Serves Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX-BF16 (stock-Qwen3.8-27B-based,
# vision-intact Qwen3_5ForConditionalGeneration, base-graft MTP head), quantized
@@ -108,7 +118,7 @@ services:
devices:
- driver: nvidia
device_ids:
- "${MOG_GPU_ID:-1}"
- "${MOG_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
@@ -123,7 +133,7 @@ services:
- homepage.group=AI - Inference
- homepage.name=M.O.G.-SEC 27B (pen-test)
- homepage.icon=mdi-shield-lock
- homepage.description=Uncensored security model, Qwen3.8-27B NVFP4+MTP, 262K — the `mog-sec` seat (ana-ml2 GPU 1)
- homepage.description=Uncensored security model, Qwen3.8-27B NVFP4+MTP, 262K — the `mog-sec` seat (ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${MOG_PORT:-8019}/docs
networks:
+59
View File
@@ -0,0 +1,59 @@
# phasefinal-web
The Phase Final, Inc. corporate site — a single static page served by nginx on
**ana-docker**, fronted by Traefik at `www.phasefinal.com`.
Built from the design brief at `docs/design/` (boothed 2026-08-29); the markup
and CSS came from a Claude Design session and are checked in verbatim under
`conf/site/` apart from the two post-export steps below.
## Two steps every fresh export needs
The design tool ships the site with **no font binaries** and the `@font-face`
block **commented out**, deliberately, to honour the brief's zero-external-
requests rule. Both must be undone on import:
1. Drop the four woff2 files into `conf/site/fonts/`
(Space Grotesk variable, IBM Plex Sans variable, IBM Plex Mono 400 + 600 —
all SIL OFL, self-hosted, no CDN).
2. Uncomment the four `@font-face` rules at the top of `conf/site/style.css`.
Without these the page silently falls back to system fonts and looks wrong
rather than broken, which is the failure mode you won't notice.
## Routing — and the healthcheck trap that looks like a routing bug
Routers come from the compose labels: `phasefinal-web` serves
`www.phasefinal.com`; `phasefinal-apex` 301s `phasefinal.com` to it.
⚠ **Traefik silently skips containers Docker reports as unhealthy.** There is no
error, no log line, and no router — it looks exactly like a broken provider. The
first deploy of this stack hit that: the healthcheck used
`http://localhost/`, which resolves to `::1` in `nginx:alpine`, and nginx listens
on IPv4 only, so the check got "Connection refused" forever and the container
never left `unhealthy`. **Use `127.0.0.1`, never `localhost`, in a healthcheck
for an IPv4-only listener** — and if a correctly-labelled container never appears
in `/api/http/routers`, check `docker inspect --format '{{.State.Health.Status}}'`
*before* suspecting Traefik.
## Deploy
```bash
scripts/deploy-stack.sh ana-docker phasefinal-web
ssh infra-ops@10.250.50.70 'cd /opt/docker/compose/phasefinal-web && docker compose up -d'
```
## DNS
`www.phasefinal.com` → A `38.120.12.44` (the Anaheim public IP, forwarded to
ana-docker's Traefik). The apex `phasefinal.com` has **no** record — the site is
www-only by operator decision. Cloudflare-proxied is the intended end state, for
edge caching and the automatic email-address obfuscation the contact block
relies on.
## Content constraints
The copy is governed by hard legal and voice constraints — client naming rules,
the never-"military"/"weapons" phrasing rule, no named individuals. **Read the
design brief before editing any text.** Edits are made surgically in
`conf/site/index.html`; there is no CMS and there will not be one.
@@ -0,0 +1,50 @@
# Cloudflare state for phasefinal.com
Not applied by any script — recorded here so the edge configuration is
reviewable and reproducible rather than living only in someone's dashboard.
## DNS
| name | type | value | proxied |
|---|---|---|---|
| `www.phasefinal.com` | A | `38.120.12.44` | yes |
| `phasefinal.com` | A | `38.120.12.44` | yes |
`38.120.12.44` is the Anaheim public IP, forwarded to ana-docker's Traefik.
The apex 301s to `www` via a Traefik middleware, not a Cloudflare rule.
## Cache ruleset
`cache-ruleset.json` is the `http_request_cache_settings` phase entrypoint.
Applied with:
```bash
TOK=$(secret get nh3-dev/.config/cloudflare/phasefinal-cache-token)
curl -X PUT "https://api.cloudflare.com/client/v4/zones/$ZONE/rulesets/phases/http_request_cache_settings/entrypoint" \
-H "Authorization: Bearer $TOK" -H 'Content-Type: application/json' \
--data @cache-ruleset.json
```
**Edge and browser TTL are both `respect_origin` on purpose.** Cache policy is
declared once, in `../conf/nginx.conf`, which is in git — 300s on the document
so a dictated edit goes live quickly, a year on the immutable assets. Setting a
fixed edge TTL here would split that policy across two systems that then have to
be kept in agreement by memory.
## Zone settings
- **Always Online: on.** This, not the cache TTL, is what actually keeps the
site reachable during an origin outage — Cloudflare serves a crawler-archived
copy. With a 300s document TTL, caching alone would only mask about five
minutes.
- **Email Address Obfuscation: on** (Cloudflare's default). The contact block
depends on it — `inquiry@phasefinal.com` is written as a plain `mailto:` in
the markup and rewritten at the edge. Do not hand-obfuscate it in the HTML;
that interferes.
## Token
`nh3-dev/.config/cloudflare/phasefinal-cache-token` in the vault. Scoped to this
zone. Note it is an **account-scoped** token, so `/user/tokens/verify` reports
"Invalid API Token" while `/accounts/<id>/tokens/verify` reports `active` —
check the account endpoint, not the user one.
@@ -0,0 +1,14 @@
{
"rules": [
{
"expression": "(http.host eq \"www.phasefinal.com\")",
"action": "set_cache_settings",
"action_parameters": {
"cache": true,
"edge_ttl": { "mode": "respect_origin" },
"browser_ttl": { "mode": "respect_origin" }
},
"description": "PFI site: cache eligible, honour origin Cache-Control (policy lives in stacks/phasefinal-web/conf/nginx.conf)"
}
]
}
+40
View File
@@ -0,0 +1,40 @@
services:
web:
image: nginx:alpine
container_name: phasefinal-web
restart: unless-stopped
volumes:
- /opt/docker/conf/phasefinal-web/site:/usr/share/nginx/html:ro
- /opt/docker/conf/phasefinal-web/nginx.conf:/etc/nginx/conf.d/default.conf:ro
networks:
- tnet
healthcheck:
# 127.0.0.1, NOT localhost: in this image `localhost` resolves to ::1 and
# nginx listens on IPv4 only, so `localhost` gives "Connection refused"
# forever. Traefik SKIPS unhealthy containers, so a wrong healthcheck here
# silently means "no route" rather than "unhealthy service".
test: ["CMD", "wget", "-qO-", "http://127.0.0.1:80/"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
labels:
- traefik.enable=true
- traefik.http.routers.phasefinal-web.rule=Host(`www.phasefinal.com`)
- traefik.http.routers.phasefinal-web.tls=true
- traefik.http.routers.phasefinal-web.tls.certresolver=anaprod
- traefik.http.services.phasefinal-web.loadbalancer.server.port=80
# apex -> www, 301
- traefik.http.routers.phasefinal-apex.rule=Host(`phasefinal.com`)
- traefik.http.routers.phasefinal-apex.tls=true
- traefik.http.routers.phasefinal-apex.tls.certresolver=anaprod
- traefik.http.routers.phasefinal-apex.service=phasefinal-web
- traefik.http.routers.phasefinal-apex.middlewares=phasefinal-to-www
- traefik.http.middlewares.phasefinal-to-www.redirectregex.regex=^https?://phasefinal\.com/(.*)
- traefik.http.middlewares.phasefinal-to-www.redirectregex.replacement=https://www.phasefinal.com/$${1}
- traefik.http.middlewares.phasefinal-to-www.redirectregex.permanent=true
networks:
tnet:
name: traefik-net
external: true
+18
View File
@@ -0,0 +1,18 @@
server {
listen 80;
server_name _;
root /usr/share/nginx/html;
index index.html;
# Static site; long cache on immutable assets, short on the document so
# a dictated edit goes live on the next deploy without a purge.
location = / { add_header Cache-Control "public, max-age=300"; try_files /index.html =404; }
location ~* \.(woff2|svg|png|ico)$ { add_header Cache-Control "public, max-age=31536000, immutable"; }
location ~* \.css$ { add_header Cache-Control "public, max-age=3600"; }
add_header X-Content-Type-Options nosniff always;
add_header X-Frame-Options DENY always;
add_header Referrer-Policy no-referrer always;
location / { try_files $uri $uri/ =404; }
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

@@ -0,0 +1,19 @@
<?xml version="1.0"?><svg xmlns="http://www.w3.org/2000/svg" xml:space="preserve" version="1.1" style="shape-rendering:geometricPrecision; text-rendering:geometricPrecision; image-rendering:optimizeQuality; fill-rule:evenodd; clip-rule:evenodd" viewBox="1000 1820 6280 1800" xmlns:xlink="http://www.w3.org/1999/xlink">
<defs>
</defs>
<g id="Foreground">
<metadata id="CorelCorpID_0Corel-Layer"></metadata>
<polygon fill="#ffffff" points="3750,3050 3655,3050 3655,2854 3418,2854 3418,3050 3322,3050 3322,2582 3418,2582 3418,2772 3655,2772 3655,2582 3750,2582 "></polygon>
<path fill="#ffffff" d="M4282 3050l-106 0 -43 -112 -205 0 -41 112 -100 0 195 -468 102 0 198 468zm-181 -194l0 0 -71 -186 -70 186 141 0z"></path>
<path fill="#ffffff" d="M4732 2911c0,40 -9,72 -26,93 -25,31 -70,46 -136,46l-259 0 0 -82 244 0c25,0 44,-3 56,-10 14,-8 21,-22 21,-43 0,-20 -5,-35 -17,-44 -11,-9 -29,-13 -54,-13l-115 0c-51,0 -89,-10 -114,-31 -27,-22 -40,-57 -40,-104 0,-49 13,-85 39,-108 26,-22 67,-33 122,-33l256 0 0 82 -240 0c-52,0 -78,21 -78,63 0,17 6,30 18,39 11,9 27,13 46,13l129 0c46,0 79,6 99,18 33,19 49,57 49,114z"></path>
<path fill="#ffffff" d="M5158 3050l-164 0c-67,0 -119,-17 -156,-52 -40,-38 -60,-96 -60,-173 0,-91 19,-155 58,-193 34,-34 88,-50 161,-50l161 0 0 79 -160 0c-35,0 -62,9 -80,27 -19,19 -29,48 -29,86l269 0 0 79 -271 0c0,44 11,75 31,94 18,16 44,24 80,24l160 0 0 79z"></path>
<path fill="#90c73e" d="M5638 2857l-278 0 0 193 -97 0 0 -215c0,-59 2,-101 7,-124 9,-42 28,-73 58,-93 36,-24 90,-36 162,-36l148 0 0 79 -156 0c-43,0 -75,11 -95,33 -18,19 -27,48 -27,84l278 0 0 79z"></path>
<polygon fill="#90c73e" points="5806,3050 5706,3050 5706,2582 5806,2582 "></polygon>
<polygon fill="#90c73e" points="6285,3050 6174,3050 5973,2694 5973,3050 5878,3050 5878,2582 6001,2582 6193,2920 6193,2582 6285,2582 "></polygon>
<path fill="#90c73e" d="M7144 3050l-139 0c-66,0 -113,-16 -141,-47 -25,-29 -38,-77 -38,-142l0 -279 95 0 0 244c0,5 0,10 0,14 0,5 0,9 0,12 0,34 5,58 15,74 16,23 47,35 92,35l116 0 0 89z"></path>
<path fill="#ffffff" d="M2656 3296c0,101 -83,183 -184,183l-1295 0 0 -1098c0,-101 82,-205 184,-205l1111 0c101,0 184,82 184,183l0 937zm-201 -542c0,-208 -169,-376 -377,-376 -2,0 -38,0 -38,0l-62 213c52,0 100,0 101,0 90,0 163,73 163,163 0,90 -73,164 -163,164 -1,0 -141,0 -193,0l-62 212c0,0 252,1 254,1 208,0 377,-169 377,-377zm-848 0c0,-90 73,-163 163,-163 0,0 73,0 125,0l62 -213c0,0 -184,0 -187,0 -207,0 -375,168 -376,375l0 0 0 555 212 0 0 -178 133 0 62 -211 -195 0 0 -165 1 0z"></path>
<path fill="#ffffff" d="M3236 2618c-31,-24 -82,-36 -155,-36l-212 0 0 82 96 0 0 0 114 0c33,0 57,5 73,15 15,10 23,29 23,55 0,47 -31,70 -91,70l-119 0 0 -95 -96 0 0 341 96 0 0 -164 120 0c2,0 4,1 6,1 4,0 8,0 12,0 43,0 79,-8 108,-23 46,-25 70,-68 70,-130 0,-53 -15,-92 -45,-116z"></path>
<polygon fill="#90c73e" points="6616,2582 6515,2582 6320,3050 6420,3050 6563,2670 6634,2856 6531,2856 6501,2938 6665,2938 6709,3050 6814,3050 "></polygon>
</g>
</svg>

After

Width:  |  Height:  |  Size: 2.9 KiB

@@ -0,0 +1,19 @@
<?xml version="1.0"?><svg xmlns="http://www.w3.org/2000/svg" xml:space="preserve" version="1.1" style="shape-rendering:geometricPrecision; text-rendering:geometricPrecision; image-rendering:optimizeQuality; fill-rule:evenodd; clip-rule:evenodd" viewBox="1000 1820 6280 1800" xmlns:xlink="http://www.w3.org/1999/xlink">
<defs>
</defs>
<g id="Foreground">
<metadata id="CorelCorpID_0Corel-Layer"></metadata>
<polygon fill="#ffffff" points="3750,3050 3655,3050 3655,2854 3418,2854 3418,3050 3322,3050 3322,2582 3418,2582 3418,2772 3655,2772 3655,2582 3750,2582 "></polygon>
<path fill="#ffffff" d="M4282 3050l-106 0 -43 -112 -205 0 -41 112 -100 0 195 -468 102 0 198 468zm-181 -194l0 0 -71 -186 -70 186 141 0z"></path>
<path fill="#ffffff" d="M4732 2911c0,40 -9,72 -26,93 -25,31 -70,46 -136,46l-259 0 0 -82 244 0c25,0 44,-3 56,-10 14,-8 21,-22 21,-43 0,-20 -5,-35 -17,-44 -11,-9 -29,-13 -54,-13l-115 0c-51,0 -89,-10 -114,-31 -27,-22 -40,-57 -40,-104 0,-49 13,-85 39,-108 26,-22 67,-33 122,-33l256 0 0 82 -240 0c-52,0 -78,21 -78,63 0,17 6,30 18,39 11,9 27,13 46,13l129 0c46,0 79,6 99,18 33,19 49,57 49,114z"></path>
<path fill="#ffffff" d="M5158 3050l-164 0c-67,0 -119,-17 -156,-52 -40,-38 -60,-96 -60,-173 0,-91 19,-155 58,-193 34,-34 88,-50 161,-50l161 0 0 79 -160 0c-35,0 -62,9 -80,27 -19,19 -29,48 -29,86l269 0 0 79 -271 0c0,44 11,75 31,94 18,16 44,24 80,24l160 0 0 79z"></path>
<path fill="#ffffff" d="M5638 2857l-278 0 0 193 -97 0 0 -215c0,-59 2,-101 7,-124 9,-42 28,-73 58,-93 36,-24 90,-36 162,-36l148 0 0 79 -156 0c-43,0 -75,11 -95,33 -18,19 -27,48 -27,84l278 0 0 79z"></path>
<polygon fill="#ffffff" points="5806,3050 5706,3050 5706,2582 5806,2582 "></polygon>
<polygon fill="#ffffff" points="6285,3050 6174,3050 5973,2694 5973,3050 5878,3050 5878,2582 6001,2582 6193,2920 6193,2582 6285,2582 "></polygon>
<path fill="#ffffff" d="M7144 3050l-139 0c-66,0 -113,-16 -141,-47 -25,-29 -38,-77 -38,-142l0 -279 95 0 0 244c0,5 0,10 0,14 0,5 0,9 0,12 0,34 5,58 15,74 16,23 47,35 92,35l116 0 0 89z"></path>
<path fill="#ffffff" d="M2656 3296c0,101 -83,183 -184,183l-1295 0 0 -1098c0,-101 82,-205 184,-205l1111 0c101,0 184,82 184,183l0 937zm-201 -542c0,-208 -169,-376 -377,-376 -2,0 -38,0 -38,0l-62 213c52,0 100,0 101,0 90,0 163,73 163,163 0,90 -73,164 -163,164 -1,0 -141,0 -193,0l-62 212c0,0 252,1 254,1 208,0 377,-169 377,-377zm-848 0c0,-90 73,-163 163,-163 0,0 73,0 125,0l62 -213c0,0 -184,0 -187,0 -207,0 -375,168 -376,375l0 0 0 555 212 0 0 -178 133 0 62 -211 -195 0 0 -165 1 0z"></path>
<path fill="#ffffff" d="M3236 2618c-31,-24 -82,-36 -155,-36l-212 0 0 82 96 0 0 0 114 0c33,0 57,5 73,15 15,10 23,29 23,55 0,47 -31,70 -91,70l-119 0 0 -95 -96 0 0 341 96 0 0 -164 120 0c2,0 4,1 6,1 4,0 8,0 12,0 43,0 79,-8 108,-23 46,-25 70,-68 70,-130 0,-53 -15,-92 -45,-116z"></path>
<polygon fill="#ffffff" points="6616,2582 6515,2582 6320,3050 6420,3050 6563,2670 6634,2856 6531,2856 6501,2938 6665,2938 6709,3050 6814,3050 "></polygon>
</g>
</svg>

After

Width:  |  Height:  |  Size: 2.9 KiB

@@ -0,0 +1,19 @@
<?xml version="1.0"?><svg xmlns="http://www.w3.org/2000/svg" xml:space="preserve" version="1.1" style="shape-rendering:geometricPrecision; text-rendering:geometricPrecision; image-rendering:optimizeQuality; fill-rule:evenodd; clip-rule:evenodd" viewBox="1000 1820 6280 1800" xmlns:xlink="http://www.w3.org/1999/xlink">
<defs>
</defs>
<g id="Foreground">
<metadata id="CorelCorpID_0Corel-Layer"></metadata>
<polygon fill="#14191f" points="3750,3050 3655,3050 3655,2854 3418,2854 3418,3050 3322,3050 3322,2582 3418,2582 3418,2772 3655,2772 3655,2582 3750,2582 "></polygon>
<path fill="#14191f" d="M4282 3050l-106 0 -43 -112 -205 0 -41 112 -100 0 195 -468 102 0 198 468zm-181 -194l0 0 -71 -186 -70 186 141 0z"></path>
<path fill="#14191f" d="M4732 2911c0,40 -9,72 -26,93 -25,31 -70,46 -136,46l-259 0 0 -82 244 0c25,0 44,-3 56,-10 14,-8 21,-22 21,-43 0,-20 -5,-35 -17,-44 -11,-9 -29,-13 -54,-13l-115 0c-51,0 -89,-10 -114,-31 -27,-22 -40,-57 -40,-104 0,-49 13,-85 39,-108 26,-22 67,-33 122,-33l256 0 0 82 -240 0c-52,0 -78,21 -78,63 0,17 6,30 18,39 11,9 27,13 46,13l129 0c46,0 79,6 99,18 33,19 49,57 49,114z"></path>
<path fill="#14191f" d="M5158 3050l-164 0c-67,0 -119,-17 -156,-52 -40,-38 -60,-96 -60,-173 0,-91 19,-155 58,-193 34,-34 88,-50 161,-50l161 0 0 79 -160 0c-35,0 -62,9 -80,27 -19,19 -29,48 -29,86l269 0 0 79 -271 0c0,44 11,75 31,94 18,16 44,24 80,24l160 0 0 79z"></path>
<path fill="#90c73e" d="M5638 2857l-278 0 0 193 -97 0 0 -215c0,-59 2,-101 7,-124 9,-42 28,-73 58,-93 36,-24 90,-36 162,-36l148 0 0 79 -156 0c-43,0 -75,11 -95,33 -18,19 -27,48 -27,84l278 0 0 79z"></path>
<polygon fill="#90c73e" points="5806,3050 5706,3050 5706,2582 5806,2582 "></polygon>
<polygon fill="#90c73e" points="6285,3050 6174,3050 5973,2694 5973,3050 5878,3050 5878,2582 6001,2582 6193,2920 6193,2582 6285,2582 "></polygon>
<path fill="#90c73e" d="M7144 3050l-139 0c-66,0 -113,-16 -141,-47 -25,-29 -38,-77 -38,-142l0 -279 95 0 0 244c0,5 0,10 0,14 0,5 0,9 0,12 0,34 5,58 15,74 16,23 47,35 92,35l116 0 0 89z"></path>
<path fill="#14191f" d="M2656 3296c0,101 -83,183 -184,183l-1295 0 0 -1098c0,-101 82,-205 184,-205l1111 0c101,0 184,82 184,183l0 937zm-201 -542c0,-208 -169,-376 -377,-376 -2,0 -38,0 -38,0l-62 213c52,0 100,0 101,0 90,0 163,73 163,163 0,90 -73,164 -163,164 -1,0 -141,0 -193,0l-62 212c0,0 252,1 254,1 208,0 377,-169 377,-377zm-848 0c0,-90 73,-163 163,-163 0,0 73,0 125,0l62 -213c0,0 -184,0 -187,0 -207,0 -375,168 -376,375l0 0 0 555 212 0 0 -178 133 0 62 -211 -195 0 0 -165 1 0z"></path>
<path fill="#14191f" d="M3236 2618c-31,-24 -82,-36 -155,-36l-212 0 0 82 96 0 0 0 114 0c33,0 57,5 73,15 15,10 23,29 23,55 0,47 -31,70 -91,70l-119 0 0 -95 -96 0 0 341 96 0 0 -164 120 0c2,0 4,1 6,1 4,0 8,0 12,0 43,0 79,-8 108,-23 46,-25 70,-68 70,-130 0,-53 -15,-92 -45,-116z"></path>
<polygon fill="#90c73e" points="6616,2582 6515,2582 6320,3050 6420,3050 6563,2670 6634,2856 6531,2856 6501,2938 6665,2938 6709,3050 6814,3050 "></polygon>
</g>
</svg>

After

Width:  |  Height:  |  Size: 2.9 KiB

@@ -0,0 +1 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 32 32"><rect width="32" height="32" rx="7" fill="#90C73E"></rect></svg>

After

Width:  |  Height:  |  Size: 124 B

+107
View File
@@ -0,0 +1,107 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Phase Final, Inc. — Systems Engineering Consultancy</title>
<meta name="description" content="Phase Final, Inc. is a systems engineering consultancy offering systems engineering, infrastructure hosting, and AI services, with a track record in hardware, firmware, and software engineering.">
<link rel="icon" href="favicon.svg" type="image/svg+xml">
<link rel="stylesheet" href="style.css">
</head>
<body>
<header id="hero">
<div class="hero-inner">
<img class="wordmark" src="assets/pfi-wordmark-duotone-dark.svg" alt="Phase Final, Inc.">
<div class="hero-copy">
<p class="lede">A systems engineering consultancy.</p>
<p class="sublede">Phase Final, Inc. has a long track record in hardware, firmware, and software engineering across an unusually wide range of industries: consumer electronics, financial services, health services, government, film and video game production, and industrial automation.</p>
</div>
<p class="spec-bar">
<span>Systems engineering</span>
<span>Infrastructure hosting</span>
<span>AI services</span>
<span class="spec-accent">Fixed-bid or retainer</span>
</p>
</div>
</header>
<main>
<section id="services">
<h2><span class="num">01</span>What we do</h2>
<div class="section-body">
<ul class="service-list">
<li>
<h3>Systems Engineering</h3>
<p>Design and integration of complete systems, hardware and software. Requirements, architecture, implementation, and verification are treated as one problem with one owner.</p>
</li>
<li>
<h3>Infrastructure Hosting</h3>
<p>Hosting and operation of client infrastructure. Systems the firm builds can run on infrastructure the firm operates, under the same accountability.</p>
</li>
<li>
<h3>AI Services &amp; Engineering</h3>
<p>Engineering and operation of AI systems. Model-backed capability is built, integrated, and run as production software, held to the same standard as everything else the firm ships.</p>
</li>
</ul>
</div>
</section>
<section id="engagement">
<h2><span class="num">02</span>How we engage</h2>
<div class="section-body">
<p>Engagements are fixed-bid or retainer. Fixed-bid where the scope can be specified: a defined deliverable at a defined price. Retainer where the work is continuous: operation, maintenance, and engineering capacity over time.</p>
<p class="referral-note">Most new work arrives by referral.</p>
</div>
</section>
<section id="track-record">
<h2><span class="num">03</span>Where we've done it</h2>
<div class="section-body">
<dl class="record-list">
<div class="record">
<dt><span class="idx">A</span> High-end consumer audio and electronics</dt>
<dd>Firmware design, software tooling, and R&amp;D engineering for <strong>Audeze</strong> and <strong>Sony</strong>.</dd>
</div>
<div class="record">
<dt><span class="idx">B</span> Portable illumination</dt>
<dd>Illumination and lighting technology for <strong>SureFire</strong>, including flashlights. The firm holds patents in lighting technology.</dd>
</div>
<div class="record">
<dt><span class="idx">C</span> Financial services</dt>
<dd>Mortgage lending underwriting engines for <strong>Fannie Mae</strong> and <strong>Freddie Mac</strong>. Credit card processing and gift card systems.</dd>
</div>
<div class="record">
<dt><span class="idx">D</span> Health services and medical industry</dt>
<dd>Document automation and billing systems.</dd>
</div>
<div class="record">
<dt><span class="idx">E</span> Government and DoD</dt>
<dd>Hardware and firmware, as a subcontractor to prime contractors.</dd>
</div>
<div class="record">
<dt><span class="idx">F</span> Film and video game production</dt>
<dd>Automated asset processing, rapid prototyping, and workflow automation.</dd>
</div>
<div class="record">
<dt><span class="idx">G</span> Industrial automation</dt>
<dd>CNC systems automation, additive manufacturing, and robotics automation.</dd>
</div>
</dl>
</div>
</section>
<section id="contact">
<h2><span class="num">04</span>Contact</h2>
<p class="contact-address"><a href="mailto:inquiry@phasefinal.com">inquiry@phasefinal.com</a></p>
<p class="contact-note">Tell us the spec and the deadline. We'll tell you what it takes to reach the final phase.</p>
</section>
<footer>
<span>Phase Final, Inc.</span>
</footer>
</main>
</body>
</html>
+255
View File
@@ -0,0 +1,255 @@
/* Phase Final, Inc. — single-page site
Brand: Volt #90C73E over cool slate neutrals, per Phase Final tokens.
Zero external requests; no JavaScript.
FONTS: the brand faces are Space Grotesk (display), IBM Plex Sans (body),
IBM Plex Mono (labels). To keep this zip free of external requests they are
referenced by local name only, with system fallbacks. To ship the exact
faces, place woff2 files in fonts/ and uncomment the @font-face blocks. */
@font-face { font-family: "Space Grotesk"; src: url("fonts/SpaceGrotesk-Var.woff2") format("woff2"); font-weight: 300 700; font-display: swap; }
@font-face { font-family: "IBM Plex Sans"; src: url("fonts/IBMPlexSans-Var.woff2") format("woff2"); font-weight: 100 700; font-display: swap; }
@font-face { font-family: "IBM Plex Mono"; src: url("fonts/IBMPlexMono-Regular.woff2") format("woff2"); font-weight: 400; font-display: swap; }
@font-face { font-family: "IBM Plex Mono"; src: url("fonts/IBMPlexMono-SemiBold.woff2") format("woff2"); font-weight: 600; font-display: swap; }
:root {
--ink: #0f1318;
--ink-2: #181d22;
--text: #0f1318;
--text-secondary: #4d5763;
--text-tertiary: #6b7682;
--paper: #f6f8f9;
--border: #c2cad2;
--border-subtle: #dde2e7;
--accent: #90c73e;
--accent-text: #588021; /* AA on light */
--volt-200: #d0e8a0;
--volt-300: #b8db72;
--volt-400: #a3d052;
--volt-500: #90c73e;
--slate-300: #c2cad2;
--slate-400: #97a2ad;
--slate-700: #394149;
--slate-800: #262c33;
--display: "Space Grotesk", "IBM Plex Sans", -apple-system, "Segoe UI", Helvetica, sans-serif;
--sans: "IBM Plex Sans", -apple-system, BlinkMacSystemFont, "Segoe UI", Helvetica, Arial, sans-serif;
--mono: "IBM Plex Mono", ui-monospace, "SF Mono", Menlo, Consolas, monospace;
}
@media (prefers-color-scheme: dark) {
:root {
--text: #f6f8f9;
--text-secondary: #c2cad2;
--text-tertiary: #97a2ad;
--paper: #0f1318;
--border: #394149;
--border-subtle: #262c33;
--accent-text: #b8db72;
--ink: #0a0d11;
--ink-2: #181d22;
}
}
* { box-sizing: border-box; }
body {
margin: 0;
background: var(--paper);
color: var(--text);
font-family: var(--sans);
font-size: 16.5px;
line-height: 1.65;
-webkit-font-smoothing: antialiased;
}
::selection { background: var(--volt-200); color: #0f1318; }
a { color: var(--accent-text); text-decoration: underline; text-underline-offset: 3px; }
p { margin: 0; text-wrap: pretty; }
/* ---- Hero: ink console surface with engineering grid ---- */
#hero {
background: var(--ink);
color: #f6f8f9;
position: relative;
overflow: hidden;
background-image:
linear-gradient(rgba(246,248,249,0.05) 1px, transparent 1px),
linear-gradient(90deg, rgba(246,248,249,0.05) 1px, transparent 1px);
background-size: 48px 48px;
border-bottom: 3px solid var(--volt-500);
}
.hero-inner {
max-width: 920px;
margin: 0 auto;
padding: clamp(56px, 9vw, 104px) 28px clamp(56px, 8vw, 96px);
display: flex;
flex-direction: column;
gap: clamp(40px, 6vw, 64px);
}
.wordmark { height: clamp(48px, 7.5vw, 72px); width: auto; display: block; align-self: flex-start; }
.hero-copy { max-width: 800px; }
.lede {
font-family: var(--display);
font-size: clamp(32px, 5.6vw, 54px);
line-height: 1.1;
letter-spacing: -0.02em;
font-weight: 500;
color: #ffffff;
text-wrap: balance;
}
.sublede {
margin-top: 24px;
font-size: clamp(16.5px, 2vw, 19px);
color: var(--slate-300);
max-width: 62ch;
}
.intro-note {
margin-top: 20px;
font-size: 16px;
color: var(--slate-400);
max-width: 62ch;
}
.spec-bar {
font-family: var(--mono);
font-size: 12.5px;
letter-spacing: 0.07em;
text-transform: uppercase;
color: var(--slate-400);
display: flex;
flex-wrap: wrap;
gap: 10px 0;
}
.spec-bar span { padding: 0 18px; border-left: 1px solid var(--slate-700); }
.spec-bar span:first-child { padding-left: 0; border-left: none; }
.spec-bar .spec-accent { color: var(--volt-300); }
/* ---- Body ---- */
main {
max-width: 920px;
margin: 0 auto;
padding: 0 28px 72px;
}
section {
border-top: 1px solid var(--border);
padding: 52px 0 64px;
display: flex;
flex-wrap: wrap;
gap: 32px;
}
#services { border-top: none; padding-top: clamp(52px, 7vw, 76px); }
h2 {
margin: 0;
flex: 0 0 208px;
font-family: var(--mono);
font-size: 12px;
font-weight: 600;
letter-spacing: 0.15em;
text-transform: uppercase;
color: var(--text-tertiary);
}
h2 .num { color: var(--accent-text); margin-right: 12px; }
.section-body { flex: 1 1 380px; max-width: 620px; }
/* Services */
.service-list {
list-style: none;
margin: 0;
padding: 0;
display: flex;
flex-direction: column;
gap: 30px;
}
.service-list h3 {
margin: 0 0 6px;
font-family: var(--display);
font-size: 22px;
line-height: 1.3;
letter-spacing: -0.01em;
font-weight: 600;
}
.service-list p { color: var(--text-secondary); }
/* Engagement */
#engagement .section-body p + p { margin-top: 18px; }
.referral-note { color: var(--text-secondary); }
/* Track record */
.record-list { margin: 0; }
.record { padding: 24px 0 26px; }
.record:first-child { padding-top: 0; }
.record + .record { border-top: 1px solid var(--border-subtle); }
.record dt {
font-family: var(--mono);
font-size: 12.5px;
letter-spacing: 0.05em;
text-transform: uppercase;
color: var(--text-tertiary);
margin-bottom: 6px;
}
.record .idx { color: var(--accent-text); margin-right: 6px; }
.record dd { margin: 0; font-size: 17.5px; }
.record strong { font-family: var(--display); font-weight: 600; }
/* Contact — ink console panel */
#contact {
display: block;
border-top: none;
background: var(--ink-2);
background-image:
linear-gradient(rgba(246,248,249,0.05) 1px, transparent 1px),
linear-gradient(90deg, rgba(246,248,249,0.05) 1px, transparent 1px);
background-size: 48px 48px;
border-left: 3px solid var(--volt-500);
border-radius: 16px;
padding: clamp(40px, 6vw, 72px);
color: #f6f8f9;
}
#contact h2 { color: var(--slate-400); margin-bottom: 20px; }
#contact h2 .num { color: var(--volt-400); }
.contact-address {
font-size: clamp(23px, 4vw, 38px);
line-height: 1.25;
}
.contact-address a {
font-family: var(--display);
font-weight: 600;
letter-spacing: -0.015em;
color: var(--volt-300);
text-decoration: none;
border-bottom: 2px solid var(--volt-500);
}
.contact-address a:hover { color: var(--volt-200); }
.contact-note {
margin-top: 20px;
color: var(--slate-300);
max-width: 56ch;
font-size: 16px;
}
/* Footer */
footer {
display: flex;
justify-content: space-between;
align-items: baseline;
gap: 16px;
flex-wrap: wrap;
padding-top: 28px;
font-family: var(--mono);
font-size: 12.5px;
color: var(--text-tertiary);
}
@media (max-width: 480px) {
main { padding: 0 20px 56px; }
.hero-inner { padding-left: 20px; padding-right: 20px; }
section { padding: 40px 0 48px; }
}