Files
vh c488eadc31 memory: snapshot — althing v3 fleet-wide at 3.1.1, sec on GPU0
The in-flight section was a day stale: it still described run 3c as the
live subject on a box where nothing had moved. Rewritten around what is
actually true now -- v3 deployed fleet-wide, the post office relocated
to nh3-docker, sec serving on GPU0, run 3c still held on power.

Six new decision entries, three of which carry findings that outlive
their incident: the inbound half of the handle-resolution bug (a stale
ALTHING_HANDLE reads another agent's mailbox and reports it empty,
which is a second route into the failure v3 exists to prevent), the OOM
attribution to Claude Code sessions, and the operator's two explicit
belays recorded so a later session does not re-raise them as new.

Auto-archival fired at 301 lines and moved exactly one entry. Three
others were old enough and every one carries a still-open deferred
pointer -- the parked CI flip, muninn-gate's submit path, and the triton
backend deferred to the Ada refresh. Held back per the guards; an
over-cap file that keeps live decisions beats a scannable one that lost
them. The entry that did move had its deferred item closed today:
nh3-extdev's staged v2.1.2 wheel is moot now that the box runs 3.1.1.
2026-08-28 22:22:32 -07:00

1.9 KiB

[2026-08-28] The sec pen-test seat moved GPU1 → GPU0 and came up — on a circuit that tripped 36h earlier

Operator-directed. sec = M.O.G.-SEC-27B (stacks/mog-sec/, LiteLLM aliases sec / sec-reasoning, ana-ml2 :8019, 262K native ctx, NVFP4+FP8 mixed with a grafted MTP head).

Why it had to move

GPU1  69,895 MiB used of 97,887 — gen 46 GB + embed 9.8 + coder 8.4 + rerank 3.5 + reward 2.1
      -> ~28 GB free, against the ~51 GB this seat reserves at MOG_GPU_MEM_UTIL=0.52
      -> it could not start on GPU1 AT ALL
GPU0  empty since run 3c was stopped 2026-08-26

One line: MOG_GPU_ID=1 -> 0 in /opt/docker/compose/mog-sec/.env. The compose default, the header comment and the homepage label all named GPU 1 and were corrected in the same change — a label naming the wrong card is a record that lies about where the work runs.

Landed state

container   vllm-mog-sec, healthy after ~400s load (22 GB model off the DEGRADED /tank)
GPU0        51,532 MiB, idle draw 16.34 W
GPU1        69,895 MiB, idle draw  6.42 W

This re-arms the two-GPU load condition that tripped the Anaheim rack breaker on 08-26. Idle draw is negligible — ~23 W across both cards. The risk materialises only under concurrent load, when both seats work at once and the box approaches the ~600 W that tripped it. The operator accepted that with the constraint stated. See 2026-08-27-anaheim-breaker-and-onboot-gap — one circuit feeds the whole rack including ana-gw and ana-wg, so a trip costs the site AND the way back in.

Deploy gotchas worth keeping

  • up -d, never restartlabels bind at container creation, so a restart keeps the old homepage label and the dashboard silently keeps showing the old GPU number.
  • Diff deployed-vs-canonical BEFORE pushing. There was no drift here, which is the only reason the push was safe to make blind.