memory: snapshot — gx10 unracked and next up as an inference+training box

Run 3c's new intended home is the GX10 rather than a power triage on ana-ml2: a ~240 W
appliance instead of the kilowatt-class box that tripped the breaker, and 121 GB unified
holds the 49 GB bf16 base comfortably where ana-ml2 was tight. The box is bare, so the
first move is a throughput probe rather than a harness port.

Also banked: the Ada migration settling on zfs send with branch (b) ruled out by irv-ml1
keeping its eight services; the Synapse 39-release upgrade with its one-way schema
migration, the appservice namespace opening and the admin API lockdown; the ratified room
alias convention; and a named failure class — a correct check aimed at the wrong object —
with six instances from one day across three sessions.
This commit is contained in:
2026-09-01 16:32:21 -07:00
parent 73866f6a7e
commit 7142657749
6 changed files with 435 additions and 105 deletions
+95 -95
View File
@@ -1,16 +1,16 @@
# Graph Report - eshpfi-management (2026-08-28)
# Graph Report - eshpfi-management (2026-09-01)
## Corpus Check
- 373 files · ~566,262 words
- 378 files · ~573,530 words
- Verdict: corpus is large enough that graph structure adds value.
## Summary
- 3845 nodes · 4087 edges · 426 communities (385 shown, 41 thin omitted)
- 3878 nodes · 4117 edges · 426 communities (385 shown, 41 thin omitted)
- Extraction: 99% EXTRACTED · 1% INFERRED · 0% AMBIGUOUS · INFERRED: 38 edges (avg confidence: 0.71)
- Token cost: 0 input · 0 output
## Graph Freshness
- Built from commit: `c488eadc`
- Built from commit: `73866f6a`
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
- Run `graphify update .` after code changes (no API cost).
@@ -157,7 +157,7 @@
- description
- 7. Critical Warnings by Model
- think-leak — reproducers for the h300 unterminated-`<think>` defect
- catalog_version
- properties
- method
- test_invocation.py
- scriberr — self-hosted transcription + diarization (ana-ml2, GPU1)
@@ -214,7 +214,7 @@
- zonos-engine — ZONOS2 native TTS engine (`:1920`, irv-ml1 3090)
- omnivoice/app.py
- mOrpheus voice-agent system prompt
- stream_chunks
- phasefinal-web
- build.sh script
- heretic2-nvfp4-quant — fast char-rp-reasoning seat (NVFP4 + MTP)
- booth/app.py
@@ -339,7 +339,7 @@
- [2026-08-23] hrafn adopted; its CI deploy reported green while deploying nothing
- optional
- properties
- adapter/server.py
- gpu.py
- step
- `[2026-08-27]` A transport failure that enters a measurement as a VALUE looks like whatever you hoped to find
- host
@@ -361,7 +361,7 @@
- FOLLOW-UP 3 (2026-08-23): WireGuard over the same internet path does 767 Mbit/s on ONE stream
- [2026-08-23] `/mnt/smithy` mounted on ana-ml2 — read-only and SOFT, deliberately not matching nh3-dev
- `[2026-08-27]` There were TWO run-3c launches, not three — and the phantom third was my reporting
- properties
- $ref
- speaches — OpenAI-compatible ASR (faster-whisper) on irv-ml1
- RESOLVED (2026-08-23): it is the UDM's software AES-CBC. The FortiGate is exonerated.
- judge-bench — evaluate a candidate LLM-as-judge seat
@@ -370,26 +370,26 @@
- `[2026-08-27]` Anaheim tripped a power breaker — and four guests including the NAS never came back
- FOLLOW-UP 2 (2026-08-23): it is NOT a capacity problem, and it IS specific to IPsec
- 2026-08-27-anaheim-breaker-and-onboot-gap.md
- parakeet/app.py
- Cloudflare state for phasefinal.com
- 6. Round-1 aborted; throughput root-caused (measured 2026-08-24 22:00 PDT)
- ChunkResult
- pfi-gx10 — ASUS Ascent GX10 (NVIDIA GB10)
- [2026-08-24] Scriberr transcription deployed on ana-ml2, GPU1
- FOLLOW-UP (2026-08-23): what the per-stream limit actually is
- Gemma-4 26B-A4B ERP/RP tune — GPU sizing adjudication
- Training throughput probes
- BaseModel
- accepted_types
- `[2026-08-27]` Run 3 gated: the rule PASSED and a k=25 follow-up found a self-harm guardrail collapse
- CLOSED OUT (2026-08-23): ACME disabled; and the "all-port VIP" alarm was FALSE
- ERP/RP tune run-01 COMPLETE — 7.36h, gate passed on the axis it was built for
- gpu.py
- stream_chunks
- CatalogField
- accepted_types
- ChunkResult
- type
- `[2026-08-28]` althing v3.0.0 flag day (U9b) — the post office replaced the P2P bus, one-way
- counted_classifier.py
- `[2026-08-26]` ERP run 2 — complete, merged, coherence-gated, and serving as `erp-tune-v2`
- invocation.py
- tts/app.py
- BaseModel
- _cli
- step1_profile.py
- build_calib
@@ -399,23 +399,23 @@
- The 8.6% MFU was an accounting artifact — attention on Ampere kernels
- NVFP4A16 serving pipeline — built, validated, and the MoE landmine it found
- Worldtree b188 + b189 bridge cutover, and the selene metadata that lied
- license
- audio_format_field
- max_length
- quant_with_gen_down.sh
- min
- output_field
- FastAPI
- mime
- `[2026-08-28]` A stale `ALTHING_HANDLE` silently reads ANOTHER agent's inbox and reports it empty
- `[2026-08-28]` The `sec` pen-test seat moved GPU1 → GPU0 and came up — on a circuit that tripped 36h earlier
- fields
- version
- generated_at
- endpoint
- index-tts/app.py
- mime_from_field
- parakeet/app.py
- tts/app.py
- adapter/server.py
- test_keep_and_unkeep_go_through_the_same_name_guard
- streamable
- timestamps_field
- lifecycle
- FastAPI
- license
- version
- FOLLOW-UP 4 (2026-08-23): a downstream WireGuard terminator costs nothing to forward through
## God Nodes (most connected - your core abstractions)
@@ -423,11 +423,11 @@
2. `_touch()` - 30 edges
3. `Recent decisions (archived 2026-08-03 batch)` - 21 edges
4. `build_command()` - 21 edges
5. ``[2026-08-28]` althing v3.0.0 flag day (U9b) — the post office replaced the P2P bus, one-way` - 18 edges
5. `VM 102 — Matrix Synapse Deployment` - 21 edges
6. `JobManager` - 18 edges
7. `[2026-08-23] Anaheim's IPsec tunnel delivers ~25% of a verified 2 Gbps circuit` - 17 edges
8. `_req()` - 17 edges
9. `VM 102 — Matrix Synapse Deployment` - 17 edges
7. ``[2026-08-28]` althing v3.0.0 flag day (U9b) — the post office replaced the P2P bus, one-way` - 18 edges
8. `[2026-08-23] Anaheim's IPsec tunnel delivers ~25% of a verified 2 Gbps circuit` - 17 edges
9. `_req()` - 17 edges
10. `Status + Open Issues` - 16 edges
## Surprising Connections (you probably didn't know these)
@@ -469,7 +469,7 @@ Nodes (29): ana-docker (VM 10.250.50.70), ana-filebot (CT 112 on pfi-pve) — fi
### Community 5 - "properties"
Cohesion: 0.15
Nodes (13): anyOf, default, title, anyOf, default, title, properties, anyOf (+5 more)
Nodes (13): anyOf, default, title, properties, anyOf, default, title, audio_field (+5 more)
### Community 6 - "1. Locked architectural decisions"
Cohesion: 0.07
@@ -480,8 +480,8 @@ Cohesion: 0.08
Nodes (23): Agent Users Always Show as Offline, Agent Virtual Users, Application Service Model, Dependencies, env.sh additions, Formatting, Operational Notes, Overview (+15 more)
### Community 8 - "VM 102 — Matrix Synapse Deployment"
Cohesion: 0.10
Nodes (20): Architecture, Components, Network, Next Steps, Overview, Port Allocation, Security Notes, Sources (+12 more)
Cohesion: 0.06
Nodes (33): Appservice namespace — why `exclusive` is false, Architecture, Components, Conventions, Current state — 2026-09-01, Network, Next Steps, Overview (+25 more)
### Community 9 - "CatalogLifecycle"
Cohesion: 0.11
@@ -796,8 +796,8 @@ Cohesion: 0.25
Nodes (7): API, Deploy, Gotchas, Switching to the 7B variant, VibeVoice, Voices, Why this stack exists
### Community 87 - "properties"
Cohesion: 0.15
Nodes (13): properties, title, type, $ref, title, type, endpoint, model (+5 more)
Cohesion: 0.14
Nodes (14): properties, anyOf, default, $ref, lifecycle, model, reproducibility, response (+6 more)
### Community 88 - "Dia / Dia2"
Cohesion: 0.29
@@ -988,8 +988,8 @@ Cohesion: 0.13
Nodes (15): [2026-08-23] Anaheim's IPsec tunnel delivers ~25% of a verified 2 Gbps circuit, Access note, Admin surfaces closed, Consequences, CORRECTION (2026-08-23): port 80 on the WAN IP is the FortiOS ACME listener, Gotcha: the two UDM vault items have DIFFERENT shapes, Immediate mitigation, no config change, LANDED (2026-08-23): AES-128 on both tunnels; FortiGate public admin closed (+7 more)
### Community 136 - "properties"
Cohesion: 0.25
Nodes (8): properties, default, title, default, section, anyOf, default, title
Cohesion: 0.18
Nodes (11): properties, default, title, title, type, default, name, section (+3 more)
### Community 137 - "category"
Cohesion: 0.50
@@ -1011,9 +1011,9 @@ Nodes (5): 7. Critical Warnings by Model, Gemma 4, Nemotron 3 Super, Qwen3-Coder
Cohesion: 0.22
Nodes (8): Did our abliteration cause it? No — the base did (~83% / ~17%), Resolution — Cold-Fusion abandoned, seat rolled back to `heresy` (2026-08-21), Scripts, The fix, The trigger is TEMPERATURE, not presence_penalty, think-leak — reproducers for the h300 unterminated-`<think>` defect, Using these on any future seat, What actually happens
### Community 142 - "catalog_version"
Cohesion: 0.50
Nodes (4): default, title, type, catalog_version
### Community 142 - "properties"
Cohesion: 0.15
Nodes (13): default, title, type, anyOf, default, title, properties, catalog_version (+5 more)
### Community 143 - "method"
Cohesion: 0.50
@@ -1167,9 +1167,9 @@ Nodes (5): Containerization plan (pending build), Live invocation (source of tru
Cohesion: 0.14
Nodes (12): Any, _base_gen_kwargs(), _pcm16(), Thin FastAPI wrapper exposing OmniVoice (k2-fsa/OmniVoice) for the fleet. Upstr, Validate the voice source and build the MODEL.generate kwargs minus `text`., Synthesize one text span -> (float32 audio [-1,1], audio_seconds)., float32 [-1,1] -> little-endian s16 PCM bytes (24 kHz mono on the wire)., WAV header. data_len=None -> streaming (0xFFFFFFFF sizes, read to EOF); an i (+4 more)
### Community 202 - "stream_chunks"
Cohesion: 0.19
Nodes (17): ClockFn, GenerateFn, ChunkConfig, _ema(), _est_gen_time(), plan_chunk(), protect_first_audio(), Split into sentence units, preserving punctuation. Whitespace-collapsed. (+9 more)
### Community 202 - "phasefinal-web"
Cohesion: 0.29
Nodes (6): Content constraints, Deploy, DNS, phasefinal-web, Routing — and the healthcheck trap that looks like a routing bug, Two steps every fresh export needs
### Community 204 - "heretic2-nvfp4-quant — fast char-rp-reasoning seat (NVFP4 + MTP)"
Cohesion: 0.29
@@ -1607,9 +1607,9 @@ Nodes (4): default, title, type, optional
Cohesion: 0.11
Nodes (19): additionalProperties, properties, required, title, type, CatalogAuditEntry, title, type (+11 more)
### Community 343 - "adapter/server.py"
Cohesion: 0.31
Nodes (6): _get_model(), OpenAI-ish /v1/audio/speech adapter in front of the Zonos Python SDK. Why this, _speaker_embedding(), speech(), JSONResponse, Zonos
### Community 343 - "gpu.py"
Cohesion: 0.42
Nodes (8): _bus_id_by_index(), _cmdline(), gpu_status(), _query_compute_apps(), _query_devices(), GPU status for arbo's device-aware scheduler (GET /gpu-status). arbo steers a l, Return `{devices: [...], tts_on_3090: bool}`. Degrades to an error field on nvid, _run()
### Community 344 - "step"
Cohesion: 0.50
@@ -1691,9 +1691,9 @@ Nodes (4): [2026-08-23] `/mnt/smithy` mounted on ana-ml2 — read-only and SOFT,
Cohesion: 0.33
Nodes (6): `[2026-08-27]` There were TWO run-3c launches, not three — and the phantom third was my reporting, Second-order cost, Step 22 vs step 24, ⚠ THE DURABLE FINDING — an event report with no timestamp is a claim about "now", The evidence, and where it lives, The outage window, pinned to two minutes
### Community 368 - "properties"
### Community 368 - "$ref"
Cohesion: 0.15
Nodes (14): $ref, properties, reproducibility_audit, section_groups, services, items, title, type (+6 more)
Nodes (13): items, title, type, $ref, fields, section_groups, services, items (+5 more)
### Community 369 - "speaches — OpenAI-compatible ASR (faster-whisper) on irv-ml1"
Cohesion: 0.18
@@ -1723,17 +1723,17 @@ Nodes (5): Both IPsec tunnels converge on the same numbers despite different far
Cohesion: 0.25
Nodes (4): [2026-08-24] ana-gw public admin surface closed to zero, ACME listener included, Final state, Port 80 was the FortiOS ACME listener, and I got it wrong first, Retracted in the same pass: the "four all-port VIPs" alarm
### Community 377 - "parakeet/app.py"
Cohesion: 0.29
Nodes (8): OfflineRecognizer, _decode(), _ensure_model_present(), _load_recognizer(), openai_transcriptions(), Thin FastAPI wrapper around sherpa-onnx's OfflineRecognizer for Parakeet-TDT. L, transcribe(), UploadFile
### Community 377 - "Cloudflare state for phasefinal.com"
Cohesion: 0.33
Nodes (5): Cache ruleset, Cloudflare state for phasefinal.com, DNS, Token, Zone settings
### Community 378 - "6. Round-1 aborted; throughput root-caused (measured 2026-08-24 22:00 PDT)"
Cohesion: 0.20
Nodes (10): 6.1 Where the step time goes, 6.1a ⚠ 8.6% MFU was an accounting artifact — real utilisation is 17–20%, 6.1b Backend eligibility, measured — every source claim confirmed, 6.2 ⚠ The attention kernels are Ampere, on a Blackwell card, 6.3 Masking is CORRECT — and padding is what costs, 6.4 The corpus is 29.9% padding — and bucketing is the biggest win available, 6.5 The chunked CE is fine — do not swap it, 6.6 MoE is ~8% — stop optimising it (+2 more)
### Community 379 - "ChunkResult"
Cohesion: 0.15
Nodes (13): ChunkConfig, ChunkResult, _chunk_config(), _log_chunk(), Streaming /tts request — chatterbox-fast-compatible wire protocol., Streaming: chunked 24 kHz mono PCM (or open-ended WAV) for live consumers., tts(), TTSStreamRequest (+5 more)
### Community 379 - "pfi-gx10 — ASUS Ascent GX10 (NVIDIA GB10)"
Cohesion: 0.33
Nodes (5): Access, Headless conversion, pfi-gx10 — ASUS Ascent GX10 (NVIDIA GB10), Relevance to Flash-Next, ⚠ The address in `ssh-target` is TEMPORARY
### Community 380 - "[2026-08-24] Scriberr transcription deployed on ana-ml2, GPU1"
Cohesion: 0.50
@@ -1751,9 +1751,9 @@ Nodes (13): 1. ⚠ QLoRA IS NOT AVAILABLE ON THIS ARCHITECTURE, 2. What the run
Cohesion: 0.50
Nodes (4): Design rules worth preserving when you adapt these, Raw evidence, The probes, Training throughput probes
### Community 384 - "BaseModel"
Cohesion: 0.19
Nodes (12): ACEStepInput, ACEStepOutput, generate_audio(), initialize_pipeline(), Generate music; respond with the audio bytes inline. Pre-2026-05-11 this re, ACEStepPipeline, SpeechRequest, BaseModel (+4 more)
### Community 384 - "accepted_types"
Cohesion: 0.50
Nodes (4): anyOf, default, title, accepted_types
### Community 385 - "`[2026-08-27]` Run 3 gated: the rule PASSED and a k=25 follow-up found a self-harm guardrail collapse"
Cohesion: 0.40
@@ -1767,17 +1767,17 @@ Nodes (4): ACME disabled — the WAN IP now exposes nothing, CLOSED OUT (2026-08
Cohesion: 0.17
Nodes (10): ERP/RP tune run-01 COMPLETE — 7.36h, gate passed on the axis it was built for, lora_B gate — PASSED, twice, The confound I built and he caught, The gate — brokkr-smithy-dev, The noise-floor near-miss — the methodology lesson, The run, ⚠⚠ But it is the WRONG AXIS — brokkr's catch, and it is the better one, Refusal retention — the axis the gate did not have, and the axis I measured wrong (+2 more)
### Community 388 - "gpu.py"
Cohesion: 0.42
Nodes (8): _bus_id_by_index(), _cmdline(), gpu_status(), _query_compute_apps(), _query_devices(), GPU status for arbo's device-aware scheduler (GET /gpu-status). arbo steers a l, Return `{devices: [...], tts_on_3090: bool}`. Degrades to an error field on nvid, _run()
### Community 388 - "stream_chunks"
Cohesion: 0.19
Nodes (17): ClockFn, GenerateFn, ChunkConfig, _ema(), _est_gen_time(), plan_chunk(), protect_first_audio(), Split into sentence units, preserving punctuation. Whitespace-collapsed. (+9 more)
### Community 389 - "CatalogField"
Cohesion: 0.33
Nodes (6): additionalProperties, description, required, title, type, CatalogField
### Community 390 - "accepted_types"
Cohesion: 0.50
Nodes (4): anyOf, default, title, accepted_types
### Community 390 - "ChunkResult"
Cohesion: 0.15
Nodes (13): ChunkConfig, ChunkResult, _chunk_config(), _log_chunk(), Streaming /tts request — chatterbox-fast-compatible wire protocol., Streaming: chunked 24 kHz mono PCM (or open-ended WAV) for live consumers., tts(), TTSStreamRequest (+5 more)
### Community 391 - "type"
Cohesion: 0.50
@@ -1799,9 +1799,9 @@ Nodes (7): `[2026-08-26]` ERP run 2 — complete, merged, coherence-gated, and s
Cohesion: 0.19
Nodes (14): Path, test_published_relative_path(), test_published_relative_path_explicit_train_id_wins(), ValueError, InvalidTrainRequest, published_relative_path(), Fixed-invocation command builder — the enforcement point for INV-T7. arbo hands, The ComfyUI-relative loras path for a succeeded LoRA (Phase 2 publish step). (+6 more)
### Community 396 - "tts/app.py"
Cohesion: 0.30
Nodes (7): _build_prompt(), _cap(), _decode(), _encode_ref(), tts(), tts_stream(), TTSReq
### Community 396 - "BaseModel"
Cohesion: 0.19
Nodes (12): ACEStepInput, ACEStepOutput, generate_audio(), initialize_pipeline(), Generate music; respond with the audio bytes inline. Pre-2026-05-11 this re, ACEStepPipeline, SpeechRequest, BaseModel (+4 more)
### Community 398 - "_cli"
Cohesion: 0.25
@@ -1827,9 +1827,9 @@ Nodes (6): Four silent defects the dry run found, ⚠ MERGED WEIGHTS ARE MANDATO
Cohesion: 0.33
Nodes (5): #411 — the debug-room failure, diagnosed twice and wrong both times first, b188 — matrix.yaml pre-sync (#406/#409/#410 closed), b189 — #407 bridge extracted to its own repo (#404 umbrella closed), selene-1-mini-8b — a config that lied about what answers, Worldtree b188 + b189 bridge cutover, and the selene metadata that lied
### Community 408 - "license"
Cohesion: 0.67
Nodes (3): title, type, license
### Community 408 - "audio_format_field"
Cohesion: 0.50
Nodes (4): anyOf, default, title, audio_format_field
### Community 409 - "max_length"
Cohesion: 0.50
@@ -1843,9 +1843,9 @@ Nodes (4): anyOf, default, title, min
Cohesion: 0.50
Nodes (4): anyOf, default, title, output_field
### Community 413 - "FastAPI"
Cohesion: 0.47
Nodes (4): FastAPI, lifespan(), sfx(), SfxRequest
### Community 413 - "mime"
Cohesion: 0.50
Nodes (4): anyOf, default, title, mime
### Community 414 - "`[2026-08-28]` A stale `ALTHING_HANDLE` silently reads ANOTHER agent's inbox and reports it empty"
Cohesion: 0.40
@@ -1855,44 +1855,44 @@ Nodes (4): `[2026-08-28]` A stale `ALTHING_HANDLE` silently reads ANOTHER agent'
Cohesion: 0.40
Nodes (4): `[2026-08-28]` The `sec` pen-test seat moved GPU1 → GPU0 and came up — on a circuit that tripped 36h earlier, Deploy gotchas worth keeping, Landed state, Why it had to move
### Community 416 - "fields"
Cohesion: 0.50
Nodes (4): items, title, type, fields
### Community 417 - "version"
### Community 416 - "endpoint"
Cohesion: 0.67
Nodes (3): version, title, type
Nodes (3): title, type, endpoint
### Community 418 - "generated_at"
Cohesion: 0.50
Nodes (4): anyOf, default, title, generated_at
### Community 419 - "index-tts/app.py"
### Community 417 - "index-tts/app.py"
Cohesion: 0.22
Nodes (8): _chunk_to_pcm_bytes(), IndexTTS-2 — minimal FastAPI wrapper. Upstream (https://github.com/index-tts/in, 44-byte RIFF/WAVE/PCM header with placeholder data length so the payload can, Normalize whatever IndexTTS-2 yields (torch tensor, numpy array, int16 or fl, _resolve(), SpeechRequest, synthesize(), _wav_header()
### Community 420 - "mime_from_field"
Cohesion: 0.50
Nodes (4): anyOf, default, title, mime_from_field
### Community 418 - "parakeet/app.py"
Cohesion: 0.29
Nodes (8): OfflineRecognizer, _decode(), _ensure_model_present(), _load_recognizer(), openai_transcriptions(), Thin FastAPI wrapper around sherpa-onnx's OfflineRecognizer for Parakeet-TDT. L, transcribe(), UploadFile
### Community 422 - "streamable"
Cohesion: 0.50
Nodes (4): streamable, default, title, type
### Community 419 - "tts/app.py"
Cohesion: 0.30
Nodes (7): _build_prompt(), _cap(), _decode(), _encode_ref(), tts(), tts_stream(), TTSReq
### Community 423 - "timestamps_field"
Cohesion: 0.50
Nodes (4): timestamps_field, anyOf, default, title
### Community 420 - "adapter/server.py"
Cohesion: 0.31
Nodes (6): _get_model(), OpenAI-ish /v1/audio/speech adapter in front of the Zonos Python SDK. Why this, _speaker_embedding(), speech(), JSONResponse, Zonos
### Community 424 - "lifecycle"
### Community 422 - "FastAPI"
Cohesion: 0.47
Nodes (4): FastAPI, lifespan(), sfx(), SfxRequest
### Community 423 - "license"
Cohesion: 0.67
Nodes (3): anyOf, default, lifecycle
Nodes (3): title, type, license
### Community 424 - "version"
Cohesion: 0.67
Nodes (3): version, title, type
### Community 425 - "FOLLOW-UP 4 (2026-08-23): a downstream WireGuard terminator costs nothing to forward through"
Cohesion: 0.67
Nodes (3): Design consequences of terminating downstream — the parts that need decisions, FOLLOW-UP 4 (2026-08-23): a downstream WireGuard terminator costs nothing to forward through, Standing recommendation
## Knowledge Gaps
- **2314 isolated node(s):** `Recent decisions (archived)`, `What landed`, `Load-bearing lessons (the whole point of this file)`, `Dead ends (tried + abandoned)`, `granite retired + gateway repoint` (+2309 more)
- **2337 isolated node(s):** `⚠ The password is NOT in this copy`, `/_synapse/admin is LAN-only`, `Upgrades`, `Active migration — docker.io 20.10 → docker-ce 29.x`, `Architecture decisions (durable)` (+2332 more)
These have ≤1 connection - possible missing edges or undocumented components.
- **41 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
@@ -1901,10 +1901,10 @@ _Questions this graph is uniquely positioned to answer:_
- **Why does `Tensor` connect `convert_hf_to_native.py` to `adapter/server.py`?**
_High betweenness centrality (0.004) - this node is a cross-community bridge._
- **Why does `properties` connect `properties` to `fields`, `version`, `streamable`, `lifecycle`, `category`, `content_type`, `description`, `$defs`, `method`, `CatalogModel`, `properties`, `properties`, `license`, `status`, `host`, `license_warning`, `estimated_latency`?**
- **Why does `_speaker_embedding()` connect `adapter/server.py` to `convert_hf_to_native.py`?**
_High betweenness centrality (0.004) - this node is a cross-community bridge._
- **What connects `Recent decisions (archived)`, `What landed`, `Load-bearing lessons (the whole point of this file)` to the rest of the system?**
_2314 weakly-connected nodes found - possible documentation gaps or missing edges._
- **What connects `⚠ The password is NOT in this copy`, `/_synapse/admin is LAN-only`, `Upgrades` to the rest of the system?**
_2337 weakly-connected nodes found - possible documentation gaps or missing edges._
- **Should `Status + Open Issues` be split into smaller, more focused modules?**
_Cohesion score 0.05555555555555555 - nodes in this community are weakly interconnected._
- **Should `CatalogReproducibility` be split into smaller, more focused modules?**
@@ -0,0 +1,75 @@
# `[2026-09-01]` Ada migration settled on `zfs send` — and branch (b) was never available
The Ada box (ComfyUI's new home, sm_89, x86-64) lands at **NH3**. irv-ml1 is in **Irvine**.
comfy-dev asked whether `/storetank` rides along or the stack is rebuilt from source.
## The answer: (a) `zfs send`. Measured, not derived.
NH3 -> irv-ml1 11-26 ms, 0% loss
throughput 99.0 MB/s (real 800 MB transfer over the WireGuard tunnel)
payload 1.38 TB -> ~3.9 hours
`zfs send` is **incremental**: snapshot now, ship the base over ~4 hours while irv-ml1 keeps
serving, then a small delta at cutover. Near-zero service interruption.
## Why (b) — physically moving the disks — was never on the table
`/storetank` is a two-disk **mirror** of Crucial MX500 2TB SATA SSDs, so it is genuinely
portable hardware. That fact is true and **irrelevant**.
**irv-ml1 is NOT being decommissioned.** It runs arbo, comfyui, tts-gateway, dots-tts,
voice-studio, waterland-studio, yt-voice-clipper and dockge — and **comfyui bind-mounts
`/storetank/arbo/models`**. Pulling those disks does not inconvenience a source box that no
longer needs them; it guts a live one.
⚠ **The lesson:** infra-ops asked whether the DATA could move and never asked whether the
SOURCE still needed it. One `docker ps` away, at any point. Recommended (b) twice on the
strength of a true-but-irrelevant fact. See [[2026-09-01-wrong-object-measurement]].
## (c) rebuild-from-source: rejected on reproducibility, not time
~6 hours of re-fetch at ~65 MB/s. The real objection is that **two of comfy-dev's pins are
already paywalled** (Big Love went permanent-paid; Moody Krea 2 Mix's newest releases are
gated while their pin is free). A rebuild today would **not reproduce today's stack**. A
fallback that provably cannot restore what it exists to restore is not a fallback.
## ⚠ The two-boxes confusion — do not repeat it
There are **TWO new machines** and infra-ops collapsed them into one:
| | |
|---|---|
| **Ada box** | ComfyUI's target. sm_89, x86-64. Lands at NH3. Not yet arrived. |
| **ASUS Ascent GX10** | Local inference + run 3c. **GB10, sm_121, aarch64.** On the operator's desk. → [[2026-09-01-pfi-gx10-onboarding]] |
infra-ops found no sm_89 part in the current inventory, saw the Ascent, and inferred the
Ascent was the incoming box — then sent comfy-dev down an aarch64 re-platform investigation
that was entirely void. **comfy-dev's original premise was correct throughout.**
Consequences of the retraction, all restored to their original state:
- Their single-arch container image, 18 GB x86-64 venv and torch pin are **fine**.
- Their nvfp4 pin is **RIGHT, not wrong** — sm_89 does not do native nvfp4 (that is
Blackwell), so pinning away from the 7.74 GB nvfp4 build to the 12.84 GB int8 build was
correct for the hardware they are actually getting.
**Not wasted:** comfy-dev's aarch64 runtime research (ecarmen16/SparkyUI — CUDA 13.0.2,
torch 2.9.1+cu130 ARM64, SageAttention compiled with `TORCH_CUDA_ARCH_LIST="12.1"`, built on
the box, ~10 min; verified from source by infra-ops) applies to the **GX10** if anything
ComfyUI-shaped ever runs there.
## Their distinction, worth keeping
> **The weights port. The runtime does not.**
safetensors are architecture-neutral and travel anywhere. Container image, torch build, venv
and attention kernels do not. "Migrate the model store" and "migrate ComfyUI" are different
jobs, and the 1.38 TB transfer is the easy half.
## Open
- **Cutover window** — operator's, not yet set.
- **comfy-dev's ~112 GB batch** — unheld by infra-ops, but they are deliberately NOT pulling
until the operator approves putting that much onto his infrastructure. Their call, correct
instinct. Manifest pinned and staged (`29324e9`).
Thread: `01M1EYBSYA4QRYK54PX0K1S8CS`.
@@ -0,0 +1,102 @@
# `[2026-09-01]` Matrix: 39-release Synapse upgrade, appservice namespace opened, admin API closed
## The upgrade
**Synapse v1.120.0 → v1.159.0** (21 months, 39 releases) and **Element-web v1.11.80 →
v1.12.27**. Schema migrations applied cleanly through schema 94. Postgres deliberately left
at 16 — changing two stateful things at once destroys failure attribution.
⚠ **Schema migrations are ONE-WAY.** v1.120 cannot start against a v1.159 database. Rollback
is restore-from-dump, not revert-the-tag. Verified pre-upgrade dump (739 TOC entries from a
33 MB database) plus all four config files at
`/opt/docker/backups/synapse-preupgrade-20260901T174625Z/`.
Reviewed every upgrade note in the range; nothing applicable bit us (PG 11/12/13 drops — we
are on 16; MSC3861/MAS items — no MAS; s3-storage-provider and worker media quarantine — not
in use).
## The appservice namespace — `exclusive: true` → `false`
The `aipa-bridge` registration claimed `@[a-z][a-z0-9_-]*:matrix.phasefinal.com` **exclusively**
— every localpart on the server. 14 of 14 accounts fell inside it; 13 were appservice-owned.
**`exclusive` governs who ELSE may act, not what the appservice may do.** On a homeserver
with registration disabled, one admin and no competing actor, it bought anti-squatting
protection against a threat that cannot occur, while locking out every other means of
account creation — admin registration returned `M_EXCLUSIVE` with no explanation.
⚠ **Do NOT narrow the users regex to a prefix** — all 13 accounts fall inside it and would be
orphaned. ⚠ **Do NOT rename the `id`** — Synapse keys account ownership on `aipa-bridge` in
the `users` table. The FILE may be renamed; the id may not.
The narrow **aliases** namespace (`#aipa-debug-*`) was left exclusive — specific, costs nothing.
⚠ The registration is named `aipa`, but the service behind it is **`wt-matrix-bridge`**, the
Worldtree PERSONAL instance on corviduo-dev `10.250.50.152:8010`. AIPA is a dead project name
on a live service, and it is why infra-ops mis-routed a provisioning request to worldtree-dev.
**Operator ruling: worldtree-dev writes the bridge code; infra-ops OPERATES this instance and
has full authority over it.**
## `/_synapse/admin` closed to the internet
Synapse mounts its admin API on the same vhost as the client API, so publishing
`matrix.phasefinal.com` published the admin surface — it **answered 200 from the open
internet**. A higher-priority router (explicit priority 100) now scopes it behind an
`ipallowlist`.
Verified from a **genuinely external vantage** — the NH3 residential egress proxy, because
testing from a fleet host sits inside the allow-list and proves nothing: admin **403**,
client API **200**, Element unaffected.
⚠ **The `10.0.0.0/8` entry matches NOTHING and that is expected.** The hostname resolves
publicly, so fleet hosts hairpin out their own WAN — a request from nh3-dev measured as
`70.230.226.88`. The rule is effectively **deny-all through Traefik**, which is intended:
admin work goes via `docker exec synapse` against `localhost:8008` and never touches Traefik.
Allow-listing the sites' WAN addresses was **rejected** — dynamic, and a stale entry either
locks us out or hands admin to whoever inherits the address.
## Conventions ratified (operator, 2026-09-01)
#<agent>-<purpose>:matrix.phasefinal.com
Mirrors the existing `@<agent>:` user-ID convention. Proposed by ledger-dev. Rationale is
**"so the room IDENTITY carries the tier"** — deliberately NOT "so the push payload carries
the room name", which is true only for clients without a notification service extension.
Pre-existing rooms are not renamed ("The High Seat", `!NiVVoMsyoHCBRPrrrn`).
## Push reality — measured, and it inverts the obvious reading
The registered pusher (`@vhoang`, Element X iOS) uses **`"format": "event_id_only"`** via
matrix.org's sygnal. That payload carries event_id, room_id and counts — **no room name, no
sender, no content**. It still produces a useful notification because `mutable-content: 1`
means Element X runs a **Notification Service Extension**: iOS wakes it with the near-empty
payload and it **fetches the event and renders the notification on the device**.
1. The tier-in-room-identity scheme works — but **via the client fetch**, not the payload.
`m.room.name` must be set at creation; the ALIAS is not what reaches the phone. Synapse
sends `ctx["name"]` (the `m.room.name` state event) and omits the key entirely if unset.
2. **`push: include_content: false` is irrelevant for clients with an NSE.** It bites clients
without one.
3. ⚠ **Server-invisible failure mode:** if the phone cannot reach the homeserver at wake time
the fetch fails and iOS shows the bare word "Notification". **Synapse records
`last_success` and sees a delivered push.**
Self-hosted sygnal **considered and rejected** — sygnal is a relay to FCM/APNs, not a
replacement, so it removes matrix.org and nothing else; and with `event_id_only` the path
already carries nothing worth protecting.
## QR sign-in — requires MAS, deferred
MSC4108 hard-requires `matrix_authentication_service`; Synapse refuses to start otherwise.
MSC4388 enables independently but is only the rendezvous **channel**, not a login flow.
Deferred: MAS is a service, a database and a migration of every account off built-in auth,
and the v1.139.0 note warns `/register` from **old appservice implementations may break under
MAS** — precisely the bridge owning 13 of 15 accounts.
## Shared-secret registration gotcha
`HMAC-SHA1(secret, nonce \0 user \0 password \0 "notadmin")` — the null **separates**, it does
not **terminate**. A trailing `\x00` yields `HMAC incorrect`. Run inside the container against
`localhost:8008`; port 8008 is not published to the host.
Full doc: `docs/pfi/vm-102-matrix-synapse.md` (`931bac8`, `73866f6`).
@@ -0,0 +1,79 @@
# `[2026-09-01]` pfi-gx10 (ASUS Ascent GX10) onboarded headless — and it is the intended new home for run 3c
ASUS Ascent GX10 arrived and was registered, converted to headless, and given a rack-move
playbook. **It was NOT racked** — the operator ran out of day. It is still on his desk, on
Wi-Fi, on a temporary DHCP lease.
| | |
|---|---|
| GPU | **NVIDIA GB10**, driver 580.173.02, **compute capability 12.1 (`sm_121`)** |
| CPU | 20 cores, **aarch64** |
| Memory | **121 GB UNIFIED** — CPU and GPU share it; not 121 GB *plus* VRAM |
| Storage | 916 GB NVMe, 6% used |
| OS | Ubuntu 24.04.4, kernel 6.17.0-1031-nvidia |
| Access | `infra-ops` NOPASSWD sudo (operator-bootstrapped). `lkraven` has key auth but needs a password to escalate — **automation must connect as `infra-ops`**. |
## Purpose (operator, 2026-09-01)
Local inference experiments **and** the failed training — run 3c. That is the whole point:
run 3c died on ana-ml2 when a **kilowatt-class** training box tripped a breaker
(2× 300 W GPUs + dual EPYC 9254). The GX10 is a ~240 W appliance, roughly a fifth of the
draw, on a different site's circuits.
**The memory arithmetic favours it strongly.** Run 3c is a LoRA (r=64, batch 2 × accum 8,
gradient checkpointing) over `gemma4-26b-a4b-it-bf16` — **49 GB of base weights**, working
set roughly 55–65 GB. The `gemma4-charrp` compose warns in capitals that "48.10 GiB of BF16
weights CANNOT be served here" on ana-ml2's shared GPU0. 121 GB unified turns that
constraint into a non-issue.
## What has NOT been established — do not assume any of it
1. **The box is bare.** No torch, no nvcc, no CUDA stack. Docker 29.2.1 and 822 GB free.
2. **aarch64 dependency risk.** torch 2.9.1+cu130 ARM64 wheels exist; transformers/TRL/PEFT
are pure Python. Compiled deps — flash-attn, bitsandbytes, liger, xformers — are open
questions per-arch.
3. **Triton has NO sm_121 support** (established independently by comfy-dev the same day).
Anything reaching for `torch.compile` or Triton-backed kernels is closed on this silicon.
4. **Throughput is unmeasured.** Run 3c ran 16.45 s/it on a 300 W RTX PRO 6000. GB10 will be
slower; how much decides whether 604 steps is an overnight run or two days.
**Measure this before porting anything** — the recommended first move is a probe: install
ARM64 torch, load the base, run ten steps, report s/it.
5. **Model transfer:** 49 GB from ana-ml2 over the Anaheim↔NH3 link at a **measured
32.4 MB/s** — about 25 minutes. Note that is a third of the irv-ml1↔NH3 link's 99 MB/s.
## The headless conversion, and the lesson inside it
`playbooks/gx10-headless.yaml` (`1b596c8`): multi-user.target, gnome-remote-desktop stopped,
sleep/suspend/hibernate **masked**, logind ignores lid and idle, sshd keepalives, hostname
corrected `gx10-a745` → `pfi-gx10`.
⚠ **`gdm` is a STATIC unit on Ubuntu** — pulled in by `display-manager.service`, never
"enabled". The first version guarded on `is-enabled | grep enabled`, which always skips, and
**the verify tested the same wrong property and passed**. Six green verifies having not
stopped the display manager. Both now test `is-active`. This is an instance of
[[2026-09-01-wrong-object-measurement]].
The playbook refuses to stop GDM while a seat session is held (`--var force_dm_stop=true` to
override) — automation should not yank a display out from under someone at the machine.
⚠ **elway's `--sudo` applies only to ad-hoc `--shell`/`--upload`.** Playbook steps run as the
connecting user and must carry their own `sudo`.
## The rack move, written but not run
`playbooks/gx10-rack-network.yaml` (`a0c5fc6`) — target settled: **`nh3-servers` VLAN 50,
static `10.100.50.60`**, clear of `.40`/`.42`/`.50`/`.90` and below the `.150` DHCP pool.
Requires nothing from the operator beyond racking it. The wired NIC has its own MAC
(`30:c5:99:3d:a7:45`, distinct from Wi-Fi `50:bb:b5:a2:00:a8`), so the post-move address AND
switch port are discoverable from the UDM rather than relayed.
**Design property worth preserving:** the playbook never leaves itself one path back. Wi-Fi
stays up while the wired interface is configured beside it; Wi-Fi teardown is explicitly a
separate later change. `netplan try`'s auto-rollback needs a TTY that elway cannot provide,
so a live Wi-Fi link is what substitutes for it. Two preconditions are asserted as steps
rather than assumed: carrier must be 1, and the MAC must match (interface names renumber
across kernels; MACs do not).
⚠ **The GX10 is NOT the Ada box.** Two separate machines — see
[[2026-09-01-ada-migration-branch-a]].
@@ -0,0 +1,42 @@
# `[2026-09-01]` A named failure class: a correct check aimed at the wrong object
Six instances surfaced across three sessions in a single day, independently, in unrelated
domains. It has a distinguishing property that makes it worth naming separately from
"a bad measurement":
> **Re-running the same check cannot catch it, because the check is correct and the object
> is wrong.** The only move that breaks it is asking what the artifact *is* before trusting
> any metric computed over it.
## The instances
| where | the metric | the artifact nobody opened |
|---|---|---|
| comfy-dev, civitai | a "~25 KB/s throttle" | a **9,685-byte login page** returned on failed auth |
| comfy-dev, render harness | reported "completed" | a 4.7 KB **all-black PNG** |
| comfy-dev, audio | RMS said healthy | degenerate audio where the **loudness WAS the noise** |
| ledger-dev, capability 5 | healthy ping, healthy container, healthy delivery | a **requirement unmet** — a mailbox, not an away-channel |
| infra-ops, gx10 headless | six green verifies | `is-enabled` on a **static unit**; gdm still running |
| infra-ops, synapse admin | an `ipallowlist` comment promising fleet access | fleet traffic **hairpins out the WAN**; 10.0.0.0/8 matches nothing |
## Related lessons banked the same day
- **A caveat plus propagation is decoration.** infra-ops flagged a sample-size problem AND
escalated the claim in the same message. If a number needs re-measuring before it can be
repeated, hold the escalation until it has been. The caveat made the uncareful thing look
examined.
- **Verify with a negative control.** A 200 means nothing without a 401 beside it. Used on
the pewpewstudio key mint; used on the `/_synapse/admin` lock (tested from a genuinely
external vantage via the NH3 residential egress proxy, because testing from a fleet host
sits inside the allow-list and proves nothing).
- **Reasoning from what is VISIBLE to what EXISTS.** infra-ops found no `sm_89` part in the
inventory, saw one new machine, and collapsed "the box I can see" into "the box that is
coming". They were two different machines. Wrote "I cannot resolve it and will not guess",
then built three messages on the guess.
## Disposition
Recommended for a row in `docs/pfi/training-throughput-playbook.md` §4 (the durable home for
"why a run LIES about itself"), with attribution to comfy-dev and ledger-dev.
**NOT YET WRITTEN — awaiting operator.** Tracking surface: this file, plus althing threads
`01M1EYBSYA4QRYK54PX0K1S8CS` (comfy-dev) and `01M1EXVPZT66SSGCRZH3W624R7` (ledger-dev).
+42 -10
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-08-28_
_Last updated: 2026-09-01_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -108,19 +108,51 @@ no longer deployed sidecars here. See Recent decisions.)
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
## Current state / in-flight
_As of 2026-08-28 16:20 PDT — **althing v3 is live fleet-wide at 3.1.1 with the post office on nh3-docker. `sec` is serving on ana-ml2 GPU0. Nothing is training.**_
_As of 2026-09-01 — **the GX10 is on the operator's desk, NOT racked. Standing it up as an inference AND training box is the next action.** Nothing is training._
- **⏸ RUN 3c STILL HELD — power capacity, unchanged.** lr `2e-4 -> 1e-5`, corpus BYTE-IDENTICAL, config `/tank/erp-tune/run-03c.json` validated, relaunch is one command. **Do not relaunch until the power triage lands.** ⚠ `sec` now occupies GPU0 (~51 GB), so a 3c relaunch needs GPU0 freed OR accepts three-way contention. Exactly TWO 3c launches, only one died. → `persistent-memory.d/2026-08-27-run3c-launch-count-reconstruction.md`
- **🟢 `sec` (M.O.G.-SEC-27B) IS UP on ana-ml2 GPU0 :8019**, operator-directed. 51,532 MiB, idle 16 W. ⚠ Re-arms the two-GPU load condition that tripped the rack breaker; the risk is concurrent load, not idle. → `persistent-memory.d/2026-08-28-sec-seat-gpu0.md`
- **🟢 althing v3.1.1 on both heralds; post office on nh3-docker `http://10.100.50.40:8390`.** Container untouched since the move. All 73 handles seeded; 5 agents push-reachable (infra-ops, forseti, and the four smithy peers via pane routes). → `persistent-memory.d/2026-08-28-althing-v3-cutover.md`
- **🔴 `/tank` DEGRADED on ana-ml2** — 7 physical NVMe where the pool expects 8, raidz2, one parity spent, no data errors. **Operator replacing it.** → `persistent-memory.d/2026-08-27-anaheim-breaker-and-onboot-gap.md`
- **⏸ WEEKEND: power triage, "probably shut down some seats"** (operator). Sheddable on ana-ml2: `vllm-gen` 46 GB, `sec` 51 GB, embed 9.8, coder 8.4, rerank-a3 3.5, reward 2.1, scriberr.
- **⏸ Owed to brokkr-smithy-dev when there is a card again:** nothing blocking. They hold the dose curve and the entanglement finding.
- **⏸ Awaiting pewpew-dev:** which other boxes/CI runners resolve `pypotrace` and need the potrace headers. nh3-dev is done; `playbooks/install-potrace-headers.yaml` makes each additional box one command.
- ⚠ **`/mnt/smithy` will be MISSING after every ana-ml2 reboot** — manual by design. → `persistent-memory.d/2026-08-23-smithy-mount-ana-ml2.md`
- **▶ NEXT: pfi-gx10 as an inference + training box.** The operator forgot to rack it; it is
still on his desk on Wi-Fi at `10.100.10.226` (temp DHCP). Already headless and registered.
**The box is BARE — no torch, no nvcc, no CUDA stack.** Purpose is local inference *and*
**run 3c**, whose 49 GB bf16 base fits 121 GB unified with room where ana-ml2 was tight.
**First move is a throughput probe, not a port**: ARM64 torch, load the base, ten steps,
report s/it — that decides whether 604 steps is an overnight run or unusable. Racking is
one command afterwards (`playbooks/gx10-rack-network.yaml`, VLAN 50, static `10.100.50.60`).
⚠ Triton has no sm_121 support; compiled deps are per-arch unknowns.
→ `persistent-memory.d/2026-09-01-pfi-gx10-onboarding.md`
- **⏸ RUN 3c STILL HELD — but the plan has changed.** Config `/tank/erp-tune/run-03c.json`
validated, relaunch is one command on ana-ml2. **It is now intended to move to the GX10
instead**, which is the power answer rather than a power triage. Do not relaunch on ana-ml2
without deciding that first. Exactly TWO 3c launches, only one died.
→ `persistent-memory.d/2026-08-27-run3c-launch-count-reconstruction.md`
- **⏸ ADA MIGRATION — strategy settled, cutover window is the operator's.** Branch (a)
`zfs send`, ~3.9 h for 1.38 TB at a measured 99 MB/s, incremental so irv-ml1 keeps serving.
Branch (b) was never available. **comfy-dev is holding a ~112 GB pull awaiting the
operator's go**, not infra-ops'. → `persistent-memory.d/2026-09-01-ada-migration-branch-a.md`
- **⏸ Worldtree `route_not_found` awaiting the operator's DEPLOY PUSH.** Approved and landed
by worldtree-dev as `ece0250c` (wire 2.5.0→2.6.0). main auto-deploys demo. Until it ships,
"trust the HTTP status, not `error_code`" still applies on running instances.
- **🟢 Matrix upgraded and hardened.** Synapse v1.159.0, Element v1.12.27, `/_synapse/admin`
LAN-only, appservice namespace opened, alias convention ratified. Miranda provisioned; her
summons reached the operator's watch. → `persistent-memory.d/2026-09-01-matrix-upgrade-and-hardening.md`
- **🟢 phasefinal.com LIVE** behind Cloudflare with edge caching + Always Online; apex 301s to
www. `inquiry@` alias added by the operator. Stack `stacks/phasefinal-web/`.
- **🟢 char-rp restored** on ana-ml2 **GPU1** (was dead 7 days on a quant-flag/checkpoint
mismatch crash-loop). ⚠ **GPU1 is now at 96,384 / 97,887 MiB — ~1.5 GB free, six tenants.**
GPU0's ~44 GB free is a DELIBERATE scratch reserve, not headroom to reclaim.
- **⏸ Deferred, no blocker:** convert the live synapse compose to read `POSTGRES_PASSWORD`
from a `.env` — until then `stacks/synapse/` and the live file have DIVERGED and
`deploy-stack.sh` must not be used on it (README says so).
- ⚠ **`/mnt/smithy` will be MISSING after every ana-ml2 reboot** — manual by design.
→ `persistent-memory.d/2026-08-23-smithy-mount-ana-ml2.md`
## Recent decisions
- `[2026-09-01]` **pfi-gx10 onboarded headless — and it is the intended new home for run 3c, which died on a tripped breaker.** GB10/sm_121/aarch64, 121 GB unified. NOT racked yet. Bare of any CUDA stack; probe throughput before porting. → `persistent-memory.d/2026-09-01-pfi-gx10-onboarding.md`
- `[2026-09-01]` **Ada migration is `zfs send` (branch a) — branch (b) was never available because irv-ml1 keeps its eight services.** 99 MB/s measured; ~3.9 h. Also records the two-boxes confusion: the Ada box and the GX10 are DIFFERENT machines. → `persistent-memory.d/2026-09-01-ada-migration-branch-a.md`
- `[2026-09-01]` **Matrix: Synapse 1.120→1.159, appservice namespace opened, `/_synapse/admin` closed to the internet, alias convention ratified.** Schema migrations are one-way; push is `event_id_only` and assembled on-device. → `persistent-memory.d/2026-09-01-matrix-upgrade-and-hardening.md`
- `[2026-09-01]` **A named failure class: a correct check aimed at the wrong object.** Six instances in one day across three sessions; re-running the same check cannot catch it. **Recommended for `docs/pfi/training-throughput-playbook.md` §4 — NOT YET WRITTEN, awaiting operator.** → `persistent-memory.d/2026-09-01-wrong-object-measurement.md`
- `[2026-09-01]` **Ops boundary ruled by the operator: worldtree-dev writes the bridge code; infra-ops OPERATES the Worldtree/Matrix instances and may change them.** Corrects a mis-route where infra-ops asked worldtree-dev to provision an account on a box it does not run. Tracked at `931bac8` + althing `01M1F4PK796EDGDCBKZ9W3JC0S`.
- `[2026-09-01]` **Idle VRAM on this fleet is a RESERVED scratch pool, not waste.** Operator declined raising `vllm-mog-sec` from `gpu-memory-utilization 0.52`: single-user dev fleet, KV headroom nobody will consume is worth less than room for ephemeral models and small training runs. vLLM's "fully utilize gpu memory" startup hint does NOT apply here. Tracked in auto-memory `feedback_idle_vram_is_reserved_not_waste`.
- `[2026-08-28]` **althing v3 flag day (U9b) executed, then six releases to 3.1.1 in one afternoon — and the post office MOVED to nh3-docker.** Every v2 command deleted; 73 handles seeded and verified by set difference; 5,043 orphaned wake FIFOs deleted (v2 named them per-session+PID, v3 per-handle). Image now registry-pulled, digest-pinned, under the `claude-bot` namespace. → `persistent-memory.d/2026-08-28-althing-v3-cutover.md`
- `[2026-08-28]` **A stale `ALTHING_HANDLE` silently reads another agent's inbox and reports it empty — a SECOND route into the failure v3 exists to prevent.** Outbound mis-signing sometimes gets caught; inbound never does. Shipped as a 3.1.1 warning. ⚠ My `session_handles.json` grounding was wrong (v2 artifact, v3 never opens it) and the same stale source had survived inside my statusline rewrite. → `persistent-memory.d/2026-08-28-handle-resolution-wrong-inbox.md`
- `[2026-08-28]` **nh3-dev's three OOM events attribute to CLAUDE CODE, and the "no kernel evidence" was a permissions artifact.** journald was persistent all along; `journalctl` silently shows only your own messages outside `adm`. Single CC sessions measured 5.4-18.4 GB, so 27 GB is 3-4 long-lived sessions. sysstat + atop now instrument the ramp. → `persistent-memory.d/2026-08-28-nh3-dev-oom-attribution.md`