ce09ac4fa698cca6cd948c4becf7cf3df4e62d37
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ce09ac4fa6 |
test(gen-seat): add a real vision battery — orcarouter scores 7/8
surface_test.py's vision check is one image and one word. It proves the tower loads; it does not prove the tower works. This battery uses generated images with known ground truth so every answer is objectively gradeable. Against orcarouter NVFP4-mixed on the `gen` alias: T1 OCR, 5 lines incl. one at 18px PASS all 5 exact T2 counting + attribute binding PASS 7 circles / 3 triangles / 1 square T3 bar chart, 6 values + max/min PASS 6/6 exact T4b occlusion, star behind rectangle PASS T4c aspect ratio of a 160x140 rectangle FAIL called it taller than wide T5 two images, which has text PASS T6 four images, the seat's cap PASS all four named T7 five images, one over the cap PASS rejected with HTTP 400 No <think> leak on any vision call. The single miss is fine-grained relative-dimension estimation on a near-square shape, and it reproduced across two runs (the longer T4 called the same rectangle "equal width and height"). Counting, OCR, chart values and occlusion ordering are all solid, so this is a precise-geometry weakness, not a broken tower. Recorded so nobody builds a feature on this model judging relative sizes. T7 earns its place separately: it confirms the per-prompt image cap fails loudly with a 400 rather than silently dropping the extra image. |
||
|
|
f85d102813 |
test(gen-seat): orcarouter passes every gate — in-band MTP head delivers +11 points
Gates run against the live seat while the operator tested in parallel. <think> leak (n=30, 4 prompt types + multi-turn) 0/30, 0 empty MTP acceptance 58.4% @ 117.11 tok/s median surface 6/6 abliteration survival 4/4 compliance deterministic quality gens coherent and correct PPL still blocked For scale on the leak gate, the abandoned h300 build scored 8/30 on this exact instrument, and its abliteration-survival samples had 2 of 4 open with "<think>Ok, let's figure this out:". Orcarouter has none. The headline is MTP acceptance. 58.4% against heresy's byte-identical base head at 47.2% is +11 points, and it sits level with our own in-band L35 at 59.1%. That is the additive in-band-vs-graft delta the entire Cold-Fusion experiment was built to measure and never cleanly delivered -- orcarouter handed it over for free because the author had already done the Robinson edit on the head. Surface 6/6 covers plain chat, vision, tool calling, the thinking split, a 36,042-token long-context retrieval, and streaming. PPL remains blocked on a spec-decode-free probe seat: it needs ~22 GB and GPU1 has ~16 GB free. Comparison target is heresy at 6.910. |
||
|
|
ba53c30192 |
feat(gen-seat): cut over to orcarouter — live, 7/7 aliases, vision intact, no think-leak
Operator directive was seat-first so he can test while the gates run. GEN_MODEL -> /tank/aimodels/qwen38-27b-orcarouter-nvfp4-mixed. Healthy in ~4 min. KV pool 401,550 tok / 1.53x. MTP drafter detected and wired, sharing embedding and lm_head with the target. 7/7 gateway aliases 200. Vision correct on the shape probe. Live decode observed at 102-133 tok/s under load. Critically, <think> does not appear in the top-20 first tokens on the live seat. That is the Cold-Fusion failure mode measured absent in production, matching the pre-quant screen on the bf16 (1.23e-06, rank 52). Rollback is one line to .env.bak-heresy-restored-20260821. Full gates were still running when this landed; PPL stays blocked on a spec-decode-free probe seat, which needs ~22 GB against GPU1's ~16 GB free. |
||
|
|
c8f128bdff |
feat(gen-seat): quant orcarouter — its MTP head is already Robinson-abliterated in-band
Pulled orcarouter/Qwen3.8-27B-Uncensored at rev 9878936b (55.5 GB, gated, our token has access) and built /tank/aimodels/qwen38-27b-orcarouter-nvfp4-mixed (23.4 GB, mixed NVFP4+FP8). Verified, not yet cut over. The operator asked whether we could apply the Robinson path to the MTP head. We cannot, because the author already did. compare_mtp_head.py against the verbatim base graft: 13 of 15 tensors byte-identical, exactly 2 differ -- mtp.layers.0.self_attn.o_proj.weight and mtp.layers.0.mlp.down_proj.weight, which are precisely the two residual writers our own abliterate.py targets (EXPECT_MTP_WRITERS = 2). Reverse-engineered the edit from the weights alone (mtp_delta.py, added here): sigma2/sigma1 = 0.0164 on BOTH tensors rank-1, a single-direction projection |cos| between the two recovered dirs = 1.0000 ONE shared direction ||delta||/||W|| = 1.42% and 1.41% a gentle, consistent projection sink energy dim 3994 = 0.0000% sink-clean; Heretic's was 6.18% That is the Robinson in-band MTP abliteration, already applied, with a direction that passes our sink screen outright. Nothing to do but preserve it, and the quant carries it byte-identically. This is the configuration the entire Cold-Fusion experiment was designed to test and never cleanly delivered. The new format screen paid for itself on its first real use: think_prior.py on the bf16 BEFORE any GPU time gave P(<think>) = 1.23e-06 at rank 52, against Cold-Fusion stock 0.1850 and h300 0.2216. Roughly 150,000x cleaner. Two durable findings about the pipeline itself: The quant needs ~17 GB, not a whole card. It ran entirely in GPU1's spare 16 GB with ZERO production seats stopped -- the h300 run's "stop BOTH GPU0 seats" was never necessary, it simply had a free card by coincidence. The first attempt OOM'd by 2.37 GiB at layer 64 of 65 with 3.57 GiB reserved-but-unallocated, which is fragmentation, and PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True closed it. post_quant.py now builds a missing output index from the safetensors headers. A sub-23 GB quant saves one bare shard with no index, and post_quant needs one; this has broken three separate rounds and been hand-fixed every time. The header is read by struct-unpacking the u64 length and parsing the JSON -- never safe_open, which mmaps the whole 22 GB shard and ENOMEMs on ZFS. Artifact verified: mixed-precision, 1968 tensors, 15 mtp, 333 visual, re:^mtp.* present in the ignore list (llm-compressor pruned it as always), preproc restored. Imatrix deferred per operator; the log confirms the usual uniform-MSE fallback, so this build stays apples-to-apples with heresy's PPL 6.910. |
||
|
|
bf65d0254d |
docs(pfi): evaluate the two gen-seat replacement candidates
preetpatel/Qwen3.8-27B-Uncensored-NVFP4 is disqualified on two independent hard
failures, both read directly off the artifacts via HTTP Range requests against the
safetensors header (about a megabyte, not a 20 GB download):
- ZERO mtp tensors. The author's recipe.yaml asks to ignore re:.*mtp.*, but the
written config.json has no mtp ignore entry while re:.*visual.* expanded to 110
explicit ones. That asymmetry is llm-compressor pruning a pattern that matched
nothing, i.e. the MTP head was never loaded. Costs roughly half our decode.
- NVFP4 W4A4, 4-bit activations. Precisely the AEON failure mode: the fidelity
gradient is W4A4 < W4+FP8 < W4+bf16, W4A4 drove ~15-20% stochastic degeneration,
and it collapses past ~30k context. The gen seat serves 262K.
orcarouter/Qwen3.8-27B-Uncensored checks out as a quant source: stock-Qwen base
rather than a reasoning-compression finetune, Arditi-style single-direction
abliteration, 15 mtp and 333 visual tensors verified present, chat template
byte-identical to the heresy build we are serving, and the gate is already accepted
on our token.
Also records the author's FP8 release as a noted-but-not-recommended third option:
far more traction, but 30.9 GB against NVFP4's 22 GB, and on a zero-sum GPU0 that
+9 GB comes out of the KV pool and breaks 262K context.
And states the imatrix constraint plainly. Our recipe has always requested
imatrix_mse and always silently fallen back to uniform MSE; playbook 3.13 warns
against assuming an imatrix would help before verifying llm-compressor can consume
external importance data at all. The W4A16 portions are data-free by construction
and cannot use it regardless.
|
||
|
|
48410a6a90 |
chore(coldfusion-abliteration): delete the Cold-Fusion bf16 weights — ~154 GB reclaimed
Operator directive following the decision to abandon the Cold-Fusion base.
Removed with explicit literal paths, one at a time:
qwen38-27b-coldfusion-bf16 stock DavidAU base
qwen38-27b-coldfusion-abliterated-L35-bf16 Robinson L35
qwen38-27b-coldfusion-h300-mtp-bf16 Heretic-300 + MTP graft
qwen38-27b-coldfusion-heretic300-bf16 raw Heretic export
Verified against ZFS used, not df: 4.48T -> 4.33T, ~154 GB. No snapshots were
holding the blocks, all four paths confirmed gone, gen seat unaffected.
The last two were hardlink twins -- same inode, links=2, because the MTP graft
hardlinked every unchanged shard -- so deleting only one would have freed
nothing. `du` across several paths in a single invocation dedupes hardlinks and
reported heretic300-bf16 as 2.5K, which would have made a size estimate wrong in
both directions. Check `stat -c %h` before sizing a delete.
Kept deliberately, so the research record outlives the weights:
qwen38-27b-coldfusion-bf16.PROVENANCE.txt pinned HF revision 9c44193f
coldfusion-abliteration/ harness, 300-trial Optuna
journal, catatonia-T260.json
With those two, every deleted build is reproducible: re-pull stock at the pin and
replay the winning config.
Held back pending an explicit call: the two NVFP4 quants, h300-nvfp4-mixed (the
only remaining servable copy of the Heretic-300 result) and L35-nvfp4-mixed. The
directive named bf16 weights; these are quants, and there is no storage pressure
arguing for haste at 4.26T free.
|
||
|
|
37e9e1ca7f |
revert(gen-seat): abandon Cold-Fusion, roll back to heresy — the leak is in the base
Operator directive, given before the result was in: if it's the base, abandon h300 and the base too. The dose-response said base (18.5% of 22.2%), so it fired. Live gen seat is /tank/aimodels/qwen38-27b-heresy-nvfp4-mixed again, restored from .env.bak-coldfusion-L35-20260820. The h300 env is preserved at .env.bak-h300-abandoned-20260821. The clincher, same probe pointed at heresy: Cold-Fusion stock P(<think>) 0.1850 Cold-Fusion L35 0.2048 Cold-Fusion h300 0.2216 heresy (restored) not in the top 20, <0.002 A >100x gap between the families, which is why no rollback inside Cold-Fusion would have helped -- stock and L35 leak at nearly the h300 rate. Verified after rollback: 0/30 leaks and 0 empty on the same instrument that scored h300 at 8/30, with the EXISTING enable_thinking:false config; KV pool 403,065 tok / 1.54x, heresy's exact documented baseline; 7/7 aliases; vision intact. No LiteLLM change was needed, so the chat_template_kwargs fix is left unapplied -- it worked, but it was a workaround for a base we no longer serve. Cost, stated plainly: 8/100 refusals becomes 29/100, a 3.6x regression on the axis the whole Heretic-300 run existed to move. Accepted deliberately. What carries forward is the methodology, none of which lived in the Cold-Fusion weights: direction_scope=0 beating per-layer on a merged base, aggression not being the lever, PR #317 silently dropping the MTP head on save, the MPOA and sink-screen reasoning, the graft/KL/catatonia/export harnesses, and the finding that a pristine MTP graft accepts as well as an in-band edit. New acceptance gate earned here: run think_prior.py on a candidate's STOCK weights before committing GPU time. It is a ~10s CPU measurement and it would have disqualified Cold-Fusion before the 300-trial study ever started. Heretic's objective has no format-compliance term at all -- the same blindness that removed the self-harm guardrail. Nothing deleted. Every Cold-Fusion artifact, the 300-trial Optuna journal and catatonia-T260.json remain on disk. Abandon means stop serving, not rm. |
||
|
|
5ee2325820 |
feat(coldfusion-abliteration): dose-response settles the <think> leak — base 83%, our abliteration 17%
Answers "how likely is it that our abliteration caused this?" with a measurement instead of a prior. P(<think>) at the first generated token, template rendered enable_thinking=false so the prompt already carries a CLOSED think pair -- the exact event behind the leak. Raw softmax, bf16, CPU-only, one process per model. Deterministic: stock reproduced to 17 significant figures across two runs. coldfusion-bf16 none (stock) 0.1850 rank 3 coldfusion-abliterated-L35-bf16 Robinson L35, mild 0.2048 rank 2 coldfusion-h300-mtp-bf16 Heretic-300, heavy 0.2216 rank 2 The stock, untouched base already puts 18.5% of first-token mass on opening a think block the template had closed. Abliteration adds a real, monotonic, dose-dependent +3.7 points -- a nudge on a pre-existing base, not the cause. Cold-Fusion is a reasoning-token-compression finetune, i.e. a model trained to think briefly, and the leak's text shape agrees: a compact correct trace with a trained transition marker, which is trained behavior rather than damage. This changes the options. Rolling back to L35 or stock does NOT fix the leak -- at 18.5% under temp 0.7 / top_p 0.8 they leak at nearly the h300 rate. Only leaving the Cold-Fusion family escapes it, at the cost of the 8/100 refusal result. The chat_template_kwargs fix is the correct lever. Durable methodology point: a forward-KL budget cannot catch this. Heretic minimizes forward KL(stock||abliterated), which is near-blind to the model putting new mass on tokens stock barely used -- that is reverse KL's job, and we measured exactly that asymmetry on L35 (reverse 1.43 vs forward 0.70). h300's KL of 0.0136 is not evidence of innocence. For any "did the abliteration break behavior X" question, measure P(token) directly. Ran CPU-only deliberately: 96 EPYC cores and 265 GB of RAM make a 27B forward pass cheap, so this cost no GPU window and no seat downtime, where the obvious route was stopping both GPU0 seats. Also normalizes two more abliteration output dirs from root-owned 0600 to llmuser 0664. The unreadable-model failure surfaces as FileNotFoundError rather than a permission error, which is worth knowing before it wastes a run. |
||
|
|
91f4cf22e1 |
fix(gen-seat): diagnose the unterminated-<think> leak — model defect, temp-triggered
Operator reported the new Heretic-300 gen seat "sends CoT but never completes
the turn" through Lobe. Diagnosed; not yet fixed (the fix changes gen's
semantics, so it is the operator's call).
The Qwen3.8 chat template appends a pre-closed <think>\n\n</think>\n\n when
enable_thinking is false. The h300 model opens a fresh <think> anyway and never
closes it. Because the prompt already closed the block, vLLM's qwen3 reasoning
parser is not in reasoning state, so the tag passes through as ordinary text --
reasoning_content empty, reasoning_tokens 0, and the whole reasoning-plus-answer
blob lands in content. Lobe then correctly treats the unterminated tag as
still-thinking and renders no answer. The client and the serving stack are both
behaving correctly; the model is not.
The trigger is TEMPERATURE, not presence_penalty (n=12 per arm):
temp 0.7, pp 1.5 (current gen) 4/12
temp 0.7, pp 0.0 4/12
temp 0.7, pp 0.5 3/12
temp 0, pp 1.5 0/12
That falsifies the standing hypothesis, recorded in the litellm config comment
and in the operator's own 2026-08-16 note, that presence_penalty 1.5 is the
first dial to move. It is not this bug's cause.
It also explains the blast radius: only the two temp-0.7 aliases leak, `gen`
and `summarizer-large`. summarizer, classifier, image-judge and qwen-image-bench
all run at temp 0 and are clean, so nevermore's summarizer path is unaffected.
Candidate fix, validated n=30 over 4 prompt types plus a 3-turn conversation:
chat_template_kwargs {enable_thinking: true, reasoning_effort: low} takes 8/30
leaks to 0/30, at ~+27% completion tokens and a ~3% empty-content residual.
The tell appears in eval_coldfusion_h300.json and in none of the aeon, heresy,
mixed or w4a16 evals, so it is new with this build -- but L35 was never evaled,
so this does not separate a Cold-Fusion base trait from a Heretic-300
abliteration artifact.
Reproducers and the full method land in bench/think-leak/. Note in particular
that the 7/7 alias smoke test run at cutover structurally could not catch this:
trivial prompts never invite reasoning, so they never sample the leaking token.
|
||
|
|
1d3b80169a |
fix(nevermore): repoint onto live aliases — its LLM pass had been dead 8 days
nevermore pinned LLAMA_SWAP_MODEL=granite-4.1-8b, an alias retired with the
granite seat on 2026-08-12. Every summarization call since then failed: 67
consecutive status=failure rows, 0 tokens, twice daily, entirely silently. The
briefing had been rendering with no LLM pass at all. Nothing alerts on
status=failure in the spend logs, so it took an unrelated question about
reranker VRAM to surface it.
It was also pinned to NEVERMORE_RERANK_MODEL=qwen3-reranker -- the incumbent
Brokkr R43 measured harming 80/90 fleet queries -- and was its ONLY caller,
while the production `reranker` alias sat at 0 calls for 4 days. The R43
cutover repointed the alias but never moved the consumer.
nevermore/.env LLAMA_SWAP_MODEL granite-4.1-8b -> summarizer
NEVERMORE_RERANK_MODEL qwen3-reranker -> reranker
(server-only; .env is excluded from the mirror both ways)
Verified against nevermore's exact call shape: summarizer returns clean content
with 0 reasoning chars at temperature 0.2 / max_tokens 4000; reranker scores
0.95 on-topic vs ~1e-5 off-topic; embedding returns dim-1024.
Retired alongside it:
vllm-rerank :8002 Qwen3-Reranker-0.6B + the qwen3-reranker alias
vllm-rerank-a4 :8014 gte-reranker-modernbert + its alias
vllm-granite :8004 Exited 8 days, dead service block
and vllm-rerank-a3 was promoted from a throwaway `docker run` into this stack
(the selection ledger's own open follow-up). Healthy in 55s. It keeps the
bake-off arm name so the ledger, memory and R43 record stay valid.
VLLM_VERSION is pinned latest -> v0.24.0. Every service in the stack shares that
one variable, so a bare `compose up -d` could have silently upgraded all of
them at once; both tags resolved to the same local image (4091d5593f77), so the
pin changed nothing at runtime.
GPU1 is down to 81,448 of 97,887 MiB -- 13.9 GB reclaimed tonight.
Correction: an earlier claim that A4 had no gateway alias was wrong. It did.
LiteLLM serves both config-defined and DB-defined models -- live showed 32
against config.yaml's 26 -- and grepping the file cannot see the difference.
/v1/models and /model/info (which flags db_model) are the ground truth. DB
models delete hot via POST /model/delete with no restart.
Left alone: reranker-a3-bge-v2-m3, a zero-call duplicate of `reranker` on the
same backend. It is Brokkr's cutover-verification handle -- redundant rather
than broken, and another agent's tooling is not mine to delete unilaterally.
|
||
|
|
b990951d80 |
chore(vllm): retire LFM2.5-2.6B permanently; audit finds nevermore on the harmful reranker
Operator directive: lfm2.5-2.6b goes down permanently.
- stacks/vllm/compose.yaml vllm-lfm25 service removed (replaced by a
tombstone comment), pushed live to ana-ml2
- ana-ml2 container docker rm -f'd, 8,721 MiB freed on GPU1
(95,388 -> 86,667 of 97,887)
- litellm config lfm2.5-2.6b alias deleted, live + canonical,
28 -> 27 models
It was an EVAL-ONLY bake-off seat against granite-4.1-8b that never received
the operator ruling it was pending; the comparator was retired from the roster
on 2026-08-15; it was deliberately never wired into any default or fallback
routing chain; and spend logs show 0 calls in the 4-day window to 2026-08-21.
Weights stay in the shared HF cache -- nothing deleted from disk.
The gateway restart that makes the alias deletion take effect is HELD so it can
batch with a pending reranker change. Until then the name is still routable
in-memory and will error against a dead backend.
Auditing the three reranker seats while answering "why do we have three" turned
up a real problem. The design is one production, one rollback, one fallback --
but the traffic is backwards:
:8013 A3 bge-v2-m3 PRODUCTION, backs `reranker` 0 calls / 4 days
:8002 Qwen3-Reranker RETIRED incumbent, rollback only 7 calls, 12-hourly
:8014 A4 gte-modernbert "fallback" no alias at all
nevermore is hard-wired to the incumbent by name (NEVERMORE_RERANK_MODEL=
qwen3-reranker), so the R43 cutover never moved it -- the cutover repointed the
`reranker` alias and correctly left `qwen3-reranker` naming the Qwen model.
Brokkr R43 measured that model harming 80/90 fleet queries, so nevermore's
twice-daily rerank pass is likely degrading its own briefing.
Fix is one line in nevermore's .env plus a nevermore restart, and it must land
before :8002 is retired. Recorded in persistent-memory with the A4 alias also
noted as absent (global CLAUDE.md names reranker-a4-gte-modernbert; it does not
exist).
|
||
|
|
e3ce713f7f |
feat(gen-seat): cut over to Heretic-300 — 7/7 aliases, vision intact, MTP 59.7%
Live GEN_MODEL is now qwen38-27b-coldfusion-h300-nvfp4-mixed (ana-ml2 GPU0 :8015). Served-name left unchanged so all 7 LiteLLM aliases route without a gateway edit. Verification: KV pool 401,550 tok / 1.53x (baseline 403k / 1.54x) LiteLLM aliases 7/7 green vision 3/3 shapes, colour+form+position correct MTP acceptance 59.7% median @ 118.37 tok/s quality gens 4/4 correct abliteration 4/4 compliance PPL NOT measured (see below) The roadmap predicted ~47% acceptance for a pristine MTP graft versus L35's 59.1% in-band edit. Measured 59.7% on the same harness: there is no acceptance penalty, which removes the throughput argument for reimplementing MPOA. A single long-prose generation read 47.5% off the same counters -- below the 8-run minimum of 49.0% -- and would have "confirmed" the prediction by coincidence. Acceptance must be read from quickbench.py, never one sample. PPL is blocked on VRAM, not on the model: eval_quality.py aborts with "prompt_logprobs look uniform" under --speculative-config, and the probe-seat workaround needs ~22 GB while both cards sit at ~96% committed. Also normalizes the quant dir from root:0600 to llmuser:llmuser 0664 to match every other model dir, and records that config.json sha256 is byte-identical across the h300 and L35 quants and is therefore useless for confirming which weights are mounted (mtime and a head-hash are the discriminating views). Rollback is one line to .env.bak-pre-h300-20260820. |
||
|
|
407ca017ae |
memory: snapshot — Heretic-300 built, quantized and verified; gen-seat cutover is the next step
8/100 refusals at KL 0.0136, hand-verified coherent, beating the absolute-heresy bar 3.6x. NVFP4 quant complete: 21 GB, 1968 tensors, MTP head grafted back after PR #317 dropped it, and re:^mtp.* re-injected into quantization_config.ignore after llm-compressor pruned it. Self-harm guardrail is gone on this build and is the operator's own next work item; the four-dwarf panel is stood down. |
||
|
|
f90a5025de |
feat(coldfusion-abliteration): Heretic-300 — 8/100 refusals at KL 0.0136, beats the heresy bar 3.6x
Ran Heretic v1.4.0's 300-trial TPE search on Cold-Fusion. Best trial scores 8/100 refusals at KL 0.0136 against a 98/100 base, versus absolute-heresy at 29/100 and our hand-tuned Robinson L35 at 72/100 / KL 0.0116 — i.e. 64 fewer refusals for the same damage. Hand-verified coherent: correct arithmetic with shown working, clean code, 66-167 word prose across nine probes. Durable findings: - direction_scope=0 (single shared direction) is decisive on this merged base: n=129, best 8/100. Per-layer directions n=131 never beat 52/100 despite a better median. Points against the multi-direction intuition for a diffuse direction (our two-template |cos| is 0.62 vs Robinson's 0.99 on stock). - Aggression is not the lever. r(KL, refusals) = -0.561 over 261 trials; the KL<0.02 band contains both the worst results (median 87/100) and the single best. A KL 0.3554 trial scored worse than one at 0.0193. - PR #317 confirmed: Heretic silently drops the MTP head on save. Source 1199 tensors -> export 1184, all 15 mtp.* gone, vision 333/333 intact, exit 0, no warning. This is also why absolute-heresy ships a byte-identical MTP head — a bug, not a design choice. Always diff tensor keys after a Heretic export. - Heretic's recovered direction carries 6.18% of its energy in sink dim 3994, versus 0.094% for our L35 and 1.97% for the L39 we rejected as brick-inducing. It survives that only because of magnitude-preserving ablation (row_normalization=FULL); our plain projection has no such protection, so the sink screen correctly refused the in-band MTP graft. Same direction, different operation. MPOA is the prerequisite for in-band MTP on a Heretic trunk. - Heretic's edit is recoverable from weights: delta is rank-1 (s2/s1 ~ 0.010), SVD gives the direction, norms give per-layer weights (1.08 -> 1.34, i.e. over-projection). Cross-layer |cos| agreement 0.9903 independently confirms the single-direction result. New tooling in services/coldfusion-abliteration/: kl_divergence.py first-token KL, class-split, zero noise floor catatonia_gate.py 12 probes x 220 tokens, prints every completion heretic_export.py PTY driver; selects by measured value, never by menu position — Heretic's resume prompt puts "delete the checkpoint and all results" one arrow-key from the target graft_mtp.py recovers the trunk direction by SVD; --pristine for the safe path when the sink screen refuses Also adds quant playbook 3.13: the NVFP4 recipe sets observer="imatrix_mse" but llm-compressor has always silently fallen back to uniform MSE for want of importance data — on this build and on the incumbent. Existing A/B comparisons stay valid since every build shares the fallback. Parked as id 42. Guardrail note: this build has lost the self-harm guardrail that the Robinson L35 build retained. Restoration is the operator's own work item. |
||
|
|
78484ac87d |
memory: GPU0 seat boot order is part of the state — restore rule + KV-pool baselines
vLLM sizes --gpu-memory-utilization against total VRAM but gates startup on free VRAM, so the GPU0 pair coexists only in its original boot order. Records the restore sequence (meromero to healthy first, then gen), the observed-not-slept rule, and the KV-pool baselines to verify a restore against — nvidia-smi used-MiB is the wrong check, it swings ~7 GB on allocator slack at identical capacity. |
||
|
|
a9d73dad41 |
fix(coldfusion-abliteration): GPU0 seat restore order is load-bearing — correct the claim and the runbook
Restoring the two GPU0 seats with `start meromero; sleep 10; start gen` put
meromero into a 7-restart crash-loop:
ValueError: Free memory on device cuda:0 (35.3/94.97 GiB) on startup is less
than desired GPU memory utilization (0.52, 49.38 GiB).
The previous commit's README claimed restore order "is not actually load-bearing"
on the grounds that both seats pass --gpu-memory-utilization as a fraction of
total VRAM. That is half right and the wrong half mattered: the fraction sets the
target, but vLLM gates startup on FREE VRAM and refuses to start unless the whole
target is available. GPU0 runs at ~96.4/97.9 GB with roughly 0.4 GiB of slack, so
the seats coexist only in the order they were originally brought up, and meromero
is the one that does not fit in the remainder. The pre-existing auto-memory note
("gen takes a fraction of free VRAM at startup and will starve meromero") was
pointing at the real effect.
Also: "first" means healthy, not ten seconds earlier. A sleep 10 against a
two-to-three minute weight load is simultaneity, not ordering — gate on observed
state.
Recovery applied: stop gen, wait for meromero healthy, start gen. Verified
against the pre-window baseline rather than against "both green":
gen KV 14.36 GiB / 403,065 tok / 1.54x -> 14.34 GiB / 401,550 tok / 1.53x
meromero KV 542,202 tok -> 542,202 tok
RestartCount 0 on both; summarizer smoke-tested through LiteLLM
Note for the next reader: raw nvidia-smi used-MiB is the wrong check here. It
reads 89,503 now vs 96,376 before, which looks like a 6.9 GB regression and is
allocator slack — serving capacity is unchanged. The anomalous boots were the
high ones (34.95 GiB KV), where gen came up on an empty card mid-window.
|
||
|
|
1b3fb270e7 |
feat(coldfusion-abliteration): first-token KL measured — 28.4x selectivity, harmless median 0.0211
Adds `kl_divergence.py`: first-token KL(stock || abliterated) over the full 248,320-token vocabulary, bf16 vs bf16, scored separately for held-out harmless and reserved-harmful prompts. Result (L35, 256 harmless / 104 harmful, answer mode): harmless median 0.0211 mean 0.0364 top-1 agreement 89.8% harmful median 0.5996 mean 0.6992 top-1 agreement 55.8% selectivity 28.4x (72.8x in think mode) Self-KL noise floor is exactly 0.0, and all 720 per-prompt values are bit-identical between a single-process and a two-process run, so the figures are signal rather than bf16 jitter. Reverse KL on harmful/answer is 1.43 vs forward 0.70 — the mass-where-stock-had-none asymmetry expected of a refusal-direction removal. Against the Heretic reference figures (0.1191 prior seat, 0.0759 the live absolute-heresy seat) this is materially gentler, but those are the other tool's optimizer output on a different base with its own harmless set and template — order-of-magnitude, not head-to-head. KL remains a fidelity number; the viability gate is still MTP acceptance (59.1%). Method notes: - Prompt classes are reported separately by design. A single averaged KL over a mixed corpus is close to meaningless, since the metric is meant to be large on harmful prompts and small on benign ones; the ratio carries the information. - The harmless evaluation set is drawn from the alpaca pool minus calibration's own draw, reconstructed by replaying that draw rather than remembered, and asserted disjoint on text. The harmful set is the reserved test split. - `render` is imported from abliterate.py rather than copied, so the measurement cannot drift from the rendering the direction was captured against. - Batch size 1 with logits_to_keep=1: no padding semantics, ~0.6 MB of logits. Three corrections to the runbook, each of which cost time: - "bf16 is 50 GB, only gen must go" was 50.10 GiB mislabelled. Text-only weights are 51,300 MiB; freeing either GPU0 seat alone leaves ~50,933 MiB. Both must stop. VRAM is now sized from the safetensors headers at run time. - A 27B model cannot be released in-process: `del` + gc + empty_cache left free VRAM at 45,287 MiB, and so did confining the model to an inner frame that exits. Only process exit returned the card (96,689 MiB). The first run completed only because the allocator hit OOM, collected, and retried. Each model now gets its own process, handing log-probs to disk between stages. - The residency gate read hf_device_map, which transformers leaves empty when the model fits on one device — it reported "(unsharded)" whether or not anything was wrong, so it could never fail. It now reads parameter devices directly. Model-agnostic lessons promoted to the quant playbook (new 3.12). |
||
|
|
8c354a0e79 |
memory: snapshot — Cold-Fusion thesis PROVEN (MTP 59.1% > incumbent 47%)
Flip the Cold-Fusion in-flight line to thesis-proven: L35 quantized to mixed NVFP4, MTP acceptance 59.1% median beats the incumbent Heretic graft's ~47%, abliteration survives quant. Not cut over — cutover is a separate operator decision. Records the two env foot-guns hardened (quant venv config-delegation drift; single-file no-index quant needs a header-built index). |
||
|
|
725c8fdf9e |
feat(coldfusion-abliteration): THESIS PROVEN — in-band-abliterated MTP head accepts 59.1% (beats incumbent ~47%)
Quantized the L35 abliterated model to mixed NVFP4 and measured MTP acceptance
end to end. The experiment's whole premise: Heretic (the incumbent gen seat)
leaves the MTP head a byte-identical base graft its wrapper never loads, whereas
Robinson abliterates the MTP head in-band — the question was whether that in-band
edit survives well enough to spec-decode. It does, better than the graft:
MTP acceptance 59.1% median (51-65%, 8 cache-busted topics) vs incumbent ~47%
decode 118.7 tok/s median (faster; image-confounded, read as not-worse)
abliteration survives quant (creative refusals drop, self-harm guardrail
intact, coherent)
Output at /tank/aimodels/qwen38-27b-coldfusion-L35-nvfp4-mixed (22.5 GB). Result
JSON in bench/. NOT cut over — the incumbent seat is untouched; making L35 the gen
seat is a separate decision needing the full Stage-3 gate + real multi-turn hold.
Two env foot-guns hardened along the way:
- quant_mixed_nvfp4.py now promotes text_config attention fields
(num_attention_heads etc.) to the top-level config for the oneshot, then
restores. transformers 5.10 / llmcompressor 0.12 (this venv moved under us
since the Aug-15 heresy quant) no longer delegate the top-level lookup, so
oneshot raised "Cannot determine num_attention_heads". Same "the fight is the
environment" pattern as the abliteration capture.
- a sub-~23GB quant saves as a single model.safetensors with no index, so the
post_quant MTP graft needed an index built first — from the safetensors header,
not safe_open (which mmaps the whole shard and ENOMEMs on ZFS).
post_quant grafted the abliterated MTP (15 tensors, 849 MB) and re-injected
re:^mtp.* into quantization_config.ignore (llm-compressor pruned it again — the
two-rounds-lost 0%-MTP bug, fired and repaired as designed). Probe served on the
pinned nightly (#51113 qwen3_5_mtp fix) to match the live seat's vLLM.
|
||
|
|
c55b1390b7 |
memory: snapshot — Cold-Fusion abliteration LANDED at layer 35
Flip the in-flight status from 'capture done, calibration expansion next' to 'landed, works'. New detail file captures the three corrected diagnoses (layer- selection metric, sharding/allocator misdiagnosis, corpus-size falsified) and the verify/quant work still owed. Supersedes the -capture.md detail file's framing. |
||
|
|
e9dbc8660b |
feat(coldfusion-abliteration): abliteration LANDS at layer 35 — separation selector, shard-surgery write, three false diagnoses corrected
The abliterated model works. A/B vs stock on a matched greedy battery: explicit sexual + graphic torture (the measured stock refusal surface) go from refused to complied/engaged, held-out AdvBench prompts loosen, the self-harm guardrail survives, coherence intact — the Robinson design point exactly. Output at /tank/aimodels/qwen38-27b-coldfusion-abliterated-L35-bf16, verified bitwise: 131/131 targets changed, 333/333 vision byte-identical (delta 0.0), 735/735 others untouched. Getting there corrected three diagnoses the prior session had backwards. 1. The layer-selection metric was wrong, and that was the whole ballgame. The recipe picks the abliteration layer by peak two-template |cos| agreement. On this heavily-merged base that metric is anti-correlated with efficacy: its argmax (layer 18) is the WORST-separating layer in the window (Cohen's d 5.51 vs 9.89 at the peak), and abliterating there was a measured behavioral no-op — stock and "abliterated" refused all six probes identically. Cause: the two renderings end in different generative modes (</think> vs <think>), so |cos| scores answer-vs-reason mode, not refusal, and on a merge the mode term dominates. Replaced selection with harmful/harmless SEPARATION (Cohen's d / AUC of the direction's projection), gated on the sink screen since separation and sink-energy both climb with depth. Picks layer 35 (d 9.35, AUC 0.9997, sink 0.094%). Agreement is kept as a printed diagnostic. 2. The "bf16 NaNs, use fp32" rule was a misdiagnosis. The NaN was never precision — it was multi-GPU sharding (the residual stream zeroes two layers past the GPU0->GPU1 boundary; the first capture's layer 22 happened to sit in the healthy region, which is why it looked fine) plus PYTORCH_CUDA_ALLOC_CONF=expandable_segments (corrupts retained tensors; the corruption MOVED between bit-identical forwards, the tell that it was memory not math). On one GPU with a plain allocator, bf16 full-64-layer is exactly deterministic and coherent, at 50 GB and 4.3x the throughput of the 111 GB fp32 it replaced. Both defects are now hard gates (residency exit 8, allocator exit 9); capture pins CUDA_VISIBLE_DEVICES=0. 3. The corpus-size hypothesis was falsified. 52x more calibration data (8->416, mlabonne/harmful_behaviors = the recipe's actual AdvBench split, already on the box) moved agreement 0.594->0.624 — nothing. Kept the 416/416 corpus anyway (calibration.py); it gives the clean separation signal. The held-out 104-prompt test split is reserved and asserted disjoint. Also: the --out write is now shard-level surgery (reads/writes the 18 safetensors directly, no model object, no GPU). This is correctness, not thrift — AutoModelForCausalLM resolves to the TEXT model, so save_pretrained would drop all 333 vision tensors AND skip the MTP head (the in-band MTP edit is the entire point of the Robinson formula). Neither failure raises. Shard surgery makes vision and the other 1068 tensors byte-identical by construction. Batched capture with a dtype-aware equivalence gate; hidden states captured via forward pre-hook (reading output_hidden_states off the returned object is unsafe here — buffers get recycled). Sharding/allocator lessons promoted to the quantization playbook (model-agnostic, sections 3.9-3.11 + superseded table); the selection-metric lesson added to the recipe doc. The dead layer-18 no-op checkpoint was removed (52 GB, confirmed identical to stock). Incumbent gen seat untouched. Full canonical refusal-probe re-profile and MTP-acceptance-on-quant still owed before this becomes a gen-seat candidate. |
||
|
|
f714f28195 |
feat(coldfusion-abliteration): Robinson's real 416-prompt corpus, batched capture, two new gates
The 8/8 calibration set gave |cos| agreement 0.594 against the recipe's 0.9925. This wires in the corpus the recipe actually used and makes a capture at that scale affordable. Corpus (calibration.py, new). The recipe's "held-out train/test split of 416/104 with overlap 0" names mlabonne/harmful_behaviors exactly — 416 train / 104 test, AdvBench-derived — and it plus harmless_alpaca were already staged in ana-ml2's HF dataset cache. Read via pyarrow, no datasets dependency, no hub access. Harmful is order-deterministic (no seed), so a re-capture is reproducible from the flags alone. The 104-prompt test split is reserved as the held-out generalization probe and asserted disjoint, so the post-write re-profile cannot silently become in-distribution. --calib builtin reproduces the legacy run. Batched capture. 832 prompts x 2 templates = 1664 forwards. Padding is on the RIGHT: in a causal stack nothing after position t reaches position t, so trailing pads cannot touch the token read, whereas left padding feeds pads into the DeltaNet recurrence ahead of the prompt — the path whose torch fallback already NaN'd once here. Means accumulate in float64; the direction is a difference of means, which is where cancellation lives on this model. Gates added, both protecting numbers rather than tensors: - batch-equivalence: proves padded-batch == single-prompt (rel 1e-3) before spending the capture window. - surgery pre-check: aborts if any of the 131 targets is absent or on the meta device. orthogonalize_ edits in place, and an in-place write to an accelerate-offloaded tensor is a silent no-op — that ships a half-abliterated model past a smoke test. Fixed a reporting bug: the agreement line printed the global agree.max() beside the window's argmax layer, so the first capture read as 0.8538 when the real in-window number was 0.5944. The global peak sits in the early layers where the dim-3994 massive activation inflates agreement for reasons unrelated to refusal. Now prints window max, a top-5, and labels the global figure informational. --max-layer truncates the decoder for capture. Exact, not approximate: a causal stack's layer-N state cannot depend on layers above N, so any value above the window top leaves the direction bit-identical while cutting fp32 residency and forward cost. 46 drops 18 of 64 layers and is what keeps fp32 off CPU offload. Refused on the write path, where it would emit a truncated checkpoint. Verified on ana-ml2 without the GPU: dry-run still 1:1 (131 tensors, all coverage gates), calibration loads 416/416 deterministically with its guards firing, both --max-layer guards exit as designed. Also confirmed against chat_template.jinja that enable_thinking=True does resolve reasoning_effort to xhigh, so the two renderings are the recipe's — template selection was not the cause of the low agreement. The re-capture itself is unrun: it needs the fp32 VRAM window and therefore production seat downtime. |
||
|
|
530f1452e8 |
memory: snapshot — Cold-Fusion abliteration in flight, capture done
Captures the session's real work as the in-flight focus: abliterating DavidAU Cold-Fusion with the Robinson formula. fp32 capture succeeded (finite direction, layer 22, sink-clean) but two-template agreement is 0.59 vs Robinson's 0.99 — calibration-set expansion is the next step. New detail file records the full saga including the transformers/DeltaNet bf16-NaN fight (fp32 fix, the causal-conv1d kernel gap, the seat-restart VRAM-greed gotcha). Supersedes the earlier "watch for DavidAU's heretic build" posture — we abliterate it ourselves. Auto-archived 4 closed entries (Recent decisions: Booth-3-features 08-05, worldtree-sdk 07-31; Tried and abandoned: containerd-race 08-03, mv-rename 08-02) to archival-memory.md; the rest of the over-cap entries are held back by the <14-day and open-deferred guards. Index 331 -> 327. |
||
|
|
7abd3011f7 |
fix(coldfusion-abliteration): capture works — fp32 forward + finite-gate
The --capture forward NaN'd repeatedly. Root cause: transformers' Qwen3.5 DeltaNet linear-attention needs the causal-conv1d fast-path kernel, which can't be built here (no nvcc, no prebuilt wheel). Its torch fallback produces nondeterministic all-NaN hidden states in bf16 -- same 11-token input finite on one forward, NaN at layer 4 on the next. bf16 and fp32 share exponent range, so it's precision-driven catastrophic cancellation, not overflow, and fp32 resolves it. Fixes: - --capture now loads fp32 (the write/surgery path stays bf16 -- no forward, no NaN). attn_implementation=sdpa pinned. - A finite-gate aborts on a non-finite direction. The sink screen alone can't catch this: nan > threshold is False, so a NaN direction "passed" it and got saved silently on the first run. Capture result (fp32, full GPU): refusal direction finite, unit-normed, layer 22, sink energy 0.0008% in dim 3994 -- clean, not sink-dominated. Saved. Caveat recorded: two-template |cos| agreement is 0.59 at layer 22 vs Robinson's 0.99, almost certainly the small 8/8 calibration set vs their 416/104. Valid but noisier than ideal; the README flags expanding the sets before the write. README documents the three environment gotchas (fp32-for-capture, the seats that must be stopped for the 110GB fp32 VRAM and how to restore them, and the fla side-dir PYTHONPATH) so the next run doesn't rediscover them. |
||
|
|
b56cb0db13 |
docs(coldfusion-abliteration): dry-run passed — recipe maps 1:1 (131 tensors)
Dry-run against the fully-staged bf16 confirms the Robinson recipe transfers onto the DavidAU Cold-Fusion checkpoint with no name drift: 1199 tensors, 333 vision preserved, down_proj=64/o_proj=16/linear_out=48/mtp=2/embed=1, coverage gate 6/6, exactly 131 tensors to orthogonalize. Harness verified-ready; the destructive write still gates on operator go. |
||
|
|
1857a8eb81 |
feat(coldfusion-abliteration): Robinson-formula harness, gated, staged
Harness to abliterate DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1 using the MTP-aware, vision-preserving recipe in docs/pfi/abliteration-recipe-qwen38.md. Motivation is measured, not assumed: the stock model's refusal profile (probed 2026-08-19, hand-verified) is ~33% on creative content, concentrated on explicit-sexual and graphic-torture, with 4/5 hard-harm refused, self-harm guardrails intact, and zero benign over-refusal. So there is a real creative-content refusal surface. The Robinson formula is chosen specifically because it abliterates the MTP head IN-BAND -- which the current gen seat's Heretic pass does not (its MTP head is a byte-identical base graft the Qwen3_5 wrapper never loads). That in-band MTP edit is the additive delta. The script refuses to brick the model. Two hard gates from the recipe halt before any write: the coverage identity o_proj(16)+linear_out(48)==64 (catches a tensor-name mismatch that would ship a half-abliterated model), and the attention-sink screen on dim 3994 (orthogonalizing a direction living there produces a model that loads, runs, and emits garbage). The direction is captured from two chat templates and the layer auto-picked by peak |cos| agreement in [18,45]. Classification is suffix-based and name-agnostic so it survives minor drift; the coverage gate is the backstop. Modes: --dry-run (enumerate + gate, no forward, no write), --capture (direction + sink screen, no write), default (write to --out). The README sequences dry-run -> capture -> write -> verify, and names the post-checks (vision byte-identical, refusal re-profile via services/refusal-probe/, MTP acceptance on the quant, PPL/coherence). bf16 staged to ana-ml2:/tank/aimodels/qwen38-27b-coldfusion-bf16 (pinned 9c44193, provenance recorded). The destructive run is NOT executed here -- dry-run verification and operator go gate it. |
||
|
|
ccb56a0a51 |
docs(pfi): capture the RobinsonLabs Qwen3.8-27B abliteration recipe
Reference recipe (not a deployed artifact) for MTP-aware, vision-preserving single-direction abliteration of Qwen3.8-27B -- the base family the gen seat runs. Captures the two things this recipe gets right that naive abliterations of this architecture miss: - The MTP head is abliterated in-band (its two residual-write matrices, glue left alone), so speculative acceptance does not collapse on the prompts abliteration exists to fix -- directly relevant to the gen seat's MTP>=40% gate. - The vision tower is preserved byte-identical (333 tensors, max delta 0). Plus the two calibration traps specific to this base: the twice-captured refusal direction (layer 26, |cos| 0.99) and the attention-sink dimension 3994 that bricks the model if orthogonalized out. Documents the coverage gate (o_proj 16 + linear_out 48 == 64 layers) that catches a half-abliterated model before it writes a byte, and the foot-gun that the GGUF imatrix does not cover the MTP block. Links into model-quantization-playbook.md for the quant half of the pipeline. |
||
|
|
7010f9a1da |
feat(booth): render .md/.txt/.log inline in the gallery, collapsible + closable
Docs used to render as a clumsy link that navigated to a separate page. They now render in place: build_gallery pre-renders each doc (markdown -> HTML, plain text raw) and the gallery shows it inside a native <details open> disclosure that spans the full grid width so prose has a readable measure. The doc bar carries: a collapse chevron (the whole <details> summary toggles, works with JS off), a full-page link (still reaches the standalone viewer), a download link, and a session-close ✕. The ✕ needed stopPropagation + preventDefault because it lives inside <summary> — otherwise its click would toggle the disclosure instead of hiding the item. Close is JS (progressive enhancement); collapse is native. Two design points: - Plain text is returned RAW from build_gallery and escaped by the template inside <pre>. Pre-escaping in Python plus Jinja autoescape would double-encode angle brackets; a test pins the single-escape. - Inlining is bounded by DOC_MAX_BYTES. A doc over the limit keeps the old link-out behaviour rather than being rendered into every index load; a test covers the fallback. The shared .markdown-body / .textview typography moved from doc.html's scoped <style> into base.html so the inline body and the full-page view render identically; doc.html keeps only its page-layout wrapper. Updated the pre-existing test_gallery_links_docs_to_view: it asserted the old link-out behaviour the operator asked to change, so it now asserts the inline render plus the surviving full-page and download affordances. 61 pass. Verified live: markdown renders with headings/table/blockquote/code, txt preserves whitespace and single-escapes, collapse and ✕-close both work. |
||
|
|
40257247b0 |
memory: IPv6 plan settled — endpoints not internal numbering; ESH has a /56
Corrects three claims that had been standing in the fleet IPv6 notes and that sent a three-arm research effort after a problem that did not exist: - ESH was recorded as having no IPv6. It has a /56 delegated and a routable WAN GUA -- substantially more prefix than NH3's single /64. - The mesh was recorded as broken by ESH's CGNAT. It is not and was not down; ESH is outbound and working. CGNAT on v4 alongside generous v6 is just the modern ISP pattern, not an outage. - IPv6 was framed as the escape hatch for that outage. The actual plan is that IPv6 carries tunnel ENDPOINTS for Site Magic and WireGuard, and LANs are not numbered in v6 at all. NH3 internal v6 was brought up on the delegated /64 and verified end-to-end (global GUA on nh3-docker, zero loss to Cloudflare and Google v6, un-NATed source address seen from outside), then reverted on operator direction: one /64 lights exactly one VLAN and that is not worth the split-brain. The AT&T prefix research is kept as reference rather than deleted -- the /60 is real but undelegatable, the living multi-prefix mechanism is multiple IA_PD in one solicit rather than the VRRP/multi-MAC recipe we were handed, and the UDM SE can express neither. That is the answer if NH3 LAN-side v6 ever earns a maintenance window; it is not on any critical path now. |
||
|
|
6770ba26d6 |
feat(booth): kept boards — a .forever sentinel and a standing link board
Agent sessions hand the operator URLs and they drown in terminal scrollback. The Booth is the right home for them — it already has the one property that decides adoption, which is that a session can publish with mkdir and cp, no API key, no schema, no deploy — but everything in it dies in 24h. So: a booth containing `.forever` is never swept, and renders in its own Kept lane at the top of the index. Opt-in per booth, so the ephemeral default is untouched and nobody inherits a cleanup chore. `rm` the sentinel and the board rejoins the sweep; the CLI verbs are sugar over exactly that, which keeps the filesystem-is-the-state model honest. The pin is deliberately NOT wired into is_expired(). That stays a pure age question feeding the `expires_in` countdown; only sweep_once() honours the sentinel. Keeping expiry arithmetic and reaper policy apart means they cannot drift into each other. Kept cards are visually separated per Australis: a 2px top edge in aurora blue, the one accent border the system sanctions. They show "kept" instead of a countdown, and they deliberately lose the one-click wipe button — a × next to the durable stuff is a footgun, so removing a kept board is a two-step act. `booth link <url> [description]` appends to the standing `links` board, creating and keeping it on first use. Entries carry provenance (handle or hostname, plus a timestamp) because a bare URL is unreadable three days later. The append is one printf of one line to an O_APPEND fd — atomic under PIPE_BUF on POSIX — which matters because many agents post to one board and interleaved half-lines would be the obvious failure mode. Seven tests cover the sentinel: detection, survival of a sweep that wipes its neighbour, the deliberate is_expired/sweep_once split, the listing flag, the sentinel not inflating item counts, and both lane-rendering directions. Two of them originally asserted on the bare strings "Kept" and "kept-grid", which passed for the wrong reason — those also appear in the inlined stylesheet served on every page — so they now assert the full class attribute. 55 pass. Also corrects the Homepage card's description, which advertised a flat 24h TTL that is no longer the whole story. |
||
|
|
23cccf5f53 |
fix(homepage): force the canvas clear of the cached wallpaper div
Removing the `background:` block from settings.yaml was not sufficient. Homepage server-renders the wallpaper as an INLINE style on `<div id="background">` and Next.js caches the rendered page, so the aurora survived both the config removal and a container restart. Only a full recreate clears that cache, and recreating this container costs an hour of missing tab bar and i18n before it heals itself. Adding `#background` to the canvas reset is deterministic and immediate, and it also keeps the canvas correct if the setting is ever re-added by accident. The existing selector missed it: the DOM is body > div#__next > div#background, so `body > div` matched the Next.js root, not the wallpaper layer. Verified live rather than locally: the served page now reports no background image, with all three canonical faces loaded and the group eyebrows rendering as JetBrains Mono in Australis cyan. |
||
|
|
b271db1f44 |
feat(homepage): rebuild the theme on canonical Australis tokens
The predecessor theme was ugly for two structural reasons, not one. It did not use the design system's colours. It built a parallel OKLCH palette "derived from the Australis philosophy" and swapped the canonical typeface for Supreme -- a fork, not a theme. Every hex here is now copied verbatim from ~/.claude/skills/australis-design/colors_and_type.css, and build.py re-checks all 19 against that file at build time and warns on drift so it cannot quietly fork again. Type is the canonical stack: Space Grotesk / Inter / JetBrains Mono, vendored as latin-subset VARIABLE woff2 (one file per family, 102 KB total against 56 KB for three static Supreme cuts, and no Google Fonts request at page load). It also carried a generated full-bleed aurora image behind the entire dashboard. Canon forbids exactly that -- "solid fills only on chrome, no full-bleed photography, no decorative gradients", and the aurora motif "never as a background fill behind text". The predecessor knew, said so in its own header, and dialled the opacity down rather than dropping it. The image is gone; the aurora survives as a 1px accent edge under the tab bar, which is where canon sanctions it. The asset stays in images/ in case it is ever revisited. Direction is instrument panel. Group headings become the Australis mono eyebrow with a hairline to the right edge -- canon calls the eyebrow a system signature, and it turns the groups into register bands instead of headings floating over a grid. Status stops shouting: the filled emerald chips read louder than the service names they annotated, so they are now a semantic dot plus a mono micro-label at tertiary contrast. Cards are bordered and opaque, per canon's border-over-shadow rule for chrome. Alignment, per operator feedback that pills and cards did not line up: - The status cluster is centred on the service name's line rather than parked in the card's top-right corner, where Homepage's `absolute top-0` left it floating ~7px above the title's optical centre. The offsets reconstruct the title line box and are documented as moving together. - Descriptions get a two-line minimum, so the common one-line/two-line mix bottom-aligns across a row. This is what made the grid look ragged. useEqualHeights stays false: it inflated short cards to match a widget card twice their height, which was the worse failure. - The status dot is flex-centred rather than nudged with vertical-align, so it stays centred if the type scale changes. Retires the Skyfall sources and the Supreme faces; theme/ now has one source of truth. |
||
|
|
b92097688c |
fix(esh-pve): hardware watchdog, and close the single-resolver DNS SPOF
esh-pve hard-froze at 03:34 on 2026-08-19 and stayed frozen ~4.5 hours until a manual power cycle. No panic, no OOM, no MCE — the journal stops mid-operation. The whole ESH site lost DNS with it, because esh-userland (VLAN 10, the PVC SSID and wired userland LAN) was handed exactly one resolver: 10.0.50.45, AdGuard on esh-docker-vm, on a different VLAN, with no secondary. Internet and routing were healthy throughout. Two fixes. 1. DNS: 10.0.10.1 (the gateway, verified resolving) added as secondary on esh-userland via the UDM Classic API. Note this is degradation cover, not clean failover — clients that query resolvers in parallel will bypass AdGuard for a share of lookups. 2. Watchdog: softdog -> iTCO_wdt under systemd (RuntimeWatchdogSec=60), watchdog-mux masked. The box looked watchdog-protected and was not: a software watchdog cannot fire when the kernel it lives in is wedged, and watchdog-mux only pets the device while an HA client is connected, which never happens on a cluster with no HA resources. Firmware does not block the TCO timer here, checked before committing to it. Also pins VM 102 off (onboot: 0). It starts with full GPU passthrough and vfio-pci enabling that device is the last thing the kernel logged, 39 minutes before the freeze. The other suspect is the kernel itself: the host ran 4.5 months on 6.8.12-16, took 6.8.12-42 in an apt batch on 08-18, and died 20 hours into the first boot on it. 6.8.12-16 is still installed as the rollback. The playbook is idempotent — a second run skips all six steps and passes all six verifies. The watchdog is confirmed armed (identity=iTCO_wdt, state=active, held by PID 1) but has NOT been observed firing; proving that needs a deliberate wedge. Memory also corrects two wrong mid-incident calls: the mgmt VLAN is routed over the site tunnel and is not firewalled off — both symptoms were the dead host generating ICMP unreachables. |
||
|
|
059f963118 |
docs(waterland-studio): note why an adopted job shows as failed
waterland-dev confirmed the mechanism: adoption marks a job failed on a sidecar saying running/queued, or on a directory with no plate.png. The pre-header-fix renders died 1.7s in with a source and no plate, so they land in the second branch. Recorded so nobody investigates adopted history as a live fault. |
||
|
|
e6907819b0 |
feat(waterland-studio): deploy b72425b — all three upstream findings fixed
One update.sh run on irv-ml1 carried both open upstream PRs, per the operator's green-light on the job-store fix: - #5 (464dfc2) declares cupy-cuda12x[ctk] on the gpu extra and takes uv out of the render path (sys.executable -m waterland.cli), retiring the runtime prune trap at the source. - #6 (b72425b) rehydrates the job index from the data volume at startup, fixing the unbounded store growth reported from this side. Verified after the update rather than assumed: healthy on backend cupy; /api/jobs went 1 -> 16 against 16 directories on disk, so API and volume agree for the first time; nothing wrongly reclaimed, correct since 16 is under RETAIN=40 and adoption only makes them visible; a real 256^2 plate render completes warm, so the kernel-cache volume survived the image swap. A subsequent render took both counts to 17. The image keeps its explicit [ctk] install and UV_NO_SYNC/UV_OFFLINE pins even though both are now redundant. The header requirement is a property of this slim base, not of the upstream extra, and the cost is measured rather than assumed: uv sync satisfies it first, so the line reports "Audited 1 package" and adds 0.3s to the build. The env pins are now cheap defence-in-depth against any future path that re-enters uv. Docs corrected in place: the README's upstream-finding section is now a resolved-finding record, and the two "bounded ~500 MB" claims say which commit made that bound hold across restarts rather than only within a process. Comment-side changes pushed to the live compose dir; no restart was needed for them. |
||
|
|
a2b6bf409e |
memory: waterland-studio upstream fixes landed, container stays pinned
waterland-dev merged PR #5 (main now 464dfc2), fixing both landmines at source: the gpu extra declares cupy-cuda12x[ctk], and the renderer spawns sys.executable -m waterland.cli instead of re-entering uv mid-job. The running container deliberately stays on 8025366. Its own [ctk] install and UV_NO_SYNC/UV_OFFLINE pins already neutralise both defects, so a rebuild would buy reliability that is already present — and the project is in wind-down. Both guards are kept rather than dropped: the header requirement is a property of this slim image, not of the upstream extra, and the uv pins are now cheap defence-in-depth against any future path that re-enters uv. Also records waterland-dev's confirmation of the unbounded job-store growth and the operator's green-light on their startup-rehydrate fix. That PR merging is the rebuild trigger: one update.sh run lands the rehydrate and 464dfc2 together. Marks the inbox drained. |
||
|
|
bc3aada73a |
memory: snapshot — .internal DNS live, waterland containerised, homepage themed
Captures a long infra session: fleet *.internal DNS (git-sourced, 42 names, three resolvers including a new colo one), waterland studio containerised on irv-ml1, Homepage cleaned up and themed with Australis Skyfall over an Arbo-generated background, and four unmanaged stacks adopted into stacks/. Four detail files added. Auto-archived 4 entries to archival-memory.md (Recent decisions 2, Tried and abandoned 2); 5 held back by the open-deferred guard rather than moved. Also records three operator-owned open items: the colo DNS repoint, the static-v6 convention, and the deliberately belayed AI-tab Dormant regrouping. |
||
|
|
b8003c73ae |
feat(dns): fleet .internal naming — git-sourced, agent-managed, three resolvers
Names for fleet hosts so addresses stop needing to be memorised. Built because IPv6 makes that hopeless — and, more to the point, because v6 addresses are derived rather than assigned, so they cannot reliably be written down once and trusted either. dns/internal.yaml source of truth: 38 hosts + 4 service aliases scripts/dns-sync.py reconciles AdGuard resolvers against it stacks/adguard-ana/ the colo's resolver, which did not exist Naming is <host>.<site>.internal with sites ana/esh/nh3 (operator's call). .internal is ICANN-reserved for this; .local is reserved for mDNS, which is why searxng.pfi.local was a collision that merely happened to work. Same posture as deploy-stack.sh: file is intent, resolvers are derived state, you see a diff before anything changes. Every name is published to every resolver, so the site label says where a host IS, not who knows about it. Two properties that matter: - Authority is scoped to the ZONE, not the resolver. ESH carries hand-made esteban.net rewrites predating this; they are read, ignored and preserved. Resolver-wide authority would have silently deleted them. - Within .internal it IS authoritative, so UI-added names get removed. That is the point — one place to look. Colo gap closed: ana-docker had no resolver at all (hosts went straight to 1.1.1.1). Its AdGuard runs API on 8053 because 8080/3000 were taken, so the port is carried per-site in the yaml rather than assumed by the script. It ships with no blocklists — a false positive on a server network breaks service-to-service calls for no upside. Auth is a dedicated infra-ops AdGuard user, not the operator's account, password vaulted at nh3-dev/adguard-infra-ops-password. Pre-change configs backed up on each resolver. Both resolvers stayed answering across the restart. searxng.pfi.local -> searxng.ana.internal, with the old Host() kept alongside so nothing breaks mid-migration. matrix.pfi.local deliberately NOT migrated: a Matrix server_name is baked into every user id, room id and signing key, so renaming it rebuilds the homeserver's identity rather than changing a DNS name. The v6 column is empty and correct — no fleet host has a global v6 address yet. The file documents why addresses must be pinned statically before they go in, since a record that silently stops matching is worse than no record. |
||
|
|
8189076daf |
docs(waterland-studio): claude-bot read grant wired, and an upstream store-growth finding
Operator granted claude-bot read on vh/waterland; verified scoped correctly (admin false, push false, pull true). Token is on irv-ml1 at /root/.config/waterland-studio/git-credentials, 0600 root-owned, wired as a REPO-SCOPED credential helper rather than a global one, and .git/config holds no token so the remote stays clean in any diff or backup. The vh site-admin token was used only for the initial clone and the grant itself and was never written to disk on that host. update.sh now runs end to end: fetch, rebuild, recreate, health. Verified the kernel-cache volume survives a recreate (warm 256^2+anim render 6.6s straight after) and the job store survives with all 16 directories intact. Records an upstream finding surfaced by that check: JobStore._jobs is memory-only and nothing scans the data dir at startup, so after a restart the API lists only new jobs while old ones persist on disk — cosmetic — but the RETAIN=40 eviction only sees in-memory jobs, so restart-orphaned directories are never reclaimed. The handover's ~500 MB bound holds per process lifetime, not across restarts. Reported to waterland-dev; upstream's call to fix. |
||
|
|
a2b5b58eee |
feat(waterland-studio): containerise the GPU render service on irv-ml1
Replaces a bare nohup on irv-ml1:8410 that would not have survived a reboot, handed over by waterland-dev. Tracks vh/waterland @ main (PR #4 merged; main HEAD is exactly the pinned 8025366). Build context is a checkout at /opt/waterland-studio/src, deliberately OUTSIDE the compose dir — deploy-stack.sh rsyncs stacks/<stack>/ with --delete and would otherwise eat it. The Dockerfile is passed out-of-context. Three landmines, all measured: 1. Both uv extras are load-bearing at build AND run. jobs.py shells the renderer out as a literal with no --extra flags, so uv would re-sync at runtime and prune cupy — silently dropping to the numpy path at ~21x wall time. UV_NO_SYNC pins it; UV_OFFLINE makes any failure loud instead of quietly slow. 2. cupy needs CUDA HEADERS for its NVRTC compile, not just the driver and the wheel's runtime libs. The host has a system CUDA toolkit so the nohup process found them by accident; a slim image does not, and every render died 1.7s in with 'Failed to find CUDA headers' printed through argparse's usage banner — which reads like a CLI bug, not a missing toolkit. Fixed with cupy-cuda12x[ctk] (hundreds of MB, vs ~6 GB for a -devel base). 3. The A6000 is host device 1 but container device 0, since compose exposes exactly one GPU. CUDA_VISIBLE_DEVICES_TARGET=0 inside; copying the host's value selects a device that does not exist. /root/.cupy is a volume because the NVRTC compile costs ~17s: verified at 23.3s cold vs 6.1s warm, and re-verified across a restart (23.2s on a fresh cache volume, 6.0s once populated). Warm 256^2+anim beats the 7.4s recorded against bare metal, so containerising cost nothing. Job store seeded with the 4 jobs from the displaced instance. Serial by design (one replica, one card) and unauthenticated, so it stays LAN/WireGuard-only. |
||
|
|
df68dd2753 |
style(homepage): tone the stat values down, run the aurora through the page
Two operator corrections in one pass. Stat values overshot: the previous commit took them from font-thin 13px to bold 22px in heading white, which went from whisper to shout. A stat only has to out-rank its own label, not the service name above it — now --text-md at medium weight in cyan, which clears the label but sits below the card title where it belongs. Colour lift, staying inside the system rather than around it: Skyfall names Aurora (blue, cyan, green) the PRIMARY families, 'used generously, in that order', while Dawn (amber, red, violet) is semantic-only. So group markers now cycle blue -> cyan -> green down the page — icons at full strength, names at 0.72 — service icons take a single cool wash, header resource icons go cyan, and latency tags move to the info family so 'how fast' stops looking like 'is it alive'. No Dawn colour is used decoratively anywhere. Also fixes selectors that never bound: Homepage emits docker-status-<state>, not status-<state>, so the green pills up to now were stock colouring rather than this file. Both forms are matched and the trap is commented. README records the iteration loop that would have caught the overshoot: CSS is served per-request, so it needs a reload, not a recreate and not the layout warm-up — and candidate CSS can be injected into the live page for a seconds-long feedback loop instead of a 10-minute one. |
||
|
|
f38cf69fe4 |
fix(homepage): invert the widget stat hierarchy — numbers lead, labels recede
Stock Homepage builds each stat as a font-thin (weight 100) 13px value above a font-bold 12px uppercase label, so the number you actually came to read is the quietest thing in the card while its label shouts. Skyfall's rule is that hierarchy comes emphatically from weight AND size, and that numbers are data. Value now renders at --text-xl bold in tabular mono at --text-heading; label drops to a --text-2xs tracked eyebrow at --text-faint. The well itself moves to --surface-input, one step DOWN from the card it sits on, so stats read as inset data rather than as another floating surface — a recess, so it takes the hairline without the shadow. Also drops .service-block from the generic .service-tag rule, which was what pinned every number to --text-2xs in the first place. Visible on Plex, Jellyfin, PaperlessNGX and Uptime Kuma. |
||
|
|
dc3e47b3a2 |
feat(heretic2-charrp-reasoning): track the NVFP4+MTP reasoning seat
The char-rp-reasoning seat on ana-ml2 GPU0 — NEO-CODE Heretic2 27B at modelopt NVFP4 with a grafted BF16 MTP head, ~77 tok/s via qwen3_5_mtp spec-decode, replacing the retired GGUF seat. It had been running untracked. Includes conf/mtp-workaround/sitecustomize.py, which is not optional: vLLM 0.24.0 does not propagate modelopt exclude_modules to the spec-decode DRAFT model, so the BF16 MTP head gets quantized and the engine dies at load. The shim force-skips mtp.* in is_layer_skipped. Both the mount and PYTHONPATH are load-bearing. Adds the two files house convention expects and the directory lacked: a .env.example naming every knob (all values are the compose defaults; the live host overrides only the three VRAM ones) and a README that points at docs/runbooks/heretic2-nvfp4-mtp-seat.md rather than duplicating it. No secrets: API_KEY is empty by default and the real .env stays on the host. |
||
|
|
45c1995d7a |
feat(homepage): Australis Skyfall theme + Arbo-generated aurora background
Replaces the previous theme attempt, which was built on a misread: the ask was to use Arbo as an IMAGE-GEN ENGINE for the background, with the operator's Australis Skyfall design system supplying the palette. theme/ holds the source — colors/layout/typography vendored verbatim from the Skyfall handoff bundle, Supreme 400/500/700 woff2, the Homepage bindings in skyfall.css.in, and build.py which inlines fonts + tokens into conf/custom.css. custom.css is GENERATED; edit the .in file and rebuild. The build exists because Homepage serves only custom.css and custom.js out of its config dir, so a @font-face pointing at a vendored woff2 would 404 — the face has to arrive as a data: URI. The background image takes the other route: /app/public/images is a real static route, so compose.yaml now mounts images/ there read-only and settings.yaml points at /images/. Bindings map Skyfall's semantic layer onto Homepage's DOM: Sea surfaces, the depth recipe (hairline AND two-layer shadow, never one alone), uppercase eyebrow group headers, the sanctioned accent-rail on the active tab rather than a glow, and semantic status colour so a green pill means the service is actually serving. Background generated by Arbo (irv-ml1:8201) workflow t2i-ui-background, job 13f0891f4e42, seed 26, flux2-klein-9b, 2048x1152 — abstract, no subject, cool-temperature aurora. 1.6 MB PNG -> 22 KB WebP. Two deviations are documented rather than hidden: Skyfall forbids imagery behind body text (held at opacity 30 as mitigation), and service icons stay full-colour vendor logos. NOT DEPLOYED — live still runs the old theme. Prototype on :5199. |
||
|
|
c3de7dbd58 |
feat(homepage): add Arbo 'Raven' theme to custom.css; correct the tab-bar note
Ports Arbo's design tokens (irv-ml1:8201) into conf/custom.css — flat raven ink #021425, card surface #112333 on #1B2E3D borders, Manrope, and mint #2FFC89 reserved for signal so a green pill means the service is actually serving. Values read off Arbo's running :root custom properties rather than sampled from a screenshot. CSS rather than settings.yaml because Homepage's color: setting only accepts built-in Tailwind ramps. NOT YET DEPLOYED — live still runs the stock theme pending an A/B decision. Prototype is at 10.0.50.45:5199; shots in ~/booth-data/homepage-cleanup/. Promoting it also means dropping the background: block from settings.yaml. Also corrects the previous commit's tab-bar claim. It is not a fixed few minutes of warm-up: a fresh container was still tab-less at 4m30s twice, and recovered on its own about an hour later. Cause remains unpinned; the README now records the measured timing and the four ruled-out causes. |
||
|
|
42c594c29f |
fix(searxng,seafile): repair wget healthcheck argv, restore seafile after 3-month outage
searxng: the healthcheck passed '--tries' and '--spider' as separate argv entries, so wget consumed '--spider' as the value of '--tries'. Spider mode never engaged and every 30s probe downloaded the response to disk; the container's working directory had accumulated 295,287 healthz.N files since April, and the directory scan to pick the next free filename is what intermittently blew the 10s timeout and flapped the dashboard card to UNHEALTHY. Restored '--tries=1'. The junk was in the writable layer, so the recreate cleared it. Now healthy, fails=0, 200 in 0.16s. seafile: none of the three services declared a restart policy, so Docker defaulted them to 'no'. The daemon stopped all three within 200ms on 2026-05-06 and nothing brought them back — a three-month outage whose only trace was an EXITED card. Exit 255 is what a container ignoring SIGTERM reports when the daemon stops it, not a crash. Added restart: unless-stopped. Stack is back up; mysql gates on its healthcheck as designed and seahub started without the race. 302 -> login page. Both stacks were running unmanaged on ana-docker and are now tracked here. homepage: AI tab reordered by clickability per operator — chat frontends, ComfyUI and the control plane on top; vLLM /docs seats and TTS endpoints below. Corrects the previous commit's UNRESOLVED tab-bar section: it was warm-up time after a recreate, not a defect. |
||
|
|
9d92c4bd21 |
fix(homepage): pin UltraSeedbox to one tab, dedupe Uptime Kuma, size columns to members
UltraSeedbox had no layout: entry, and Homepage renders an untabbed group on every tab — eight full-width bookmark bars repeated four times. Pinned to Main with a row layout. Uptime Kuma rendered twice: a manual services.yaml block under Monitoring plus homepage.group=Apps on the container. Dropped the manual block, moved the label to Monitoring, added homepage.siteMonitor. Adopted the previously unmanaged uptimekuma stack into stacks/ so the label is version-controlled. Column counts declared more columns than groups had members, leaving the last row of several groups mostly empty. Columns now track member counts. Also records an UNRESOLVED regression: since the container was recreated the client render has lost its tab bar, wallpaper and i18n. Ruled out the config changes (committed pre-cleanup config reproduces it) and v2.0.0 (v1.13.2 reproduces it). Server HTML still carries the tab markup, so the loss is client-side. Details in the stack README. |
||
|
|
084ad924f0 |
memory: snapshot — ESH fiber live, esh-pve-nas migrated+patched, v6 mapped
Session captured: ESH cut over to Cityside 2Gb symmetric fiber and was fully provisioned on it; esh-pve-nas completed its ZFS-root migration off the USB DOM and took its 225-package security backlog with the reboot deferred; ESH<->colo IPsec was rebuilt as a dialup tunnel with NAT-T after CGNAT broke the statically-pinned one; IPv6 was mapped across all three sites. Auto-archival fired at the soft cap: 7 entries moved to archival-memory.md (Recent decisions 3, Tried and abandoned 4), all verified-complete arcs, with two detail files moved and removed. The remaining pre-Aug-05 entries were held back by the open-deferred-work guard, so the index stays slightly over cap at 313 lines rather than losing a live pointer. Also records the one self-inflicted outage of the session (missing --make-rslave on a chroot rbind) and that three long-dead things surfaced incidentally: pvestatd down 82 days, a vzdump hung 126 days, and a VM sitting in prelaunch for four months. |
||
|
|
34d3f42bf5 |
docs: park the BGW210 v6 work pending the Device Access Code
Operator will retrieve the BGW210 Device Access Code from the NH3 office and vault it, after which the IPv6 LAN settings page can be driven remotely. Parked on the henge as reclaim-nh3-s-7-unclaimed-ipv6-64s-from-the with everything needed to resume cold: the verified facts about the /60 split and the seven unclaimed prefixes, why AT&T cannot fix it, the exact page to start at, what to look for in priority order, the other settings pages behind the same login, and the wpa_supplicant fallback with its warning about modifying NH3's only uplink. Suggested vault path unifi/bgw210-nh3-device-access-code, matching the existing unifi/* credentials. |
||
|
|
57e080319b |
docs: NH3 v6 root cause is the BGW210, and seven /64s are unclaimed
Operator suggested checking the BGW on its 192.x management address, which turned out to give the whole picture from unauthenticated status pages. The CPE is a BGW210-700 on firmware 4.28.7 at 192.168.1.254. AT&T does hand it a /60 -- c110 through c11f. The BGW keeps c110-c117 for itself and re-delegates up to eight individual /64s on c118-c11f, top-down. Our UDM holds c11f, delegation number eight. So the earlier conclusion that AT&T only grants a /64 was right about the symptom and wrong about the cause. Seven further /64s are available and simply never solicited, because UniFi exposes a single wan_dhcpv6_pd_size integer with no field for how many prefixes to request. The documented workaround is repeated -P flags to dhclient, which the UniFi UI cannot express. This also settles that an AT&T ticket cannot help: the rationing is CPE firmware behaviour, not provisioning. Records the two real options -- accept one /64, or bypass the BGW entirely with wpa_supplicant EAP-TLS on the UDM to negotiate the full /60 -- with the warning that the latter modifies NH3's only uplink and needs a planned window. |
||
|
|
1b6c26ce58 |
docs: AT&T PD is a hard /64 at NH3, tested on the wire
AT&T support guessed 'I do not believe att will do that' from a DNS provisioning desk. The guess was correct, but it needed proving rather than accepting, so the UDM solicited DHCPv6-PD at /48, /56 and /60. All three returned the same single /64. This is not a case of nobody having asked -- the ask was made three ways. Proven by temporarily enabling PD on nh3-iot, the only NH3 VLAN with zero clients, then setting ipv6_pd_prefixid to 0, 15 and 16. All three returned an identical c11f prefix, which only happens when exactly one /64 is delegated; with a larger block the prefix-id moves the LAN within it. Records a mistake worth not repeating: I first read the gap between the WAN address (c110) and the delegated prefix (c11f) as evidence of a /60. It is not -- AT&T assigns those from different parts of their pool and the spread means nothing. NH3 was fully restored afterwards, with rollback artifacts kept. Also captures the concrete ask for AT&T Business, phrased as something their provisioning team can verify against their own DHCPv6 logs, and the consequence if refused: NH3 can host exactly one v6 segment against ESH's 256. |
||
|
|
fddf7f587f |
docs: record the pending Cogent IPv6 provisioning request for the colo
Operator opened a ticket with Cogent for v6 at Anaheim, which closes the one thing tonight's investigation could not resolve from our side. Captures the diagnosis so the ticket has evidence behind it: a single RA in 90 seconds of sniffing wan1, from fe80::ea0a:b9ff:fe3b:2c16, proving an IPv6-capable router sits one hop away on the circuit terminating 38.120.12.42/29 -- but SLAAC obtained no global address across multiple RA intervals and ping6 to Cloudflare and Google both returned 100% loss. Router present, circuit unprovisioned. Also records that the FortiGate v6 config was fully reverted after testing, the FortiOS gotcha that SLAAC is 'set autoconf enable' rather than an ip6-mode, and the ask to make when it lands: a /56 or better, since NH3 only receives a single /64 from AT&T. Notes the consequence worth planning around -- once provisioned, the colo becomes the only site with both a static public v4 and routable v6, which makes it the natural v6 hub given ESH is CGNAT'd and NH3 is prefix-constrained. |
||
|
|
fb91ea759e |
docs(pfi): lesson 10 -- v6 collapses two exposure controls into one
Operator's framing, and it is a better argument than the terminology correction that preceded it. Under v4, exposing a host needed two affirmative acts -- a DNAT and an accept rule -- so missing either left the host dark. There is no v4 misconfiguration that exposes an internal host by accident. NAT was load-bearing security whether or not anyone designed it that way. v6 removes the first control entirely. The path exists inherently, so the firewall is the only thing left, and the failure mode inverts from fail-closed to fail-open. Rule-ordering slips, rulesets that silently match only one address family, new VLANs added without policy, and re-delegated prefixes unmatching address-literal rules all become exposure events rather than no-ops. Records the practical consequences: key rules on interface/zone rather than address literals, treat enabling v6 on a segment as requiring policy to exist first, and verify default-deny from off-net rather than by reading the ruleset -- which is lesson 3's assert-the-effective-value discipline applied to firewall policy. Also corrects my own claim from the previous commit that the pending firewall pass was 'smaller' than I had implied. It is not smaller, it is different in kind. |
||
|
|
707a8cbcce |
docs: correct addressable vs reachable in the ESH v6 entry
I wrote that enabling SLAAC would give LAN devices 'globally reachable addresses'. Wrong word, and the wrong word in a persistent-memory entry a future session inherits as fact. Addressable is a property of the address. Reachable is a policy decision the firewall makes. v6 removes NAT; it does not remove the firewall, and treating those as the same thing is exactly how v6 gets mischaracterised as automatic exposure. Records the operator's position while correcting it: no 1:1 inbound pass-through. The pending firewall-policy pass is about writing explicit default-deny inbound rules per v6 segment, not about deciding what to expose. |
||
|
|
8be8a51437 |
docs: re-gloss esh-iot, re-spell esh-mgmt
esh-iot keeps the identical eight digits -- 4DBAD107, rendering 4dba:d107 -- and only the reading changes: 4 is FOR rather than A, so it parses 'FOR DA BAD IOT', which describes what the segment is actually for. esh-mgmt genuinely changes: 115D:B055 becomes 15DA:B055. The leading I is dropped and DA is spelled in full, giving 'IS DA BOSS' with the network as subject rather than speaker. Still eight digits. DA written out needs no substitution since D and A are both native hex; spelling it as a single D the way esh-iot does would have yielded seven digits and broken the house pattern. |
||
|
|
18c683b399 |
docs: reserve 4411:DBAD for a future DMZ
'FOR ALL DA BAD' -- 4=FOR, 411=ALL, D=DA, BAD=BAD. Eight digits, house style, renders 4411:dbad. No DMZ network exists on the ESH UDM today; this is a name claimed against the day one is built. Pairs deliberately with esh-iot's 4DBA:D107 -- IoT is 'a da bad', the DMZ is 'for all da bad', which is the correct relationship between the two segments. |
||
|
|
959bb6ee05 |
docs: server network gets 4411:B105 -- ESH v6 naming scheme complete
'FOR ALL BIOS' -- 4=FOR, 411=ALL, B105=BIOS. Eight digits like the rest. Completes the set. All six ESH networks now carry an 8-hex-digit phrase in a consistent first-person/declarative voice: Default 4BA5:3417 A BASE FOR IT esh-mgmt 115D:B055 I IS DA BOSS esh-server 4411:B105 FOR ALL BIOS esh-userland CAFE:4411 CAFE FOR ALL esh-iot 4DBA:D107 A DA BAD IOT esh-cameras 1533:FACE5 I SEE FACES Still a documentation convention rather than wire-level configuration -- UniFi has no static-v6 client assignment and the gateway address is platform-fixed -- but these are the values to use whenever ESH LAN v6 is switched on and hosts get hand-assigned addresses. |
||
|
|
a264e001ae |
docs: default network settles on 4BA5:3417
Same phrase, 'A BASE FOR IT', but written as a plain 8-digit string rather than forcing the article into its own group. 4=A, BA53=BASE, 4=FOR, 17=IT renders as 4ba5:3417 -- two groups, matching every other network in the scheme, with the words straddling the colon exactly the way 4DBA:D107 does. Corrects the previous commit, which claimed this needed nine digits and a third group. It is eight, and always was. |
||
|
|
e5bba048c8 |
docs: revise default network to A:BA53:0417
'A BASE FOR IT' -- the article makes it a full sentence, matching the voice of the other five. BA53 uses 3=E rather than the 5E spelling used in the previous BA5E version. Nine hex digits rather than eight, so unlike the others it does not fit two groups and renders across three as a:ba53:0417. |
||
|
|
309a240fa8 |
docs: default network gets BA5E:0417
'BASE FOR IT' -- BA5E=BASE, 4=FOR, 17=IT. The foundation segment everything else hangs off, which is what the default network is, and it doubles as 'base for IT'. Note the trailing group zero-pads: it renders as ba5e:0417, not ba5e:417. |
||
|
|
e58cfde7fd |
docs: userland network gets CAFE:4411
'CAFE FOR ALL' -- CAFE, 4=FOR, 411=ALL. Eight digits like the others, and the only one so far that splits on its own phrase boundary, so it renders legibly as cafe:4411. Bonus reading: 411 is US directory assistance, which is a fitting second joke for the segment the humans actually live on. |
||
|
|
ab8481907d |
docs: iot network gets 4DBAD107
'A DA BAD IOT' -- 4=A, D=DA, BAD=BAD, 107=IOT. Eight hex digits to match the cameras and mgmt picks. Worth noting it renders as 4dba:d107, so unlike the other two the phrase does not split on its word boundaries and reads as noise unless you know it is there -- which is arguably right for the untrusted segment. |
||
|
|
805fa6ff22 |
docs: mgmt network gets 115D:B055
'I IS DA BOSS' -- 1=I, 15=IS, D=DA, B055=BOSS. Eight hex digits like the cameras pick, so it renders as 2607:73c0:402:1d??::115d:b055 with the same two-group split and room for host numbering. Pairs structurally with 1533:FACE5 on cameras: both eight digits, both first-person, and the network that is actually in charge gets to say so to the one that is merely watching. |
||
|
|
35e7ecbadb |
docs: cameras network gets 1533:FACE5
'I SEE FACES' -- 1->I, 5->S, 3->E, 3->E then FACES. Eight hex digits splitting cleanly across two groups, so it renders as 2607:73c0:402:1d00::1533:face5 with room left for host numbering. Still a documentation convention rather than anything on the wire, per the constraints recorded in the same entry, but this one is good enough that it should survive to whenever ESH LAN v6 actually gets switched on. |
||
|
|
fe3d765873 |
docs: record the ESH IPv6 naming scheme as a docs convention, not wire-level
Picked six hexspeak names for the ESH LANs during a wind-down moment (FACE/B055/B105/CAFE/DEAD/BASE), then checked whether any of it could actually land on the wire before implementing anything. It can't, for three independent reasons: a network's only nameable slot is its /64 prefix id, which is 2 hex digits and can't spell a 4-char word; the gateway's own address is fixed at ::1 by the UniFi platform with no field to override it; and UniFi has no IPv6 equivalent of use_fixedip/fixed_ip, confirmed directly against the client schema, so individual devices can't be pinned to a chosen v6 address either -- SLAAC devices self-assign via EUI-64 or privacy extension. So this stays a documentation mnemonic. Recorded as such rather than implied as something live, since I'd already started suggesting a static-camera-assignment plan that the schema check ruled out. |
||
|
|
50d13f57cb |
docs: park the ipsec_local_ip watcher until ESH fiber is up
The ESH<->colo tunnel is restored and the FortiGate end is permanently address-agnostic, but the UniFi end still needs a literal ipsec_local_ip and so drops on any ESH WAN change -- Cox reclaiming WAN1, the fiber cutover, or a DHCP renewal. Operator's call not to build the self-healing watcher yet, which is right: it would be written against the 5G failover address, which is about to be replaced, and the fiber may reshape the topology anyway. Parked as self-healing-ipsec-local-ip-watcher-for-the-esh with the trigger recorded, plus the follow-up to retire the old ana-to-eshudm tunnel whose distance-10 route would otherwise reclaim traffic if Cox returned on the old address. Records the manual stopgap in persistent memory so the gap is cheap to cover by hand in the meantime: read wan_ip from the UDM's health endpoint, PUT it into esh-ana.ipsec_local_ip. |
||
|
|
8a742f59b8 |
fix(ana-gw): restore ESH<->colo IPsec as a dialup tunnel with NAT-T
The link died when ESH lost its public IP during the fiber cutover. Two independent causes, and the second would have defeated the obvious fix: - phase1 ana-to-eshudm was type static, pinned to 70.181.90.232, an address that no longer exists. - nattraversal was disable, so ESP could not have crossed NAT even with the peer IP corrected. pfi-ana-nh3 shares that setting and survives only because NH3 is publicly addressed, which is why the two tunnels diverged. FortiOS refuses `set type dynamic` on an existing tunnel -- "Cannot change tunnel type once configured" -- and rolled back cleanly, so the fix could not be an edit. Rather than delete and recreate, which cascades into the phase2, two static routes and ten policies, the replacement was built alongside: new phase1+phase2 ana-eshudm-dyn (type dynamic, ikev2, aes256-sha1, dh14, NAT-T on, PSK read from the ESH UDM API so neither side needed a new key), static route id 10 at distance 20, and two consolidated multi-zone policies 73/74. The old tunnel is left in place, dead and harmless, as rollback. Verified up: ana-eshudm-dyn_0 97.170.236.56:4500 selectors 1/1 -- the _0 suffix is a dialup child, :4500 is NAT-T, and the address is the carrier's, which is precisely what could never have been pinned. ESH reaches all four colo hosts at 40-56ms, the colo reaches all three ESH hosts, and traceroute drops from eight hops leaking into the carrier network to three hops fully encapsulated. Config was backed up before any write (1.17MB, 36903 lines, off-box). Residual fragility recorded: the UDM's ipsec_local_ip demands a literal address -- empty is rejected as api.err.InvalidPayload -- so it still needs updating when the fiber changes ESH's WAN address. The gateway end is now address-agnostic; the UniFi end is not. |
||
|
|
9407e7f144 |
docs: correct the persistent-memory IPv6 entry to match the evidence
The prior edit missed its anchor and left the over-broad version in place. The entry now separates the two inter-site links rather than treating them as one: Site Magic (WireGuard, NH3<->ESH) survives arbitrary NAT and is proven to; IPsec (colo<->ESH via ana-gw) does not and is currently down, with traffic leaking unencapsulated to the carrier. IPv6 keeps its justification on the IPsec link specifically. |
||
|
|
dec4ba45db |
docs: scope the NAT refutation to WireGuard; IPsec to colo is broken
Correcting an over-generalisation from earlier today. Proving that NAT does not break Site Magic, I wrote it up as "no addressing outcome threatens the inter-site tunnel." That is wrong: the fleet has two inter-site links with opposite NAT behaviour. - NH3<->ESH is Site Magic, i.e. WireGuard. It survives arbitrary NAT, proven live on RFC1918 double-NAT (192.168.200.111) with nh3-dev and nh3-docker reachable at ~40ms. It dials out to NH3's public edge and never needs inbound reachability. - colo<->ESH is IPsec on the ana-gw FortiGate, and it is broken right now under those same conditions. ana-docker, pfi-pve and pbs-ana all fail from esh-pve-nas, and traceroute shows packets for 10.250.x leaving the UDM to the 5G modem and then wandering the carrier network before dying -- not encapsulated at all, so no SA is up and the traffic falls through to the default route. Site-to-site IPsec pins a peer IP and ESH no longer has a routable one. So the IPv6 work keeps its justification, but on the IPsec link specifically rather than on the tunnels generally. Operator caught the over-generalisation. Adds lesson 8 -- a result proven for one protocol does not transfer to another -- and corrects the superseded-claims row rather than replacing it, since the original claim was half right and the halves are the point. Also records my own over-broad claim as its own superseded row. |
||
|
|
40a4121a43 |
docs(pfi): add lesson 7 — test a 'this will break X' premise before building on it
Seeded by the Site Magic / CGNAT premise, which justified a body of IPv6 work and turned out to be false the first time anything actually tested it. The mechanism was discoverable in advance: Site Magic is WireGuard and the far side has a public endpoint, so the NAT'd side dials out and never needs inbound reachability. NAT breaks inbound; it does not break outbound-initiated tunnels with keepalives. Also fills the first row of the superseded-claims table, which is what that table exists for -- the claim is corrected with a date rather than quietly deleted, so older references to it resolve instead of misleading. |
||
|
|
78cc760ef6 |
docs: refute the CGNAT-breaks-Site-Magic premise with a live test
The fleet IPv6 work was justified primarily by the expectation that ESH fiber landing behind CGNAT would break Site Magic on IPv4, making v6 the escape hatch. The fiber cutover provided a free natural experiment and the premise does not hold. Cox was unplugged, ESH failed over to the 5G WAN (already configured failover-only, so this needed no intervention), and the resulting WAN address is 192.168.200.111 -- RFC1918, double-NAT, no inbound path at all, which is strictly worse than the CGNAT that was feared. Site Magic stayed up throughout: all four ESH hosts reachable, ssh and command exec working, 20MB pulled over the tunnel, latency 15ms -> ~46ms as expected for cable to 5G. The mechanism is visible on the device: magic_site_to_site_vpn holds only `enabled` plus a WireGuard keypair, with peer orchestration in the UniFi cloud and no WAN binding of any kind. NH3's edge is publicly reachable, so the NAT'd side dials out and never needs reachability. Consequence: no addressing outcome on the new fiber -- public, CGNAT or double-NAT -- threatens the inter-site tunnel. IPv6 stays worth doing on its own merits but stops being urgent, and stops gating anything. Also worth recording that Site Magic cannot be pinned to a WAN. It rides whichever uplink is active, so the only lever is failover priority -- which moves all site traffic, not just the tunnel. The existing failover-only config on WAN2 already handles a primary-WAN outage correctly and needed no change. |
||
|
|
668b63a398 |
feat(esh-pve-nas): install the 225-package backlog; reboot deferred
pve-manager 8.4.11 -> 8.4.20, corosync 3.1.9 -> 3.1.10-pve2, and kernel
6.8.12-42 staged on the /boot LV. dpkg clean, nothing outstanding for
apt -f install, all PVE services active, cluster quorate, no unapplied
conffiles. Reboot deliberately deferred at operator request, so the host
still runs 6.8.12-13 until a chosen window.
This validates the GRUB fix from
|
||
|
|
0559e12a2d |
docs(pfi): add an ops-lessons playbook for the transferable failures
Sibling to model-quantization-playbook.md, and it exists for the same reason that one does: hard-won lessons were dying inside per-host runbooks where nobody finds them until after repeating the mistake. Six entries seeded from the esh-pve-nas migration, all of which would bite identically on any other host: 1. mount --rbind into a chroot needs --make-rslave, and losing cgroup2 impersonates failing root-disk I/O closely enough that it was misdiagnosed as exactly that. 2. A reboot is not confirmed until the host is observed DOWN; "never rebooted" and "rebooted fast" are indistinguishable otherwise. 3. Assert the effective value, not the presence of a substring. Grep proves presence; only evaluation proves effect. 4. Ask the server who its clients are -- documented dependent lists rot. Plus the corollary that an idle hard NFS mount blocks and resumes, so quiescing means stopping consumers, not always unmounting. 5. The scoped-looking command can be the dangerous one; setting a ZFS cachefile on one pool of three would have stopped the other two from importing at boot. 6. Long uptime hides breakage, and a forced look is worth more than it appears -- one migration surfaced an 82-day-dead pvestatd, a 126-day hung vzdump, a VM in prelaunch for four months, and an undocumented cluster, none of them caused by the work. Carries a superseded-claims table so corrections are dated rather than silently edited, same discipline as the quantization playbook. The ESH runbook now links here so the general rules are reachable from the specific story and vice versa. |
||
|
|
7d27ec9d41 |
feat(esh-pve): upgrade to 8.4.20 and reboot onto 6.8.12-42
171 packages, pve-manager 8.4.11 -> 8.4.20, kernel 6.8.12-16 -> 6.8.12-42, corosync 3.1.9 -> 3.1.10-pve2. dpkg clean, no unapplied conffiles, no failed units, cluster quorate with both nodes visible after the reboot. Adds a reusable pve-node-upgrade playbook (upgrade only -- reboot stays a separate deliberate step, since it has cluster and NFS consequences the playbook cannot see). It guards on quorum and free space, snapshots /etc/pve and friends first, uses --force-confdef/--force-confold, and surfaces any .dpkg-dist files that policy left unapplied so they are not silently ignored. The reboot needed a forced guest stop, operator-authorised after the risk was surfaced. Two obstacles, only one of them ours: - A vzdump had been hung since 14 April -- 126 days, stalled at 0% of 256 GiB -- holding lock: backup on VM 102, which had therefore been sitting in QEMU prelaunch that entire time. Killed by explicit PID; 102 is now cleanly stopped rather than half-alive. - esh-vm-db would not shut down: its guest agent had died and ACPI went unanswered. Most likely ours -- it hard-mounts /mnt/backup from CT 103, which we deliberately left mounted through the NAS reboots. PostgreSQL survived the hard stop. It had checkpointed five minutes prior, so recovery replayed 56 bytes of WAL in 0.02s and came up ready; all four databases present and queryable. That was lucky timing as much as anything -- a hard stop mid-checkpoint on a busy database would not read the same way. The reboot also repaired esh-vm-db, which had silently lost sshd, mongod and its guest agent. All three are back. |
||
|
|
061c4b7712 |
fix(esh-pve-nas): stop the boot default pinning a single kernel
The cutover left saved_entry=pve-zfs-root, a hand-authored 40_custom entry hardcoding /vmlinuz-6.8.12-13-pve. The pending upgrade installs proxmox-kernel-6.8.12-42, which made that a trap with two exits: if -13 were autoremoved the default entry would point at a missing kernel and the host would need console recovery it has no IPMI for; if -13 survived the host would silently keep booting the old kernel, so 161 security updates including a kernel would install and never run. That entry was written as a one-time cutover target. It was never fit to be the standing default across kernel upgrades, and this is remediation of that, caught before the upgrade rather than after. Fix is to stop hand-authoring the ZFS entry: GRUB_DEFAULT=0 boots the first auto-generated entry, which grub-mkconfig regenerates for the newest kernel on every install, and which /etc/default/grub.d/zfs-root.cfg already corrects to the pool-qualified root=ZFS=nvme/ROOT/pve-1. grubenv is cleared so nothing overrides it. The rollback entry stays pinned, which is correct rather than an oversight: it boots the untouched ext4 root on the DOM, whose /boot is never regenerated because update-initramfs writes only to the /boot LV. That kernel genuinely never changes. Also adds a reusable safe-reboot playbook for this host, carrying the constraints that are easy to forget: quiesce the hard-NFS clients first, the other cluster node goes read-only while this one is down (quorum 2, no qdevice), and a failed boot has no auto-fallback and no remote console. |
||
|
|
b84f8a996f |
docs(esh): record the 2-node cluster, and that a dark tile is pvestatd
Two findings from chasing node `pve` showing dark in the UI. esh-pve and esh-pve-nas are a 2-node cluster (esh-pve-cluster, expected votes 2, quorum 2, no qdevice). This was undocumented, and last night's migration rebooted one of the members without accounting for it. Nothing broke -- quorum is intact and both nodes report the same ring id -- but that was luck. Rebooting either node drops the survivor below quorum and makes its /etc/pve read-only until the partner returns. It matters for the pending confirmation reboot and the 225-package upgrade, both of which take a node down. The dark tile itself was NOT last night's doing. pvestatd SEGV'd on 2026-05-28 and had been dead 82 days; the journal has nothing between that crash and the restart today. It is only the reporting daemon, so the node stayed quorate and healthy with all services active and all three guests running the whole time -- the UI simply had nothing telling it the node was alive. Fourth SEGV in that unit's history, so treat a recurrence as expected and consider a watchdog: nothing alerts on it, and the sole symptom is cosmetic enough to go unnoticed for months. |
||
|
|
5f11d1b3cb |
feat(esh-pve-nas): cut PVE root over to ZFS on the mirrored NVMe
Root is now nvme/ROOT/pve-1. The USB DOM keeps the ESP and /boot but is out of the runtime I/O path, so a bus reset can no longer drop root from under a running hypervisor. All five guests healthy, three pools ONLINE, system running, ext4 pve-root intact and unmounted as the rollback with its own kernel and initrd. zfs-import-cache is now the active import path -- the all-three-pools cachefile fix doing its job. The window cost an unplanned outage, and the cause was this repo's own tooling rather than the migration. The staging chroot ran `mount --rbind /dev` and /sys with no --make-rslave. On systemd `/` has shared propagation, so the cutover's `umount -R` propagated back into the live host and removed the real /sys/fs/cgroup, /dev/pts and /dev/shm. With cgroup2 gone systemd-logind could not create a session: ping fine, TCP fine, SSH authentication succeeded, resident daemons kept serving -- and every new exec hung, including /sbin/reboot, so the reboot never ran at all. It impersonates failing root-disk I/O almost perfectly, and I called it as the DOM dying. That was wrong. dmesg had the answer throughout: the DOM attached cleanly with no errors, and the last log timestamp was 12114881s -- 140 days -- meaning this was still the original boot. A down-detector had also never reported the host down, which I read as a fast reboot rather than as no reboot. Fixes and guards: - --make-rslave after every rbind, plus a guard that refuses to proceed while any chroot bind still reports shared propagation. - Confirm a reboot by observing the host DOWN, not by watching for it to come back. Those two states are indistinguishable otherwise. - Blast radius now measured from the server: `ss` inside CT 103 found five NFS clients, not the two documented. The new one that mattered is esh-vm-db, hard-mounted and unreachable by ssh. Left mounted on purpose and it came through read-write. - grub-reboot's one-shot does NOT work here: grubenv sits on an LVM LV which GRUB reads but cannot write, so next_entry survived the boot that consumed it. Steady state is saved_entry=pve-zfs-root with no next_entry. There is no auto-fallback on this host and no IPMI. Recovery needed no console: an idempotent cgroup2/devpts/shm remount landed in the brief windows where exec succeeded. No data was lost, and neither the DOM nor any pool was ever at risk. |
||
|
|
b637947ffd |
docs(park): record the API-key ruling — leave it as is
Operator ruled 2026-08-18 on the 5-character PARK_API_KEY flagged during the v1.0.0-beta.2 deploy: leave it. The henge is LAN/WG-internal and never internet-exposed. Written down so the next audit does not re-raise a question that has already been answered. |
||
|
|
ec1c482bd5 |
deploy(park): stonehenge-park v1.0.0-beta.2 on ana-docker
Operator-directed request from park-dev. Rebuilt from tag v1.0.0-beta.2 (commit 2c258f7) and redeployed; park-data volume preserved (28 items, 9 comments verified present after the recreate). Build source is now exported per-tag to ~/deploy-src/stonehenge-park-v1.0.0-beta.2 rather than overwriting the single mirror, so the previous tag's tree stays on the host as a rollback. The mirror was never a git checkout, so the source comes from `git archive <tag>` against a box that has the repo -- which also leaves park-dev's working tree untouched. Verification, and two things worth writing down: - HEAD 401s on EVERY route, including /healthz and /. So park-dev's suggested check `curl -sI .../ui/assets/favicon.svg` reports a false failure. GET is 200 with content-type image/svg+xml; the packaging is fine and all nine assets are in the wheel. App-wide and pre-existing, not a beta.2 regression -- /healthz predates this release. - /park/due-count returning 0 is not a data-loss signal; it counts what is due now, and /park/due is empty across overdue/today/stale. Items survived: GET /park returns all 28. README now says to check that instead of the due counters. Also corrected the stack README, which still told the reader to build on nh3-docker and verify against 10.100.50.40 -- the host decommissioned for this stack on 2026-08-13 and the one park-dev explicitly asked us not to deploy to. |
||
|
|
c4b2278e7d |
feat(esh-pve-nas): stage the PVE root migration off the USB DOM
Everything but the reboot. Two rerunnable elway playbooks; the host is
still running from the ext4 root and its boot path is byte-identical to
the last 140 days, because grub-install is deliberately held back to the
cutover window.
Phase 1 (esh-pve-nas-stage-zfs-root.yaml): carve a 512 MB /boot LV out
of the 768 MB swap LV, populate it, rsync the 4.3 GB ext4 root into
nvme/ROOT/pve-1, write the copy's fstab.
Phase 2 (esh-pve-nas-stage-bootloader.yaml): ZFS initramfs, grub.cfg,
explicit pve-zfs-root and pve-ext4-rollback entries with stable ids,
grubenv pinned to the rollback so cutover's grub-reboot is a one-shot.
Three landmines the plan did not predict, all caught by verify steps
asserting effective state rather than by reading the plan:
- The /boot LV had nowhere to live. VG pve had 4 MB free and mounted
ext4 cannot shrink; freeing space from root needs a rescue boot, which
costs the one-reboot property. Space came from swap (768M -> 256M).
- The runbook's `zpool set cachefile=... nvme` would have broken the
NAS. Populating a cachefile flips the host from import-by-scan to
import-by-cache, so a one-pool cache leaves ssd and tank unimported --
and CT 103 esh-nas has twelve bind mounts spanning all three pools.
Set on all three instead, verified in the resulting cache.
- update-grub silently emitted a pool-less root=ZFS=/ROOT/pve-1, which
boots to an initramfs prompt. Debian's 10_linux builds ${rpool}${bootfs}
and rpool comes from grub-probe --target=fs_label, which returns empty
because GRUB's ZFS reader cannot open a pool with encryption,
large_dnode and zstd_compress -- the same feature set that forced /boot
to stay ext4. The probe failure is swallowed by `2>/dev/null || true`.
Fixed with a /etc/default/grub.d drop-in plus explicit menu entries.
The transferable lesson: the original verify grepped for the correct
root= string appearing somewhere in grub.cfg, which passes while every
menu entry is still broken. Assert the effective value, not the presence
of a substring.
|
||
|
|
d3e1cc4a41 |
docs(lobe-chat): TTS works with zero client-side settings now
infra-ops aliased tts-1, tts-1-hd and gpt-4o-mini-tts onto ext-tts's upstream and extended the lobe-chat-esh key allow-list 20 -> 23 models, so the manual "set the TTS model to ext-tts, per browser" step this file described a few hours ago is obsolete. Lobe's stock three-field payload now returns 200 audio/mpeg — verified from the host with this stack's own .env. Adds the coupling that the fix introduces: the three new names are independent LiteLLM DB rows carrying their own copy of the upstream URL, so a future repoint of ext-tts must move all four or stock clients land on a dead engine without any error on the gateway side. Comment/doc only — no functional change, no redeploy. |
||
|
|
637ed3bd89 |
memory: Lobe TTS fixed by aliasing stock OpenAI model names at the gateway
Lobe's TTS had never worked. It sends model:"tts-1" and LiteLLM resolves the model name before routing, so it 403'd against the scoped key's allow-list and never reached :8198 -- our belief that an unknown model routes to the gateway default was true of the gateway and false of the LiteLLM path, which is what hid it. Fixed at the gateway rather than the client: tts-1, tts-1-hd and gpt-4o-mini-tts aliased to the same upstream as ext-tts, and added to the lobe-chat-esh allow-list. Verified with Lobe's exact payload on Lobe's own key. The obsolete 'one-time human UI pass' follow-up is dropped. Banks two durable facts: those aliases are independent DB rows that must move if ext-tts repoints, and the infra-ops key has admin rights for /model/new and /key/update so this class of work does not need sk-corvid. |
||
|
|
356752d99c | memory: snapshot — heresy gen seat live, irv-ml1 cleared, homepage repo'd, esh-pve-nas DOM planned | ||
|
|
8ddc87c852 |
docs(esh-pve-nas): record the blocked-patching driver and the upgrade ordering
The operator-visible symptom is that PVE cannot be updated on this box for lack of room. Measured: 225 packages pending, 161 carrying deb12uN/Debian-Security bumps including ssh, against esh-pve's 8.4.14 versus this host's 8.4.11 and 20 weeks of uptime. Records the ordering explicitly -- migrate first, upgrade after. The pending set includes proxmox-kernel-6.8.12-42-pve-signed, roughly 250 MB of kernel plus initramfs landing in /boot which is on root with 1.3 GB free. Unpacking 225 packages including dpkg and perl into that headroom risks filling the disk mid-transaction and wedging dpkg on a hypervisor running five guests. Notes the apt archive-dir redirect as a partial escape hatch if patching cannot wait, and that zfs-initramfs 2.2.8 is fully capable of root-on-ZFS so there is no need to upgrade ZFS before migrating. |
||
|
|
3e311756d7 |
docs(esh-pve-nas): split boot from root instead of reinstalling
Operator's proposal, and it is strictly better than the reinstall plan. Boot and root do not have to share a device. Keep the ESP and /boot on the DOM as ext4 -- so GRUB never has to read ZFS, which matters because the nvme pool has encryption, large_dnode and zstd_compress enabled and GRUB cannot read those -- and move root to nvme/ROOT/pve-1. The initramfs imports the pool and pivots. What this buys over the reinstall: the nvme pool survives, so no guest migration, no export/import of ssd and tank, no reinstall. Downtime is one reboot rather than half a day. Rollback is a GRUB menu entry, because the ext4 root stays on the DOM untouched. And it retires the actual top risk -- with root on NVMe, a USB bus reset mid-run no longer takes the running system down; the DOM becomes read-mostly, written only on kernel updates. Preconditions verified and already met: UEFI with grub-efi, zfs-initramfs 2.2.8-pve1 installed with 76 ZFS files already in the running initrd, root only 4.3 GB to copy, swap negligible against 125 GB RAM. Two traps recorded: canmount=noauto on the root dataset or ZFS tries to mount over the running root; and cachefile is currently none with a 0-byte zpool.cache, so the pool imports by scan today and must be given a cachefile before the initramfs is rebuilt. The reinstall plan is retained as the fallback. |
||
|
|
2275e11be0 |
docs(esh-pve-nas): plan the migration off the USB DOM; flag the NFS blast radius
PVE root on esh-pve-nas is a USB Disk-on-Module: 6 GB ext4 with the host's only ESP. A DOM is SLC/pSLC so wear is not the driver -- the problems are that it is on the USB bus (a reset drops root under a running hypervisor), has no headroom, and is unmirrored while 928 GB of mirrored NVMe sits 96% empty. Runbook targets a fresh PVE install to ZFS RAID1 across both NVMes. In-place conversion is unsupported, and adding an ESP to the existing NVMes is impossible -- both are whole-disk ZFS members with 1.7 MiB free and proxmox-boot-tool manages nothing today. The headline risk is not on the host being rebuilt: CT 103 esh-nas IS the NAS at 10.0.50.50, and both esh-docker-vm and esh-pve mount it hard. Taking this box down stalls esh-pve's storage layer and wedges esh-docker-vm into the D-state whose only remedy is a host reboot -- the incident shape already on record. Quiescing those clients is step one of the window, and the README now warns against casual reboots. Config snapshot captured off-box to nh3-dev (0600) with /etc/pve, network and fstab config plus zpool/zfs/disk-by-id/guest state; the newest on-disk copy before this was June 2024. |
||
|
|
ca8c0a318e |
docs(lobe-chat): correct the TTS notes — the deploy's TTS never worked
Two claims in this stack's docs were reasoned from the wrong hop, and one of
them hid a dead feature since deploy. Re-measured from esh-docker-vm against
the live .env:
1. "An unknown `model` routes to the gateway default" — true of :8198, false of
the path Lobe takes. LiteLLM resolves the model name first, so Lobe's default
`tts-1` returns 403 (`key not allowed to access model`) and never reaches the
gateway. `ext-tts` returns 200 + audio. The endpoint inheriting
OPENAI_PROXY_URL is necessary but not sufficient: Settings -> TTS -> OpenAI
TTS model -> `ext-tts` is a required one-time step per browser, and removing
it needs a LiteLLM alias plus a key allow-list entry (both master-key, so
infra-ops).
2. "`response_format: mp3` ... set it in the UI" — not possible. Lobe's OpenAI
TTS client sends `{input, model, voice}` and nothing else (server bundle
chunks/29685.js), so format is not selectable from this stack at any level.
The deploy gets the fleet gateway's default (WAV, ~23.5 MB for a 245 s turn),
relabelled `audio/mpeg` by LiteLLM. That is tts-dev's fence, not this one's.
Comment/doc only — no functional change, so the host copy needs no redeploy.
|
||
|
|
d1f4f1cb96 |
docs(homepage): record the ALLOWED_HOSTS fix, .env lockdown, and verification method
HOMEPAGE_ALLOWED_HOSTS now carries the IP:port form; direct access to http://10.0.50.45:5100/ returns 200 and the host-validation errors are gone from the container log. The .env was mode 644 holding the Plex and Jellyfin API keys; now 600. It is root-owned, so editing it needs the infra-ops identity -- lkraven has only password-sudo on that host. Also records that Homepage renders client-side, so grepping the served HTML to verify a config change is the wrong instrument (it gave a stale prerender and then an empty page). GET /api/services is the honest check, and config changes need a recreate rather than a restart. |
||
|
|
c5beeac32d |
feat(homepage): bring the fleet dashboard under version control
Homepage on esh-docker-vm:5100 was the one stack whose config lived only on the host, edited in place. Its version history was six hand-rolled services.yaml.bak-* files. Now canonical here and deployed with deploy-stack.sh like everything else; the .bak files are gone. Corrections from the audit: - ANA-Firewall described a 'Fortigate 81F'. It is a FortiGate-80F running FortiOS 7.2.10, verified live against the device. - NH3-Ansible pointed at 10.100.50.42 as an 'Ansible control node'. That host is nh3-extdev, the manager/external-dev successor after nh3-ansible was retired. Renamed and re-described. - Dropped the UltraSeedbox layout group: nothing provides it, so it only ever rendered empty. Adds .env.example and a README documenting the two-path service model (docker label discovery across five engines vs manual entries), the labels-only-apply- on-recreate rule, and the foot-guns found: HOMEPAGE_ALLOWED_HOSTS matches host AND port so a bare IP does not cover IP:port; :2375 is plaintext and unauthenticated on all five engines; ping: cards can only be judged from the dashboard host. Verified after deploy via /api/services: 105 cards across 19 groups, both corrections live, ana-docker discovery intact. |
||
|
|
4b6daadb16 |
memory: operator confirms the heresy gen seat working well in real use
Records the one signal the synthetic gates cannot provide -- multi-turn degeneration is stochastic and invisible to probes, and four synthetic tests once validated three non-fixes on this exact seat. Not yet the 60k-token bar the prior seat cleared, so the rollback weights stay in place. |
||
|
|
d676a1375b |
memory: watch DavidAU's heretic Qwen3.8, not Cold-Fusion-GAIN V1.1
Cold-Fusion-GAIN V1.1 examined and not adopted -- it is a capability finetune of stock Qwen3.8 and every bench row is labelled [non heretic], so adopting it would reintroduce base refusals the current seat does not have. Records why it reads as uncensored at a glance: DavidAU's back catalog is almost entirely Uncensored-Heretic builds, so the naming pattern implies it. The heretic stage for this one is still in progress from base, and that is the release worth watching. Also banks what makes it interesting when the heretic build lands -- real third-party benchmark gains over stock, claimed MTP acceptance well above ours, thinking tokens cut to a fraction -- and the two caveats: the MTP numbers are GGUF/llama.cpp not vLLM, and a trained MTP head means the free CPU-hash gate would not apply. Also drops the now-stale 'primary until the DavidAU Qwen3.8 lands' clause from the superseded seat entry. |
||
|
|
11b688ff68 |
feat(gen-seat): promote absolute-heresy to the live gen seat
MuXodious/Qwen3.8-27B-absolute-heresy (Heretic v1.4.0 + SOMPOA, trial T377, pin c2374593) quantized through our mixed NVFP4+FP8 recipe and promoted after passing the full gate on the probe port. Gate vs incumbent -- MTP acceptance 47.2% (48.2%), decode 103.5 tok/s (96.4), prefill 6618/5403 at 6.7k/27k (6334/5085), TTFT 27k 5.00s (5.31s), perplexity 6.910 (7.059, 2.1% better), surface 6/6, abliteration compliance 4/4. On our battery-instruct arm -- the framing that actually elicits refusals -- 0/55 with zero EMPTY, so no catatonia at the hard edge. Speed deltas are image-confounded: the probe ran the seat's pinned nightly while the incumbent's stored numbers came from an earlier image. Read as not worse. Acceptance, perplexity, surface and refusal are apples-to-apples. All 7 LiteLLM aliases verified end-to-end. GPU0 at 91.3/97.9 GB with meromero healthy -- more headroom than the previous build. Incumbent weights untouched and .env.bak-heresy-20260817 in place for rollback. Candidate is a 2-day-old RC1 with ~348 downloads; watch real multi-turn use. |
||
|
|
993421bf59 |
fix(post-quant): handle sources that keep mtp.* inside a numbered shard
post_quant assumed the source ships a standalone model-mtp.safetensors, which is how JonathanColetti's grafted head is packaged. MuXodious/absolute-heresy is an unmodified full checkpoint, so its mtp.* lives in model-00012-of-00012 -- the copy silently did nothing while the index was still rewritten to point at model-mtp.safetensors, leaving 15 unresolvable tensors. Tensor counts looked correct; the checkpoint would have failed at load. The existing FAILED-CHECKS assertion caught it, which is the design working. Now extracts from the numbered shard when the standalone file is absent. Verified on the heresy build: 1968 tensors, all resolvable, 15 mtp, 333 visual, no missing shards, no orphans. |
||
|
|
b0c2d3d1c4 |
fix(bench): serve_probe must mirror the live seat -- image, parsers, context
Three defects, each of which produced a false read on the candidate: 1. Hardcoded vllm/vllm-openai:latest. The Qwen3.8-27B gen seat is pinned to a nightly carrying the #51113 qwen3_5_mtp x GDN fix; probing on :latest reproduces the multi-turn corruption we already diagnosed and reads as a candidate failure. Now PROBE_IMAGE, defaulting to :latest for older seats. 2. --speculative-config JSON died twice on quoting. The inner double quotes are stripped by the outer double-quoted ssh string, and then bash BRACE EXPANSION splits {"a":1,"b":2} on the comma. Needs escaped quotes AND remote-side single quotes; both traps documented inline. 3. No --tool-call-parser/--enable-auto-tool-choice/--reasoning-parser. Without them surface_test reported tool calling as a 400 and measured a thinking split of reasoning=0ch -- both probe-config artifacts, not model defects. Re-running with the seat's flags took the candidate from 5/6 to 6/6. Also adds PROBE_MAXLEN; the hardcoded 32768 rejected prefill_bench's ~27k prompt. |
||
|
|
2c3602869f |
fix(gen-seat): hash bf16 tensors via uint8 reinterpret, not numpy
numpy has no bfloat16, so .numpy().tobytes() raised 'TypeError: Got unsupported ScalarType BFloat16' on real checkpoints. Flatten then view(torch.uint8) before hashing. Result on the heresy candidate: VERDICT IDENTICAL -- all 15 mtp.* tensors byte-identical to the incumbent's verbatim base graft, the head already measured at 47.7% acceptance in production. The ~56 GB bf16 acceptance gate is redundant, so no second seat comes down. |
||
|
|
254c588921 |
feat(gen-seat): CPU-only MTP head check so the gate costs no second seat
Operator ruled the probe port for validation; runbook updated to match. The bf16 MTP acceptance gate is ~56 GB resident, which on a full 97.9 GB card means downing meromero-charrp as well as gen -- freeing gen's 0.43 (~42 GB) alone is not enough. Two seats down to answer one question. compare_mtp_head.py answers the common case for free. The Qwen3_5ForConditionalGeneration wrapper never loads the MTP head, so PEFT merges, Heretic runs, and llm-compressor passes all leave mtp.* as it came from the base. It hashes a candidate's 15 mtp.* tensors against the incumbent's grafted-verbatim head -- the one measured at 47.7% acceptance in production through this exact pipeline. Identical means the acceptance question is already answered; different means the head was edited and the real gate is warranted; missing means it was dropped. CPU only, reads just the shard holding mtp.*. The runbook states the residual risk plainly: an identical head proves the head is intact, not that the abliterated body still drafts well with it -- which the Stage-3 acceptance measurement on the 22 GB quantized build catches anyway. |
||
|
|
7997f111b0 |
docs(gen-seat): runbook for the absolute-heresy swap
Candidate MuXodious/Qwen3.8-27B-absolute-heresy (Heretic v1.4.0 + SOMPOA, trial T377), pinned c2374593. Beats the incumbent on both axes: refusals 2/101 vs 12/100, first-token KL 0.0759 vs 0.1191. Structurally a clean full checkpoint (1199 tensors, 15 mtp.*, 333 visual.*, lm_head), so the existing mixed NVFP4+FP8 recipe applies with no graft-and-reconstruct. Runbook carries the bf16 MTP-acceptance gate before any quant spend, the llm-compressor ignore-pruning foot-gun, the three measurement traps (cache-busting, unseeded prefill nonce, PPL with spec off), the .env 0600 sudo trap, the GPU0 co-tenant starvation risk, and rollback. Flags that the candidate is a 2-day-old RC1 whose own card carries a broken GGUF benchmark block (RC1 and RC2 report identical at-chance scores across three benchmarks), so its numbers are claims rather than measurements. |
||
|
|
c18f5c5d33 |
memory: WT #401 closed on demo verify; host ulimit floor staged not active
worldtree-dev closed #401 on our demo verification. Records the two-layer state (their e41b139 compose pin verified on demo, covered-not-verified on personal/pinned; our daemon floor staged), the measured fact that default-ulimits is not SIGHUP-reloadable on Docker 29.4.3, the explicit no-dockerd-restart decision, live-restore parked as a separate call, and the one ping we still owe once worldtree-personal recreates. |
||
|
|
7f3f265384 |
feat(corviduo-dev): stage a host-wide docker nofile floor (65536) for WT #401
Worldtree #401: a slow fd accrual in worldtree-personal hit the 1024 soft nofile ceiling and converted into a hard deadlock. Operator authorized the raise 2026-08-17 (relayed via worldtree-dev); sizing 65536 agreed. Applied at the daemon layer rather than compose because /opt/worldtree-*/ compose.yaml on corviduo-dev is written by the team CI deploy identity -- a host-side compose edit reverts on the next deploy and would leave a false 'raised' record. Daemon config is infra-ops-owned and covers all 13 containers on the box. worldtree-dev shipped a redundant compose-level pin (e41b139) as the belt to this braces. daemon.json is written and valid, but the floor is STAGED, NOT ACTIVE: default-ulimits is not in dockerd's SIGHUP-reloadable set. Measured on 29.4.3 -- the post-reload 'Reloaded configuration' log enumerates the live config without default-ulimits, and a fresh container still reports ulimit -n 1024. Activation needs a full dockerd restart, which bounces every container; not taken, since #401 is not urgent at fd ~100 and the compose pin already covers worldtree. The playbook documents this and its verify step 3 fails by design until a restart happens. |
||
|
|
2686042106 |
memory: fleet IPv6 state + verified VPN topology; ana-wg key material locked down
Durable capture ahead of the ESH fiber install (2026-08-18) that puts the house behind CGNAT and breaks Site Magic on IPv4 -- IPv6 becomes the escape hatch and the likely first consumer of fleet v6. Topology verified rather than assumed: Site Magic between UniFi units, IPsec IKEv2 colo<->UniFi, and WireGuard as a remote-access convention only, host-based on ana-wg behind a FortiGate UDP VIP. The FortiGate port-forwards and never terminates WireGuard, so FortiOS 7.2's lack of native WG is a non-issue. IPv6 today: NH3 WAN live, colo and ESH none. AT&T delegates exactly one /64 at NH3 -- established by forcing the prefix ID from auto to 0 and observing the subnet not move, since the c110/c11f pattern otherwise reads as a /60. PD enabled on nh3-iot to measure, then reverted; all five NH3 LANs are back to ipv6_interface_type=none. Also fixed on ana-wg: wg0.conf, keys/*_priv, keys/*_psk and the client configs were mode 644 with private key material in them. Now 600, with keys/ and configs/ at 700. wg-quick@wg0 stayed active, three peers intact. Corrects two stale in-flight rows: the DS regeneration is retired, not queued, and SPEC-ds-regeneration.md is deleted rather than untracked. |