84349d7a0e4ac23656cfa659ade1c9016377bc5d
135 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
84349d7a0e |
docs(althing): check the hook list, not the version string
A version number cannot tell you what a stale plugin cost. 0.0.1 and 0.1.1 differ by two hooks and a script, so the runbook now carries a check that compares hook lists across cached versions and looks for pane-route.sh directly. Also records why this hid for five days, which is the more transferable half. A missing deploy surface does not present as an error -- it presents as "the migration needs manual work", and there was a ready explanation for that, because four of five seats were non-Claude and genuinely did need hand-holding. The seat that falsified the story was our own: a Claude Code seat that should have self-declared and did not, and it looked exactly like the other four. Nobody asked why the automatic path had not fired on the one seat it was built for. So: when a migration needs manual intervention, verify the automatic path was actually deployed before concluding it does not apply to your case. |
||
|
|
d0882fb830 |
feat(althing): four-surface deploy script + runbook
Deploying althing touches four independent surfaces on nh3-dev. Three were known. The fourth -- the plugin -- had no step in any runbook and drifted for five days before anyone noticed. The plugin chain is repo plugin/ to the marketplace directory to Claude Code's cache, and neither hop was automated. The marketplace directory was a frozen copy from Aug 28 carrying only the UserPromptSubmit hook, with no SessionStart, no SessionEnd and no pane-route.sh at all. So the claim that CC seats re-declare their pane route automatically at session start was never true on this box, which is why every seat had to be hand-declared with a pid measured by hand. The script backs up the marketplace directory before syncing, re-stamps its marketplace.json from the repo's plugin.json, and uses `claude plugin update` for the cache rather than hand-editing installed_plugins.json -- that is Claude Code's own bookkeeping and a subtle mistake there breaks the plugin in a way that looks like an upstream bug. The runbook also carries the two things most likely to waste someone's afternoon: `uv tool install .` without --force is a silent no-op that exits 0 having done nothing, and a live waiter reporting mode:pull is a seat that will never be poked, with the audit loop for finding them. |
||
|
|
931bac8f68 |
docs(matrix): current state, upgrade procedure, alias convention, push findings
Synapse v1.120.0 -> v1.159.0 and Element-web v1.11.80 -> v1.12.27 (2026-09-01). The existing build steps date from the AIPA era and are now marked as provenance rather than as instructions. Records what only existed in a session transcript: - Schema migrations are one-way; rollback is restore-from-dump. Pre-upgrade pg_dump procedure, with a pg_restore --list verification step. - Why the appservice user namespace is now exclusive: false. exclusive governs who ELSE may act, not what the appservice may do, so on a closed single-admin server it locked out all other account creation to prevent squatting that cannot occur. Includes the two things not to do: narrow the regex (orphans 13 accounts) or rename the id (Synapse keys ownership on it). - The shared-secret registration HMAC takes no trailing null after notadmin. - Room alias convention #<agent>-<purpose>, operator-ratified, with its cost accepted deliberately and its rationale stated as room-identity-carries-tier rather than push-payload-carries-room-name. - Push reality: the pusher is event_id_only, so the notification is assembled on-device by Element X's service extension. Records the resulting server-invisible failure mode when the phone cannot reach the homeserver. - QR sign-in requires Matrix Authentication Service and why it is deferred. Ops ownership recorded: worldtree-dev writes the bridge, infra-ops operates this instance. |
||
|
|
1a36e60d3a |
docs(quant-playbook): §3.7's APC-off mitigation was reverted nine days ago and the section never said so
Found while answering a question from the operator, relayed via brokkr-smithy-dev, about whether a recorded Qwen3.8 degeneracy at ~1,700 tokens relates to a length sensitivity just measured on the tuned Gemma-4. The record is §3.7 and the number is ~2,000 -- but reading it to answer that question surfaced that the section is stale. §3.7 presented "disable prefix caching, keep MTP" as THE MITIGATION, resolved 2026-08-17, and stated the gen seat runs that config. It does not and has not since that same day: APC-off passed a synthetic 7-turn probe and the operator still saw severe degeneration in real use, so it was reverted. The multi-day hunt resolved to the AEON W4A4 quant being defective, with MTP / prefix-caching / gateway merely amplifying it (§3.8 records the corrected causal story; §3.7 was never updated to match). Verified against the live container rather than against the compose file alone: vllm-gen runs --enable-prefix-caching with qwen3_5_mtp / num_speculative_tokens 3. stacks/gen-seat/compose.yaml carries the full corrected history inline and is the current authority. §3.7's superseded text is kept and fenced rather than deleted -- it is the history of a mitigation that looked right and was not. Added a dated row to §7 per the standing rule that a wrong playbook claim gets a superseded-claims entry, not just a fix. The lesson inside the lesson is worth more than the correction: §3.7's own standing rule is "gate MTP on a multi-turn coherence probe, not just single-shot acceptance." The APC-off mitigation was gated on exactly that probe, passed it, and still failed in real use -- the multi-turn probe was itself too small to gate on. A passing probe is not sufficient evidence at any size that has not been calibrated against real use. |
||
|
|
1a4ef5c7a1 |
docs(training-playbook): 4.6.3 was wrong twice — correct it, and keep the retraction visible
The entry reported an 'output-stability regression' as a novel run-2 finding. Both halves were false and the corrections are more instructive than the original conclusion, so they stay in-line rather than being edited over. Not new: run 1's own gate record already carried the same effect with a caveat attached and unresolved. Two runs across two different base models makes it a property of the RECIPE, not of the base swap -- which also means a third run that changes the base again will not fix it. Not degeneracy, and not a separate finding: all 46 flags were too_short rp turns of 3-14 words, and the two collapse guards fired ZERO times on any run. It is the left tail of a length distribution that had been measured and reported in the same message. Truncation is the same mechanism mirrored on the story side. Both are thresholds calibrated on the base's output shape applied to a model with a different one -- 4.6.1, which both parties had written down and neither applied. The surviving lesson is sharper: a short-answer gate cannot see length behaviour AT ALL, and because it could not, the effect went two full runs before anyone named it. The cost of a gate-set blind spot is measured in runs. Adds 4.6.3.1 on trip points inside the serving stack's jitter -- same seed, same weights, rate moves 9.6% -> 12.6%, sd 1.77pp. Not 'the gate is non-deterministic' but 'the trip point sits inside the jitter', because the fix follows from the precise statement. Includes the split-design rule for measuring such a rate, and the rule that a measured rate must carry its corpus in its name. |
||
|
|
37d3189622 |
docs(erp-dpo): the clip hypothesis is falsified — the distribution is bimodal
The output-side test ran on the live seat. There is no shoulder at 123: the 120-139 bin holds three of ninety-six and is a TROUGH, and 17.7% of generations cross a cap PIPPA can never cross. The clip-as-boundary reading is dead, killed by the test that could have confirmed it. Corrects this document's own earlier read, which compared the tuned MEAN (88.5) to PIPPA's MEDIAN (67) and concluded 'comfortably inside the upper body'. Median to median it is 62 against 67. Mixing statistics across a comparison produced a more reassuring answer than the data supports. What the data shows instead is bimodality -- a mode at 20-39, a trough, a second mode astride PIPPA's centre, a tail to 505, against a base with no such shape. The tune changed rp length's SHAPE rather than its centre: roots whose length distributions do not overlap learned as distinct modes rather than blended into an average. And the skew is rp-ONLY, which localises it to the family the clipped root lives in and is the strongest support the turn-share mechanism gets from the output side. Consequence for pair generation: chosen/rejected sampled from a bimodal generator inherit the mixture, not a mean, and naive sampling over-draws the short mode. Also records that the degeneracy rate is NOT yet a usable baseline -- same arm, same seed, VOID flipped no->YES across a re-run because the 10% budget sits at the noise boundary. A guard whose trip point is at the noise floor produces disagreement between honest observers rather than silence. Replicates running. |
||
|
|
5171f19e16 |
docs(erp-dpo): the PIPPA length clip, measured — DPO pairs would inherit it
The run-2 gate found tuned rp turns 36% shorter than base. brokkr hypothesised the mix was teaching PIPPA's 2023 Character.AI product clip; the corpus side is now measured and confirmed. PIPPA's max is 123 words EXACTLY, 100% at or under it, and 0.00% in the 124-130 band -- a wall, not a preference. Every other root crosses its own p99 smoothly. The mechanism is sharper than 'PIPPA is in the mix'. PIPPA is 70.3% of bot TURNS but only 37.5% of bot WORDS, precisely because its turns are clipped -- and length is learned per turn, not per token. By loss tokens it looks like a third of the dialogue signal; by end-of-turn demonstrations it is seven in ten from a source that cannot exceed 123 words. Generalises: a length-clipped root is over-represented in the length signal by exactly the ratio its clipping creates. Counter-evidence recorded too: the tune landed near PIPPA's MEDIAN (67), not its CAP, which is central tendency rather than learning the boundary. Weaker claim than the hypothesis, and not demonstrated either way. Filed here rather than only in the gate record because preference pairs generated FROM this tune inherit its length distribution in both chosen and rejected -- DPO would train an artifact in as an explicit objective. Settle the length question before generating pairs. |
||
|
|
0bb9ee7777 |
docs(training-playbook): 4.6.3 — a short-answer gate cannot see a long-form defect
Run 2's reasoning battery reported zero truncations and zero degenerates on both arms across four passes. The same tune, measured on long-form generation in the same session: truncated 0/384 -> 38/384, degenerate 0/384 -> 19/384. A real output-stability regression, structurally invisible to that gate because its answers are short. Not a bug in the battery -- a coverage property. An instrument measures the regime it samples, and output length is a regime. Generalises to context length, conversation depth, and any axis where the gate's operating point is narrower than production's. The actionable form: enumerate the regimes your gate set spans, name the ones it does not, and decide deliberately rather than discovering the gap downstream. Corollary on sequencing -- put a long-form generation in the gate and put it early, because a length-dependent regression is exactly the one you want found before four clean short-task passes make everyone comfortable. |
||
|
|
a0f59d2778 |
docs(training-playbook): 4.6.2 — a null result needs a positive control
From the run-2 gate. A memorisation probe reporting 0.00% across all 72 items is the correct output for a model that has not seen the corpus, and is also the exact output of a probe that is not firing. Nothing in the number distinguishes them. brokkr-smithy-dev drove the overlap function with known-answer inputs (identical 100%, half-verbatim 65.38%, unrelated 0%, empty 0%) before trusting the null, which is what converts a suspicious zero into evidence. This is 4.5's inert gate wearing a different face: there a check that could not return 'fail', here a measurement that cannot return non-zero. A clean null is the most reassuring output an instrument produces and the least self-evidencing. Same section records the identical-on-both-arms variant: the diversity battery's rp family froze zero markers, so its attractor hit rate read 0.0 on base AND tuned. That reads as a clean result and means the instrument cannot discriminate on that family. Report as a bounded limitation, never as a delta of zero -- a check returning the same value for every input is not measuring. Checklist gains the line. |
||
|
|
d54f25605f |
docs(training-playbook): 4.4.1 sample identity at launch; 4.6.1 calibrate gates against correct input
4.4.1 -- the dirty-tree case was only half of the harness_commit problem. Run 2 launched CLEAN at 1909d86 and recorded 460f372, because three commits landed on the same checkout during its seven hours and _git_commit() was called at save time. Commit AHEAD of the code that ran, naming changes it never executed -- including the provenance fields this section prompted. Same defect as run 1's BEHIND, opposite sign: the identity was sampled at the wrong moment. Sample at launch, carry it, and record the dirty flag beside the commit rather than instead of it. Generalises to every run-scoped identity: anything read at save time describes the world at save time. 4.6.1 -- the inverse of the inert gate, and it costs trust rather than correctness. A coherence gate false-rejected 'The capital of Portugal is Lisbon' as degenerate against a global 15-word floor. The floor was calibrated against the wrong reference, not set too strict. Lowering it globally would blunt the check where short output genuinely is degeneration; the fix is a floor per prompt. Write the positive test alongside the negative one. |
||
|
|
bcf63db527 |
docs(erp-dpo): readiness survey for the DPO stage
Run 2 is an SFT on the official instruct base, so it will refuse at near-stock rates by design; targeted DPO is where refusals get pruned on chosen axes. That was the trade accepted when the stock base was picked over a third-party abliteration. Surveys what is on disk against what the stage needs. Ready: the merged tune, the SFT adapter, GPU0 once the eval seat comes down, the whole non-loss half of the SFT harness, two unvetted Gutenberg preference sets, and the LitBench-RM judge. Missing, in order of pain: preference data for the refusal axes (nothing on disk targets it -- the Gutenberg sets are prose-quality), the axis list itself, and a DPO trainer (trl is not installed). The gating item is not technical: WHICH refusal axes are in scope and which are explicitly kept. Data generation, pair counts, the held-out split and the success probe are all functions of that list, so nobody should generate a pair before it is written down. Flags that the domain-compliance probe should measure run 2 BEFORE pruning, since the pre-number is the only baseline that will ever exist. Also records the operational trap: do the trl install AFTER a run finishes, never during one -- a resolution that upgrades transformers under a live process can break its save path. |
||
|
|
dbca9a3c66 |
docs(training-playbook): audit the whole manifest against the pairing rule
§4.3's generalisation was stated and then not applied to the manifest that prompted it. brokkr-smithy-dev did the audit: most fields are intent-only, and the one pairing that would have caught the §4.1 cache failure -- the mask's sha against the loss-token delta -- existed by accident, because someone had asked for an encode report for unrelated reasons. Adds the audit table, and the rider that matters more than the table: put the observed check where it can actually FAIL. chat_template_sha256's pair is the sha of the string the tokenizer carries, but asserting that in the parent one line after assigning the file to the tokenizer compares a value to itself. It belongs in the encode worker -- a different process, across a pickle boundary, where an unset config key silently leaves every worker rendering through the checkpoint's own template. |
||
|
|
c1db188e6a |
docs(training-playbook): §4.3 records an OBSERVED consequence, not just a config string
brokkr-smithy-dev pointed §4.5's own test at §4.3's remedy: recording `attn_implementation_resolved` is a check that cannot fail on the axis the failure lives on. A silent Dynamo fallback to uncompiled flex leaves `config._attn_implementation == "flex_attention"` untouched while the run computes at ~20x the cost and, per torch's own docs, does not work correctly through the backward pass. The field records the request's RESOLUTION, not its SURVIVAL. On the failure mode that matters it reports success either way. So the section now requires the step-time distribution beside it -- n, min, p50, p99, max -- which is the check that can actually fail. Compiled sits at p50 ~20 s; a fallback at ~400 s. One perf_counter() in on_step_end buys it. Distribution rather than a mean, because a mean hides exactly the bimodality a PARTIAL fallback produces. Generalised past this instance: any provenance field recording a CONFIGURED value is a claim about intent. If the failure you fear is the configuration silently not taking effect, you need a second field recording an OBSERVED consequence, and the pairing is the check. A settings dump alone is decorative. Two implementation details are called out because both were wrong in the first draft -- percentiles nearest-rank so every reported value is a real observation, and exclude the FIRST step rather than the slowest, since step 1 carries compilation but is not reliably the maximum on a variable-width run. New §4.7.1: rotate the log on relaunch. Run 2's first attempt died on the warmup_ratio TypeError and the relaunch appended, so the traceback sat at line 15 of a file whose live run began at line 39 -- and a `tail -n +1 -F` monitor replayed the dead traceback as a fresh event. One file describes one run. Checklist gains both lines. |
||
|
|
dae6ede8e2 |
docs(training-playbook): §4 — when the artifact lies about itself
The playbook covered why a run is SLOW. It did not cover the more expensive
failure: a run that COMPLETES, reports plausible numbers, and is wrong about
itself. Seven of those turned up on the Gemma-4 ERP/RP tune between 08-24 and
08-26 and not one raised an error.
New §4, seven landmines plus a pre-launch checklist:
4.1 a cache key must cover the MEANING of the cached thing. The encode
cache missed the impersonation mask; run 2 would have reused run 1's
unmasked encodings and written impersonation_mask_sha256 into its own
manifest while doing it. No error, no count change, normal loss curve.
4.2 validating a VALUE is not validating the PARAMETER. warmup_ratio was
in range and deleted from transformers 5. Build kwargs as data and
diff the NAMES against the installed signature -- you cannot check the
argument list of a call you have already made.
4.3 record what the run RESOLVED to, never what it requested. Run 1
recorded no attention backend, so an MFU panel profiled the serving
seat under sdpa and recommended adopting flex_attention for a run that
was already using it.
4.4 never train from a dirty tree; harness_commit will name a commit that
does not describe the run. Annotate afterwards, never edit the shipped
artifact -- and state what is NOT wrong, or the note casts doubt on
every field it omits.
4.5 a watchdog whose pgrep pattern appears in its own argv can only ever
return "alive". The inert-gate shape in a liveness check.
4.6 an instrument nobody runs is not an instrument. Mutation-check any
test guarding a property that fails silently.
4.7 fix a stale measurement at the source. "~4.3 HOURS to rebuild the
encode cache" (really 145.5 s) was copied into a new launcher by the
same person who had just measured the real number.
4.8 the pre-launch honesty checklist, ten minutes.
Also:
- Header and framing widened. The file is now a training playbook with a
throughput half and an integrity half; the filename stays for inbound links.
- Sections 4-7 renumbered to 5-8. External refs are all to §1.1 and §3.4 and
are unaffected.
- Four rows added to the superseded-claims table, including the kernel table /
68% quadratic / 8.6% MFU set, which describe the serving seat rather than
the training run.
- gemma4-erp-tune-sizing.md §6 carries a correction banner with the explicit
falls/survives split, because that is the doc someone actually reads before
a run.
|
||
|
|
64bf9d313f |
docs(training-playbook): measure refusal retention on the abliteration's OWN axis
§3.13, plus the probe that produced it. Two lessons, both about measuring the wrong thing confidently. First: a tune applied AFTER an abliteration can walk it back, and a reasoning/craft/memorisation gate cannot see that. brokkr-smithy-dev's preregistered gate measured none of it — a tune that gains 41 items of contradiction detection and quietly restores refusals passes every check. The compliance axis has to be added explicitly. Second, and this is the trap: measure the axis the abliteration was actually FOR. Ours was run so the model engages explicit fiction. The probe reached for mlabonne/harmful_behaviors — weapons, malware, fraud — because it was cached and carried a recorded baseline. Different refusal surface entirely, and a model moves on them independently. 29/100 general-harm refusals on a tune whose prose the operator was praising at the time is not obviously a defect and may be desirable: general-harm refusals returning while domain compliance holds is close to the ideal shape for an internal creative seat. The measurement was real; its relevance was assumed. Also recorded, because both were nearly missed: - Read the interesting cell. In 29 hard / 0 deflect / 71 comply, the load-bearing number is 71. Stock refused 100/100; near that would mean the abliteration was undone. 71 complying means partially walked back on one axis — a different finding, and only one of the two threatens the seat. - A baseline from a different harness is not a baseline. The recorded 3/100 came from the abliteration tool's scorer, which reads first-token probability distributions; a probe that generates and regexes is a different instrument. Run your own against both arms on the same seat or report the number alone. - A refusal regex undercounts, so classify hard/deflect/comply — and the free discriminator: if both arms return zero deflections the model is binary; if only one does, the regex is fine. An artifact does not care which arm it runs against. |
||
|
|
a696b49e2a |
docs(training-playbook): merging a tune back toward stock can UNDO an abliteration
§3.12. brokkr-smithy-dev caught and retracted his own recommendation mid-thread; recording it before it reads back later as advice. A common remedy for an overfit tune is a partial merge back toward the base to recover general capability. The published recipes that recommend it merge into the STOCK instruct checkpoint. On an abliterated base, following that literally re-introduces the exact refusal directions the abliteration was run to remove — and it is silent, because the merged model looks healthier on general benchmarks while the property the seat exists for quietly returns. Rule: any merge-back targets the SAME base the LoRA was trained against, never the upstream stock weights however similar the name. The wider lesson is about recipe-card provenance. Community cards are per-checkpoint artifacts and do not transfer across dense-vs-MoE, stock-vs-abliterated, or size variants. The worked example: a recommendation carried from a card for a DENSE STOCK 31B onto a MoE ABLITERATED 26B-A4B on the strength of a shared family name. The overfitting warning on that card happened to come from the right architecture; the pipeline, reward stacks and merge-back came from the wrong one. Same family, three axes apart. So: before quoting a recipe card at a decision, state which checkpoint it was written for and which axes differ. "Same family" is not an answer. |
||
|
|
2ec8f42297 |
docs(training-playbook): base-viability pre-flight, three greps before you pick
§3.11. Three consecutive "what about X as a base?" questions in one session,
each answerable in minutes, none of which had been asked before a 7-hour
training window was committed. Writing the check down so it runs first.
1. does it fit for TRAINING - BF16 weights against the real measured peak,
not the weight figure (Gemma-4 is 48.1 GiB of weights and peaks at
79.7 GiB at mb2/seq-16k). Model-line names lie: "Mistral Small 4" is 119 B,
238 GB in BF16, more than both cards combined. QLoRA is not an escape
hatch for MoE - bitsandbytes walks nn.Linear and fused 3-D experts are not
that.
2. if MoE - does the serving engine implement get_expert_mapping. Zero means
LoRA cannot be served at all. gemma4*.py -> 0; deepseek_v2, mixtral,
glm4_moe, ernie45_moe -> present.
3. does the model class support LoRA - and GREP THE CLASS, NOT THE FILE.
Point 3 has teeth and I nearly got it wrong twice in one turn. mistral.py greps
as SupportsLoRA=0 and is fully LoRA-capable via LlamaForCausalLM.
mistral_large_3.py greps as 0 for both and inherits get_expert_mapping from
DeepseekV3ForCausalLM. Capability is inherited; a file-level grep misses it and
only MRO resolution answers it. Same class of error as asserting a substring
instead of an effective value.
Worked results recorded for the three candidates evaluated:
Gemma-4 26B-A4B fits, no expert mapping -> trainable, MERGE-ONLY
Mistral Small 4 119B 238 GB, has mapping -> servable, NOT trainable here
Ministral 3 14B ~28 GB, dense, inherited -> passes all three
Adds a fourth glance at architecture shape, since it predicts how much of this
playbook applies at all: uniform head_dim <= 128 with no sliding window keeps
both flash and cuDNN reachable and makes §3.1/§3.3 moot, while mixed head dims
plus a sliding window is exactly what forces dense O(n^2) attention onto
Ampere-generation kernels for 65% of the step.
|
||
|
|
96731bb090 |
docs(training-playbook): prove the serving path before spending the window
§3.10. The quantization playbook already says prove your targets before spending GPU time; this is the same rule one step later, and easier to skip. A ~7h LoRA run was built assuming the adapter could be hot-swapped onto a quantized base at serve time. The sizing doc flagged serving as unsettled and said the requirement was needed "while he is early, not after the run" — the concern was identified correctly and then the check was deferred. Tested afterwards, vLLM refuses outright: gemma4's model class implements zero occurrences of get_expert_mapping, which process_packed_modules_mapping requires for any MoE model. One grep, available months earlier. Two generalisations recorded: - Feature support is per-architecture, not per-family. LoRA works for the DENSE sibling of this same model family and not the MoE one, so "model X is supported" says nothing about X's variants. - A capability gap in the serving engine cannot be worked around from the training side. The adapter here never touched experts and was refused anyway, because the refusal keys on the model being MoE, not on what the adapter targets. Includes the mechanical check: grep the engine's model class for the capability, then start the engine with the feature flag alone — no adapter required, since --enable-lora forces the machinery to initialise and that is where it fails. The recovery is cheap here (merge, ~35 min per tune). The cost of finding out late is that it forecloses an architecture choice after the training window has already been spent. |
||
|
|
8de5f7a73c |
docs(gemma4-erp-tune): merged weights are mandatory — vLLM cannot LoRA any Gemma-4
The §5 open question was whether LoRA-on-NVFP4 hot-swap still silently no-ops as it did on vLLM 0.24.0 (#47639), with merged weights as the fallback if it did. Retested on vllm/vllm-openai:latest against the NVFP4A16 base plus the live run's checkpoint adapter. It does not no-op. It refuses to start: AttributeError: To support LoRA for MoE model, 'get_expert_mapping' must be implemented And the reason is bigger than the quant. The check is in vllm/lora/utils.py::process_packed_modules_mapping and branches on whether the model is MoE — quantization is not in the condition. gemma4.py, gemma4_mm.py, gemma4_mtp.py and gemma4_unified.py contain zero occurrences of get_expert_mapping, while deepseek_v2, glm4_moe and ernie45_moe do implement it. So vLLM cannot serve a LoRA on Gemma-4 at all, BF16 or quantized. Merging is not a workaround for a quantization limitation; it is the only path for this architecture. This holds even though the adapter never touches experts — validate_adapter_parameters forbids per-expert params, so all 205 targets are attention and dense MLP. The refusal is about the model being MoE, not about what the adapter targets. Worth recording that the current behaviour is an improvement: a loud refusal beats the 0.24.0 silent no-op, which would ship a base model wearing the tune's name and pass every check that does not compare against base. |
||
|
|
6a8582936e |
feat(erp-tune): NVFP4A16 serving pipeline, and the MoE landmine it uncovered
Merge + quantize path for turning the Gemma-4 26B-A4B ERP/RP LoRA into a
servable NVFP4A16 seat, plus a playbook entry for the defect found while
validating it.
The landmine (playbook §3.15): a `targets=["Linear"]` NVFP4 recipe silently
misses every MoE expert on this architecture. Gemma-4 stores each layer's 128
experts as two fused 3-D nn.Parameter tensors, not nn.Linear modules, so the
recipe resolves 205 of 427 modules and ZERO experts — 22.84 B params, 88.5% of
the model, left in BF16 with no warning. This is the same blind spot that
killed QLoRA here via bitsandbytes; the tool changed, the checkpoint layout did
not.
before linearize_moe: 427 Linears, 205 targeted, experts 0
after linearize_moe: 11,947 Linears, 11,725 targeted, experts 11,520
(30 layers x 128 experts x 3 projections)
llmcompressor's linearize_moe unfuses them; no registration needed because
Gemma-4 satisfies FusedExpertsProtocol structurally. Caught by an §4.1 dry run
that asserts the expert count before any GPU spend, which is now the documented
requirement rather than an optional step.
Scheme is NVFP4A16, deviating from the playbook's mixed-W4A4 default on
measured grounds: brokkr-smithy-dev benched the W4A4 quant of this checkpoint
at 12% on contradiction detection with CoT off against gen's 81%, the signature
of 4-bit input activations on a reasoning-dense task, and W4A4 KLD degrades
2-4x past ~10k ctx on sm_120. This is a 16,384-ctx RP seat. Marlin's prefill
cost is accepted.
Two further silent-failure guards, both from prior hard-won lessons:
- the merged model ships the UPSTREAM chat template, not the trainee base's
stale 365-line one, because training rendered through upstream and the
mismatch would present as a tuning failure
- calibration reads the run's own encode cache rather than re-tokenizing, which
sidesteps §3.14 (a fast tokenizer mutated by truncation=True and persisted by
save_pretrained clamps every prompt forever)
Merge-then-quantize rather than LoRA hot-swap, since hot-swap onto NVFP4 was a
silent no-op on vLLM 0.24.0 (#47639). merge_lora.py asserts sampled target
weights actually changed, so an inert adapter cannot ship as a tune.
|
||
|
|
7b5fd91d3c |
docs(gemma4-erp-tune): root-cause the 8.6% MFU — attention on Ampere kernels, 29.9% padding
Run-01 was killed at step 19 by operator instruction to root-cause before spending a ~13.9h window. Two independent methods now agree on where the step time went, and neither was the hypothesis the consult panel converged on. Scaling fit (3 points, 2 params, residuals <3ms over an 8x range): A = 6.87e-4 s/token, B = 8.85e-8 s/token^2 quadratic share 20.9% @ w=2048 -> 67.8% @ w=16384 No fixed term was needed, which refutes launch-bound outright. Profiler kernel table (device rows only): attention 22,835.8 ms 65.2% fmha_cutlass*_sm80 dense GEMM 2,774.0 ms 7.9% other 5,739.0 ms 16.4% The attention kernels are sm80 — Ampere-generation CUTLASS running on an sm_120 Blackwell card, with the forward on the gmem fallback tier. That is the mechanism behind 100% SM utilisation at 27 of 304 available TFLOPS. Correctness cleared separately: the sliding mask asserts at max 1024 allowed/row, so the 25 windowed layers were genuinely windowed. The same probe found that right-padding is what pins the 5 global layers to an explicit 4D mask and off the is_causal fast path — measured at 9.4% slower for 24% less loss work at fixed width. The largest available win is not the attention kernel. The corpus is 29.9% padding, and bucket-to-pair + shuffle-to-mix takes it to 0.0% for >=35.5% wall clock, no new dependency, unchanged peak memory. Bucket size turned out not to be a diversity knob — roots per accumulation window are flat across a 256x range, so the global micro-batch shuffle does that work alone and the bucket should be tight. Adds docs/pfi/training-throughput-playbook.md as the durable model-agnostic home (sibling to the quantization playbook), the four probes under scripts/training-probes/ with raw output kept for re-derivation, and a §6 to the sizing doc carrying the Gemma-4-specific numbers and round-2 restart parameters. Measured negatives recorded so they are not re-chased: grouped_mm (0.9% slower, and MoE is only 7.9% of the step), CUDA graphs / torch.compile over the expert loop (no fixed cost to amortise), liger fused CE (~1-3% lever), FA4 on sm_120. Round-1 state preserved: 609MB encode cache, order manifest, truncation report, resume script. No checkpoints — it died at step 19 and the first was due at 100, so the lora_B inert-adapter gate never ran and moves to the restart. |
||
|
|
33433e0d1e |
docs(gemma4-erp-tune): replace the estimates with measurements — they were 3x optimistic
Ran the loss path on the real checkpoint on GPU0 with synthetic tokens.
The arithmetic held for parameter counts and was badly wrong for
activation memory.
naive CE bsz1 seq 8192 81.93 GiB
naive CE bsz1 seq16384 OOM
chunked CE bsz1 seq16384 65.66 GiB
chunked CE bsz2 seq16384 79.71 GiB <- the run config
chunked CE bsz4 seq16384 OOM
The marginal cost of an extra 16,384-token sequence is ~14 GiB, not the
~5 GiB estimated: the estimate modelled gradient checkpointing as
storing layer inputs plus a modest recompute peak, and the real MoE
recompute peak with top-8-of-128 routing and its scatter/gather buffers
is far heavier. Dense-model intuition does not size an MoE run.
Two predictions landed exactly — 205 target modules and 74,342,400
trainable params at r64 — which is why the rest of the model of the
thing is still worth trusting.
The headline is that chunked CE at seq 16384 costs 16 GiB less than
naive CE at seq 8192, so chunking is what makes brokkr's 16384
recommendation reachable rather than an optimisation on top of it.
max_seq_len moves 8192 -> 16384 on his truncation finding: the cap
drops 6.2% of samples but 22.4% of tokens, concentrated entirely in
dialogue, which is 60% of the mix.
Also records the four harness changes this required (eitri-smithy
62b556b), including the inert-adapter trap: without
enable_input_require_grads() alongside gradient checkpointing on a
frozen base, no gradient reaches the adapters, every one stays at its
initialisation, and the run completes successfully having learned
nothing.
|
||
|
|
c9943b1507 |
docs(gemma4-erp-tune): whole-card placement — gen moves to GPU1, sec stands down
Operator chose a third placement over the two the sizing offered: rather than train beside gen on GPU0 or on GPU1 in mog-sec's slot, move gen to GPU1 and empty GPU0 completely. The tune gets 95.60 GiB with no co-tenant and gen never goes dark beyond its own restart. Revised run parameters, since a whole card changes them: - micro-batch 8 (71.8 GiB of 95.60) rather than 4, grad-accum 1, giving 888 optimizer steps instead of 444. At one epoch the step count is worth having, and 8 x 8192 tokens puts ~4,096 rows through each expert per step against ~512 at micro-batch 1 — a far healthier GEMM on 704-wide experts. - Gradient checkpointing stays ON. Dropping it takes ~17% off wall-clock but pushes activations to ~24 GiB per sequence, which forces micro-batch 1 and costs 8x on MoE efficiency. Wide beats shallow. - Scriberr stays on GPU1. The previous revision suggested moving it to GPU0, which was correct only while training was going to live on GPU1. Records the ordering constraint in both directions, the elway identity requirement, and that sec's aliases should be allowed to fail at the gateway rather than be substituted with another model. |
||
|
|
c507db9ac0 |
docs(gemma4-erp-tune): size the run against the checkpoint — QLoRA is structurally unavailable
The proposed shape was QLoRA r64. It cannot be run as specified. The checkpoint stores each layer's 128 experts as two fused 3-D nn.Parameter tensors (experts.gate_up_proj [128,1408,2816], experts.down_proj [128,2816,704] — no .weight suffix, so they are parameters, not modules). bitsandbytes 4-bit replacement walks nn.Linear only, so 22.84B params / 42.54 GiB — 88.5% of the model — is skipped and stays BF16. load_in_4bit saves ~3.1 GiB of 48.07 and does not error while doing it. Verdict: plain LoRA on BF16, ~57.6 GiB at micro-batch 1, +2.5 GiB per additional 8192-token sequence. Two sizing items were absent from the brief and both are load-bearing: - vocab 262,144 x seq 8,192 = 2.147B logits, with final_logit_softcapping 30.0 adding a saved pre-cap tensor. Naive HF cross-entropy peaks at ~28-30 GiB transient at batch 1, which puts the run at ~85.6 GiB on a 95.6 GiB card — it starts, then OOMs on the first long sample. Fused or chunked linear CE is mandatory and must be smoke-proven before a window is booked, since Liger may not carry a Gemma-4 MoE patch. - v_proj does not exist on layers 5/11/17/23/29 (attention_k_eq_v on the full-attention layers). A v_proj target silently produces no adapter there, and k_proj adapts K and V simultaneously. 45.96M trainable at r64 across q/k/v/o. Placement, measured: GPU0 has 53.46 GiB free beside gen, ~4 GiB short, and gen's footprint grows with uptime. Stopping mog-sec frees 74.29 GiB on GPU1, which holds micro-batch 4 at 61.8 GiB with margin for Scriberr. Recommend standing down sec (2 aliases, last request ~5h ago) rather than gen (7 aliases, 765 busy-engine log lines in 24h). Estimated 1.28e18 FLOPs for the epoch at ~3.67B active params; 4-10 hours at 10-25% MFU. 7,104 packed sequences is only 444 optimizer steps at effective batch 16, which makes the wall-clock-checkpointing amendment concrete rather than hypothetical. Package as a uv venv on /tank: root is 91% full (36 GB) with /var/lib/docker on it. |
||
|
|
d127e29fac |
docs(ipv6): retire the next-candidates line now that all three are done
Replaces it with why the remaining segments have no eligible hosts: two are appliance-only and three have no IPv6 enabled yet, pending the firewall-policy pass that SLAAC on a client segment would require. |
||
|
|
4e83395ddf |
docs(ipv6): all three esh-server hosts now carry the segment name
esh-pve-nas and esh-vm-db join esh-docker-vm on 4411:b105, at :50:55 and :50:60 respectively, each applied by the same prefix-deriving if-up.d hook so the last two groups read straight off the IPv4 address. Two obstacles are recorded because both will recur. The Proxmox node had link-local only despite every relevant sysctl appearing correct, because its bridge carries per-interface forwarding and the kernel ignores router advertisements on a forwarding interface unless accept_ra is explicitly two rather than one. The fix takes the advertised prefix while declining the default route, so the hypervisor gains an address without any change to how it routes; this was verified after applying, with the v6 default route count still at zero. The database VM refuses key authentication for the privileged accounts and its unprivileged login cannot escalate without a password, so the hook went in through the QEMU guest agent from the hypervisor, which executes as root inside the guest. The document notes the base64 indirection needed to get a multi-line script through intact. |
||
|
|
ffb7fba346 |
docs: give the ESH IPv6 naming scheme a home, and make it real on one host
The scheme has existed since August as a single line of persistent memory, which a snapshot then deleted. It is a naming convention rather than temporal state, so it now lives in docs/pfi as a proper document, and the memory entry is reduced to a pointer at it. The document carries the full table, the address structure, the reasoning about which slots can and cannot hold a name, and the recipe for applying one to a host. It also corrects the conclusion the original note ended on. That note held that these names could never appear on the wire, which is true of everything UniFi is able to assign but not of what a host can assign to itself, and the distinction is the whole difference between a joke and an address. AdGuard on esh-docker-vm now holds the esh-server name, at 2607:73c0:402:1d02:4411:b105:50:45, where the segment identity and the IPv4 address are both legible. It is applied by an if-up.d hook that derives the prefix at runtime rather than hardcoding it, backgrounds itself with a retry so it cannot stall interface bring-up, and adds nothing to the existing interface configuration. This is load-bearing rather than decorative. The gateway advertises an IPv6 resolver to clients, macOS prefers it over the IPv4 one, and it previously pointed at an address derived from that host's MAC. |
||
|
|
ca3c984f93 |
feat(selene): retire the seat; chat-judge -> gen, selene-1-mini-8b 404s by design
Benchmarked selene against gen on selene's own job: 24 designed judge items with checkable ground truth, pairwise + absolute modes, 3 repeats, on BOTH a neutral JSON prompt and Selene's native Atla template. 288 calls, all free local. neutral JSON selene 20/24 (83%) gen 23/24 (96%) native Atla selene 21/24 (88%) gen 22/24 (92%) gen won on both templates and selene's BEST sat below gen's WORST. Selene was given its own fine-tuned template as a fairness check; it gained one point, not the three it needed. Decisive defect: selene cannot emit "tie" -- 0/2 on both templates, forcing a winner on every equivalent pair. For eval work that is the case that matters most. gen returned tie correctly on the JSON template. Selene also compressed the 1-5 scale (clustered at 2s and 4s) where gen used it fully. Selene's only win was ~3x latency, unexercised at ~60 calls/day with zero queueing. TWO NAMES, TWO DIFFERENT TREATMENTS, deliberately: - chat-judge -> repointed to gen. It is a ROLE alias and ADR-0012 says consumers bind the capability, not a concrete model. Sampler profile copied from image-judge (temp 0, top_p 1.0, top_k 1, thinking off) so the served config matches the benchmarked condition. - selene-1-mini-8b -> REMOVED. It 404s. It was NOT aliased to gen. A served-name is a contract about what the model IS; answering it with a different model hides a material change behind a stable string. Operator ruling: "never repoint a named model at a different model's endpoint -- that is intentionally misleading." Verified: the gateway now returns HTTP 400 "Invalid model name" for it. Reclaimed 17.2 GiB on ana-ml2 GPU 1 (free 1,818 -> 19,450 MiB) on a card that had under 2 GiB of headroom. gen already runs on GPU 0, so the judge role moved onto an existing seat rather than allocating anything new. Canonical litellm config synced from the host; ana-ml2 README and recommended-model-settings updated. compose.yaml kept for reference, not deployed. |
||
|
|
ab3a0ca5bc |
docs(quant-playbook): acceptance is not throughput -- always run the depth control
Measured 2026-08-22 on one target with one instrument: raising MTP num_speculative_tokens from 3 to 7 improved accepted length from 2.753 to 3.041 per forward pass while throughput fell from 114.9 to 74.0 tok/s. Reporting acceptance alone would have recommended a 36% regression. The cause is architectural rather than model-specific. A single-module MTP head has no depth of its own, so vLLM runs it autoregressively and k draft tokens cost k sequential forward passes. Past a shallow depth the drafting cost exceeds what the extra accepted tokens save. Records the comparison rule that follows: match k when comparing two speculative methods, or the measurement is of depth rather than method. A parallel-drafting drafter at k=7 against an autoregressive MTP at k=3 is not a method comparison. In the case that produced this, the depth control showed most of the apparent acceptance advantage was depth, while the throughput advantage was real and came from parallel drafting -- our MTP was better at position 0 and still lost overall. Only the measured, model-agnostic result is recorded here. The DFlash2-specific findings, the hypotheses that remain unproven, and the wrong turns taken along the way live in persistent-memory.d/2026-08-22-dflash2-spec-decode.md with explicit epistemic labels, deliberately kept out of the playbook. |
||
|
|
0755ba7d00 |
fix(quant): stop baking the calibration truncation cap into the shipped tokenizer
load_calib tokenizes with tok(..., truncation=True, max_length=seqlen). For a
fast tokenizer that mutates the Rust backend's truncation state in place, and
the subsequent tok.save_pretrained() persisted it, so every mixed-NVFP4 build
shipped a tokenizer.json carrying
"truncation": {"direction": "Right", "max_length": 2048, ...}
against a source whose value is null. Every prompt was clamped at the
calibration length, permanently.
It hid because older transformers does not enforce the text-vs-ids count
check. On a newer one the seat dies at startup with a message that names
images and never mentions tokenizers:
ValueError: Mismatch in `image` token count between text and `input_ids`.
Got ids=[2047] and text=[16384].
The cap also silently limited image resolution well before it killed
anything -- at 2048 the largest servable image is about 1448x1448, since
(edge/patch)^2 / merge^2 image tokens have to fit under it.
Fix saves a pristine tokenizer re-read from the source rather than the
mutated calibration object, and then asserts truncation is null so the
defect fails the build instead of shipping again.
Playbook gains section 3.14 with the symptom, the cause, the audit one-liner
and a table of which builds were affected, plus a fourth mandatory post-step.
The transferable lesson is called out: this is the third case of an artifact
carrying config authored against an older transformers that a newer one
begins enforcing, so an image bump is a config-compatibility event rather
than just a version change.
|
||
|
|
36c173c6a1 |
feat(mog-sec): quant + serve M.O.G.-SEC pen-test seat; PPL on gen; retire fable
Autonomous overnight run under the operator's full-autonomy grant. End state:
fleet up, gen seat untouched, a new verified pen-test seat serving where fable was.
PPL on the orcarouter gen seat (fable downed to free GPU1 for a nospec probe,
probe torn down after): mean 7.07 / median 5.76, within noise of heresy 6.910 /
5.625 and identical to our recipe's usual 7.059. The gen-seat search is settled.
M.O.G.-SEC: chose Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX-BF16 (rev deede677)
over the pre-made ModelOpt NVFP4, which was disqualified on W4A4 4-bit activations
(the AEON degradation mode, catastrophic on a 1M-context model), zero MTP tensors,
and ModelOpt format. Pulled, format-screened (P(<think>) 1.11e-05, clean), quanted
in-house to mixed NVFP4+FP8 (23.4 GB, MTP + vision preserved), and served in the
retired fable slot.
stacks/mog-sec ana-ml2 GPU1 :8019, KV 418,218 tok / 1.60x @ 262K
aliases mog-sec (non-thinking), mog-sec-reasoning (thinking)
gates surface 6/6, MTP 55.3%, format 0/15 leak, vision 7/3/1,
capability 4/4 (delivers offensive-security content)
Served at native 262K, NOT the card's 1M -- the 1M needs YaRN (absent from the
weights' config) plus the SGLang/DFlash2 path the repo ships a deployment kit for,
neither of which is our vLLM surface. A real 1M seat is a separate SGLang project.
Retired char-rp-reasoning + char-rp-fable (zero traffic, pointed at the downed
fable :8019; now 404 cleanly, not repointed -- a security model is not an RP model).
char-rp (meromero) untouched. Vision preprocessor built from the model's own
image_processor block, same trick as the MeroMero seat.
GPU0 seats (gen, meromero) were untouched and healthy throughout. The quant ran in
GPU1 free space with no production seat stopped except fable, which was replaced.
|
||
|
|
bf65d0254d |
docs(pfi): evaluate the two gen-seat replacement candidates
preetpatel/Qwen3.8-27B-Uncensored-NVFP4 is disqualified on two independent hard
failures, both read directly off the artifacts via HTTP Range requests against the
safetensors header (about a megabyte, not a 20 GB download):
- ZERO mtp tensors. The author's recipe.yaml asks to ignore re:.*mtp.*, but the
written config.json has no mtp ignore entry while re:.*visual.* expanded to 110
explicit ones. That asymmetry is llm-compressor pruning a pattern that matched
nothing, i.e. the MTP head was never loaded. Costs roughly half our decode.
- NVFP4 W4A4, 4-bit activations. Precisely the AEON failure mode: the fidelity
gradient is W4A4 < W4+FP8 < W4+bf16, W4A4 drove ~15-20% stochastic degeneration,
and it collapses past ~30k context. The gen seat serves 262K.
orcarouter/Qwen3.8-27B-Uncensored checks out as a quant source: stock-Qwen base
rather than a reasoning-compression finetune, Arditi-style single-direction
abliteration, 15 mtp and 333 visual tensors verified present, chat template
byte-identical to the heresy build we are serving, and the gate is already accepted
on our token.
Also records the author's FP8 release as a noted-but-not-recommended third option:
far more traction, but 30.9 GB against NVFP4's 22 GB, and on a zero-sum GPU0 that
+9 GB comes out of the KV pool and breaks 262K context.
And states the imatrix constraint plainly. Our recipe has always requested
imatrix_mse and always silently fallen back to uniform MSE; playbook 3.13 warns
against assuming an imatrix would help before verifying llm-compressor can consume
external importance data at all. The W4A16 portions are data-free by construction
and cannot use it regardless.
|
||
|
|
f90a5025de |
feat(coldfusion-abliteration): Heretic-300 — 8/100 refusals at KL 0.0136, beats the heresy bar 3.6x
Ran Heretic v1.4.0's 300-trial TPE search on Cold-Fusion. Best trial scores 8/100 refusals at KL 0.0136 against a 98/100 base, versus absolute-heresy at 29/100 and our hand-tuned Robinson L35 at 72/100 / KL 0.0116 — i.e. 64 fewer refusals for the same damage. Hand-verified coherent: correct arithmetic with shown working, clean code, 66-167 word prose across nine probes. Durable findings: - direction_scope=0 (single shared direction) is decisive on this merged base: n=129, best 8/100. Per-layer directions n=131 never beat 52/100 despite a better median. Points against the multi-direction intuition for a diffuse direction (our two-template |cos| is 0.62 vs Robinson's 0.99 on stock). - Aggression is not the lever. r(KL, refusals) = -0.561 over 261 trials; the KL<0.02 band contains both the worst results (median 87/100) and the single best. A KL 0.3554 trial scored worse than one at 0.0193. - PR #317 confirmed: Heretic silently drops the MTP head on save. Source 1199 tensors -> export 1184, all 15 mtp.* gone, vision 333/333 intact, exit 0, no warning. This is also why absolute-heresy ships a byte-identical MTP head — a bug, not a design choice. Always diff tensor keys after a Heretic export. - Heretic's recovered direction carries 6.18% of its energy in sink dim 3994, versus 0.094% for our L35 and 1.97% for the L39 we rejected as brick-inducing. It survives that only because of magnitude-preserving ablation (row_normalization=FULL); our plain projection has no such protection, so the sink screen correctly refused the in-band MTP graft. Same direction, different operation. MPOA is the prerequisite for in-band MTP on a Heretic trunk. - Heretic's edit is recoverable from weights: delta is rank-1 (s2/s1 ~ 0.010), SVD gives the direction, norms give per-layer weights (1.08 -> 1.34, i.e. over-projection). Cross-layer |cos| agreement 0.9903 independently confirms the single-direction result. New tooling in services/coldfusion-abliteration/: kl_divergence.py first-token KL, class-split, zero noise floor catatonia_gate.py 12 probes x 220 tokens, prints every completion heretic_export.py PTY driver; selects by measured value, never by menu position — Heretic's resume prompt puts "delete the checkpoint and all results" one arrow-key from the target graft_mtp.py recovers the trunk direction by SVD; --pristine for the safe path when the sink screen refuses Also adds quant playbook 3.13: the NVFP4 recipe sets observer="imatrix_mse" but llm-compressor has always silently fallen back to uniform MSE for want of importance data — on this build and on the incumbent. Existing A/B comparisons stay valid since every build shares the fallback. Parked as id 42. Guardrail note: this build has lost the self-harm guardrail that the Robinson L35 build retained. Restoration is the operator's own work item. |
||
|
|
1b3fb270e7 |
feat(coldfusion-abliteration): first-token KL measured — 28.4x selectivity, harmless median 0.0211
Adds `kl_divergence.py`: first-token KL(stock || abliterated) over the full 248,320-token vocabulary, bf16 vs bf16, scored separately for held-out harmless and reserved-harmful prompts. Result (L35, 256 harmless / 104 harmful, answer mode): harmless median 0.0211 mean 0.0364 top-1 agreement 89.8% harmful median 0.5996 mean 0.6992 top-1 agreement 55.8% selectivity 28.4x (72.8x in think mode) Self-KL noise floor is exactly 0.0, and all 720 per-prompt values are bit-identical between a single-process and a two-process run, so the figures are signal rather than bf16 jitter. Reverse KL on harmful/answer is 1.43 vs forward 0.70 — the mass-where-stock-had-none asymmetry expected of a refusal-direction removal. Against the Heretic reference figures (0.1191 prior seat, 0.0759 the live absolute-heresy seat) this is materially gentler, but those are the other tool's optimizer output on a different base with its own harmless set and template — order-of-magnitude, not head-to-head. KL remains a fidelity number; the viability gate is still MTP acceptance (59.1%). Method notes: - Prompt classes are reported separately by design. A single averaged KL over a mixed corpus is close to meaningless, since the metric is meant to be large on harmful prompts and small on benign ones; the ratio carries the information. - The harmless evaluation set is drawn from the alpaca pool minus calibration's own draw, reconstructed by replaying that draw rather than remembered, and asserted disjoint on text. The harmful set is the reserved test split. - `render` is imported from abliterate.py rather than copied, so the measurement cannot drift from the rendering the direction was captured against. - Batch size 1 with logits_to_keep=1: no padding semantics, ~0.6 MB of logits. Three corrections to the runbook, each of which cost time: - "bf16 is 50 GB, only gen must go" was 50.10 GiB mislabelled. Text-only weights are 51,300 MiB; freeing either GPU0 seat alone leaves ~50,933 MiB. Both must stop. VRAM is now sized from the safetensors headers at run time. - A 27B model cannot be released in-process: `del` + gc + empty_cache left free VRAM at 45,287 MiB, and so did confining the model to an inner frame that exits. Only process exit returned the card (96,689 MiB). The first run completed only because the allocator hit OOM, collected, and retried. Each model now gets its own process, handing log-probs to disk between stages. - The residency gate read hf_device_map, which transformers leaves empty when the model fits on one device — it reported "(unsharded)" whether or not anything was wrong, so it could never fail. It now reads parameter devices directly. Model-agnostic lessons promoted to the quant playbook (new 3.12). |
||
|
|
e9dbc8660b |
feat(coldfusion-abliteration): abliteration LANDS at layer 35 — separation selector, shard-surgery write, three false diagnoses corrected
The abliterated model works. A/B vs stock on a matched greedy battery: explicit sexual + graphic torture (the measured stock refusal surface) go from refused to complied/engaged, held-out AdvBench prompts loosen, the self-harm guardrail survives, coherence intact — the Robinson design point exactly. Output at /tank/aimodels/qwen38-27b-coldfusion-abliterated-L35-bf16, verified bitwise: 131/131 targets changed, 333/333 vision byte-identical (delta 0.0), 735/735 others untouched. Getting there corrected three diagnoses the prior session had backwards. 1. The layer-selection metric was wrong, and that was the whole ballgame. The recipe picks the abliteration layer by peak two-template |cos| agreement. On this heavily-merged base that metric is anti-correlated with efficacy: its argmax (layer 18) is the WORST-separating layer in the window (Cohen's d 5.51 vs 9.89 at the peak), and abliterating there was a measured behavioral no-op — stock and "abliterated" refused all six probes identically. Cause: the two renderings end in different generative modes (</think> vs <think>), so |cos| scores answer-vs-reason mode, not refusal, and on a merge the mode term dominates. Replaced selection with harmful/harmless SEPARATION (Cohen's d / AUC of the direction's projection), gated on the sink screen since separation and sink-energy both climb with depth. Picks layer 35 (d 9.35, AUC 0.9997, sink 0.094%). Agreement is kept as a printed diagnostic. 2. The "bf16 NaNs, use fp32" rule was a misdiagnosis. The NaN was never precision — it was multi-GPU sharding (the residual stream zeroes two layers past the GPU0->GPU1 boundary; the first capture's layer 22 happened to sit in the healthy region, which is why it looked fine) plus PYTORCH_CUDA_ALLOC_CONF=expandable_segments (corrupts retained tensors; the corruption MOVED between bit-identical forwards, the tell that it was memory not math). On one GPU with a plain allocator, bf16 full-64-layer is exactly deterministic and coherent, at 50 GB and 4.3x the throughput of the 111 GB fp32 it replaced. Both defects are now hard gates (residency exit 8, allocator exit 9); capture pins CUDA_VISIBLE_DEVICES=0. 3. The corpus-size hypothesis was falsified. 52x more calibration data (8->416, mlabonne/harmful_behaviors = the recipe's actual AdvBench split, already on the box) moved agreement 0.594->0.624 — nothing. Kept the 416/416 corpus anyway (calibration.py); it gives the clean separation signal. The held-out 104-prompt test split is reserved and asserted disjoint. Also: the --out write is now shard-level surgery (reads/writes the 18 safetensors directly, no model object, no GPU). This is correctness, not thrift — AutoModelForCausalLM resolves to the TEXT model, so save_pretrained would drop all 333 vision tensors AND skip the MTP head (the in-band MTP edit is the entire point of the Robinson formula). Neither failure raises. Shard surgery makes vision and the other 1068 tensors byte-identical by construction. Batched capture with a dtype-aware equivalence gate; hidden states captured via forward pre-hook (reading output_hidden_states off the returned object is unsafe here — buffers get recycled). Sharding/allocator lessons promoted to the quantization playbook (model-agnostic, sections 3.9-3.11 + superseded table); the selection-metric lesson added to the recipe doc. The dead layer-18 no-op checkpoint was removed (52 GB, confirmed identical to stock). Incumbent gen seat untouched. Full canonical refusal-probe re-profile and MTP-acceptance-on-quant still owed before this becomes a gen-seat candidate. |
||
|
|
ccb56a0a51 |
docs(pfi): capture the RobinsonLabs Qwen3.8-27B abliteration recipe
Reference recipe (not a deployed artifact) for MTP-aware, vision-preserving single-direction abliteration of Qwen3.8-27B -- the base family the gen seat runs. Captures the two things this recipe gets right that naive abliterations of this architecture miss: - The MTP head is abliterated in-band (its two residual-write matrices, glue left alone), so speculative acceptance does not collapse on the prompts abliteration exists to fix -- directly relevant to the gen seat's MTP>=40% gate. - The vision tower is preserved byte-identical (333 tensors, max delta 0). Plus the two calibration traps specific to this base: the twice-captured refusal direction (layer 26, |cos| 0.99) and the attention-sink dimension 3994 that bricks the model if orthogonalized out. Documents the coverage gate (o_proj 16 + linear_out 48 == 64 layers) that catches a half-abliterated model before it writes a byte, and the foot-gun that the GGUF imatrix does not cover the MTP block. Links into model-quantization-playbook.md for the quant half of the pipeline. |
||
|
|
fb91ea759e |
docs(pfi): lesson 10 -- v6 collapses two exposure controls into one
Operator's framing, and it is a better argument than the terminology correction that preceded it. Under v4, exposing a host needed two affirmative acts -- a DNAT and an accept rule -- so missing either left the host dark. There is no v4 misconfiguration that exposes an internal host by accident. NAT was load-bearing security whether or not anyone designed it that way. v6 removes the first control entirely. The path exists inherently, so the firewall is the only thing left, and the failure mode inverts from fail-closed to fail-open. Rule-ordering slips, rulesets that silently match only one address family, new VLANs added without policy, and re-delegated prefixes unmatching address-literal rules all become exposure events rather than no-ops. Records the practical consequences: key rules on interface/zone rather than address literals, treat enabling v6 on a segment as requiring policy to exist first, and verify default-deny from off-net rather than by reading the ruleset -- which is lesson 3's assert-the-effective-value discipline applied to firewall policy. Also corrects my own claim from the previous commit that the pending firewall pass was 'smaller' than I had implied. It is not smaller, it is different in kind. |
||
|
|
8a742f59b8 |
fix(ana-gw): restore ESH<->colo IPsec as a dialup tunnel with NAT-T
The link died when ESH lost its public IP during the fiber cutover. Two independent causes, and the second would have defeated the obvious fix: - phase1 ana-to-eshudm was type static, pinned to 70.181.90.232, an address that no longer exists. - nattraversal was disable, so ESP could not have crossed NAT even with the peer IP corrected. pfi-ana-nh3 shares that setting and survives only because NH3 is publicly addressed, which is why the two tunnels diverged. FortiOS refuses `set type dynamic` on an existing tunnel -- "Cannot change tunnel type once configured" -- and rolled back cleanly, so the fix could not be an edit. Rather than delete and recreate, which cascades into the phase2, two static routes and ten policies, the replacement was built alongside: new phase1+phase2 ana-eshudm-dyn (type dynamic, ikev2, aes256-sha1, dh14, NAT-T on, PSK read from the ESH UDM API so neither side needed a new key), static route id 10 at distance 20, and two consolidated multi-zone policies 73/74. The old tunnel is left in place, dead and harmless, as rollback. Verified up: ana-eshudm-dyn_0 97.170.236.56:4500 selectors 1/1 -- the _0 suffix is a dialup child, :4500 is NAT-T, and the address is the carrier's, which is precisely what could never have been pinned. ESH reaches all four colo hosts at 40-56ms, the colo reaches all three ESH hosts, and traceroute drops from eight hops leaking into the carrier network to three hops fully encapsulated. Config was backed up before any write (1.17MB, 36903 lines, off-box). Residual fragility recorded: the UDM's ipsec_local_ip demands a literal address -- empty is rejected as api.err.InvalidPayload -- so it still needs updating when the fiber changes ESH's WAN address. The gateway end is now address-agnostic; the UniFi end is not. |
||
|
|
dec4ba45db |
docs: scope the NAT refutation to WireGuard; IPsec to colo is broken
Correcting an over-generalisation from earlier today. Proving that NAT does not break Site Magic, I wrote it up as "no addressing outcome threatens the inter-site tunnel." That is wrong: the fleet has two inter-site links with opposite NAT behaviour. - NH3<->ESH is Site Magic, i.e. WireGuard. It survives arbitrary NAT, proven live on RFC1918 double-NAT (192.168.200.111) with nh3-dev and nh3-docker reachable at ~40ms. It dials out to NH3's public edge and never needs inbound reachability. - colo<->ESH is IPsec on the ana-gw FortiGate, and it is broken right now under those same conditions. ana-docker, pfi-pve and pbs-ana all fail from esh-pve-nas, and traceroute shows packets for 10.250.x leaving the UDM to the 5G modem and then wandering the carrier network before dying -- not encapsulated at all, so no SA is up and the traffic falls through to the default route. Site-to-site IPsec pins a peer IP and ESH no longer has a routable one. So the IPv6 work keeps its justification, but on the IPsec link specifically rather than on the tunnels generally. Operator caught the over-generalisation. Adds lesson 8 -- a result proven for one protocol does not transfer to another -- and corrects the superseded-claims row rather than replacing it, since the original claim was half right and the halves are the point. Also records my own over-broad claim as its own superseded row. |
||
|
|
40a4121a43 |
docs(pfi): add lesson 7 — test a 'this will break X' premise before building on it
Seeded by the Site Magic / CGNAT premise, which justified a body of IPv6 work and turned out to be false the first time anything actually tested it. The mechanism was discoverable in advance: Site Magic is WireGuard and the far side has a public endpoint, so the NAT'd side dials out and never needs inbound reachability. NAT breaks inbound; it does not break outbound-initiated tunnels with keepalives. Also fills the first row of the superseded-claims table, which is what that table exists for -- the claim is corrected with a date rather than quietly deleted, so older references to it resolve instead of misleading. |
||
|
|
0559e12a2d |
docs(pfi): add an ops-lessons playbook for the transferable failures
Sibling to model-quantization-playbook.md, and it exists for the same reason that one does: hard-won lessons were dying inside per-host runbooks where nobody finds them until after repeating the mistake. Six entries seeded from the esh-pve-nas migration, all of which would bite identically on any other host: 1. mount --rbind into a chroot needs --make-rslave, and losing cgroup2 impersonates failing root-disk I/O closely enough that it was misdiagnosed as exactly that. 2. A reboot is not confirmed until the host is observed DOWN; "never rebooted" and "rebooted fast" are indistinguishable otherwise. 3. Assert the effective value, not the presence of a substring. Grep proves presence; only evaluation proves effect. 4. Ask the server who its clients are -- documented dependent lists rot. Plus the corollary that an idle hard NFS mount blocks and resumes, so quiescing means stopping consumers, not always unmounting. 5. The scoped-looking command can be the dangerous one; setting a ZFS cachefile on one pool of three would have stopped the other two from importing at boot. 6. Long uptime hides breakage, and a forced look is worth more than it appears -- one migration surfaced an 82-day-dead pvestatd, a 126-day hung vzdump, a VM in prelaunch for four months, and an undocumented cluster, none of them caused by the work. Carries a superseded-claims table so corrections are dated rather than silently edited, same discipline as the quantization playbook. The ESH runbook now links here so the general rules are reachable from the specific story and vice versa. |
||
|
|
5f11d1b3cb |
feat(esh-pve-nas): cut PVE root over to ZFS on the mirrored NVMe
Root is now nvme/ROOT/pve-1. The USB DOM keeps the ESP and /boot but is out of the runtime I/O path, so a bus reset can no longer drop root from under a running hypervisor. All five guests healthy, three pools ONLINE, system running, ext4 pve-root intact and unmounted as the rollback with its own kernel and initrd. zfs-import-cache is now the active import path -- the all-three-pools cachefile fix doing its job. The window cost an unplanned outage, and the cause was this repo's own tooling rather than the migration. The staging chroot ran `mount --rbind /dev` and /sys with no --make-rslave. On systemd `/` has shared propagation, so the cutover's `umount -R` propagated back into the live host and removed the real /sys/fs/cgroup, /dev/pts and /dev/shm. With cgroup2 gone systemd-logind could not create a session: ping fine, TCP fine, SSH authentication succeeded, resident daemons kept serving -- and every new exec hung, including /sbin/reboot, so the reboot never ran at all. It impersonates failing root-disk I/O almost perfectly, and I called it as the DOM dying. That was wrong. dmesg had the answer throughout: the DOM attached cleanly with no errors, and the last log timestamp was 12114881s -- 140 days -- meaning this was still the original boot. A down-detector had also never reported the host down, which I read as a fast reboot rather than as no reboot. Fixes and guards: - --make-rslave after every rbind, plus a guard that refuses to proceed while any chroot bind still reports shared propagation. - Confirm a reboot by observing the host DOWN, not by watching for it to come back. Those two states are indistinguishable otherwise. - Blast radius now measured from the server: `ss` inside CT 103 found five NFS clients, not the two documented. The new one that mattered is esh-vm-db, hard-mounted and unreachable by ssh. Left mounted on purpose and it came through read-write. - grub-reboot's one-shot does NOT work here: grubenv sits on an LVM LV which GRUB reads but cannot write, so next_entry survived the boot that consumed it. Steady state is saved_entry=pve-zfs-root with no next_entry. There is no auto-fallback on this host and no IPMI. Recovery needed no console: an idempotent cgroup2/devpts/shm remount landed in the brief windows where exec succeeded. No data was lost, and neither the DOM nor any pool was ever at risk. |
||
|
|
c4b2278e7d |
feat(esh-pve-nas): stage the PVE root migration off the USB DOM
Everything but the reboot. Two rerunnable elway playbooks; the host is
still running from the ext4 root and its boot path is byte-identical to
the last 140 days, because grub-install is deliberately held back to the
cutover window.
Phase 1 (esh-pve-nas-stage-zfs-root.yaml): carve a 512 MB /boot LV out
of the 768 MB swap LV, populate it, rsync the 4.3 GB ext4 root into
nvme/ROOT/pve-1, write the copy's fstab.
Phase 2 (esh-pve-nas-stage-bootloader.yaml): ZFS initramfs, grub.cfg,
explicit pve-zfs-root and pve-ext4-rollback entries with stable ids,
grubenv pinned to the rollback so cutover's grub-reboot is a one-shot.
Three landmines the plan did not predict, all caught by verify steps
asserting effective state rather than by reading the plan:
- The /boot LV had nowhere to live. VG pve had 4 MB free and mounted
ext4 cannot shrink; freeing space from root needs a rescue boot, which
costs the one-reboot property. Space came from swap (768M -> 256M).
- The runbook's `zpool set cachefile=... nvme` would have broken the
NAS. Populating a cachefile flips the host from import-by-scan to
import-by-cache, so a one-pool cache leaves ssd and tank unimported --
and CT 103 esh-nas has twelve bind mounts spanning all three pools.
Set on all three instead, verified in the resulting cache.
- update-grub silently emitted a pool-less root=ZFS=/ROOT/pve-1, which
boots to an initramfs prompt. Debian's 10_linux builds ${rpool}${bootfs}
and rpool comes from grub-probe --target=fs_label, which returns empty
because GRUB's ZFS reader cannot open a pool with encryption,
large_dnode and zstd_compress -- the same feature set that forced /boot
to stay ext4. The probe failure is swallowed by `2>/dev/null || true`.
Fixed with a /etc/default/grub.d drop-in plus explicit menu entries.
The transferable lesson: the original verify grepped for the correct
root= string appearing somewhere in grub.cfg, which passes while every
menu entry is still broken. Assert the effective value, not the presence
of a substring.
|
||
|
|
8ddc87c852 |
docs(esh-pve-nas): record the blocked-patching driver and the upgrade ordering
The operator-visible symptom is that PVE cannot be updated on this box for lack of room. Measured: 225 packages pending, 161 carrying deb12uN/Debian-Security bumps including ssh, against esh-pve's 8.4.14 versus this host's 8.4.11 and 20 weeks of uptime. Records the ordering explicitly -- migrate first, upgrade after. The pending set includes proxmox-kernel-6.8.12-42-pve-signed, roughly 250 MB of kernel plus initramfs landing in /boot which is on root with 1.3 GB free. Unpacking 225 packages including dpkg and perl into that headroom risks filling the disk mid-transaction and wedging dpkg on a hypervisor running five guests. Notes the apt archive-dir redirect as a partial escape hatch if patching cannot wait, and that zfs-initramfs 2.2.8 is fully capable of root-on-ZFS so there is no need to upgrade ZFS before migrating. |
||
|
|
3e311756d7 |
docs(esh-pve-nas): split boot from root instead of reinstalling
Operator's proposal, and it is strictly better than the reinstall plan. Boot and root do not have to share a device. Keep the ESP and /boot on the DOM as ext4 -- so GRUB never has to read ZFS, which matters because the nvme pool has encryption, large_dnode and zstd_compress enabled and GRUB cannot read those -- and move root to nvme/ROOT/pve-1. The initramfs imports the pool and pivots. What this buys over the reinstall: the nvme pool survives, so no guest migration, no export/import of ssd and tank, no reinstall. Downtime is one reboot rather than half a day. Rollback is a GRUB menu entry, because the ext4 root stays on the DOM untouched. And it retires the actual top risk -- with root on NVMe, a USB bus reset mid-run no longer takes the running system down; the DOM becomes read-mostly, written only on kernel updates. Preconditions verified and already met: UEFI with grub-efi, zfs-initramfs 2.2.8-pve1 installed with 76 ZFS files already in the running initrd, root only 4.3 GB to copy, swap negligible against 125 GB RAM. Two traps recorded: canmount=noauto on the root dataset or ZFS tries to mount over the running root; and cachefile is currently none with a 0-byte zpool.cache, so the pool imports by scan today and must be given a cachefile before the initramfs is rebuilt. The reinstall plan is retained as the fallback. |
||
|
|
2275e11be0 |
docs(esh-pve-nas): plan the migration off the USB DOM; flag the NFS blast radius
PVE root on esh-pve-nas is a USB Disk-on-Module: 6 GB ext4 with the host's only ESP. A DOM is SLC/pSLC so wear is not the driver -- the problems are that it is on the USB bus (a reset drops root under a running hypervisor), has no headroom, and is unmirrored while 928 GB of mirrored NVMe sits 96% empty. Runbook targets a fresh PVE install to ZFS RAID1 across both NVMes. In-place conversion is unsupported, and adding an ESP to the existing NVMes is impossible -- both are whole-disk ZFS members with 1.7 MiB free and proxmox-boot-tool manages nothing today. The headline risk is not on the host being rebuilt: CT 103 esh-nas IS the NAS at 10.0.50.50, and both esh-docker-vm and esh-pve mount it hard. Taking this box down stalls esh-pve's storage layer and wedges esh-docker-vm into the D-state whose only remedy is a host reboot -- the incident shape already on record. Quiescing those clients is step one of the window, and the README now warns against casual reboots. Config snapshot captured off-box to nh3-dev (0600) with /etc/pve, network and fstab config plus zpool/zfs/disk-by-id/guest state; the newest on-disk copy before this was June 2024. |
||
|
|
2f2bbce73d |
docs(quant): correct 3.8 — TWO real causes, not a lone defective quant
Operator correction to the prior 3.8 framing (
|
||
|
|
d28a371049 |
fix(gen-seat): AEON W4A4 was the defect — purged; mixed FP8-attn build is primary gen
Root cause of the multi-day degeneration hunt, operator-confirmed: the AEON NVFP4 W4A4 quant (sakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4, full W4A4 incl. attention) went degenerate ~15-20% of generations in real multi-turn use and forced regenerates. MTP, prefix-caching, and the gateway all merely AMPLIFIED it, which is why MTP-off, APC-off, and the vLLM #51113 fix each 'helped' a synthetic probe without fixing it -- three plausible false root-causes, each passing one clean run then failing in real use. The fix was the WEIGHTS: the in-house JonathanColetti/Heretic mixed NVFP4+FP8 build (qwen38-27b-uncensored-nvfp4-mixed, FP8 attention not W4A4, same base, same MTP) is coherent through long multi-turn with MTP ON. W4A4 *attention* was the defect; FP8 attention is not. This commit: - GEN_MODEL -> the mixed FP8-attn build (primary gen until DavidAU 3.8 lands) - GEN_IMAGE pinned to vllm/vllm-openai:nightly-311b3513... (v0.27.2rc1.dev150, carries #51113; pinned by sha so it does not drift on the next pull) - AEON weights PURGED from /tank (no-good), safety-checked not-in-use first - playbook 3.8: the stochastic-W4A4-degeneration lesson + isolate-weights-early + do-not-declare-a-fix-from-one-probe (it validated three non-fixes) AEON is re-pullable from HF if ever needed, but the operator ruled it no-good. |
||
|
|
63a3cb2d86 |
fix(gen-seat): MTP mitigation — disable prefix caching, keep MTP (speed restored, multi-turn clean)
The qwen3_5_mtp corruption (playbook 3.7) is gated on MTP x prefix-caching
TOGETHER (vllm#43559 / #47194), per both cross-frontier peers. Disabling
prefix caching (--no-enable-prefix-caching; vLLM V1 defaults it ON, so the
explicit --no- form is required) forces the GDN cache into a mode where the
partial-accept align-path bug is inert, so MTP can stay on.
Verified on our stack (AEON W4A4): MTP on + prefix-caching off -> the 7-turn
varied series stays coherent through 3.9k tokens, zero cross-turn bleed, at
104.6 tok/s / 53.6% acceptance -- the FULL MTP speedup restored (vs ~half
with MTP off), losing only prefix-cache reuse. All 7 aliases route.
Ruled out on the way: num_speculative_tokens=1 (corruption is
depth-independent, n=1 and n=2 both corrupt); switching to SGLang (vLLM /
SGLang / llama.cpp mainline all share the architectural GDN-rollback bug).
Proper upstream fix (#51113) is in main / v0.27.2rc0 only, not stable, so we
hold at APC-off rather than jump the fleet gateway to an RC.
Supersedes the MTP-off config from
|
||
|
|
a8ed6e7428 |
docs(quant): record the MTP-corrupts-Qwen3.8-multi-turn lesson (playbook 3.7)
The single hardest bug of the night, and invisible to the existing acceptance gate: a LOADED, healthy-accepting MTP head still corrupts Qwen3.8-27B multi-turn output past ~2k cumulative tokens (length collapse + cross-turn content bleed), while single-turn is perfect. Model-independent across all three of our Qwen3.8 quants; Qwen3.6 on the same qwen3_5_mtp method is clean; disabling MTP fixes it. New rule: gate MTP on a multi-turn coherence probe, not just single-shot acceptance. |