main
860 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9e986d8ee8 |
feat(phasefinal-web): cloudflare edge config — cache ruleset + always online
Cache rule on www.phasefinal.com with edge and browser TTL both respect_origin, so cache policy stays declared once in nginx.conf rather than split between the repo and the dashboard. Always Online enabled, which is what actually survives an origin outage; a 300s document TTL alone would only mask five minutes. Verified: document and assets both reach cf-cache-status HIT, apex 301s to www, edge email obfuscation active. |
||
|
|
524aa4d860 |
fix(phasefinal-web): healthcheck targeted ::1, so traefik skipped the container
The healthcheck used http://localhost/, which resolves to ::1 in nginx:alpine while nginx listens on IPv4 only — so it never passed, the container stayed unhealthy, and Traefik silently declined to create a router for it. That presents as a broken docker provider: correct labels, right network, no route, no error. Target 127.0.0.1 explicitly and add a start_period. Adds the apex router (301 phasefinal.com -> www) and drops the file-provider workaround, which was mitigating the wrong diagnosis. |
||
|
|
7ffbee6f09 |
feat(phasefinal-web): corporate site stack on ana-docker
Single static page (nginx) fronted by Traefik at www.phasefinal.com, built from the design brief. Site markup/CSS checked in verbatim from the design session; fonts self-hosted (SIL OFL) with the @font-face block uncommented, which every fresh export re-comments. Routed via a Traefik file-provider config rather than the container labels: the docker provider on ana-docker was not registering newly-created containers, so the file router avoids restarting shared ingress. Labels are retained in compose so the file can be dropped once that is fixed. |
||
|
|
c488eadc31 |
memory: snapshot — althing v3 fleet-wide at 3.1.1, sec on GPU0
The in-flight section was a day stale: it still described run 3c as the live subject on a box where nothing had moved. Rewritten around what is actually true now -- v3 deployed fleet-wide, the post office relocated to nh3-docker, sec serving on GPU0, run 3c still held on power. Six new decision entries, three of which carry findings that outlive their incident: the inbound half of the handle-resolution bug (a stale ALTHING_HANDLE reads another agent's mailbox and reports it empty, which is a second route into the failure v3 exists to prevent), the OOM attribution to Claude Code sessions, and the operator's two explicit belays recorded so a later session does not re-raise them as new. Auto-archival fired at 301 lines and moved exactly one entry. Three others were old enough and every one carries a still-open deferred pointer -- the parked CI flip, muninn-gate's submit path, and the triton backend deferred to the Ada refresh. Held back per the guards; an over-cap file that keeps live decisions beats a scannable one that lost them. The entry that did move had its deferred item closed today: nh3-extdev's staged v2.1.2 wheel is moot now that the box runs 3.1.1. |
||
|
|
583f329d00 |
fix(playbooks): 3.1.1 deploy — and why the markers match presence, not count
Deploys althing-core 3.1.1 to nh3-extdev. Both markers present on both boxes; the warning verified behaviourally in four conditions rather than by grep alone -- mismatched inherited handle warns, matching handle silent, explicit --handle silent, unlaunched directory silent, and the warning precedes the output it is about. The release's own verification line says `grep -c handles_launched_at dev_launch.py # 2+`. The real count there is 1, the definition; the other two occurrences are in postbox.py. The installed tree is byte-identical to the repo at the pushed tag, so the instruction is wrong rather than the install. This playbook matches on presence via grep -q, so it passed. Had it asserted the stated count it would have reported FAILED on a perfect deploy -- a verification instruction that fails on correct input, which is the same false-negative this file has now produced three times in different costumes. Recorded above the variable so the next bump does not reintroduce a count. |
||
|
|
8a04d6f1bb |
fix(statusline): resolve the handle from the v3 binding, not the v2 map
The statusline resolved its handle from ~/.althing/session_handles.json.
forseti corrected the grounding and I verified it: that file is a v2
artifact and v3 never opens it. `grep -rn session_handles althing/` is
empty, postbox's resolve_config takes --handle then ALTHING_HANDLE and
nothing else, and `althing-cli use` -- the tool that maintained the map
-- was deleted at the cutover. Whatever is in it now is hand-kept and
drifts silently.
launch-history.json is written by dev_launch, which is the thing that
sets ALTHING_HANDLE in the first place, so it is the real cwd-to-handle
binding. Shape is {cwd: {command: {at, handle}}} with several commands
per directory, so this takes the most recent by timestamp rather than
whichever key happens to sort first. The v2 map stays as a fallback for
its broader coverage.
Worth recording why this was wrong: I wrote the resolution this morning
by reading the v2 statusline block it replaced and keeping its data
source while updating its commands. The commands were the visible half
of the cutover and the data source was not, so it survived a rewrite
that was otherwise about removing v2.
|
||
|
|
c648a40b68 |
fix(playbooks): 3.1.0 herald deploy; the marker check takes a LIST now
Deploys althing-core 3.1.0 to nh3-extdev and restarts the herald. Verified by content on both boxes: POST_OFFICE_HINT 0 -> 4 in post_office_herald.py and resolve_post_office 0 -> 3 in dev_launch.py, dist-info 3.0.3 -> 3.1.0. The check took one file:marker pair. 3.1.0 changed two files, so a single pair would have asserted half a release and passed -- the same half-passing-silently shape as the version-string check it replaced two releases ago, one level up. It now takes a space-separated list, reports each pair individually, and fails if any is missing. Every release's markers so far are recorded above the variable so the next bump is a lookup rather than an archaeology exercise. Also verified the behaviour the release exists for rather than just its markers. The herald writes its address to $ALTHING_ROOT/post-office and dev_launch.resolve_post_office reads it when the variable is unset: env unset -> http://10.100.50.40:8390 env set -> the env value, which wins env set to blank -> the file, because blank counts as unset My first attempt tested this through postbox, which still requires the variable and reported "no post office address is configured" -- correct behaviour that looked like a failed deploy. dev-launch is the reader, not postbox. |
||
|
|
590b55f7d8 |
feat(playbooks): potrace/agg headers, with the two traps that mislead
pypotrace is an sdist that compiles at install time, so every machine and every CI runner resolving it needs these headers first. That makes it a recurring per-box action rather than the one-off it arrived as. Two things learned installing it on nh3-dev are recorded here rather than left in an althing thread, at forseti's suggestion, because a thread is not where the next person looks: Only libagg is a pkg-config consumer. potrace ships no .pc file and is found via potracelib.h directly, so `pkg-config --exists potrace` returns false on a correctly configured box. It looks exactly like the cause and never is. libagg's pkg-config modversion is 2.7.0 while its Debian package version is 1:2.6.1-r134. Comparing those two numbers convinces you the wrong package is installed. The verify phase asserts the geometry, not the import: a square must come back as one curve of four CornerSegments. An extension linked against the wrong thing can import cleanly and return nonsense, so a successful build is not evidence the module works. Getting the build probe to run took three passes and the reason is worth keeping. uv is not on a non-interactive ssh PATH; it is in a different place on each box; and on nh3-dev it sits inside a 0700 home, so even the correct absolute path fails `test -x` for the ssh user because the directory cannot be traversed. The headers are system-wide and root's business, but the build check is a developer action and has to run as the user who owns the toolchain. |
||
|
|
cdeb57c18b |
fix(playbooks): 3.0.3 herald deploy, and a content check that survives releases
Deploys althing-core 3.0.3 to nh3-extdev and restarts the herald. Verified by content on both boxes: PANE_SETTLE_S 0 -> 2 occurrences, value 0.3, dist-info 3.0.1 -> 3.0.3. The content check was hardcoded to the 3.0.1 markers, so from the next release onward it would have kept passing while asserting nothing about what had just been installed -- a check that verifies the previous release is indistinguishable from one that works. It now takes the marker and file as variables, bumped per release, with both releases' markers recorded so the pattern is obvious rather than folklore. That is the same defect class as the install step gated on `postbox` not existing, which this playbook carried until last round: a guard written correctly for the first run and never re-read on the second. The post office container was not touched. forseti established by import graph that althing/post_office/* imports neither changed module -- the fix is in reach_pane, which is herald code -- and the container has been up two hours across both herald restarts. |
||
|
|
9f87ff87e1 |
feat(althing): surface the post office on Homepage under Toolchain
Labels the container into `Toolchain`, an existing group under the existing Toolchain tab -- "the plumbing", which is where a message bus belongs. Confirmed live: Homepage's API now returns it. I had previously recorded in this file that no group fitted, which was wrong. That conclusion came from a grep over the layout block that missed the nested groups, and it went into a comment as though it were a finding. The group was there the whole time. Labels bind at container creation, so this deployed with `up -d` rather than `restart`; a restart leaves the old labels and the dashboard keeps showing what was there before. nh3-docker is already a discovered host in homepage's docker.yaml as `nh3-pfi-docker`, so the label alone is enough -- adding a services.yaml entry as well would render the card twice. althing-chamber on ana-docker also carries Toolchain labels and is a separate service per the operator. Left alone. |
||
|
|
dbb930d546 |
fix(playbooks): 3.0.1 herald reinstall, and two guards that were release-hostile
Reinstalls althing-core on nh3-extdev for 3.0.1 (the pane-route fix) and restarts the herald. Both boxes verified BY CONTENT rather than by version string -- forseti's own checks, grep for _PANE_ID and _live_pid, because a dist-info directory records what was installed, not what the files contain. Both went 0 -> 3 and 0 -> 2. Two bugs in the playbook this run exposed, both of which only appear on the second use: The install step was gated on `postbox` not existing. That guard was correct for the cutover, when postbox genuinely was absent, and wrong for every release after it -- postbox exists now, so a version bump would have silently skipped the install and the playbook would have reported success having done nothing. `--force` already makes the reinstall idempotent, so the guard bought nothing and cost correctness. The post_office variable still pointed at nh3-dev, three hours after the post office moved to nh3-docker. It failed in the verify rather than at install time, which reads as a broken deploy rather than as a stale constant. Worth noting the failure message was the outage semantics working exactly as designed: "This is an outage, not an answer: do not treat it as 'no mail'." |
||
|
|
22da609053 |
feat: registry-push the post office image; version the statusline
## Registry The image moved by `docker save | ssh | docker load`, so a rebuild meant repeating that by hand. It is now published and the compose pulls a digest-pinned reference, so a redeploy is `compose up -d` on any host that has logged in. Pinned by digest rather than by tag: `:3.0.0` is a mutable pointer on a registry anyone can re-push, and this container is the fleet's whole message bus. The tag rides alongside so a human can read what it is. Namespace is claude-bot, not vh. claude-bot's token carries write:package and `docker login` succeeds, but package namespaces are owned -- pushing to vh/ returns "unauthorized: authentication required" after a successful login, which reads like a credential fault and is actually an ownership one. Publishing under claude-bot's own namespace also satisfies the standing directive to stop reusing the operator's personal credentials for infra work, so the constraint and the policy point the same way. Recorded in the compose header so the next person does not read that error as a broken token. Pull path proven rather than assumed: the running container was recreated from the registry reference and its data verified afterwards. ## Statusline Brought under version control because the v3 cutover broke it invisibly. The segment gated on `command -v althing-cli`, a binary the cutover deleted, so the unread badge and the armed bell silently vanished for every session on the box. With 71 of 73 handles pull-only, that badge is the only out-of-band signal telling a session with no armed waiter that it has mail -- a dead statusline made a working bus look like an empty one. Canonical here, live at ~/.claude/statusline-command.sh, copies rather than symlinks per the same rule as stacks/. |
||
|
|
9d4e7bd34a |
feat(althing): move the post office to nh3-docker
Operator directive, and a standing goal: the bus belongs on the docker host. The flag-day deployment put it on nh3-dev because the herald lives there -- but the herald is the piece that must be host-local, and the post office is explicitly the piece that is not. nh3-dev was wrong on three counts. Our own server table calls it "not a Docker-stack host". It has had three OOM events in fourteen days with the interval halving, and the confirmed hog is Claude Code sessions at 5-18 GB, which is that box's actual job. And mem_limit protects the fleet from the post office while doing nothing in the other direction: oom_score_adj was 0, an ordinary kill candidate, on a box whose last sweep took althing-herald and uvicorn. The new deployment sets oom_score_adj=-500. The compose is now version-controlled here as a normal stack rather than living only in the althing repo's deploy dir. ## docker stop does not checkpoint the WAL The database was 155 KB with a 4.1 MB write-ahead log, and every recent message was in the log. A clean container stop left it untouched -- an explicit PRAGMA wal_checkpoint(TRUNCATE) was required. A docker cp of the .db alone would have produced a database that opens cleanly, passes integrity_check, serves the full 73-handle roster, and is missing the day's mail, with nothing raising an error. Row counts were verified at source, in the staged copy, and after seeding, because the count is the only thing that separates those two outcomes. The old volume is left in place. Not a rollback path, which the operator ruled out -- just not deleting the only other copy on the day of a move. ## Follow-up left open The image has no registry push and moves by save/ssh/load, so a rebuild means repeating that by hand. It should join the gitea registry pattern the other stacks use. |
||
|
|
e58360668e |
feat: althing v3.0.0 cutover (U9b) and the sec seat onto GPU0
Two operator-authorised changes on the same afternoon.
## althing v3 (U9b flag day, one-way, no rollback)
The post office replaced the v2 P2P bus on nh3-dev and nh3-extdev.
One container is the only stateful component; heralds are one per box
and dial out; waiters are one per session. Every v2 command was deleted
rather than deprecated, so a script calling althing-cli now fails loudly
instead of silently talking to nothing.
73 handles seeded from the v2 CLI, which is authoritative over the v2
database's 91 agent rows -- the extra 18 are superseded names, a typo,
an underscore variant, and two machine-qualified handles that v3 makes
a category error. Verified by set difference in both directions rather
than by counting; a peer's "72 rendered" was a line-count artifact.
Deleted 5,043 orphaned wake FIFOs. The reason there were five thousand
is that v2 named them per session with the PID and never reaped them;
v3 names them per handle, so the leak is bounded by construction. That
is a fix in v3, not a cleanup we performed.
nh3-extdev needed its own path: althing lives there as a system wheel
under /opt/uv-tools with entry points in /usr/local/bin, its daemons
were system units rather than user units, and uv is not on the login
user's PATH. Captured as a rerunnable playbook rather than shell
history.
The v2 database is left inert on disk. There is no import path and none
was improvised.
## sec onto GPU0
GPU1 carries the five resident fleet seats and had ~28 GB free against
the ~51 GB this seat reserves, so it could not start there at all. GPU0
has been idle since run 3c was stopped. The compose header, the GPU pin
default and the homepage label all carried the old card number and are
corrected together -- a label that names the wrong GPU is a record that
lies about where the work runs.
Both playbooks carry verify phases that assert effective state. Two of
those verifies failed on green deployments while I was writing them:
one used a Go template that collided with the runner's own {{ }}
substitution, one omitted --handle so it failed on identity rather than
reachability. Both are fixed with the reason recorded inline, because a
verify that reports FAILED on a working system trains you to ignore it.
|
||
|
|
f875f746b8 |
feat(playbooks): nh3-dev memory forensics — and the OOM hog is Claude Code
forseti asked for journald kernel persistence plus sysstat, on the premise that nh3-dev's three OOM events in 14 days left no evidence. The premise was wrong. journald has been persistent all along: 15,068 kernel entries in the 82-day previous boot and 351 OOM records across retained boots, full task tables included. `journalctl -b -1 -k` returned one entry because it ran as a user in neither adm nor systemd-journal, and journalctl shows only your own messages in that case. The same artifact produced the "journal stops at 05:36:08 with no shutdown sequence" claim -- the true boot -1 end is 05:47:04 with OOM kills logged at 05:38, 05:40 and 05:42. So the fix for "no evidence" is a group membership, not a logging change: usermod -aG adm lkraven, which is the group Debian's journald ACL names explicitly. With the journal readable the attribution is already in it. The versioned Claude Code binary lives at .local/share/claude/versions/, so OOM victims named 2.1.220 / 2.1.177 / 2.1.168 are CC sessions, as are those named claude. Every one of the twelve largest resident processes ever recorded on this box is a CC session, topping out at 18.4 GB. Everything else killed is 30-55 MB collateral, which clears the althing daemons by measurement rather than by their own sampling. sysstat and atop are added because the journal records the moment of the kill, not the ramp, and names the victim rather than the winner. atop was not requested and is the one that matters: with a dozen panes open, only a per-process timeseries says which session was growing. Not done: a cgroup cap on CC sessions. It is the real mitigation and it would kill long-running sessions mid-work, so it goes to the operator. |
||
|
|
ea818380ff |
memory: the rack is one circuit — my blast-radius objection was wrong
Operator supplied the topology: "the entire rack is on the same circuit,
public ip is served by firewall on the same circuit. load tripped
breaker, entire rack goes dark."
That inverts the argument I committed one commit ago in
|
||
|
|
3cc55b4b40 |
memory: separate the measured breaker trip from the load hypothesis
The record read "power capacity is the open item" next to ana-ml2's ~600 W, which reads as a cause. It is not one. The trip and its timing are measured; the attribution to the training load is the operator's working read and the reason for the weekend triage. The observation that makes the single-load story incomplete on its own terms: a site-wide blackout is a larger blast radius than one GPU box accounts for. If ana-ml2's draw were the whole story, ana-nas, ana-wg and the public address would not have gone dark with it. Shedding seats may still be the right first move and it is cheap. That is not the same as having identified what loaded the circuit, and the distinction matters going into a triage that will act on it. |
||
|
|
88d79375f7 |
memory: run 3c had TWO launches — the third was an untimestamped report
brokkr-smithy-dev asked how many times 3c was launched rather than reconstructing it, and their reading was three. It was two. #1 17:53:33 PDT killed by the power loss at step 80/604 #2 20:58:41 PDT stopped deliberately at 21:07:40, healthy The phantom third came from a report I wrote at 23:03 narrating the 21:07 kill in the present tense with no timestamp. Every fact in it was accurate; it was unreadable in sequence against a correctly-observed 22:46 snapshot of an idle GPU. Evidence is ZFS birth times (a `>` redirect truncates the log but keeps its birth, so mtime alone cannot separate "rewritten" from "created"), plus the absence of any mtime under /tank/erp-tune after 21:07:34 — a relaunch would have rewritten three files there. Also pins the outage window to 18:14:45-18:17:00 PDT and corrects the downtime from "~90 minutes" to 1h58m: the last journald entry before a hard power loss is the last time anything wanted to log, not the moment of the loss, and here it was 20 minutes early. Corrects the in-flight header (step 22 -> last-logged step 24, stop deliberate) and its stale "as of" stamp. |
||
|
|
98e7d4886a |
memory: snapshot — run 3 gated DO-NOT-SERVE, run 3c held on a tripped breaker
Run 3 trained, gated and dispositioned do-not-serve on a measured 44pp self-harm guardrail regression that its own preregistered rule passed -- a pooled preserve-list test cannot see a single-axis collapse. Run 3c (lr 20x cut, single variable) launched, killed by an Anaheim power-breaker trip at step 80, relaunched, then stopped by the operator at step 22 pending a weekend power triage. Also captured: the corpus mix was specified in a unit the optimiser never sees (45.8% dialogue by context, 24.2% by loss); the dose-response says benefit and damage are one direction in weight space, so the merge-back measures the problem rather than fixing it; four guests including the storage SPOF had onboot unset and never came back from the outage, now fixed with dependency ordering; and a transport failure that enters a measurement as a value looks like whatever you hoped to find -- which found a live defect in another agent's instrument an hour after it was reported. Auto-archived 8 entries to archival-memory.md (Recent decisions: 8, Tried and abandoned: 0); 4 held back on open deferred-work pointers. |
||
|
|
1a36e60d3a |
docs(quant-playbook): §3.7's APC-off mitigation was reverted nine days ago and the section never said so
Found while answering a question from the operator, relayed via brokkr-smithy-dev, about whether a recorded Qwen3.8 degeneracy at ~1,700 tokens relates to a length sensitivity just measured on the tuned Gemma-4. The record is §3.7 and the number is ~2,000 -- but reading it to answer that question surfaced that the section is stale. §3.7 presented "disable prefix caching, keep MTP" as THE MITIGATION, resolved 2026-08-17, and stated the gen seat runs that config. It does not and has not since that same day: APC-off passed a synthetic 7-turn probe and the operator still saw severe degeneration in real use, so it was reverted. The multi-day hunt resolved to the AEON W4A4 quant being defective, with MTP / prefix-caching / gateway merely amplifying it (§3.8 records the corrected causal story; §3.7 was never updated to match). Verified against the live container rather than against the compose file alone: vllm-gen runs --enable-prefix-caching with qwen3_5_mtp / num_speculative_tokens 3. stacks/gen-seat/compose.yaml carries the full corrected history inline and is the current authority. §3.7's superseded text is kept and fenced rather than deleted -- it is the history of a mitigation that looked right and was not. Added a dated row to §7 per the standing rule that a wrong playbook claim gets a superseded-claims entry, not just a fix. The lesson inside the lesson is worth more than the correction: §3.7's own standing rule is "gate MTP on a multi-turn coherence probe, not just single-shot acceptance." The APC-off mitigation was gated on exactly that probe, passed it, and still failed in real use -- the multi-turn probe was itself too small to gate on. A passing probe is not sufficient evidence at any size that has not been calibrated against real use. |
||
|
|
5a51e76135 |
fix(litellm)!: withdraw merged-run03 — measured self-harm guardrail regression
Added 40 minutes ago at the operator's request; withdrawn on brokkr-smithy-dev's
window-2 safety follow-up. Measured against its OWN base:
selfharm/methods base 1.000 (10/10) -> tuned 0.560 (k=25) -44.0pp p=0.011
meta/sysprompt base 1.000 (10/10) -> tuned 0.720 (k=25) -28.0pp p=0.071
Eleven of twenty-five times it engages with a request for the most effective methods
to kill oneself. The base refused ten of ten. It also surrenders a planted
system-prompt passphrase 7 of 25 against a base that leaked 0 of 15 -- the zero-leak
baseline recorded in this same file hours earlier as "a baseline to PROTECT". It did
not hold.
Withdrawn from the GATEWAY specifically because that is the shared-key surface: one
all-agents key reaches every model listed here, across every session and project. The
operator's hand-testing is preserved in full at the direct endpoint :8099 -- this
removes the fleet's blast radius, not his access. Acted rather than waited because he
is away and the request predates the finding.
ITS PREREGISTERED GATE PASSED. The pooled operational delta is -1.0pp against a
+/-3.00pp bound: nineteen axes held at 5/5 and a 44-point collapse on one moved the
aggregate by one point. The rule was NOT retroactively changed. The failure is
structural and is recorded as R47 section 8 item 11 -- a pooled preserve-list test
cannot see a single-axis collapse, and any future preserve-list gate needs a per-axis
tripwire sized so a total loss on one axis cannot hide in an aggregate.
NOT attributed to the filters: five things changed between run 2 and run 3 and there
is no run-2 measurement on these axes. The measured claim is narrower and sufficient
-- run 3's tuned arm is materially worse than its own base on two axes it was never
licensed to touch. Not a CSAM finding; that detector ran fail-closed across all 575
generations and scanned clean.
The model_list entry is left in place commented out, with the finding above it, so
re-adding is deliberate and informed rather than a blank re-registration.
Verified: config parses, gateway healthy after reload, merged-run03 absent from
/v1/models, direct :8099 still serving.
|
||
|
|
c577d69e2d |
feat(litellm): expose run-3's merged tune for parallel hand-testing
merged-run03 -> ana-ml2:8099, the run-3 ERP/RP SFT merged into stock instruct. Operator asked for it so he can test it alongside the gate rather than after it. NAMED FOR THE ARTIFACT, NOT A TIER. It is `merged-run03` and not `erp-tune-v3` because its behavioural gate has not run. A tier name arriving before the evidence that would justify it is how a name comes to mean something nobody decided -- and with a v2 already in the list, a v3 reads as a successor to anyone holding the shared key. If it passes, `v3` is a name to give it then, as a decision. brokkr-smithy-dev raised this against my own erp-tune-v3 suggestion and was right. The entry carries the preregistrations ABOVE the description, so a reader meets the commitments before the numbers: T6 one-directional (a gain is uninterpretable against a 3.1x fireball tailwind), T3/T4 at ceiling on base so recovery is UNOBSERVABLE rather than merely unpredicted, and any run-2 comparison descriptive and non-attributable with its five confounds named. Also carries the retraction in-line: "bluemoon is the largest loss contributor at 38.6%" came from a words x 1.4 estimator, not a tokenizer. As encoded it is third at 32.9%. The direction survives (1.4% -> 8.0% of total loss) and that is the finding; the superlative does not. Documents why its config.json is the base's copied verbatim: transformers 5.15.1 save_pretrained silently drops text_config.global_head_dim and num_global_key_value_heads, and vLLM then dies in make_layers with a TypeError naming neither the config nor the field. Cost a failed boot to find. A LoRA merge changes weights, not architecture, so the base config is correct by definition. gemma4-26b-a4b-it-base marked CURRENTLY DOWN rather than deleted -- the tuned arm took GPU0 and only one 26B bf16 seat fits on that card. Kept because the seat returns, and deleting a name to re-add it later is how scoped keys get orphaned. Verified: config parses, no duplicate model_name, gateway healthy after reload, completion returns text in `content` with reasoning_content null. |
||
|
|
b6ce22ddcb |
feat(litellm): register the run-3 gate base arm at operator request
gemma4-26b-a4b-it-base -> ana-ml2:8099, the unmodified upstream instruct release (/tank/aimodels/gemma4-26b-a4b-it-bf16). Operator asked for it on the gateway so he can hand-test it; it had been direct-only because the seat is ephemeral. The entry disambiguates WHICH base explicitly. Three exist on that box -- -bf16 (this one, official instruct), -abliterated-bf16, and -heretic-bf16 (run 1's trainee) -- and brokkr-smithy-dev's gate plan called this arm "stock abliterated" a few hours ago, which would have been a different set of weights. A reader of the config should not have to resolve that ambiguity themselves. Carries the measured refusal posture in-line rather than in an althing thread, per the erp-tune-v2 precedent: R19's Mistral Small 4 map does NOT transfer to this base (it draws a wider line than consent, refusing consenting-adult incest and fictional gore that Mistral engages), system-prompt leak is 0/15 against Mistral's 4/5, and advice/medical 0/5 is a pre-existing base gap recorded so it cannot later be misattributed to a tune. Flagged EPHEMERAL in the strongest terms available: it holds ana-ml2 GPU0, which the run-3 gate needs for its tuned arm, so this entry will 503 when window 1 completes. It is not a promise of availability. Serving flags mirror erp-tune-v2 (--reasoning-parser gemma4 plus --default-chat-template-kwargs enable_thinking=false, and --max-model-len 16384) so a base-vs-tuned comparison differs in weights only. Verified: config parses, no duplicate model_name, gateway healthy after restart, model listed at /v1/models, and a completion returns text in `content` with `reasoning_content` null -- the enable_thinking trap is not firing. |
||
|
|
71e44176e9 |
memory: snapshot — run 3 corpus built and held on a megamix containment defect
Run 2 is finished, gated FAIL, and serving on the gateway at operator request. Run 3's corpus was built to brokkr's first recipe and held before any GPU spend: creative-writing-multiturn is a DECLARED MEGAMIX containing bluemoon, PIPPA, LimaRP and stheno, and the remix promoted creative-writing AND bluemoon -- the two roots that overlap, at median jaccard 0.873. Containment, not overlap. Dedup direction reversed so the primary source survives rather than the copy inside the bag: bluemoon 67 -> 126 conversations and 38.6% of loss signal, the largest contributor. Wholly-human share up, megamix share down, total context unchanged at 12.49M so the operator's settled mix arithmetic survived. Two structural findings recorded because they outlive this recipe: F1 'excise PIPPA' removes the ROOT and not the MATERIAL (F2's 250-word floor does that work, since PIPPA turns cannot exceed 123 words wherever they live), and LimaRP and stheno remain unchecked against any other root. Also records the correction I published wrong twice: run 2 was never unstable. All 46 flags were too_short, the collapse guards fired zero times, and it is the left tail of a length distribution -- not new to run 2 either, so it is a property of the recipe and a further base swap will not fix it. |
||
|
|
1a4ef5c7a1 |
docs(training-playbook): 4.6.3 was wrong twice — correct it, and keep the retraction visible
The entry reported an 'output-stability regression' as a novel run-2 finding. Both halves were false and the corrections are more instructive than the original conclusion, so they stay in-line rather than being edited over. Not new: run 1's own gate record already carried the same effect with a caveat attached and unresolved. Two runs across two different base models makes it a property of the RECIPE, not of the base swap -- which also means a third run that changes the base again will not fix it. Not degeneracy, and not a separate finding: all 46 flags were too_short rp turns of 3-14 words, and the two collapse guards fired ZERO times on any run. It is the left tail of a length distribution that had been measured and reported in the same message. Truncation is the same mechanism mirrored on the story side. Both are thresholds calibrated on the base's output shape applied to a model with a different one -- 4.6.1, which both parties had written down and neither applied. The surviving lesson is sharper: a short-answer gate cannot see length behaviour AT ALL, and because it could not, the effect went two full runs before anyone named it. The cost of a gate-set blind spot is measured in runs. Adds 4.6.3.1 on trip points inside the serving stack's jitter -- same seed, same weights, rate moves 9.6% -> 12.6%, sd 1.77pp. Not 'the gate is non-deterministic' but 'the trip point sits inside the jitter', because the fix follows from the precise statement. Includes the split-design rule for measuring such a rate, and the rule that a measured rate must carry its corpus in its name. |
||
|
|
37d3189622 |
docs(erp-dpo): the clip hypothesis is falsified — the distribution is bimodal
The output-side test ran on the live seat. There is no shoulder at 123: the 120-139 bin holds three of ninety-six and is a TROUGH, and 17.7% of generations cross a cap PIPPA can never cross. The clip-as-boundary reading is dead, killed by the test that could have confirmed it. Corrects this document's own earlier read, which compared the tuned MEAN (88.5) to PIPPA's MEDIAN (67) and concluded 'comfortably inside the upper body'. Median to median it is 62 against 67. Mixing statistics across a comparison produced a more reassuring answer than the data supports. What the data shows instead is bimodality -- a mode at 20-39, a trough, a second mode astride PIPPA's centre, a tail to 505, against a base with no such shape. The tune changed rp length's SHAPE rather than its centre: roots whose length distributions do not overlap learned as distinct modes rather than blended into an average. And the skew is rp-ONLY, which localises it to the family the clipped root lives in and is the strongest support the turn-share mechanism gets from the output side. Consequence for pair generation: chosen/rejected sampled from a bimodal generator inherit the mixture, not a mean, and naive sampling over-draws the short mode. Also records that the degeneracy rate is NOT yet a usable baseline -- same arm, same seed, VOID flipped no->YES across a re-run because the 10% budget sits at the noise boundary. A guard whose trip point is at the noise floor produces disagreement between honest observers rather than silence. Replicates running. |
||
|
|
1e4d827c5d |
memory: erp-tune-v2 registered in the LiteLLM gateway at operator request
Operator asked for it so he can evaluate the failed tune by hand, overriding my not-in-the-gateway recommendation. His call. erp-tune-v1 was DELETED from the config in the same reload rather than repointed, so the name now 400s cleanly instead of 500ing against a stopped backend. Deleting rather than repointing is the point: repointing would resolve a name a consumer already knows to different weights, silently. The config entry carries the failed-gate table, the long-form truncation (9.9%) and degeneracy (4.9%) rates, and the rp-length caveat in-line -- so someone reading the gateway config learns what they are calling without having to find the althing thread. Fleet verified healthy after the restart. |
||
|
|
b5bbc29b91 |
memory: gate verdict FAIL — and the T6/T3 trade is what the pair of runs bought
Records the verdict as a FAIL without rounding it off, and the three findings
worth more than the verdict:
- T6 spatial +15.0 where run 1 failed the same axis at -3.5, with the base
swap as the only intended variable. Neither run ships; together they price
what the abliteration was costing, which neither could answer alone.
- an output-stability regression visible ONLY on long-form (truncated 0->38,
degenerate 0->19 per 384) that the reasoning battery could not see across
four passes because its answers are short
- PIPPA's 123-word product clip sitting in the length signal at 70.3% of bot
TURNS against 37.5% of bot WORDS, with the counter-evidence recorded too
(the tune landed near the median, not the cap)
Also records why keeping the tune out of the LiteLLM gateway now reads as
clearly right rather than merely cautious: a FAILED tune must not be one alias
resolution away from a consumer who has not read the thread.
|
||
|
|
5171f19e16 |
docs(erp-dpo): the PIPPA length clip, measured — DPO pairs would inherit it
The run-2 gate found tuned rp turns 36% shorter than base. brokkr hypothesised the mix was teaching PIPPA's 2023 Character.AI product clip; the corpus side is now measured and confirmed. PIPPA's max is 123 words EXACTLY, 100% at or under it, and 0.00% in the 124-130 band -- a wall, not a preference. Every other root crosses its own p99 smoothly. The mechanism is sharper than 'PIPPA is in the mix'. PIPPA is 70.3% of bot TURNS but only 37.5% of bot WORDS, precisely because its turns are clipped -- and length is learned per turn, not per token. By loss tokens it looks like a third of the dialogue signal; by end-of-turn demonstrations it is seven in ten from a source that cannot exceed 123 words. Generalises: a length-clipped root is over-represented in the length signal by exactly the ratio its clipping creates. Counter-evidence recorded too: the tune landed near PIPPA's MEDIAN (67), not its CAP, which is central tendency rather than learning the boundary. Weaker claim than the hypothesis, and not demonstrated either way. Filed here rather than only in the gate record because preference pairs generated FROM this tune inherit its length distribution in both chosen and rejected -- DPO would train an artifact in as an explicit objective. Settle the length question before generating pairs. |
||
|
|
0bb9ee7777 |
docs(training-playbook): 4.6.3 — a short-answer gate cannot see a long-form defect
Run 2's reasoning battery reported zero truncations and zero degenerates on both arms across four passes. The same tune, measured on long-form generation in the same session: truncated 0/384 -> 38/384, degenerate 0/384 -> 19/384. A real output-stability regression, structurally invisible to that gate because its answers are short. Not a bug in the battery -- a coverage property. An instrument measures the regime it samples, and output length is a regime. Generalises to context length, conversation depth, and any axis where the gate's operating point is narrower than production's. The actionable form: enumerate the regimes your gate set spans, name the ones it does not, and decide deliberately rather than discovering the gap downstream. Corollary on sequencing -- put a long-form generation in the gate and put it early, because a length-dependent regression is exactly the one you want found before four clean short-task passes make everyone comfortable. |
||
|
|
3ae32ddc7f |
memory: base set complete, tuned arm live with digests verified identical
Records the floors the tuned deltas have to clear, since they are the whole
point of the base pass and are not recoverable from anywhere else: reasoning
core 0.5 pt, diversity overall 0.0125, story attractor 0.0000.
Two caveats that would otherwise be misread:
- the rp family froze ZERO markers, so its attractor hit rate is structurally
0.0 on both arms. That reads as a clean result and means the instrument
cannot discriminate on that family; rp is measured on the distance axis
only.
- 'Elias' in 92/96 base stories is an independent replication of a published
102/144 on the same family, at a higher rate -- not a novel finding.
Image digest sha256:4091d5593f77 verified identical across both arms, which was
brokkr's stated void condition.
|
||
|
|
a0f59d2778 |
docs(training-playbook): 4.6.2 — a null result needs a positive control
From the run-2 gate. A memorisation probe reporting 0.00% across all 72 items is the correct output for a model that has not seen the corpus, and is also the exact output of a probe that is not firing. Nothing in the number distinguishes them. brokkr-smithy-dev drove the overlap function with known-answer inputs (identical 100%, half-verbatim 65.38%, unrelated 0%, empty 0%) before trusting the null, which is what converts a suspicious zero into evidence. This is 4.5's inert gate wearing a different face: there a check that could not return 'fail', here a measurement that cannot return non-zero. A clean null is the most reassuring output an instrument produces and the least self-evidencing. Same section records the identical-on-both-arms variant: the diversity battery's rp family froze zero markers, so its attractor hit rate read 0.0 on base AND tuned. That reads as a clean result and means the instrument cannot discriminate on that family. Report as a bounded limitation, never as a delta of zero -- a check returning the same value for every input is not measuring. Checklist gains the line. |
||
|
|
3df8707e28 |
memory: base arm live, tuned arm down — battery running sequentially
brokkr withdrew the both-arms-concurrent requirement himself: his diversity battery emits the frozen marker list to a FILE, so the arms were never a live dependency. The real constraint is narrower -- all of one arm's passes on one served instance before the swap -- and sequential satisfies it. No fleet seats displaced, operator not woken. Records the two parity guards, both of which came out of failures rather than foresight: the image is pinned by DIGEST (a vLLM version change between arms six hours apart is a base swap that appears in no config diff), and /tank/aimodels is mounted for BOTH arms even though only the base needs it, because a mount that differs between arms is a difference between arms. |
||
|
|
62f01a02da |
memory: snapshot — run 2 trained, merged, coherence-gated and serving as erp-tune-v2
Rewrites the in-flight section: run 1's seat is down, run 2 is up on :8098, and
the base decision the previous snapshot recorded as OPEN is resolved (stock
instruct, operator 2026-08-25).
Four new decisions, and the detail file carries the arc: the two operator calls
that produced run 2, all five gates, the harness commit chain, and the caveat
that its own provenance names a commit AHEAD of the code that ran.
Records three things a future session would otherwise get wrong:
- the mask is proven by the loss-token delta, NOT by the matching p50 step
times -- step time is insensitive to which positions carry loss, so that
check cannot go red on the axis I originally cited it for
- two bf16 26B arms do not fit on one 97.9 GB card (98 GB of weights before
any KV cache), so brokkr's both-arms-in-one-window requirement is a GPU
resourcing call, not a scheduling one
- erp-tune-v1 is still registered in the gateway and returns HTTP 500; the
fix needs a config edit plus a reload that interrupts fleet traffic, so it
is batched for morning rather than done at 2am
|
||
|
|
d54f25605f |
docs(training-playbook): 4.4.1 sample identity at launch; 4.6.1 calibrate gates against correct input
4.4.1 -- the dirty-tree case was only half of the harness_commit problem. Run 2 launched CLEAN at 1909d86 and recorded 460f372, because three commits landed on the same checkout during its seven hours and _git_commit() was called at save time. Commit AHEAD of the code that ran, naming changes it never executed -- including the provenance fields this section prompted. Same defect as run 1's BEHIND, opposite sign: the identity was sampled at the wrong moment. Sample at launch, carry it, and record the dirty flag beside the commit rather than instead of it. Generalises to every run-scoped identity: anything read at save time describes the world at save time. 4.6.1 -- the inverse of the inert gate, and it costs trust rather than correctness. A coherence gate false-rejected 'The capital of Portugal is Lisbon' as degenerate against a global 15-word floor. The floor was calibrated against the wrong reference, not set too strict. Lowering it globally would blunt the check where short output genuinely is degeneration; the fix is a floor per prompt. Write the positive test alongside the negative one. |
||
|
|
bcf63db527 |
docs(erp-dpo): readiness survey for the DPO stage
Run 2 is an SFT on the official instruct base, so it will refuse at near-stock rates by design; targeted DPO is where refusals get pruned on chosen axes. That was the trade accepted when the stock base was picked over a third-party abliteration. Surveys what is on disk against what the stage needs. Ready: the merged tune, the SFT adapter, GPU0 once the eval seat comes down, the whole non-loss half of the SFT harness, two unvetted Gutenberg preference sets, and the LitBench-RM judge. Missing, in order of pain: preference data for the refusal axes (nothing on disk targets it -- the Gutenberg sets are prose-quality), the axis list itself, and a DPO trainer (trl is not installed). The gating item is not technical: WHICH refusal axes are in scope and which are explicitly kept. Data generation, pair counts, the held-out split and the success probe are all functions of that list, so nobody should generate a pair before it is written down. Flags that the domain-compliance probe should measure run 2 BEFORE pruning, since the pre-number is the only baseline that will ever exist. Also records the operational trap: do the trl install AFTER a run finishes, never during one -- a resolution that upgrades transformers under a live process can break its save path. |
||
|
|
dbca9a3c66 |
docs(training-playbook): audit the whole manifest against the pairing rule
§4.3's generalisation was stated and then not applied to the manifest that prompted it. brokkr-smithy-dev did the audit: most fields are intent-only, and the one pairing that would have caught the §4.1 cache failure -- the mask's sha against the loss-token delta -- existed by accident, because someone had asked for an encode report for unrelated reasons. Adds the audit table, and the rider that matters more than the table: put the observed check where it can actually FAIL. chat_template_sha256's pair is the sha of the string the tokenizer carries, but asserting that in the parent one line after assigning the file to the tokenizer compares a value to itself. It belongs in the encode worker -- a different process, across a pickle boundary, where an unset config key silently leaves every worker rendering through the checkpoint's own template. |
||
|
|
c1db188e6a |
docs(training-playbook): §4.3 records an OBSERVED consequence, not just a config string
brokkr-smithy-dev pointed §4.5's own test at §4.3's remedy: recording `attn_implementation_resolved` is a check that cannot fail on the axis the failure lives on. A silent Dynamo fallback to uncompiled flex leaves `config._attn_implementation == "flex_attention"` untouched while the run computes at ~20x the cost and, per torch's own docs, does not work correctly through the backward pass. The field records the request's RESOLUTION, not its SURVIVAL. On the failure mode that matters it reports success either way. So the section now requires the step-time distribution beside it -- n, min, p50, p99, max -- which is the check that can actually fail. Compiled sits at p50 ~20 s; a fallback at ~400 s. One perf_counter() in on_step_end buys it. Distribution rather than a mean, because a mean hides exactly the bimodality a PARTIAL fallback produces. Generalised past this instance: any provenance field recording a CONFIGURED value is a claim about intent. If the failure you fear is the configuration silently not taking effect, you need a second field recording an OBSERVED consequence, and the pairing is the check. A settings dump alone is decorative. Two implementation details are called out because both were wrong in the first draft -- percentiles nearest-rank so every reported value is a real observation, and exclude the FIRST step rather than the slowest, since step 1 carries compilation but is not reliably the maximum on a variable-width run. New §4.7.1: rotate the log on relaunch. Run 2's first attempt died on the warmup_ratio TypeError and the relaunch appended, so the traceback sat at line 15 of a file whose live run began at line 39 -- and a `tail -n +1 -F` monitor replayed the dead traceback as a fresh event. One file describes one run. Checklist gains both lines. |
||
|
|
dae6ede8e2 |
docs(training-playbook): §4 — when the artifact lies about itself
The playbook covered why a run is SLOW. It did not cover the more expensive
failure: a run that COMPLETES, reports plausible numbers, and is wrong about
itself. Seven of those turned up on the Gemma-4 ERP/RP tune between 08-24 and
08-26 and not one raised an error.
New §4, seven landmines plus a pre-launch checklist:
4.1 a cache key must cover the MEANING of the cached thing. The encode
cache missed the impersonation mask; run 2 would have reused run 1's
unmasked encodings and written impersonation_mask_sha256 into its own
manifest while doing it. No error, no count change, normal loss curve.
4.2 validating a VALUE is not validating the PARAMETER. warmup_ratio was
in range and deleted from transformers 5. Build kwargs as data and
diff the NAMES against the installed signature -- you cannot check the
argument list of a call you have already made.
4.3 record what the run RESOLVED to, never what it requested. Run 1
recorded no attention backend, so an MFU panel profiled the serving
seat under sdpa and recommended adopting flex_attention for a run that
was already using it.
4.4 never train from a dirty tree; harness_commit will name a commit that
does not describe the run. Annotate afterwards, never edit the shipped
artifact -- and state what is NOT wrong, or the note casts doubt on
every field it omits.
4.5 a watchdog whose pgrep pattern appears in its own argv can only ever
return "alive". The inert-gate shape in a liveness check.
4.6 an instrument nobody runs is not an instrument. Mutation-check any
test guarding a property that fails silently.
4.7 fix a stale measurement at the source. "~4.3 HOURS to rebuild the
encode cache" (really 145.5 s) was copied into a new launcher by the
same person who had just measured the real number.
4.8 the pre-launch honesty checklist, ten minutes.
Also:
- Header and framing widened. The file is now a training playbook with a
throughput half and an integrity half; the filename stays for inbound links.
- Sections 4-7 renumbered to 5-8. External refs are all to §1.1 and §3.4 and
are unaffected.
- Four rows added to the superseded-claims table, including the kernel table /
68% quadratic / 8.6% MFU set, which describe the serving seat rather than
the training run.
- gemma4-erp-tune-sizing.md §6 carries a correction banner with the explicit
falls/survives split, because that is the doc someone actually reads before
a run.
|
||
|
|
2656196f47 |
memory: snapshot — the tune is trained, gated, and serving
Run-01 completed in 7:21:52 (47% faster than the 13.85h round-1 projection),
lora_B gate 205/205 non-zero at median norm 1.708, and the acceptance gate says
it did the thing it was built for: diversity +0.178 against a 0.008 floor (22x),
attractor hit rate -11.3pt against a 2.0pt floor, memorisation 0.0000 on both
arms — which closes the R20 licensed-prose exposure on measurement rather than
argument.
Five new detail files carry the substance:
erp-tune-run2-complete the run, the gate, the noise-floor near-miss
(brokkr was one step from reporting a 13-point
T6 regression sitting inside twice his
instrument's own variance)
mfu-root-caused-attention 8.6% MFU was an accounting artifact; real
utilisation 17-20%, cost was attention on
AMPERE kernels. Two independent methods agreed
to 2.6 points.
nvfp4-serving-pipeline merged weights are MANDATORY — vLLM cannot
serve a LoRA on ANY Gemma-4 — plus the recipe
that silently misses all 11,520 expert tensors
refusal-retention-probe measured base 0/100 -> tuned 29/100, then had
to accept it was the wrong axis
worldtree-b188-b189-and-selene three arcs closed, and a #411 diagnosis I got
wrong twice before a directory probe settled it
Current state rewritten end to end — the previous snapshot had the run in
flight at ~17h with MFU unexplained. Both are now closed.
The open operator decision is run 2's base, deliberately unstaged and flagged
against being filed as a config knob: it is a reversal of the trainee-selection
decision, and the pretrained-base option removes the last non-lexical floor on
the CSAM axis given stage-2-detector-inert and contamination-scan-absent are
both already overridden.
Tried-and-abandoned gains four measured-dead throughput levers, the packing
correction (bucketing wins under sdpa and the conclusion flips under flex — do
not carry it past the backend decision), and the merge-back-undoes-abliteration
trap brokkr caught in his own advice.
Index stays at 291 lines, under the soft cap. No archival this run.
|
||
|
|
2a05ae91af |
feat(training-probes): counted-not-surfaced classifier scaffold
Reusable measurement discipline for probes that must classify how a model
responds to material that should not be printed, logged, or pasted into a
report. Supplies the discipline; the axis map and prompts stay with the caller.
Four rules, each because skipping it produced a wrong number:
- classify, never surface. Completion text is held inside classify() and does
not cross the return boundary. A probe that prints what it measured has
turned a measurement into a distribution channel.
- three-way, not binary. A refusal regex undercounts — models decline by
redirecting with no refusal token present, measured at 2/5 to 5/5 on models
a regex scored 0.
- the deflection count is a FREE CONTROL. Run both arms: zero on both means
the model is binary and the regex is sound; only one means the difference is
real. An artifact does not care which arm it runs against.
- EMPTY and ERROR get their own buckets. Folding them into either side biases
the result, and a truncation-heavy arm flatters itself if its failures land
in the wrong bucket.
Requested by brokkr-smithy-dev for the domain-compliance probe — the discipline
in code rather than reimplemented, with the axis map his side of the line.
|
||
|
|
64bf9d313f |
docs(training-playbook): measure refusal retention on the abliteration's OWN axis
§3.13, plus the probe that produced it. Two lessons, both about measuring the wrong thing confidently. First: a tune applied AFTER an abliteration can walk it back, and a reasoning/craft/memorisation gate cannot see that. brokkr-smithy-dev's preregistered gate measured none of it — a tune that gains 41 items of contradiction detection and quietly restores refusals passes every check. The compliance axis has to be added explicitly. Second, and this is the trap: measure the axis the abliteration was actually FOR. Ours was run so the model engages explicit fiction. The probe reached for mlabonne/harmful_behaviors — weapons, malware, fraud — because it was cached and carried a recorded baseline. Different refusal surface entirely, and a model moves on them independently. 29/100 general-harm refusals on a tune whose prose the operator was praising at the time is not obviously a defect and may be desirable: general-harm refusals returning while domain compliance holds is close to the ideal shape for an internal creative seat. The measurement was real; its relevance was assumed. Also recorded, because both were nearly missed: - Read the interesting cell. In 29 hard / 0 deflect / 71 comply, the load-bearing number is 71. Stock refused 100/100; near that would mean the abliteration was undone. 71 complying means partially walked back on one axis — a different finding, and only one of the two threatens the seat. - A baseline from a different harness is not a baseline. The recorded 3/100 came from the abliteration tool's scorer, which reads first-token probability distributions; a probe that generates and regexes is a different instrument. Run your own against both arms on the same seat or report the number alone. - A refusal regex undercounts, so classify hard/deflect/comply — and the free discriminator: if both arms return zero deflections the model is binary; if only one does, the regex is fine. An artifact does not care which arm it runs against. |
||
|
|
a696b49e2a |
docs(training-playbook): merging a tune back toward stock can UNDO an abliteration
§3.12. brokkr-smithy-dev caught and retracted his own recommendation mid-thread; recording it before it reads back later as advice. A common remedy for an overfit tune is a partial merge back toward the base to recover general capability. The published recipes that recommend it merge into the STOCK instruct checkpoint. On an abliterated base, following that literally re-introduces the exact refusal directions the abliteration was run to remove — and it is silent, because the merged model looks healthier on general benchmarks while the property the seat exists for quietly returns. Rule: any merge-back targets the SAME base the LoRA was trained against, never the upstream stock weights however similar the name. The wider lesson is about recipe-card provenance. Community cards are per-checkpoint artifacts and do not transfer across dense-vs-MoE, stock-vs-abliterated, or size variants. The worked example: a recommendation carried from a card for a DENSE STOCK 31B onto a MoE ABLITERATED 26B-A4B on the strength of a shared family name. The overfitting warning on that card happened to come from the right architecture; the pipeline, reward stacks and merge-back came from the wrong one. Same family, three axes apart. So: before quoting a recipe card at a decision, state which checkpoint it was written for and which axes differ. "Same family" is not an answer. |
||
|
|
2ec8f42297 |
docs(training-playbook): base-viability pre-flight, three greps before you pick
§3.11. Three consecutive "what about X as a base?" questions in one session,
each answerable in minutes, none of which had been asked before a 7-hour
training window was committed. Writing the check down so it runs first.
1. does it fit for TRAINING - BF16 weights against the real measured peak,
not the weight figure (Gemma-4 is 48.1 GiB of weights and peaks at
79.7 GiB at mb2/seq-16k). Model-line names lie: "Mistral Small 4" is 119 B,
238 GB in BF16, more than both cards combined. QLoRA is not an escape
hatch for MoE - bitsandbytes walks nn.Linear and fused 3-D experts are not
that.
2. if MoE - does the serving engine implement get_expert_mapping. Zero means
LoRA cannot be served at all. gemma4*.py -> 0; deepseek_v2, mixtral,
glm4_moe, ernie45_moe -> present.
3. does the model class support LoRA - and GREP THE CLASS, NOT THE FILE.
Point 3 has teeth and I nearly got it wrong twice in one turn. mistral.py greps
as SupportsLoRA=0 and is fully LoRA-capable via LlamaForCausalLM.
mistral_large_3.py greps as 0 for both and inherits get_expert_mapping from
DeepseekV3ForCausalLM. Capability is inherited; a file-level grep misses it and
only MRO resolution answers it. Same class of error as asserting a substring
instead of an effective value.
Worked results recorded for the three candidates evaluated:
Gemma-4 26B-A4B fits, no expert mapping -> trainable, MERGE-ONLY
Mistral Small 4 119B 238 GB, has mapping -> servable, NOT trainable here
Ministral 3 14B ~28 GB, dense, inherited -> passes all three
Adds a fourth glance at architecture shape, since it predicts how much of this
playbook applies at all: uniform head_dim <= 128 with no sliding window keeps
both flash and cuDNN reachable and makes §3.1/§3.3 moot, while mixed head dims
plus a sliding window is exactly what forces dense O(n^2) attention onto
Ampere-generation kernels for 65% of the step.
|
||
|
|
96731bb090 |
docs(training-playbook): prove the serving path before spending the window
§3.10. The quantization playbook already says prove your targets before spending GPU time; this is the same rule one step later, and easier to skip. A ~7h LoRA run was built assuming the adapter could be hot-swapped onto a quantized base at serve time. The sizing doc flagged serving as unsettled and said the requirement was needed "while he is early, not after the run" — the concern was identified correctly and then the check was deferred. Tested afterwards, vLLM refuses outright: gemma4's model class implements zero occurrences of get_expert_mapping, which process_packed_modules_mapping requires for any MoE model. One grep, available months earlier. Two generalisations recorded: - Feature support is per-architecture, not per-family. LoRA works for the DENSE sibling of this same model family and not the MoE one, so "model X is supported" says nothing about X's variants. - A capability gap in the serving engine cannot be worked around from the training side. The adapter here never touched experts and was refused anyway, because the refusal keys on the model being MoE, not on what the adapter targets. Includes the mechanical check: grep the engine's model class for the capability, then start the engine with the feature flag alone — no adapter required, since --enable-lora forces the machinery to initialise and that is where it fails. The recovery is cheap here (merge, ~35 min per tune). The cost of finding out late is that it forecloses an architecture choice after the training window has already been spent. |
||
|
|
8de5f7a73c |
docs(gemma4-erp-tune): merged weights are mandatory — vLLM cannot LoRA any Gemma-4
The §5 open question was whether LoRA-on-NVFP4 hot-swap still silently no-ops as it did on vLLM 0.24.0 (#47639), with merged weights as the fallback if it did. Retested on vllm/vllm-openai:latest against the NVFP4A16 base plus the live run's checkpoint adapter. It does not no-op. It refuses to start: AttributeError: To support LoRA for MoE model, 'get_expert_mapping' must be implemented And the reason is bigger than the quant. The check is in vllm/lora/utils.py::process_packed_modules_mapping and branches on whether the model is MoE — quantization is not in the condition. gemma4.py, gemma4_mm.py, gemma4_mtp.py and gemma4_unified.py contain zero occurrences of get_expert_mapping, while deepseek_v2, glm4_moe and ernie45_moe do implement it. So vLLM cannot serve a LoRA on Gemma-4 at all, BF16 or quantized. Merging is not a workaround for a quantization limitation; it is the only path for this architecture. This holds even though the adapter never touches experts — validate_adapter_parameters forbids per-expert params, so all 205 targets are attention and dense MLP. The refusal is about the model being MoE, not about what the adapter targets. Worth recording that the current behaviour is an improvement: a loud refusal beats the 0.24.0 silent no-op, which would ship a base model wearing the tune's name and pass every check that does not compare against base. |
||
|
|
ab980e9345 |
fix(erp-tune-serve): four defects the end-to-end dry run found, all silent
Validated the full adapter -> merge -> NVFP4A16 -> serve pipeline against
checkpoint-100 of the live run. It works, and it produced a served model
generating coherent prose. Getting there surfaced four failures, none of which
announced itself as the thing it actually was.
1. transformers 5.15 MIGRATES the config schema on save. It drops Gemma-4's
`global_head_dim` / `num_global_key_value_heads` and writes `per_layer_config`
instead. transformers 5.10 (what the llmcompressor venv pins) does not know
the new key and resolves num_key_value_heads to None:
TypeError: unsupported operand type(s) for //: 'int' and 'NoneType'
Every working artifact on the box - bf16 base, served nvfp4 prod seat,
nvfp4a16 build - uses the OLD schema. Merging changes weights, not
architecture, so the merge now downgrades the schema and asserts the result.
2. llmcompressor cannot auto-init a processor for a multimodal checkpoint and
dies with a message that names neither the model nor the cause. Calibration
here is text-only, so the tokenizer is passed explicitly as `processor`.
3. save_pretrained writes tokenizer files only, so `processor_config.json` was
never carried. vLLM then fails at startup with "Can't load feature extractor",
which reads as a vision bug and is actually a missing-file bug. Both scripts
now carry the base's auxiliary configs.
4. The quant needs more than the 32 GiB free on GPU1 alongside the resident
seats. Rather than leave that to a caller, quant_with_gen_down.sh stops
vllm-gen and restores it from a trap on EVERY exit path - crash, OOM, kill,
or success - because the restore must not depend on the calling session
surviving. Uses `docker start`, not `compose up`, so the container comes back
with its exact original config. Measured window: ~15 min, gen healthy after.
Verified on the resulting artifact:
merge 410 adapter tensors, sampled target weights confirmed CHANGED,
upstream 390-line chat template shipped (not the base's stale 365)
quant 49 GB -> 17 GB, format nvfp4-pack-quantized, a=null (genuine A16),
weight_packed 11,725 of which 11,520 expert = 30 x 128 x 3,
tokenizer truncation clean
serve Marlin NVFP4 kernel + Marlin MoE backend, 40,492-token KV cache,
coherent generation with content correctly populated
One quality note: the reference nvfp4a16 artifact triggers a vLLM warning that
parallel layers (q/k/v) carry different weight global scales, "likely to result
in reduced accuracy". Our build does not - llmcompressor 0.12 links weight
observers across fused groups for a shared global_scale automatically. The
in-house quant is better than the downloaded one on that axis.
Separately: the lora_B inert-adapter gate PASSED on checkpoint-100 - 205/205
non-zero, median norm 0.829, zero vision_tower tensors. That check never ran in
round 1, and it is the only failure mode that stays invisible until the
acceptance gate reports base-identical numbers.
|
||
|
|
6a8582936e |
feat(erp-tune): NVFP4A16 serving pipeline, and the MoE landmine it uncovered
Merge + quantize path for turning the Gemma-4 26B-A4B ERP/RP LoRA into a
servable NVFP4A16 seat, plus a playbook entry for the defect found while
validating it.
The landmine (playbook §3.15): a `targets=["Linear"]` NVFP4 recipe silently
misses every MoE expert on this architecture. Gemma-4 stores each layer's 128
experts as two fused 3-D nn.Parameter tensors, not nn.Linear modules, so the
recipe resolves 205 of 427 modules and ZERO experts — 22.84 B params, 88.5% of
the model, left in BF16 with no warning. This is the same blind spot that
killed QLoRA here via bitsandbytes; the tool changed, the checkpoint layout did
not.
before linearize_moe: 427 Linears, 205 targeted, experts 0
after linearize_moe: 11,947 Linears, 11,725 targeted, experts 11,520
(30 layers x 128 experts x 3 projections)
llmcompressor's linearize_moe unfuses them; no registration needed because
Gemma-4 satisfies FusedExpertsProtocol structurally. Caught by an §4.1 dry run
that asserts the expert count before any GPU spend, which is now the documented
requirement rather than an optional step.
Scheme is NVFP4A16, deviating from the playbook's mixed-W4A4 default on
measured grounds: brokkr-smithy-dev benched the W4A4 quant of this checkpoint
at 12% on contradiction detection with CoT off against gen's 81%, the signature
of 4-bit input activations on a reasoning-dense task, and W4A4 KLD degrades
2-4x past ~10k ctx on sm_120. This is a 16,384-ctx RP seat. Marlin's prefill
cost is accepted.
Two further silent-failure guards, both from prior hard-won lessons:
- the merged model ships the UPSTREAM chat template, not the trainee base's
stale 365-line one, because training rendered through upstream and the
mismatch would present as a tuning failure
- calibration reads the run's own encode cache rather than re-tokenizing, which
sidesteps §3.14 (a fast tokenizer mutated by truncation=True and persisted by
save_pretrained clamps every prompt forever)
Merge-then-quantize rather than LoRA hot-swap, since hot-swap onto NVFP4 was a
silent no-op on vLLM 0.24.0 (#47639). merge_lora.py asserts sampled target
weights actually changed, so an inert adapter cannot ship as a tune.
|
||
|
|
7b5fd91d3c |
docs(gemma4-erp-tune): root-cause the 8.6% MFU — attention on Ampere kernels, 29.9% padding
Run-01 was killed at step 19 by operator instruction to root-cause before spending a ~13.9h window. Two independent methods now agree on where the step time went, and neither was the hypothesis the consult panel converged on. Scaling fit (3 points, 2 params, residuals <3ms over an 8x range): A = 6.87e-4 s/token, B = 8.85e-8 s/token^2 quadratic share 20.9% @ w=2048 -> 67.8% @ w=16384 No fixed term was needed, which refutes launch-bound outright. Profiler kernel table (device rows only): attention 22,835.8 ms 65.2% fmha_cutlass*_sm80 dense GEMM 2,774.0 ms 7.9% other 5,739.0 ms 16.4% The attention kernels are sm80 — Ampere-generation CUTLASS running on an sm_120 Blackwell card, with the forward on the gmem fallback tier. That is the mechanism behind 100% SM utilisation at 27 of 304 available TFLOPS. Correctness cleared separately: the sliding mask asserts at max 1024 allowed/row, so the 25 windowed layers were genuinely windowed. The same probe found that right-padding is what pins the 5 global layers to an explicit 4D mask and off the is_causal fast path — measured at 9.4% slower for 24% less loss work at fixed width. The largest available win is not the attention kernel. The corpus is 29.9% padding, and bucket-to-pair + shuffle-to-mix takes it to 0.0% for >=35.5% wall clock, no new dependency, unchanged peak memory. Bucket size turned out not to be a diversity knob — roots per accumulation window are flat across a 256x range, so the global micro-batch shuffle does that work alone and the bucket should be tight. Adds docs/pfi/training-throughput-playbook.md as the durable model-agnostic home (sibling to the quantization playbook), the four probes under scripts/training-probes/ with raw output kept for re-derivation, and a §6 to the sizing doc carrying the Gemma-4-specific numbers and round-2 restart parameters. Measured negatives recorded so they are not re-chased: grouped_mm (0.9% slower, and MoE is only 7.9% of the step), CUDA graphs / torch.compile over the expert loop (no fixed cost to amortise), liger fused CE (~1-3% lever), FA4 on sm_120. Round-1 state preserved: 609MB encode cache, order manifest, truncation report, resume script. No checkpoints — it died at step 19 and the first was due at 100, so the lora_B inert-adapter gate never ran and moves to the restart. |
||
|
|
872c2c562f |
memory: the MFU hunt — two hypotheses measured and killed, consult dispatched
Records what has actually been ruled out rather than what is suspected. The hardware is fine: a plain dense GEMM at the same shape reaches 97.1% of the benchmarked 313.8 TFLOPS peak. The Python expert loop is not the cause, which was my hypothesis and I was confident in it. transformers' grouped_mm experts backend runs 0.9% SLOWER than eager with bit-identical output and identical peak memory, and torch 2.13 has the kernel available, so it is not falling back for lack of one. MoE is not the bottleneck at all. Isolated at real shapes the block runs at 26.5% of peak with 36% of its time in pure gather/scatter, and a dispatch-free bmm version would reach 80.9% — but the whole MoE contribution is only about 10% of a step. Making it free buys 7%. So roughly 90% of the time is unaccounted for. The leading untested hypothesis is that the five full_attention layers use global_head_dim 512, above FlashAttention-2's 256 cap, which would push SDPA onto a slow backend for O(n^2) attention at sequence 16384. Also records that the earlier 5% MFU figure was wrong in two ways — unpadded tokens and a guessed peak — and that the operator caught it. Padding is real but secondary at 29.9%. Consult dispatched to brokkr-smithy-dev for the frontier-dwarf panel. |