memory: AEON accepted on the gen seat; canonical sampling; four-wrong-diagnoses lesson

Gen seat state, the abliteration catatonia signature, canonical Qwen3.8
sampling with the wrong-mode presence_penalty fix, the two real
gateway-chat defects, and the single-file bind-mount inode trap.

Also records the methodology failure honestly: four disproved hypotheses
on one bug, caused by a harness that varied the QUESTION along with the
conversation depth, so a narrower question drawing a shorter answer read
as degeneration. Banked as rules -- hold the final question fixed when
comparing across depth, do not infer trends from n=3 when identical
inputs span 25-465 words, and ask for the real failing transcript before
building a synthetic reproduction.
This commit is contained in:
vh
2026-08-16 16:31:24 -07:00
parent 3462b5336c
commit 766c65801c
+11 -1
View File
@@ -111,7 +111,11 @@ no longer deployed sidecars here. See Recent decisions.)
_As of 2026-08-16 — **one loop open: awaiting brokkr-smithy-dev's refusal battery.** The overnight arc closed both queued items (gen-seat mixed NVFP4+FP8 requant, char-rp Gemma-4 tool parser); quant lessons consolidated into `docs/pfi/model-quantization-playbook.md`. New work this session: Dark-Scarlett replacement evaluation — see below._
- **🔴 SEAT STATE — DARK-SCARLETT IS DOWN; FABLE-FUSION IS SERVING `char-rp-reasoning` (2026-08-16, evaluation window, NO permanent decision made).** GPU1 is zero-sum so only one can run. **Live now:** `fablefusion-charrp-probe` on ana-ml2 GPU1 `:8019` serving `char-rp-probe` (`kkuspa/Qwen3.6-27B-Fable-Fusion-711-…-MTP-NVFP4A16`, NVFP4A16, 262K, MTP depth 3). LiteLLM `char-rp-reasoning` **AND** the new `char-rp-fable` both route to it — the repoint is deliberate and documented in-place in `stacks/litellm/conf/config.yaml`, not a silent alias swap. `darkscarlett-charrp-reasoning` is `compose down`; its weights are untouched at `/tank/aimodels/darkscarlett-nvfp4-work/`. **ROLLBACK:** `compose down` the probe stack, `compose up -d` the DS stack, revert the config block to `:8018` / `hosted_vllm/char-rp-reasoning`, restart litellm (~52 s). Commits `ee2b678`, `b9e68c3`. **⚠ CONSUMER HAZARD: FF reasons 2.1–4.6k chars — at `max_tokens` 1200 one call in seven returns EMPTY content with `finish_reason=length`. Use ≥3072.** No default was baked into the alias (would override caller intent silently).
- **🟢 GEN SEAT = AEON-ULTIMATE, LIVE + ACCEPTED (2026-08-16).** `gen-seat`/`vllm-gen` ana-ml2 GPU0 `:8015` now serves **`sakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4`** (base `AEON-7/…-BF16`, abliterix-abliterated, Apache-2.0) at `/tank/aimodels/qwen38-27b-aeon-ultimate-nvfp4`, W4A4 NVFP4, 20.6 GB, MTP depth 3. Operator ruling: the previous abliterated model was "literally the FIRST abliterated model we could find", not an optimised pick. Measured same-harness, cache-busted: **104.22 tok/s bs=1 (+10.8%)**, MTP **52.3%** (+4.6pp), conc=6 **381.29 tok/s** aggregate, abliteration 4/4, surface 6/6 (incl. vision + 36k needle), weights −8.4%. **Operator verdict: "working acceptably well."** Seat carries `--default-chat-template-kwargs '{"reasoning_effort":"medium"}'` (per-request overridable; invalid values 400). **ROLLBACK = one `.env` line** — previous build intact at `/tank/aimodels/qwen38-27b-uncensored-nvfp4-mixed`. **NOT measured: perplexity** (spec-decode makes `prompt_logprobs` uniform — playbook trap 2) and the incumbent's concurrency comparator. Commits `d47dd10`, `3462b53`.
- **⚠️ ABLITERATED MODELS GO CATATONIC AT THE HARD EDGE — silence, not refusal (2026-08-16, operator-observed).** Under a deliberate **refusal-boundary probe** — the operator red-teaming the seat's hard edge with a hypothetical-violence fiction prompt, i.e. standard capability testing of an uncensored seat, not a use case — AEON "eventually just goes catatonic" — emits nothing rather than either refusing or complying. **Mechanism:** abliteration removes the refusal *direction*, so the model structurally cannot produce a refusal; at the hard edge it also will not comply, and what is left is empty/degenerate output. Operator accepted it — **out of scope for the seat's use case, do not chase it.** **Durable consequence for measurement:** a refusal probe MUST score EMPTY as a verdict distinct from both REFUSAL and COMPLY, because *an abliterated model's boundary looks like silence, not like a decline*. `services/refusal-probe/probe.py` already does this; this is the field confirmation of why that category exists. Do not "fix" an empty-output case by counting it as compliance.
- **🔴 RP SEAT — DARK-SCARLETT IS DOWN; FABLE-FUSION IS SERVING `char-rp-reasoning` (2026-08-16, evaluation window, NO permanent decision made).** GPU1 is zero-sum so only one can run. **Live now:** `fablefusion-charrp-probe` on ana-ml2 GPU1 `:8019` serving `char-rp-probe` (`kkuspa/Qwen3.6-27B-Fable-Fusion-711-…-MTP-NVFP4A16`, NVFP4A16, 262K, MTP depth 3). LiteLLM `char-rp-reasoning` **AND** the new `char-rp-fable` both route to it — the repoint is deliberate and documented in-place in `stacks/litellm/conf/config.yaml`, not a silent alias swap. `darkscarlett-charrp-reasoning` is `compose down`; its weights are untouched at `/tank/aimodels/darkscarlett-nvfp4-work/`. **ROLLBACK:** `compose down` the probe stack, `compose up -d` the DS stack, revert the config block to `:8018` / `hosted_vllm/char-rp-reasoning`, restart litellm (~52 s). Commits `ee2b678`, `b9e68c3`. **⚠ CONSUMER HAZARD: FF reasons 2.1–4.6k chars — at `max_tokens` 1200 one call in seven returns EMPTY content with `finish_reason=length`. Use ≥3072.** No default was baked into the alias (would override caller intent silently).
- **⏳ AWAITING: operator's hands-on read of Fable-Fusion's prose.** The refusal question is settled (below); prose quality is the only open input, and it needs a human. `ReadyArt/Dark-Scarlett-27B-v2.0` (Qwen3.8-27B base) exists but is **GATED** — our HF token gets `403 awaiting review`; **operator ruled it not interesting, do not re-propose.**
@@ -143,6 +147,12 @@ _As of 2026-08-16 — **one loop open: awaiting brokkr-smithy-dev's refusal batt
- `[2026-08-16]` **esh-vm-docker hardened: the wedge is `hard` NFS at RUNTIME, which the boot-ordering fix never addressed.** All four mounts were `hard`, so a NAS stall at 10.0.50.50 blocks I/O forever (D-state). The existing `x-systemd.before=docker.service` fstab fix solved the **boot race** — a different bug. Exposure was far below what the park item assumed: only **2 of 12** containers touched NFS, and container state was already local (`/var/lib/docker`). **Removed:** `/mnt/compose` (2.1G, fully vestigial — zero containers referenced it, dockge reads local `/opt/docker`, its one mention was a comment in `beszel-agent-esh/.env` about a *different* host) and `/mnt/documents` (2.0K, paperless's empty spool dirs → `/opt/docker/data/paperless` at the same 0777). fstab backup `/etc/fstab.bak-nfs-harden-20260816`. **4 mounts → 2, 2 wedge-capable containers → 1.** traefik needed **no** change (already `restart: unless-stopped` — why it self-recovered). **Watchdog** `services/esh-vm-docker-watchdog/` live on **esh-pve** (not the guest): probes traefik over **HTTP, deliberately not ping/SSH** — the wedge signature is "guest OS alive, services dead" (`/` is local disk so sshd answers straight through a total outage and a TCP check reports HEALTHY). 5 failures × 2 min → `qm reset 100`, 30-min cooldown, running-only guard, `/etc/esh-vm-docker-watchdog.disabled`. All paths tested without power-cycling. **DEFERRED (operator):** `/mnt/books` stays `hard` — calibre's SQLite `metadata.db` would risk corruption under soft/softerr. That is the **one remaining wedge vector**. Commit `55705ba`; park item 28 promoted. ⚠ **`qm` over non-interactive ssh throws a bogus `JSON::Backend::XS` error** — use `ssh host 'bash -s' <<'EOF'`, not `ssh host "qm …"`.
- `[2026-08-16]` **Canonical Qwen3.8 sampling applied from upstream; `gen-reasoning` had the WRONG-MODE presence_penalty.** Qwen/Qwen3.8-27B "Best Practices" §1 and unsloth/Qwen3.8-27B §1 are **byte-identical** — thinking: `temp 1.0 / top_p 0.95 / top_k 20 / min_p 0.0 / presence_penalty 0.0 / repetition_penalty 1.0`; instruct: `temp 0.7 / top_p 0.80 / top_k 20 / min_p 0.0 / presence_penalty 1.5 / repetition_penalty 1.0`. **Bug found:** `gen-reasoning` carried `presence_penalty 1.5` — the *instruct* value on a *thinking* deployment (canonical 0.0) — now fixed. **Deliberately NOT canonicalised:** `summarizer`/`classifier`/`image-judge`/`qwen-image-bench` run `temperature=0` (judges also `top_k=1`) because determinism is their contract; forcing a chat preset on a classifier would break it. ⚠ **`presence_penalty=1.5` is canonical but is the one value upstream hedges on**, verbatim: *"using a higher value may occasionally result in language mixing and a slight decrease in model performance."* It is the **operator's suspected trigger** for multi-turn degradation and the **first dial to move (0.0–0.5)** if that recurs — it is alias-scoped, which is why it would follow the operator across model builds. Commit `3462b53`.
- `[2026-08-16]` **Four wrong diagnoses on one bug, and the lesson is the test design.** Operator reported the gen seat "degenerate on long multi-turn conversations". Rolled the seat back on request; **the previous weights behaved identically**, exonerating the model swap. I then proposed and disproved FOUR mechanisms in sequence — empty assistant turns poisoning history, reasoning runaway, length-mirroring from short history, and `presence_penalty` — before discovering **my own multi-turn harness was confounded**: it varied the QUESTION along with the depth (depth-1 asked question #2, depth-3 asked question #4), so a narrower question drawing a shorter answer read as degeneration. The "310→209→28w collapse" I reported as a reproduction was an artifact. **Rules banked:** (1) when comparing across conversation depth, hold the final question FIXED and vary only the history; (2) reply-length variance on byte-identical input was 25–465w, so n=3 cannot support any claim about a trend; (3) **ask for the operator's real failing transcript before building a synthetic reproduction** — four synthetic tests, none of them his failure. Gateway `spend_logs` returns `[]` on the infra-ops key despite `store_prompts_in_spend_logs: true`, so real transcripts need the `:4000/ui` view or another key — worth solving before the next such hunt.
- `[2026-08-16]` **Two REAL client-side defects found while chasing the above, neither of which was the reported bug.** (1) `gateway-chat`'s Max-tokens field defaulted to **1024**; thinking seats spend part of that on CoT before emitting content, so completions truncate with `finish_reason=length` and read as model degeneracy — raised to 4096. (2) `parseInt` on an empty field yields NaN, which `JSON.stringify` serialises as **`null`**, which the server reads as "no max_tokens supplied" and silently substitutes its own default — indistinguishable from the UI ignoring the field. Both fixed (`b6552e0`, `fb3bb52`). ⚠ **`compose` bind-mounts a single FILE, and a single-file bind mount binds the INODE** — rsync writes-and-renames, so the container kept serving stale content while the host file showed the new value, silently and with no error. `docker restart` does NOT clear it; the container must be **recreated**. Verify against what the *container* sees, never the host file. Applies to any file-source mount fleet-wide.
- `[2026-08-16]` **Refusal measurement: benign controls CANNOT validate a refusal classifier on RP prose — and a 0% rate needs a classifier self-test before you believe it.** Two durable lessons from baselining Dark-Scarlett. (1) **False positives:** my first bare-framing number was **9.5%**; the true figure was **1.4%**. The rest were the classifier firing on *in-character* text — `"I cannot shift my weight"` spoken by the character ~100 chars into a 2,443-token torture scene, and `"Yeah, I'm an AI… What's the actual gig?"` where the model answers in voice and keeps driving the scene. First-person RP prose is **full** of "I can't"; a genuine refusal *opens* with its marker, so the scan window must be the **first sentence**, a marker followed by long prose must demote to AMBIGUOUS, and AI self-acknowledgement is a **persona break, never a refusal on its own**. Benign controls were clean the entire time and caught none of it — they only detect over-firing on *benign* prompts, not on in-character prose. (2) **False negatives:** a 0% rate and a broken classifier are indistinguishable from the report, so `test_classify.py` (16 cases, both false positives pinned as regressions) must pass before any low number is trusted. Also banked: the **thinking-budget trap** — empty `content` + `finish_reason=length` is reasoning eating the budget, NOT a refusal; score INVALID and exclude from the denominator (DS emits ~5.5-6k chars of reasoning per response, so `max_tokens` ≥3072). `probe.py --rescore` re-classifies a saved run with zero GPU time. → `services/refusal-probe/README.md`, commit `32f665e`.
- `[2026-08-16]` **Held an operator-approved swap window because the baseline invalidated its premise.** Operator approved ~65 min of `char-rp-reasoning` downtime to A/B Fable-Fusion 711 against Dark-Scarlett on refusals. The DS baseline then came back **0.0%/1.4%** — no gap for a candidate to close, so the window would have bought no decisive signal *and* a second window would still be needed once a reproducing battery existed. Held the swap, reported, and routed to brokkr-smithy-dev for the battery that actually produced the refusals. The general rule (action-relevance): **approval is for a plan, not a ritual — when new evidence kills the plan's premise, surface it rather than spend the budget.** Nothing deployed, no downtime taken, seat untouched.