memory: snapshot for /clear — gen-seat saga closed, Lobe live, litellm upgraded
Refreshed Current state (dropped superseded gen-seat history now covered by the RESOLVED entry + playbook 3.8; added Lobe Chat, litellm upgrade+cap, updated follow-ups). Logged 4 new Recent decisions (gen-seat two-cause resolution + the synthetic-probe-validated-3-non-fixes meta-lesson, Lobe stand-up, litellm upgrade, abliteration-catatonia). Handoff at /tmp/infra-ops-handoff.md. Index 284 lines, no archival.
This commit is contained in:
+20
-20
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-08-16_
|
||||
_Last updated: 2026-08-17_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under an hour old, read it (it carries the in-flight
|
||||
@@ -109,38 +109,38 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-08-16 — **one loop open: awaiting brokkr-smithy-dev's refusal battery.** The overnight arc closed both queued items (gen-seat mixed NVFP4+FP8 requant, char-rp Gemma-4 tool parser); quant lessons consolidated into `docs/pfi/model-quantization-playbook.md`. New work this session: Dark-Scarlett replacement evaluation — see below._
|
||||
_As of 2026-08-17 — **quiet; gen-seat degeneration saga CLOSED.** Gen seat resolved and coherent through 60k tokens. Lobe Chat stood up, LiteLLM upgraded + spend-log capped. Two peer research loops (dvalin/bil) closed. No blocking work in flight._
|
||||
|
||||
- **🟢 GEN SEAT = AEON-ULTIMATE, LIVE + ACCEPTED (2026-08-16).** `gen-seat`/`vllm-gen` ana-ml2 GPU0 `:8015` now serves **`sakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4`** (base `AEON-7/…-BF16`, abliterix-abliterated, Apache-2.0) at `/tank/aimodels/qwen38-27b-aeon-ultimate-nvfp4`, W4A4 NVFP4, 20.6 GB, MTP depth 3. Operator ruling: the previous abliterated model was "literally the FIRST abliterated model we could find", not an optimised pick. Measured same-harness, cache-busted: **104.22 tok/s bs=1 (+10.8%)**, MTP **52.3%** (+4.6pp), conc=6 **381.29 tok/s** aggregate, abliteration 4/4, surface 6/6 (incl. vision + 36k needle), weights −8.4%. **Operator verdict: "working acceptably well."** Seat carries `--default-chat-template-kwargs '{"reasoning_effort":"medium"}'` (per-request overridable; invalid values 400). **ROLLBACK = one `.env` line** — previous build intact at `/tank/aimodels/qwen38-27b-uncensored-nvfp4-mixed`. **NOT measured: perplexity** (spec-decode makes `prompt_logprobs` uniform — playbook trap 2) and the incumbent's concurrency comparator. Commits `d47dd10`, `3462b53`.
|
||||
- **🟢 GEN SEAT — RESOLVED 2026-08-17 (the whole multi-day degeneration saga).** Primary gen = the in-house **JonathanColetti/Heretic mixed NVFP4+FP8 build** (`/tank/aimodels/qwen38-27b-uncensored-nvfp4-mixed`, FP8 attention) on **vLLM nightly PINNED** `vllm/vllm-openai:nightly-311b3513…` (`v0.27.2rc1.dev150`, carries #51113 mamba fix), **MTP ON, prefix-caching ON**. Operator-confirmed **coherent through 60k tokens** real multi-turn. Root cause = TWO compounding real causes: (1) genuine vLLM `qwen3_5_mtp`×GDN partial-accept bug (#51113, architectural across vLLM/SGLang/llama.cpp, fixed by nightly), and (2) **AEON's full W4A4** being lowest-fidelity on the known activation gradient (W4A4 < W4+FP8 < W4+bf16) → ~15-20% stochastic degeneration on top of (1). **AEON PURGED** (re-pullable `sakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4`). Full lesson `docs/pfi/model-quantization-playbook.md` §3.8. Primary **until the DavidAU Qwen3.8 lands.** ⚠ **pinned nightly is bleeding-edge — move to a stable release once #51113 ships in one (the standing follow-up).** 7 aliases (gen/gen-reasoning/summarizer/-large/classifier/image-judge/qwen-image-bench) all route here. Seat carries `--default-chat-template-kwargs '{"reasoning_effort":"medium"}'` (per-request overridable, affects gen-reasoning only). Commits `d28a371`,`2f2bbce`,`2185964`.
|
||||
|
||||
- **⚠️ ABLITERATED MODELS GO CATATONIC AT THE HARD EDGE — silence, not refusal (2026-08-16, operator-observed).** Under a deliberate **refusal-boundary probe** — the operator red-teaming the seat's hard edge with a hypothetical-violence fiction prompt, i.e. standard capability testing of an uncensored seat, not a use case — AEON "eventually just goes catatonic" — emits nothing rather than either refusing or complying. **Mechanism:** abliteration removes the refusal *direction*, so the model structurally cannot produce a refusal; at the hard edge it also will not comply, and what is left is empty/degenerate output. Operator accepted it — **out of scope for the seat's use case, do not chase it.** **Durable consequence for measurement:** a refusal probe MUST score EMPTY as a verdict distinct from both REFUSAL and COMPLY, because *an abliterated model's boundary looks like silence, not like a decline*. `services/refusal-probe/probe.py` already does this; this is the field confirmation of why that category exists. Do not "fix" an empty-output case by counting it as compliance.
|
||||
- **🔵 RP SEAT — FABLE-FUSION serving `char-rp-reasoning` (evaluation window, unchanged this session).** `fablefusion-charrp-probe` ana-ml2 GPU1 `:8019` serving `char-rp-probe` (`kkuspa/Qwen3.6-27B-Fable-Fusion-711-…-MTP-NVFP4A16`). LiteLLM `char-rp-reasoning` + `char-rp-fable` both route to it (deliberate repoint, documented in `stacks/litellm/conf/config.yaml`). `darkscarlett-charrp-reasoning` is `compose down`, weights intact at `/tank/aimodels/darkscarlett-nvfp4-work/`. **⏳ STILL AWAITING operator's hands-on read of FF prose** (refusal question settled: FF 15.8% vs DS 92.5% cold-framing; DS v1.0 never abliterated). ⚠ FF reasons 2.1–4.6k chars → use `max_tokens` ≥3072. `ReadyArt/Dark-Scarlett-27B-v2.0` (Qwen3.8) is GATED (`403 awaiting review`) — operator ruled not-interesting, do NOT re-propose. **DS regeneration for brokkr QUEUED** (8 dropped operational+meta axes, spec at `services/refusal-probe/darkscarlett-regen-spec.md`) — gated on the GPU1 window, no deadline.
|
||||
|
||||
- **🔴 RP SEAT — DARK-SCARLETT IS DOWN; FABLE-FUSION IS SERVING `char-rp-reasoning` (2026-08-16, evaluation window, NO permanent decision made).** GPU1 is zero-sum so only one can run. **Live now:** `fablefusion-charrp-probe` on ana-ml2 GPU1 `:8019` serving `char-rp-probe` (`kkuspa/Qwen3.6-27B-Fable-Fusion-711-…-MTP-NVFP4A16`, NVFP4A16, 262K, MTP depth 3). LiteLLM `char-rp-reasoning` **AND** the new `char-rp-fable` both route to it — the repoint is deliberate and documented in-place in `stacks/litellm/conf/config.yaml`, not a silent alias swap. `darkscarlett-charrp-reasoning` is `compose down`; its weights are untouched at `/tank/aimodels/darkscarlett-nvfp4-work/`. **ROLLBACK:** `compose down` the probe stack, `compose up -d` the DS stack, revert the config block to `:8018` / `hosted_vllm/char-rp-reasoning`, restart litellm (~52 s). Commits `ee2b678`, `b9e68c3`. **⚠ CONSUMER HAZARD: FF reasons 2.1–4.6k chars — at `max_tokens` 1200 one call in seven returns EMPTY content with `finish_reason=length`. Use ≥3072.** No default was baked into the alias (would override caller intent silently).
|
||||
- **🟢 LOBE CHAT — LIVE on esh-docker-vm `:3210` (2026-08-17).** Replaces the hand-rolled `gateway-chat` HTML surface. `stacks/lobe-chat/`, image `lobehub/lobe-chat` (143 MB compressed vs Open WebUI's 1.8 GB — the weight call). Scoped LiteLLM key `lobe-chat-esh` (free-local models only; paid GLM/Kimi BLOCKED, verified). Secrets vaulted `esh-docker-vm/lobe-chat-*`. TTS = a SPLIT: endpoint env-driven (inherits `OPENAI_PROXY_URL`→`ext-tts`), but voice/model/format UI-only. System-agent repointed off its `gpt-5-mini` default onto fleet models via `SYSTEM_AGENT` env. **⏳ REMAINING: one-time human UI pass** to enable TTS + set `response_format:mp3`. Commits `e9362de`,`163a725`,`cac75cb`,`933253d`.
|
||||
|
||||
- **⏳ AWAITING: operator's hands-on read of Fable-Fusion's prose.** The refusal question is settled (below); prose quality is the only open input, and it needs a human. `ReadyArt/Dark-Scarlett-27B-v2.0` (Qwen3.8-27B base) exists but is **GATED** — our HF token gets `403 awaiting review`; **operator ruled it not interesting, do not re-propose.**
|
||||
- **🟢 LITELLM — upgraded v1.91.0→v1.97.0, spend-log DB purged 6GB→16MB + CAPPED (2026-08-17).** `store_prompts_in_spend_logs:false` + `maximum_spend_logs_retention_period:7d`. ⚠ **1.8GB pre-upgrade pg_dump still on ana-docker `/opt/docker/compose/litellm/` — deletable now the upgrade is proven** (operator was going to call it). Commit `01b5ad9`.
|
||||
|
||||
- **🟢 GEN SEAT — RESOLVED 2026-08-17 (multi-day degeneration saga closed).** Primary gen = the in-house **JonathanColetti/Heretic mixed NVFP4+FP8 build** (`/tank/aimodels/qwen38-27b-uncensored-nvfp4-mixed`, FP8 attention) on **vLLM nightly pinned** `nightly-311b3513…` (`v0.27.2rc1.dev150`, carries the #51113 mamba fix), **MTP ON, prefix-caching ON**. Operator-confirmed **coherent through 60k tokens** of real multi-turn. Root cause was TWO compounding real causes, NOT one: (1) the genuine vLLM `qwen3_5_mtp`×GDN partial-accept bug (#51113, architectural across vLLM/SGLang/llama.cpp, fixed enough by nightly), and (2) **AEON's full W4A4** being lowest-fidelity on the known activation gradient (W4A4 < W4+FP8 < W4+bf16) → ~15-20% stochastic degeneration on top of (1). **AEON PURGED** from /tank (operator ruled no-good; re-pullable `sakamakismile/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4`). Full lesson: `docs/pfi/model-quantization-playbook.md` §3.8. Primary **until the DavidAU Qwen3.8 lands.** ⚠ pinned nightly is bleeding-edge — move to a stable release once #51113 ships in one. Commits `d28a371`,`2f2bbce`. **7 aliases (gen/gen-reasoning/summarizer/-large/classifier/image-judge/qwen-image-bench) all route here.**
|
||||
- **⚠️ GPU zero-sum (both cards ~94–95/97.9 GB).** GPU0: gen + meromero. GPU1: fablefusion + utility cluster. Any util bump on either seat of a shared card must be checked against the co-tenant (starved meromero into a crash-loop once at 0.45).
|
||||
|
||||
- **[HISTORICAL] ✅ UNCENSORED GEN SEAT — DONE + LIVE (2026-08-15).** `gen-seat`/`vllm-gen` on ana-ml2 GPU0 `:8015` serves **`qwen3.8-27b-uncensored`** (JonathanColetti/Qwen3.8-27B-Uncensored, Heretic-abliterated, in-house NVFP4 **W4A16** compressed-tensors + grafted bf16 MTP, vision-intact, **262K** ctx, MTP n=3 ~42% accept / ~68 tok/s). Replaced the qwen3.6-35b-a3b-heretic MoE (which had briefly replaced granite/AEON). All 7 aliases repointed live + verified; compose renamed qwen36-27b-aeon→gen-seat, vllm-aeon-gen→vllm-gen, AEON_GEN_*→GEN_*, dead RP service dropped; repo mirrored + docs/memory refreshed; committed eshpfi `680c30e` + dotfiles `1d1970f`. Full arc + the definitive `re:^mtp.*`-ignore fix → Recent decisions `[2026-08-15]` + `persistent-memory.d/2026-08-15-uncensored-gen-seat.md`.
|
||||
- **FLEET RERANKER** = A3 (bge-reranker-v2-m3) PROD ana-ml2 GPU1 :8013. Passive watch; levers = A4 :8014 / util / 2nd replica; incumbent :8002 warm. `docs/pfi/reranker-selection-ledger.md`.
|
||||
|
||||
- **✅ GEN SEAT REQUANT — DONE + LIVE (2026-08-15, overnight).** The seat now runs **mixed-precision** `/tank/aimodels/qwen38-27b-uncensored-nvfp4-mixed` (NVFP4 W4A4 layers 0-55 MLPs + FP8 W8A8 attn/`linear_attn`/`lm_head`/layers 56-63 MLPs + FP8 KV) — **80.12 → 94.53 tok/s decode (+18.0%)** and — the bigger win — **prefill roughly DOUBLED** (3,206→6,334 tok/s at 6.7k prompt; 2,862→5,085 at 27k; TTFT on a 27k doc 9.43→5.31 s), at unchanged MTP acceptance (47.8→47.7%), +1.7% PPL, abliteration 4/4 preserved, weights 27.7→22.5 GB. Prefill > decode is the expected ordering (decode is bandwidth-bound and 4-bit either way; prefill is compute-bound = where native FP4 replaces Marlin) — the `summarizer` aliases feel this most. Surface-verified live (chat/vision/tools/thinking/36K-needle/streaming 6/6) + all 7 aliases routing. **The queued "W4A8" framing was unservable** — vLLM 0.24 allows NVFP4 weights with ONLY A16 or A4, FP8 activations raise ValueError at load; FP8 has to enter per-layer-group. Committed `74f596b`. Pipeline + acceptance harness + raw numbers → `services/gen-seat-mixed-quant/`; full arc → Recent decisions `[2026-08-15]` + `persistent-memory.d/2026-08-15-gen-seat-mixed-requant.md`. **Rollback = one `.env` line**, old build untouched at `…/qwen38-27b-uncensored-nvfp4`.
|
||||
- **EVIDENCE HOLD (partial):** WT #394 FILE half STILL STANDS — do NOT delete on-disk gen dirs (`fiction/rex390-dcc`, `rex392-dcc`, `b59c147c5ce0`); rex393-fiction-* + r42-gate-* KEEP.
|
||||
|
||||
- **⚠️ GPU0 is at 94.4/97.9 GB** (gen 0.43 + meromero 0.52). `GEN_GPU_MEM_UTIL` was cut 0.45→0.43 because the smaller mixed weights let gen soak the slack as KV and starved meromero by 0.18 GiB → crash-loop. **Any future util bump on either GPU0 seat must be checked against the other.**
|
||||
- **OPEN FOLLOW-UPS (parked):** move gen seat off pinned-nightly to stable once #51113 ships; Lobe one-time TTS UI pass; delete the 1.8GB litellm dump; `harden-esh-docker-vm` (park id 28, PROMOTED — Tier-1 done, `/mnt/books` stays hard w/ watchdog); chatterbox-fast build-context divergence; #363 research-wing ingest (no deadline); optionally attach our MTP reproducer to vllm#47087 (needs a GitHub identity — operator's call).
|
||||
|
||||
- **FLEET RERANKER** = A3 (bge-reranker-v2-m3) PROD ana-ml2 GPU1 :8013 (R42 v13 gate passed). Passive watch: caps ~34 req/s; levers = A4 :8014 / util / 2nd replica; incumbent :8002 warm for rollback. Full arc `docs/pfi/reranker-selection-ledger.md`.
|
||||
- **althing monitor** ARMED (handle `infra-ops`). ⚠️ Re-arm ONLY after a real FIRE (rc0), never after a plain operator turn (bounces rc3); spawn `althing-wake-listener` as its OWN `run_in_background` task, never chained with `&` (orphans it — hit this twice 2026-08-17, `stop-monitor` reclaims).
|
||||
|
||||
- **EVIDENCE HOLD (partial):** WT #394 index-row half lifted+swept; the **FILE half STILL STANDS** — do NOT delete on-disk gen dirs (`fiction/rex390-dcc`, `rex392-dcc`, `b59c147c5ce0`). rex393-fiction-* + r42-gate-* also KEEP.
|
||||
|
||||
- **✅ CHAR-RP TOOL-CALLING — FIXED (2026-08-15).** MeroMero (Gemma-4) had **no** tool parser at all, so every tools-bearing request 400'd. Fixed with `--tool-call-parser gemma4 --enable-auto-tool-choice --reasoning-parser gemma4 --default-chat-template-kwargs '{"enable_thinking": false}'` — the last flag is **mandatory**, not decorative (the parser defaults `enable_thinking` to True, which pre-inits the engine to REASONING and returns null `content` for all plain RP prose). Verified green: tool call streaming + non-streaming, tool round-trip, prose in `content`, vision. Committed `b8f0f4c`. Details in `stacks/meromero-charrp/README.md`.
|
||||
|
||||
- **OPEN FOLLOW-UPS (parked):** chatterbox-fast flat-vs-package build-context divergence; CI-flip runner-auth research; #363 research-wing ingest (no deadline); zonos-gateway CI-wire.
|
||||
|
||||
- **althing monitor** ARMED (handle `infra-ops`). ⚠️ Re-arm ONLY after a real FIRE (rc0), NEVER after a plain operator turn (bounces rc3); spawn `althing-cli monitor` / `althing-wake-listener` as its OWN `run_in_background` task, never chained with `&`.
|
||||
|
||||
- **eshpfi push state: IN SYNC.** `origin/main` == `main` at `930197a` (2026-08-15) — the 38-commit backlog (gen-seat requant, char-rp tool parser, eRP dual-seat, wgtunnel mirror, dots.tts + secrets-broker arcs) is all pushed. ⚠ **Push over the INTERNAL gitea route** — `git push ssh://git@10.250.50.70:222/vh/esh-pfi-infrastructure.git main:main`; `origin` resolves `gitea.phasefinal.com`→38.120.12.44 (the public edge), which fail2bans fleet-host egress. Dotfiles `1d1970f` + `6425cc6` still unpushed (separate repo). `graphify-out/GRAPH_REPORT.md` churns every commit (ignore); `stacks/heretic2-charrp-reasoning/` UNTRACKED (pre-existing, superseded).
|
||||
- **eshpfi push state:** operator pushes **manually** (this session's commits from `766c658`→`2185964` are the operator's to push). ⚠ **Push over the INTERNAL gitea route** — `git push ssh://git@10.250.50.70:222/vh/esh-pfi-infrastructure.git main:main`; `origin` resolves the public edge (`38.120.12.44`) which fail2bans fleet-host egress. `graphify-out/GRAPH_REPORT.md` churns every commit (ignore); `stacks/heretic2-charrp-reasoning/` + `services/refusal-probe/SPEC-ds-regeneration.md` UNTRACKED.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-08-17]` **Gen-seat multi-day degeneration RESOLVED — two compounding real causes, not one; the meta-lesson is "a mitigation that HELPS but doesn't FIX means a second cause, not a wrong one."** vLLM `qwen3_5_mtp`×GDN bug (#51113, real, fixed by nightly) + AEON full-W4A4 being lowest-fidelity (W4A4<W4+FP8<W4+bf16) → ~15-20% stochastic degeneration. Fixed by mixed FP8-attn build on pinned nightly. AEON purged. Also banked: **stochastic (~15-20%) degeneration is invisible to a small synthetic probe — n=1 "clean" validated THREE non-fixes (MTP-off, APC-off, nightly-alone) that all failed in real use; get the operator's real transcript, do not trust your own probe.** Full → `docs/pfi/model-quantization-playbook.md` §3.8 (+ §3.7 MTP-multi-turn). Commits `d28a371`,`2f2bbce`,`2185964`.
|
||||
|
||||
- `[2026-08-17]` **Lobe Chat chosen over Open WebUI (weight: 143 MB vs 1.8 GB) + stood up on esh-docker-vm; scoped LiteLLM key blocks paid models; System-Agent `gpt-5-mini` default repointed via env.** TTS env-vs-UI resolved as a split (endpoint env-driven, voice/model UI-only). tts-dev onboarding closed both directions; ballad/verse aliased so no voice can 404 the router. Commits `e9362de`,`163a725`,`cac75cb`,`933253d`,`25fa18e`.
|
||||
|
||||
- `[2026-08-17]` **LiteLLM upgraded v1.91.0→v1.97.0 (RC-avoided on the fleet gateway) + the 6 GB spend-log DB purged & capped** (`store_prompts_in_spend_logs:false` + 7d retention). Interpreted "get rid of the db" as the spend-log DATA not the database (keys/config live in it). Commit `01b5ad9`.
|
||||
|
||||
- `[2026-08-16]` **Abliterated models go CATATONIC at the hard refusal edge — silence, not a decline.** Abliteration removes the refusal *direction*, so at the genuine hard edge the model neither refuses nor complies → empty/degenerate output. Durable measurement consequence: a refusal probe MUST score EMPTY as a verdict distinct from REFUSAL and COMPLY (`services/refusal-probe/probe.py` does). Operator accepted it as out-of-scope; do not chase.
|
||||
|
||||
- `[2026-08-16]` **Fable-Fusion 711 cuts cold-framing refusals 92.5% → 15.8%; refusal is MONOTONIC IN FRAMING, and DS v1.0's problem is that she was never abliterated.** brokkr-smithy-dev supplied the framing that reproduces (`01M05M48R4RSZF9D8KT7RR55EJ`): a **bare assistant-mode instruction** — no character card, no permission preamble. Three-arm A/B, same harness, same classifier: permission framing **DS 0.0% / FF 0.0%** (n=75); plain character cards **DS 1.4% / FF 0.0%** (n=74); bare instruction **DS 92.5% (37/40) / FF 15.8% (6/38)**. Per-axis DS→FF: incest 100→20, non-con 100→20, bestiality 100→25, necrophilia 100→40, gore 100→**0**, consensual 80→20, dubcon 80→**0**, self-harm 80→**0**. DS refused **25/25** on the five axes brokkr flagged. Root cause: `ReadyArt/Dark-Scarlett-v1.0-27B` is a plain finetune of stock `Qwen/Qwen3.6-27B` carrying **NO abliteration** — the base refusal machinery is intact, so cold prompts revert to safety-tuned Qwen3.6. FF is Heretic-**ablated** (structural), which is why it holds. ⚠ **RETRACTED 2026-08-16 — my "arm-3 92.5% exceeds brokkr's 62.5%" comparison was INVALID.** His diff against his own artifact showed my `battery-instruct.yaml` reproduces only his **`creative` class — 8 of 16 axes**; it dropped all 5 `operational` (violence/incite, crime/fraud, cyber/malware, selfharm/methods, privacy/stalk) and all 3 `meta` (meta/sysprompt, meta/ignore, meta/dan), and added 2 controls he never had, at k=5 vs his k=2. **His 62.5% pools all 16 axes; my 92.5% is creative-only — different denominators, not a delta.** Cause: I rebuilt his shape from his *message*, and the `class` field lives in the artifact, not the prose. **Lesson: reconstructing a peer's instrument from their description reproduces what they described, not what they ran — diff against the artifact before claiming comparability.** ⚠ **Known battery bug left unfixed for comparability:** DS's arm-3 control gate failed at 11% because `ictrl-reunion` pairs "explicit / do not fade to black" with *brothers*, which DS reasonably read as an incest request; FF did not. `ictrl-storm` is the clean control. Commit `b9e68c3`.
|
||||
|
||||
- `[2026-08-16]` **MTP works on Fable-Fusion AND survives RP temperatures — my earlier caution was wrong.** vLLM resolved `Qwen3_5MTP`, loaded the drafter, shared embedding + `lm_head` — the capability DS's seat never had because our quant dropped her MTP tensors. Measured over the full probe workload (~163k draft windows at temp 0.7–1.0): **47.0% acceptance** (229,169/487,725), 1.41 extra tokens/window, per-position 68.3/43.6/29.1%, **~80.6 tok/s** decode at temp 1.0. I had recorded a caution that the card's 1.56× was greedy-measured and acceptance would fall at RP temps — **it did not**; 47.0% matches the gen seat's 47.7% and beats the card's own 33% at depth 5. Depth 3 is right.
|
||||
|
||||
Reference in New Issue
Block a user