diff --git a/persistent-memory.md b/persistent-memory.md index 84e12a9..f2cde11 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-07-05_ +_Last updated: 2026-07-06_ ## Repo purpose @@ -102,7 +102,17 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-07-05:_ +_As of 2026-07-06:_ + +- **`gen` is now AEON Qwen3.6-27B (qwopus DISPLACED 2026-07-06).** Dual co-located NVFP4 serves on + ana-ml2 GPU0: `vllm-aeon-gen` :8015 (MTP off) + `vllm-aeon-rp` :8016 (native MTP), 256K, multimodal, + **vLLM 0.24.0 + LiteLLM v1.91.0**. Gateway gen/gen-reasoning/summarizer-large→:8015, + char-rp/char-rp-reasoning→:8016; retired qwen3.5-122-a10b[-reasoning]+qwen-large[-reasoning]; qwopus + container stopped (rollback). The reasoning-trace "bug" was the LiteLLM shared-config-mutation footgun, + FIXED durably via distinct `-thinking` served-names (→ distinct LiteLLM deployments). Worldtree personal + wired: character→char-rp, thoughtful-character→char-rp-reasoning. Full detail: auto-memory + `reference_aeon_27b_gen`; `stacks/qwen36-27b-aeon/` (commit e6ab51c). T1 (qwopus E-RP LoRA) is now a + SEPARATE parked effort — qwopus stays the LoRA base, but is no longer the live gen. - **T1 TRAIN — venue decision live; NEXT ACTION = the ana-ml2 SMOKE.** T1 (qwopus E-RP writing LoRA) datasets + recipe + smoke set are all staged (mtf-dev, on `nh3-dev:data/export/qwen-3.5-122b-erp-lora/` @@ -115,12 +125,10 @@ _As of 2026-07-05:_ both paths in `reference_t1_cloud_train_plan`. Post-train serve-path (swappable-LoRA-on-NVFP4 test) still queued: `reference_gen_qwopus_122b`. -- **LitBench-RM UP (irv-ml1 A6000); comfyui DISPLACED.** Respun the R19 creative-writing reward - (`vllm-litbench-rm` :8202, ~19.6G) for T1's reward ensemble; stopped comfyui to free the A6000 - (operator is sole comfyui consumer + holding image-gen until LitBench done). Score scale - re-verified from nh3-dev (literary 0.95 > flat 0.44 > slop 0.12). **TEARDOWN on "litbench done":** - `ssh irv-ml1 'docker rm -f vllm-litbench-rm && docker start comfyui'` — ⚠️ then watch the mmartial - boot trap (root-owned venv → crash-loop; `chown -R 1000:1000 /comfy/mnt/venv`). `reference_litbench_rm_irv_ml1`. +- **LitBench-RM TORN DOWN; comfyui RESTORED (2026-07-06).** Operator called litbench done → + `docker rm -f vllm-litbench-rm && docker start comfyui` on irv-ml1; comfyui booted clean (no + mmartial venv trap — venv intact, HTTP 200 :8188, A6000 reclaimed). Respin litbench on demand + per `reference_litbench_rm_irv_ml1` (weights staged, ~90s) — ⚠️ that re-displaces comfyui. - **character-rp role — SHIPPED; one QUEUED spot-check.** Proved per-request `extra_body` (top_k/repetition_penalty) forwards through the `gen-reasoning` alias to the vLLM sampler (no