memory: snapshot — AEON Qwen3.6-27B is now gen (qwopus displaced), litbench torn down/comfyui restored
- gen := AEON dual NVFP4 serves (vLLM 0.24 + LiteLLM v1.91.0); reasoning-trace bug was the LiteLLM shared-config mutation, fixed durably via distinct -thinking served-names. - Worldtree personal character/thoughtful-character repointed to char-rp/char-rp-reasoning. - LitBench-RM torn down, comfyui restored on irv-ml1.
This commit is contained in:
+16
-8
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-07-05_
|
||||
_Last updated: 2026-07-06_
|
||||
|
||||
## Repo purpose
|
||||
|
||||
@@ -102,7 +102,17 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-07-05:_
|
||||
_As of 2026-07-06:_
|
||||
|
||||
- **`gen` is now AEON Qwen3.6-27B (qwopus DISPLACED 2026-07-06).** Dual co-located NVFP4 serves on
|
||||
ana-ml2 GPU0: `vllm-aeon-gen` :8015 (MTP off) + `vllm-aeon-rp` :8016 (native MTP), 256K, multimodal,
|
||||
**vLLM 0.24.0 + LiteLLM v1.91.0**. Gateway gen/gen-reasoning/summarizer-large→:8015,
|
||||
char-rp/char-rp-reasoning→:8016; retired qwen3.5-122-a10b[-reasoning]+qwen-large[-reasoning]; qwopus
|
||||
container stopped (rollback). The reasoning-trace "bug" was the LiteLLM shared-config-mutation footgun,
|
||||
FIXED durably via distinct `-thinking` served-names (→ distinct LiteLLM deployments). Worldtree personal
|
||||
wired: character→char-rp, thoughtful-character→char-rp-reasoning. Full detail: auto-memory
|
||||
`reference_aeon_27b_gen`; `stacks/qwen36-27b-aeon/` (commit e6ab51c). T1 (qwopus E-RP LoRA) is now a
|
||||
SEPARATE parked effort — qwopus stays the LoRA base, but is no longer the live gen.
|
||||
|
||||
- **T1 TRAIN — venue decision live; NEXT ACTION = the ana-ml2 SMOKE.** T1 (qwopus E-RP writing
|
||||
LoRA) datasets + recipe + smoke set are all staged (mtf-dev, on `nh3-dev:data/export/qwen-3.5-122b-erp-lora/`
|
||||
@@ -115,12 +125,10 @@ _As of 2026-07-05:_
|
||||
both paths in `reference_t1_cloud_train_plan`. Post-train serve-path (swappable-LoRA-on-NVFP4
|
||||
test) still queued: `reference_gen_qwopus_122b`.
|
||||
|
||||
- **LitBench-RM UP (irv-ml1 A6000); comfyui DISPLACED.** Respun the R19 creative-writing reward
|
||||
(`vllm-litbench-rm` :8202, ~19.6G) for T1's reward ensemble; stopped comfyui to free the A6000
|
||||
(operator is sole comfyui consumer + holding image-gen until LitBench done). Score scale
|
||||
re-verified from nh3-dev (literary 0.95 > flat 0.44 > slop 0.12). **TEARDOWN on "litbench done":**
|
||||
`ssh irv-ml1 'docker rm -f vllm-litbench-rm && docker start comfyui'` — ⚠️ then watch the mmartial
|
||||
boot trap (root-owned venv → crash-loop; `chown -R 1000:1000 /comfy/mnt/venv`). `reference_litbench_rm_irv_ml1`.
|
||||
- **LitBench-RM TORN DOWN; comfyui RESTORED (2026-07-06).** Operator called litbench done →
|
||||
`docker rm -f vllm-litbench-rm && docker start comfyui` on irv-ml1; comfyui booted clean (no
|
||||
mmartial venv trap — venv intact, HTTP 200 :8188, A6000 reclaimed). Respin litbench on demand
|
||||
per `reference_litbench_rm_irv_ml1` (weights staged, ~90s) — ⚠️ that re-displaces comfyui.
|
||||
|
||||
- **character-rp role — SHIPPED; one QUEUED spot-check.** Proved per-request `extra_body`
|
||||
(top_k/repetition_penalty) forwards through the `gen-reasoning` alias to the vLLM sampler (no
|
||||
|
||||
Reference in New Issue
Block a user