feat(coder-seat): Qwen2.5-Coder-1.5B copy on nh3-ml1 (not yet in the gateway)

Same model revision (df3ce67, blob sha matches fv-ml1), vLLM v0.24.0 and flags;
0.33 of the Ada (~5.4 GB, fv-ml1's absolute budget). Healthy, KV 79,808 tokens.

Parity (teacher-forced true code, 40 FIM prompts, 2,215 tokens, cache-salted):
each host is bit-exact with itself; cross-host |dlogprob| median 0.049, top-1
agreement 0.966 (Blackwell vs Ada + fp8 KV); against ground truth no quality
difference (mean logprob diff +0.008 +/- 0.019 SE). Speed on-box: ~5x slower
(64-tok FIM p50 ~1.0 s vs ~0.2 s; 63 vs 338 tok/s). Gateway coder-fast left on
fv-ml1 pending Prime's call.
This commit is contained in:
vh
2026-09-25 23:56:34 -07:00
parent c923727cf2
commit 7f066a4b79
4 changed files with 164 additions and 0 deletions
+7
View File
@@ -33,6 +33,13 @@ Prime's call (recommendation: load-share; see below).
VRAM ~2.7 GB for both, so ~13 GB is free.
**Also here since 2026-09-25: a copy of the code-completion seat.**
`stacks/coder-seat`, `vllm-coder` (Qwen2.5-Coder-1.5B, vLLM v0.24.0) on `:8020`,
about 5 GB. It is **not behind the gateway yet**: on this card it runs about 5×
slower than on fv-ml1 (a 64-token FIM completion takes ~1.0 s vs ~0.2 s), so
Prime is deciding. The GPU total with TEI is ~7.6 of 16 GB. Details:
`stacks/coder-seat/README.md`.
## Parity and speed vs esh-ml1 (2026-09-25 ~1540 PT)
Client on nh3-dev, calling both seats directly. Corpus: 1,120 paragraphs from