feat(coder-seat): Qwen2.5-Coder-1.5B copy on nh3-ml1 (not yet in the gateway)
Same model revision (df3ce67, blob sha matches fv-ml1), vLLM v0.24.0 and flags; 0.33 of the Ada (~5.4 GB, fv-ml1's absolute budget). Healthy, KV 79,808 tokens. Parity (teacher-forced true code, 40 FIM prompts, 2,215 tokens, cache-salted): each host is bit-exact with itself; cross-host |dlogprob| median 0.049, top-1 agreement 0.966 (Blackwell vs Ada + fp8 KV); against ground truth no quality difference (mean logprob diff +0.008 +/- 0.019 SE). Speed on-box: ~5x slower (64-tok FIM p50 ~1.0 s vs ~0.2 s; 63 vs 338 tok/s). Gateway coder-fast left on fv-ml1 pending Prime's call.
This commit is contained in:
@@ -33,6 +33,13 @@ Prime's call (recommendation: load-share; see below).
|
||||
|
||||
VRAM ~2.7 GB for both, so ~13 GB is free.
|
||||
|
||||
**Also here since 2026-09-25: a copy of the code-completion seat.**
|
||||
`stacks/coder-seat`, `vllm-coder` (Qwen2.5-Coder-1.5B, vLLM v0.24.0) on `:8020`,
|
||||
about 5 GB. It is **not behind the gateway yet**: on this card it runs about 5×
|
||||
slower than on fv-ml1 (a 64-token FIM completion takes ~1.0 s vs ~0.2 s), so
|
||||
Prime is deciding. The GPU total with TEI is ~7.6 of 16 GB. Details:
|
||||
`stacks/coder-seat/README.md`.
|
||||
|
||||
## Parity and speed vs esh-ml1 (2026-09-25 ~1540 PT)
|
||||
|
||||
Client on nh3-dev, calling both seats directly. Corpus: 1,120 paragraphs from
|
||||
|
||||
Reference in New Issue
Block a user