revert(coder-seat): keep Qwen2.5-Coder on fv-ml1; nh3-ml1 copy removed (5x slower, same quality)

This commit is contained in:
vh
2026-09-25 23:58:08 -07:00
parent 7f066a4b79
commit ddd67df2e9
3 changed files with 15 additions and 10 deletions
+7 -4
View File
@@ -17,11 +17,14 @@ Blackwells). It was chosen for Zed edit-predictions by a deep-research pass on
| revision | `df3ce67c0e24480f20468b6ef2894622d69eb73b` (pinned; same as fv-ml1 and HF main on 2026-09-25) |
| memory | 0.33 of 16,380 MiB ≈ 5.4 GB, the same absolute budget as on fv-ml1 |
## ⏸ Status 2026-09-25 2355: running, NOT behind the gateway (Prime deciding)
## ❌ NOT DEPLOYED. Decision (Prime, 2026-09-26 0000): the seat stays on fv-ml1
The copy is up and parity-checked, but the gateway still points `coder-fast` at
fv-ml1. On the 50 W Ada the same model is **about 5× slower**, which Zed
edit-predictions would notice.
The nh3-ml1 copy was built, checked for parity, measured, and then removed (container,
image, model cache). It was the same quality but **about 5× slower** on the 50 W
Ada. That lag would reach Zed edit-predictions, and moving it would only have freed
6.3 GB on an fv-ml1 GPU that already had about 20 GB spare. The live seat is still
`stacks/vllm` on fv-ml1. This stack is kept as the tested recipe in case the
calculation changes (a faster card, or speculative decoding).
**Speed, measured on each box** (3 interleaved reps each; fv-ml1 GPU 1 read 0%
before and after):