revert(coder-seat): keep Qwen2.5-Coder on fv-ml1; nh3-ml1 copy removed (5x slower, same quality)
This commit is contained in:
@@ -17,11 +17,14 @@ Blackwells). It was chosen for Zed edit-predictions by a deep-research pass on
|
||||
| revision | `df3ce67c0e24480f20468b6ef2894622d69eb73b` (pinned; same as fv-ml1 and HF main on 2026-09-25) |
|
||||
| memory | 0.33 of 16,380 MiB ≈ 5.4 GB, the same absolute budget as on fv-ml1 |
|
||||
|
||||
## ⏸ Status 2026-09-25 2355: running, NOT behind the gateway (Prime deciding)
|
||||
## ❌ NOT DEPLOYED. Decision (Prime, 2026-09-26 0000): the seat stays on fv-ml1
|
||||
|
||||
The copy is up and parity-checked, but the gateway still points `coder-fast` at
|
||||
fv-ml1. On the 50 W Ada the same model is **about 5× slower**, which Zed
|
||||
edit-predictions would notice.
|
||||
The nh3-ml1 copy was built, checked for parity, measured, and then removed (container,
|
||||
image, model cache). It was the same quality but **about 5× slower** on the 50 W
|
||||
Ada. That lag would reach Zed edit-predictions, and moving it would only have freed
|
||||
6.3 GB on an fv-ml1 GPU that already had about 20 GB spare. The live seat is still
|
||||
`stacks/vllm` on fv-ml1. This stack is kept as the tested recipe in case the
|
||||
calculation changes (a faster card, or speculative decoding).
|
||||
|
||||
**Speed, measured on each box** (3 interleaved reps each; fv-ml1 GPU 1 read 0%
|
||||
before and after):
|
||||
|
||||
Reference in New Issue
Block a user