feat: Zed edit-predictions keyless FIM route (Qwen2.5-Coder-1.5B / coder-fast)
Deep-research-picked Qwen2.5-Coder-1.5B (BASE, Apache-2.0, native FIM) as a low-latency inline-completion seat: - stacks/vllm: vllm-coder service (ana-ml2 GPU1 :8020) + granite shrunk (util 0.27->0.13, max-len 131072->16384, seqs 1024->256; granite phasing out) to free GPU1 room. - stacks/litellm: coder-fast alias -> :8020 (mode: completion, /v1/completions). - stacks/zed-fim-proxy (NEW): keyless /v1/completions front door on ana-docker :4141 for Zed (which can't send an auth header) — POST + path + model allowlist, injects a coder-fast-scoped virtual key -> LiteLLM :4000. Anon /ping liveness. Verified keyless FIM end-to-end. Zed api_url = http://10.250.50.70:4141/v1, model coder-fast, prompt_format qwen. Source-IP allowlist off pending the Mac's observed source IP.
This commit is contained in:
@@ -98,7 +98,10 @@ GRANITE_SERVED_NAME=granite-4.1-8b
|
||||
# short parallel calls, so the 64K cap is ample.
|
||||
# 131072 — RESTORED 2026-07-16 (native max) for full-chapter summarization; GPU-1
|
||||
# freed by the image-bench evict + qwen36-vl gone, so the 128K ctx fits again.
|
||||
GRANITE_MAX_MODEL_LEN=131072
|
||||
# 16384 — SHRUNK 2026-07-27 (operator: granite is being phased out) to free GPU-1
|
||||
# room for vllm-coder (the Zed FIM seat). Full-chapter ctx dropped; util 0.13 holds
|
||||
# 16K cleanly (32768 @ util 0.12 crash-looped: KV est-max was only 19376 tokens).
|
||||
GRANITE_MAX_MODEL_LEN=16384
|
||||
# FP8 KV cache (native on Blackwell cc 12.0). At 50K ≈ ~4.2 GB (vs ~8.4 GB at fp16).
|
||||
GRANITE_KV_CACHE_DTYPE=fp8
|
||||
# util 0.35 (~33.6 GB) — tuned 2026-06-13 to leave ~3.5 GB free on GPU 1 alongside
|
||||
@@ -118,8 +121,24 @@ GRANITE_KV_CACHE_DTYPE=fp8
|
||||
# 0.27 — RE-GROWN 2026-07-16 (same session) after char-rp moved onto GPU-1: spend the
|
||||
# leftover room on full-chapter context (max-len 131072). KV 15.0 GiB = 196,560 tokens
|
||||
# = 1.50x @ 131072; GPU-1 lands ~6.7 GB headroom (char-rp 30 + granite 27 + selene 17 + trio).
|
||||
GRANITE_GPU_MEM_UTIL=0.27
|
||||
# Concurrency cap — set VERY HIGH 2026-07-16 (was vLLM default 128) so the KV pool is
|
||||
# the only bound. granite is the fleet fan-out summarizer/classifier (many concurrent
|
||||
# SHORT calls); default 128 capped below the KV bound (~192 @ 1K-tok). VRAM-neutral.
|
||||
GRANITE_MAX_NUM_SEQS=1024
|
||||
# 0.13 — SHRUNK 2026-07-27 (was 0.27) with the max-len drop; frees ~14 GB on GPU-1 for
|
||||
# vllm-coder. granite phasing out, so no longer worth the big KV pool. (0.12 was too
|
||||
# small for even 16K KV; 0.13 gives ~1.4x @ 16384.)
|
||||
GRANITE_GPU_MEM_UTIL=0.13
|
||||
# Concurrency cap. Was 1024 (fan-out summarizer); DROPPED to 256 on 2026-07-27 with the
|
||||
# phasing-out shrink (256 is ample for the reduced summarizer load; smaller sched state).
|
||||
GRANITE_MAX_NUM_SEQS=256
|
||||
|
||||
# Qwen2.5-Coder-1.5B (BASE) — FIM code-completion seat for Zed editor inline
|
||||
# edit-predictions (deep-research pick 2026-07-27, Apache-2.0). GPU-1, alongside the
|
||||
# small-model trio + (phasing-out) granite. Reached via LiteLLM `coder-fast` alias
|
||||
# and the keyless zed-fim-proxy (stacks/zed-fim-proxy). util 0.06 (~5.7 GB) holds
|
||||
# the 1.5B fp16 + fp8 KV; 8192 ctx is ample for FIM (KV = 13.75x concurrency).
|
||||
CODER_PORT=8020
|
||||
CODER_GPU_ID=1
|
||||
CODER_MODEL=Qwen/Qwen2.5-Coder-1.5B
|
||||
CODER_SERVED_NAME=qwen2.5-coder-1.5b
|
||||
CODER_MAX_MODEL_LEN=8192
|
||||
CODER_KV_CACHE_DTYPE=fp8
|
||||
CODER_GPU_MEM_UTIL=0.06
|
||||
CODER_MAX_NUM_SEQS=32
|
||||
|
||||
Reference in New Issue
Block a user