feat: Zed edit-predictions keyless FIM route (Qwen2.5-Coder-1.5B / coder-fast)
Deep-research-picked Qwen2.5-Coder-1.5B (BASE, Apache-2.0, native FIM) as a low-latency inline-completion seat: - stacks/vllm: vllm-coder service (ana-ml2 GPU1 :8020) + granite shrunk (util 0.27->0.13, max-len 131072->16384, seqs 1024->256; granite phasing out) to free GPU1 room. - stacks/litellm: coder-fast alias -> :8020 (mode: completion, /v1/completions). - stacks/zed-fim-proxy (NEW): keyless /v1/completions front door on ana-docker :4141 for Zed (which can't send an auth header) — POST + path + model allowlist, injects a coder-fast-scoped virtual key -> LiteLLM :4000. Anon /ping liveness. Verified keyless FIM end-to-end. Zed api_url = http://10.250.50.70:4141/v1, model coder-fast, prompt_format qwen. Source-IP allowlist off pending the Mac's observed source IP.
This commit is contained in:
@@ -128,6 +128,8 @@ _As of 2026-07-25 — **DONE: infra-ops Worldtree config-as-code repo BUILT + PU
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-07-27]` **Zed edit-predictions: keyless FIM-completion route SHIPPED end-to-end.** Operator wants Zed's inline edit-prediction (which CANNOT send an auth header) to reach a FIM coder via `/v1/completions`. **Deep-research (106-agent workflow) picked `Qwen/Qwen2.5-Coder-1.5B`** (BASE, Apache-2.0; native FIM `<|fim_prefix|>/<|fim_suffix|>/<|fim_middle|>` IDs 151659/60/61; Zed `prompt_format:"qwen"`). Runner-up 3B = non-commercial Qwen-Research license; **no small dense Qwen3-Coder exists (all MoE, smallest 30B)**. **Stood up `vllm-coder`** on ana-ml2 **GPU1 :8020** (served-name `qwen2.5-coder-1.5b`, 8192 ctx, util 0.06, fp8 KV). To fit, **shrank granite (phasing out, operator-directed):** util 0.27→0.13, max-len 131072→16384, seqs 1024→256 (freed ~14 GB; the KV-≥-1×-max-len rule crash-looped it at util 0.12/32768 → settled 0.13/16384). **LiteLLM alias `coder-fast`** → :8020 (`mode: completion`). **Minted a `coder-fast`-SCOPED virtual key** (verified 403 on `gen` — the real blast-radius bound). **Built `zed-fim-proxy`** (ana-docker **:4141**, `network_mode: host`, stdlib-python, `stacks/zed-fim-proxy`): keyless POST `/v1/completions`, model-allowlist `coder-fast`, injects the scoped key → LiteLLM :4000; `GET /ping` anon liveness; wrong-model→403, wrong-path→404, `/chat/completions` rejected. Verified keyless FIM end-to-end ('a + b', finish `stop`). **Zed `api_url` = `http://10.250.50.70:4141/v1`, model `coder-fast`, prompt_format `qwen`.** ⚠ **source-IP allowlist is OFF** (`ZED_ALLOWED_IPS` empty) — verify the Mac's real source IP once it connects (`docker logs zed-fim-proxy`), then tighten (site-to-site NAT may mask `10.0.10.83`). Canonical: `stacks/vllm` (coder + granite shrink), `stacks/litellm` (coder-fast), `stacks/zed-fim-proxy` (NEW). Server vllm compose.yaml has benign stale-comment drift vs canonical (didn't overwrite the newer canonical).
|
||||
|
||||
- `[2026-07-27]` **Muninn ingestion-watcher sidecar deployed on PERSONAL Worldtree (#377).** worldtree-dev request (research-wing ingest arc, personal-only per the 2026-07-16 topology ruling); operator-approved. Added a `worldtree-muninn` **compose sidecar** to `/opt/worldtree-personal/compose.yaml` — `<<: *worldtree-common` anchor inherits the api's image + full env + config/state/kb mounts; `command: python -m core.muninn --watch`; `restart: unless-stopped`; `stop_grace_period: 1h` (INV-377-7: max 2 concurrent × worst-case job, SIGTERM-drains). **Pinned to the running SHA `773866084af9`** (b146, ≥ b143 — dodges both the `:latest` trap AND the "pre-b143 ref resurrects deleted dispatch.py from stale bytecode" warning). Verified: running / 0 restarts / flock sole-runner (no rc3) / heartbeat live at `{ingestion_root=/data/state/ingestion}/.watcher-heartbeat` (poll 30s). Container `worldtree-personal-worldtree-muninn-1`; backup `compose.yaml.bak-muninn-20260727-081920`. **DURABILITY RESOLVED (worldtree-dev, same day):** Q1 was a LIVE FOOTGUN — `deploy-personal.yml` scp's the REPO compose.yaml over the box's + runs `up -d --remove-orphans`, so the box-local sidecar would've been clobbered AND orphan-removed at the next staging tag. worldtree-dev fixed at source: moved the sidecar into their repo compose.yaml gated behind a **`muninn` compose profile** (commit 5d7f6bd) — shared compose stays instance-identical, `.env` `COMPOSE_PROFILES` differentiates (demo watcher-less). **My action:** added `COMPOSE_PROFILES=muninn` to `/opt/worldtree-personal/.env` (backup `.bak-muninn-profile-20260727-082541`; no-op vs the current unprofiled box-local sidecar → seamless handover at next deploy). Q2: their deploy `up -d`'s the whole stack w/ `WORLDTREE_IMAGE` exported → sidecar version-tracks the api, no drift. **CONFIG-AS-CODE EXTENSION:** mirrored the non-secret delta as `personal/env.public` in `vh/worldtree-instance-configs` (repo `a9d091e`) — FIRST extension beyond config.yaml files to env-level config; the secret-laden `.env` stays box-only, `env.public` records only non-secret infra-ops-owned env deltas (record, not a deploy source — `deploy-wt-config` globs `*.yaml`). **BOUNDARY CLARIFIED:** compose.yaml = worldtree-dev's (their repo, instance-identical, scp'd on deploy); per-instance `.env` = infra-ops's differentiator. Deploy step of the #363/#377 arc. **#377 CLOSED — acceptance PASSED 2026-07-27:** worldtree-dev enqueued a test job via muninn-dispatch 0.1.0 in a one-shot ephemeral container (no docker-exec); the sidecar claimed it within one 30s poll, drove it to terminal (structure→summarize→complete), zero restarts/rc3, heartbeat fresh throughout — whole loop (request→deploy→durability fix→acceptance) in <2h. (Pre-existing pipeline bug #379 surfaced — `output.kb_notes=false` ignored → 1 inert test note in the research wing — worldtree-dev owns it, nothing infra-ops-side.) **⚠ OPERATOR-SURFACE (open):** the `env.public` overlay mechanism is a repo-scope call to bless/adjust. [[reference_worldtree_deploys_cicd]] [[reference_worldtree_instance_configs_repo]] [[project_worldtree_research_wing_ingest]]
|
||||
|
||||
- `[2026-07-27]` **jackdaw-compose.service DECOMMISSIONED** (jackdaw-dev request; the JackDAW AI Composer was cut from v1 by operator decision 2026-07-27). Stopped + disabled the nh3-dev `:8787` user service (no client calls it — ai/server/AiChat deleted from main, `/compose` proxy removed); unit **archived not deleted** → `~/.config/systemd/user/jackdaw-compose.service.decommissioned-20260727` (revival = rename + `daemon-reload`). **No credential revoked** — the unit used the SHARED all-agents LiteLLM key (`sk-eA_XOd…`, model `gen`), not a dedicated one. Code preserved on jackdaw `origin/ai-composer-preserved`; treat as permanent. The `:4500` HTTPS audition bench is untouched. (Supersedes the 2026-07-23 stand-up line below.)
|
||||
|
||||
@@ -257,6 +257,24 @@ model_list:
|
||||
model_info:
|
||||
mode: rerank
|
||||
|
||||
# --- coder-fast → Qwen2.5-Coder-1.5B (BASE), FIM code-completion seat (ana-ml2
|
||||
# GPU1 :8020, vLLM; deep-research pick 2026-07-27). For Zed editor inline
|
||||
# edit-predictions via the LEGACY /v1/completions endpoint with Qwen FIM
|
||||
# markers (<|fim_prefix|>/<|fim_suffix|>/<|fim_middle|>). BASE not -Instruct
|
||||
# (FIM is a pretraining objective; base completions are cleaner). Apache-2.0.
|
||||
# mode: completion — this is text-completion, not chat. Reached KEYLESS from
|
||||
# Vuong's Mac via the zed-fim-proxy (separate port on ana-docker) which injects
|
||||
# a coder-fast-scoped virtual key; the proxy's model-allowlist + the scoped key
|
||||
# bound the blast radius. Runner-up was Qwen2.5-Coder-3B (higher HumanEval-FIM,
|
||||
# non-commercial Qwen-Research license). ---
|
||||
- model_name: coder-fast
|
||||
litellm_params:
|
||||
model: hosted_vllm/qwen2.5-coder-1.5b
|
||||
api_base: http://10.250.50.54:8020/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
model_info:
|
||||
mode: completion
|
||||
|
||||
# --- z.ai GLM (cloud API) — fronted for unified logging across local
|
||||
# + cloud inference. Explicit entries, so they win over the "*"
|
||||
# wildcard below (no collision with llama-swap's glm4.7-flash etc.
|
||||
|
||||
@@ -98,7 +98,10 @@ GRANITE_SERVED_NAME=granite-4.1-8b
|
||||
# short parallel calls, so the 64K cap is ample.
|
||||
# 131072 — RESTORED 2026-07-16 (native max) for full-chapter summarization; GPU-1
|
||||
# freed by the image-bench evict + qwen36-vl gone, so the 128K ctx fits again.
|
||||
GRANITE_MAX_MODEL_LEN=131072
|
||||
# 16384 — SHRUNK 2026-07-27 (operator: granite is being phased out) to free GPU-1
|
||||
# room for vllm-coder (the Zed FIM seat). Full-chapter ctx dropped; util 0.13 holds
|
||||
# 16K cleanly (32768 @ util 0.12 crash-looped: KV est-max was only 19376 tokens).
|
||||
GRANITE_MAX_MODEL_LEN=16384
|
||||
# FP8 KV cache (native on Blackwell cc 12.0). At 50K ≈ ~4.2 GB (vs ~8.4 GB at fp16).
|
||||
GRANITE_KV_CACHE_DTYPE=fp8
|
||||
# util 0.35 (~33.6 GB) — tuned 2026-06-13 to leave ~3.5 GB free on GPU 1 alongside
|
||||
@@ -118,8 +121,24 @@ GRANITE_KV_CACHE_DTYPE=fp8
|
||||
# 0.27 — RE-GROWN 2026-07-16 (same session) after char-rp moved onto GPU-1: spend the
|
||||
# leftover room on full-chapter context (max-len 131072). KV 15.0 GiB = 196,560 tokens
|
||||
# = 1.50x @ 131072; GPU-1 lands ~6.7 GB headroom (char-rp 30 + granite 27 + selene 17 + trio).
|
||||
GRANITE_GPU_MEM_UTIL=0.27
|
||||
# Concurrency cap — set VERY HIGH 2026-07-16 (was vLLM default 128) so the KV pool is
|
||||
# the only bound. granite is the fleet fan-out summarizer/classifier (many concurrent
|
||||
# SHORT calls); default 128 capped below the KV bound (~192 @ 1K-tok). VRAM-neutral.
|
||||
GRANITE_MAX_NUM_SEQS=1024
|
||||
# 0.13 — SHRUNK 2026-07-27 (was 0.27) with the max-len drop; frees ~14 GB on GPU-1 for
|
||||
# vllm-coder. granite phasing out, so no longer worth the big KV pool. (0.12 was too
|
||||
# small for even 16K KV; 0.13 gives ~1.4x @ 16384.)
|
||||
GRANITE_GPU_MEM_UTIL=0.13
|
||||
# Concurrency cap. Was 1024 (fan-out summarizer); DROPPED to 256 on 2026-07-27 with the
|
||||
# phasing-out shrink (256 is ample for the reduced summarizer load; smaller sched state).
|
||||
GRANITE_MAX_NUM_SEQS=256
|
||||
|
||||
# Qwen2.5-Coder-1.5B (BASE) — FIM code-completion seat for Zed editor inline
|
||||
# edit-predictions (deep-research pick 2026-07-27, Apache-2.0). GPU-1, alongside the
|
||||
# small-model trio + (phasing-out) granite. Reached via LiteLLM `coder-fast` alias
|
||||
# and the keyless zed-fim-proxy (stacks/zed-fim-proxy). util 0.06 (~5.7 GB) holds
|
||||
# the 1.5B fp16 + fp8 KV; 8192 ctx is ample for FIM (KV = 13.75x concurrency).
|
||||
CODER_PORT=8020
|
||||
CODER_GPU_ID=1
|
||||
CODER_MODEL=Qwen/Qwen2.5-Coder-1.5B
|
||||
CODER_SERVED_NAME=qwen2.5-coder-1.5b
|
||||
CODER_MAX_MODEL_LEN=8192
|
||||
CODER_KV_CACHE_DTYPE=fp8
|
||||
CODER_GPU_MEM_UTIL=0.06
|
||||
CODER_MAX_NUM_SEQS=32
|
||||
|
||||
@@ -278,6 +278,70 @@ services:
|
||||
- homepage.description=Granite 4.1 8B FP8 via vLLM (ana-ml2)
|
||||
- homepage.href=http://10.250.50.54:${GRANITE_PORT}/docs
|
||||
|
||||
vllm-coder:
|
||||
image: vllm/vllm-openai:${VLLM_VERSION}
|
||||
container_name: vllm-coder
|
||||
restart: unless-stopped
|
||||
ipc: host
|
||||
ports:
|
||||
- "${CODER_PORT}:8000"
|
||||
volumes:
|
||||
- /tank/aimodels/huggingface:/hfcache
|
||||
environment:
|
||||
- HF_HOME=/hfcache
|
||||
- HF_HUB_CACHE=/hfcache/hub
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_API_KEY=${API_KEY:-}
|
||||
command:
|
||||
# Qwen2.5-Coder-1.5B (BASE) — FIM code-completion seat for Zed edit-predictions
|
||||
# (deep-research pick 2026-07-27). Native fill-in-the-middle: <|fim_prefix|> /
|
||||
# <|fim_suffix|> / <|fim_middle|> (IDs 151659/151660/151661); Zed sends the
|
||||
# FIM-formatted prompt to /v1/completions and vLLM passes it through (the FIM
|
||||
# special tokens live in the tokenizer). BASE not -Instruct (FIM is a
|
||||
# pretraining objective; base completions are cleaner). Apache-2.0. Runner-up =
|
||||
# Qwen2.5-Coder-3B (higher HumanEval-FIM but non-commercial Qwen-Research license).
|
||||
- ${CODER_MODEL}
|
||||
- --served-model-name
|
||||
- ${CODER_SERVED_NAME}
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- "8000"
|
||||
- --gpu-memory-utilization
|
||||
- ${CODER_GPU_MEM_UTIL}
|
||||
- --max-model-len
|
||||
- ${CODER_MAX_MODEL_LEN}
|
||||
- --max-num-seqs
|
||||
- ${CODER_MAX_NUM_SEQS}
|
||||
- --dtype
|
||||
- auto
|
||||
- --kv-cache-dtype
|
||||
- ${CODER_KV_CACHE_DTYPE}
|
||||
- --enable-prefix-caching
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids:
|
||||
- "${CODER_GPU_ID}"
|
||||
capabilities:
|
||||
- gpu
|
||||
healthcheck:
|
||||
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
start_period: 300s
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI - Inference
|
||||
- homepage.name=vLLM Qwen2.5-Coder 1.5B (FIM)
|
||||
- homepage.icon=mdi-code-braces
|
||||
- homepage.description=Qwen2.5-Coder-1.5B FIM code-completion (ana-ml2, Zed edit-predictions)
|
||||
- homepage.href=http://10.250.50.54:${CODER_PORT}/docs
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# zed-fim-proxy tunables — copy to `.env` on ana-docker (the real .env holds the
|
||||
# scoped key and is server-only / gitignored). See conf/proxy.py + README.md.
|
||||
|
||||
# Port the keyless route listens on (host network).
|
||||
ZED_PORT=4141
|
||||
|
||||
# LiteLLM gateway to forward to (localhost:4000 via network_mode: host).
|
||||
ZED_UPSTREAM=http://localhost:4000
|
||||
|
||||
# The ONLY model this route will forward (proxy rejects any other "model").
|
||||
ZED_ALLOWED_MODEL=coder-fast
|
||||
|
||||
# Best-effort source-IP allowlist (comma-separated). Empty = allow all — fine on the
|
||||
# internal network, but TIGHTEN to the Mac's observed source IP once it connects
|
||||
# (watch `docker logs zed-fim-proxy` for the real source). Only meaningful if the
|
||||
# proxy sees the real client IP (network_mode: host); a site-to-site NAT may mask it.
|
||||
ZED_ALLOWED_IPS=
|
||||
|
||||
# A LiteLLM virtual key SCOPED TO ZED_ALLOWED_MODEL ONLY (the real blast-radius
|
||||
# bound). Mint: POST /key/generate {"models":["coder-fast"]}. NEVER commit the value.
|
||||
ZED_SCOPED_KEY=
|
||||
@@ -0,0 +1,47 @@
|
||||
# zed-fim-proxy
|
||||
|
||||
A **keyless `/v1/completions` front door** on ana-docker for Zed's editor
|
||||
edit-prediction (inline FIM completion) feature, which **cannot send an
|
||||
`Authorization` header**. Runs on a separate port from LiteLLM and forwards to it
|
||||
with an injected, model-scoped key.
|
||||
|
||||
- **Host:** ana-docker `10.250.50.70`, port **4141** (`network_mode: host`).
|
||||
- **Backs:** the `coder-fast` model (Qwen2.5-Coder-1.5B FIM seat, `stacks/vllm`
|
||||
`vllm-coder` on ana-ml2:8020) via LiteLLM `:4000`.
|
||||
- **Zed config** (`edit_predictions.open_ai_compatible_api`): `api_url:
|
||||
http://10.250.50.70:4141/v1`, `model: coder-fast`, `prompt_format: qwen`,
|
||||
`max_output_tokens: <n>`. (The proxy also accepts `http://10.250.50.70:4141` —
|
||||
it matches both `/v1/completions` and `/completions`.)
|
||||
|
||||
## Security model (stdlib proxy in `conf/proxy.py`)
|
||||
|
||||
Four guards + a scoped key — a keyless route that injects a working key is only
|
||||
safe if it can't be pivoted:
|
||||
|
||||
1. **POST + path** `/v1/completions` (or `/completions`) only. `GET /ping` is an
|
||||
anonymous liveness (`{"service":"ok"}`). `/v1/chat/completions` is rejected.
|
||||
2. **Model allowlist** — request body `model` must equal `ZED_ALLOWED_MODEL`
|
||||
(`coder-fast`); anything else → 403.
|
||||
3. **Injected scoped key** — a LiteLLM virtual key scoped to `coder-fast` ONLY
|
||||
(`POST /key/generate {"models":["coder-fast"]}`). Even if guards 1–2 were
|
||||
bypassed, the key reaches nothing else (verified: 403 on `gen`). **This is the
|
||||
real blast-radius bound.**
|
||||
4. **Best-effort source-IP allowlist** (`ZED_ALLOWED_IPS`) — only enforceable if
|
||||
the proxy sees the real client IP (hence `network_mode: host`; docker
|
||||
port-publish would NAT it away). A site-to-site NAT may still mask the Mac's
|
||||
`10.0.10.83` — verify against `docker logs zed-fim-proxy` and tighten. Internal
|
||||
network only; no public exposure.
|
||||
|
||||
## Deploy
|
||||
|
||||
```
|
||||
# conf/proxy.py -> /opt/docker/conf/zed-fim-proxy/proxy.py
|
||||
# compose.yaml -> /opt/docker/compose/zed-fim-proxy/compose.yaml
|
||||
# .env (from .env.example, with ZED_SCOPED_KEY filled) -> same dir, mode 600
|
||||
cd /opt/docker/compose/zed-fim-proxy && docker compose up -d
|
||||
# verify keyless:
|
||||
curl -s http://localhost:4141/v1/completions -H 'Content-Type: application/json' \
|
||||
-d '{"model":"coder-fast","prompt":"def add(a,b):\n return","max_tokens":16,"temperature":0.2}'
|
||||
```
|
||||
|
||||
Stdlib-only proxy (no pip) in a bare `python:3.12-slim` container — no build.
|
||||
@@ -0,0 +1,36 @@
|
||||
name: zed-fim-proxy
|
||||
|
||||
# Keyless /v1/completions front door for Zed edit-predictions (coder-fast only).
|
||||
# See conf/proxy.py for the guard model. Runs network_mode: host so it (a) sees the
|
||||
# real client source IP for the best-effort allowlist (docker port-publish would NAT
|
||||
# it away) and (b) reaches the LiteLLM gateway on localhost:4000. Stdlib-only proxy
|
||||
# in a bare python image — no build, no pip.
|
||||
|
||||
services:
|
||||
zed-fim-proxy:
|
||||
image: python:3.12-slim
|
||||
container_name: zed-fim-proxy
|
||||
restart: unless-stopped
|
||||
network_mode: host
|
||||
command: ["python", "/app/proxy.py"]
|
||||
volumes:
|
||||
- /opt/docker/conf/zed-fim-proxy/proxy.py:/app/proxy.py:ro
|
||||
environment:
|
||||
- ZED_PORT=${ZED_PORT:-4141}
|
||||
- ZED_UPSTREAM=${ZED_UPSTREAM:-http://localhost:4000}
|
||||
- ZED_ALLOWED_MODEL=${ZED_ALLOWED_MODEL:-coder-fast}
|
||||
# best-effort source-IP allowlist (comma-separated); empty = allow all.
|
||||
- ZED_ALLOWED_IPS=${ZED_ALLOWED_IPS:-}
|
||||
- ZED_SCOPED_KEY=${ZED_SCOPED_KEY}
|
||||
healthcheck:
|
||||
test: ["CMD", "python", "-c", "import urllib.request,os; urllib.request.urlopen('http://localhost:%s/ping' % os.environ.get('ZED_PORT','4141'))"]
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
retries: 3
|
||||
start_period: 10s
|
||||
labels:
|
||||
- homepage.group=AI - Gateways & Chat
|
||||
- homepage.name=Zed FIM Proxy
|
||||
- homepage.icon=mdi-code-braces-box
|
||||
- homepage.description=Keyless /v1/completions for Zed edit-predictions (coder-fast, ana-docker)
|
||||
- homepage.href=http://10.250.50.70:4141/ping
|
||||
@@ -0,0 +1,102 @@
|
||||
"""zed-fim-proxy — a keyless front door for Zed's edit-prediction feature.
|
||||
|
||||
Zed's edit_predictions.open_ai_compatible_api provider cannot send an
|
||||
Authorization header, so it needs a route where POST /v1/completions succeeds
|
||||
with no key. This proxy is that route, on a separate port from LiteLLM, with
|
||||
four narrow guards + an injected model-scoped key:
|
||||
|
||||
1. POST only, path exactly /v1/completions (GET /ping is an anonymous liveness).
|
||||
2. source-IP allowlist (best-effort — set ZED_ALLOWED_IPS; empty = allow all).
|
||||
NOTE: only meaningful if this process sees the real client IP — run the
|
||||
container with network_mode: host (docker port-publish would NAT the source
|
||||
to the bridge gateway and defeat it). Behind a site-to-site NAT it may still
|
||||
see a mesh IP, not the Mac — verify against the access log at deploy.
|
||||
3. request body "model" must equal ZED_ALLOWED_MODEL (default coder-fast).
|
||||
4. injects Authorization: Bearer $ZED_SCOPED_KEY (a LiteLLM virtual key scoped
|
||||
to that ONE model) and forwards to $ZED_UPSTREAM. The scoped key is the real
|
||||
blast-radius bound: even if guards 1-3 were bypassed, the key can reach
|
||||
nothing but coder-fast.
|
||||
|
||||
Stdlib only (no pip) — runs in a bare python:slim container.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
UPSTREAM = os.environ.get("ZED_UPSTREAM", "http://localhost:4000").rstrip("/")
|
||||
SCOPED_KEY = os.environ["ZED_SCOPED_KEY"]
|
||||
ALLOWED_MODEL = os.environ.get("ZED_ALLOWED_MODEL", "coder-fast")
|
||||
ALLOWED_IPS = set(x.strip() for x in os.environ.get("ZED_ALLOWED_IPS", "").split(",") if x.strip())
|
||||
PORT = int(os.environ.get("ZED_PORT", "4141"))
|
||||
TIMEOUT = float(os.environ.get("ZED_TIMEOUT", "60"))
|
||||
|
||||
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
server_version = "zed-fim-proxy/1.0"
|
||||
|
||||
def _json(self, code, obj):
|
||||
body = json.dumps(obj).encode()
|
||||
self.send_response(code)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(body)))
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
|
||||
def do_GET(self):
|
||||
if self.path.rstrip("/") == "/ping":
|
||||
return self._json(200, {"service": "ok"})
|
||||
return self._json(404, {"error": "not found"})
|
||||
|
||||
def do_POST(self):
|
||||
src = self.client_address[0]
|
||||
if ALLOWED_IPS and src not in ALLOWED_IPS:
|
||||
return self._json(403, {"error": f"source {src} not allowed"})
|
||||
# Accept both /v1/completions and /completions so the Zed api_url can be
|
||||
# set to either http://host:4141/v1 or http://host:4141 (Zed appends
|
||||
# /completions). /v1/chat/completions is NOT in the set → stays rejected.
|
||||
if self.path.split("?")[0].rstrip("/") not in ("/v1/completions", "/completions"):
|
||||
return self._json(404, {"error": "only POST /v1/completions or /completions"})
|
||||
try:
|
||||
length = int(self.headers.get("Content-Length", 0))
|
||||
except ValueError:
|
||||
return self._json(400, {"error": "bad content-length"})
|
||||
raw = self.rfile.read(length)
|
||||
try:
|
||||
body = json.loads(raw)
|
||||
except Exception:
|
||||
return self._json(400, {"error": "invalid json body"})
|
||||
if body.get("model") != ALLOWED_MODEL:
|
||||
return self._json(403, {"error": f"model must be '{ALLOWED_MODEL}'"})
|
||||
req = urllib.request.Request(
|
||||
UPSTREAM + "/v1/completions",
|
||||
data=raw,
|
||||
method="POST",
|
||||
headers={
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {SCOPED_KEY}",
|
||||
},
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=TIMEOUT) as r:
|
||||
data, code, ctype = r.read(), r.status, r.headers.get("Content-Type", "application/json")
|
||||
except urllib.error.HTTPError as e:
|
||||
data, code, ctype = e.read(), e.code, e.headers.get("Content-Type", "application/json")
|
||||
except Exception as e: # noqa: BLE001
|
||||
return self._json(502, {"error": f"upstream error: {e}"})
|
||||
self.send_response(code)
|
||||
self.send_header("Content-Type", ctype)
|
||||
self.send_header("Content-Length", str(len(data)))
|
||||
self.end_headers()
|
||||
self.wfile.write(data)
|
||||
|
||||
def log_message(self, fmt, *args):
|
||||
sys.stderr.write("%s - %s\n" % (self.client_address[0], fmt % args))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(f"zed-fim-proxy on :{PORT} -> {UPSTREAM} (model={ALLOWED_MODEL}, "
|
||||
f"ip-allowlist={'set' if ALLOWED_IPS else 'OFF'})", flush=True)
|
||||
ThreadingHTTPServer(("0.0.0.0", PORT), Handler).serve_forever()
|
||||
Reference in New Issue
Block a user