diff --git a/persistent-memory.md b/persistent-memory.md index 7b8c908..c6c693b 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -128,6 +128,8 @@ _As of 2026-07-25 — **DONE: infra-ops Worldtree config-as-code repo BUILT + PU ## Recent decisions +- `[2026-07-27]` **Zed edit-predictions: keyless FIM-completion route SHIPPED end-to-end.** Operator wants Zed's inline edit-prediction (which CANNOT send an auth header) to reach a FIM coder via `/v1/completions`. **Deep-research (106-agent workflow) picked `Qwen/Qwen2.5-Coder-1.5B`** (BASE, Apache-2.0; native FIM `<|fim_prefix|>/<|fim_suffix|>/<|fim_middle|>` IDs 151659/60/61; Zed `prompt_format:"qwen"`). Runner-up 3B = non-commercial Qwen-Research license; **no small dense Qwen3-Coder exists (all MoE, smallest 30B)**. **Stood up `vllm-coder`** on ana-ml2 **GPU1 :8020** (served-name `qwen2.5-coder-1.5b`, 8192 ctx, util 0.06, fp8 KV). To fit, **shrank granite (phasing out, operator-directed):** util 0.27→0.13, max-len 131072→16384, seqs 1024→256 (freed ~14 GB; the KV-≥-1×-max-len rule crash-looped it at util 0.12/32768 → settled 0.13/16384). **LiteLLM alias `coder-fast`** → :8020 (`mode: completion`). **Minted a `coder-fast`-SCOPED virtual key** (verified 403 on `gen` — the real blast-radius bound). **Built `zed-fim-proxy`** (ana-docker **:4141**, `network_mode: host`, stdlib-python, `stacks/zed-fim-proxy`): keyless POST `/v1/completions`, model-allowlist `coder-fast`, injects the scoped key → LiteLLM :4000; `GET /ping` anon liveness; wrong-model→403, wrong-path→404, `/chat/completions` rejected. Verified keyless FIM end-to-end ('a + b', finish `stop`). **Zed `api_url` = `http://10.250.50.70:4141/v1`, model `coder-fast`, prompt_format `qwen`.** ⚠ **source-IP allowlist is OFF** (`ZED_ALLOWED_IPS` empty) — verify the Mac's real source IP once it connects (`docker logs zed-fim-proxy`), then tighten (site-to-site NAT may mask `10.0.10.83`). Canonical: `stacks/vllm` (coder + granite shrink), `stacks/litellm` (coder-fast), `stacks/zed-fim-proxy` (NEW). Server vllm compose.yaml has benign stale-comment drift vs canonical (didn't overwrite the newer canonical). + - `[2026-07-27]` **Muninn ingestion-watcher sidecar deployed on PERSONAL Worldtree (#377).** worldtree-dev request (research-wing ingest arc, personal-only per the 2026-07-16 topology ruling); operator-approved. Added a `worldtree-muninn` **compose sidecar** to `/opt/worldtree-personal/compose.yaml` — `<<: *worldtree-common` anchor inherits the api's image + full env + config/state/kb mounts; `command: python -m core.muninn --watch`; `restart: unless-stopped`; `stop_grace_period: 1h` (INV-377-7: max 2 concurrent × worst-case job, SIGTERM-drains). **Pinned to the running SHA `773866084af9`** (b146, ≥ b143 — dodges both the `:latest` trap AND the "pre-b143 ref resurrects deleted dispatch.py from stale bytecode" warning). Verified: running / 0 restarts / flock sole-runner (no rc3) / heartbeat live at `{ingestion_root=/data/state/ingestion}/.watcher-heartbeat` (poll 30s). Container `worldtree-personal-worldtree-muninn-1`; backup `compose.yaml.bak-muninn-20260727-081920`. **DURABILITY RESOLVED (worldtree-dev, same day):** Q1 was a LIVE FOOTGUN — `deploy-personal.yml` scp's the REPO compose.yaml over the box's + runs `up -d --remove-orphans`, so the box-local sidecar would've been clobbered AND orphan-removed at the next staging tag. worldtree-dev fixed at source: moved the sidecar into their repo compose.yaml gated behind a **`muninn` compose profile** (commit 5d7f6bd) — shared compose stays instance-identical, `.env` `COMPOSE_PROFILES` differentiates (demo watcher-less). **My action:** added `COMPOSE_PROFILES=muninn` to `/opt/worldtree-personal/.env` (backup `.bak-muninn-profile-20260727-082541`; no-op vs the current unprofiled box-local sidecar → seamless handover at next deploy). Q2: their deploy `up -d`'s the whole stack w/ `WORLDTREE_IMAGE` exported → sidecar version-tracks the api, no drift. **CONFIG-AS-CODE EXTENSION:** mirrored the non-secret delta as `personal/env.public` in `vh/worldtree-instance-configs` (repo `a9d091e`) — FIRST extension beyond config.yaml files to env-level config; the secret-laden `.env` stays box-only, `env.public` records only non-secret infra-ops-owned env deltas (record, not a deploy source — `deploy-wt-config` globs `*.yaml`). **BOUNDARY CLARIFIED:** compose.yaml = worldtree-dev's (their repo, instance-identical, scp'd on deploy); per-instance `.env` = infra-ops's differentiator. Deploy step of the #363/#377 arc. **#377 CLOSED — acceptance PASSED 2026-07-27:** worldtree-dev enqueued a test job via muninn-dispatch 0.1.0 in a one-shot ephemeral container (no docker-exec); the sidecar claimed it within one 30s poll, drove it to terminal (structure→summarize→complete), zero restarts/rc3, heartbeat fresh throughout — whole loop (request→deploy→durability fix→acceptance) in <2h. (Pre-existing pipeline bug #379 surfaced — `output.kb_notes=false` ignored → 1 inert test note in the research wing — worldtree-dev owns it, nothing infra-ops-side.) **⚠ OPERATOR-SURFACE (open):** the `env.public` overlay mechanism is a repo-scope call to bless/adjust. [[reference_worldtree_deploys_cicd]] [[reference_worldtree_instance_configs_repo]] [[project_worldtree_research_wing_ingest]] - `[2026-07-27]` **jackdaw-compose.service DECOMMISSIONED** (jackdaw-dev request; the JackDAW AI Composer was cut from v1 by operator decision 2026-07-27). Stopped + disabled the nh3-dev `:8787` user service (no client calls it — ai/server/AiChat deleted from main, `/compose` proxy removed); unit **archived not deleted** → `~/.config/systemd/user/jackdaw-compose.service.decommissioned-20260727` (revival = rename + `daemon-reload`). **No credential revoked** — the unit used the SHARED all-agents LiteLLM key (`sk-eA_XOd…`, model `gen`), not a dedicated one. Code preserved on jackdaw `origin/ai-composer-preserved`; treat as permanent. The `:4500` HTTPS audition bench is untouched. (Supersedes the 2026-07-23 stand-up line below.) diff --git a/stacks/litellm/conf/config.yaml b/stacks/litellm/conf/config.yaml index 0751b48..5c64ce9 100644 --- a/stacks/litellm/conf/config.yaml +++ b/stacks/litellm/conf/config.yaml @@ -257,6 +257,24 @@ model_list: model_info: mode: rerank + # --- coder-fast → Qwen2.5-Coder-1.5B (BASE), FIM code-completion seat (ana-ml2 + # GPU1 :8020, vLLM; deep-research pick 2026-07-27). For Zed editor inline + # edit-predictions via the LEGACY /v1/completions endpoint with Qwen FIM + # markers (<|fim_prefix|>/<|fim_suffix|>/<|fim_middle|>). BASE not -Instruct + # (FIM is a pretraining objective; base completions are cleaner). Apache-2.0. + # mode: completion — this is text-completion, not chat. Reached KEYLESS from + # Vuong's Mac via the zed-fim-proxy (separate port on ana-docker) which injects + # a coder-fast-scoped virtual key; the proxy's model-allowlist + the scoped key + # bound the blast radius. Runner-up was Qwen2.5-Coder-3B (higher HumanEval-FIM, + # non-commercial Qwen-Research license). --- + - model_name: coder-fast + litellm_params: + model: hosted_vllm/qwen2.5-coder-1.5b + api_base: http://10.250.50.54:8020/v1 + api_key: os.environ/VLLM_API_KEY + model_info: + mode: completion + # --- z.ai GLM (cloud API) — fronted for unified logging across local # + cloud inference. Explicit entries, so they win over the "*" # wildcard below (no collision with llama-swap's glm4.7-flash etc. diff --git a/stacks/vllm/.env.example b/stacks/vllm/.env.example index f72a5ef..b36ad07 100644 --- a/stacks/vllm/.env.example +++ b/stacks/vllm/.env.example @@ -98,7 +98,10 @@ GRANITE_SERVED_NAME=granite-4.1-8b # short parallel calls, so the 64K cap is ample. # 131072 — RESTORED 2026-07-16 (native max) for full-chapter summarization; GPU-1 # freed by the image-bench evict + qwen36-vl gone, so the 128K ctx fits again. -GRANITE_MAX_MODEL_LEN=131072 +# 16384 — SHRUNK 2026-07-27 (operator: granite is being phased out) to free GPU-1 +# room for vllm-coder (the Zed FIM seat). Full-chapter ctx dropped; util 0.13 holds +# 16K cleanly (32768 @ util 0.12 crash-looped: KV est-max was only 19376 tokens). +GRANITE_MAX_MODEL_LEN=16384 # FP8 KV cache (native on Blackwell cc 12.0). At 50K ≈ ~4.2 GB (vs ~8.4 GB at fp16). GRANITE_KV_CACHE_DTYPE=fp8 # util 0.35 (~33.6 GB) — tuned 2026-06-13 to leave ~3.5 GB free on GPU 1 alongside @@ -118,8 +121,24 @@ GRANITE_KV_CACHE_DTYPE=fp8 # 0.27 — RE-GROWN 2026-07-16 (same session) after char-rp moved onto GPU-1: spend the # leftover room on full-chapter context (max-len 131072). KV 15.0 GiB = 196,560 tokens # = 1.50x @ 131072; GPU-1 lands ~6.7 GB headroom (char-rp 30 + granite 27 + selene 17 + trio). -GRANITE_GPU_MEM_UTIL=0.27 -# Concurrency cap — set VERY HIGH 2026-07-16 (was vLLM default 128) so the KV pool is -# the only bound. granite is the fleet fan-out summarizer/classifier (many concurrent -# SHORT calls); default 128 capped below the KV bound (~192 @ 1K-tok). VRAM-neutral. -GRANITE_MAX_NUM_SEQS=1024 +# 0.13 — SHRUNK 2026-07-27 (was 0.27) with the max-len drop; frees ~14 GB on GPU-1 for +# vllm-coder. granite phasing out, so no longer worth the big KV pool. (0.12 was too +# small for even 16K KV; 0.13 gives ~1.4x @ 16384.) +GRANITE_GPU_MEM_UTIL=0.13 +# Concurrency cap. Was 1024 (fan-out summarizer); DROPPED to 256 on 2026-07-27 with the +# phasing-out shrink (256 is ample for the reduced summarizer load; smaller sched state). +GRANITE_MAX_NUM_SEQS=256 + +# Qwen2.5-Coder-1.5B (BASE) — FIM code-completion seat for Zed editor inline +# edit-predictions (deep-research pick 2026-07-27, Apache-2.0). GPU-1, alongside the +# small-model trio + (phasing-out) granite. Reached via LiteLLM `coder-fast` alias +# and the keyless zed-fim-proxy (stacks/zed-fim-proxy). util 0.06 (~5.7 GB) holds +# the 1.5B fp16 + fp8 KV; 8192 ctx is ample for FIM (KV = 13.75x concurrency). +CODER_PORT=8020 +CODER_GPU_ID=1 +CODER_MODEL=Qwen/Qwen2.5-Coder-1.5B +CODER_SERVED_NAME=qwen2.5-coder-1.5b +CODER_MAX_MODEL_LEN=8192 +CODER_KV_CACHE_DTYPE=fp8 +CODER_GPU_MEM_UTIL=0.06 +CODER_MAX_NUM_SEQS=32 diff --git a/stacks/vllm/compose.yaml b/stacks/vllm/compose.yaml index 527aa2f..3461041 100644 --- a/stacks/vllm/compose.yaml +++ b/stacks/vllm/compose.yaml @@ -278,6 +278,70 @@ services: - homepage.description=Granite 4.1 8B FP8 via vLLM (ana-ml2) - homepage.href=http://10.250.50.54:${GRANITE_PORT}/docs + vllm-coder: + image: vllm/vllm-openai:${VLLM_VERSION} + container_name: vllm-coder + restart: unless-stopped + ipc: host + ports: + - "${CODER_PORT}:8000" + volumes: + - /tank/aimodels/huggingface:/hfcache + environment: + - HF_HOME=/hfcache + - HF_HUB_CACHE=/hfcache/hub + - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-} + - VLLM_API_KEY=${API_KEY:-} + command: + # Qwen2.5-Coder-1.5B (BASE) — FIM code-completion seat for Zed edit-predictions + # (deep-research pick 2026-07-27). Native fill-in-the-middle: <|fim_prefix|> / + # <|fim_suffix|> / <|fim_middle|> (IDs 151659/151660/151661); Zed sends the + # FIM-formatted prompt to /v1/completions and vLLM passes it through (the FIM + # special tokens live in the tokenizer). BASE not -Instruct (FIM is a + # pretraining objective; base completions are cleaner). Apache-2.0. Runner-up = + # Qwen2.5-Coder-3B (higher HumanEval-FIM but non-commercial Qwen-Research license). + - ${CODER_MODEL} + - --served-model-name + - ${CODER_SERVED_NAME} + - --host + - 0.0.0.0 + - --port + - "8000" + - --gpu-memory-utilization + - ${CODER_GPU_MEM_UTIL} + - --max-model-len + - ${CODER_MAX_MODEL_LEN} + - --max-num-seqs + - ${CODER_MAX_NUM_SEQS} + - --dtype + - auto + - --kv-cache-dtype + - ${CODER_KV_CACHE_DTYPE} + - --enable-prefix-caching + deploy: + resources: + reservations: + devices: + - driver: nvidia + device_ids: + - "${CODER_GPU_ID}" + capabilities: + - gpu + healthcheck: + test: ["CMD", "curl", "-f", "http://localhost:8000/health"] + interval: 30s + timeout: 10s + retries: 3 + start_period: 300s + networks: + - tnet + labels: + - homepage.group=AI - Inference + - homepage.name=vLLM Qwen2.5-Coder 1.5B (FIM) + - homepage.icon=mdi-code-braces + - homepage.description=Qwen2.5-Coder-1.5B FIM code-completion (ana-ml2, Zed edit-predictions) + - homepage.href=http://10.250.50.54:${CODER_PORT}/docs + networks: tnet: name: traefik-net diff --git a/stacks/zed-fim-proxy/.env.example b/stacks/zed-fim-proxy/.env.example new file mode 100644 index 0000000..c2214fc --- /dev/null +++ b/stacks/zed-fim-proxy/.env.example @@ -0,0 +1,21 @@ +# zed-fim-proxy tunables — copy to `.env` on ana-docker (the real .env holds the +# scoped key and is server-only / gitignored). See conf/proxy.py + README.md. + +# Port the keyless route listens on (host network). +ZED_PORT=4141 + +# LiteLLM gateway to forward to (localhost:4000 via network_mode: host). +ZED_UPSTREAM=http://localhost:4000 + +# The ONLY model this route will forward (proxy rejects any other "model"). +ZED_ALLOWED_MODEL=coder-fast + +# Best-effort source-IP allowlist (comma-separated). Empty = allow all — fine on the +# internal network, but TIGHTEN to the Mac's observed source IP once it connects +# (watch `docker logs zed-fim-proxy` for the real source). Only meaningful if the +# proxy sees the real client IP (network_mode: host); a site-to-site NAT may mask it. +ZED_ALLOWED_IPS= + +# A LiteLLM virtual key SCOPED TO ZED_ALLOWED_MODEL ONLY (the real blast-radius +# bound). Mint: POST /key/generate {"models":["coder-fast"]}. NEVER commit the value. +ZED_SCOPED_KEY= diff --git a/stacks/zed-fim-proxy/README.md b/stacks/zed-fim-proxy/README.md new file mode 100644 index 0000000..58fd5dc --- /dev/null +++ b/stacks/zed-fim-proxy/README.md @@ -0,0 +1,47 @@ +# zed-fim-proxy + +A **keyless `/v1/completions` front door** on ana-docker for Zed's editor +edit-prediction (inline FIM completion) feature, which **cannot send an +`Authorization` header**. Runs on a separate port from LiteLLM and forwards to it +with an injected, model-scoped key. + +- **Host:** ana-docker `10.250.50.70`, port **4141** (`network_mode: host`). +- **Backs:** the `coder-fast` model (Qwen2.5-Coder-1.5B FIM seat, `stacks/vllm` + `vllm-coder` on ana-ml2:8020) via LiteLLM `:4000`. +- **Zed config** (`edit_predictions.open_ai_compatible_api`): `api_url: + http://10.250.50.70:4141/v1`, `model: coder-fast`, `prompt_format: qwen`, + `max_output_tokens: `. (The proxy also accepts `http://10.250.50.70:4141` — + it matches both `/v1/completions` and `/completions`.) + +## Security model (stdlib proxy in `conf/proxy.py`) + +Four guards + a scoped key — a keyless route that injects a working key is only +safe if it can't be pivoted: + +1. **POST + path** `/v1/completions` (or `/completions`) only. `GET /ping` is an + anonymous liveness (`{"service":"ok"}`). `/v1/chat/completions` is rejected. +2. **Model allowlist** — request body `model` must equal `ZED_ALLOWED_MODEL` + (`coder-fast`); anything else → 403. +3. **Injected scoped key** — a LiteLLM virtual key scoped to `coder-fast` ONLY + (`POST /key/generate {"models":["coder-fast"]}`). Even if guards 1–2 were + bypassed, the key reaches nothing else (verified: 403 on `gen`). **This is the + real blast-radius bound.** +4. **Best-effort source-IP allowlist** (`ZED_ALLOWED_IPS`) — only enforceable if + the proxy sees the real client IP (hence `network_mode: host`; docker + port-publish would NAT it away). A site-to-site NAT may still mask the Mac's + `10.0.10.83` — verify against `docker logs zed-fim-proxy` and tighten. Internal + network only; no public exposure. + +## Deploy + +``` +# conf/proxy.py -> /opt/docker/conf/zed-fim-proxy/proxy.py +# compose.yaml -> /opt/docker/compose/zed-fim-proxy/compose.yaml +# .env (from .env.example, with ZED_SCOPED_KEY filled) -> same dir, mode 600 +cd /opt/docker/compose/zed-fim-proxy && docker compose up -d +# verify keyless: +curl -s http://localhost:4141/v1/completions -H 'Content-Type: application/json' \ + -d '{"model":"coder-fast","prompt":"def add(a,b):\n return","max_tokens":16,"temperature":0.2}' +``` + +Stdlib-only proxy (no pip) in a bare `python:3.12-slim` container — no build. diff --git a/stacks/zed-fim-proxy/compose.yaml b/stacks/zed-fim-proxy/compose.yaml new file mode 100644 index 0000000..7284e75 --- /dev/null +++ b/stacks/zed-fim-proxy/compose.yaml @@ -0,0 +1,36 @@ +name: zed-fim-proxy + +# Keyless /v1/completions front door for Zed edit-predictions (coder-fast only). +# See conf/proxy.py for the guard model. Runs network_mode: host so it (a) sees the +# real client source IP for the best-effort allowlist (docker port-publish would NAT +# it away) and (b) reaches the LiteLLM gateway on localhost:4000. Stdlib-only proxy +# in a bare python image — no build, no pip. + +services: + zed-fim-proxy: + image: python:3.12-slim + container_name: zed-fim-proxy + restart: unless-stopped + network_mode: host + command: ["python", "/app/proxy.py"] + volumes: + - /opt/docker/conf/zed-fim-proxy/proxy.py:/app/proxy.py:ro + environment: + - ZED_PORT=${ZED_PORT:-4141} + - ZED_UPSTREAM=${ZED_UPSTREAM:-http://localhost:4000} + - ZED_ALLOWED_MODEL=${ZED_ALLOWED_MODEL:-coder-fast} + # best-effort source-IP allowlist (comma-separated); empty = allow all. + - ZED_ALLOWED_IPS=${ZED_ALLOWED_IPS:-} + - ZED_SCOPED_KEY=${ZED_SCOPED_KEY} + healthcheck: + test: ["CMD", "python", "-c", "import urllib.request,os; urllib.request.urlopen('http://localhost:%s/ping' % os.environ.get('ZED_PORT','4141'))"] + interval: 30s + timeout: 5s + retries: 3 + start_period: 10s + labels: + - homepage.group=AI - Gateways & Chat + - homepage.name=Zed FIM Proxy + - homepage.icon=mdi-code-braces-box + - homepage.description=Keyless /v1/completions for Zed edit-predictions (coder-fast, ana-docker) + - homepage.href=http://10.250.50.70:4141/ping diff --git a/stacks/zed-fim-proxy/conf/proxy.py b/stacks/zed-fim-proxy/conf/proxy.py new file mode 100644 index 0000000..abec38e --- /dev/null +++ b/stacks/zed-fim-proxy/conf/proxy.py @@ -0,0 +1,102 @@ +"""zed-fim-proxy — a keyless front door for Zed's edit-prediction feature. + +Zed's edit_predictions.open_ai_compatible_api provider cannot send an +Authorization header, so it needs a route where POST /v1/completions succeeds +with no key. This proxy is that route, on a separate port from LiteLLM, with +four narrow guards + an injected model-scoped key: + + 1. POST only, path exactly /v1/completions (GET /ping is an anonymous liveness). + 2. source-IP allowlist (best-effort — set ZED_ALLOWED_IPS; empty = allow all). + NOTE: only meaningful if this process sees the real client IP — run the + container with network_mode: host (docker port-publish would NAT the source + to the bridge gateway and defeat it). Behind a site-to-site NAT it may still + see a mesh IP, not the Mac — verify against the access log at deploy. + 3. request body "model" must equal ZED_ALLOWED_MODEL (default coder-fast). + 4. injects Authorization: Bearer $ZED_SCOPED_KEY (a LiteLLM virtual key scoped + to that ONE model) and forwards to $ZED_UPSTREAM. The scoped key is the real + blast-radius bound: even if guards 1-3 were bypassed, the key can reach + nothing but coder-fast. + +Stdlib only (no pip) — runs in a bare python:slim container. +""" +import json +import os +import sys +import urllib.error +import urllib.request +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer + +UPSTREAM = os.environ.get("ZED_UPSTREAM", "http://localhost:4000").rstrip("/") +SCOPED_KEY = os.environ["ZED_SCOPED_KEY"] +ALLOWED_MODEL = os.environ.get("ZED_ALLOWED_MODEL", "coder-fast") +ALLOWED_IPS = set(x.strip() for x in os.environ.get("ZED_ALLOWED_IPS", "").split(",") if x.strip()) +PORT = int(os.environ.get("ZED_PORT", "4141")) +TIMEOUT = float(os.environ.get("ZED_TIMEOUT", "60")) + + +class Handler(BaseHTTPRequestHandler): + server_version = "zed-fim-proxy/1.0" + + def _json(self, code, obj): + body = json.dumps(obj).encode() + self.send_response(code) + self.send_header("Content-Type", "application/json") + self.send_header("Content-Length", str(len(body))) + self.end_headers() + self.wfile.write(body) + + def do_GET(self): + if self.path.rstrip("/") == "/ping": + return self._json(200, {"service": "ok"}) + return self._json(404, {"error": "not found"}) + + def do_POST(self): + src = self.client_address[0] + if ALLOWED_IPS and src not in ALLOWED_IPS: + return self._json(403, {"error": f"source {src} not allowed"}) + # Accept both /v1/completions and /completions so the Zed api_url can be + # set to either http://host:4141/v1 or http://host:4141 (Zed appends + # /completions). /v1/chat/completions is NOT in the set → stays rejected. + if self.path.split("?")[0].rstrip("/") not in ("/v1/completions", "/completions"): + return self._json(404, {"error": "only POST /v1/completions or /completions"}) + try: + length = int(self.headers.get("Content-Length", 0)) + except ValueError: + return self._json(400, {"error": "bad content-length"}) + raw = self.rfile.read(length) + try: + body = json.loads(raw) + except Exception: + return self._json(400, {"error": "invalid json body"}) + if body.get("model") != ALLOWED_MODEL: + return self._json(403, {"error": f"model must be '{ALLOWED_MODEL}'"}) + req = urllib.request.Request( + UPSTREAM + "/v1/completions", + data=raw, + method="POST", + headers={ + "Content-Type": "application/json", + "Authorization": f"Bearer {SCOPED_KEY}", + }, + ) + try: + with urllib.request.urlopen(req, timeout=TIMEOUT) as r: + data, code, ctype = r.read(), r.status, r.headers.get("Content-Type", "application/json") + except urllib.error.HTTPError as e: + data, code, ctype = e.read(), e.code, e.headers.get("Content-Type", "application/json") + except Exception as e: # noqa: BLE001 + return self._json(502, {"error": f"upstream error: {e}"}) + self.send_response(code) + self.send_header("Content-Type", ctype) + self.send_header("Content-Length", str(len(data))) + self.end_headers() + self.wfile.write(data) + + def log_message(self, fmt, *args): + sys.stderr.write("%s - %s\n" % (self.client_address[0], fmt % args)) + + +if __name__ == "__main__": + print(f"zed-fim-proxy on :{PORT} -> {UPSTREAM} (model={ALLOWED_MODEL}, " + f"ip-allowlist={'set' if ALLOWED_IPS else 'OFF'})", flush=True) + ThreadingHTTPServer(("0.0.0.0", PORT), Handler).serve_forever()