7bc648672f
Symptom: granite-4-small and qwen3.6-27b were evicting each other when called in alternation. granite is the news-digest curator (fires twice daily on cron) — being evicted means a cold reload (~5s) on every digest tick, plus visible churn whenever the user uses 27b concurrently. Added granite-4-small to the `pinned` group as a persistent member. ~5-6 GB at Q4_K_M + 120K KV ≈ comfortable inside the existing pin budget (qwen3.5-9b ~6 GB → ~12 GB total persistent). Single RTX 6000 Ada is 48 GB, leaves ~36 GB headroom for whichever non-pinned model the user invokes (qwen3.6-27b at ~30 GB fits cleanly). Updated the pinned group's docstring to capture the current member set + VRAM math + the historical context (qwen3.6-35-a3b was here, was too heavy, got removed yesterday). Marked the granite ttl: 0 with the matching "pinned — never unloads" comment as the other group members.
llama-swap
GGUF model server with on-demand model swapping. Served via llama.cpp's llama-server under the llama-swap proxy.
Server: ana-ml2
Port: 9292 (configurable via .env)
GPU: both (unpinned — runtime: nvidia grants access to all devices; per-model GPU selection happens inside config.yaml)
Files
compose.yaml— canonical compose. Deployed to/opt/docker/compose/llama-swap/compose.yamlon ana-ml2..env.example— template for the per-host.env. Copy to.envon the server and tweak.config.yaml— model definitions and groups. Deployed to/opt/docker/conf/llama-swap/config.yamlon the server.
Homepage labels are in the compose file under the AI Systems group, matching the convention used by vllm-qwen3 and infinity.
Deploy a fresh install
scripts/deploy-stack.sh ana-ml2 llama-swap
ssh ana-ml2 '
cd /opt/docker/compose/llama-swap && \
cp -n .env.example .env && \
docker compose config && \
docker compose up -d && \
docker compose logs --tail=30
'
Model reference conventions
- Modern entries: use
-hf <user>/<repo>[:<quant>]— reads from the shared HF cache, nothing to pre-stage outsidehf download - Legacy entries: use
--model /models/<dir>/<file>.gguf— reads GGUFs from/tank/aimodels/llm/(pre-HF-cache era, gradually being migrated)
New models should prefer the -hf pattern.
Deploy updates to config only
# After editing config.yaml here:
scp config.yaml ana-ml2:/opt/docker/conf/llama-swap/config.yaml
ssh ana-ml2 'cd /opt/docker/compose/llama-swap && docker compose restart'
Deploy updates to compose only
# After editing compose.yaml or .env.example here:
scripts/deploy-stack.sh ana-ml2 llama-swap
ssh ana-ml2 'cd /opt/docker/compose/llama-swap && docker compose up -d'