news-digest: fixes from first deploy on ana-docker
Three iterations to get end-to-end:
1. Dockerfile missed COPY run-digest.sh — cron's exec target wasn't
in the image, every fire failed. Added COPY + chmod.
2. Jinja template used {{ list|sum(attribute='items') }} which
sum()s lists with start=0 → TypeError int+list. Switched to
computing reddit_total / tech_total in Python and passing as
template args.
3. LLM defaulted to qwen3.5-35-a3b which (a) is broken in
llama-swap (model process exits on launch), (b) when working,
defaults to extended-thinking mode that eats the entire token
budget without producing any visible content. Same pattern with
qwen3.6-35-a3b. Switched default to granite-4-small — small (4B),
fast (~1s/call), no thinking-mode pathology, returns clean JSON.
Whole pipeline now runs in ~35s total across 8 sources.
Also hardened the LLM response parser to fall back to
reasoning_content when content is empty — catches the thinking-mode
case if anyone ever points the digest at one of those models. Plus
the deploy playbook gained DOCKER_BUILDKIT=0 because ana-docker is
on docker 20.10 which doesn't carry the buildx driver versions our
newer client expects ("client version 1.52 is too new"). Real fix is
upgrading docker on the fleet — separate workstream.
This commit is contained in:
@@ -21,7 +21,13 @@ NEWS_DIGEST_TZ=America/Los_Angeles
|
||||
# Model picked for one-shot summarization quality + low VRAM impact.
|
||||
# qwen3.5-35-a3b is loaded in the persistent group on ana-ml2.
|
||||
LLAMA_SWAP_URL=http://10.250.50.54:9292
|
||||
LLAMA_SWAP_MODEL=qwen3.5-35-a3b
|
||||
# granite-4-small — small (~4B), fast (~1s/call), no extended-thinking
|
||||
# phase that eats the token budget like qwen3.x do. Plenty of capability
|
||||
# for the one-sentence-tldr + one-word-tag task. To swap to a larger
|
||||
# model later, ones currently working: gemma4-26b-a4b, granite-4-small.
|
||||
# Avoid: qwen3.5-35-a3b (model file broken — process exits on launch),
|
||||
# qwen3.6-35-a3b (defaults to thinking mode, eats budget without output).
|
||||
LLAMA_SWAP_MODEL=granite-4-small
|
||||
LLAMA_SWAP_TIMEOUT=180
|
||||
|
||||
# ── miniflux (feed source for Tech aggregators + subreddit list) ─────
|
||||
|
||||
Reference in New Issue
Block a user