Codifies the previously-manual workflow described in
stacks/llama-swap/README.md: install hf CLI via pipx (one-time),
inject hf_transfer for fast multi-connection downloads,
`hf download` into the shared HF cache at /tank/aimodels/huggingface
with optional --include filter.
Model-format-agnostic by design — same playbook handles GGUFs for
llama-swap and safetensors for vLLM (both stacks read the same cache
dir via HF_HOME=/hfcache). Does NOT edit any consumer's config.yaml;
per-model run params (ctx-size, sampler defaults, quant choice,
chat template, etc.) stay human-curated.
Usage:
scripts/elway ana-ml2 --playbook playbooks/pull-hf-model.yaml \
--var hf_repo=<user>/<repo> \
[--var allow_patterns='*Q6_K*']
Idempotent: hf CLI skips already-cached blobs; re-runs are
sub-second when the snapshot is already complete.
Smoke-tested 2026-05-13 against:
- mradermacher/Selene-1-Mini-Llama-3.1-8B-GGUF (Q6_K, ~6.5 GB)
- Skywork/Skywork-Reward-V2-Llama-3.1-8B (full safetensors, ~16 GB)