docs(ana-ml2): correct GPU spec Ada -> RTX PRO 6000 Blackwell (96GB, cc 12.0)
ana-ml2 was upgraded 2026-06 from dual RTX 6000 Ada (46GB, cc 8.9) to
dual RTX PRO 6000 Blackwell Max-Q (96GB, cc 12.0 / sm_120). Update the
stale hardware facts across the workspace:
- CLAUDE.md servers table row
- servers/ana-ml2/README.md hardware spec (+ refreshed system-details.txt)
- stacks/vllm compose + .env.example FP8/KV comments (Ada cc 8.9 -> Blackwell cc 12.0)
- stacks/llama-swap config VRAM-budget comment (48GB -> 96GB, GPU-0 pin)
Also corrects the adjacent stale 'Phi-4-mini' comment in the granite
service block (the service has been Granite 4.1 8B since 34a43a0).
Doc/comment-only; no runtime change.
This commit is contained in:
@@ -636,9 +636,9 @@ groups:
|
||||
#
|
||||
# Current pins:
|
||||
# qwen3.5-9b — ~6 GB at Q4 + KV. General-purpose chat baseline.
|
||||
# VRAM budget: ~6 GB persistent in the pin slot. Single RTX 6000 Ada
|
||||
# is 48 GB, so this leaves ~40 GB for whichever non-pinned model the
|
||||
# user invokes alongside.
|
||||
# VRAM budget: ~6 GB persistent in the pin slot. llama-swap is pinned to
|
||||
# GPU 0 (a single RTX PRO 6000 Blackwell, 96 GB), so this leaves ~90 GB
|
||||
# for whichever non-pinned model the user invokes alongside.
|
||||
#
|
||||
# granite-4-small WAS pinned here; removed 2026-06-04 — superseded by
|
||||
# phi4-mini (vLLM FP8, stacks/vllm → vllm-phi4). Freed ~24 GB (120K KV).
|
||||
|
||||
Reference in New Issue
Block a user