docs(gemma4-erp-tune): size the run against the checkpoint — QLoRA is structurally unavailable
The proposed shape was QLoRA r64. It cannot be run as specified. The checkpoint stores each layer's 128 experts as two fused 3-D nn.Parameter tensors (experts.gate_up_proj [128,1408,2816], experts.down_proj [128,2816,704] — no .weight suffix, so they are parameters, not modules). bitsandbytes 4-bit replacement walks nn.Linear only, so 22.84B params / 42.54 GiB — 88.5% of the model — is skipped and stays BF16. load_in_4bit saves ~3.1 GiB of 48.07 and does not error while doing it. Verdict: plain LoRA on BF16, ~57.6 GiB at micro-batch 1, +2.5 GiB per additional 8192-token sequence. Two sizing items were absent from the brief and both are load-bearing: - vocab 262,144 x seq 8,192 = 2.147B logits, with final_logit_softcapping 30.0 adding a saved pre-cap tensor. Naive HF cross-entropy peaks at ~28-30 GiB transient at batch 1, which puts the run at ~85.6 GiB on a 95.6 GiB card — it starts, then OOMs on the first long sample. Fused or chunked linear CE is mandatory and must be smoke-proven before a window is booked, since Liger may not carry a Gemma-4 MoE patch. - v_proj does not exist on layers 5/11/17/23/29 (attention_k_eq_v on the full-attention layers). A v_proj target silently produces no adapter there, and k_proj adapts K and V simultaneously. 45.96M trainable at r64 across q/k/v/o. Placement, measured: GPU0 has 53.46 GiB free beside gen, ~4 GiB short, and gen's footprint grows with uptime. Stopping mog-sec frees 74.29 GiB on GPU1, which holds micro-batch 4 at 61.8 GiB with margin for Scriberr. Recommend standing down sec (2 aliases, last request ~5h ago) rather than gen (7 aliases, 765 busy-engine log lines in 24h). Estimated 1.28e18 FLOPs for the epoch at ~3.67B active params; 4-10 hours at 10-25% MFU. 7,104 packed sequences is only 444 optimizer steps at effective batch 16, which makes the wall-clock-checkpointing amendment concrete rather than hypothetical. Package as a uv venv on /tank: root is 91% full (36 GB) with /var/lib/docker on it.
This commit is contained in:
@@ -0,0 +1,235 @@
|
|||||||
|
# Gemma-4 26B-A4B ERP/RP tune — GPU sizing adjudication
|
||||||
|
|
||||||
|
_Measured 2026-08-24 on `ana-ml2` against
|
||||||
|
`/tank/aimodels/gemma4-26b-a4b-it-heretic-bf16` (llmfan46 abliterated trainee)._
|
||||||
|
|
||||||
|
Division of labour for this run: **Eitri writes the harness, brokkr-smithy-dev
|
||||||
|
audits, infra-ops owns the GPU window and executes.** This document is the
|
||||||
|
sizing infra-ops owes; it is arithmetic against the real checkpoint and the
|
||||||
|
real card, not an estimate.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. ⚠ QLoRA IS NOT AVAILABLE ON THIS ARCHITECTURE
|
||||||
|
|
||||||
|
**The proposed shape was QLoRA r64. It cannot be run as specified**, and the
|
||||||
|
reason is structural rather than a tuning preference.
|
||||||
|
|
||||||
|
The checkpoint stores each layer's 128 experts as **two fused 3-D
|
||||||
|
`nn.Parameter` tensors**, not as 128 `nn.Linear` modules:
|
||||||
|
|
||||||
|
model.language_model.layers.N.experts.gate_up_proj BF16 [128, 1408, 2816]
|
||||||
|
model.language_model.layers.N.experts.down_proj BF16 [128, 2816, 704]
|
||||||
|
|
||||||
|
Note the absence of a `.weight` suffix — compare `mlp.down_proj.weight`
|
||||||
|
(an `nn.Linear`) against `experts.down_proj` (a bare parameter). That is the
|
||||||
|
tell, and it is decisive: **`bitsandbytes` 4-bit replacement walks `nn.Linear`
|
||||||
|
modules.** A fused 3-D parameter is not one, so it is skipped and stays BF16.
|
||||||
|
|
||||||
|
What `load_in_4bit=True` would actually buy on this model:
|
||||||
|
|
||||||
|
| block | params | BF16 | after bnb NF4 | saved |
|
||||||
|
|---|---:|---:|---:|---:|
|
||||||
|
| **MoE experts** (fused 3-D — **NOT quantized**) | 22.84 B | 42.54 GiB | **42.54 GiB** | **0** |
|
||||||
|
| lm attention (`nn.Linear`) | 1.11 B | 2.07 GiB | 0.52 GiB | 1.55 |
|
||||||
|
| dense shared MLP (`nn.Linear`) | 0.54 B | 1.00 GiB | 0.25 GiB | 0.75 |
|
||||||
|
| vision tower (`nn.Linear`) | 0.57 B | 1.06 GiB | 0.27 GiB | 0.79 |
|
||||||
|
| embed (tied, normally kept BF16) | 0.74 B | 1.38 GiB | 1.38 GiB | 0 |
|
||||||
|
| router + norms | 0.01 B | 0.02 GiB | 0.02 GiB | 0 |
|
||||||
|
| **total** | **25.81 B** | **48.07 GiB** | **~44.98 GiB** | **~3.1 GiB** |
|
||||||
|
|
||||||
|
**88.5% of the model is in tensors bitsandbytes cannot touch.** "QLoRA" here
|
||||||
|
means paying the NF4 dequant tax on 6% of the weights to save 6% of the
|
||||||
|
footprint. The premise does not survive contact with the checkpoint.
|
||||||
|
|
||||||
|
> **Eitri: do not hard-code a `BitsAndBytesConfig` / `load_in_4bit` path.**
|
||||||
|
> It will not error loudly — it will load, report a 4-bit model, and quietly
|
||||||
|
> leave 42.5 GiB in BF16. Same silent-failure shape as the stale chat template.
|
||||||
|
|
||||||
|
**The one thing that could overturn this** is a third-party fork shipping
|
||||||
|
custom grouped-GEMM 4-bit MoE kernels for this specific architecture (Unsloth
|
||||||
|
is the candidate). **Not chased, deliberately** — see §4, where the run fits in
|
||||||
|
BF16 without displacing anything the fleet depends on, which collapses QLoRA's
|
||||||
|
value to zero. If it is ever revisited, it must be *before* the harness
|
||||||
|
hard-codes a quantization path, not after.
|
||||||
|
|
||||||
|
**Verdict: plain LoRA on BF16 weights.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. What the run actually costs
|
||||||
|
|
||||||
|
Adapter targeting `q_proj,k_proj,v_proj,o_proj` at r64, computed from the real
|
||||||
|
tensor shapes:
|
||||||
|
|
||||||
|
| | layers | per layer | total |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| sliding-attention (q 4096, kv 2048, o 4096) | 25 | 1,507,328 | 37,683,200 |
|
||||||
|
| full-attention (q 8192, kv 1024, o 8192) | 5 | 1,654,784 | 8,273,920 |
|
||||||
|
| **trainable** | | | **45,957,120** (0.178% of base) |
|
||||||
|
|
||||||
|
⚠ **`v_proj` DOES NOT EXIST ON LAYERS 5, 11, 17, 23, 29.** Those are the
|
||||||
|
`full_attention` layers, and `attention_k_eq_v: true` means one projection
|
||||||
|
serves both K and V. Consequences the harness must respect:
|
||||||
|
|
||||||
|
- PEFT matches by name suffix, so a `v_proj` target **silently produces no
|
||||||
|
adapter** on those five layers. Do not assert a fixed adapter count.
|
||||||
|
- Adapting `k_proj` on a global layer **adapts K and V simultaneously** — a
|
||||||
|
different intervention than on the sliding layers. If that asymmetry matters
|
||||||
|
to the recipe, say so explicitly rather than discovering it in the loss curve.
|
||||||
|
|
||||||
|
### Memory budget, batch 1, `max_seq_len` 8192
|
||||||
|
|
||||||
|
| item | GiB | note |
|
||||||
|
|---|---:|---|
|
||||||
|
| base weights BF16 | 48.07 | measured: 25,805,936,206 params × 2 B |
|
||||||
|
| adapters + grads + AdamW fp32 m/v | 0.75 | 45.96 M trainable — rounding error |
|
||||||
|
| checkpointed layer inputs | 1.29 | 30 × 8192 × 2816 × 2 B |
|
||||||
|
| recompute peak, one layer | ~2.5 | 8192 tok × top-8 of 128, `moe_intermediate 704` |
|
||||||
|
| loss head, **fused/chunked CE** | ~2.0 | see the warning below |
|
||||||
|
| CUDA context + cuBLAS + fragmentation | ~3.0 | the item `--gpu-memory-utilization` never covered |
|
||||||
|
| **total** | **~57.6** | |
|
||||||
|
|
||||||
|
Marginal cost per extra sequence in the micro-batch: **~2.5 GiB.**
|
||||||
|
|
||||||
|
| micro-batch | GiB |
|
||||||
|
|---:|---:|
|
||||||
|
| 1 | 54.3 |
|
||||||
|
| 2 | 56.8 |
|
||||||
|
| **4** | **61.8** |
|
||||||
|
| 6 | 66.8 |
|
||||||
|
| 8 | 71.8 |
|
||||||
|
|
||||||
|
### ⚠ The loss head is the whole ballgame, and it is not in the brief
|
||||||
|
|
||||||
|
`vocab_size` is **262,144** and `final_logit_softcapping` is **30.0**. One
|
||||||
|
8192-token sequence produces **2.147 billion logits**. Through a naive HF
|
||||||
|
`ForCausalLM` loss that is:
|
||||||
|
|
||||||
|
BF16 logits 4.0 GiB
|
||||||
|
fp32 upcast 8.0 GiB
|
||||||
|
softcap tanh saved 8.0 GiB (autograd keeps the pre-cap tensor)
|
||||||
|
softmax + grad 8.0 GiB
|
||||||
|
------------------------------
|
||||||
|
~28-30 GiB transient, at BATCH 1
|
||||||
|
|
||||||
|
Naive CE at batch 1 lands the run at **~85.6 GiB on a 95.6 GiB card** — it will
|
||||||
|
appear to work and then OOM on the first long sample. At micro-batch 4 it is
|
||||||
|
~120 GiB and never starts. **Fused/chunked linear cross-entropy is mandatory,
|
||||||
|
not an optimization.**
|
||||||
|
|
||||||
|
⚠ Honest uncertainty: Liger ships per-architecture patches and Gemma-4 MoE with
|
||||||
|
softcapping may not have one. Three ways out, in order of preference —
|
||||||
|
(a) generic `LigerFusedLinearCrossEntropyLoss` wired against the lm_head with
|
||||||
|
softcapping applied inside the chunk; (b) `cut-cross-entropy`; (c) hand-rolled
|
||||||
|
sequence-chunked CE. **This must be proven on a 10-step smoke run before the
|
||||||
|
window is booked**, because everything else in this document assumes it works.
|
||||||
|
|
||||||
|
### Step count
|
||||||
|
|
||||||
|
58.2 M tokens / 20,576 samples = **2,829 tokens/sample average** — well under
|
||||||
|
8192, so packing matters.
|
||||||
|
|
||||||
|
- Packed to 8192: **7,104 sequences.** At micro-batch 4 × grad-accum 4
|
||||||
|
(effective 16) → **444 optimizer steps for the whole epoch.**
|
||||||
|
- ⚠ That is a *small* step count. A "checkpoint every 100 steps" default gives
|
||||||
|
four checkpoints across a multi-hour run. This is exactly why the amendment
|
||||||
|
asked for **wall-clock-interval checkpointing, not step-count** — the case is
|
||||||
|
now concrete, not hypothetical.
|
||||||
|
- ⚠ **Packing must use `position_ids` + varlen/block-diagonal attention.** Naive
|
||||||
|
concatenation bleeds samples into each other. `sliding_window` is 1024 on 25
|
||||||
|
of 30 layers so the damage is bounded there — but the 5 `full_attention`
|
||||||
|
layers see the entire packed sequence.
|
||||||
|
|
||||||
|
**Open question for brokkr/Eitri:** what fraction of the 20,576 samples exceed
|
||||||
|
8192 tokens? Below ~2%, 8192 is right. A long tail means truncation is cutting
|
||||||
|
the ends off RP scenes, which is where the signal lives.
|
||||||
|
|
||||||
|
### Runtime
|
||||||
|
|
||||||
|
Active parameters per token ≈ **3.67 B** (2.93 B routed + attention, plus the
|
||||||
|
0.74 B tied lm_head matmul). Forward + backward + gradient-checkpoint recompute
|
||||||
|
≈ 6 × active × tokens = **1.28e18 FLOPs** for the epoch.
|
||||||
|
|
||||||
|
At 10–25% MFU on a 300 W-capped Max-Q card — HF MoE paths with 704-wide experts
|
||||||
|
are not efficient — **4 to 10 hours, most likely ~6.** Treat as a band, not a
|
||||||
|
number; it will be measured on the smoke run.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Where it fits (measured 2026-08-24, 18:20 PDT)
|
||||||
|
|
||||||
|
Card total: 97,887 MiB = **95.60 GiB** each.
|
||||||
|
|
||||||
|
| | GPU0 | GPU1 |
|
||||||
|
|---|---|---|
|
||||||
|
| resident | `vllm-gen` 42,508 MiB (up 3 h) | `vllm-mog-sec` 56,624 MiB + `embed` 3,304 + `reward` 9,512 + `coder` 6,158 + `rerank-a3` 2,170; Scriberr pinned here, loads on demand |
|
||||||
|
| free | 54,741 MiB = **53.46 GiB** | 19,446 MiB = **18.99 GiB** |
|
||||||
|
|
||||||
|
- **GPU0 beside `gen`: does not fit.** 53.46 GiB free against ~57.6 GiB needed —
|
||||||
|
short by ~4 GiB. And `gen` is only three hours old: measured footprint runs
|
||||||
|
38.5 GiB fresh → 42.5 GiB at 3 h → 45.6 GiB at 3 days. Budgeting against the
|
||||||
|
current number is budgeting against a moving one.
|
||||||
|
- **GPU1 as-is: nowhere near.** 18.99 GiB.
|
||||||
|
- **GPU1 with `vllm-mog-sec` stopped: 76,070 MiB = 74.29 GiB free.** Fits
|
||||||
|
micro-batch 4 (61.8 GiB) with ~12 GiB clear even after reserving ~6 GiB for
|
||||||
|
Scriberr's on-demand whisper load.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Recommendation
|
||||||
|
|
||||||
|
**Run on GPU1. Stand down `sec`, not `gen`.**
|
||||||
|
|
||||||
|
| | `gen` | `sec` |
|
||||||
|
|---|---|---|
|
||||||
|
| aliases | 7 (`gen`, `gen-reasoning`, `chat-judge`, `image-judge`, summarizer/classifier family) | 2 (`sec`, `sec-reasoning`) |
|
||||||
|
| standing role | the fleet's general seat; a documented always-available dependency in global `CLAUDE.md` | M.O.G.-SEC, niche |
|
||||||
|
| measured traffic | 765 busy-engine log lines in 24 h — continuously in use | bursty; peak 8 concurrent, **last request ~5 h ago** |
|
||||||
|
|
||||||
|
`sec` is genuinely in use and this is not free — but it is the smaller blast
|
||||||
|
radius by a wide margin, and it is the difference between a 4–10 hour window
|
||||||
|
that nobody outside the security work notices and one that takes the fleet's
|
||||||
|
default model offline for a working day.
|
||||||
|
|
||||||
|
Also fold in, cheaply and reversibly:
|
||||||
|
|
||||||
|
- **Move Scriberr to GPU0 for the window** (`device_ids: ["0"]`, one compose
|
||||||
|
edit + `up -d`). GPU0 will be sitting on ~53 GiB free with `gen` up, and it
|
||||||
|
removes contention from the training card entirely.
|
||||||
|
- **Micro-batch 4, grad-accum 4** (effective 16, 444 steps). Micro-batch 6 is
|
||||||
|
the measured ceiling; the gap is deliberate margin. A seat crash-looped
|
||||||
|
earlier the same day on 0.6 GiB of assumed headroom — this document does not
|
||||||
|
repeat that.
|
||||||
|
- **Package as a `uv` venv on `/tank`, not a Docker image.** Root is at **91%
|
||||||
|
(36 GB free)** and `/var/lib/docker` lives on it; a PyTorch training image
|
||||||
|
would come close to filling it. `/tank` has 4.0 TB.
|
||||||
|
|
||||||
|
### If `sec` may not be stood down
|
||||||
|
|
||||||
|
Fallback is GPU0 with `gen` stopped — the run fits with ~38 GiB to spare, which
|
||||||
|
buys micro-batch 8 and a shorter wall-clock. Make it **resumable and split** in
|
||||||
|
that case: INV-T7 plus wall-clock checkpointing already allows the window to be
|
||||||
|
broken into two or three shorter stretches with `gen` restored between them,
|
||||||
|
rather than one long outage.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Standing warnings that apply to this run
|
||||||
|
|
||||||
|
- **Never render training examples through the base's own
|
||||||
|
`chat_template.jinja`.** Every third-party Gemma-4 derivative ships a stale
|
||||||
|
one; the trainee's is 365 lines against upstream's 390. Use
|
||||||
|
`/tank/aimodels/gemma4-26b-a4b-it-bf16/chat_template.jinja`. Training through
|
||||||
|
the wrong template is train/serve skew with no error — it presents as a
|
||||||
|
tuning failure.
|
||||||
|
- **Base path and chat-template path are config keys, not constants.** The
|
||||||
|
trainee base already moved once (stock BF16 → `-heretic-bf16`).
|
||||||
|
- **`--gpu-memory-utilization` sizes the KV cache only.** It does not cover CUDA
|
||||||
|
context, graphs, or non-torch overhead — the same misreading that OOM'd the
|
||||||
|
char-rp seat.
|
||||||
|
- **Serving the result is not settled.** LoRA-on-NVFP4 hot-swap was a silent
|
||||||
|
no-op on vLLM 0.24.0 (#47639, proven quant-agnostic). Retest on the tagged
|
||||||
|
`vllm/vllm-openai:v0.27.1` already on disk. **If it still no-ops, the harness
|
||||||
|
must emit merged weights** — and Eitri needs that requirement while he is
|
||||||
|
early, not after the run.
|
||||||
@@ -110,7 +110,7 @@ no longer deployed sidecars here. See Recent decisions.)
|
|||||||
|
|
||||||
_As of 2026-08-24 (late) — a very long ops session. The homepage arc and the char-rp arc both closed. **The live thread is the ERP/RP trainee: QLoRA sizing is the next conversation and brokkr-smithy-dev is waiting on it.**_
|
_As of 2026-08-24 (late) — a very long ops session. The homepage arc and the char-rp arc both closed. **The live thread is the ERP/RP trainee: QLoRA sizing is the next conversation and brokkr-smithy-dev is waiting on it.**_
|
||||||
|
|
||||||
- **🔴 NEXT UP — QLoRA SIZING, AND I OWN THE RUN.** Operator set the division of labour: **Eitri writes the training harness, brokkr audits, infra-ops owns the GPU window AND executes the run.** Harness contract requires headless-from-a-config (no notebook, no interactive steps) and INV-T7 resumable-from-checkpoint. I asked for one amendment: **checkpoint on a wall-clock interval, not only step count**, since step time under contention is not knowable in advance. Rough shape: QLoRA r64, attention-only adapters, max_seq_len 8192, 1 epoch, ~58.2M tokens / 20,576 samples. **I have NOT sized it and must not guess** — GPU0 carries `gen` (~38.5 GiB fresh, ~45.6 GiB after days up) and GPU1 carries `sec` + Scriberr with ~19.4 GiB free. The window will require standing something down; which seat is the arithmetic I owe.
|
- **🔴 SIZING DONE — AND IT KILLED THE QLoRA PREMISE. AWAITING THE SEAT CALL.** → `docs/pfi/gemma4-erp-tune-sizing.md` (measured, not estimated). **QLoRA is structurally unavailable on this architecture:** the checkpoint stores each layer's 128 experts as two fused 3-D `nn.Parameter` tensors (`layers.N.experts.gate_up_proj` `[128,1408,2816]`, `.down_proj` `[128,2816,704]` — note the missing `.weight` suffix), and `bitsandbytes` 4-bit replacement walks `nn.Linear` modules only. **88.5% of the model (22.84B params / 42.54 GiB) is untouchable; `load_in_4bit` saves ~3.1 GiB of 48.07 and does NOT error.** ⚠ **Eitri must not hard-code `BitsAndBytesConfig`** — it loads, reports 4-bit, and silently leaves 42.5 GiB BF16. Verdict: **plain LoRA on BF16**, ~57.6 GiB at micro-batch 1, +2.5 GiB per extra 8192-seq. ⚠ **The loss head is the real driver and was NOT in the brief:** vocab 262,144 × 8192 = 2.147B logits, plus `final_logit_softcapping 30.0` → **~28-30 GiB transient at BATCH 1** through naive HF CE (= ~85.6 GiB total, an OOM-on-first-long-sample). **Fused/chunked linear CE is mandatory and must be smoke-proven before the window is booked** (Liger may lack a Gemma-4 MoE patch). ⚠ **`v_proj` DOES NOT EXIST on layers 5/11/17/23/29** (`attention_k_eq_v` on the full-attention layers) — a `v_proj` target silently no-ops there and `k_proj` adapts K and V at once; 45.96M trainable at r64. ⚠ 7,104 packed seqs → **only 444 optimizer steps** at effective batch 16, which is why wall-clock checkpointing matters concretely. **RECOMMENDATION: run on GPU1, stand down `sec` (2 aliases, last request ~5h ago), NOT `gen` (7 aliases, 765 busy-engine lines/24h).** Stopping mog-sec frees 74.29 GiB; micro-batch 4 = 61.8 GiB. Est. **4-10h, likely ~6.** Also: move Scriberr to GPU0 for the window, and package as a `uv` venv on `/tank` — **root is 91% full (36 GB)** and `/var/lib/docker` is on it.
|
||||||
- **⚠ TELL EITRI BEFORE HE HARD-CODES: the trainee base changed.** Contract still names the stock BF16. It is now `/tank/aimodels/gemma4-26b-a4b-it-heretic-bf16` (llmfan46). **Base path AND chat-template path must be config keys, not constants** — and the template must point at upstream's (`gemma4-26b-a4b-it-bf16/chat_template.jinja`), never the base's own, or training renders a different prompt than production serves.
|
- **⚠ TELL EITRI BEFORE HE HARD-CODES: the trainee base changed.** Contract still names the stock BF16. It is now `/tank/aimodels/gemma4-26b-a4b-it-heretic-bf16` (llmfan46). **Base path AND chat-template path must be config keys, not constants** — and the template must point at upstream's (`gemma4-26b-a4b-it-bf16/chat_template.jinja`), never the base's own, or training renders a different prompt than production serves.
|
||||||
- **🟢 char-rp seat = Gemma-4 26B-A4B MoE NVFP4** on `:8016`, both aliases on ONE backend. **Currently DOWN by operator instruction** to hold GPU0 headroom for the tune. `gen` is UP and verified. MeroMero-v2 retained stopped in `created` state for rollback (stop-then-start; both bind :8016). → `persistent-memory.d/2026-08-24-charrp-gemma4-moe-swap-and-trainee.md`
|
- **🟢 char-rp seat = Gemma-4 26B-A4B MoE NVFP4** on `:8016`, both aliases on ONE backend. **Currently DOWN by operator instruction** to hold GPU0 headroom for the tune. `gen` is UP and verified. MeroMero-v2 retained stopped in `created` state for rollback (stop-then-start; both bind :8016). → `persistent-memory.d/2026-08-24-charrp-gemma4-moe-swap-and-trainee.md`
|
||||||
- **🟢 THREE trainee-relevant model dirs on `/tank/aimodels/`, NOT interchangeable:** `gemma4-26b-a4b-it-bf16` (stock, 49 GB — its chat_template is the canonical upstream one), `gemma4-26b-a4b-it-heretic-bf16` (llmfan46 abliterated, the trainee), `gemma4-26b-a4b-it-abliterated-bf16` (TrevorJS, KL 0.09, alternate). Plus `-nvfp4` (served) and `-nvfp4a16` (activation control). ⚠ **BF16 cannot coexist with `gen`** — 48.07 GiB of weights on a 94.97 GiB card. Every BF16 window means gen stops.
|
- **🟢 THREE trainee-relevant model dirs on `/tank/aimodels/`, NOT interchangeable:** `gemma4-26b-a4b-it-bf16` (stock, 49 GB — its chat_template is the canonical upstream one), `gemma4-26b-a4b-it-heretic-bf16` (llmfan46 abliterated, the trainee), `gemma4-26b-a4b-it-abliterated-bf16` (TrevorJS, KL 0.09, alternate). Plus `-nvfp4` (served) and `-nvfp4a16` (activation control). ⚠ **BF16 cannot coexist with `gen`** — 48.07 GiB of weights on a 94.97 GiB card. Every BF16 window means gen stops.
|
||||||
|
|||||||
Reference in New Issue
Block a user