dd627b3b31
Dark-Scarlett v1.0 refuses too much on the char-rp-reasoning seat. Root cause is visible on its card: ReadyArt/Dark-Scarlett-v1.0-27B is a plain finetune of stock Qwen/Qwen3.6-27B, tagged unaligned/nsfw/erp but carrying no abliteration -- the base model's refusal machinery is intact, so off-distribution prompts revert to safety-tuned Qwen3.6 behaviour. Candidate kkuspa/Qwen3.6-27B-Fable-Fusion-711-...-NVFP4A16 is refusal-ablated (Heretic), a structural edit rather than a behavioural preference. Verified before pulling: Qwen3_5ForConditionalGeneration wrapper class, 15 mtp.* tensors in a separate bf16 shard AND individually enumerated in quantization_config.ignore, NVFP4A16 with null input_activations, FP8 KV scales shipped, 262K context, Apache-2.0. Staged byte-verified at /tank/aimodels/fable-fusion-711-nvfp4a16 (28.55 GB). services/refusal-probe: deterministic marker-based classifier (LLM judge only breaks AMBIGUOUS ties, never overrides), intensity-graded battery so the report renders a refusal curve rather than an average, benign controls that gate run validity, and explicit handling of the thinking-budget trap -- empty content with finish_reason=length is reasoning exhausting the budget, not a refusal, and is excluded from the denominator. stacks/fablefusion-charrp-probe: throwaway :8019 seat serving as char-rp-probe, never aliased to char-rp-reasoning. MTP depth 3 rather than the card's 5 -- its 1.56x was measured greedy, and acceptance degrades at the temp 1.0 this seat is probed at. GPU1 is zero-sum at 94.9/97.9 GB, so this seat takes Dark-Scarlett's vacated slot; the A/B is sequential.
8 lines
253 B
Bash
8 lines
253 B
Bash
# ana-ml2 GPU1 THROWAWAY probe seat (Fable-Fusion 711). Real .env lives on the host.
|
|
# Takes Dark-Scarlett's vacated slot — DS must be down first (GPU1 is zero-sum).
|
|
FF_GPU_MEM_UTIL=0.44
|
|
FF_MAX_MODEL_LEN=262144
|
|
FF_GPU_ID=1
|
|
FF_PORT=8019
|
|
FF_MTP_DEPTH=3
|