fix(gateway-chat): raise max_tokens default 1024 -> 4096; document single-file bind-mount trap
Operator reported the gen seat 'cutting off'. It is not the seat. The chat UI's max_tokens field defaults to 1024, and every thinking seat spends part of that budget on CoT before emitting content, so the completion truncates mid-sentence with finish_reason=length and reads as model degeneracy. Measured through the gateway: gen 1024 -> finish=stop, 716w (survives, but marginally) gen-reasoning 1024 -> finish=length, cut mid-word <-- the symptom gen-reasoning 4096 -> finish=stop, 839w Seat itself is clean: direct long-form generations return finish=stop with complete sentences and a max repeated 6-gram of 1 (no degeneration), and enable_thinking:false still holds on every non-thinking alias, so the AEON swap did not cause this. Also documents a trap that made the fix look like it had not applied: compose bind-mounts a single FILE, and a single-file bind mount binds the INODE. rsync writes-and-renames, producing a new inode, so the container kept serving the old content while the host file showed the new value -- silently, with no error. docker restart does NOT clear it; the container must be recreated. Verify against what the container sees, never the host file. Applies to any file-source mount; directory mounts are unaffected.
This commit is contained in:
@@ -65,7 +65,13 @@
|
||||
<textarea id="sys" placeholder="You are a helpful assistant."></textarea>
|
||||
<div class="row">
|
||||
<div><label>Temperature</label><input id="temp" type="number" step="0.05" value="1"></div>
|
||||
<div><label>Max tokens</label><input id="max" type="number" step="1" value="1024"></div>
|
||||
<!-- 4096, not 1024: every thinking seat on this gateway (gen-reasoning,
|
||||
char-rp-reasoning) spends part of the completion budget on CoT before
|
||||
emitting any content, so a 1024 cap truncates mid-sentence with
|
||||
finish_reason=length and reads as the MODEL being degenerate. Measured
|
||||
2026-08-16: gen-reasoning at 1024 -> finish=length cut mid-word; at
|
||||
4096 -> clean stop at 839 words. -->
|
||||
<div><label>Max tokens</label><input id="max" type="number" step="1" value="4096"></div>
|
||||
</div>
|
||||
<label style="margin-top:.5rem">🔊 TTS (quoted text)</label>
|
||||
<label class="hint" style="display:flex;gap:.45rem;align-items:center;margin:.15rem 0"><input id="ttsOn" type="checkbox" style="width:auto"> Auto-play quoted dialogue</label>
|
||||
|
||||
Reference in New Issue
Block a user