revert(chatterbox-fast): drop context-priming (§1.6) — discard-cut leaks context

Revert the priming feature from d707439. Live A/B caught an audible artifact: the
context-priming discard-cut left part of the throwaway prefix in the output, so a
clause ("...without a trace of sarcasm,") was spoken an extra time.

Root cause is structural: generate() returns one finished waveform with no marker
for where the prefix ends, and the model renders the same prefix with different
timing when followed by content than when generated solo — so the duration-estimate
+ energy-minimum cut is a guess and can leave a sliver (or a whole clause) of prefix
in. A reliable cut would need token-level access (the abandoned native-streaming
arc) or a per-chunk ASR/alignment pass (heavy, still imperfect, eats the latency
budget). Fails the agreed bar: "keep only if it closes the gap without a seam."

Kept from d707439: the .gitignore (build artifacts). NOT re-applied: the bundled
margin_first fix — wiring it would shrink chunk 1 (more joins = worse coherence),
against the operator's priority, and margin=0.8 there is already starvation-safe.

Coherence loss at joins stays an accepted limitation; cold streaming was judged
"really good". Phase 1 + Phase 2 parity/perf untouched. Next: Phase 3 deploy.
This commit is contained in:
vh
2026-06-01 23:26:56 -07:00
parent d707439041
commit 090e70aed5
4 changed files with 22 additions and 190 deletions
+3 -8
View File
@@ -33,10 +33,8 @@ DEFAULT_TEXT = (
)
def run(host: str, text: str, out: str, *, oneshot: bool, voice: str | None,
prime: bool, prime_n: int) -> None:
payload = {"text": text, "format": "pcm", "stream": not oneshot,
"prime": prime, "prime_first_n": prime_n}
def run(host: str, text: str, out: str, *, oneshot: bool, voice: str | None) -> None:
payload = {"text": text, "format": "pcm", "stream": not oneshot}
if voice:
payload["voice"] = voice
req = urllib.request.Request(
@@ -100,11 +98,8 @@ def main() -> None:
ap.add_argument("--out", default="/refs/_fast.wav")
ap.add_argument("--voice", default=None)
ap.add_argument("--oneshot", action="store_true", help="whole-text one-shot baseline")
ap.add_argument("--prime", action="store_true", help="context-prime the early joins")
ap.add_argument("--prime-n", type=int, default=2, help="how many early joins to prime")
args = ap.parse_args()
run(args.host, args.text, args.out, oneshot=args.oneshot, voice=args.voice,
prime=args.prime, prime_n=args.prime_n)
run(args.host, args.text, args.out, oneshot=args.oneshot, voice=args.voice)
if __name__ == "__main__":