feat(chatterbox-fast): context-priming at joins (§1.6, opt-in)
Prime early joins by prepending the prior sentence as backward prosodic context,
generating context+content together, then discarding the context audio. The cut
snaps to the inter-sentence pause (energy-minimum search around the context's
solo duration) with a 5ms fade-in to kill any seam click (app: _cut_at_pause /
_fade_in / Engine.generate_primed). Opt-in via request `prime` (default off).
Scheduler: priming is AFFORDABILITY-GATED so it can never starve. A primed chunk
costs ~(2·context + content)/rtf (a 2nd context-solo pass); a chunk is only primed
when buffer ≥ prime_buffer_factor (1.5) × that cost, else it falls back to a cold
generate. Consequences proven in the GPU-free sim (17 tests):
- fires on early joins for any GPU at/above rtf_prior (3.4 = 3090; A6000 ~3.8-4.0)
- self-skips (degrades to cold) on a slower-than-fleet GPU rather than starving
- never primes chunk 0 (latency-critical)
Also fixed a latent Phase-1 bug: margin_first was applied at chunk 0 (budget always
0 there) so it never did anything — now applied at chunk 1 (the first transition).
Live A/B on irv-ml1 (A6000, GLaDOS): TTFB unaffected (445 vs 467ms), no starvation;
priming fired on chunk 2 (gen 1.6s for the doubled pass). On typical text exactly
ONE early join safely primes — priming chunk 2 flattens the buffer so later/larger
chunks no longer clear the safety gate. Samples: ~/chatterbox-ab/_p2_{cold,primed}.wav.
This commit is contained in:
@@ -33,8 +33,10 @@ DEFAULT_TEXT = (
|
||||
)
|
||||
|
||||
|
||||
def run(host: str, text: str, out: str, *, oneshot: bool, voice: str | None) -> None:
|
||||
payload = {"text": text, "format": "pcm", "stream": not oneshot}
|
||||
def run(host: str, text: str, out: str, *, oneshot: bool, voice: str | None,
|
||||
prime: bool, prime_n: int) -> None:
|
||||
payload = {"text": text, "format": "pcm", "stream": not oneshot,
|
||||
"prime": prime, "prime_first_n": prime_n}
|
||||
if voice:
|
||||
payload["voice"] = voice
|
||||
req = urllib.request.Request(
|
||||
@@ -98,8 +100,11 @@ def main() -> None:
|
||||
ap.add_argument("--out", default="/refs/_fast.wav")
|
||||
ap.add_argument("--voice", default=None)
|
||||
ap.add_argument("--oneshot", action="store_true", help="whole-text one-shot baseline")
|
||||
ap.add_argument("--prime", action="store_true", help="context-prime the early joins")
|
||||
ap.add_argument("--prime-n", type=int, default=2, help="how many early joins to prime")
|
||||
args = ap.parse_args()
|
||||
run(args.host, args.text, args.out, oneshot=args.oneshot, voice=args.voice)
|
||||
run(args.host, args.text, args.out, oneshot=args.oneshot, voice=args.voice,
|
||||
prime=args.prime, prime_n=args.prime_n)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
Reference in New Issue
Block a user