Files
esh-pfi-infrastructure/stacks/csm
vh a4b8c2a9f4 stacks: add csm (Sesame Conversational Speech Model) to the TTS bench
Sesame CSM-1B via phildougherty/sesame_csm_openai — OpenAI-compat
/v1/audio/speech, context-aware conversational speech (voice-agent
layer, not a plain reader). Port 8201 on irv-ml1. Gated model:
requires CSM_HF_TOKEN (license acceptance) — placeholder in .env.example,
real token only in host .env.
2026-05-31 10:59:22 -07:00
..

Sesame CSM

Sesame's Conversational Speech Model (the engine behind the "Maya/Miles" demo) — a context-aware speech generator (Llama backbone + Mimi audio codec) — served via phildougherty/sesame_csm_openai, an OpenAI-compat wrapper.

Server: irv-ml1 (Irvine, WireGuard-only) Port: 8201 (container listens on 8000) GPUs: pins to device 0 (RTX 3090) by default; ~8 GB VRAM Image: local/csm:v1 — built locally from a pinned git SHA of the wrapper via docker buildx's git URL context Upstream wrapper: phildougherty/sesame_csm_openai (MIT) Upstream model: sesame/csm-1b (gated — Sesame's own license) + unsloth/Llama-3.2-1B tokenizer + Mimi codec

Why this stack exists — and what it is not

CSM is a conversational speech model, not a plain reader. It conditions on prior conversation turns (text and audio) to choose prosody and tone — it's designed to be the speech layer of a voice agent, where a separate LLM produces the words and CSM voices them with context-aware delivery.

  • As pure TTS it works standalone (OpenAI-compat POST /v1/audio/speech, voices alloy/echo/fable/onyx/nova/shimmer plus cloned IDs).
  • Its real edge — contextual prosody across turns — only pays off in interactive / voice-agent use, not monologue narration. For skaldsong's reader, the emotive single-voice engines (Fish S2-Pro, IndexTTS-2) and Dia (dialogue) remain the better fits; CSM is here for voice-agent experiments.

⚠️ Gated model — token required before first boot

sesame/csm-1b is gated. Before deploying:

  1. Accept the license at https://huggingface.co/sesame/csm-1b.
  2. Put a HF token (read scope) in CSM_HF_TOKEN in the host .env (never in .env.example).

Without it the first-boot model download 401s.

Deploy

# from this workstation (irv-ml1 is WG-only — routes via ana-wg):
scripts/deploy-stack.sh irv-ml1 csm
# then on irv-ml1, first run builds from the pinned SHA:
#   docker compose up -d --build

First boot pulls csm-1b + the Llama-3.2-1B tokenizer + Mimi into CSM_CACHE_DIR; the 600 s start_period covers it.

Notes

  • Pin CSM_SHA to a full 40-char commit before relying on this — .env.example ships main, which is not reproducible.
  • The wrapper can also serve Dia-1.6B, but we run Dia from its own dia stack — keep this one CSM-only to avoid overlap.
  • Endpoints: /v1/audio/speech (OpenAI-compat), /v1/audio/voices, /v1/audio/models. No dedicated /health route — the healthcheck probes /v1/audio/voices (only answers once the model is loaded).