Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-15-svos-miranda-plugin-validation.md
T
vh b9b14b5baf feat(parakeet): stand up Parakeet STT on fv-ml1 GPU 3 + LiteLLM ext-stt/whisper-1
Retargets the existing sherpa-onnx stack from irv-ml1 to fv-ml1's utility card
and puts it behind the gateway. GPU 3 was the only card with room: 0/1/2 carry
the vLLM seats at 84-95.5 GB of 96.

Changes:

- compose: pin GPU via `device_ids: ["3"]` (the dead on-host stub used
  `count: all`, which would have handed a 0.6B ASR seat all four cards);
  join traefik-net; port 8300; homepage href to the live FV address.
- .env.example: default to the v3 int8 model (25 European languages, 464 MiB)
  rather than English-only v2; models to /tank/parakeet/models.
- app.py: warm the recognizer at startup before uvicorn accepts traffic.

The warmup is not an optimisation. ONNX Runtime's CUDA EP compiles and autotunes
lazily on the FIRST DECODE, and on sm_120 that measured 45.7s cold (reproduced at
45.1s on a second container) against ~0.50s warm. A 45s first request is
indistinguishable from a hang and LiteLLM's default timeout abandons it long
before it returns. Decoding 1s of silence at load moves the cost inside the
healthcheck's 300s start_period; first real request after restart is now 0.65s.

Verification, because "provider=cuda" in the log is only an echo of the env var:
ORT falls back to CPU silently and still returns correct text, so the service
being up and the transcript being right establishes nothing. The discriminator is
a process on GPU 3 (922 MiB), confirmed. Controls both directions — a known TTS
sentence transcribes near-exactly (positive), 3s of digital silence returns
empty (null). Warm throughput 0.50s median on an 8.52s clip, n=5, spread
0.47-0.65s, single-stream, one clip: a smoke measurement with its harness
stated, not a benchmark.

Gateway aliases `ext-stt` (engine-neutral, mirrors ext-tts) and `whisper-1`
(OpenAI-compatible drop-in) registered via POST /model/new, i.e. LiteLLM's
Postgres store where the ext-tts family already lives — no gateway restart, and
config.yaml is consequently not a complete picture of what the gateway serves.
Both verified end to end.

The aliases use a raw IP deliberately: ana-docker resolves no .internal names at
all (resolv.conf points at 1.1.1.1), and LiteLLM only reaches irv-ml1 through a
hand-pinned extra_hosts entry. A second hosts entry would mean recreating the
container and bouncing the gateway for every consumer.

Also records the svos_miranda plugin validation pass and its structural findings,
and notes that the irv-ml1 parakeet is still running — there are two now, and
retiring the old one is the operator's call.
2026-09-15 01:41:41 -07:00

3.4 KiB

svos_miranda Hermes plugin — validation pass (2026-09-15)

svos-dev asked infra-ops to run hermes plugins validate → doctor → compat on /home/lkraven/development/svos/hermes_plugin/ and report before enabling. Hermes Agent v0.21.1 (2026.9.7), local b88e6776, on nh3-dev.

The blocker (found, fixed by svos-dev at c964e64)

Three absolute intra-package imports — from hermes_plugin._vendored, .forward, .jwt — pinned the package to its source directory name. The documented install renames it to svos_miranda, and Hermes loads directory plugins under the hermes_plugins.<dir> namespace; in neither case does a top-level hermes_plugin exist. Fix: three relative imports.

⚠ The harness hid this from three different readers. My first validate passed the import only because my cwd was the SVOS repo root. svos-dev's test suite imports hermes_plugin.* from that same root, and their editable install resolves the name from anywhere on the box — it only reproduced for them once sys.path was stripped. Same class as feedback_filters_that_silently_narrow_the_window: the instrument carried the result.

⚠ Two of the three commands CANNOT pass this plugin, ever

Neither is fixable from the plugin side. Both are now documented in its README.

  • validate — two independent causes. Its RecordingContext.get_config (hermes_cli/plugin_validate.py:219-222) returns the default for every key, ignoring config.yaml entirely, so dispatch_key is always "". And its register_tool returns None, which the plugin's INV-P6 guard correctly reads as a name collision — so even with a key supplied it raises on the first tool.
  • doctor — runs register() under a temp HERMES_HOME with sockets blocked, so no config exists there either.

Tool-level gotchas worth remembering

  • ⚠ hermes plugins doctor exits 0 even when it prints ERROR. Needs --ci.
  • ⚠ hermes plugins compat <nonexistent-path> prints ✓ and exits 0. A typo'd path reads as a pass. (The instrument itself is sound — verified with a throwaway plugin importing a real deprecated path, which it flagged with file:line, exit 1.)
  • ⚠ doctor's sandbox registry starts EMPTY — 0 entries, no built-ins. So doctor cannot detect tool-name collisions at all. validate's separate static "built-in tool collisions" check is what covers that.
  • The real PluginContext.register_tool (hermes_cli/plugins.py:449-491) returns a truthy PluginRegistration on success — confirmed against the live runtime.

Roster verified another way

Since neither command can supply config, a probe mirroring validate's context but returning real settings and a truthy handle gave: 8 tools with repo_read_enabled: true, 7 with false or omitted, names matching plugin.yaml exactly, zero hooks/middleware/commands. All nine settings-validation controls (quoted booleans, "90 s", zero/negative timeouts, empty/whitespace strings) raise errors naming their own key.

A false finding I caught on myself

A probe registering read_file got back a PluginRegistration instead of the expected refusal — which looked like the plugin's collision reading was wrong. It was not: doctor's sandbox holds no built-ins, so nothing was claimed and my positive case was not positive. Reported as untested rather than as a finding.

Where enabling stands

Blocked on the operator only. See the OPEN row in the index for the sequence.