26 KiB
Archival memory — eshpfi-management
Entries moved out of persistent-memory.md to keep the active file scannable. Read this when researching historical decisions or revisiting past foot-guns.
Recent decisions (archived)
-
[2026-05-12]corviduo-dev (Worldtree-team dev VM, 10.250.50.152, CT 106 on pfi-pve) added toservers/inventory. Treat like SF client hosts: PFI hosts + provides emergency-ops backstop; Worldtree team owns OS config + deploys + backup decisions. Archived 2026-05-27. -
[2026-05-12]Worldtree:latesttag drift bug — fixed by health-gated:latestadvance in vh/worldtree's deploy workflow (architect commit8ef3801): only tag:latestAFTER the new container's/healthprobe passes. Build-on-host stacks here don't have this problem because the playbook always builds the SHA-tagged image from agit reset --hard <ref>checkout. Archived 2026-05-27. -
[2026-05-12]asset-engine stack scaffolded LAN-direct athttp://10.250.50.70:8200. Initially included Traefik labels for public hostname; user pulled them out (internal tool, no public TLS surface needed). Pattern: internal tools default LAN-direct; Traefik wiring only when external/TLS required. Archived 2026-05-27. -
[2026-05-12]asset-engine catalog gainslifecycle: { stack, vram_gb, gpu_device_id }per irv-ml1 service for the orchestrator feature. SSH keypair scaffolded atana-docker:/opt/docker/conf/asset-engine/ssh/for asset-engine container → irv-ml1 orchestration via dedicated ed25519 key. Archived 2026-05-27. -
[2026-05-13]pull-hf-repo.yamlis the canonical HF-fetch playbook on ana-ml2. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadcalls. Archived 2026-05-27. -
[2026-05-13]Selene-1-Mini-Llama-3.1-8B added to llama-swap as judge model. mradermacheri1-Q6_Kimatrix quant (~6.5GB). AtlaAI reward/eval model — temp 0.01, ctx 32K, q8_0 KV cache. New JUDGE / EVAL MODELS section instacks/llama-swap/conf/config.yaml. Archived 2026-05-27. -
[2026-05-13]vllm-qwen3→vllmstack rename. Addedvllm-rewardservice (Skywork-Reward-V2-Llama-3.1-8B-AWQ classifier). Three vLLM services share GPU 1 (embed 0.20, rerank 0.20, reward 0.30 utilization; 30% headroom). All use--runner pooling; classification drives via model'sarchitectures: [LlamaForSequenceClassification]in config.json, NOT--task classify(deprecated in vLLM 0.19.1). Archived 2026-05-27. -
[2026-05-13]/tend-docs first pass deletions:stacks/infinity/removed (retired by vllm). Archiveddocs/asset-engine/design-brief.md→docs/archive/asset-engine/with archival header. Fixedpfi-pveVM list to fullqm listenumeration. Dropped stale weak-password section frompfi-postgres(rotation done 2026-04-23). Archived 2026-05-27. -
[2026-05-14]althing-chamber stack scaffolded: chamber + forseti. Internal LAN-only at port 7881 (chamber default 7878 collides with task-board). Two-service compose, shared SQLite bind-mount, build-on-host pattern via vh/althing's gitea-workflow. Forseti is the canonical dev for this stack (galdrabok is on a different project). Archived 2026-05-31. -
[2026-05-16]althing-chamber Phase 2: addedalthing-agent-runneras third compose service (worldtree-driver agent dispatcher). All three althing services use the same image;command:selects entrypoint. Safe to enable preemptively (sleeps when no driver=worldtree handles declared). Archived 2026-05-31. -
[2026-05-17]Phase 3.1 cross-process streaming uses Valkey 8 alpine as a sibling compose service instacks/althing-chamber/, redis-protocol pub/sub for high-volumemsg_delta/msg_thinking/msg_start/msg_completeevent kinds. DB bridge keepsmsg_curated+floor_grant(structured / canonical). Two-channel architecture, no overlap. chamber + agent-runnerdepends_on: valkey: service_healthy. Archived 2026-05-31. -
[2026-05-17]Worldtree admin workflow shift (per vh): infra-ops gets its own permanent admin-tier key (61419c92, stored atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin). Future admin ops route through this key, not the bootstrap admin via docker-as-root. Archived 2026-05-31. -
[2026-05-17]Worldtree env-var addition checklist: anytime introducingos.environ.get("FOO")in worldtree code, update BOTH.env.exampleANDcompose.yaml's&worldtree-envanchor in the same PR. Same Z_AI_API_KEY-shape footgun bitBIFROST_CLIENT_ALLOWED_HOSTS(#170) until worldtree-dev added the passthrough line in08f02b2. Archived 2026-05-31. -
[2026-05-18]Volva systemd install complete after three-stage debug. Final unit at/etc/systemd/system/volva.serviceruns asUser=lkravenwithProtectHome=read-only+ReadWritePaths=/home/lkraven/.althing /home/lkraven/.codexcarve-outs for state writes.VOLVA_ALTHING_CLI=/home/lkraven/ .local/bin/althing-cli+ALTHING_HANDLE=volvaboth pinned in env.sh. Archived 2026-05-31. -
[2026-05-19]Worldtree CD disk-hygiene strategy: watermark gate (env-tunable threshold + window, fail-loud on still-low post-prune)- eager post-deploy prune (only after
:latestadvance succeeds, usesdocker image prune -a --filter "until=24h"which respects in-use semantic — protects pinned + personal images automatically). Combined: demo VM holds ~24h of deploy history instead of unbounded accumulation. Shipped in vh/Worldtree PR #184 (306cd61+613dac2+bd91df5). Archived 2026-05-31.
- eager post-deploy prune (only after
-
[2026-05-19]Skaldsong CD shape: shape (1) of three operator options — container + Gitea registry + pull-restart, matching Worldtree's pattern. Target host ana-docker (NOT nh3-dev where skaldsong-dev runs for iteration). SHA-pin only for now; health-gated:latestadvance is a follow-up once/healthexercises Worldtree- Kokoro reachability. Archived 2026-05-31.
-
[2026-05-19]Skaldsong prod (ana-docker) switched from demo Worldtree (:8080) to personal (:8081). Sameuser_id=skaldsongas the nh3-dev hand-launch key — shared Heimdall agent slot (skaldsong:wizard-v2), differentkey_ids for independent rotation. Demo Worldtree stays for isolation; personal becomes the multi-consumer dev iteration instance. Archived 2026-05-31. -
[2026-05-19]mead-hall Bifrost v0.3 end-to-end smoke green. Closed task #32 (althing thread01KRV1M2KW6N6HBEXGTH72QXCA). Wire layer (handshake + binding + dispatch) + data-flow (per-dispatch JWT claims →ctx.session_idpopulated → real session-scoped data) + agent-loop (LLM reads + quotes back) all proven. Resolves the "stalled mid-Worldtree" state from the 2026-05-17 snapshot. Archived 2026-05-31. -
[2026-05-25]v0.25.3 lofn tuning:temperature 0.6 → 1.0+repetition_penalty 1.0 → 1.15on default+fast profiles. Heretic-abliterated qwen3.6 was locking into degenerate attractors at the model's thinking-mode floor (0.6). Pattern: abliterated/uncensored Qwen variants need higher temp + non-trivial rep-penalty than base, NOT the model-card's documented floors. Archived 2026-05-31. -
[2026-05-25]Worldtree #205 v0.25.2 ships/app/config/as bind-mount + root-then-drop entrypoint shim (gosu). Operators get persistent per-instance config without container-rebuild. Same bind-mount pattern hit twice subsequently in v0.27.0 (selene) and v0.29.9 (echo) — bind-mount shadows image-baked defaults, so every new required key surfaces as a crash-loop on existing deployments. The v0.29.12 canonical example files close this loop. Archived 2026-05-31. -
[2026-05-26]Worldtree v0.27.0/v0.27.1 fixes Tier 3 GET visibility.available_agents()helper was over-applied toGET /agents/<id>, masking ALL Tier 3 agents regardless of row state. Bug only visible as "agent not found" via GET; storage was fine (silent-2xx PATCHes had persisted correctly). v0.27.1 added fail-fast hardening for the startup pre-resolve fragility class. Archived 2026-05-31. -
[2026-05-26]Skaldsong v0.30.7 defensive 409→PATCH fallback. v0.30.6's GET-then-define-or-patch path crash-looped against pre-v0.27.0 Worldtree's GET-visibility bug (GET 404 phantom → define 409 conflict). v0.30.7 catches the 409 and falls through to PATCH (which silently 2xx'd on the pre-fix Worldtree). Archived 2026-05-31. -
[2026-05-27]Worldtree v0.29.x landed full saga→echo refactor + config-validator hardening (v0.29.10 create_provider family-before-regex; v0.29.11 collect-then-raise echo startup validators; v0.29.12 ships providers.yaml.example/defaults.yaml.example canonical configs; v0.29.13 reasoning_content extraction + catalog family lookup). Operator-asked, worldtree-dev-shipped, CI-deployed. Archived 2026-06-01. -
[2026-05-27]artemis-31b-v1i added to llama-swap + worldtree personal. BeaverAI Gemma 4 31B Q6_K (~28.6GB), 128K ctx,--reasoning-format deepseek(gemma format unsupported in deployed llama.cpp). Worldtree catalogfamily: gemmaso GemmaProvider routes reasoning tokens. Archived 2026-06-01. -
[2026-05-27]Skaldsong streaming TTS v0.32.0→v0.32.2: chunked-batch SSE (one Kokoro POST per paragraph); defensive event_stream catch-all; NDJSON parsing for Kokoro /dev/captioned_speech multi-line responses. Archived 2026-06-01. -
[2026-05-31]Dia2 deployed as two fixed-model instances (dia2-2b:8200,dia2-1b:8202) fromlocal/dia:v2, retiring legacy Dia 1.6B; catalogdiaentry removed → dia2-2b + dia2-1b (breaking for asset-engine). Rationale: the devnen wrapper is single-model and IGNORES the OpenAImodelfield (verified on its live OpenAPI), so the only way to offer both Dia2 models as real per-request asset-engine choices is one fixed endpoint per model.3139e81(deploy),db15638(catalog swap). Archived 2026-06-03. -
[2026-05-31]Both dia2 catalog entries route to the wrapper's richer/ttsendpoint (not/v1/audio/speech) to expose the full control surface (cfg_scale/temperature/top_p/cfg_filter_top_k/voice_mode/clone); all defaults sourced from the wrapper'sCustomTTSRequestPydantic blessed values. Voice default isvoice_mode: clone+clone_reference_filename: Abigail.wavso a stable (non-random-gender) voice is the out-of-box behavior.55602b7,5c47843. Archived 2026-06-03. -
[2026-05-31]Zonos REST adapter (stacks/zonos/adapter/,local/zonos-api) — thin OpenAI-ish/v1/audio/speechFastAPI in front of the Gradio-only Zonos SDK; JSON-envelope{audio, audio_format, seed}(Zonos is the fleet's first seedable TTS). Port 8203 (moved off 8201 — collided with csm). Built; NOT deployed (stack down for VRAM). Also fixed the upstream image's missing CMD (71df6f7).81efa8d. Archived 2026-06-03. -
[2026-05-31]Catalog schema regenerated: addedCatalogLifecycle+reproducibility.seed_field(b7b2130). Resolves the stale-schema hand-off; catalog now validates clean. (asset_enginecatalog.pyPydantic already supported both — schema file was just behind.) Archived 2026-06-03. -
[2026-05-31]TTS bench expanded withstacks/{dia,zonos,csm}(666f7f3dia+zonos,a4b8c2acsm). The bench already had Fish S2-Pro / Chatterbox-Turbo / IndexTTS-2 / CosyVoice3 / Kokoro / VibeVoice / Qwen3-TTS / Kyutai. (csm since removed 2026-06-01.) Archived 2026-06-03. -
[2026-05-31]Remote browser/iPad/Vision-Pro driver seat for the agent-fleet zellij sessionClaudestood up on nh3-dev (ttyd behind Caddy, network-gated). Out of this repo — full architecture + the HTTP2/OSC52/Safari-auth gotchas in auto-memoryreference_ttyd_fleet_seat. Archived 2026-06-03. -
[2026-05-30]esh-docker-vm NFS boot-ordering fix:playbooks/fix-esh-nfs-boot-ordering.yaml(c0458d9, +53157b1drop-in filename-collision fix) adds_netdev,nofailto the four 10.0.50.50 NFS mounts + a dockerAfter=remote-fs.targetdrop-in — resolves paperlessExited(255)on reboot. traefik also gainedrestart: unless-stopped. Full incident → auto-memoryincident_esh_docker_nfs_boot_race. Archived 2026-06-03.
Tried and abandoned (archived)
-
[2026-04-30]task-board workflow withcontainer: image: debian:bookworm-slim— fails:actions/checkout@v4needsnodeat runtime, slim image lacks it. Switched tonode:20-bookworm-slim(has node + apt) or runner-label default. (Pattern revisited 2026-05-17 for skaldsong-dev: container override needsnodejsapt-installed unless it IS the default.) Archived 2026-05-27. -
[2026-04-30]Dropping thecontainer:directive before runner re-registration with docker-schema labels — runner silently falls back to host mode (jobs run inside the alpineact_runnercontainer itself, no apt). The:hostsuffix in startup logs (labels updated to: [pfi-fleet:host ana-docker:host]) is the giveaway. Fix: register withpfi-fleet:docker://<image>schema labels. Archived 2026-05-27. -
[2026-04-30]Updating runner labels by editing.envand bouncing — doesn't take. The.runnerregistration cache pins labels at first registration; env-var updates are read each start but the stored token + UUID are tied to the original label set on the gitea side. Fix: stop runner, delete.runner, generate new admin registration token, redeploy. Archived 2026-05-27. -
[2026-04-30]git reset --hard origin/<sha>indeploy-task-board.yaml(and the in-repo nevermore playbook before fix) — invalid syntax:origin/prefix only works for branch refs. SHAs needgit reset --hard <sha>directly. Resolved withgit rev-parse --verify --quiet "origin/{{ ref }}^{commit}"first, then bare"{{ ref }}^{commit}"fallback. Archived 2026-05-27. -
[2026-04-30]AssumingDEPLOY_SSH_KEYwas at user scope after task-board wiring — it was actually only repo-scope onvh/task-board. vor's first CI run failed with empty SSH key (printf '%s\n' "" > ~/.ssh/id_ed25519). Fix: copy secret to user scope atgitea.phasefinal.com/user/settings/actions/secrets. Archived 2026-05-27. -
[2026-04-30]grep -vE "^(#|$)"to inspect.envfor sanity — leaked the fullMINIFLUX_PASSWORDline into the transcript. Then a follow-up redaction attempt withsed -E "s/=(.{4}).*$/=\1<redacted>/"still leaked the first 4 chars. Lesson: when probing secret-bearing files, use field-by-field SELECTIVE inspection (grep -E "^(KEY1|KEY2)=") rather than negative filters; for any password line,grep -c(existence) ortest -n "$(...)"(non-empty), nevercator value-printing. Archived 2026-05-27. -
[2026-05-08]Filtering Traefik's UTC access log by Gitea-local-PDT timestamp substrings (grep "2026/05/08 15:1[2-7]") returned zero matches and led to a wrong "no /v2/ traffic in 12 days" conclusion. Gitea logs in PDT, Traefik logs in UTC — same host, different timezones. Always normalize timezones (UTC) when correlating logs across services on the same box. Cost: ~30 min in the wrong direction. Archived 2026-05-27. -
[2026-05-08]Bumping GiteaPER_WRITE_TIMEOUT/PER_WRITE_PER_KB_TIMEOUTto addressunexpected EOFon/v2/.../blobs/uploads/PATCH — wrong direction. Both govern response writes, not request body reads.unexpected EOFfrom Go's HTTP server means the client closed mid-body-upload; not a knob Gitea exposes server-side. Archived 2026-05-27. -
[2026-05-12]Defaulting asset-engine to Traefik-routed (asset-engine.phasefinal.comwithanaprodcert resolver) on first scaffold — user pulled it back to LAN-direct. Internal tools default LAN-direct; only add Traefik when an external/TLS surface is actually needed. Archived 2026-05-31. -
[2026-05-12]Routing althing thread replies throughgaldrabokwhen the actual dev handle isforseti— bus rejectedto=forsetiinitially because thread participants list was[galdrabok, infra]. Solved by starting a new thread withforsetias the direct recipient. Lesson: when the bus auto-resolves a sender handle that doesn't match the actual dev role, start a fresh thread rather than fighting the participant list. Archived 2026-05-31. -
[2026-05-13]Initial Voxtral default voicealloy(OpenAI-compat naming) — vLLM-Omni serving Voxtral does NOT translate aliases. Native presets are<register>_<gender>shape (neutral_female,casual_male, etc.). Always live-probe/v1/audio/voicesfor the exact wrapper-deployed preset names before setting a catalog default. Same caveat for Qwen3-TTS (wrapper exposes 15 voices: 9 Qwen presets + 6 OpenAI aliases) and Kyutai-TTS (NillPointer wrapper has NO voice-listing endpoint at all; voices are filesystem paths under thekyutai/tts-voicesHF repo). Archived 2026-05-31. -
[2026-05-17]--task classifyfor Skywork in vLLM 0.19.1 — flag was deprecated. Use--runner pooling; the model'sarchitectures: [LlamaForSequenceClassification]in config.json drives the classification head. Surfaced asvllm: error: unrecognized arguments: --task classifyin container logs. Archived 2026-05-31. -
[2026-05-17]Trusting that.envedit alone propagates a new env var into a worldtree container —compose.yaml's&worldtree-envanchor must explicitly declare the passthrough or the value silently doesn't land. Same footgun bitZ_AI_API_KEY(2026-05-12) ANDBIFROST_CLIENT_ALLOWED_HOSTS(2026-05-17). Cost ~10 min of "why is env empty?" diagnosis each time. Worldtree-side fix invh/worldtree@08f02b2. Archived 2026-05-31. -
[2026-05-17]--force-recreate --pull neverfrom the docker:cli sandbox without explicit-e WORLDTREE_IMAGE=<sha>re-pins the container to:latest, even when a newer SHA-tagged image is on disk. Symptom: container "recreated" but actually reverted to a stale image. Pass-e WORLDTREE_IMAGE=...:<sha>to the docker run invocation. Worldtree-dev's8ef3801health-gated:latestadvance is the long-term fix. Archived 2026-05-31. -
[2026-05-18]Volva env.sh.template$HOMEin commented examples — systemd'sEnvironmentFile=parser doesn't expand$HOME; uncommenting lands the literal$HOME/...string. Volva-dev'sf4dda73swapped to/home/<svc-user>/...placeholders. Archived 2026-05-31. -
[2026-05-18]Initial Volva systemd unit'sProtectHome=read-onlywithoutReadWritePaths=— althing-cli's SQLite (~/.althing/ althing.db) and codex's session state (~/.codex/) both need to write. Container started but every poll failed with "db path not writable". Surgical fix:ReadWritePaths=/home/lkraven/.althing /home/lkraven/.codex(preserves the hardening intent, only carves out the specific dirs). Archived 2026-05-31. -
[2026-05-18]Trusting that env.sh'sexport VOLVA_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"template line works under systemd —EnvironmentFile=parser aborts on the first unparseable line (command substitution), andVOLVA_ALTHING_CLIdeclared below silently never lands. Symptom:Environment=property empty, daemon error "althing-cli not found at 'althing-cli'". Fix: replace command-substitution with literal path. Volva-dev'sd436c3cdropped VOLVA_ROOT entirely upstream. Archived 2026-05-31. -
[2026-05-19]Naivedocker rmi worldtree:<old-sha> --forcefor CD SHA cleanup — would untag pinned/personal worldtree images since all three deployments share corviduo-dev. Usedocker image prune -a --filter "until=Xh"instead — respects in-use semantic (Docker won't remove an image referenced by any container on the host), so pinned/personal protected automatically. Archived 2026-05-31. -
[2026-05-19]Skaldsong CD first attempt:docker pullstep failed with 401 unauthorized. ana-docker had nodocker loginforgitea.phasefinal.com. My playbook prereq note ("docker login has been done at least once") was an unverified assumption. One-time manual login persists in~/.docker/config.json; architectural fix (workflow-sidessh ana-docker 'docker login ...'step usingREGISTRY_USER/REGISTRY_TOKENsecrets) flagged as v2. Archived 2026-05-31. -
[2026-05-19]SKALDSONG_HOST_CORS_ORIGINS=http://10.250.50.70:8300as a bare URL — pydantic-settings parses complex env vars viajson.loads(); first-boot crashloop withSettingsError: error parsing value for field "cors_origins". Must be JSON array literal:SKALDSONG_HOST_CORS_ORIGINS=["http://..."]. Archived 2026-05-31. -
[2026-05-19]SKALDSONG_HOST_STATIC_ASSETS_PATH=/app/web/distin compose — mismatched Dockerfile reality. The Dockerfile COPYs SvelteKit build output flat into/app/spa(not/app/spa/dist). Lifted the path from skaldsong-dev's CD-ask message ("/app/web/dist") rather than verifying against the actual Dockerfile they shipped. Lesson: when encoding container-internal paths in compose, verify against the Dockerfile, not the design-doc. Archived 2026-05-31. -
[2026-05-19]Playbook verify stepdocker ps | grep healthyracing the container'sstart_period(30s in compose's healthcheck). Verify ran 0.09s aftercompose up -d --force-recreate— well before docker's healthcheck could flip the status from(health: starting)to(healthy). False-negative; container was operationally up (the earlier/healthpoll verify already confirmed). Fix: grep^Upnothealthy. /health-200 IS the liveness check; docker's(healthy)is just a delayed echo. Archived 2026-05-31. -
[2026-05-20]SKALDSONG_DB_PATH+SKALDSONG_RUNS_DIRin compose env block — names skaldsong's app doesn't read. App readsSKALDSONG_HOST_SQLITE_PATH+SKALDSONG_HOST_RUNS_ROOT(per Dockerfile ENV defaults). Wrong names = silently no-op; app fell back to Dockerfile defaults pointing at/app/data/...which the compose's bind mount did NOT cover (target was/app/state/...). Result: every--force-recreatewiped the SQLite DB. Caught by skaldsong-dev (althing thread01KS4DPF6SXTBP4Q360JZVWPNT). Fix in52e98fa. Lesson: verify env var NAMES against the Dockerfile/app, not against design-doc shorthand. Archived 2026-05-31. -
[2026-05-25]First selene-block patch put the block undersaga_allowed_models:instead of top-levelmodels:— usedtext.replace("models:\n", ...)which substring-matched thesaga_allowed_models:\nline first. Caused YAML parse error. Fix: anchored regexre.compile(r"^models:\n", re.MULTILINE). Pattern: substring replace on YAML top-level keys WILL match suffix-containing keys. Archived 2026-05-31. -
[2026-05-27]docker compose up -dinside thedocker:clisandbox:${VAR:-./config}defaults resolve./configto the sandbox CWD, but the Docker daemon interprets the path against the HOST filesystem → auto-creates an empty dir → entrypoint reseeded image-baked defaults (lost host-side providers.yaml patches). Fix: pass-e WORLDTREE_CONFIG_DIR=/abs/path. Folded into the docker-as-root convention note. Archived 2026-06-01. -
[2026-05-27]:latest-pinned compose + private gitea registry + sandboxed pull = recreate on ancient cached:latest(deploy pulls by SHA so the tag never advances; sandbox can't pull). Fix: retag SHA→:lateston host, then--pull never. Better: pin SHA in.env, advance in CI. Archived 2026-06-01. -
[2026-05-27]Container recreate during in-flight skaldsong gen kills the runner. With deploys every ~10min and stories >5min, structural not incidental. Roadmap (skaldsong-dev): pre-shutdown signal handler, per-scene resume-from-checkpoint, /api/admin/quiesce. None shipped. Archived 2026-06-01. -
[2026-05-27]--reasoning-format gemmaon artemis-31b-v1i — unsupported in the deployed llama.cpp (accepts none|deepseek|deepseek-legacy).deepseekpopulates thereasoning_contentSSE delta Worldtree GemmaProvider checks. Archived 2026-06-01. -
[2026-05-27]head -c Npiped after a streaming curl SIGPIPEs the curl, killing the request early. Use file-write + separate read. Archived 2026-06-01. -
[2026-05-31]Building the dia2-capable image surfaced THREE upstream packaging quirks: (1)pip install -e nari-labs/dia2fails — no PEP 660build_editablehook; (2) plainpip installbuilds an emptyUNKNOWN-0.0.0wheel (base setuptools 59.6 < dia2's required ≥70); (3)--no-depsleavestransformers/sphn/whisper-timestampedmissing. Fix (local/dia:v2): copy the pure-pythondia2/package into site-packages + install ONLY those 3 deps; base torch/numpy already satisfy Dia2. Archived 2026-06-03. -
[2026-05-31]Dia2 predefined voices (43, baked at/app/voices) are NOT reachable from the/ttsclone path — it resolvesclone_reference_filenameagainst the reference_audio dir ONLY. The OpenAI/v1/audio/speechvoiceparam auto-resolves them (separate code path), which masked the gap. Fix: stage/app/voices/*into/worktank/dia/reference_audio. Lesson: verify on the endpoint the catalog ACTUALLY targets. Archived 2026-06-03. -
[2026-05-31]voice_mode=clonewith an emptyclone_reference_filename→ asset-engine serializes it as the literal string"undefined"→/tts404. First observed on dia2; worked around in the catalog (default the field to a real voice). [2026-06-01] root cause found — the Kokoro voice-blend widget reading Shoelace.valuebefore hydration (see Current state); the real fix is asset-engine-side and is escalated. Archived 2026-06-03. -
[2026-05-31]asset-engineservices.schema.jsonis DERIVED (regen from the Pydantic model viadump_schema.py) and had DRIFTED — rejected thelifecyclefield 12/14 services use. RESOLVED: regenerated withCatalogLifecycle+reproducibility.seed_field(b7b2130). Lesson: hand-editingservices.yamlshape without regenerating re-introduces drift. Archived 2026-06-03. -
[2026-05-31]ttyd-over-TLS forces HTTP/2 (kills ttyd's terminal WebSocket → blank screen); Safari/WebKit never sends HTTP basic-auth on WS upgrades. Both solved for the fleet seat (Caddy forces HTTP/1.1; auth → network-gating) — detail in auto-memoryreference_ttyd_fleet_seat. Archived 2026-06-03. -
[2026-05-30]esh-docker-vm:hardNFS mounts from 10.0.50.50 froze a container worker in UNKILLABLE D-state when the NAS stalled — only a host reboot clears it. Separately,fstab defaults(no_netdev) made NFS-bind containersExited(255)on reboot. → auto-memoryincident_esh_docker_nfs_boot_race. Archived 2026-06-03.