feat(litellm): map reasoning_effort high/max -> xhigh for gen-reasoning

The gen-reasoning seat accepts only xhigh/medium/low and 400s on anything
else — including `high`, which is the default of several clients, so the
seat presented as broken rather than as one enum value out of step. The
DeepSeek Harness failed every request on its default setting, and the only
working client value was `low`: the seat's WEAKEST reasoning tier, while
its own default is xhigh.

conf/reasoning_effort_map.py is a pre-call hook in the same shape as the
existing strip_empty_tools hook. It is scoped to one model group, measured
rather than assumed: gen-reasoning rejects `high`; gen, sec,
char-rp-reasoning and summarizer all accept it and are left alone. Paid
passthroughs were not probed, because probing them spends vendor credits,
and are not mapped.

Verified after deploy: high and max now succeed on gen-reasoning, low and
xhigh still work, a request with no effort param still works, `gen` with
`high` still passes through unmapped, and the harness completes a real
file-edit task at full reasoning.

Needed a compose change as well as a conf push — callbacks are bind-mounted
per file, so the volume only attaches on container create. Recreated the
litellm service by name so the DB was not bounced with it.
This commit is contained in:
vh
2026-09-02 15:43:22 -07:00
parent e9605df6ff
commit 926fc2fb7a
4 changed files with 104 additions and 1 deletions
+29
View File
@@ -56,6 +56,35 @@ full Langfuse traces later:
No re-architecture: the gateway and every consumer stay pointed here.
## ⚠ `reasoning_effort` is not a universal vocabulary
`gen-reasoning` accepts **only** `xhigh` (its default), `medium` and `low`, and
returns HTTP 400 on anything else:
Unexpected reasoning effort high. Supported types are xhigh (default),
medium, and low.
That is the *default* value of several clients, so the seat presents as broken
rather than as one enum value out of step. `conf/reasoning_effort_map.py` is a
pre-call hook that maps `high` and `max` onto `xhigh` for that model group only.
Measured 2026-09-02 across every local seat before scoping it:
| model | `reasoning_effort: high` |
|---|---|
| `gen-reasoning` | **rejected** → mapped |
| `gen`, `sec`, `char-rp-reasoning`, `summarizer` | accepted → untouched |
Paid passthroughs (`gen-frontier*`, `glm*`, `kimi*`) were deliberately **not**
probed — they spend vendor credits — and are not mapped. **Add a model to
`EFFORT_MAP` only after measuring that it actually rejects the value.**
⚠ **A hook file needs a compose change, not just a conf push.** Callbacks are
bind-mounted per-file beside `config.yaml`, so a new hook requires a new volume
line and `docker compose up -d litellm` (a `restart` will not pick it up — the
volume only attaches at container creation). Target the service by name; a bare
`up -d` bounces the DB too.
## Deploy
```bash