feat(litellm): map reasoning_effort high/max -> xhigh for gen-reasoning
The gen-reasoning seat accepts only xhigh/medium/low and 400s on anything else — including `high`, which is the default of several clients, so the seat presented as broken rather than as one enum value out of step. The DeepSeek Harness failed every request on its default setting, and the only working client value was `low`: the seat's WEAKEST reasoning tier, while its own default is xhigh. conf/reasoning_effort_map.py is a pre-call hook in the same shape as the existing strip_empty_tools hook. It is scoped to one model group, measured rather than assumed: gen-reasoning rejects `high`; gen, sec, char-rp-reasoning and summarizer all accept it and are left alone. Paid passthroughs were not probed, because probing them spends vendor credits, and are not mapped. Verified after deploy: high and max now succeed on gen-reasoning, low and xhigh still work, a request with no effort param still works, `gen` with `high` still passes through unmapped, and the harness completes a real file-edit task at full reasoning. Needed a compose change as well as a conf push — callbacks are bind-mounted per file, so the volume only attaches on container create. Recreated the litellm service by name so the DB was not bounced with it.
This commit is contained in:
@@ -56,6 +56,35 @@ full Langfuse traces later:
|
||||
|
||||
No re-architecture: the gateway and every consumer stay pointed here.
|
||||
|
||||
## ⚠ `reasoning_effort` is not a universal vocabulary
|
||||
|
||||
`gen-reasoning` accepts **only** `xhigh` (its default), `medium` and `low`, and
|
||||
returns HTTP 400 on anything else:
|
||||
|
||||
Unexpected reasoning effort high. Supported types are xhigh (default),
|
||||
medium, and low.
|
||||
|
||||
That is the *default* value of several clients, so the seat presents as broken
|
||||
rather than as one enum value out of step. `conf/reasoning_effort_map.py` is a
|
||||
pre-call hook that maps `high` and `max` onto `xhigh` for that model group only.
|
||||
|
||||
Measured 2026-09-02 across every local seat before scoping it:
|
||||
|
||||
| model | `reasoning_effort: high` |
|
||||
|---|---|
|
||||
| `gen-reasoning` | **rejected** → mapped |
|
||||
| `gen`, `sec`, `char-rp-reasoning`, `summarizer` | accepted → untouched |
|
||||
|
||||
Paid passthroughs (`gen-frontier*`, `glm*`, `kimi*`) were deliberately **not**
|
||||
probed — they spend vendor credits — and are not mapped. **Add a model to
|
||||
`EFFORT_MAP` only after measuring that it actually rejects the value.**
|
||||
|
||||
⚠ **A hook file needs a compose change, not just a conf push.** Callbacks are
|
||||
bind-mounted per-file beside `config.yaml`, so a new hook requires a new volume
|
||||
line and `docker compose up -d litellm` (a `restart` will not pick it up — the
|
||||
volume only attaches at container creation). Target the service by name; a bare
|
||||
`up -d` bounces the DB too.
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
|
||||
Reference in New Issue
Block a user