# Vision battery for a VL gen seat The `surface_test.py` vision check is one image and one word ("Blue"). It proves the tower loads; it does not prove the tower *works*. This battery does, against images generated with known ground truth so every answer is objectively gradeable. Generate the fixtures with the PIL snippet in this repo's history (or any images whose content you know exactly), then: uv run --with requests python vistest.py gen uv run --with requests python vis2.py ## Results — orcarouter NVFP4-mixed on the `gen` alias, 2026-08-21 | test | what it exercises | result | |---|---|---| | T1 OCR | 5 lines, mixed case, digits, punctuation, one line at 18px | **PASS** — all 5 exact | | T2 counting + attribute binding | 7 yellow circles / 3 purple triangles / 1 red square, scattered | **PASS** — 7/3/1 | | T3 chart reading | 6 labelled bars, plus highest/lowest | **PASS** — 6/6 values, max/min right | | T4b occlusion | a green star partly behind a grey rectangle | **PASS** | | T4c aspect ratio | a 160x140 rectangle: wider, taller, or square? | **FAIL** — said taller | | T5 two images | describe each, say which has text | **PASS** | | T6 four images | the seat's `--limit-mm-per-prompt` ceiling | **PASS** — all 4 named | | T7 five images | one over the cap | **PASS** — rejected with HTTP 400 | 7 of 8. No `` leak on any vision call. **The one miss is worth keeping.** T4c is fine-grained aspect-ratio estimation on a near-square shape (160x140, a 14% difference); the model called it taller, and in the longer T4 run it called the same shape "a rectangle with equal width and height". Counting, OCR down to 18px, chart values and occlusion ordering are all solid, so this is a precise-geometry weakness rather than a broken tower. **Do not build a feature on this model's estimate of relative dimensions** — ask it what shapes are present and where, not how big they are relative to each other. T7 is worth keeping for a different reason: it confirms the per-prompt image cap fails **loudly** with a 400 rather than silently dropping the fifth image.