surface_test.py's vision check is one image and one word. It proves the tower
loads; it does not prove the tower works. This battery uses generated images with
known ground truth so every answer is objectively gradeable.
Against orcarouter NVFP4-mixed on the `gen` alias:
T1 OCR, 5 lines incl. one at 18px PASS all 5 exact
T2 counting + attribute binding PASS 7 circles / 3 triangles / 1 square
T3 bar chart, 6 values + max/min PASS 6/6 exact
T4b occlusion, star behind rectangle PASS
T4c aspect ratio of a 160x140 rectangle FAIL called it taller than wide
T5 two images, which has text PASS
T6 four images, the seat's cap PASS all four named
T7 five images, one over the cap PASS rejected with HTTP 400
No <think> leak on any vision call.
The single miss is fine-grained relative-dimension estimation on a near-square
shape, and it reproduced across two runs (the longer T4 called the same rectangle
"equal width and height"). Counting, OCR, chart values and occlusion ordering are
all solid, so this is a precise-geometry weakness, not a broken tower. Recorded so
nobody builds a feature on this model judging relative sizes.
T7 earns its place separately: it confirms the per-prompt image cap fails loudly
with a 400 rather than silently dropping the extra image.