Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
63 lines
1.4 KiB
JSON
63 lines
1.4 KiB
JSON
{
|
|
"url": "http://10.251.50.54:8032",
|
|
"health": {
|
|
"status": "ok",
|
|
"semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a",
|
|
"model": {
|
|
"source": "Qwen/Qwen3.5-4B",
|
|
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
|
|
"dtype": "bfloat16",
|
|
"device": "cuda:0",
|
|
"torch_version": "2.10.0+cu128",
|
|
"transformers_version": "5.17.0",
|
|
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
|
|
"allocated_gib": 7.84,
|
|
"reserved_gib": 8.12
|
|
},
|
|
"vram_cap_gib": 12.0,
|
|
"max_tokens": 4096,
|
|
"max_decisions": 64,
|
|
"workloads": []
|
|
},
|
|
"1_parity_vs_upstream": {
|
|
"rows": 144,
|
|
"top_choice_agree": 144,
|
|
"max_abs_prob_gap": 0.049096847865749194
|
|
},
|
|
"1_prompt_sha256_equal": 144,
|
|
"2_noise_floor_a_vs_b": {
|
|
"rows": 144,
|
|
"top_choice_agree": 144,
|
|
"max_abs_prob_gap": 0.0
|
|
},
|
|
"3_negative_rotated_options": {
|
|
"rows": 144,
|
|
"top_choice_agree": 14,
|
|
"max_abs_prob_gap": 0.9985836128543327
|
|
},
|
|
"4_shared_vs_direct": {
|
|
"groups": 36,
|
|
"rows": 72,
|
|
"top_choice_agree": 72,
|
|
"max_abs_prob_gap": 0.046346781311727814
|
|
},
|
|
"5_speed_21_binary": {
|
|
"prefix_tokens": 62,
|
|
"shared_s": {
|
|
"runs": [
|
|
0.139055563005968,
|
|
0.1375489159981953,
|
|
0.13897942200128455
|
|
],
|
|
"median": 0.13897942200128455
|
|
},
|
|
"sequential_decide_s": {
|
|
"runs": [
|
|
0.9357447380025405,
|
|
0.9384197259932989,
|
|
0.939685705001466
|
|
],
|
|
"median": 0.9384197259932989
|
|
}
|
|
}
|
|
} |