feat(intern-decision-serve): 0.1.1 adds POST /v1/systemone (Jev wire shape)

Straight passthrough to the checkpoint's own DecisionEngine.predict — never the
semif mapping, whose different prompt would change the answers. Reuses the one
inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new
concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a
prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises
the surface. Response 'model' is a string name@revision (JevBench's runner
hashes it; a dict broke its manifest step).

Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench
v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard
83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4;
controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1
per-process peak 9,866 MiB under the largest accepted request (budget 9,876);
/decide/shared positive control unchanged. 120 tests green.
This commit is contained in:
vh
2026-09-30 12:58:19 -07:00
parent 579b1f5d42
commit ff552abf1f
13 changed files with 3449 additions and 13 deletions
@@ -22,6 +22,32 @@ def score_by_description(field: str, question: dict) -> list[float]:
return [len(d) + (0.5 if i == 0 else 0.0) for i, d in enumerate(descs)]
def answer_for(field: str, question: dict, scorer) -> dict:
"""One Jev answer in the shape DecisionEngine.predict returns, for any of the three types."""
kind = question["type"]
if kind == "noul":
probs = {"yes": 0.75, "no": 0.25}
return {"type": "noul", "probabilities": probs, "confidence": 0.75, "noul": 0.75,
"source": "local", "decision": "yes"}
criteria = question["criteria"]
if isinstance(criteria, dict):
keys = list(criteria)
descriptions = list(criteria.values())
else: # score: a list of levels
keys = [str(i) for i in range(len(criteria))]
descriptions = [str(level) for level in criteria]
probs = dict(zip(keys, softmax(scorer(field, {**question, "criteria": dict(zip(keys, descriptions))}))))
best = min(keys, key=lambda k: (-probs[k], k))
answer = {"type": kind, "probabilities": probs, "confidence": probs[best], "source": "local", "decision": best}
if kind == "choice":
answer["choice"] = best
else:
answer["score"] = sum(float(k) * p for k, p in probs.items())
answer["legend"] = {k: (criteria[i] if isinstance(criteria, list) else v)
for i, (k, v) in enumerate(zip(keys, criteria.values() if isinstance(criteria, dict) else criteria))}
return answer
class FakeEngine:
def __init__(self, scorer=score_by_description, tokens_per_question: int = 100):
self.scorer = scorer
@@ -35,13 +61,8 @@ class FakeEngine:
def predict(self, request: dict) -> tuple[dict, str]:
self.calls.append(copy.deepcopy(request))
answers = {}
for field, question in request["questions"].items():
ids = list(question["criteria"])
probs = dict(zip(ids, softmax(self.scorer(field, question))))
best = min(ids, key=lambda i: (-probs[i], i))
answers[field] = {"type": "choice", "probabilities": probs, "confidence": probs[best],
"choice": best, "source": "local", "decision": best}
answers = {field: answer_for(field, question, self.scorer)
for field, question in request["questions"].items()}
response = {"answers": answers,
"usage": {"input_tokens": self.tokens_per_question * len(answers),
"output_tokens": len(answers), "decision_count": len(answers)},