The candidate made one exclusive request/field/value choice per evidence line, then used unchanged code to execute the policy. It reached 34/72 correct decisions, versus 31/72 for each fixed direct control. This is a post-hoc development comparison on 12 constructed parent groups, with no independent human label audit.
Every method, including the zero-call controls
| Path | Correct /72 | False commitments /24 | Determined correct /48 | Needless deferrals /48 | Wrong determined actions /48 | Calls | Input tokens |
|---|
Call-matched direct uses 228 calls but 19,034 fewer input tokens than joint routing. Input-token-matched direct uses 40 fewer tokens but 58 more calls. Direct methods reuse one 286-query union; actual execution was 514 unique scientific forwards plus two warmups. The grammar-specific and always-defer controls are offline computations, not new model outputs. Their CPU time was not benchmarked.
Keep the request-scope denominators separate
Blue is an exact route match to the program construction record; orange is an error. This describes returned labels and their consequences, not the model's internal reasoning. There are 228 correlated line judgments. Correct field statuses were 130/168; exact fact vectors were 34/72.
Inspect every case and saved query
Overview counts above always describe all 72 views. The controls below filter only the case inspector. The initial case illustrates a different request's rejection creating a target conflict.
Exact visible state
Exact policy
Target-field statuses from joint routing
| Field | Model status | Program reference | Selected lines |
|---|
A selected span identifies the chosen line; it does not certify its meaning. MISSING means no line selected, not certified absence. References below are program construction labels, not independent human annotations.
Expand a query to see its exact prompt, chosen route/action, candidate logits and probabilities. Probabilities are conditional on the listed answer slots, not calibrated correctness probabilities. Joint and direct questions have different candidate spaces. Exact token arrays remain in the raw JSONL.
What this result can support
Against call-matched direct, joint corrected 20 wrong views and regressed 17 correct views. At the 12-parent unit, four groups improved, four regressed and four tied. The net three-view gain does not establish significance or generalization. The next method must preserve request scope and determined coverage; this run does not test a remedy.
The 72 single-direct inputs, predictions, logits and candidate probabilities reproduced the earlier local N1 run exactly. Both matched direct methods predicted the same actions on all 72 views; their aggregate 31/72 matches the single-call score, but single-call predictions differ on three views: one correction, one regression and one wrong-to-wrong action change. Additional option-order calls did not raise the observed total.
Weights and all 504 adapter tensors matched the pinned historical checkpoint. This run used no new training, model download, paid API or cloud compute. The unrelated multi-fact packing audit remains software-only. The pure-code success is specific to this declared grammar. None of these results establishes whether Jev is necessary for broader tasks.