← Original Jev scope challenge
Completed historical Kev-LoRA N1 pilot · program labels · no new Jev run

Another request.
Still counted as evidence.

Joint routing got 160 of 162 target-request lines right. It also assigned 60 of 66 other-request lines to target fields. The fixed development screen failed; all saved queries are inspectable below.

Other-request records misbound60 / 66

These are line judgments from 72 correlated synthetic views, not independent examples or a real-world error estimate.

The candidate made one exclusive request/field/value choice per evidence line, then used unchanged code to execute the policy. It reached 34/72 correct decisions, versus 31/72 for each fixed direct control. This is a post-hoc development comparison on 12 constructed parent groups, with no independent human label audit.

Both screening conditions failed. The rule required ≤7/24 false commitments and ≥34/48 correct determined decisions. Joint routing produced 11/24 and 21/48. Its current form should not advance unchanged. No rewrite, packed-line or reserved input was scored.

Every method, including the zero-call controls

PathCorrect /72False commitments /24Determined correct /48Needless deferrals /48Wrong determined actions /48CallsInput tokens

Call-matched direct uses 228 calls but 19,034 fewer input tokens than joint routing. Input-token-matched direct uses 40 fewer tokens but 58 more calls. Direct methods reuse one 286-query union; actual execution was 514 unique scientific forwards plus two warmups. The grammar-specific and always-defer controls are offline computations, not new model outputs. Their CPU time was not benchmarked.

Keep the request-scope denominators separate

160 / 162
Target-request lines assigned the exact field and value; two were assigned the wrong field.
6 / 66
Other-request lines correctly excluded; 60 were assigned to a target field.

Blue is an exact route match to the program construction record; orange is an error. This describes returned labels and their consequences, not the model's internal reasoning. There are 228 correlated line judgments. Correct field statuses were 130/168; exact fact vectors were 34/72.

Inspect every case and saved query

Overview counts above always describe all 72 views. The controls below filter only the case inspector. The initial case illustrates a different request's rejection creating a target conflict.

Exact visible state

Exact policy

Target-field statuses from joint routing

FieldModel statusProgram referenceSelected lines

A selected span identifies the chosen line; it does not certify its meaning. MISSING means no line selected, not certified absence. References below are program construction labels, not independent human annotations.

Expand a query to see its exact prompt, chosen route/action, candidate logits and probabilities. Probabilities are conditional on the listed answer slots, not calibrated correctness probabilities. Joint and direct questions have different candidate spaces. Exact token arrays remain in the raw JSONL.

What this result can support

Against call-matched direct, joint corrected 20 wrong views and regressed 17 correct views. At the 12-parent unit, four groups improved, four regressed and four tied. The net three-view gain does not establish significance or generalization. The next method must preserve request scope and determined coverage; this run does not test a remedy.

The 72 single-direct inputs, predictions, logits and candidate probabilities reproduced the earlier local N1 run exactly. Both matched direct methods predicted the same actions on all 72 views; their aggregate 31/72 matches the single-call score, but single-call predictions differ on three views: one correction, one regression and one wrong-to-wrong action change. Additional option-order calls did not raise the observed total.

Weights and all 504 adapter tensors matched the pinned historical checkpoint. This run used no new training, model download, paid API or cloud compute. The unrelated multi-fact packing audit remains software-only. The pure-code success is specific to this declared grammar. None of these results establishes whether Jev is necessary for broader tasks.

Post-hoc ID-gate replay: 70/72, two regressions, zero new callsComplete result report and runtimeAll 514 raw forwardsFrozen pre-inference protocolSeparate interface stressKev, Qwen and related-work credits