Your turn first.
For each message, decide whether to KEEP or CANCEL the target. Only the mobile plan matters; phone insurance is a distractor. Your selections stay in this browser and are not collected.
0 of 4 decisions selected
The model cells show the primary round, both A/B candidate mappings for the same four texts. A slash separates cancel-first and keep-first. These are saved historical outputs, not predictions generated by your clicks.
The twist: a grammar-specific Python reference also got 8/8 on S01 and 12/12 complete cases across the full probe. Its four S01 outputs count under both candidate mappings only for a matched denominator. Its rules were frozen before live calls, but know the construction grammar. It does not establish a general customer-message parser—or that Jev is necessary for this task. Inspect the code.
Across the full predefined 12-case probe, Jev completed 12/12 cases and this frozen 4B completed 8/12. S01 is one illustrated case, not a new benchmark or a measure of current models. The repeat gave the same S01 counts; use the full explorer for every case, mapping and round.
A later Kev-LoRA line reader made fewer wrong actions but far more needless deferrals. Compare it with always-defer and grammar-specific code, then change which deferrals cost effort. The outcome depends on the cost assumption. This is a separate development study with no new Jev results.
Can 548 calls beat always deferring? →A subsequent joint-routing pilot assigned 60 of 66 other-request lines to target fields and failed its fixed screen. Inspect every saved query in the completed follow-up →
Target-switch inputs hold every record fixed and change only the requested ID. Same records, different request: explore all 24 pairs → This input explorer shows program references, with no model predictions. The local diagnostic is now complete: read all results, failure cases and limitations →
This is original synthetic English text with program-derived, AI-reviewed reference actions and no independent human label audit. The page embeds 32 saved S01 model decisions across two backends, two mappings and two rounds; four code-reference decisions are recomputed from the declared grammar. Copying its short prompt into another model is an informal test with a different interface; it cannot be compared numerically to the frozen run without matching that protocol. The page makes no API call and sends no choices anywhere.