Inspect 12 pairs from ShARC with Jev 1.13.0 and Qwen3.5-4B. Within each pair, one answer in the conversation history changes while the rule, question and scenario stay fixed. Some source decisions change; others stay the same.
The new Qwen letter readout reaches 13/24 and 10/24 source agreement, versus 11/24 and 12/24 for its earlier JSON generation. It uses one forward per decision and zero generated tokens. This is an outcome-aware change to prompt, verbalizer and readout together, on the same inspected inputs. One exact logit tie remains invalid. Frozen control and full results.
Repeated views of the same 24 inputs. Neither option orders nor alternative API readouts add independent samples.
| Condition | Finished | Item agreement | Both pair members | Invalid |
|---|
These five controls use no model. They do not interpret natural-language rules; compare them with the model rows before drawing a capability conclusion.
| Control | Item agreement | Both pair members |
|---|
The highlighted history row is the only changed visible field. Green decisions match the source; rust-colored decisions differ. Neither color certifies correctness. “ASK” represents a follow-up question in the source; it is not an invalid output. A missing or truncated model answer is shown separately. The source includes no Irrelevant reference in this selected cohort.
Jev native choice and displayed-probability argmax use exactly the same API response. They are shown together to expose the readout disagreement, not to select whichever answer matches the source. For Qwen, the fixed direct and thinking caps are 256 and 2,048 generated tokens respectively.
All 12 pairs appear in frozen reference order, without sorting by model performance. Source labels and the full source records were kept out of the model state: only rule snippet, question, scenario and conversation history were supplied.