← Research atlas
JEV × QWEN · PUBLIC-SOURCE EXPLORATION

One answer changes.
Does the decision follow?

Inspect 12 pairs from ShARC with Jev 1.13.0 and Qwen3.5-4B. Within each pair, one answer in the conversation history changes while the rule, question and scenario stay fixed. Some source decisions change; others stay the same.

12 source trees24 inputs2 option orders0 project human reviews
Exploratory source agreement, not verified truth. Public dataset labels are the comparison reference. The Jev native-choice result follows two disclosed interface repairs. Qwen outputs stopped at the cap remain invalid; reasoning text is not substituted for a final decision. These conditions do not establish which model architecture is better. An audit of all 12 source pairs flags a title-only rule and other evidence-sufficiency concerns without changing the reference labels.

Can one prefill replace a generated decision?

The new Qwen letter readout reaches 13/24 and 10/24 source agreement, versus 11/24 and 12/24 for its earlier JSON generation. It uses one forward per decision and zero generated tokens. This is an outcome-aware change to prompt, verbalizer and readout together, on the same inspected inputs. One exact logit tie remains invalid. Frozen control and full results.

Every condition, including failed outputs

Repeated views of the same 24 inputs. Neither option orders nor alternative API readouts add independent samples.

ConditionFinishedItem agreementBoth pair membersInvalid

How much does a simple shortcut explain?

These five controls use no model. They do not interpret natural-language rules; compare them with the model rows before drawing a capability conclusion.

ControlItem agreementBoth pair members

Rule shared by both inputs

How to read a contrast

The highlighted history row is the only changed visible field. Green decisions match the source; rust-colored decisions differ. Neither color certifies correctness. “ASK” represents a follow-up question in the source; it is not an invalid output. A missing or truncated model answer is shown separately. The source includes no Irrelevant reference in this selected cohort.

Jev native choice and displayed-probability argmax use exactly the same API response. They are shown together to expose the readout disagreement, not to select whichever answer matches the source. For Qwen, the fixed direct and thinking caps are 256 and 2,048 generated tokens respectively.

All 12 pairs appear in frozen reference order, without sorting by model performance. Source labels and the full source records were kept out of the model state: only rule snippet, question, scenario and conversation history were supplied.