← Back to the four-line Jev challenge
A separate Kev-LoRA development study · no Jev run here

Can 548 calls beat
always deferring?

The line reader made 14 fewer wrong actions than a whole-state reader, with 29 more needless deferrals. Add an always-defer baseline and change which deferrals cost effort: a pairwise win can disappear.

The missing comparator0 calls

Always choose INSUFFICIENT: zero wrong actions, 48 needless deferrals, 72 total deferrals on these same synthetic views.

Compare three historical paths through one pinned Kev-LoRA N1 model and two offline code controls on 72 synthetic views. The controls are newly derived from the same visible texts; no new model ran. Reference actions are program-derived with no independent human label audit. Interpreting INSUFFICIENT as a deferral is a hypothetical operational choice.

Whole-state facts + code
20 wrong actions · 8 needless deferrals

168 forwards; 63,467 input tokens on the original development texts.

Line-by-line facts + code
6 wrong actions · 37 needless deferrals

548 forwards; 133,977 input tokens on those same texts. Its prewritten screening rule failed.

Choose a hypothetical cost ratio

Correct uncertainty can still consume a human's time. Choose whether to charge only avoidable deferrals or every deferral, then assign a hypothetical wrong-action cost. Zero ignores wrong actions; it is a mathematical edge case.

Which deferrals cost one unit?
0×2×4×6×8×10×

Exact cost ranges, including always-defer

Wrong-action cost ratioLowest loss among model paths and always-defer

Neighboring paths tie at the shared interval boundaries. The grammar-specific control is shown in the chart above and evaluated separately from this four-path table.

What the counts mean

A wrong action is any non-INSUFFICIENT prediction that disagrees with the program-derived reference. That includes committing on an uncertain input and choosing the opposite action on a determined input. A needless deferral is INSUFFICIENT when the reference has a determined action. Every one of the 72 views belongs to exactly one of these errors or to correct decisions.

ComparatorCorrect /72Wrong actionsNeedless deferralsAll deferralsModel forwardsModel input tokensSummed model forward time
The 72 views come from 12 six-view synthetic parent groups. Prompts and budgets differ. Model forward time sums individual calls; it is not end-to-end latency. Zero model calls for the controls does not mean measured zero CPU runtime; their CPU time was not benchmarked. Monetary cost, actual human handling quality, harm weights and future generalization are unmeasured. Computation is excluded from both loss formulas.

The pairwise crossover is not the optimum

With needless deferrals charged, whole-state loss is 20r + 8, line loss 6r + 37, and always-defer loss 48. The line/whole-state crossover remains r = 29/14 ≈ 2.07. However, for r ≤ 2 the line method is at least one loss unit above whole-state; for r ≥ 2 it is at least one unit above always-defer. It never minimizes this loss for any nonnegative ratio.

Charge every deferral and the intercepts change to 18, 55 and 72. The line method is the unique minimum among those generic paths for 37/14 < r < 17/6, approximately 2.64–2.83, and ties at the endpoints. The declared-grammar parser has no outcome errors and 24 correct INSUFFICIENT outputs, so its loss is 0 in the first mode and 24 in the second. It beats the line method throughout that narrow interval. Its advantage is specific to this known grammar.

The previous page compared only the three model paths and emphasized their pairwise crossover. Adding the two zero-model-call controls changes that interpretation; the saved predictions and the prewritten screening rule are unchanged.

The exact intervals are computed with rational arithmetic from saved counts, not a slider grid. These post-hoc utility calculations do not establish a new inference algorithm.