Compare three historical paths through one pinned Kev-LoRA N1 model and two offline code controls on 72 synthetic views. The controls are newly derived from the same visible texts; no new model ran. Reference actions are program-derived with no independent human label audit. Interpreting INSUFFICIENT as a deferral is a hypothetical operational choice.
168 forwards; 63,467 input tokens on the original development texts.
548 forwards; 133,977 input tokens on those same texts. Its prewritten screening rule failed.
Choose a hypothetical cost ratio
Correct uncertainty can still consume a human's time. Choose whether to charge only avoidable deferrals or every deferral, then assign a hypothetical wrong-action cost. Zero ignores wrong actions; it is a mathematical edge case.
Exact cost ranges, including always-defer
| Wrong-action cost ratio | Lowest loss among model paths and always-defer |
|---|
Neighboring paths tie at the shared interval boundaries. The grammar-specific control is shown in the chart above and evaluated separately from this four-path table.
What the counts mean
A wrong action is any non-INSUFFICIENT prediction that disagrees with the program-derived reference. That includes committing on an uncertain input and choosing the opposite action on a determined input. A needless deferral is INSUFFICIENT when the reference has a determined action. Every one of the 72 views belongs to exactly one of these errors or to correct decisions.
| Comparator | Correct /72 | Wrong actions | Needless deferrals | All deferrals | Model forwards | Model input tokens | Summed model forward time |
|---|
The pairwise crossover is not the optimum
With needless deferrals charged, whole-state loss is 20r + 8, line loss 6r + 37, and always-defer loss 48. The line/whole-state crossover remains r = 29/14 ≈ 2.07. However, for r ≤ 2 the line method is at least one loss unit above whole-state; for r ≥ 2 it is at least one unit above always-defer. It never minimizes this loss for any nonnegative ratio.
Charge every deferral and the intercepts change to 18, 55 and 72. The line method is the unique minimum among those generic paths for 37/14 < r < 17/6, approximately 2.64–2.83, and ties at the endpoints. The declared-grammar parser has no outcome errors and 24 correct INSUFFICIENT outputs, so its loss is 0 in the first mode and 24 in the second. It beats the line method throughout that narrow interval. Its advantage is specific to this known grammar.
The previous page compared only the three model paths and emphasized their pairwise crossover. Adding the two zero-model-call controls changes that interpretation; the saved predictions and the prewritten screening rule are unchanged.
The exact intervals are computed with rational arithmetic from saved counts, not a slider grid. These post-hoc utility calculations do not establish a new inference algorithm.