one_minus_min_top2_margin is the highest observed test ranking result across candidate signals.
Full100 exploratory test interpretation
BizHallu: methods and evidence limits
The historical token-signal comparison is exploratory. B1 adds simple reference predictions and paired uncertainty; B2 checks a dev-annotation omission. Candidate winners were selected after test comparison, so this is not a confirmatory model-selection result.
Historical B1 reference checks
103 test spans from 18 questions, with AI-assisted provisional labels. AP is tied-score-aware average precision, not trapezoidal PR area. Scores use saved-trace precision; historical files are unchanged.
| Signal / reference | AP | F1 | Balanced accuracy | MCC |
|---|---|---|---|---|
| Top-2 margin | 0.835 | 0.752 | 0.753 | 0.498 |
| Token entropy | 0.809 | 0.779 | 0.673 | 0.381 |
| Flag every span | 0.592 | 0.744 | 0.500 | 0.000 |
| Flag no spans | 0.592 | 0.000 | 0.500 | 0.000 |
| Dev fact-type prior (composition control) | 0.760 | 0.836 | 0.714 | 0.555 |
Entropy minus flag-every-span F1: +0.0355; exploratory paired question-bootstrap 95% interval [-0.0356, 0.1047]. This interval crosses zero; it is conditional on fixed dev thresholds and provisional labels, not adjusted for test-based signal selection or all shared-period dependence.
The fact-type prior is an annotation-composition control, not an information-matched detector: supplied type names can contain correctness hints. Its higher F1 is not a new headline result. The two overlapping-period components are too few for reliable cluster inference.
Same-step selected energy gap equals token NLL. Probability outside the top two choices is a concentration control, not an independent replication of Spilled Energy. No superiority claim follows from the historical 0.835 / 0.779 maxima.
September 5 update. B2 sensitivity adds 9 assistant-provisional q_0048 dev atoms. Dev-only refitting changes entropy F1 from 0.779 to 0.752 on the same 103 old test spans; top-2 margin is unchanged. Original B1 results remain intact. This is not fresh held-out evidence or independent review. B2 methods appendix.
B2: one omitted dev answer changes some thresholds
September 5, 2026. Separate retrospective sensitivity, not a replacement benchmark. The original 205 provisional spans and all historical scores remain unchanged. Nine assistant-provisional, correct-key-fact atoms from q_0048 extend dev from 17 questions / 102 spans to 18 questions / 111 spans. The augmented package has 214 spans; the same 18 test questions / 103 test spans are reused.
All 12 existing signals and three references use the unchanged B1 dev-only fitting policy. The selected rows below explain the existing comparison, not a new winner. Both precision arms and all candidates are retained in the expanded table.
| Signal / reference | Threshold Before / after | Test F1 Before / after | Test AP Before / after | Changed test flags |
|---|---|---|---|---|
| Mean entropy | 0.00481712 0.0818659 | 0.7794 0.7520 | 0.8095 0.8095 | 11 |
| Minimum top-2 margin risk | 0.334333 0.334333 | 0.7523 0.7523 | 0.8351 0.8351 | 0 |
| Flag all spans | 0.5 0.5 | 0.7439 0.7439 | 0.5922 0.5922 | 0 |
| Dev fact-type prior (composition control) | 0.575 0.5 | 0.8356 0.8143 | 0.7598 0.7389 | 6 |
Displayed values are rounded; fitting uses saved-trace precision. Entropy changes 11 test decisions: TP / FP / TN / FN moves from 53 / 22 / 20 / 8 to 47 / 17 / 25 / 14. That removes 5 false alarms while missing 6 more provisionally labeled errors. Top-2 margin is unchanged in this sensitivity; this is not evidence of general stability.
AP for fixed internal scores cannot change merely because a threshold changes. The dev-fitted type prior is refitted, so its ranking can change. Its supplied fact types contain correctness hints; it is not an information-matched independent detector.
All candidates and both precision arms
stored_six_decimal_scores
| Signal / reference | Threshold Before / after | Test F1 Before / after | Test AP Before / after | Changed test flags |
|---|---|---|---|---|
| mean_token_nll | 0.037151 0.044049 | 0.7414 0.7434 | 0.8320 0.8320 | 3 |
| Mean entropy | 0.004817 0.081866 | 0.7794 0.7520 | 0.8095 0.8095 | 11 |
| max_token_entropy | 0.466101 0.687115 | 0.7458 0.7222 | 0.8053 0.8053 | 10 |
| one_minus_mean_top2_margin | 0.061265 0.061265 | 0.7193 0.7193 | 0.8240 0.8240 | 0 |
| Minimum top-2 margin risk | 0.334333 0.334333 | 0.7523 0.7523 | 0.8351 0.8351 | 0 |
| mean_spilled_energy_abs_delta | 0.793505 0.793505 | 0.7375 0.7375 | 0.5877 0.5877 | 0 |
| max_spilled_energy_abs_delta | 0.793505 0.793505 | 0.7375 0.7375 | 0.5481 0.5481 | 0 |
| mean_spilled_energy_delta | -2.39671 -2.39671 | 0.7439 0.7439 | 0.7190 0.7190 | 0 |
| negative_mean_spilled_energy_delta | -5.59079 -5.59079 | 0.7439 0.7439 | 0.5757 0.5757 | 0 |
| mean_spilled_probability_mass_after_top1 | 0.038461 0.040476 | 0.7321 0.7321 | 0.8212 0.8212 | 0 |
| mean_spilled_probability_mass_after_top2 | 0.000225 0.000225 | 0.7727 0.7727 | 0.8080 0.8080 | 0 |
| max_selected_step_energy_gap | 0.195013 0.195013 | 0.7387 0.7387 | 0.8305 0.8305 | 0 |
| Flag all spans | 0.5 0.5 | 0.7439 0.7439 | 0.5922 0.5922 | 0 |
| all_negative | 0.5 0.5 | 0.0000 0.0000 | 0.5922 0.5922 | 0 |
| Dev fact-type prior (composition control) | 0.575 0.5 | 0.8356 0.8143 | 0.7598 0.7389 | 6 |
saved_trace_precision
| Signal / reference | Threshold Before / after | Test F1 Before / after | Test AP Before / after | Changed test flags |
|---|---|---|---|---|
| mean_token_nll | 0.0371515 0.0440487 | 0.7414 0.7434 | 0.8321 0.8321 | 3 |
| Mean entropy | 0.00481712 0.0818659 | 0.7794 0.7520 | 0.8095 0.8095 | 11 |
| max_token_entropy | 0.466101 0.687115 | 0.7458 0.7222 | 0.8053 0.8053 | 10 |
| one_minus_mean_top2_margin | 0.0612654 0.0612654 | 0.7193 0.7193 | 0.8239 0.8239 | 0 |
| Minimum top-2 margin risk | 0.334333 0.334333 | 0.7523 0.7523 | 0.8351 0.8351 | 0 |
| mean_spilled_energy_abs_delta | 0.793505 0.793505 | 0.7375 0.7375 | 0.5877 0.5877 | 0 |
| max_spilled_energy_abs_delta | 0.793505 0.793505 | 0.7375 0.7375 | 0.5481 0.5481 | 0 |
| mean_spilled_energy_delta | -2.39671 -2.39671 | 0.7439 0.7439 | 0.7190 0.7190 | 0 |
| negative_mean_spilled_energy_delta | -5.59079 -5.59079 | 0.7439 0.7439 | 0.5757 0.5757 | 0 |
| mean_spilled_probability_mass_after_top1 | 0.0384607 0.0404755 | 0.7321 0.7321 | 0.8213 0.8213 | 0 |
| mean_spilled_probability_mass_after_top2 | 0.000224908 0.000224908 | 0.7727 0.7727 | 0.8079 0.8079 | 0 |
| max_selected_step_energy_gap | 0.195013 0.195013 | 0.7387 0.7387 | 0.8305 0.8305 | 0 |
| Flag all spans | 0.5 0.5 | 0.7439 0.7439 | 0.5922 0.5922 | 0 |
| all_negative | 0.5 0.5 | 0.0000 0.0000 | 0.5922 0.5922 | 0 |
| Dev fact-type prior (composition control) | 0.575 0.5 | 0.8356 0.8143 | 0.7598 0.7389 | 6 |
q_0048 evidence, answer and nine provisional boundaries
The source values are Netherlands GBP 39,655.81 and EIRE GBP 12,147.92. Their difference is GBP 27,507.89. Approximately GBP 27,508 is correct whole-pound rounding, not exact pence equality. These are net transaction values under the historical ledger scope.
In August 2011, the Netherlands generated more net revenue (39,655.81 GBP) compared to EIRE (12,147.92 GBP). The Netherlands' net revenue was higher by approximately 27,508 GBP.
| Text | Fact type | Character range |
|---|---|---|
| August 2011 | month | [3, 14) |
| Netherlands | country | [20, 31) |
| generated more net revenue | comparison_direction | [32, 58) |
| 39,655.81 GBP | currency_amount | [60, 73) |
| EIRE | country | [87, 91) |
| 12,147.92 GBP | currency_amount | [93, 106) |
| Netherlands | country | [113, 124) |
| higher | comparison_direction | [142, 148) |
| 27,508 GBP | currency_amount | [166, 176) |
The sequence and recipe were saved before this supplement's token scores were read; prior familiarity with the historical study remains. This is not independently blind annotation or prospective preregistration. Repeated comparison and country facts are correlated.
Limits: no independent human review, new model run, confirmation access, new uncertainty interval or superiority claim. The all-correct single-answer supplement supports flag counts, not an estimate of hallucination recall. Original test reuse cannot establish generalization. Owner calibration remains separate and incomplete.
Two different correctness questions
In q_0064 and q_0069, the quoted product-amount pairs match their source rows, but some stated ranks do not. These are evidence-binding errors, not proof that the amounts were fabricated.
A list marker is generated before the following product and amount. Its low token uncertainty does not establish that the model was confident about the completed business relationship. Raw teacher-forced token scores and completed-answer evidence checks have different information budgets.
Positive and negative transaction values are sign-based accounting quantities. Negative value is not a verified measure of physical returns; fees and adjustments may be included. Merchandise scope uses a stock-code heuristic, and the historical product grain is stock code plus description.
Only q_0048 is absent from the original dev annotations; B2 adds it as a separate assistant-provisional sensitivity. All 18 historical test questions are covered. The 15 selected presentation spans are not an independently reviewed replacement for the 205 provisional labels.
Historical candidate maxima
Preserved for provenance, not a superiority claim.
mean_token_entropy is a different winning signal with a dev-selected threshold.
Probability mass outside the top two choices is the strongest energy-family row.
Test false positives and false negatives for two retrospectively selected signals.
How to read it
AUPRC is ranking quality; F1 is a threshold tradeoff.
AUPRC story
one_minus_min_top2_margin reaches AUPRC 0.835. It is precise at the selected threshold, but misses 20 positive spans.
F1 story
mean_token_entropy reaches F1 0.779. It catches more positives, but specificity drops to 0.476.
Energy story
The best energy-family F1 is 0.773, but it comes from probability mass outside top two choices, not pure Spilled Energy.
Some token signals rank the provisionally labeled spans usefully in this retrospective sample. This does not establish a reliable deployment threshold, stable superiority over controls, or a relation-level confidence estimate for a complete business answer.
Error review
The two selected baselines fail in different ways.
Simple best-AUPRC: one_minus_min_top2_margin
This retrospectively selected row is more selective: 7 false positives and 20 false negatives on test spans.
- Top false-negative fact typecurrency_amount: 9
- Top false-negative question typetop3_products_month: 6
- Precision / recall0.854 / 0.672
Energy best-F1: mean_spilled_probability_mass_after_top2
This retrospectively selected row catches more positives: 20 false positives and 10 false negatives on test spans.
- Top false-positive fact typecurrency_amount: 8
- Top false-positive question typemonthly_revenue_change: 6
- Precision / recall0.718 / 0.836
Grouped misses
Where the detectors need stronger business context.
Simple false negatives by fact type
- currency_amount9
- comparison_direction3
- country3
- ranking3
Energy false positives by fact type
- currency_amount8
- month4
- comparison_direction2
- percentage2
Main takeaway
This is a credible negative-plus-positive result.
Research direction
Define the relationship before comparing detectors.
A useful next comparison needs a clear relation annotation unit, independent judgments and matched information availability. Fifteen selected presentation spans received additional assistant review; the original 205-span package remains AI-assisted and provisional, with no independent human annotation or inter-annotator agreement.
Return to the April example ยท Inspect historical error examples
Historical annotation record
Selected-span assistant review packet. This earlier record is retained for provenance and is not independent human annotation.