BHBizHallu

Full100 exploratory test interpretation

BizHallu: methods and evidence limits

The historical token-signal comparison is exploratory. B1 adds simple reference predictions and paired uncertainty; B2 checks a dev-annotation omission. Candidate winners were selected after test comparison, so this is not a confirmatory model-selection result.

Historical B1 reference checks

103 test spans from 18 questions, with AI-assisted provisional labels. AP is tied-score-aware average precision, not trapezoidal PR area. Scores use saved-trace precision; historical files are unchanged.

Signal / referenceAPF1Balanced accuracyMCC
Top-2 margin0.8350.7520.7530.498
Token entropy0.8090.7790.6730.381
Flag every span0.5920.7440.5000.000
Flag no spans0.5920.0000.5000.000
Dev fact-type prior (composition control)0.7600.8360.7140.555

Entropy minus flag-every-span F1: +0.0355; exploratory paired question-bootstrap 95% interval [-0.0356, 0.1047]. This interval crosses zero; it is conditional on fixed dev thresholds and provisional labels, not adjusted for test-based signal selection or all shared-period dependence.

The fact-type prior is an annotation-composition control, not an information-matched detector: supplied type names can contain correctness hints. Its higher F1 is not a new headline result. The two overlapping-period components are too few for reliable cluster inference.

Same-step selected energy gap equals token NLL. Probability outside the top two choices is a concentration control, not an independent replication of Spilled Energy. No superiority claim follows from the historical 0.835 / 0.779 maxima.

September 5 update. B2 sensitivity adds 9 assistant-provisional q_0048 dev atoms. Dev-only refitting changes entropy F1 from 0.779 to 0.752 on the same 103 old test spans; top-2 margin is unchanged. Original B1 results remain intact. This is not fresh held-out evidence or independent review. B2 methods appendix.

B2: one omitted dev answer changes some thresholds

September 5, 2026. Separate retrospective sensitivity, not a replacement benchmark. The original 205 provisional spans and all historical scores remain unchanged. Nine assistant-provisional, correct-key-fact atoms from q_0048 extend dev from 17 questions / 102 spans to 18 questions / 111 spans. The augmented package has 214 spans; the same 18 test questions / 103 test spans are reused.

All 12 existing signals and three references use the unchanged B1 dev-only fitting policy. The selected rows below explain the existing comparison, not a new winner. Both precision arms and all candidates are retained in the expanded table.

Signal / referenceThreshold
Before / after
Test F1
Before / after
Test AP
Before / after
Changed test flags
Mean entropy0.00481712
0.0818659
0.7794
0.7520
0.8095
0.8095
11
Minimum top-2 margin risk0.334333
0.334333
0.7523
0.7523
0.8351
0.8351
0
Flag all spans0.5
0.5
0.7439
0.7439
0.5922
0.5922
0
Dev fact-type prior (composition control)0.575
0.5
0.8356
0.8143
0.7598
0.7389
6

Displayed values are rounded; fitting uses saved-trace precision. Entropy changes 11 test decisions: TP / FP / TN / FN moves from 53 / 22 / 20 / 8 to 47 / 17 / 25 / 14. That removes 5 false alarms while missing 6 more provisionally labeled errors. Top-2 margin is unchanged in this sensitivity; this is not evidence of general stability.

AP for fixed internal scores cannot change merely because a threshold changes. The dev-fitted type prior is refitted, so its ranking can change. Its supplied fact types contain correctness hints; it is not an information-matched independent detector.

All candidates and both precision arms

stored_six_decimal_scores

Signal / referenceThreshold
Before / after
Test F1
Before / after
Test AP
Before / after
Changed test flags
mean_token_nll0.037151
0.044049
0.7414
0.7434
0.8320
0.8320
3
Mean entropy0.004817
0.081866
0.7794
0.7520
0.8095
0.8095
11
max_token_entropy0.466101
0.687115
0.7458
0.7222
0.8053
0.8053
10
one_minus_mean_top2_margin0.061265
0.061265
0.7193
0.7193
0.8240
0.8240
0
Minimum top-2 margin risk0.334333
0.334333
0.7523
0.7523
0.8351
0.8351
0
mean_spilled_energy_abs_delta0.793505
0.793505
0.7375
0.7375
0.5877
0.5877
0
max_spilled_energy_abs_delta0.793505
0.793505
0.7375
0.7375
0.5481
0.5481
0
mean_spilled_energy_delta-2.39671
-2.39671
0.7439
0.7439
0.7190
0.7190
0
negative_mean_spilled_energy_delta-5.59079
-5.59079
0.7439
0.7439
0.5757
0.5757
0
mean_spilled_probability_mass_after_top10.038461
0.040476
0.7321
0.7321
0.8212
0.8212
0
mean_spilled_probability_mass_after_top20.000225
0.000225
0.7727
0.7727
0.8080
0.8080
0
max_selected_step_energy_gap0.195013
0.195013
0.7387
0.7387
0.8305
0.8305
0
Flag all spans0.5
0.5
0.7439
0.7439
0.5922
0.5922
0
all_negative0.5
0.5
0.0000
0.0000
0.5922
0.5922
0
Dev fact-type prior (composition control)0.575
0.5
0.8356
0.8143
0.7598
0.7389
6

saved_trace_precision

Signal / referenceThreshold
Before / after
Test F1
Before / after
Test AP
Before / after
Changed test flags
mean_token_nll0.0371515
0.0440487
0.7414
0.7434
0.8321
0.8321
3
Mean entropy0.00481712
0.0818659
0.7794
0.7520
0.8095
0.8095
11
max_token_entropy0.466101
0.687115
0.7458
0.7222
0.8053
0.8053
10
one_minus_mean_top2_margin0.0612654
0.0612654
0.7193
0.7193
0.8239
0.8239
0
Minimum top-2 margin risk0.334333
0.334333
0.7523
0.7523
0.8351
0.8351
0
mean_spilled_energy_abs_delta0.793505
0.793505
0.7375
0.7375
0.5877
0.5877
0
max_spilled_energy_abs_delta0.793505
0.793505
0.7375
0.7375
0.5481
0.5481
0
mean_spilled_energy_delta-2.39671
-2.39671
0.7439
0.7439
0.7190
0.7190
0
negative_mean_spilled_energy_delta-5.59079
-5.59079
0.7439
0.7439
0.5757
0.5757
0
mean_spilled_probability_mass_after_top10.0384607
0.0404755
0.7321
0.7321
0.8213
0.8213
0
mean_spilled_probability_mass_after_top20.000224908
0.000224908
0.7727
0.7727
0.8079
0.8079
0
max_selected_step_energy_gap0.195013
0.195013
0.7387
0.7387
0.8305
0.8305
0
Flag all spans0.5
0.5
0.7439
0.7439
0.5922
0.5922
0
all_negative0.5
0.5
0.0000
0.0000
0.5922
0.5922
0
Dev fact-type prior (composition control)0.575
0.5
0.8356
0.8143
0.7598
0.7389
6
q_0048 evidence, answer and nine provisional boundaries

The source values are Netherlands GBP 39,655.81 and EIRE GBP 12,147.92. Their difference is GBP 27,507.89. Approximately GBP 27,508 is correct whole-pound rounding, not exact pence equality. These are net transaction values under the historical ledger scope.

In August 2011, the Netherlands generated more net revenue (39,655.81 GBP) compared to EIRE (12,147.92 GBP). The Netherlands' net revenue was higher by approximately 27,508 GBP.
TextFact typeCharacter range
August 2011month[3, 14)
Netherlandscountry[20, 31)
generated more net revenuecomparison_direction[32, 58)
39,655.81 GBPcurrency_amount[60, 73)
EIREcountry[87, 91)
12,147.92 GBPcurrency_amount[93, 106)
Netherlandscountry[113, 124)
highercomparison_direction[142, 148)
27,508 GBPcurrency_amount[166, 176)

The sequence and recipe were saved before this supplement's token scores were read; prior familiarity with the historical study remains. This is not independently blind annotation or prospective preregistration. Repeated comparison and country facts are correlated.

Limits: no independent human review, new model run, confirmation access, new uncertainty interval or superiority claim. The all-correct single-answer supplement supports flag counts, not an estimate of hallucination recall. Original test reuse cannot establish generalization. Owner calibration remains separate and incomplete.

Two different correctness questions

In q_0064 and q_0069, the quoted product-amount pairs match their source rows, but some stated ranks do not. These are evidence-binding errors, not proof that the amounts were fabricated.

A list marker is generated before the following product and amount. Its low token uncertainty does not establish that the model was confident about the completed business relationship. Raw teacher-forced token scores and completed-answer evidence checks have different information budgets.

Positive and negative transaction values are sign-based accounting quantities. Negative value is not a verified measure of physical returns; fees and adjustments may be included. Merchandise scope uses a stock-code heuristic, and the historical product grain is stock code plus description.

Only q_0048 is absent from the original dev annotations; B2 adds it as a separate assistant-provisional sensitivity. All 18 historical test questions are covered. The 15 selected presentation spans are not an independently reviewed replacement for the 205 provisional labels.

Historical candidate maxima

Preserved for provenance, not a superiority claim.

Exploratory max test AUPRC 0.835

one_minus_min_top2_margin is the highest observed test ranking result across candidate signals.

Exploratory max test F1 0.779

mean_token_entropy is a different winning signal with a dev-selected threshold.

Best energy F1 0.773

Probability mass outside the top two choices is the strongest energy-family row.

Error rows 57

Test false positives and false negatives for two retrospectively selected signals.

How to read it

AUPRC is ranking quality; F1 is a threshold tradeoff.

AUPRC story

one_minus_min_top2_margin reaches AUPRC 0.835. It is precise at the selected threshold, but misses 20 positive spans.

F1 story

mean_token_entropy reaches F1 0.779. It catches more positives, but specificity drops to 0.476.

Energy story

The best energy-family F1 is 0.773, but it comes from probability mass outside top two choices, not pure Spilled Energy.

Report wording

Some token signals rank the provisionally labeled spans usefully in this retrospective sample. This does not establish a reliable deployment threshold, stable superiority over controls, or a relation-level confidence estimate for a complete business answer.

Error review

The two selected baselines fail in different ways.

Simple best-AUPRC: one_minus_min_top2_margin

This retrospectively selected row is more selective: 7 false positives and 20 false negatives on test spans.

  • Top false-negative fact typecurrency_amount: 9
  • Top false-negative question typetop3_products_month: 6
  • Precision / recall0.854 / 0.672

Energy best-F1: mean_spilled_probability_mass_after_top2

This retrospectively selected row catches more positives: 20 false positives and 10 false negatives on test spans.

  • Top false-positive fact typecurrency_amount: 8
  • Top false-positive question typemonthly_revenue_change: 6
  • Precision / recall0.718 / 0.836

Grouped misses

Where the detectors need stronger business context.

Simple false negatives by fact type

  • currency_amount9
  • comparison_direction3
  • country3
  • ranking3

Energy false positives by fact type

  • currency_amount8
  • month4
  • comparison_direction2
  • percentage2

Main takeaway

This is a credible negative-plus-positive result.

ClaimEvidenceRiskUse in reportStatus
Simple uncertainty has real signal. AUPRC 0.835; F1 0.779 Not causal or semantic. Exploratory result with selection caveat. Exploratory
Pure Spilled Energy is not the winner. 4 all-positive-like rows flagged. Can overstate specificity. Guardrail result. Draft
Business context is still missing. Top-3 and currency misses dominate error review. Token confidence can be high for wrong evidence binding. Motivation for next method. Draft

Research direction

Define the relationship before comparing detectors.

A useful next comparison needs a clear relation annotation unit, independent judgments and matched information availability. Fifteen selected presentation spans received additional assistant review; the original 205-span package remains AI-assisted and provisional, with no independent human annotation or inter-annotator agreement.

Return to the April example ยท Inspect historical error examples

Research questions
Historical annotation record

Selected-span assistant review packet. This earlier record is retained for provenance and is not independent human annotation.