BHBizHallu

Interactive demo v2

BizHallu: evidence binding in business analysis

A real product and its correct revenue can still be the wrong answer to a ranking question. Two curated walkthroughs separate source-row accuracy from ranking accuracy.

Historical selected spans and detector readouts

These 15 selected spans across nine cases received additional assistant review; they are not an independent human benchmark. Detector scores apply to pre-identified spans and do not discover claims automatically. Filters below affect only this historical span table, not the complete case walkthrough.

Historical B1 reference checks

103 test spans from 18 questions, with AI-assisted provisional labels. AP is tied-score-aware average precision, not trapezoidal PR area. Scores use saved-trace precision; historical files are unchanged.

Signal / referenceAPF1Balanced accuracyMCC
Top-2 margin0.8350.7520.7530.498
Token entropy0.8090.7790.6730.381
Flag every span0.5920.7440.5000.000
Flag no spans0.5920.0000.5000.000
Dev fact-type prior (composition control)0.7600.8360.7140.555

Entropy minus flag-every-span F1: +0.0355; exploratory paired question-bootstrap 95% interval [-0.0356, 0.1047]. This interval crosses zero; it is conditional on fixed dev thresholds and provisional labels, not adjusted for test-based signal selection or all shared-period dependence.

The fact-type prior is an annotation-composition control, not an information-matched detector: supplied type names can contain correctness hints. Its higher F1 is not a new headline result. The two overlapping-period components are too few for reliable cluster inference.

Same-step selected energy gap equals token NLL. Probability outside the top two choices is a concentration control, not an independent replication of Spilled Energy. No superiority claim follows from the historical 0.835 / 0.779 maxima.

September 5 update. B2 sensitivity adds 9 assistant-provisional q_0048 dev atoms. Dev-only refitting changes entropy F1 from 0.779 to 0.752 on the same 103 old test spans; top-2 margin is unchanged. Original B1 results remain intact. This is not fresh held-out evidence or independent review. B2 methods appendix.

Historical maxima and data bundle
Open JSON data bundle