BizHallu Portfolio Demo: historical case readout
Two historical cases and 7 selected spans with additional assistant review. This page preserves the earlier span-level readout; Demo v2 is the primary interactive entry.
q_0064: product amounts and ranking relations
What were the top 3 products by merchandise net revenue in April 2011? Give the rank, stock code, product name, and net revenue.
Original Qwen answer; historical selected spans
The top 3 products by merchandise net revenue in April 2011 were: 1. **REGENCY CAKESTAND 3 TIER** (rank 1) with a net revenue of GBP 14,280.90. 2. **PARTY BUNTING** (rank 2) with a net revenue of GBP 10,323.87. 3. **WOODEN UNION JACK BUNTING** (rank 3) with a net revenue of GBP 4,173.18.
Highlight colors preserve the earlier context-bound span judgments. A red amount is not necessarily a number copied from another product. These marks do not assign new relationship labels.
Current interpretation: row fidelity is separate from rank
| Generated product | Generated value | Amount fidelity | Stated rank | Rank in shown evidence |
|---|---|---|---|---|
| REGENCY CAKESTAND 3 TIER | GBP 14,280.90 | Matches own row | 1 | 1 |
| PARTY BUNTING | GBP 10,323.87 | Matches own row | 2 | 2 |
| WOODEN UNION JACK BUNTING | GBP 4,173.18 | Matches own row | 3 | 7 |
Ranking within the eight rows shown to Qwen; historical corpus scope is separately documented. Both answers omit the requested stock codes. The six curated relationships are not a whole-answer correctness score.
Original evidence order and historical gold
Top 3 products in April 2011: 1. 22423 (REGENCY CAKESTAND 3 TIER), GBP 14,280.90; 2. 47566 (PARTY BUNTING), GBP 10,323.87; 3. 22084 (PAPER CHAIN KIT EMPIRE), GBP 6,619.51.
Headers below use current sign-based terminology; original values and row order are unchanged.
| Stock code | Product | Net value | Positive value | Negative value |
|---|---|---|---|---|
| 22423 | REGENCY CAKESTAND 3 TIER | GBP 14,280.90 | GBP 14,812.95 | GBP -532.05 |
| 22499 | WOODEN UNION JACK BUNTING | GBP 4,173.18 | GBP 4,191.03 | GBP -17.85 |
| 47566 | PARTY BUNTING | GBP 10,323.87 | GBP 10,778.47 | GBP -454.60 |
| 85099B | JUMBO BAG RED RETROSPOT | GBP 4,870.92 | GBP 4,878.08 | GBP -7.16 |
| 85123A | WHITE HANGING HEART T-LIGHT HOLDER | GBP 5,078.59 | GBP 10,017.79 | GBP -4,939.20 |
| 22084 | PAPER CHAIN KIT EMPIRE | GBP 6,619.51 | GBP 6,634.26 | GBP -14.75 |
| 84879 | ASSORTED COLOUR BIRD ORNAMENT | GBP 4,174.24 | GBP 4,174.24 | GBP 0.00 |
| 22087 | PAPER BUNTING WHITE LACE | GBP 3,512.24 | GBP 3,579.69 | GBP -67.45 |
Historical detector readout for the selected spans
These outcomes use the original atomic labels and fixed dev thresholds; they are not predictions of the completed relationships above. Residual mass is a probability concentration control.
| Annotation | Text | Historical label | Top-2 margin | Entropy | Top-2 residual mass |
|---|---|---|---|---|---|
| ann_full100_draft_q_0064_004 | GBP 14,280.90 | correct_key_fact | false alarm 0.799101 | false alarm 0.155575 | false alarm 0.016624 |
| ann_full100_draft_q_0064_006 | PARTY BUNTING | correct_key_fact | false alarm 0.769828 | false alarm 0.189533 | false alarm 0.012085 |
| ann_full100_draft_q_0064_008 | 3. | hallucinated_key_fact | missed 6e-06 | missed 3.6e-05 | missed 1e-06 |
| ann_full100_draft_q_0064_010 | GBP 4,173.18 | hallucinated_key_fact | missed 1.9e-05 | missed 5.9e-05 | missed 1e-06 |
q_0069: product amounts and ranking relations
What were the top 3 products by merchandise net revenue in September 2011? Give the rank, stock code, product name, and net revenue.
Original Qwen answer; historical selected spans
The top 3 products by merchandise net revenue in September 2011 were: 1. **JUMBO BAG RED RETROSPOT** (rank 1) with a net revenue of GBP 8630.45. 2. **PAPER CHAIN KIT 50'S CHRISTMAS** (rank 2) with a net revenue of GBP 5997.25. 3. **REGENCY CAKESTAND 3 TIER** (rank 3) with a net revenue of GBP 9315.03.
Highlight colors preserve the earlier context-bound span judgments. A red amount is not necessarily a number copied from another product. These marks do not assign new relationship labels.
Current interpretation: row fidelity is separate from rank
| Generated product | Generated value | Amount fidelity | Stated rank | Rank in shown evidence |
|---|---|---|---|---|
| JUMBO BAG RED RETROSPOT | GBP 8630.45 | Matches own row | 1 | 3 |
| PAPER CHAIN KIT 50'S CHRISTMAS | GBP 5997.25 | Matches own row | 2 | 8 |
| REGENCY CAKESTAND 3 TIER | GBP 9315.03 | Matches own row | 3 | 2 |
Ranking within the eight rows shown to Qwen; historical corpus scope is separately documented. Both answers omit the requested stock codes. The six curated relationships are not a whole-answer correctness score.
Original evidence order and historical gold
Top 3 products in September 2011: 1. 23243 (SET OF TEA COFFEE SUGAR TINS PANTRY), GBP 9,971.51; 2. 22423 (REGENCY CAKESTAND 3 TIER), GBP 9,315.03; 3. 85099B (JUMBO BAG RED RETROSPOT), GBP 8,630.45.
Headers below use current sign-based terminology; original values and row order are unchanged.
| Stock code | Product | Net value | Positive value | Negative value |
|---|---|---|---|---|
| 85123A | WHITE HANGING HEART T-LIGHT HOLDER | GBP 6,880.76 | GBP 6,945.66 | GBP -64.90 |
| 22720 | SET OF 3 CAKE TINS PANTRY DESIGN | GBP 6,250.48 | GBP 6,285.13 | GBP -34.65 |
| 22423 | REGENCY CAKESTAND 3 TIER | GBP 9,315.03 | GBP 9,619.53 | GBP -304.50 |
| 23355 | HOT WATER BOTTLE KEEP CALM | GBP 6,229.79 | GBP 6,254.54 | GBP -24.75 |
| 23243 | SET OF TEA COFFEE SUGAR TINS PANTRY | GBP 9,971.51 | GBP 9,981.41 | GBP -9.90 |
| 22086 | PAPER CHAIN KIT 50'S CHRISTMAS | GBP 5,997.25 | GBP 5,997.25 | GBP 0.00 |
| 85099B | JUMBO BAG RED RETROSPOT | GBP 8,630.45 | GBP 8,880.17 | GBP -249.72 |
| 47566 | PARTY BUNTING | GBP 6,113.70 | GBP 6,383.15 | GBP -269.45 |
Historical detector readout for the selected spans
These outcomes use the original atomic labels and fixed dev thresholds; they are not predictions of the completed relationships above. Residual mass is a probability concentration control.
| Annotation | Text | Historical label | Top-2 margin | Entropy | Top-2 residual mass |
|---|---|---|---|---|---|
| ann_full100_draft_q_0069_005 | 2. | hallucinated_key_fact | missed 1.6e-05 | missed 7.9e-05 | missed 2e-06 |
| ann_full100_draft_q_0069_008 | 3. | hallucinated_key_fact | missed 7e-06 | missed 4.8e-05 | missed 2e-06 |
| ann_full100_draft_q_0069_010 | GBP 9315.03 | hallucinated_key_fact | missed 3.2e-05 | missed 5.9e-05 | missed 2e-06 |
What the case readouts do not prove
A list marker is generated before the following product and amount. Its low token uncertainty does not establish that the model was confident about the completed business relationship. Raw teacher-forced token scores and completed-answer evidence checks have different information budgets.
Positive and negative transaction values are sign-based accounting quantities. Negative value is not a verified measure of physical returns; fees and adjustments may be included. Merchandise scope uses a stock-code heuristic, and the historical product grain is stock code plus description.
There is no independent human benchmark, automatic claim extraction or completed-answer detector here. Seven selected examples of atomic outcomes cannot establish detector superiority, error prevalence or measured business loss.
Historical thresholds: one_minus_min_top2_margin >= 0.334333; mean_token_entropy >= 0.004817; mean_spilled_probability_mass_after_top2 >= 0.000225.
Exploratory max test AUPRC 0.835073; exploratory max test F1 0.779412. These are different test-selected signals, not preregistered winners.
Historical B1 reference checks
103 test spans from 18 questions, with AI-assisted provisional labels. AP is tied-score-aware average precision, not trapezoidal PR area. Scores use saved-trace precision; historical files are unchanged.
| Signal / reference | AP | F1 | Balanced accuracy | MCC |
|---|---|---|---|---|
| Top-2 margin | 0.835 | 0.752 | 0.753 | 0.498 |
| Token entropy | 0.809 | 0.779 | 0.673 | 0.381 |
| Flag every span | 0.592 | 0.744 | 0.500 | 0.000 |
| Flag no spans | 0.592 | 0.000 | 0.500 | 0.000 |
| Dev fact-type prior (composition control) | 0.760 | 0.836 | 0.714 | 0.555 |
Entropy minus flag-every-span F1: +0.0355; exploratory paired question-bootstrap 95% interval [-0.0356, 0.1047]. This interval crosses zero; it is conditional on fixed dev thresholds and provisional labels, not adjusted for test-based signal selection or all shared-period dependence.
The fact-type prior is an annotation-composition control, not an information-matched detector: supplied type names can contain correctness hints. Its higher F1 is not a new headline result. The two overlapping-period components are too few for reliable cluster inference.
Same-step selected energy gap equals token NLL. Probability outside the top two choices is a concentration control, not an independent replication of Spilled Energy. No superiority claim follows from the historical 0.835 / 0.779 maxima.
September 5 update. B2 sensitivity adds 9 assistant-provisional q_0048 dev atoms. Dev-only refitting changes entropy F1 from 0.779 to 0.752 on the same 103 old test spans; top-2 margin is unchanged. Original B1 results remain intact. This is not fresh held-out evidence or independent review. B2 methods appendix.