BizHallu Portfolio Narrative

A guided account of the business problem, inspectable evidence and limits of the retrospective study. The cases were selected for explanation, not as a fresh evaluation sample.

Download the editable presentation (PPTX) or read the research brief for the proposed annotation pilot.

90-second pitch

I am a current MS student at Johns Hopkins Carey Business School, pursuing the Master of Science in Business Analytics and Artificial Intelligence. My accounting and supply-management background motivates BizHallu: can AI copy the right numbers and still reach the wrong business conclusion?

One retail example makes that distinction concrete. Qwen lists WOODEN UNION JACK BUNTING third at 4,173.18 pounds. That product and amount match a source row, but the product is seventh among the eight rows shown to the model. The copying is correct; the ranking is wrong.

I directed the project with AI-assisted implementation and review. The pipeline contains 100 deterministic questions, local Qwen3-0.6B answers, and 205 provisional fact spans aligned to saved token traces. It checks whether uncertainty signals identify the labeled errors.

The original entropy F1 was 0.779, versus 0.744 for flagging every span. Adding one omitted dev answer changes entropy F1 to 0.752 on the same old test. That sensitivity and provisional labels rule out a stable superiority claim.

The next study will compare internal uncertainty with an independent evidence checker on complete product-ranking relationships. I am seeking feedback on relation annotation and a fair comparison design before expanding the evaluation.

Five-minute walkthrough

Ten segments with a suggested total of five minutes. Adjust the timing to the discussion.

1. BizHallu 0:00-0:20

I am a current MS student at Johns Hopkins Carey Business School, pursuing the Master of Science in Business Analytics and Artificial Intelligence. BizHallu asks whether AI-generated business conclusions follow from transaction evidence. I directed the project with AI-assisted implementation and review. My accounting and supply-management background motivates the distinction between a number that reconciles and a relationship that supports a decision.

2. April: a correct amount with an incorrect rank 0:20-0:55

Here is April 2011. The answer places WOODEN UNION JACK BUNTING third with revenue of 4,173.18 pounds. We can locate that exact product and amount in the source table. But six shown products have larger values. The correct third product is PAPER CHAIN KIT EMPIRE at 6,619.51 pounds. A check that only asks whether the amount appears in the table would miss the ranking error.

3. September: the same relationship error 0:55-1:25

The September answer repeats the pattern across all three listed products. Their stated ranks are one, two and three, while their ranks in the eight shown evidence rows are three, eight and two. Every product still matches its own amount. I use these cases to distinguish row fidelity from relationship correctness. They are selected historical examples, not evidence of a general error rate or an automatic verifier.

4. Research question: complete business relationships 1:25-1:55

The research question is what uncertainty signals miss when copied facts form an incorrect business relationship. A rank only makes sense with its product, metric and comparison scope. The proposed comparison asks whether an evidence checker adds useful information under matched conditions. Business definitions remain part of the task: negative transaction value can include fees and adjustments, so it cannot establish physical returns. The historical merchandise definition uses a documented stock-code heuristic.

5. Historical evaluation and its sample 1:55-2:30

The pipeline turns retail transactions into deterministic questions and evidence tables, then captures local Qwen answers and token traces. B1 retains the original 205 provisional spans across 35 development and test answers. A separate B2 sensitivity adds nine assistant-provisional atoms for the omitted dev answer, without changing the 103 old test spans. Fifteen selected original spans received additional assistant review. None of this is independent human annotation or end-to-end claim extraction.

6. The F1 advantage remains uncertain 2:30-3:15

The chart preserves B1's original test results. Entropy F1 0.779; flag-every-span F1 0.744. Paired F1 difference +0.0355; exploratory 95% interval [-0.0356, 0.1047]. That interval is conditional on fixed thresholds and is not a B2 interval. B2 adds one previously omitted dev answer with nine provisional correct facts. Refitting on dev lowers entropy F1 to 0.752 on the same old test, with five fewer false alarms and six more missed errors. Margin is unchanged in this check. We report sensitivity, not a new winner or fresh confirmation.

7. A fair comparison needs the same information 3:15-3:45

One methodological correction is especially important. A list marker is generated before the product and amount that follow it. Low uncertainty on that marker cannot establish that the completed relationship was confidently wrong. The saved raw teacher-forced logits also differ from generation probabilities after sampling transformations. Some nominally different features are mathematical aliases: same-step energy gap is token negative log probability. These distinctions prevent an unfair comparison with a checker that sees the whole answer.

8. What the current evidence supports 3:45-4:10

The results describe a limited retrospective study. The annotation queue depended on generated-answer status, development and test share contexts, and independent human agreement has not been measured. The six curated relationships clarify particular cases; they do not replace the 205 labels. No result here establishes performance on larger models, unseen periods, automatically extracted claims or production business reports.

9. Proposed comparison: product-ranking verification 4:10-4:35

The next implementation targets product rankings. It will use the question, generated answer, metric contract and evidence rows to return a support decision, evidence reference and abstention when needed. Gold answers and evaluation labels stay outside prediction. Two independent reviewers will establish relation labels. The comparison will report extraction coverage and abstentions, because a checker can appear accurate by avoiding difficult claims. This verifier is proposed work, not a completed result.

10. A focused discussion with a research mentor 4:35-5:00

My request is a focused discussion of the annotation unit and comparison design. The immediate deliverable would be a small independently reviewed relation set and one evidence checker. The existing 48-context design offers a later evaluation path, but it remains unexecuted and estimation-focused. Model and prompt settings, reviewers, and the prediction protocol must be ready first. A second dataset or model would follow only after this smaller comparison is credible.

Historical B1 reference checks

103 test spans from 18 questions, with AI-assisted provisional labels. AP is tied-score-aware average precision, not trapezoidal PR area. Scores use saved-trace precision; historical files are unchanged.

Signal / referenceAPF1Balanced accuracyMCC
Top-2 margin0.8350.7520.7530.498
Token entropy0.8090.7790.6730.381
Flag every span0.5920.7440.5000.000
Flag no spans0.5920.0000.5000.000
Dev fact-type prior (composition control)0.7600.8360.7140.555

Entropy minus flag-every-span F1: +0.0355; exploratory paired question-bootstrap 95% interval [-0.0356, 0.1047]. This interval crosses zero; it is conditional on fixed dev thresholds and provisional labels, not adjusted for test-based signal selection or all shared-period dependence.

The fact-type prior is an annotation-composition control, not an information-matched detector: supplied type names can contain correctness hints. Its higher F1 is not a new headline result. The two overlapping-period components are too few for reliable cluster inference.

Same-step selected energy gap equals token NLL. Probability outside the top two choices is a concentration control, not an independent replication of Spilled Energy. No superiority claim follows from the historical 0.835 / 0.779 maxima.

September 5 update. B2 sensitivity adds 9 assistant-provisional q_0048 dev atoms. Dev-only refitting changes entropy F1 from 0.779 to 0.752 on the same 103 old test spans; top-2 margin is unchanged. Original B1 results remain intact. This is not fresh held-out evidence or independent review. B2 methods appendix.

Presentation guardrails

The historical assistant_full_review field applies to 15 selected spans, not independent review of all 205 labels. These artifacts support discussion of business analytics and AI reliability, not a production claim.

Reader paths

q_0064 and q_0069 show source-row fidelity versus ranking. The career package contains resume wording and interview FAQ. The research brief proposes a limited feedback request and preserves alternative comparison methods.