BizHallu

Methodology Hardening v1

What the current result proves, and what it does not.

This audit preserves the published experiment while separating reproducible exploratory evidence from the design of a future confirmation study.

Study classExploratory
Annotated questions35 / 36
Provisional spans205
Share statusWith caveats
Bottom line. The 0.835 AUPRC and 0.779 F1 values are real outputs of the committed pipeline, but they are exploratory maxima on an error-enriched subset with context overlap and provisional labels. They should motivate a fresh confirmation study, not be presented as its result.

Evidence audit

Five boundaries now made explicit

FindingObserved evidenceInterpretation
Outcome-informed annotation subsethigh Score rows cover 35 of 36 dev/test questions. The queue code prioritizes held-out answers when generated-answer auto-status is not likely_correct. The 205-span evaluation is conditional on a high-priority error-enriched subset; it is not a representative estimate over all generated business answers.
Post-hoc headline signal selectionhigh Each signal threshold was selected on dev, but the public AUPRC and F1 winners were identified after comparing candidate signals on test. The observed maxima are useful exploratory summaries, not unbiased confirmation-set estimates.
Context overlap across splitshigh Dev and test share 5 business periods. There are 9 exact gold evidence-row fingerprint groups crossing splits, involving 28 questions. The split is question-level, not evidence-context independent; results do not establish generalization to unseen months or evidence payloads.
Provisional labels without independent agreementhigh The 205 spans are AI-assisted provisional labels. Fifteen selected spans received additional assistant review; no independent human annotation or IAA is available. Label reliability is not yet independently estimated.
Oracle span boundarymedium Detector signals are scored on pre-identified business-fact spans. The current experiment evaluates span scoring, not automatic claim extraction or an end-to-end audit system.

Selection and split diagnosis

The current test is held out for thresholds, not for the full research decision.

What is valid

Each candidate signal uses a threshold chosen on dev spans and reuses that fixed threshold on test spans. Character offsets, token alignment, and score rows are reproducibly validated.

What remains exploratory

The headline signal was selected after viewing test results, and the annotated subset was chosen through an answer-quality triage queue. This prevents confirmatory interpretation.

Missing held-out item

q_0048 is the sole dev/test question without detector score rows. It asks which of Netherlands or EIRE generated more net revenue in August 2011.

Context overlap

Dev and test share 5 months: 2011-04, 2011-05, 2011-06, 2011-09, 2011-10. Split assignment is periodic within question type, not grouped by evidence context.

Evidence hashSplitsQuestion IDsPeriod
4692b67666 dev, train q_0006, q_0014 2011-06
eca7f3ec5b dev, test q_0009, q_0015 2011-09
4a4475b083 dev, train q_0020, q_0063, q_0075, q_0085 2011-03
c9f021e5ec dev, test q_0021, q_0064, q_0076 2011-04
12f8478dfb test, train q_0022, q_0065, q_0077 2011-05
88790eb00f dev, train q_0023, q_0066, q_0078, q_0086 2011-06
bbf5b3e686 dev, train q_0025, q_0068, q_0080 2011-08
451a40a768 dev, test q_0026, q_0069, q_0081, q_0087 2011-09
41b64219a1 test, train q_0027, q_0070, q_0082 2011-10

These are exact fingerprints of the gold evidence.rows payload. They are not claims that prompts are byte-identical; they show that the same underlying evidence table can support questions assigned to different splits.

Public claim boundary

Use the result, but name its evidence level.

Accurate wording

  • exploratory maximum test AUPRC across candidate signals
  • exploratory maximum test F1 across candidate signals
  • AI-assisted provisional span labels
  • evaluation of pre-identified business-fact spans
  • results conditional on the selected high-priority question subset

Claims not supported yet

  • confirmatory generalization estimate
  • unbiased held-out model-selection result
  • representative error rate across all business questions
  • independent human-labeled benchmark
  • end-to-end automatic hallucination detection
  • production-ready detector

Protocol v1

Design the next run before seeing its answers.

  1. Fresh contexts. fresh business contexts, model outputs, or a second dataset not used for exploratory signal selection.
  2. Context-separated splits. Group by evidence source plus business period plus canonical evidence-row payload; never place identical evidence payloads across development and confirmation.
  3. Independent labels. two independent reviewers for the confirmation subset; report binary label agreement and span-boundary agreement before adjudication.
  4. Freeze decisions. Primary metric is AUPRC; detector families, thresholds, extraction rules, prompts, and verifier logic are fixed before confirmation-set access.
  5. Separate tasks. Report oracle-span diagnostics, claim extraction, and end-to-end verification as different evaluations.
Research direction. The next study compares internal uncertainty, literature-grounded baselines, and an independently implemented evidence-aware verifier. It does not assume one family will win.

Locked current record

No result was recomputed in this phase.

Max test AUPRC0.835
Max test F10.779
Test spans103
Independent IAANot done