BizHallu

Prospective protocol design

Confirmation Set v1 starts before the model answers.

This design replaces answer-driven sampling with a frozen context manifest, independent human review, sealed confirmation labels, and separate evaluations for claim extraction and evidence verification.

Current stateDesign only
Main questions84
Human reviewers2
Execution gates4 pending
Not execution-ready. The source, context-manifest, and question-design gates are complete. Private, Git-ignored manifests now fix 48 complete-week contexts, the 6/15/27 split, and 96 deterministic questions with gold answers and evidence payloads. Prompts, model outputs, labels, detector scores, and Confirmation Set v1 metrics do not exist yet.
Active business-definition amendment: v1.1. Positive and negative transaction value include eligible non-merchandise lines; negative value is not a measure of confirmed physical returns. The same 48 contexts, split and 96 numeric/entity gold answers are preserved. Legacy v1 fields below are historical provenance; future inputs must follow active_business_definition in the protocol. Gate 3 has been revalidated; no model run is authorized.

Why this exists

Correct the current study's selection and split limits prospectively.

Outcome-blind sampling

All contexts and question IDs are selected before generation. Answers are never retained, dropped, or rebalanced because they look correct, incorrect, easy, or difficult.

Context-separated evaluation

The frozen manifest uses one unique complete week per context across every family and split. All 96 full payloads and all 96 normalized evidence-table contents have unique fingerprints, with zero cross-split overlap. The wrapper-independent content comparison matches 0 of 66 unique historical full100 evidence contents.

Independent labels

Two human reviewers annotate every in-scope business-fact claim without detector scores or each other's labels. Agreement is reported before adjudication.

Sealed research decisions

Detector families, thresholds, extraction logic, verifier rules, prompts, and analysis code are frozen before confirmation-label access.

Dataset decision

Use a staged path instead of overstating one dataset.

The source, context-manifest, and deterministic question-design gates are complete. A private, Git-ignored question manifest now fixes 96 questions, gold answers, and evidence payloads under a public SHA-256 commitment, with zero exact cross-split evidence-content overlap and zero exact content overlap against historical full100. Next freeze model, tokenizer, prompt, decoding, detector-family, and metric configurations before protocol-pilot generation. A separately scoped second-dataset replication remains required for stronger comparative or generalization claims.

OptionSource and roleAdvantagesLimitations
same_dataset_new_contextsprotocol_development_only UCI Online Retail cleaned lineage already used by BizHallu
  • lowest implementation cost
  • keeps the existing business and accounting definitions
  • allows context-grouped sampling before any new model output is viewed
  • same underlying dataset as the exploratory study
  • reuses the same transaction population even when question families are new
  • cannot support cross-dataset or unseen-domain generalization
same_source_prior_periodprospective_temporal_internal_replication UCI Online Retail II restricted to 2009-12-01 inclusive through 2010-12-01 exclusive
  • preserves the current schema and accounting definitions
  • moves the study to unseen transaction rows and earlier business periods
  • supports a strict cutoff before the current Online Retail source begins
  • same retailer and source lineage as the exploratory study
  • the full workbook overlaps the current source, so the strict prior-period cutoff and frozen record-level fingerprint proof must remain in force
  • the verified capacity is limited to three source-feasible families; customer concentration remains blocked by missing Customer ID coverage
second_public_transaction_datasetexternal_replication Complete Journey is the preferred external shortlist candidate, subject to acquisition, license-lineage, completeness, and metric-definition review
  • stronger external-validity claim
  • tests whether findings transfer beyond one retail table
  • requires a new source audit and cleaning pipeline
  • metric definitions may not map exactly to Online Retail
jhu_domain_extensiondomain_transfer_extension A future JHU course, capstone, operations, or healthcare dataset with permitted use
  • strong alignment with program coursework and faculty outreach
  • tests the business-fact grounding framework outside retail
  • availability and permissions are not yet known
  • should not delay the near-term internal replication
Near-term source decision. The strict window is conditionally suitable for aggregate business analysis after exact-duplicate, description, customer-coverage, cancellation, and value controls. Zero historical record overlap and 48-slot aggregate capacity are verified across three allowed families. The precision review rejected a superiority design and retained an estimation-only scope. The 48-context manifest and 6/15/27 split are now frozen. Customer concentration remains blocked, and a broader claim still requires the second-public-dataset arm. Read the source audit. Read the quality profile. Read the overlap proof. Read the capacity proof. Read the precision review. Read the manifest commitment.

Sampling architecture

96 generations, but only 84 enter the main study.

Protocol pilot12
Development30
Confirmation54
Main total84

Candidate business tasks

Keep accounting and supply-management relevance visible.

net_revenue_reconciliation_by_period

verified

accounting reconciliation of gross positive revenue, cancellations, and net revenue

customer_revenue_concentration

blocked_by_preexisting_quality_control

customer exposure and concentration risk

The source audit records 100207 missing Customer IDs. This family is excluded from the 48-context capacity target rather than introducing a complete-case denominator after inspecting capacity.

product_return_rate_comparison

verified

returns and product-performance monitoring

country_product_exposure

verified

market and product concentration across business entities

Three families pass source-capacity checks; customer concentration is blocked by the pre-existing Customer ID coverage rule. Six templates now create two deterministic questions per context: weekly reconciliation and cancellation hotspot, two disjoint product-return comparisons, or two country-product exposure questions. Gold calculations and evidence payload fingerprints are frozen before prompt generation.

Pre-generation diagnostics. One selected reconciliation row has a cancellation invoice prefix but positive quantity and revenue; the sign-based formula retains it in gross positive line revenue and does not call it a negative reduction. One of 64 selected product evidence rows has a recorded return-to-positive-sales unit ratio above 100% (maximum 156.25%). It is retained because same-week negative units are not linked to their original sales; this metric is a recorded ratio, not a causal return rate. Hash ordering places the correct country-product candidate across positions 1-5, and one of 32 candidate tables happens to be fully revenue-descending by chance. The ordering algorithm never reads the gold rank.

Evaluation architecture

Do not hide oracle spans inside an end-to-end claim.

1. Oracle-span diagnostics

Compare frozen internal, literature-grounded, evidence-aware, and optional hybrid families on adjudicated spans.

2. Claim extraction

Measure exact and overlap span precision, recall, F1, and fact-type classification separately.

3. Evidence verification

Produce independent statuses, a continuous unsupported-risk score, abstentions, and matched evidence references. Gold labels cannot be used as predictions.

4. End-to-end audit

Combine extraction and verification, then report missed claims, abstentions, and business-fact errors.

Metric freeze

AUPRC is primary; uncertainty is clustered by context.

Execution gates

3 gates are complete; 4 remain pending.

#GateStatus
1dataset_source_selected_and_auditedcomplete
2context_manifest_split_and_precision_review_frozencomplete
3question_templates_and_gold_calculations_validatedcomplete
4model_prompt_and_detector_configs_frozenpending
5two_independent_human_reviewers_assignedpending
6claim_extraction_and_verifier_protocols_implemented_on_non_confirmation_datapending
7sealed_confirmation_run_authorizedpending
Next authorized action. Freeze the model revision, tokenizer revision, prompt template, decoding settings, detector-family feasibility decisions, and metric implementations. Do not run Qwen, annotate outputs, score detectors, or report new metrics yet.

Historical boundary

The current 0.835 / 0.779 values remain exploratory context.

The existing numbers are not Confirmation Set v1 baselines or targets. This design introduces no new detector metric and does not retroactively upgrade the current evidence level.