Outcome-blind sampling
All contexts and question IDs are selected before generation. Answers are never retained, dropped, or rebalanced because they look correct, incorrect, easy, or difficult.
Prospective protocol design
This design replaces answer-driven sampling with a frozen context manifest, independent human review, sealed confirmation labels, and separate evaluations for claim extraction and evidence verification.
Why this exists
All contexts and question IDs are selected before generation. Answers are never retained, dropped, or rebalanced because they look correct, incorrect, easy, or difficult.
The frozen manifest uses one unique complete week per context across every family and split. All 96 full payloads and all 96 normalized evidence-table contents have unique fingerprints, with zero cross-split overlap. The wrapper-independent content comparison matches 0 of 66 unique historical full100 evidence contents.
Two human reviewers annotate every in-scope business-fact claim without detector scores or each other's labels. Agreement is reported before adjudication.
Detector families, thresholds, extraction logic, verifier rules, prompts, and analysis code are frozen before confirmation-label access.
Dataset decision
The source, context-manifest, and deterministic question-design gates are complete. A private, Git-ignored question manifest now fixes 96 questions, gold answers, and evidence payloads under a public SHA-256 commitment, with zero exact cross-split evidence-content overlap and zero exact content overlap against historical full100. Next freeze model, tokenizer, prompt, decoding, detector-family, and metric configurations before protocol-pilot generation. A separately scoped second-dataset replication remains required for stronger comparative or generalization claims.
| Option | Source and role | Advantages | Limitations |
|---|---|---|---|
| same_dataset_new_contextsprotocol_development_only | UCI Online Retail cleaned lineage already used by BizHallu |
|
|
| same_source_prior_periodprospective_temporal_internal_replication | UCI Online Retail II restricted to 2009-12-01 inclusive through 2010-12-01 exclusive |
|
|
| second_public_transaction_datasetexternal_replication | Complete Journey is the preferred external shortlist candidate, subject to acquisition, license-lineage, completeness, and metric-definition review |
|
|
| jhu_domain_extensiondomain_transfer_extension | A future JHU course, capstone, operations, or healthcare dataset with permitted use |
|
|
Sampling architecture
Candidate business tasks
verified
accounting reconciliation of gross positive revenue, cancellations, and net revenue
blocked_by_preexisting_quality_control
customer exposure and concentration risk
The source audit records 100207 missing Customer IDs. This family is excluded from the 48-context capacity target rather than introducing a complete-case denominator after inspecting capacity.
verified
returns and product-performance monitoring
verified
market and product concentration across business entities
Three families pass source-capacity checks; customer concentration is blocked by the pre-existing Customer ID coverage rule. Six templates now create two deterministic questions per context: weekly reconciliation and cancellation hotspot, two disjoint product-return comparisons, or two country-product exposure questions. Gold calculations and evidence payload fingerprints are frozen before prompt generation.
Evaluation architecture
Compare frozen internal, literature-grounded, evidence-aware, and optional hybrid families on adjudicated spans.
Measure exact and overlap span precision, recall, F1, and fact-type classification separately.
Produce independent statuses, a continuous unsupported-risk score, abstentions, and matched evidence references. Gold labels cannot be used as predictions.
Combine extraction and verification, then report missed claims, abstentions, and business-fact errors.
Metric freeze
contradicted and unmatched are positive unsupported claims; unresolved needs_review items are counted and excluded only after an adjudication attempt.one_minus_min_top2_margin is a frozen candidate because of the exploratory study, not because it is already confirmed.Execution gates
| # | Gate | Status |
|---|---|---|
| 1 | dataset_source_selected_and_audited | complete |
| 2 | context_manifest_split_and_precision_review_frozen | complete |
| 3 | question_templates_and_gold_calculations_validated | complete |
| 4 | model_prompt_and_detector_configs_frozen | pending |
| 5 | two_independent_human_reviewers_assigned | pending |
| 6 | claim_extraction_and_verifier_protocols_implemented_on_non_confirmation_data | pending |
| 7 | sealed_confirmation_run_authorized | pending |
Historical boundary
The existing numbers are not Confirmation Set v1 baselines or targets. This design introduces no new detector metric and does not retroactively upgrade the current evidence level.