Confirmation Set v1 · Prospective freeze checkpoint
The 48-context inventory and 6/15/27 split are now frozen.
A deterministic, outcome-blind assignment fixed one complete source week per context before any new question, prompt, model answer, label, detector score, or empirical metric was created.
Selected contexts48
Pilot / development6 / 15
Sealed confirmation27
Unassigned reserve2
Gate completed. All 48 selected periods are unique, every period belongs to one family only, and all 48 context-level canonical evidence-pool hashes are distinct.
Execution remains blocked. This is a manifest commitment, not a model experiment or a new performance result. The next gate is deterministic question templates and gold calculations.
Frozen allocation
Each business family receives the same 2 / 5 / 9 split.
Question family
Protocol pilot
Development
Confirmation
Total
net_revenue_reconciliation_by_period
2
5
9
16
product_return_rate_comparison
2
5
9
16
country_product_exposure
2
5
9
16
All families
6
15
27
48
Each context is planned to support two questions later, preserving the frozen 96-question target. No question identifiers have been assigned in this step.
Outcome-blind assignment
The mapping uses only source eligibility, the frozen seed, and domain-separated hashes.
Family matching
A deterministic hash-ordered bipartite matching fills 16 slots for each of the three eligible families. No generated-answer field or detector outcome is an input.
Split assignment
Within each family, an independent seeded hash order assigns 2 contexts to pilot, 5 to development, and 9 to sealed confirmation.
Period separation
All 48 selected complete weeks are unique across every family and split. The two unmatched weeks remain reserve and cannot replace difficult outputs later.
Blocked family
Customer revenue concentration remains excluded because the pre-existing data-quality audit found substantial missing Customer ID coverage.
Public commitment
The public hash commits to a private, Git-ignored manifest.
The private file contains the selected weeks, family and split assignments, eligible scope entities, context IDs, and context evidence-pool hashes. Those values are not published. The public artifact records only aggregate counts and the canonical manifest commitment:
The private manifest is Git-ignored and is not part of the GitHub Pages bundle.
Regenerating the same canonical manifest reproduces the same commitment; changing any selected period, entity, split, or evidence hash changes it.
Evidence separation
Context pools are fixed now; question payloads are checked next.
Context-pool hashes48
Unique pool hashes48
Historical canonical overlap rows0
Date-blind overlap rows0
Important distinction. Context-level source pools are now frozen and disjoint. Exact question-level evidence payload fingerprints remain pending because question templates and deterministic gold payloads do not exist yet. That check belongs to the next gate and has not been silently treated as complete.
What did not happen
No new empirical result was created.
No questions, gold answers, prompts, Qwen outputs, annotations, verifier decisions, or detector scores were created.
The historical 0.835 AUPRC and 0.779 F1 remain exploratory maxima from the earlier outcome-informed span subset.
Confirmation Set v1 remains estimation-focused and does not authorize a detector-superiority claim.
This remains a same-retailer, same-lineage temporal internal replication design, not external validation.
Next gate
Define and validate questions and gold calculations, without running Qwen.
Define deterministic question templates, scope-entity selection rules, gold calculations, and question-level evidence payload fingerprints without running the model.
Still prohibited: prompt generation, model execution, annotation, detector or verifier scoring, and new empirical metrics.