Confirmation Set v1 · Outcome-blind precision review

The 15-context comparison plan was too fragile, so the claim was narrowed instead of the thresholds.

Synthetic clustered scenarios were fixed before the formal run. No candidate met every strong-comparison rule. BizHallu therefore preserves the failed review, increases the sealed confirmation plan to 27 balanced contexts, and limits the study to estimation with uncertainty rather than a detector-superiority claim.

Strong-design candidates passed0 of 4
Revised confirmation contexts27
Total period-disjoint contexts48
Unassigned source periods2
The original precision gate did not pass. Its thresholds were not relaxed, its random seed was not replaced, and the result was not rewritten as a success.
A narrower study remains feasible. The revised 6/15/27 plan has an independently recomputed 48/48 aggregate matching, with no period assignment retained.

What was simulated

Design sensitivity, not detector performance.

Outcome-blind inputs

Only synthetic prevalence, spans per context, label ICC, score separation, and aggregate source capacity were used. At this review checkpoint, confirmation periods, labels, answers, and detector outcomes did not exist.

Independent unit

Whole synthetic evidence contexts were resampled together. Both detector scores used the same resampled context indices for the paired comparison.

Unequal cluster sizes

Each scenario used a fixed multiplier pattern around 8 or 12 mean spans per context. Stress scenarios lowered density and increased within-context label correlation.

Interpretation boundary. Interval half-widths below are Monte Carlo planning summaries. They are not observed AUPRC, F1, prevalence, power, or model-quality results.

Frozen decision rules

No candidate satisfied the complete strong-comparison gate.

A preferred central scenario required median AUPRC half-width at most 0.10 and median paired-difference half-width at most 0.08. A design needed at least 75% preferred central scenarios plus the frozen worst-case central limits.

Confirmation contextsPer familySource reservePreferred centralMedian AUPRC half-widthWorst central AUPRC half-widthWorst paired half-widthStrong gate
155145/160.1030.1500.126did not pass
217810/160.0950.1360.107did not pass
24859/160.0870.1270.102did not pass
279210/160.0830.1210.098did not pass

The 27-context candidate came closest, but passed only 10 of 16 preferred central scenarios and had a worst central median AUPRC half-width of 0.121219, just above the frozen 0.12 limit.

Scope amendment

Estimate transparently; do not declare a winner.

Primary objective

Estimate the AUPRC of each frozen detector family on the sealed confirmation set with context-resampled uncertainty intervals.

Paired comparison

Report the paired AUPRC difference and its interval as a descriptive secondary estimate, not as a binary winner or superiority test.

Subgroups

Question-family and fact-type results are descriptive only because the number of independent contexts per subgroup remains small.

Prohibited claim. Do not claim that evidence-aware verification, internal uncertainty, or a hybrid detector is statistically superior on Confirmation Set v1.

Statistical caution

Twenty-seven clusters improve precision, but do not make this publication-grade.

Fewer than 30 independent contexts remains a small-cluster setting. Report intervals and avoid a binary superiority claim when the interval is wide or interval methods disagree.

The future sealed analysis plans 5,000 context-level bootstrap replicates, a BCa interval when leave-one-context-out estimates are defined, and a paired percentile interval as sensitivity analysis. This method still requires implementation and validation on non-confirmation data.

Subsequent gate

The authorized manifest freeze is complete; questions come next.

  1. The later manifest gate froze 48 period-disjoint contexts with a 6/15/27 seeded split.
  2. Selected periods and scope entities remain local; the public manifest report exposes only aggregate proof and a commitment hash.
  3. The next gate may define deterministic question templates, gold calculations, and question-level evidence fingerprints.
  4. Prompts, Qwen execution, annotation, detector or verifier scoring, and new empirical metrics remain forbidden.
Checkpoint versus current state. At the precision-review checkpoint no manifest or split existed. The subsequent outcome-blind freeze completed those two artifacts only. No question, prompt, model output, new label, detector score, or new empirical metric exists.