Outcome-blind inputs
Only synthetic prevalence, spans per context, label ICC, score separation, and aggregate source capacity were used. At this review checkpoint, confirmation periods, labels, answers, and detector outcomes did not exist.
Confirmation Set v1 · Outcome-blind precision review
Synthetic clustered scenarios were fixed before the formal run. No candidate met every strong-comparison rule. BizHallu therefore preserves the failed review, increases the sealed confirmation plan to 27 balanced contexts, and limits the study to estimation with uncertainty rather than a detector-superiority claim.
What was simulated
Only synthetic prevalence, spans per context, label ICC, score separation, and aggregate source capacity were used. At this review checkpoint, confirmation periods, labels, answers, and detector outcomes did not exist.
Whole synthetic evidence contexts were resampled together. Both detector scores used the same resampled context indices for the paired comparison.
Each scenario used a fixed multiplier pattern around 8 or 12 mean spans per context. Stress scenarios lowered density and increased within-context label correlation.
Frozen decision rules
A preferred central scenario required median AUPRC half-width at most 0.10 and median paired-difference half-width at most 0.08. A design needed at least 75% preferred central scenarios plus the frozen worst-case central limits.
| Confirmation contexts | Per family | Source reserve | Preferred central | Median AUPRC half-width | Worst central AUPRC half-width | Worst paired half-width | Strong gate |
|---|---|---|---|---|---|---|---|
| 15 | 5 | 14 | 5/16 | 0.103 | 0.150 | 0.126 | did not pass |
| 21 | 7 | 8 | 10/16 | 0.095 | 0.136 | 0.107 | did not pass |
| 24 | 8 | 5 | 9/16 | 0.087 | 0.127 | 0.102 | did not pass |
| 27 | 9 | 2 | 10/16 | 0.083 | 0.121 | 0.098 | did not pass |
The 27-context candidate came closest, but passed only 10 of 16 preferred central scenarios and had a worst central median AUPRC half-width of 0.121219, just above the frozen 0.12 limit.
Scope amendment
Estimate the AUPRC of each frozen detector family on the sealed confirmation set with context-resampled uncertainty intervals.
Report the paired AUPRC difference and its interval as a descriptive secondary estimate, not as a binary winner or superiority test.
Question-family and fact-type results are descriptive only because the number of independent contexts per subgroup remains small.
Statistical caution
Fewer than 30 independent contexts remains a small-cluster setting. Report intervals and avoid a binary superiority claim when the interval is wide or interval methods disagree.
The future sealed analysis plans 5,000 context-level bootstrap replicates, a BCa interval when leave-one-context-out estimates are defined, and a paired percentile interval as sensitivity analysis. This method still requires implementation and validation on non-confirmation data.
Subsequent gate