A real product and its correct revenue can still be the wrong answer to a ranking question. Two curated walkthroughs separate source-row accuracy from ranking accuracy.
Historical selected spans and detector readouts
These 15 selected spans across nine cases received additional assistant review; they are not an independent human benchmark. Detector scores apply to pre-identified spans and do not discover claims automatically. Filters below affect only this historical span table, not the complete case walkthrough.
Historical B1 reference checks
103 test spans from 18 questions, with AI-assisted provisional labels. AP is tied-score-aware average precision, not trapezoidal PR area. Scores use saved-trace precision; historical files are unchanged.
Signal / reference
AP
F1
Balanced accuracy
MCC
Top-2 margin
0.835
0.752
0.753
0.498
Token entropy
0.809
0.779
0.673
0.381
Flag every span
0.592
0.744
0.500
0.000
Flag no spans
0.592
0.000
0.500
0.000
Dev fact-type prior (composition control)
0.760
0.836
0.714
0.555
Entropy minus flag-every-span F1: +0.0355; exploratory paired question-bootstrap 95% interval [-0.0356, 0.1047]. This interval crosses zero; it is conditional on fixed dev thresholds and provisional labels, not adjusted for test-based signal selection or all shared-period dependence.
The fact-type prior is an annotation-composition control, not an information-matched detector: supplied type names can contain correctness hints. Its higher F1 is not a new headline result. The two overlapping-period components are too few for reliable cluster inference.
Same-step selected energy gap equals token NLL. Probability outside the top two choices is a concentration control, not an independent replication of Spilled Energy. No superiority claim follows from the historical 0.835 / 0.779 maxima.
September 5 update. B2 sensitivity adds 9 assistant-provisional q_0048 dev atoms. Dev-only refitting changes entropy F1 from 0.779 to 0.752 on the same 103 old test spans; top-2 margin is unchanged. Original B1 results remain intact. This is not fresh held-out evidence or independent review. B2 methods appendix.