What is valid
Each candidate signal uses a threshold chosen on dev spans and reuses that fixed threshold on test spans. Character offsets, token alignment, and score rows are reproducibly validated.
Methodology Hardening v1
This audit preserves the published experiment while separating reproducible exploratory evidence from the design of a future confirmation study.
Evidence audit
| Finding | Observed evidence | Interpretation |
|---|---|---|
| Outcome-informed annotation subsethigh | Score rows cover 35 of 36 dev/test questions. The queue code prioritizes held-out answers when generated-answer auto-status is not likely_correct. | The 205-span evaluation is conditional on a high-priority error-enriched subset; it is not a representative estimate over all generated business answers. |
| Post-hoc headline signal selectionhigh | Each signal threshold was selected on dev, but the public AUPRC and F1 winners were identified after comparing candidate signals on test. | The observed maxima are useful exploratory summaries, not unbiased confirmation-set estimates. |
| Context overlap across splitshigh | Dev and test share 5 business periods. There are 9 exact gold evidence-row fingerprint groups crossing splits, involving 28 questions. | The split is question-level, not evidence-context independent; results do not establish generalization to unseen months or evidence payloads. |
| Provisional labels without independent agreementhigh | The 205 spans are AI-assisted provisional labels. Fifteen selected spans received additional assistant review; no independent human annotation or IAA is available. | Label reliability is not yet independently estimated. |
| Oracle span boundarymedium | Detector signals are scored on pre-identified business-fact spans. | The current experiment evaluates span scoring, not automatic claim extraction or an end-to-end audit system. |
Selection and split diagnosis
Each candidate signal uses a threshold chosen on dev spans and reuses that fixed threshold on test spans. Character offsets, token alignment, and score rows are reproducibly validated.
The headline signal was selected after viewing test results, and the annotated subset was chosen through an answer-quality triage queue. This prevents confirmatory interpretation.
q_0048 is the sole dev/test question without detector score rows. It asks which of Netherlands or EIRE generated more net revenue in August 2011.
Dev and test share 5 months: 2011-04, 2011-05, 2011-06, 2011-09, 2011-10. Split assignment is periodic within question type, not grouped by evidence context.
| Evidence hash | Splits | Question IDs | Period |
|---|---|---|---|
4692b67666 |
dev, train | q_0006, q_0014 | 2011-06 |
eca7f3ec5b |
dev, test | q_0009, q_0015 | 2011-09 |
4a4475b083 |
dev, train | q_0020, q_0063, q_0075, q_0085 | 2011-03 |
c9f021e5ec |
dev, test | q_0021, q_0064, q_0076 | 2011-04 |
12f8478dfb |
test, train | q_0022, q_0065, q_0077 | 2011-05 |
88790eb00f |
dev, train | q_0023, q_0066, q_0078, q_0086 | 2011-06 |
bbf5b3e686 |
dev, train | q_0025, q_0068, q_0080 | 2011-08 |
451a40a768 |
dev, test | q_0026, q_0069, q_0081, q_0087 | 2011-09 |
41b64219a1 |
test, train | q_0027, q_0070, q_0082 | 2011-10 |
These are exact fingerprints of the gold evidence.rows payload. They are not claims that prompts are byte-identical; they show that the same underlying evidence table can support questions assigned to different splits.
Public claim boundary
Protocol v1
Locked current record