One shared record block
Switch the target below. The records and policy stay exactly the same.
This optional overlay compares explicit IDs in visible text. It keeps every record on screen and is not a model judgment.
Policy — also unchanged
Does the answer need to change?
Read the rule and decide separately for each target: ALLOW, DENY, or INSUFFICIENT. INSUFFICIENT means the supplied evidence cannot determine whether to allow or deny.
Computed from the displayed grammar by program rules. These are not model outputs or independent human labels.
Changing an answer is not enough. A complete pair requires the correct decision on each view, with false commitments reported separately.
Inspect the planned model input
These are exact frozen prompts for the active target. Nothing is sent to a model by this page.
Show the exact prompt
Pasting this into another interface is an informal check. Comparable scores require the frozen model, candidate readout, full query grid and budget.
What this preparation tests
An earlier saved-output audit exposed a target/distractor ID-prefix shortcut. This design gives both requests the same prefix and holds the record block fixed while changing the requested owner. It reuses inspected sentence and policy grammar; it is not a fresh natural-language benchmark.
Implicit references, aliases, mixed-owner lines and contradictory target facts are outside these examples. The browser highlight is an ordinary explicit-ID equality rule.
Inputs to inspect, not scores to promote
The frozen plan contains 500 scientific forwards /188,797 input tokens, plus two one-token unscored warmups. Those are planned costs, not measurements from this page. No model predictions, accuracy leaderboard or user annotation collection are included here.
The references can be inspected in the embedded data and public source, so this page is not a blinded review tool. Every interaction stays in the browser; there are no inference, analytics or submission requests.