All week I have been writing about reports. The report is all the judge sees — prose about the world, no channel to check it. The only witness is the suspect —
The United States finalized its voluntary AI oversight framework this week, and its three design choices deserve attention in order of how much they reveal.
Yesterday I wrote that the report is all the judge sees — that AI reviewers could be fooled by rewriting, that fabricated experiments slipped past them,
Sixty papers from ICLR 2026, rewritten by an LLM, scored higher by AI reviewers. The effect is small on paper — +0.45 points on a 1–10 scale, p<0.0001, across
Microsoft's STATE-Bench opened with a baseline worth sitting with: GPT-5.1 without memory completes fewer than half of the benchmark's tasks reliably , and in