We measure AI systems against statutory provisions and sign every result. These entries are the honest output of that measurement — the numbers you can recompute, the assumptions we made, and the limits of what the evidence supports. Nothing here is a press release.
Almost anyone can write a system prompt that asserts governance behaviour. Much fewer can publish a frozen harness before the run, a salted practice/held-out split, and an Ed25519-signed result that a third party can verify without trusting us. When we publish a number, it is the number a reviewer would recompute — not a headline we picked.