Benchmarks

Every figure on this page is a measurement on a published harness, not a certification and not an opinion. Read the artefacts, re-run the harness, get the same number. Generated from signed benchmark artefacts at 2026-08-09 05:26 UTC.

EAT refusal suite — deterministic care gate (76 items)

The care-floor battery: 57 harmful items that must be refused and 19 benign/benign-near items that must be answered. The gate is deterministic (no model in the refuse path). Measured 2026-08-09: recall 100% · precision 100% · over-block 0%. Every one of the 57 harms is refused; zero benign requests are refused.

DifficultyItemsHarm recallBenign over-block
adversarial00
benign_near00
euphemism20100%0
fragmented11100%0
indirection7100%0
plain19100%0

Fragmented and euphemism items (the phrasings a rubber-stamp scorer flunks) are fully caught: a gate that only matches the word "harm" would score 0.68 here. The battery itself is open and versioned; deleting hard items to make recall read 1.0 is caught by a denominator floor in the regression suite.

Flywheel model board — two-sided refusal (arXiv 2512.12066)

Single-sample refusal scores are noise; refusal is unstable across seeds and temperature. The flywheel now reports two sides: refusal_tpr (catches must-refuse harm) and false_refusal_fpr (refuses must-allow). A model at TPR 1.0 and FPR 1.0 is refuse-everything — not safe, useless. Practice split; the fuel path is held-out-stripped and leak-guarded.

ModelnAccuracyRefusal TPRFalse-refusal FPRtokens/correctoverfit gap
qwen2.5:0.5b100.500.380.00376.6-0.5
qwen2.5:1.5b100.400.380.50441.8-0.6

tokens per correct verdict is our production number — nobody else publishes it. Cheap and right beats expensive and right; the scorecard prices that.

Fleet & integrity artefacts

What we do not do