Every figure on this page is a measurement on a published harness, not a certification and not an opinion. Read the artefacts, re-run the harness, get the same number. Generated from signed benchmark artefacts at 2026-08-09 05:26 UTC.
The care-floor battery: 57 harmful items that must be refused and 19 benign/benign-near items that must be answered. The gate is deterministic (no model in the refuse path). Measured 2026-08-09: recall 100% · precision 100% · over-block 0%. Every one of the 57 harms is refused; zero benign requests are refused.
| Difficulty | Items | Harm recall | Benign over-block |
|---|---|---|---|
| adversarial | 0 | — | 0 |
| benign_near | 0 | — | 0 |
| euphemism | 20 | 100% | 0 |
| fragmented | 11 | 100% | 0 |
| indirection | 7 | 100% | 0 |
| plain | 19 | 100% | 0 |
Fragmented and euphemism items (the phrasings a rubber-stamp scorer flunks) are fully caught: a gate that only matches the word "harm" would score 0.68 here. The battery itself is open and versioned; deleting hard items to make recall read 1.0 is caught by a denominator floor in the regression suite.
Single-sample refusal scores are noise; refusal is unstable across seeds and temperature. The
flywheel now reports two sides: refusal_tpr (catches must-refuse
harm) and false_refusal_fpr (refuses must-allow). A model at TPR 1.0
and FPR 1.0 is refuse-everything — not safe, useless. Practice split; the fuel path is
held-out-stripped and leak-guarded.
| Model | n | Accuracy | Refusal TPR | False-refusal FPR | tokens/correct | overfit gap |
|---|---|---|---|---|---|---|
qwen2.5:0.5b | 10 | 0.50 | 0.38 | 0.00 | 376.6 | -0.5 |
qwen2.5:1.5b | 10 | 0.40 | 0.38 | 0.50 | 441.8 | -0.6 |
tokens per correct verdict is our production number — nobody else publishes it.
Cheap and right beats expensive and right; the scorecard prices that.