MEASURED

GSPC-AGI — safety

Axis: safety · items: 14 · status: MEASURED

Refusal behaviour on an adversarial battery: should the model refuse, and does it? Scored in both directions, because a model that refuses the benign half is broken too.

Dataset: csoai/gspc-agi · Runnable Space: csoai-gspc-agi · Licence Apache-2.0

Run DefBench yourself

The same 14 items sov34 answered, graded by the same deterministic rule. sov34 scored 0.653 macro-F1 on these. No sign-up, nothing leaves your browser.

Items: csoai/gspc-agi · grading is a regex label read plus macro-F1, identical to the published harness · measurement, not certification, and not legal advice.

What this is, and is not. CSOAI measures. It issues no conformity marks, holds no accreditation, and has no enforcement powers — those are conferred by statute on market-surveillance authorities and the AI Office. Nothing on this page is a certification, an attestation of compliance, or legal advice.

Denominator. This axis has 14 items, below the usable_n = 30 floor CSOAI requires before attaching a confidence interval to a number. No score is quoted here, for any model including our own.