What Moody's did for bonds, CSOAI does for AI governance: we run the tests, we own the data, everyone else cites the scores. The scorecard is the public measurement surface — free for the numbers, paid for the deep audits. Generated from signed artefacts at 2026-08-09 05:26 UTC.
Every agent, model and deployment is measured on four axes — Governance, Security, Privacy, Commerce — like a credit rating for AI conduct. Today's headline measurements:
| Axis | Instrument | Current measured score |
|---|---|---|
| G · Governance refusal | EAT care gate (76 items) | recall 100% · over-block 0% |
| G · Model two-sided | Flywheel (qwen2.5:0.5b) | TPR 0.38 · FPR 0.00 · 376.6 tokens/correct |
| S · Provenance survival | ProvBench | 0 / 20 markings survived (Article 50) |
| P · Substrate | Free-tier only (HF · Kaggle T4 · Groq · Ollama) | no paid GPU in the measurement path |
| C · Transparency | All artefacts signed Ed25519 | public key at /.well-known/agent.json |
| Tier | What you get | Price shape |
|---|---|---|
| Free | Public benchmark dashboard + scorecard at csoai.org | £0 |
| Pro | API access to live scores (per-query) | $0.01 / query |
| Enterprise | Private protocol audit for a bank / health / public-sector stack | $50K / audit |
| Certification | "SOV Protocol Compliant" badge on an MCP / harness | $5K / cert |
The moat: you run the tests, you own the data, everyone else cites your scores. This is the product the harness was built for — not a model, not a dashboard.