2026-08-08 · CSOAI — measurement, not claim
We ran 14 sovereign/open models against the 12-axe GSPC governance battery (care, gov, art5, agi, asi, mcp, oss, prv, mach, det, swarm, xr), scoring each on exact-label classification with frozen prompt + public split. This is the measurement that tells a router where to send a request — not a leaderboard for its own sake.
Mean accuracy across the 12 axes, highest first:
Two axes — gspc-swarm and gspc-art5 — are easy for almost the whole fleet (0.94–0.97). The interesting axes are the hard ones where models genuinely diverge: gspc-asi, gspc-gov, and the small-cap labels. On those, the sovereign family does not collapse to a single winner — the differences are real enough that routing each axis to its measured best beats picking one model for everything.
That is the working hypothesis the sandwich-router architecture is built on: no single model is the best at all 12 governance sub-tasks, so the system routes per-dimension and signs each hop. This soak is the evidence that the hypothesis is testable before it is claimed.
gpt-oss:20b returned no parseable label across the battery (all nulls). That is a
measurement limitation, not a score of zero in the "fails governance" sense. We report it as
unmeasured rather than inflating or hiding it. The full per-model × per-axis table is available on
request / in the benchmark artefact store.