Arena — 12 greenfields

Every greenfield, and how much of its chain is actually complete: items published and loadable, a card with a licence, a Hugging Face Space, a runnable Space, Kaggle, a page on this site, a runnable page, an lm-eval task, an Inspect task, and a measurement against a named model. Measured live, not asserted.

63/130 = 48% chain complete. The macro-F1 column is the best score any measured run reached, and it names the model that reached it — open a row to see every run on that greenfield. Colour follows that score; grey means nothing is measured yet. Scored against 3 models: falcon3:7b, qwen2.5:1.5b, sov34:latest.

greenfieldnmacro-F1itemscardSpaceSpace toolKagglepagepage toollm-evalInspectmeasuredchaintool
GovBench
governance · Brussels
237 MEASURED 8/10 Run it →
7 scored runs · 3 dropped
modelmacro-F1accuracy 95% CIunreadablenharness
sov34:latest0.3810.515 [0.451, 0.578]0%237measure_full.py
falcon3:7b0.3530.426 [0.365, 0.490]0%237measure_full.py
qwen2.5:1.5b0.2910.430 [0.369, 0.494]0%237measure_full.py
falcon3:7b0.468n<300%24measure_robust2.py
sov34:latest0.345n<3025%24measure_robust2.py
qwen2.5:1.5b0.343n<300%24measure_robust2.py
sov34:latest0.386n<304%measure.py

3 runs dropped — sov34:latest, falcon3:7b, qwen2.5:1.5b at 97–100% instrument error (all URLError). A dropped connection is not a wrong answer, so these contribute no score.

ProvBench
provenance · San Francisco
32 MEASURED 8/10 Run it →
7 scored runs · 3 dropped
modelmacro-F1accuracy 95% CIunreadablenharness
falcon3:7b0.574n<3025%16measure_full.py
sov34:latest0.348n<3012%16measure_full.py
qwen2.5:1.5b0.304n<300%16measure_full.py
falcon3:7b0.643n<3013%15measure_robust2.py
sov34:latest0.512n<3027%15measure_robust2.py
qwen2.5:1.5b0.348n<300%15measure_robust2.py
sov34:latest0.273n<3073%measure.py

3 runs dropped — sov34:latest, falcon3:7b, qwen2.5:1.5b at 97–100% instrument error (all URLError). A dropped connection is not a wrong answer, so these contribute no score.

DefBench
safety · London
36 MEASURED 7/10 Run it →
7 scored runs · 3 dropped
modelmacro-F1accuracy 95% CIunreadablenharness
qwen2.5:1.5b1.000n<300%14measure_full.py
qwen2.5:1.5b0.775n<300%14measure_robust2.py
falcon3:7b0.733n<307%14measure_full.py
sov34:latest0.523n<3064%14measure_full.py
falcon3:7b0.300n<307%14measure_robust2.py
sov34:latest0.111n<3086%14measure_robust2.py
sov34:latest0.653n<307%measure.py

3 runs dropped — sov34:latest, falcon3:7b, qwen2.5:1.5b at 97–100% instrument error (all URLError). A dropped connection is not a wrong answer, so these contribute no score.

PQCBench
continuity · Gaithersburg
33 MEASURED 7/10 Run it →
7 scored runs · 3 dropped
modelmacro-F1accuracy 95% CIunreadablenharness
falcon3:7b0.3920.455 [0.298, 0.620]0%33measure_full.py
sov34:latest0.2930.333 [0.198, 0.504]18%33measure_full.py
qwen2.5:1.5b0.2830.364 [0.222, 0.534]0%33measure_full.py
qwen2.5:1.5b0.489n<300%13measure_robust2.py
falcon3:7b0.369n<300%13measure_robust2.py
sov34:latest0.056n<3023%13measure_robust2.py
sov34:latest0.217n<3031%measure.py

3 runs dropped — sov34:latest, falcon3:7b, qwen2.5:1.5b at 97–100% instrument error (all URLError). A dropped connection is not a wrong answer, so these contribute no score.

MCPBench
conformance · Seattle
35 MEASURED 7/10 Run it →
7 scored runs · 3 dropped
modelmacro-F1accuracy 95% CIunreadablenharness
falcon3:7b0.718n<300%11measure_robust2.py
falcon3:7b0.718n<300%11measure_full.py
qwen2.5:1.5b0.686n<300%11measure_robust2.py
qwen2.5:1.5b0.686n<300%11measure_full.py
sov34:latest0.451n<3018%11measure_robust2.py
sov34:latest0.451n<3018%11measure_full.py
sov34:latest0.667n<309%measure.py

3 runs dropped — sov34:latest, falcon3:7b, qwen2.5:1.5b at 97–100% instrument error (all URLError). A dropped connection is not a wrong answer, so these contribute no score.

OSSBench
openness · Portland
32 MEASURED 7/10 Run it →
7 scored runs · 3 dropped
modelmacro-F1accuracy 95% CIunreadablenharness
falcon3:7b0.812n<300%16measure_full.py
sov34:latest0.683n<3019%16measure_full.py
qwen2.5:1.5b0.600n<300%16measure_full.py
falcon3:7b0.764n<300%13measure_robust2.py
sov34:latest0.675n<300%13measure_robust2.py
qwen2.5:1.5b0.575n<300%13measure_robust2.py
sov34:latest0.500n<3015%measure.py

3 runs dropped — sov34:latest, falcon3:7b, qwen2.5:1.5b at 97–100% instrument error (all URLError). A dropped connection is not a wrong answer, so these contribute no score.

CareBench
care · London
1491 MEASURED 4/10 no tool yet
ConductBench
conduct · Brussels
36 MEASURED 4/10 Run it →
MachBench
machinery · Brussels
33 MEASURED 4/10 Run it →
SwarmBench
swarm · unanchored
40 MEASURED 3/10 Run it →
XRBench
cross-reality · unanchored
32 MEASURED 2/10 Run it →
DetBench
detector · Brussels
33 MEASURED 2/10 Run it →
JailBench
runtime-containment · (unanchored — Atlantic)
18 UNMEASURED 0/10 no tool yet
What the numbers do not say. No axis reaches usable_n = 30, so no confidence interval is publishable on any of them — including by us. A greenfield marked SPEC or DRAFT has no score at all: the protocol is published so a harness can consume it, and that is the whole claim. A run whose instrument error rate exceeded threshold is listed as dropped and contributes no score, because a dropped connection is not a wrong answer and must never be published as one. Measurement, not certification, and not legal advice.

Generated from arena.json on 2026-08-06. CSOAI Ltd · UK 16939677 · huggingface.co/csoai