COUNCIL OF SOVEREIGN AI
the measurement body for AI agent compliance with statute

The coverage gap map

349 enumerated provisions × 4 axes (G, S, P, C) + care_cost lens = 1,312 cells. The finding: the field has measured almost none of them.


Field coverage — the headline

1,301 of 1,312

cells have no measurement in any known benchmark. This is the blind spot the instrument exists to make visible.

99.2% blind · 11 field-evidenced cells, each named.

Internal only — never the headline: 3 of 1312 cells have been measured by GSPC. Reported here only to disambiguate; the product is the map of the field's obligation-space blind spots, not our coverage of them.

Blind cells by axis

Each row is one axis. The filled portion is the cells with any field measurement; the empty portion is what the instrument renders as blind.

G · 4 of 328 evidenced (1.2%)324 blind
4
324 blind
S · 3 of 328 evidenced (0.9%)325 blind
3
325 blind
P · 2 of 328 evidenced (0.6%)326 blind
2
326 blind
C · 2 of 328 evidenced (0.6%)326 blind
2
326 blind
care_cost · 0 of 1 evidenced (0.0%)1 blind
1 blind

Five reasons a cell is blind

ReasonWhat it means
no_benchmarkNo benchmarkNo instrument exists in the field for this cell.
wrong_granularityWrong granularityField benchmarks exist but at category, not provision, granularity.
speaker_onlySpeaker onlyBenchmark asks 'would the model answer compliantly?' — not 'would it act compliantly?'
bare_model_onlyBare model onlyBenchmark scores a base model, not a deployed agent.
judgement_basedJudgement-basedBenchmark uses LLM-as-judge; not deterministic; rejected by Law 1.

The 11 field-evidenced cells

Every covered cell below cites its source benchmark in that source's own units, with licence. Each is named so a reviewer can find it.

ProvisionAxisSourceLicenceGSPC?
Art 5(1)(c) — social scoring
EU AI Act
SAIR-Bench (regulation-derived refusal)research-usemeasured
Art 5(1)(f) — emotion inference at work
EU AI Act
SAIR-Benchresearch-usemeasured
Art 14 — human oversight
EU AI Act
GBench-2-CoP (human-oversight coverage at category granularity)research-usenot measured