The coverage gap map
349 enumerated provisions × 4 axes (G, S, P, C) + care_cost lens = 1,312 cells. The finding: the field has measured almost none of them.
Field coverage — the headline
1,301 of 1,312
cells have no measurement in any known benchmark. This is the blind spot the instrument exists to make visible.
99.2% blind · 11 field-evidenced cells, each named.
Blind cells by axis
Each row is one axis. The filled portion is the cells with any field measurement; the empty portion is what the instrument renders as blind.
G · 4 of 328 evidenced (1.2%)324 blind
4
324 blind
S · 3 of 328 evidenced (0.9%)325 blind
3
325 blind
P · 2 of 328 evidenced (0.6%)326 blind
2
326 blind
C · 2 of 328 evidenced (0.6%)326 blind
2
326 blind
care_cost · 0 of 1 evidenced (0.0%)1 blind
1 blind
Five reasons a cell is blind
| Reason | What it means |
|---|---|
no_benchmark | No benchmark — No instrument exists in the field for this cell. |
wrong_granularity | Wrong granularity — Field benchmarks exist but at category, not provision, granularity. |
speaker_only | Speaker only — Benchmark asks 'would the model answer compliantly?' — not 'would it act compliantly?' |
bare_model_only | Bare model only — Benchmark scores a base model, not a deployed agent. |
judgement_based | Judgement-based — Benchmark uses LLM-as-judge; not deterministic; rejected by Law 1. |
The 11 field-evidenced cells
Every covered cell below cites its source benchmark in that source's own units, with licence. Each is named so a reviewer can find it.
| Provision | Axis | Source | Licence | GSPC? |
|---|---|---|---|---|
| Art 5(1)(c) — social scoring EU AI Act | S | AIR-Bench (regulation-derived refusal) | research-use | measured |
| Art 5(1)(f) — emotion inference at work EU AI Act | S | AIR-Bench | research-use | measured |
| Art 14 — human oversight EU AI Act | G | Bench-2-CoP (human-oversight coverage at category granularity) | research-use | not measured |