Erza measures Skill efficacy - not how capable a model is, but how much a curated Agent Skill changes the outcome. Classic agent benchmarks report one absolute score per model. Erza runs the same task twice on the same model, harness and container - once with the Skill injected, once without - and reports the paired difference (Δ). Efficacy is measured, never assumed: Δ is a reported property of every task, never a gate for inclusion. Tasks with zero or negative Δ are kept and labelled, so the benchmark can honestly measure whether training closes the gap.
- Paired evaluation - with-Skills vs no-Skills on an identical model, harness and container
- Deterministic verifiers with a 100%-passing oracle and network-sealed grading
- Skill-to-solution leakage audit - Skills teach procedure, never answers
- Δ reported for every task, including zero and negative - no selection on outcome
- Built as an RL environment: train, and the measured Δ should shrink
One hard bundle, in context
The first published Erza bundle sits in the hard tier - a Natural-Science interlaboratory-metrology task (ERZA-RB1 robust consensus) the model solves 0 of 3 runs unaided, and 3 of 3 with the curated Skill. Below: how Skills move pass rates across Erza domains, and this bundle's real agent cost per run.
Head-to-head, three frontier models
With-Skills pass rate shows how far each model gets when handed the right procedure; the Δ column isolates how much the Skill itself contributed, independent of raw model strength. A strong model with a small Δ was already close; a large Δ means the Skill did real work.
| Model | Provider | Paired runs | No-Skills | With-Skills | Mean score | Δ |
|---|
Where the Skill earns its keep
| Tier |
|---|
| Tier |
|---|
Same task, twice - the Δ is the finding
Each task bundle ships a prompt, a container, a curated Skill, an oracle and a
deterministic verifier. Erza runs the agent twice under matched conditions and scores
both arms with the same network-sealed verifier. The reported result is
Δ = with-Skills − no-Skills.
Task bundle
prompt · container · oracle · verifier
No-Skills run
agent solves it - Skill withheld
With-Skills run
same agent - Skill injected
Verifier
deterministic · network-sealed
Δ efficacy
with-Skills − no-Skills
Where Skills help - by domain
Erza's paired methodology spans a broad domain taxonomy. Across these task domains, curated Skills lift pass rates most in knowledge-dense fields like Natural Science and Media production. Cybersecurity and Software Engineering are excluded here.
| Domain | N | No-Skills | With-Skills | Δ (pp) |
|---|---|---|---|---|
| Natural Science | 14 | 42.0% | 70.8% | +28.8 pp |
| Media & Content Production | 5 | 23.3% | 47.4% | +24.1 pp |
| Industrial & Physical Systems | 14 | 23.9% | 39.6% | +15.7 pp |
| Finance & Economics | 9 | 19.1% | 33.3% | +14.2 pp |
| Office & White Collar | 14 | 40.5% | 53.0% | +12.6 pp |
| Mathematics & OR | 8 | 45.7% | 55.4% | +9.7 pp |
Browse the task bundles
Filter and search the delivery. Expand a row to see per-model no-Skills vs with-Skills runs and the resulting Δ, plus links to the raw trajectory and task file.
| Task | Codebase | Language | Tier | No-Skills | With-Skills | Mean score | Δ |
|---|