erza-samples
Agent-Skills efficacy · Δ measured

Measuring Agent-Skill efficacy at the granularity of paired tasks.

Erza measures Skill efficacy - not how capable a model is, but how much a curated Agent Skill changes the outcome. Classic agent benchmarks report one absolute score per model. Erza runs the same task twice on the same model, harness and container - once with the Skill injected, once without - and reports the paired difference (Δ). Efficacy is measured, never assumed: Δ is a reported property of every task, never a gate for inclusion. Tasks with zero or negative Δ are kept and labelled, so the benchmark can honestly measure whether training closes the gap.

  • Paired evaluation - with-Skills vs no-Skills on an identical model, harness and container
  • Deterministic verifiers with a 100%-passing oracle and network-sealed grading
  • Skill-to-solution leakage audit - Skills teach procedure, never answers
  • Δ reported for every task, including zero and negative - no selection on outcome
  • Built as an RL environment: train, and the measured Δ should shrink
With-Skills vs no-Skills pass rate across Erza task domains.

One hard bundle, in context

The first published Erza bundle sits in the hard tier - a Natural-Science interlaboratory-metrology task (ERZA-RB1 robust consensus) the model solves 0 of 3 runs unaided, and 3 of 3 with the curated Skill. Below: how Skills move pass rates across Erza domains, and this bundle's real agent cost per run.

Skill efficacy (Δ) by task domain - Erza.
Agent cost per run - no-Skills vs with-Skills (this bundle, 3 trials each).

Head-to-head, three frontier models

With-Skills pass rate shows how far each model gets when handed the right procedure; the Δ column isolates how much the Skill itself contributed, independent of raw model strength. A strong model with a small Δ was already close; a large Δ means the Skill did real work.

Model Provider Paired runs No-Skills With-Skills Mean score Δ

Where the Skill earns its keep

Tier
Tier

Same task, twice - the Δ is the finding

Each task bundle ships a prompt, a container, a curated Skill, an oracle and a deterministic verifier. Erza runs the agent twice under matched conditions and scores both arms with the same network-sealed verifier. The reported result is Δ = with-Skills − no-Skills.

1

Task bundle

prompt · container · oracle · verifier

2

No-Skills run

agent solves it - Skill withheld

3

With-Skills run

same agent - Skill injected

4

Verifier

deterministic · network-sealed

5

Δ efficacy

with-Skills − no-Skills

Where Skills help - by domain

Erza's paired methodology spans a broad domain taxonomy. Across these task domains, curated Skills lift pass rates most in knowledge-dense fields like Natural Science and Media production. Cybersecurity and Software Engineering are excluded here.

Domain N No-Skills With-Skills Δ (pp)
Natural Science 14 42.0% 70.8% +28.8 pp
Media & Content Production 5 23.3% 47.4% +24.1 pp
Industrial & Physical Systems 14 23.9% 39.6% +15.7 pp
Finance & Economics 9 19.1% 33.3% +14.2 pp
Office & White Collar 14 40.5% 53.0% +12.6 pp
Mathematics & OR 8 45.7% 55.4% +9.7 pp

Browse the task bundles

Filter and search the delivery. Expand a row to see per-model no-Skills vs with-Skills runs and the resulting Δ, plus links to the raw trajectory and task file.

- tasks
Task Codebase Language Tier No-Skills With-Skills Mean score Δ

Explore the sources