MIKAZUKI

ADVERSARIAL CROSS-MODAL AGENT BENCHMARK

Measuring whether OpenClaw can trust its eyes over its database: cross-modal, multi-tool reconciliation under adversarial conditions.

OpenClaw-Bench Samples is a curated 20-task multimodal evaluation set for measuring whether an agent can trust perceptual truth over a stale digital record. Each task pits an atomic cross-modal artefact (a photo, PDF, or audio clip) against out-of-date SaaS state inside a 101-mock-service sandbox, then grades the agent under a dual-signal reward: a rubric LLM-judge council plus a pytest verifier suite, blended into a single score.

  • 20 tasks split enterprise (12) / prosumer (8) across 8 personas
  • Dual-signal grading: rubric LLM-judge council + pytest verifier suite
  • Adversarial by construction: stale-state traps, poison-pill distractors, injection surface, human-in-the-loop guardrails
  • Reference agent: Claude Opus 4.8 (single-run baseline per task)
  • 7-file persona memory shared across tasks: IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS.md

Four numbers that define the scope of OpenClaw-Bench.

Total Tasks

20

12 enterprise · 8 prosumer

Rubric Criteria

364

250 positive · 114 negative (31%)

Mock SaaS Services

101

full catalog available every task

Personas

8

4 enterprise · 4 prosumer · 7-file memory

Four design pressures every OpenClaw-Bench task exerts on the agent.

01

Stale-State Traps

Digital records disagree with the world. A photo, PDF, or audio clip is the source of truth; the SaaS record is out-of-date.

02

Poison-Pill Distractors

Requests carry plausible-looking off-topic sub-goals. The agent must stay inside the scope red-line to preserve reward.

03

Injection Surface

Prompt content, file metadata, and tool responses may attempt to redirect the agent. Rubric penalises taking the bait.

04

Human-in-the-Loop

Side-effects (sends, deletes, payments) are gated. The agent must ask before acting on irreversible operations.

The method

Three phases turn a real request into a blended reward.

Phase 01

Task Design

  • Author persona-anchored requests with 7-file memory (IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS)
  • Embed stale digital state and poison-pill distractors
  • Ship §Focal Event / §Solve Path / §Value Lock ground truth
  • Gate side-effects behind human-in-the-loop approval

Phase 02

Agent Execution

  • Deploy the agent against the 101-service mock-SaaS catalog with its persona memory
  • Atomic multimodal ask: one artefact contradicts the digital record
  • Enterprise segments load higher-difficulty domain contexts; prosumer skews personal
  • Reference baseline: Claude Opus 4.8 (single run per task)

Phase 03

Blended Reward

  • Rubric LLM-judge council scores against 364 criteria across 6 dimensions
  • Weights: +5 / +3 / +1 for positive · -1 / -3 / -5 for negative
  • Pytest verifier suite runs against the task workspace
  • Combined reward aggregated across tasks and segments

We ran Claude Opus 4.8 across all 20 tasks as the reference baseline.

Numbers below are recomputed dynamically from the dataset.

Enterprise · Claude Opus 4.8 - Loading...
Prosumer · Claude Opus 4.8 - Loading...
Fig. 1. Reference reward distribution across the 20-task set (Claude Opus 4.8, single run per task).
Hard -
- - Enterprise
- - Prosumer
- - Combined
Medium -
- - Enterprise
- - Prosumer
- - Combined
Easy -
- - Enterprise
- - Prosumer
- - Combined
Hard -
- Enterprise
- Prosumer
- Mean reward
Medium -
- Enterprise
- Prosumer
- Mean reward
Easy -
- Enterprise
- Prosumer
- Mean reward
Loading...
Task Segment Difficulty Persona / Domain Reference reward Rubric

Head-to-head shape of Enterprise vs. Prosumer under the Claude Opus 4.8 reference run.

Enterprise

Tasks
12
Reference mean
47.4
Personas
4 (Shanice Giles · Joyce Padilla · Michelle Brooks · Dorothy Langston)
SaaS per task
101 (full catalog)
Skew
domain contexts, higher-difficulty asks

Prosumer

Tasks
8
Reference mean
53.0
Personas
4 (Jessica Webb · Amanda Tran · Carlton McGee · Nicole King)
SaaS per task
101 (full catalog)
Skew
personal-life asks, medium-difficulty skew