§01 · Overview
OpenClaw-Bench Samples is a curated 20-task multimodal evaluation set for measuring whether an agent can trust perceptual truth over a stale digital record. Each task pits an atomic cross-modal artefact (a photo, PDF, or audio clip) against out-of-date SaaS state inside a 101-mock-service sandbox, then grades the agent under a dual-signal reward: a rubric LLM-judge council plus a pytest verifier suite, blended into a single score.
20 tasks split enterprise (12) / prosumer (8) across 8 personas
Dual-signal grading: rubric LLM-judge council + pytest verifier suite
Adversarial by construction: stale-state traps, poison-pill distractors, injection surface, human-in-the-loop guardrails
Reference agent: Claude Opus 4.8 (single-run baseline per task)
7-file persona memory shared across tasks: IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS.md
At a glance
Tasks 20
Segments 12 enterprise · 8 prosumer
Personas 8 across 8 domains
Mock SaaS 101 per task
Rubric criteria 364 (31% negative)
Reference Claude Opus 4.8
§02 · Key metrics
Four numbers that define the scope of OpenClaw-Bench.
Total Tasks
20
12 enterprise · 8 prosumer
Rubric Criteria
364
250 positive · 114 negative (31%)
Mock SaaS Services
101
full catalog available every task
Personas
8
4 enterprise · 4 prosumer · 7-file memory
§03 · Adversarial pillars
Four design pressures every OpenClaw-Bench task exerts on the agent.
01
Stale-State Traps
Digital records disagree with the world. A photo, PDF, or audio clip is the source of truth; the SaaS record is out-of-date.
02
Poison-Pill Distractors
Requests carry plausible-looking off-topic sub-goals. The agent must stay inside the scope red-line to preserve reward.
03
Injection Surface
Prompt content, file metadata, and tool responses may attempt to redirect the agent. Rubric penalises taking the bait.
04
Human-in-the-Loop
Side-effects (sends, deletes, payments) are gated. The agent must ask before acting on irreversible operations.
§04 · Evaluation pipeline
The method
Three phases turn a real request into a blended reward.
Phase 01
Task Design
Author persona-anchored requests with 7-file memory (IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS)
Embed stale digital state and poison-pill distractors
Ship §Focal Event / §Solve Path / §Value Lock ground truth
Gate side-effects behind human-in-the-loop approval
Phase 02
Agent Execution
Deploy the agent against the 101-service mock-SaaS catalog with its persona memory
Atomic multimodal ask: one artefact contradicts the digital record
Enterprise segments load higher-difficulty domain contexts; prosumer skews personal
Reference baseline: Claude Opus 4.8 (single run per task)
Phase 03
Blended Reward
Rubric LLM-judge council scores against 364 criteria across 6 dimensions
Weights: +5 / +3 / +1 for positive · -1 / -3 / -5 for negative
Pytest verifier suite runs against the task workspace
Combined reward aggregated across tasks and segments
§05 · Results
We ran Claude Opus 4.8 across all 20 tasks as the reference baseline.
Numbers below are recomputed dynamically from the dataset.
Enterprise · Claude Opus 4.8
-
Loading...
Prosumer · Claude Opus 4.8
-
Loading...
Click to expand
Fig. 1. Reference reward distribution across the 20-task set (Claude Opus 4.8, single run per task).
Hard
-
-
-
Enterprise
-
-
Prosumer
-
-
Combined
Medium
-
-
-
Enterprise
-
-
Prosumer
-
-
Combined
Easy
-
-
-
Enterprise
-
-
Prosumer
-
-
Combined
§07 · Dataset viewer
Hard
-
-
Enterprise
-
Prosumer
-
Mean reward
Medium
-
-
Enterprise
-
Prosumer
-
Mean reward
Easy
-
-
Enterprise
-
Prosumer
-
Mean reward
All segments
Enterprise
Prosumer
All difficulties
Hard
Medium
Easy
Reference reward
Task ID (A-Z)
Segment
Difficulty
Rubric size
↓
Loading...
Task
Segment
Difficulty
Persona / Domain
Reference reward
Rubric
§08 · Segment comparison
Head-to-head shape of Enterprise vs. Prosumer under the Claude Opus 4.8 reference run.
Enterprise
Tasks 12
Reference mean 47.4
Personas 4 (Shanice Giles · Joyce Padilla · Michelle Brooks · Dorothy Langston)
SaaS per task 101 (full catalog)
Skew domain contexts, higher-difficulty asks
Prosumer
Tasks 8
Reference mean 53.0
Personas 4 (Jessica Webb · Amanda Tran · Carlton McGee · Nicole King)
SaaS per task 101 (full catalog)
Skew personal-life asks, medium-difficulty skew