§01 · Overview
OpenClaw SFT Samples is a 20-task long-horizon dataset for supervised fine-tuning of AI agents on the kind of messy, cross-modal, multi-tool work knowledge workers actually do all day. Each task ships a several-KB, turn-marked prompt, a focused 19-to-32-service SaaS sandbox drawn from a 67-service catalog, and structured §Focal-Event / §Solve-Path / §Value-Lock ground truth. Agents are graded under a dual-signal reward: a rubric LLM-judge council plus a pytest verifier suite, blended into a single score.
20 tasks split enterprise (12) / prosumer (8) across 9 unique personas
Turn-marked prompts (--- TURN 1 ---) with scope red-lines the agent must honor
Structured ground truth: §Focal Event / §Solve Path / §Value Lock (canonical + decoy values)
Dual-signal grading: rubric LLM-judge council + pytest verifier suite
Reference agent: Claude Opus 4.8 (single-run baseline per task)
7-file persona memory shared across tasks: IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS.md
At a glance
Tasks 20 (long-horizon SFT)
Segments 12 enterprise · 8 prosumer
Personas 9 unique
Mock SaaS 19–32 per task (67-svc catalog)
Rubric criteria 486 (18% negative)
Reference Claude Opus 4.8
§02 · Key metrics
Four numbers that define the scope of OpenClaw SFT Samples.
Total Tasks
20
12 enterprise · 8 prosumer
Rubric Criteria
486
400 positive · 86 negative (18%)
Mock SaaS Services
19–32
focused subset · 67-service catalog
Personas
9
enterprise + prosumer · 7-file memory
§03 · Adversarial pillars
Four design pressures every OpenClaw SFT Samples task exerts on the agent.
01
Stale-State Traps
Digital records disagree with the world. A photo, PDF, or audio clip is the source of truth; the SaaS record is out-of-date.
02
Poison-Pill Distractors
Requests carry plausible-looking off-topic sub-goals. The agent must stay inside the scope red-line to preserve reward.
03
Injection Surface
Prompt content, file metadata, and tool responses may attempt to redirect the agent. Rubric penalises taking the bait.
04
Human-in-the-Loop
Side-effects (sends, deletes, payments) are gated. The agent must ask before acting on irreversible operations.
§04 · Evaluation pipeline
The method
Three phases turn a real request into a blended reward.
Phase 01
Task Design
Author persona-anchored requests with 7-file memory (IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS)
Embed stale digital state and poison-pill distractors
Ship §Focal Event / §Solve Path / §Value Lock ground truth
Gate side-effects behind human-in-the-loop approval
Phase 02
Agent Execution
Deploy the agent against a focused 19–32-service SaaS subset per task, drawn from a 67-service catalog
Long-horizon workstream ask: several-KB turn-marked prompt, cross-modal artefacts contradict stale SaaS state
Enterprise skews institutional (board, compliance, program); prosumer skews owner-operator
Reference baseline: Claude Opus 4.8 (single run per task)
Phase 03
Blended Reward
Rubric LLM-judge council scores against 486 criteria across 6 dimensions
Weights: +5 / +3 / +1 for positive · -1 / -3 / -5 for negative
Pytest verifier suite runs against the task workspace
Combined reward aggregated across tasks and segments
§05 · Results
We ran Claude Opus 4.8 across all 20 tasks as the reference baseline.
Numbers below are recomputed dynamically from the dataset.
Enterprise · Claude Opus 4.8
-
Loading...
Prosumer · Claude Opus 4.8
-
Loading...
Hard
-
-
-
Enterprise
-
-
Prosumer
-
-
Combined
Medium
-
-
-
Enterprise
-
-
Prosumer
-
-
Combined
Easy
-
-
-
Enterprise
-
-
Prosumer
-
-
Combined
§07 · Dataset viewer
Hard
-
-
Enterprise
-
Prosumer
-
Mean reward
Medium
-
-
Enterprise
-
Prosumer
-
Mean reward
Easy
-
-
Enterprise
-
Prosumer
-
Mean reward
All segments
Enterprise
Prosumer
All difficulties
Hard
Medium
Easy
Reference reward
Task ID (A-Z)
Segment
Difficulty
Rubric size
↓
Loading...
Task
Segment
Difficulty
Persona / Domain
Reference reward
Rubric
§08 · Segment comparison
Head-to-head shape of Enterprise vs. Prosumer under the Claude Opus 4.8 reference run.
Enterprise
Tasks 12
Reference mean 51.5
Personas 5 (Andrew Santos · Colette Mullins · Carlos Moore · Stuart Myer · Daniel Brooks)
SaaS per task 19–32 (focused subset)
Skew institutional workstreams, compliance and program scope
Prosumer
Tasks 8
Reference mean 53.9
Personas 5 (Jake Thornton · Jessica Taylor · Stuart Myer · Jessica Spencer · Catherine Conway)
SaaS per task 19–32 (focused subset)
Skew owner-operator, solo-practitioner, studio work