DOJIGIRI

LONG-HORIZON SFT SAMPLES · CROSS-MODAL AGENT DATASET

Can OpenClaw hold a multi-turn workstream together across dense, turn-marked context, honoring scope red-lines while reconciling perceptual truth against a stale digital record?

OpenClaw SFT Samples is a 20-task long-horizon dataset for supervised fine-tuning of AI agents on the kind of messy, cross-modal, multi-tool work knowledge workers actually do all day. Each task ships a several-KB, turn-marked prompt, a focused 19-to-32-service SaaS sandbox drawn from a 67-service catalog, and structured §Focal-Event / §Solve-Path / §Value-Lock ground truth. Agents are graded under a dual-signal reward: a rubric LLM-judge council plus a pytest verifier suite, blended into a single score.

  • 20 tasks split enterprise (12) / prosumer (8) across 9 unique personas
  • Turn-marked prompts (--- TURN 1 ---) with scope red-lines the agent must honor
  • Structured ground truth: §Focal Event / §Solve Path / §Value Lock (canonical + decoy values)
  • Dual-signal grading: rubric LLM-judge council + pytest verifier suite
  • Reference agent: Claude Opus 4.8 (single-run baseline per task)
  • 7-file persona memory shared across tasks: IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS.md

Four numbers that define the scope of OpenClaw SFT Samples.

Total Tasks

20

12 enterprise · 8 prosumer

Rubric Criteria

486

400 positive · 86 negative (18%)

Mock SaaS Services

19–32

focused subset · 67-service catalog

Personas

9

enterprise + prosumer · 7-file memory

Four design pressures every OpenClaw SFT Samples task exerts on the agent.

01

Stale-State Traps

Digital records disagree with the world. A photo, PDF, or audio clip is the source of truth; the SaaS record is out-of-date.

02

Poison-Pill Distractors

Requests carry plausible-looking off-topic sub-goals. The agent must stay inside the scope red-line to preserve reward.

03

Injection Surface

Prompt content, file metadata, and tool responses may attempt to redirect the agent. Rubric penalises taking the bait.

04

Human-in-the-Loop

Side-effects (sends, deletes, payments) are gated. The agent must ask before acting on irreversible operations.

The method

Three phases turn a real request into a blended reward.

Phase 01

Task Design

  • Author persona-anchored requests with 7-file memory (IDENTITY, HEARTBEAT, MEMORY, USER, SOUL, AGENTS, TOOLS)
  • Embed stale digital state and poison-pill distractors
  • Ship §Focal Event / §Solve Path / §Value Lock ground truth
  • Gate side-effects behind human-in-the-loop approval

Phase 02

Agent Execution

  • Deploy the agent against a focused 19–32-service SaaS subset per task, drawn from a 67-service catalog
  • Long-horizon workstream ask: several-KB turn-marked prompt, cross-modal artefacts contradict stale SaaS state
  • Enterprise skews institutional (board, compliance, program); prosumer skews owner-operator
  • Reference baseline: Claude Opus 4.8 (single run per task)

Phase 03

Blended Reward

  • Rubric LLM-judge council scores against 486 criteria across 6 dimensions
  • Weights: +5 / +3 / +1 for positive · -1 / -3 / -5 for negative
  • Pytest verifier suite runs against the task workspace
  • Combined reward aggregated across tasks and segments

We ran Claude Opus 4.8 across all 20 tasks as the reference baseline.

Numbers below are recomputed dynamically from the dataset.

Enterprise · Claude Opus 4.8 - Loading...
Prosumer · Claude Opus 4.8 - Loading...
Hard -
- - Enterprise
- - Prosumer
- - Combined
Medium -
- - Enterprise
- - Prosumer
- - Combined
Easy -
- - Enterprise
- - Prosumer
- - Combined
Hard -
- Enterprise
- Prosumer
- Mean reward
Medium -
- Enterprise
- Prosumer
- Mean reward
Easy -
- Enterprise
- Prosumer
- Mean reward
Loading...
Task Segment Difficulty Persona / Domain Reference reward Rubric

Head-to-head shape of Enterprise vs. Prosumer under the Claude Opus 4.8 reference run.

Enterprise

Tasks
12
Reference mean
51.5
Personas
5 (Andrew Santos · Colette Mullins · Carlos Moore · Stuart Myer · Daniel Brooks)
SaaS per task
19–32 (focused subset)
Skew
institutional workstreams, compliance and program scope

Prosumer

Tasks
8
Reference mean
53.9
Personas
5 (Jake Thornton · Jessica Taylor · Stuart Myer · Jessica Spencer · Catherine Conway)
SaaS per task
19–32 (focused subset)
Skew
owner-operator, solo-practitioner, studio work