.eval.ts file that explicitly prepares a Sandbox workspace, sends the task, and asserts the result.
For a runnable reference, see coding-agent-skill (a separate repository).
Recommended directory structure
workspaces/ holds the starting project visible to the Agent. .eval.ts controls when the Workspace uploads, the task prompt, and result checks. experiments/ select the Agent, model, and whether to inject a Skill or plugin.
Prepare the workspace
Upload the starting project to the Sandbox in the eval:t.sandbox.writeText() to add seed files in one eval, which suits small tasks or security probes.
Write the task prompt
A prompt should read like a real work order. Describe the goal and constraints, but do not disclose the verification answer.Verify files and code
t.sandbox.fileChanged() is a scoped assertion: after the run ends, it uses the Agent-attributed diff to determine whether the target file truly changed.
Run project tests or probe scripts
For Coding Agent tasks, real commands are usually more reliable than text matching:Use Experiments for a baseline
Do not put an Agent name, plugin name, or URL in positional arguments. The Experiment selects the Agent and fixes run configuration. CLI eval positionals filter eval IDs only.experiments/baseline.ts: an Agent with no extra configuration.experiments/with-skill.ts: the same Agent with a setup Hook that injectsCLAUDE.md.- Both groups use the same eval set, model, attempts, and budget.
When to use a Sandbox Fixture
- The Agent needs to modify real files.
- You need to run the project’s tests, build, or lint.
- You need to compare diffs.
- You need to evaluate a Coding Agent such as Claude Code or Codex.
- You need to compare whether a Skill, plugin, or Hook improves real tasks.