Skip to main content
When you evaluate a Coding Agent, checking one reply is not enough. You normally give it a real project, let it read and write files, run commands, and commit changes, then check whether the artifact is correct. NiceEval’s current recommended approach is an ordinary .eval.ts file that explicitly prepares a Sandbox workspace, sends the task, and asserts the result. For a runnable reference, see coding-agent-skill (a separate repository).
workspaces/ holds the starting project visible to the Agent. .eval.ts controls when the Workspace uploads, the task prompt, and result checks. experiments/ select the Agent, model, and whether to inject a Skill or plugin.

Prepare the workspace

Upload the starting project to the Sandbox in the eval:
You can also use t.sandbox.writeText() to add seed files in one eval, which suits small tasks or security probes.

Write the task prompt

A prompt should read like a real work order. Describe the goal and constraints, but do not disclose the verification answer.
Leave implicit quality requirements such as path-traversal protection or reuse of existing tools for the verification phase. That is how you test whether a Skill or plugin actually helps the Agent fill in context on its own.

Verify files and code

t.sandbox.fileChanged() is a scoped assertion: after the run ends, it uses the Agent-attributed diff to determine whether the target file truly changed.

Run project tests or probe scripts

For Coding Agent tasks, real commands are usually more reliable than text matching:
You can also run the project’s own test, lint, or build scripts:

Use Experiments for a baseline

Do not put an Agent name, plugin name, or URL in positional arguments. The Experiment selects the Agent and fixes run configuration. CLI eval positionals filter eval IDs only.
A typical A/B setup:
  • experiments/baseline.ts: an Agent with no extra configuration.
  • experiments/with-skill.ts: the same Agent with a setup Hook that injects CLAUDE.md.
  • Both groups use the same eval set, model, attempts, and budget.

When to use a Sandbox Fixture

  • The Agent needs to modify real files.
  • You need to run the project’s tests, build, or lint.
  • You need to compare diffs.
  • You need to evaluate a Coding Agent such as Claude Code or Codex.
  • You need to compare whether a Skill, plugin, or Hook improves real tasks.
If you only evaluate a direct Agent application, a normal direct Adapter is lighter.