Built-in Sandbox Agents
claude-code
Runs the Anthropic Claude Code CLI and requires
ANTHROPIC_API_KEY.codex
Runs the OpenAI Codex CLI and requires
CODEX_API_KEY.bub
Runs the bub coding agent and requires
BUB_API_KEY plus BUB_API_BASE.Run a built-in Agent
Setsandbox on an Eval or Experiment to choose where NiceEval creates the isolated environment:
There is no matching CLI flag or project-wide default provider. If neither the Eval nor Experiment contributes a template-bearing layer, link planning fails before NiceEval creates a Sandbox.
NICEEVAL_CLAUDE_CODE_E2B_TEMPLATE for Claude Code or NICEEVAL_BUB_E2B_TEMPLATE for bub. Each constant is a complete, version-pinned reference whose version follows the Agent in that template. See Sandbox Providers: Prebuilt Environments and Runtime Checkpoints to add system packages, binaries, or model caches.
Built-in Agents exported from niceeval/adapter are factories. Put authentication, proxies, MCP, or GitHub skills in the factory arguments. Keep the model in the Experiment’s model field and the provider in its sandbox field:
Agent environment variables
These are each Agent’s official variable names. CI can keep only
OPENAI_API_KEY and OPENAI_BASE_URL as shared secrets, but the Experiment must explicitly map them through apiKey, baseUrl, or apiBase. That mapping does not change the public environment variable names that each Adapter reads.
Workflow
test(t). The Agent can see only files you wrote into the Sandbox.
t.sandbox.fileChanged() and the other attribution assertions cover only files the Agent changed during t.send(). NiceEval records workspace state before and after each t.send(), so the changes between those points belong to the Agent. Starting files you uploaded and verification materials you write after t.send() are not included. fileChanged("src/app.ts") passes only when the Agent really changed that file.
Requirements for collecting file changes
NiceEval creates a private Git repository in the Sandbox. It records workspace state before and after eacht.send(); the difference is the Agent’s work. Collection requires all of the following:
- The Sandbox has Git and a POSIX shell. Built-in provider images include them. Verify them for a custom image. The project under test does not need to be a Git repository: NiceEval’s private repository lives outside the workdir, does not touch your
.git, and is not visible to the Agent. - Files are written in the workdir. Collection covers
sandbox.workdir. OmittargetDirorcwdwhen writing files. For an absolute path, readsandbox.workdir; do not hard-code/workspace. Files outside the workdir are neither visible to the Agent nor collected. - The workdir root has no nested Git repository. Put the checkout directly in the workdir root. A submodule or another clone there is an execution error that lists the path, because normal file changes inside it would disappear from the record. Exclude a repository that does not participate in scoring with
diff.ignore. - The change happens during
t.send(). Files written after the finalt.send()are your verification material, not Agent work.
.git, node_modules, __pycache__, Python virtual environments, common build output, and package-manager caches by default. Otherwise one npm install could produce tens of thousands of paths. Adjust that list with an Eval’s diff field:
/ matches a name at every depth; one with / starts at the workdir root; a trailing / denotes a directory. The project’s own .gitignore does not participate, so project-ignored files are still recorded. NiceEval freezes the list at the first snapshot; a later Agent edit to .gitignore cannot affect it.
One t.send() window can scan at most 100,000 paths and 256 MiB of deduplicated text. Within that scan, File Changes retains at most 10,000 changes; when the candidate set is larger, it saves the deterministic prefix and marks the collection partial with collection-cap-reached. Exceeding the scan limit is an execution error, never an empty change set that could let a file assertion pass. Narrow unusually large workspaces with diff.include or diff.ignore.
Experiments, Eval Groups, Evals, and Sandbox Agents can all contribute command-only .before() actions. When an Agent writes its own .env, put that action on the Agent’s sandbox field and give it a larger changeFrequency so fixed tools and Fixtures sort first. Files written by these actions are environment state, not part of the Agent diff. See Sandbox Providers: Prepare the Sandbox in a deterministic order for the API and rules.
Create a custom Sandbox Agent
createNpmCliInstaller prepares the npm package on the host, sends it into the Sandbox, and installs an ordinary package there with npm. Such a Sandbox needs Node.js and npm; only a self-contained platformPackage can omit them. The complete workflow is in Create a Custom Sandbox Agent.
setup, every send, and teardown each receive feedback methods for their own scope:
ctx.progress({ message, current?, total? })updates short-lived status. Use it for CLI installation, a Turn, or transcript reading. Do not call it for every token or JSONL frame.ctx.diagnostic({ code, level, message, data?, dedupeKey? })records protocol degradation, missing transcripts, and cleanup problems. It enters the terminal’s permanent event stream and is committed as a channel event with the Attempt.- Throw when execution cannot continue. The Runner writes the stage, code, message, cause, and stack to a diagnostic channel, then forms an
erroredVerdict in theniceeval.verdictchannel.
console.log or console.error, or write process.stdout or process.stderr, from an Adapter. That breaks up the human dashboard and CI log order.
When a Sandbox or Adapter error appears during a run, the terminal prints an Attempt identity:
niceeval show @<attempt-locator> to inspect the structured error, diagnostics, and completed lifecycle stages. For AI, scripts, and CI, first use niceeval query discover, then send the fixed attempt.trace request to query explain or query run. The fixed result reads timing facts including queueing, Sandbox startup, shell work in setup and teardown, Agent CLI installation and startup, each send, correlated OTel model or tool work, and cleanup. It identifies the layer where an error or timeout occurred and where earlier time went.
When the tree has more than 80 detail nodes, it retains failures, slow points, and head and tail samples, then reports how many nodes it omitted. Use the complete fixed trace operation to audit every node. Sandbox creation can fail before telemetry starts, so error review does not depend on a trace.