Sandbox interface, so an Adapter does not need to know whether the current provider is local Docker, a Vercel micro-VM, E2B, or a third-party cloud service.
The Sandbox interface
Common Adapter operations include:
The two
OrThrow methods include a sanitized, redacted, and truncated stderr tail in the error message. When stderr is empty, they fall back to stdout. For troubleshooting, start with that summary and read SandboxCommandExitError.result when you need complete output.
Choose a provider
Set thesandbox field in the experiment:
A missing SDK does not fail silently: NiceEval errors the moment it creates the sandbox and prints the install command above, for example
Docker sandbox requires 'dockerode'. Install it with: pnpm add dockerode @types/dockerode.
Docker Compose: workspaceService names the main Sandbox
When a task needs more than one container — an app container plus a database, say — declare the whole Compose environment with dockerComposeSandbox:
workspaceService names which service in the Compose file is the main Sandbox. The agent, t.sandbox commands and file operations, workdir, and the change diff all land on that one container. Every other service in the Compose file — a database, a mock server, whatever — is an accompanying resource that never goes through the Sandbox interface: you cannot runCommand against it, upload or read files on it, or see its diff. To reach an accompanying service, rely on the task’s own networking — for example, reaching db:5432 from the client container over Compose’s built-in DNS.
Let an Agent run Docker itself
When an Agent needs to rundocker build, docker run, or docker compose, you can explicitly mount an existing Unix socket or choose raw privileged or managed rootless Docker-in-Docker. For the use cases, Dockerfiles, and complete configuration for all three modes, see Let a Sandbox Use Docker.
Prepare the Sandbox in a deterministic order
Every Provider factory returns an immutableSandboxLayer. Use .before() for work that must finish before the Agent starts and .after() for unconditional cleanup:
changeFrequency is any finite non-negative number. Dependencies are satisfied first; among ready actions the lowest number runs first. The presets are rare = 10, normal = 100, and frequent = 1000. Experiment, Eval Group, Eval, and Sandbox Agent actions share this one DAG. Owner order only breaks equal-frequency ties, so a low-frequency Agent action can run before an Experiment or Eval action.
All current SandboxLayer actions have attempt occurrence. Promoting a proven prefix to a physical-instance occurrence is deferred performance work; author code cannot request or infer that promotion.
Define a reusable Action
The built-incommand(), shell(), writeText(), writeBytes(), uploadFile(), uploadDirectory(), and gitCheckout() functions are all SandboxAction families. Projects can define a family through the same public protocol:
cache: { fingerprint?: JsonValue } option. Their .after() forms do not cache and do not accept cache.
Ordinary JSON, strings, and text are author-declared non-sensitive data. NiceEval does not promise taint analysis for process.env, closures, or arbitrary strings. Keep secrets and credentials in runtime callbacks or Provider-private bindings.
Write, upload, and check out Git content
uploadFile() and uploadDirectory() take a module-relative URL, not a path that changes with the process working directory. gitCheckout() requires a credential-free HTTPS URL and a full commit object id. Bytes and manifests enter the automatic fingerprint.
A frozen pnpm install must register the lockfile content through inputs, as in the first example. An unlocked apt index, latest, a moving URL, time, randomness, or another network-dependent operation is not deterministic input; put that work in an opaque callback barrier.
Dynamic callbacks and cleanup
Use.before(async (sandbox, context) => ...) when a decision can only be made at runtime. The callback always executes and acts as an opaque cache barrier. Register acquired resources immediately with context.onCleanup():
sandbox.sandboxId is the Provider-native ID of the current physical Sandbox. Treat it as opaque unless this layer is deliberately bound to a known Provider; trusted host code may then pass it to that Provider SDK to acquire Attempt-scoped auxiliary resources. Those resources are outside NiceEval’s managed Case topology and cache, so register idempotent cleanup immediately after acquisition. A setup prefix may replace the instance before the first callback starts; once it starts, that Attempt’s cleanup and after callbacks observe the same ID.
Use .after() for unconditional, idempotent cleanup:
after is registered when its Attempt occurrence begins. Dynamic cleanup joins that same stack at the moment it is registered, and the complete stack runs in LIFO order. Cleanup keeps running after an earlier cleanup failure. context.progress() reports short-lived status; context.diagnostic() preserves a problem for later review. Throw when the environment cannot continue.
A Direct Agent created with defineAgent has no Sandbox preparation path. A Sandbox Agent may contribute a command-only layer, but cannot select or replace the Provider template.
Prebuilt environments and runtime checkpoints
NiceEval references prebuilt environments through a typed spec, but does not offer a fake universal build command:Build on the official baselines to speed things up
Content that is stable, large, and identical for every attempt — system packages, Agent CLIs, compiled binaries, large model caches — should be baked into the provider’s publishable artifact before you run evals, so every attempt starts from a prebuilt environment and skips the runtime install. All three built-in providers can derive from an official baseline, so you never have to install the agent from a blank environment. Their build tooling, credentials, and publishing semantics differ, though, so NiceEval only unifies how you consume the artifact ID (image / snapshotId / template) and does not fabricate a cross-provider build DSL.
Adapters and Sandboxes do not guess each other’s configuration: the Adapter checks for the CLI it needs, and the Sandbox spec picks the provider and the artifact. Claude Code and Codex fall back to a runtime install when the CLI is missing, so baking them in is purely a speed-up. The default Bub 0.4.0 recipe pins any-llm-sdk==1.17.0 and openai==2.31.0 as well as the OTel plugin. Its install fingerprint covers the Bub requirement, those client overrides, the OTel plugin, and the normalized Python plugin set, so command -v bub alone is not enough to skip installation. A prebuilt environment with an older marker is reinstalled once. Explicit non-default Bub versions keep their caller-owned dependency closure. For all three providers, the build runs once when environment dependencies change, and the artifact takes a new versioned name. Do not reinstall it in every Attempt’s .before().
E2B: derive from official and public templates
E2B already provides aclaude template for Claude Code and a codex template for Codex. NiceEval ships a thin E2B-specific wrapper that lets you keep chaining the native E2B API from those two official starting points. E2B has no official Bub template yet, so the default Bub branch uses the same fully pinned client closure and OTel plugin as bubAgent().
NiceEval also publishes three public templates that any E2B team can reference. NiceEval itself maintains the full namespace and the verified release tag; downstream code just takes the complete reference:
2.1.207, Codex 0.144.1, Bub 0.4.0 (its install fingerprint moves with the Bub requirement, client overrides, OTel plugin, and Python plugin pins). A template’s version follows the agent it ships (0.144.1-r1: agent version plus NiceEval’s recipe revision); the three agents are released independently, and NiceEval’s own library version never appears in the tag. So use the constants directly instead of copying these strings or tracking versions — including when a derived template needs to record which base it came from.
Among this current set of public templates, the Claude Code template follows E2B’s official /usr npm prefix, so an ordinary user cannot write global modules there directly. When you need to add global Node tools such as pnpm or yarn at runtime, explicitly install them in /usr/local, which both templates already put on PATH:
sudo npm install -g. Root global packages already present in the template can conflict on files. NiceEval will unify this default prefix in a future template release. Until release notes explicitly say a new release meets that condition, keep the explicit --prefix /usr/local.
scripts/build-e2b-template.ts
e2bCodingAgentTemplate("claude-code" | "codex" | "bub") returns a native TemplateBuilder, not a private NiceEval build DSL. You can keep using .aptInstall(), .runCmd(), .copy(), and the rest of E2B’s capabilities. Building your own alias freezes the official starting point and your project’s dependencies into a single reproducible artifact; rebuild and pick a new versioned alias when dependencies change.
If the Bub Adapter is configured with pythonPlugins, pass the same set of packages to the factory when you build the template. Only then does the plugin set enter the compatibility fingerprint and actually hit the preinstalled environment:
Docker: use NiceEval-maintained images, or derive from the official node base image
To run NiceEval’s built-inclaude-code, codex, or bub Adapters directly, use the matching public image: niceeval/claude-code, niceeval/codex, or niceeval/bub. Each image contains only its own Agent CLI and publishes a manifest for linux/amd64 and linux/arm64. Its tag matches the corresponding E2B public template — the version position is the version of the agent inside the image. Stable CI should use the named constant or a digest, not the moving latest:
library/* Official Image. A new tag is published when the agent version or the build recipe changes, independently of NiceEval’s own release cadence.
If you only need one agent, or you also need to add project-specific dependencies, write a Dockerfile that derives from Docker’s official baseline, node:24-slim:
Dockerfile
USER, so the container runs commands as root by default, and /usr/local/bin is already on root’s PATH, so global binaries installed with npm install -g are visible out of the box. When you need a non-root identity (for example, Claude Code refuses --dangerously-skip-permissions under root), add USER node to the Dockerfile (the node:24-slim image already has that user), or override it explicitly with dockerSandbox({ source: { type: "image", image }, user: "node" }). If an agent installs somewhere else (for example into ~/.local/bin), remember to put that directory on the PATH. dockerSandbox requires an explicit image; stable CI should reference an immutable tag.
Vercel: snapshot the official runtime
Vercel has no E2B-style template registry and no Dockerfile; a Sandbox snapshot is taken from a microVM that is already running. Use the Vercel SDK to start a Sandbox from the official runtime (node24), install the Agent CLI, call .snapshot() to get a snap_..., then hand it to vercelSandbox({ snapshotId }):
scripts/build-vercel-snapshot.ts
snap_7sIjfs71xfmVly0WEUTGhTBoMGeL, but it is not a cross-account public ID.
Runtime checkpoints
createCheckpoint() / restoreCheckpoint() are a different thing. They pack the Linux paths you name into a Buffer, so you can restore a slice of the file system into an already-created Sandbox:
.before() callback. After restoration succeeds, register the paired createCheckpoint() through context.onCleanup(). The pair runs for the current Attempt and participates in the same global cleanup stack.
sandboxReuse: true only keeps one physical instance alive within the current Invocation. When several Invocations read and write the same checkpoint, declare sharedState: { key } at the Experiment top level. NiceEval acquires that lease before Experiment setup or Sandbox creation and releases it after checkpoint save, the Provider finalizer, and Experiment teardown. It provides mutual exclusion, not checkpoint storage, transaction recovery, or cross-machine coordination.
Transient error retries
When a built-in provider creates a Sandbox, transient failures — rate limiting,fetch failed, connection resets, 5xx, temporary network unreachability — automatically retry with exponential backoff. Configuration errors, such as a missing template or missing credentials, fail on the first try. Once retries are exhausted, the attempt is recorded as errored. defineSandbox’s custom provider create is your own function; NiceEval does not retry it for you.
readText, readBytes, downloadFile, uploadFile, file writes, and directory uploads automatically make a bounded number of retries on transient transport errors: 429, 5xx, fetch failed, and connection resets. Missing files, permission errors, cancellation, and a terminated Sandbox are not retried.
runCommand and runShell do not retry automatically. A command may already have produced side effects, so retry it explicitly in a hook or an eval only when you can confirm it is safe to repeat.
Docker
Docker is good for local development and standard CI. It is simple, controllable, and has no cloud dependency. The tradeoff is limited machine resources, plus slower cold starts and dependency installs.Vercel Sandbox
Vercel Sandbox is good when you want cloud isolation, more resources, or a more stable environment. It requires the right token or OIDC setup.Custom provider
UsedefineSandbox to plug in another service. A custom Provider author is responsible for filesystem, process, network, and credential isolation; NiceEval does not add those safety boundaries to an arbitrary Sandbox implementation. The feedback given to create is already bound to the sandbox.create phase, so it can report on allocating an instance, pulling an image, or restoring a Sandbox snapshot:
Sandbox interface. Do not write the provider SDK’s raw logs straight to the host process’s stdout / stderr. Short-lived status goes through feedback.progress, problems worth keeping go through feedback.diagnostic, and if the environment cannot be created, throw. That keeps the human dashboard from being torn apart by logs and keeps CI output single and ordered.
Permissions and execution identity
Providers differ in what they allow around execution identity, networking, the file system, and process lifecycle. When writing Fixtures, avoid depending on the host machine’s environment; put dependencies inpackage.json or in Fixture setup.
Performance advice
- Bake stable, heavy dependencies into an image/template/snapshot instead of reinstalling them in every Attempt’s
.before(). - Keep Fixture dependencies small.
- Use small, explicit caches or preflight checks for dynamic content.
- Tune
maxConcurrency(the experiment field or--max-concurrency) so local Docker does not run out of resources. - Split slow tests into required gates and optional soft checks.
Reuse Provider startup across Attempts
When Provider startup dominates feedback time, declaresandboxReuse: true in an Experiment. Several Attempts then share one Sandbox serially: creation is paid once per reused instance, every real Attempt still has an attempt .before() / .after() plan, and NiceEval resets the working directory between tasks. Declarative Docker actions can separately reuse a matching preparation prefix.