Skip to main content
When a batch of Evals can safely share one running environment, creating a new Sandbox for every Attempt repeats Provider startup. Declare sandboxReuse: true in the Experiment to have Attempts share a Sandbox serially. Sandbox creation runs once per reused instance; every real Attempt still gets its own .before() / .after() plan, and NiceEval resets the working directory between tasks. Declarative Docker actions can reuse a matching preparation prefix by fingerprint.
This page covers Experiment-level sandboxReuse. When only several compatible Eval groups need their own reuse boundaries, use Eval Groups: each group reuses serially while different groups remain concurrent. An Experiment that selects Eval Groups cannot also declare sandboxReuse: true.
The between-task reset does not reset the whole Sandbox. NiceEval restores only the workdir according to its ledger. State outside the workdir—such as /opt, $HOME, and /tmp—plus global installations, package caches, and background processes remains. For large persistent builds or caches, the author must choose capacity limits, threshold alerts, cleanup or rotation policies. You also need a strategy to retire an unsafe Sandbox.
Understand the trade-offs before enabling it:
  • Results from a reused run still enter the cache. When a pair has a stable carry identity and a terminal-result fingerprint matches, NiceEval adopts it directly and creates no Sandbox. A Plugin’s attachment owner, name, instance key, behavior revision, declared identity, host lifecycle shape, and Sandbox action declarations all belong to carry identity; changing any of them makes the slot fresh. Callback function bodies remain opaque. When behavior changes, update declared identity too or use --rerun all. Only unadopted Attempts execute in the shared Sandbox for this Invocation.
  • Interruption does not roll back external state. Resuming is the same Experiment trajectory only when cross-Attempt state can return to the boundary of the last terminal commit. Otherwise, rebuild a clean cohort from the beginning.
  • State outside the workdir remains. $HOME, /tmp, global installations, and background processes survive the between-task reset. Eval setup code must account for this, as described in Place preparation by what changes.
  • It is incompatible with --keep-sandbox.
Good use cases include locally smoke-testing a batch of Evals, running the same task N times to assess stability, and validating an integration. You get speed while retaining result adoption and the choice of --rerun.

Enable it: three Experiment settings

timeoutMs and lifetimeMs are two clocks that measure different objects. The former limits how long one Attempt can run; the latter limits how long one Sandbox can live. To keep a Sandbox longer, increase lifetimeMs, not timeoutMs: increasing the latter also weakens protection against a stuck Agent. A Provider account tier determines the upper limit for lifetimeMs—for example, the E2B free tier allows one hour—and the Provider reports its own error during creation when it is exceeded.

Sandboxes are not shared between terminals

Sandbox reuse occurs only within one Invocation. When two terminals run one Experiment at the same time, each has its own Run and Sandbox pool. NiceEval does not give a running Sandbox handle to another process. Both Invocations can publish results concurrently to the same project’s .niceeval/record.sqlite. Normal concurrent runs do not need separate Records to avoid writer conflicts. Copy or archive .niceeval/record.sqlite only after an Invocation completes its portable gate successfully; a SQLite file during a run is not a portable snapshot. When a Sandbox retains only its own temporary state, both terminals can run different Evals at the same time. When a dynamic .before() callback restores and saves the same external checkpoint, declare a stable, non-secret key for the Experiment:
NiceEval acquires this key before Experiment setup or Sandbox creation, and releases it only after registered Sandbox cleanup, the Provider finalizer, and Experiment teardown. A waiter creates no Sandbox. After the lease is released, it continues its own plan; it does not read or adopt results from the other Run. The key enters the result’s configuration identity. Changing it starts a different state cohort, so old results do not mix in. This lease protects only the external checkpoint; it is independent of the project Record, which normal Invocations can publish to concurrently. Different machines or working copies still need an external distributed lock.

Lifecycles: how often each runs

After every Attempt, NiceEval runs git reset --hard and git clean on workdir, returning it to the reset point before it starts the next Attempt. That reset covers only workdir. /opt, $HOME, /tmp, global installations, package caches, and background processes do not disappear. Either an Eval’s teardown must clean them up, or the author must deliberately retain them for the next task. Large persistent builds or caches cannot rely on unbounded growth. The author should declare a capacity limit and a threshold before reaching it. Record normal size and hit behavior with facts, issue a diagnostic warning only at a risky threshold, and provide a cleanup, rotation, or Sandbox-retirement strategy.

Place preparation by what changes

  • Heavy dependencies every Experiment needs—an Agent CLI or language runtime—go into a Provider image, template, or snapshot, or into the lowest-frequency .before() actions.
  • Preparation the whole batch shares—installing a toolchain, cloning a common repository, warming a build cache—belongs in low-frequency .before() actions. Docker can reuse the matching preparation prefix across Invocations. The current occurrence remains attempt; author code cannot request physical-instance promotion.
  • Material only this task needs—its own repository, data, or dependencies—belongs in that Eval’s setup or test(t). It is replayed after every task reset, so replaying it must stay correct. When every task clones, use one named temporary clone directory for the whole Experiment, such as .niceeval-task-clone, and add it to diff.ignore. Before cloning, run rm -rf .niceeval-task-clone to clear only the previous task’s leftovers. Never remove workdir’s root .git: NiceEval needs it for the between-task git reset, and it survives that reset.
  • A background process or occupied port must be cleaned up by the Eval’s own teardown; the between-task reset does not kill processes.

Idempotence is mandatory: one bad pattern and one good one

Both Sandbox callbacks and Eval-level setup can encounter partial state from a previous run. These two approaches differ:
The test is simple: setup code describes the target state; it does not describe whether to do something. Detection-and-skip guards treat partial state as complete state. They fail precisely under reuse, where a one-off local run never exposes them.

Native Agent Plugins: installation converges in the Adapter

You can declare codexAgent({ plugins: [...] }), claudeCodeAgent({ plugins: [...] }), and sandboxReuse together. Native Agent Plugins install in $HOME, which survives the between-task reset, but you do not need to handle that: before every Attempt, the Adapter converges Plugin installation on the configuration you declared. A same-named marketplace registration and Plugin left by the prior Attempt are replaced by a fresh installation using the declared source and ref. The same rule absorbs a Plugin script that rewrites marketplace registration, such as changing it to a hosted source. This describes native Plugins in an Agent factory. Top-level plugins fields on Evals, Experiments, and Eval Groups carry NiceEval conditions; see Compose Lifecycles with Plugins for their division of responsibilities. You still own two things:
  • A postSetup script must be idempotent. It reruns on a $HOME with leftovers for every Attempt. A script that registers a hook in global configuration must converge on the same configuration when it runs again. Before enabling reuse, run two or three tasks down one Sandbox lane; the second task onward is the real test.
  • Where Plugin data lives. The Plugin installation directory is reinstalled and overwritten for every Attempt. Runtime data that must last until the next task can only live outside that directory. Whether retained data contaminates the next task is a question your Experiment design must answer.
Convergence is not free: Plugin installation is still paid for on every Attempt. Reuse saves Sandbox creation; declarative .before() actions can separately reuse a matching Docker preparation prefix. If a marketplace fetch is too slow, use sparse to fetch only the Plugin path, or bake the Plugin into the template and remove its declaration from plugins. The cost is that the installation manifest no longer records the Plugin and resolved version.

A batch with different Sandbox layers

When selected Evals declare different template-bearing sandbox layers, you do not need to split the command. Link planning resolves each Eval × Experiment pair independently. Attempts with different Provider plans share neither Sandboxes nor preparation prefixes, while the Experiment’s maxConcurrency limits all selected work together.

A typical rhythm

Step 3 must use another Experiment: --keep-sandbox and sandboxReuse are incompatible, and a reused Sandbox belongs to the whole batch, so keeping it would not be faithful to one task. Do the same when “the chain fails but an isolated run passes.” Do not debug that problem inside a shared batch. See Keep a Sandbox for Debugging for the workflow.

Cannot do it? Use concurrency, not reuse

When setup cannot become idempotent, or an Eval depends on a fresh $HOME—a memory-oriented subject under test or a stateful service—do not declare sandboxReuse. Use concurrency for speed instead:
Reuse saves only the portion of time spent on one Sandbox creation plus common preparation. When that time is much shorter than one Attempt’s execution time, concurrency already gets almost all wall-clock benefit without assuming reuse’s state semantics. Start with concurrency, measure whether common preparation truly dominates, then consider reuse.