sandboxReuse: true in the Experiment to have Attempts share a Sandbox serially. Sandbox creation runs once per reused instance; every real Attempt still gets its own .before() / .after() plan, and NiceEval resets the working directory between tasks. Declarative Docker actions can reuse a matching preparation prefix by fingerprint.
This page covers Experiment-level
sandboxReuse. When only several compatible Eval groups need their own reuse boundaries, use Eval Groups: each group reuses serially while different groups remain concurrent. An Experiment that selects Eval Groups cannot also declare sandboxReuse: true.- Results from a reused run still enter the cache. When a pair has a stable carry identity and a terminal-result fingerprint matches, NiceEval adopts it directly and creates no Sandbox. A Plugin’s attachment owner, name, instance key, behavior revision, declared identity, host lifecycle shape, and Sandbox action declarations all belong to carry identity; changing any of them makes the slot fresh. Callback function bodies remain opaque. When behavior changes, update declared identity too or use
--rerun all. Only unadopted Attempts execute in the shared Sandbox for this Invocation. - Interruption does not roll back external state. Resuming is the same Experiment trajectory only when cross-Attempt state can return to the boundary of the last terminal commit. Otherwise, rebuild a clean cohort from the beginning.
- State outside the workdir remains.
$HOME,/tmp, global installations, and background processes survive the between-task reset. Eval setup code must account for this, as described in Place preparation by what changes. - It is incompatible with
--keep-sandbox.
--rerun.
Enable it: three Experiment settings
timeoutMs and lifetimeMs are two clocks that measure different objects. The former limits how long one Attempt can run; the latter limits how long one Sandbox can live. To keep a Sandbox longer, increase lifetimeMs, not timeoutMs: increasing the latter also weakens protection against a stuck Agent. A Provider account tier determines the upper limit for lifetimeMs—for example, the E2B free tier allows one hour—and the Provider reports its own error during creation when it is exceeded.
Sandboxes are not shared between terminals
Sandbox reuse occurs only within one Invocation. When two terminals run one Experiment at the same time, each has its own Run and Sandbox pool. NiceEval does not give a running Sandbox handle to another process. Both Invocations can publish results concurrently to the same project’s.niceeval/record.sqlite.
Normal concurrent runs do not need separate Records to avoid writer conflicts. Copy or archive .niceeval/record.sqlite only after an Invocation completes its portable gate successfully; a SQLite file during a run is not a portable snapshot.
When a Sandbox retains only its own temporary state, both terminals can run different Evals at the same time. When a dynamic .before() callback restores and saves the same external checkpoint, declare a stable, non-secret key for the Experiment:
teardown. A waiter creates no Sandbox. After the lease is released, it continues its own plan; it does not read or adopt results from the other Run.
The key enters the result’s configuration identity. Changing it starts a different state cohort, so old results do not mix in. This lease protects only the external checkpoint; it is independent of the project Record, which normal Invocations can publish to concurrently. Different machines or working copies still need an external distributed lock.
Lifecycles: how often each runs
After every Attempt, NiceEval runs
git reset --hard and git clean on workdir, returning it to the reset point before it starts the next Attempt. That reset covers only workdir. /opt, $HOME, /tmp, global installations, package caches, and background processes do not disappear. Either an Eval’s teardown must clean them up, or the author must deliberately retain them for the next task.
Large persistent builds or caches cannot rely on unbounded growth. The author should declare a capacity limit and a threshold before reaching it. Record normal size and hit behavior with facts, issue a diagnostic warning only at a risky threshold, and provide a cleanup, rotation, or Sandbox-retirement strategy.
Place preparation by what changes
- Heavy dependencies every Experiment needs—an Agent CLI or language runtime—go into a Provider image, template, or snapshot, or into the lowest-frequency
.before()actions. - Preparation the whole batch shares—installing a toolchain, cloning a common repository, warming a build cache—belongs in low-frequency
.before()actions. Docker can reuse the matching preparation prefix across Invocations. The current occurrence remainsattempt; author code cannot request physical-instance promotion. - Material only this task needs—its own repository, data, or dependencies—belongs in that Eval’s
setuportest(t). It is replayed after every task reset, so replaying it must stay correct. When every task clones, use one named temporary clone directory for the whole Experiment, such as.niceeval-task-clone, and add it todiff.ignore. Before cloning, runrm -rf .niceeval-task-cloneto clear only the previous task’s leftovers. Never removeworkdir’s root.git: NiceEval needs it for the between-taskgit reset, and it survives that reset. - A background process or occupied port must be cleaned up by the Eval’s own
teardown; the between-task reset does not kill processes.
Idempotence is mandatory: one bad pattern and one good one
Both Sandbox callbacks and Eval-level setup can encounter partial state from a previous run. These two approaches differ:Native Agent Plugins: installation converges in the Adapter
You can declarecodexAgent({ plugins: [...] }), claudeCodeAgent({ plugins: [...] }), and sandboxReuse together. Native Agent Plugins install in $HOME, which survives the between-task reset, but you do not need to handle that: before every Attempt, the Adapter converges Plugin installation on the configuration you declared. A same-named marketplace registration and Plugin left by the prior Attempt are replaced by a fresh installation using the declared source and ref. The same rule absorbs a Plugin script that rewrites marketplace registration, such as changing it to a hosted source.
This describes native Plugins in an Agent factory. Top-level plugins fields on Evals, Experiments, and Eval Groups carry NiceEval conditions; see Compose Lifecycles with Plugins for their division of responsibilities.
You still own two things:
- A
postSetupscript must be idempotent. It reruns on a$HOMEwith leftovers for every Attempt. A script that registers a hook in global configuration must converge on the same configuration when it runs again. Before enabling reuse, run two or three tasks down one Sandbox lane; the second task onward is the real test. - Where Plugin data lives. The Plugin installation directory is reinstalled and overwritten for every Attempt. Runtime data that must last until the next task can only live outside that directory. Whether retained data contaminates the next task is a question your Experiment design must answer.
.before() actions can separately reuse a matching Docker preparation prefix. If a marketplace fetch is too slow, use sparse to fetch only the Plugin path, or bake the Plugin into the template and remove its declaration from plugins. The cost is that the installation manifest no longer records the Plugin and resolved version.
A batch with different Sandbox layers
When selected Evals declare different template-bearingsandbox layers, you do not need to split the command. Link planning resolves each Eval × Experiment pair independently. Attempts with different Provider plans share neither Sandboxes nor preparation prefixes, while the Experiment’s maxConcurrency limits all selected work together.
A typical rhythm
--keep-sandbox and sandboxReuse are incompatible, and a reused Sandbox belongs to the whole batch, so keeping it would not be faithful to one task. Do the same when “the chain fails but an isolated run passes.” Do not debug that problem inside a shared batch. See Keep a Sandbox for Debugging for the workflow.
Cannot do it? Use concurrency, not reuse
When setup cannot become idempotent, or an Eval depends on a fresh$HOME—a memory-oriented subject under test or a stateful service—do not declare sandboxReuse. Use concurrency for speed instead:
Related reading
- Eval Groups — explicitly declare members and ordering within a group while retaining concurrency between groups.
- Rerun and Carry Results — see how reused Sandboxes adopt terminal results and when to use
--rerun. - Concurrency and Execution Order — get the same speedup through concurrency when you do not reuse.
- Sandbox Provider Configuration — create Provider templates, images, and snapshots.