experiments/. CLI positional arguments only select Evals; they do not temporarily change an Agent or runtime configuration.
The smallest Experiment
agent holds an already configured Agent instance. The URL, authentication, and protocol details of the subject under test usually go to the Adapter factory. The Runner does not store a separate agentConfig field.
Compare System Prompt effects on an Agent
Use flags to configure two Experiments with different prompts and compare them. The two cells differ by one variable only—here,flags.promptVariant—and every other field is byte-for-byte the same. The fewer differences there are, the more clearly a score difference can be attributed:
flags in both cells, or to pin the same model in a model comparison—including giving both the same environment-variable default. The run completes and the report appears, but the two columns compare the same configuration and reveal nothing.
model reaches the Adapter as ctx.model. If your Agent supports model selection, construct the request yourself.
flags reaches the Adapter as ctx.flags and is also available to an Eval as t.flags.
Their semantics match a product A/B-test feature flag. Write the Adapter to send flags to your Agent, then have the Agent switch its System Prompt or behavior from them:
Write a group of Experiments
One Experiment file is one configuration cell. To compare several models, Agents, or flag values, write several files. Directories only determine IDs:Common fields
Pass a runtime address without invalidating cached results
flags record Experiment conditions, and the whole bag participates in cache eligibility. Do not put a tunnel address in flags: it gets a new URL every time the tunnel restarts, which would invalidate every completed result on every restart.
The address is available only after setup starts it. Keep it in the lifecycle closure, then let the Agent factory or a Sandbox callback use it. It is not an Experiment condition, and no general API writes it as a Report field:
flags: { memoryVersion: "0.10.39" }. Changing a version can change behavior, so it should rerun results. The test is whether it is a comparable Experiment condition: if so, put it in flags; if it is only this Invocation’s connection coordinate, keep it in lifecycle code or external-service configuration. There is no general callback, custom data family, or registration point for writing arbitrary values into results. Use labels for annotations used only to group reports that neither an Adapter nor an Eval needs to read.
Start a service shared by an Experiment
Some resources have one instance per Experiment and are shared by all Attempts: a tunnel to an internal memory service, an Experiment-specific mock server, or a license lease. Put them in a pair of Experiment-levelsetup / teardown Hooks, each of which runs at most once. setup runs before the Experiment’s first Attempt that needs dispatch. teardown runs after all Attempts finish, including an interrupted run, if and only if the setup point was reached. A thrown setup still reaches teardown, so cleanup must defend against variables that may not have been assigned. When every prior result is adopted and this Experiment has no Attempt that must actually run, neither setup nor teardown runs:
setup runs, the terminal’s ACTIVE section shows experiment setup · <experiment-id>. The ctx.progress(...) message updates at the end of that row. Its Attempts count as queued while they wait; this is not a hang. In CI or Agent output, the start and finish of setup and teardown each append one line.
When setup throws, every Attempt in this Experiment is recorded as errored with code experiment-setup-failed and appears individually in the report. Other Experiments in the batch continue normally: an environment that cannot start should neither pretend to be green nor take unrelated work down with it.
Releasing resources in teardown is mandatory. Put it behind try/finally so it runs no matter whether earlier observability code fails. Observability work—a health check or metrics reporting—is best effort. Give it a short timeout, do not let failure block release, and skip it when ctx.signal.aborted. On an interruption path, one possibly stuck observation must not stand before tearing down a tunnel or returning a lease:
setup / teardown own one service per Experiment on your machine. To prepare an environment inside a Sandbox per Experiment before the Agent runs—install binaries, warm it, load and save state across Attempts—attach actions to the spec in the sandbox field instead:
maxConcurrency: 1 serializes only this Invocation. sharedState.key protects the checkpoint across Invocations in the same project Coordination domain. The lease starts before Experiment setup or Sandbox creation and ends after registered Sandbox cleanup, Provider finalizers, and Experiment teardown. It provides mutual exclusion only; storage, atomic write-back, forced-kill recovery, and cross-machine coordination remain the author’s responsibility.
Bake fixed Agent CLIs, system packages, and large model caches into an image, template, or snapshot first, or declare them as low-frequency .before() actions. NiceEval can reuse an unchanged Docker preparation prefix instead of rebuilding fixed installation work when an Eval or final .env action changes. For how to derive a prebuilt environment from official Docker images, E2B templates, or a Vercel runtime, see Sandbox Provider: Build on the Official Baselines to Speed Things Up.
For action ordering, caching, dynamic callbacks, and cleanup semantics, see Sandbox Provider: Prepare the Sandbox in a deterministic order.
Cooperate with Sandbox callbacks
An Experiment-level Hook starts a host-side service. A Sandbox callback writes its coordinates before an Attempt and registers state-saving cleanup immediately. The two layers connect through module variables in the same file, and the Runner guarantees their order: the Experiment-levelsetup is earlier than any Sandbox callback for that Experiment, so the callback always reads an assigned variable:
sandbox callback and read the Experiment-level output. How an Agent connects to the subject under test and an Eval’s Fixtures belong in the Agent definition and EvalDef, not the Experiment file.
Share one lifecycle implementation across Experiments
A comparison group often has several Experiments aimed at the same kind of infrastructure: for example, Claude and Codex cells for one memory product, with identical startup and shutdown mechanics. Put that work in a factory function that returns a full kit sharing one closure: the Experiment-level Hook pair, a getter Agents and MCP factories use to read coordinates, and a Sandbox command that writes the coordinates into a Sandbox. Each Experiment file calls the factory separately, so it has the same code but its own instance and coordinates:- The factory creates only a closure at import time. It does no I/O and reads no configuration. Experiment files are imported during
niceeval expdiscovery; an import-time throw would take unrelated Experiments in the same batch down with it. Leave hard failures forsetup. - Runtime coordinates live in the factory closure, not a module-level singleton. Two concurrent Experiments in one batch each have their own copy and cannot overwrite each other. Coordinates exist only after
setup.
teardown runs only when its same-layer setup point was reached, and it still runs when setup throws. refs cannot leak.
The lifecycle boundary is this batch. A service that should survive across runs—start it once, then run niceeval exp repeatedly—still belongs to external orchestration such as docker compose; pass its URL through an environment variable.
Give different Evals different prebuilt environments
A batch of real tasks can need different runtime and dependency versions. Put the specific template-bearing factory directly on each Eval so the task and execution environment remain one declaration:ctx.model and ctx.flags, see Adapter.