Skip to main content
An Experiment is check-in-able runtime configuration. It records which Adapter and model a group of Evals targets, which flags are on, how many times to run, and the budget, all under experiments/. CLI positional arguments only select Evals; they do not temporarily change an Agent or runtime configuration.

The smallest Experiment

agent holds an already configured Agent instance. The URL, authentication, and protocol details of the subject under test usually go to the Adapter factory. The Runner does not store a separate agentConfig field.

Compare System Prompt effects on an Agent

Use flags to configure two Experiments with different prompts and compare them. The two cells differ by one variable only—here, flags.promptVariant—and every other field is byte-for-byte the same. The fewer differences there are, the more clearly a score difference can be attributed:
The most common mistake is to write the same flags in both cells, or to pin the same model in a model comparison—including giving both the same environment-variable default. The run completes and the report appears, but the two columns compare the same configuration and reveal nothing. model reaches the Adapter as ctx.model. If your Agent supports model selection, construct the request yourself. flags reaches the Adapter as ctx.flags and is also available to an Eval as t.flags. Their semantics match a product A/B-test feature flag. Write the Adapter to send flags to your Agent, then have the Agent switch its System Prompt or behavior from them:

Write a group of Experiments

One Experiment file is one configuration cell. To compare several models, Agents, or flag values, write several files. Directories only determine IDs:
The report compares these selected Experiments directly; it needs no additional grouping field. To validate one configuration only, use its full ID as a positional argument:
This is useful for per-configuration debugging. You can confirm whether one cell’s change reaches the target without first running the whole group or moving other configuration files out of the directory.

Common fields

Pass a runtime address without invalidating cached results

flags record Experiment conditions, and the whole bag participates in cache eligibility. Do not put a tunnel address in flags: it gets a new URL every time the tunnel restarts, which would invalidate every completed result on every restart. The address is available only after setup starts it. Keep it in the lifecycle closure, then let the Agent factory or a Sandbox callback use it. It is not an Experiment condition, and no general API writes it as a Report field:
When the URL changes and you rerun, completed work is still adopted and only missing work runs:
The address serves only this Invocation. Adopted results are not rewritten with its new value. Long-lived conditions you need to compare must be declared as stable Experiment configuration. A service version is different: write flags: { memoryVersion: "0.10.39" }. Changing a version can change behavior, so it should rerun results. The test is whether it is a comparable Experiment condition: if so, put it in flags; if it is only this Invocation’s connection coordinate, keep it in lifecycle code or external-service configuration. There is no general callback, custom data family, or registration point for writing arbitrary values into results. Use labels for annotations used only to group reports that neither an Adapter nor an Eval needs to read.

Start a service shared by an Experiment

Some resources have one instance per Experiment and are shared by all Attempts: a tunnel to an internal memory service, an Experiment-specific mock server, or a license lease. Put them in a pair of Experiment-level setup / teardown Hooks, each of which runs at most once. setup runs before the Experiment’s first Attempt that needs dispatch. teardown runs after all Attempts finish, including an interrupted run, if and only if the setup point was reached. A thrown setup still reaches teardown, so cleanup must defend against variables that may not have been assigned. When every prior result is adopted and this Experiment has no Attempt that must actually run, neither setup nor teardown runs:
While setup runs, the terminal’s ACTIVE section shows experiment setup · <experiment-id>. The ctx.progress(...) message updates at the end of that row. Its Attempts count as queued while they wait; this is not a hang. In CI or Agent output, the start and finish of setup and teardown each append one line. When setup throws, every Attempt in this Experiment is recorded as errored with code experiment-setup-failed and appears individually in the report. Other Experiments in the batch continue normally: an environment that cannot start should neither pretend to be green nor take unrelated work down with it. Releasing resources in teardown is mandatory. Put it behind try/finally so it runs no matter whether earlier observability code fails. Observability work—a health check or metrics reporting—is best effort. Give it a short timeout, do not let failure block release, and skip it when ctx.signal.aborted. On an interruption path, one possibly stuck observation must not stand before tearing down a tunnel or returning a lease:
setup / teardown own one service per Experiment on your machine. To prepare an environment inside a Sandbox per Experiment before the Agent runs—install binaries, warm it, load and save state across Attempts—attach actions to the spec in the sandbox field instead:
maxConcurrency: 1 serializes only this Invocation. sharedState.key protects the checkpoint across Invocations in the same project Coordination domain. The lease starts before Experiment setup or Sandbox creation and ends after registered Sandbox cleanup, Provider finalizers, and Experiment teardown. It provides mutual exclusion only; storage, atomic write-back, forced-kill recovery, and cross-machine coordination remain the author’s responsibility. Bake fixed Agent CLIs, system packages, and large model caches into an image, template, or snapshot first, or declare them as low-frequency .before() actions. NiceEval can reuse an unchanged Docker preparation prefix instead of rebuilding fixed installation work when an Eval or final .env action changes. For how to derive a prebuilt environment from official Docker images, E2B templates, or a Vercel runtime, see Sandbox Provider: Build on the Official Baselines to Speed Things Up. For action ordering, caching, dynamic callbacks, and cleanup semantics, see Sandbox Provider: Prepare the Sandbox in a deterministic order.

Cooperate with Sandbox callbacks

An Experiment-level Hook starts a host-side service. A Sandbox callback writes its coordinates before an Attempt and registers state-saving cleanup immediately. The two layers connect through module variables in the same file, and the Runner guarantees their order: the Experiment-level setup is earlier than any Sandbox callback for that Experiment, so the callback always reads an assigned variable:
Read an Experiment file from top to bottom as a complete run description. Host resources that run once for the whole Experiment are in the Experiment-level Hook pair. Per-Attempt Sandbox writes and registered cleanup live in the chained sandbox callback and read the Experiment-level output. How an Agent connects to the subject under test and an Eval’s Fixtures belong in the Agent definition and EvalDef, not the Experiment file.

Share one lifecycle implementation across Experiments

A comparison group often has several Experiments aimed at the same kind of infrastructure: for example, Claude and Codex cells for one memory product, with identical startup and shutdown mechanics. Put that work in a factory function that returns a full kit sharing one closure: the Experiment-level Hook pair, a getter Agents and MCP factories use to read coordinates, and a Sandbox command that writes the coordinates into a Sandbox. Each Experiment file calls the factory separately, so it has the same code but its own instance and coordinates:
In an Experiment file, changing Agents changes only the Agent lines; four lifecycle lines complete the connection:
Two disciplines keep several Experiments running the same code concurrently from stepping on each other:
  • The factory creates only a closure at import time. It does no I/O and reads no configuration. Experiment files are imported during niceeval exp discovery; an import-time throw would take unrelated Experiments in the same batch down with it. Leave hard failures for setup.
  • Runtime coordinates live in the factory closure, not a module-level singleton. Two concurrent Experiments in one batch each have their own copy and cannot overwrite each other. Coordinates exist only after setup.
When starting several service instances is too expensive and the Experiments in one batch must share one instance, replace per-Experiment instances with reference counting: start on the first entrant and stop after the last leaves.
The paired trigger rule keeps the count balanced: teardown runs only when its same-layer setup point was reached, and it still runs when setup throws. refs cannot leak. The lifecycle boundary is this batch. A service that should survive across runs—start it once, then run niceeval exp repeatedly—still belongs to external orchestration such as docker compose; pass its URL through an environment variable.

Give different Evals different prebuilt environments

A batch of real tasks can need different runtime and dependency versions. Put the specific template-bearing factory directly on each Eval so the task and execution environment remain one declaration:
There is no profile registry or Sandbox creator table registered by source kind. Each factory owns both its support declaration and implementation. Physical planning validates every selected Eval before it creates any Sandbox. Share repeated templates with ordinary TypeScript helpers. One Experiment can still cover every Eval; link planning pairs each Eval layer with the Experiment layer in turn. For comparison design across configurations, see Experiment Matrices. For how an Adapter uses ctx.model and ctx.flags, see Adapter.