Skip to main content

The defineEval shape

tags and sandbox

tags are available to --tag and an Experiment’s evals predicate. Declare sandbox on the Eval when the task itself owns the starting environment. If the Experiment owns the template-bearing Sandbox layer instead, omit this field; each selected Eval × Experiment pair must have exactly one template-bearing layer.

Single-turn evals

t.send() drives one interaction. t.succeeded() and t.calledTool() are scoped Assertions, while t.check() records a value Assertion immediately.

The Turn object

Multi-turn evals

Use t.newSession() when you need parallel independent sessions.

Data-driven testing (dataset fan-out)

One file can export an array of evals:
This generates IDs such as sql/0000 and sql/0001. See Data-driven Testing.

Sandbox Workspace

Coding Agent evals remain ordinary .eval.ts files. Their test prepares a Sandbox Workspace, sends a task, and checks file results:
See Fixtures.

Report information

setup handles this eval’s Fixture. Its second argument is bound to the eval setup phase. Feedback inside test(t) is bound to the eval run phase:
progress updates only short-lived status while execution is running and never enters the result. diagnostic is written into the current Attempt’s result.json, but it neither replaces an Assertion nor changes the Verdict automatically. Business conclusions still come from t.check(value, match), a scoped Assertion, or a Judge measurement. Throw when infrastructure cannot continue.

Eval lifecycle

setup(sandbox, ctx) prepares a Fixture. Pair it with teardown(sandbox, ctx) for cleanup; together they form the Fixture for this eval and each runs once per Attempt. Its position in the full Attempt is fixed: The right column says “once per what”: four Hook layers use the same paired shape but run at different frequencies. Without sandboxReuse, every Attempt owns one Sandbox, so “once per Sandbox” and “per Attempt” coincide. They separate after you declare reuse; see Reuse Sandboxes. Most Fixtures do not need teardown: starting files written into the Sandbox and dependencies installed there disappear automatically when the Sandbox is destroyed. You need teardown for Fixtures outside the Sandbox—temporary resources created for this Attempt in a shared external service, such as a temporary repository, bucket, or queue topic. Otherwise they leak. Multiple Attempts for the same eval—when attempts is greater than 1, or several Experiments in a batch run the same eval—execute concurrently and share a module. A setup handle cannot live in a regular module variable because a later concurrent Attempt would overwrite it. Key it by the sandbox instance with a WeakMap: a Sandbox and Attempt map one-to-one, making it a natural per-Attempt key:
teardown runs if and only if this Attempt reached the setup point. A thrown error in setup does not exempt cleanup, so teardown code must guard against resources that may not have been created. When teardown throws or exceeds its 30-second cleanup limit, NiceEval records only a teardown-failed diagnostic and does not change the Verdict the Attempt already produced. To make cleanup affect the result, throw from setup or test; do not expect teardown to change the Verdict.

Naming conventions

Filename

Only .eval.ts files are discovered by the Runner.

ID namespace

evals/billing/refund.eval.ts has the ID billing/refund.

Dataset

Good for many cases with the same structure and different inputs.

Sandbox Workspace

Good for Coding Agents that need a real file system, commands, and diffs.