defineEvalGroup() puts compatible evals
inside one physical reuse boundary. Attempts that really dispatch within a group run serially in stable order and reuse one Sandbox;
other groups and ungrouped evals can still use other concurrency slots.
Eval groups schedule only Attempts that need to execute for real in this run. Carried results never enter the group’s Sandbox;
--rerun, attempts, and early exit keep their existing Experiment behavior.
Choose the right reuse mode first
One Experiment cannot both select eval groups and declare
sandboxReuse: true. Eval groups already own their reuse boundary;
when both appear, NiceEval reports eval-group-sandbox-reuse-conflict before the Provider creates a Sandbox.
Step 1: Put the group file next to its members
Placeeval-group.ts in a named directory under evals/. The directory path becomes the eval-group ID:
evals/**/eval-group.ts. Do not name it *.eval-group.ts and do not put it at
evals/eval-group.ts; the latter has no usable group ID.
Step 2: Use a factory result to declare members
Each member continues to default-export the result ofdefineEval() or defineScoreEval(). The group file imports those results,
then lists compatible members in one closed set:
evals must not be empty. It accepts only objects actually returned by the factories—not eval IDs, directory prefixes,
glob, tag, or selector. NiceEval does not collect files from the directory automatically either. Adding, removing, or reordering members is explicit in the
eval-group.ts diff.
Every member must have the same evaluation kind: all defineEval() results or all defineScoreEval() results. Discovery rejects a mixed group before Experiment filtering or Sandbox planning, and lists both sets of Eval IDs so you can split the group.
The evals array declares members, not business order. The Runner always sorts normalized eval IDs stably; changing
only the array positions does not change scheduling behavior or the group fingerprint. Put a result dependency such as “build first, then query” into
one eval. Eval groups do not implement a business-order API.
onUnavailable is required. "stop-group" stops later dispatch in the group when the physical Sandbox cannot be created, reset, or prepared;
"replace-sandbox" retires the current instance first, then lets the next slot try to establish a replacement instance.
Omitting the policy fails while loading eval-group.ts, so the Runner does not have to guess the cost and side effects after failure for its author.
An eval can belong to at most one group, and it cannot appear more than once in the same group. Experiments and the CLI still select evals:
selecting workflow/02-query does not pull workflow/01-index into this run.
Step 3: Let the Experiment provide a reusable Sandbox
Most projects let an Experiment choose the Provider andtemplate; the eval group owns only the reuse queue:
sandboxReuse: true here. maxConcurrency: 4 caps the whole Experiment; it does not run four Attempts from one group at once.
It lets other eval groups or ungrouped evals use the remaining concurrency slots.
Eval groups support only Sandbox Agents and Providers that support reuse. Direct Agents cannot run eval groups.
Step 4: Check the group ID and selection result first
Confirm selection with--dry; it creates no Sandbox:
attempts is greater than 1, consecutive Attempts for one member
enter the group lane before the next member. This is a stable scheduling rule, not a cross-eval data-dependency contract.
Slots that early exit or result carry do not dispatch are skipped directly.
Put common preparation on the eval group
One Sandbox plan can receive declarations from the Experiment, eval group, and eval, but only one of the three layers can provide a template. The usual split is:
An eval group can provide its own
SandboxLayer:
ctx.evalGroup.id and
ctx.evalGroup.definitionHash. Shared functions can also serve ungrouped evals, so the type keeps evalGroup
optional. When you need to isolate a cache or service namespace outside workdir, first check that the field exists, then derive a key from the group ID.
Members cannot provide a template, but they can contribute command-only .before() / .after() actions. Experiment, group, member, and Agent actions enter the same dependency graph. Once dependencies are satisfied, the smallest changeFrequency runs first, so a stable group action can precede a more volatile Experiment action.
Do not treat eval groups as task dependency graphs
Between Attempts, NiceEval resetsworkdir to the state after common preparation. A file written to workdir by one task is not an input to the next;
$HOME, /tmp, global installations, and background processes can remain. Do not reuse a Sandbox if you cannot accept those leftovers.
An eval group also does not guarantee that the preceding task executes. Result carry, CLI filtering, budget exhaustion, and interruption can leave members out of this run’s Sandbox.
When a later step must read a file from an earlier step, put both steps in one eval. Put evals in the same reusable group only when their shared state belongs outside workdir and every eval can independently produce a valid Verdict.
Common errors
Continue reading
- Reuse Complete Evaluation Conditions with Plugins—combine declarations needed by a group, Experiment, or member into an explicit occurrence while preserving the eval group’s Docker Sandbox reuse boundary.
- Reuse Sandboxes—the lifecycle and leftover-state boundary for ordinary batches that suit
sandboxReuse: true. - Tune Concurrency—how eval groups work with global and Experiment concurrency caps.
- Configure Sandbox Providers—choose a
template, set lifetime, and write aSandboxLayer. - Rerun and Carry Results—which Attempts enter this run’s group queue.