Basic shape
One Experiment file = one configuration (one Agent × one model). The path produces an Experiment ID and supports prefix selection:directory path/file name). You do not need to batch-run the directory first:
gpt-5.4-mini.ts, use a filename prefix to select it:
defineExperiment fields and how flags flow through.
Comparable run dimensions
- Different Adapters.
- Different models. Tier 1 integration is enough: the application interface exposes model choice and forwards
modelthroughctx.model. - Different prompts or feature flags. This requires Tier 3 integration because the variant is internal to the application; it must expose it as an Experiment-selectable flag and forward it through
flags→ctx.flags. - Different Sandbox Providers.
- Different runtime environment conditions, such as whether a memory-tool binary is installed or state is preseeded. Write environment differences in the
.before()/.after()actions of thesandboxspec, one Experiment file per variant. See Sandbox Providers · Prepare the Sandbox in a deterministic order. - Pass@N for the same task.
View results
Experiment output usually appears by(agent, model, eval):
evals selection. A function form iterates over every discovered eval:
eval.id is a project-local ID derived from the file path, not an absolute path, so you can test it with startsWith / includes. In Authoring Evals, the coding/fix-button example satisfies both predicates. The parsed result becomes this run’s plan. Fixed Inspection operations read actual physical Attempts and coverage facts.
Advice for Experiment design
- Keep the eval set stable so a comparison does not mix in new variables.
- Run several Attempts per cell, especially for nondeterministic coding Agents.
- Make budget and concurrency explicit.
- Classify failure causes instead of looking only at the total score.
Relation to ordinary runs
npx niceeval exp <path-or-id> runs by file identity. Each file’s evals determines which evals it dispatches; the Run records actual physical Attempts and the known-task denominator.