Skip to main content
defineEval is the main entry point for authoring an Eval. Each Eval file calls it once, supplies a description and test(t), and default-exports the result.
Do not provide id or name. NiceEval derives the Eval ID from the file path.

defineEval options

description

A one-sentence description shown in niceeval list and view. It is explanatory only and does not affect scheduling or scoring.

tags

Tags for CLI --tag filtering and view classification. They are an independent filtering dimension from ID-prefix filtering.

sandbox

The Sandbox declaration layer contributed by this Eval. Omission is an empty command-only layer and provides no implicit template. Every actual Eval × Experiment match must have exactly one side provide a template-bearing layer.

plugins

Explicit, immutable Eval Plugin occurrences. There is no directory inheritance.

judge

Declares Judge capability. true inherits from Experiment/Config; an object both declares it and overrides them.

reporters

Overrides or adds to project-level Config.reporters for this Eval only.

timeoutMs

Overrides the project-level or CLI timeout for one Attempt, in milliseconds, for this Eval only.

metadata

Arbitrary extra metadata stored as Attempt Provenance. It does not affect scheduling or scoring. Fixed Inspection operations may expose only declared, sealed facts; metadata does not create a custom reporting surface.

diff

Adjusts the exclusion list for Agent-diff attribution, only for Sandbox Agents. See docs/feature/eval/README.md. Both arrays contain gitignore-style globs relative to workdir. The default excludes .git, node_modules, build output, and package-manager caches. ignore adds exclusions to the default list; include has highest priority and explicitly adds matching paths back. The composition rule is fixed: default ∪ ignore, with holes cut by include. The list freezes at the ledger anchor.

test

Test context: t

t (TestContext) is the high-level context an Eval author receives. Every author-facing entry point registers one Assertion when called. t.check(subject, match) checks explicit material; t.check(contextualMatch) reads facts from the current context. t.judge(material, match) takes exactly two arguments and passes a managed ScoreMatch to the same check receiver. Root, Session, and Turn judge methods all require explicit material. The four built-in Judges with explicit material construct a Match and pass it to that receiver. The shared closeQA(materialMatch, question, options?) selects material across applications to answer an acceptance question. Agent additionally offers closeQA(question, options?) for complete scope material and closeQA(domainMatch, question, options?) with an EventMatch or ToolMatch. All forms compose a scoring Match and register it through the same check. usedNoTools() encapsulates its own zero-occurrence criterion and takes no arguments, including no Match. In a pass/fail Eval, Boolean assertions enter the Attempt Verdict by default. In a score-style Eval, Boolean assertions are record-only by default. Both Eval kinds include measurements in failed with .gate(minimum); a score-style Eval can combine a gate and score on the same entry. A Boolean .orStop() uses its own condition. A measurement uses .orStop(minimum), or a parameterless .orStop() after a gate is set. defineScoreEval’s ScoreTestContext additionally offers t.score(n) to directly register a contribution; n must be finite and non-negative. An existing Assertion contributes a score through .score(n), where n must be finite and non-negative; zero is still an explicit contribution. For the complete contract for calledTool and notCalledTool, see Scoped assertions. All members follow:

evaluationKind

send

sendFile

requireInputRequest

respond

respondAll

reply

sessionId

events

newSession

signal

model

reasoningEffort

flags

progress

diagnostic

log

skip

group

check

sandbox

o11y

usage

succeeded

usedNoTools

maxToolCalls

noFailedActions

event

notEvent

maxTokens

maxCost

judge

Judge measurement

Declare the Judge model or Provider in the Eval or its resolved runtime, then pass a Judge definition to t.judge:
Judge material accepts bounded JSON values and explicit judgeImage() values. NiceEval snapshots it when the Assertion is registered. The model configuration resolves from the project’s judgeRuntime, then the Eval’s judge option and Experiment’s judgeRuntime. A full Provider replaces the service and execution configuration; a model string only changes the selected Provider’s model. The Eval’s judge option selects a model or Provider for this question; it is not a Match allowlist. Pass/fail and score-style Evals can call .gate(minimum); score-style Evals can also call .score(n), retaining the contribution when a gate fails.

Image material

Import judgeImage from niceeval/judge. It synchronously copies the original image bytes. PNG and JPEG are supported, with a maximum of 4 MiB and 16,777,216 pixels per image. Each Assertion and model request accepts at most 4 image references and 8 MiB in total; repeated references count toward the limits. Each Attempt retains at most 32 MiB of image data. Input validation checks the signature, header bounds, and dimensions without fully decoding or transcoding the image. OpenAIProvider, VercelProvider, and OpenRouterProvider must explicitly set supportsImages: true and select a model that supports both images and tool calls. Image support is disabled by default; TypesafeProvider does not support images. When image capability is unavailable, NiceEval returns judge-capability-unavailable instead of degrading the image to text.

Built-in Judges and custom Matches

defineJudge and the built-in Judge presets return ScoreMatch values that you can pass to t.check. t.judge is a convenience entry point on the same check path, so registration, material snapshotting, budget, evaluation, audit, sealing, and the handle all happen once. Agent also accepts closeQA(question, options?) using the current Turn, Session, or Attempt’s complete recorded history. Pass an EventMatch to select event views or a ToolMatch to select tool facts, including inputs and receipts. Custom MaterialMatch readers provide business material across applications. Each form uses the same scoring Match and check path. For example, use instructionFollowing() directly:
faithfulness extracts facts, evaluates them individually, and calculates the share in code. An incomplete extraction has no valid score. TypesafeProvider does not support extract, so it returns unavailable instead of estimating the whole output or changing the denominator. pairwisePreference compares one candidate with one reference in a fixed order. closeQA sends all matched items to one Judge in their original order to answer a question based only on the material. A complete empty collection produces 0 with no model call. Partial, unknown, or oversized evidence is unavailable; NiceEval does not truncate it before scoring. For a workflow with multiple model steps, use score(value, context) with the advanced defineScoreMatch API. context.llm provides Effect-returning score, classify, extract, and batchClassify methods. They share call budgets, timeouts, and audit records with built-in Judges. Static configuration checks make no network requests; a missing Provider, model, or key also makes no request. Chat services use forced-function requests, while TypeSafe uses /systemone; both validate responses strictly. Transport failures and timeouts are unavailable; HTTP 400 and protocol incompatibilities are errored. See Custom Matches for the full contract.

Turn return type

t.send(...) returns a TurnHandle: convenience fields derived from the event stream plus a complete set of Assertions scoped to this turn. calledTool and notCalledTool are also Turn-scoped Assertions. Their signatures and Match rules are defined only in Scoped assertions.

events

toolCalls

status

message

data

usage

succeeded

toolOrder

usedNoTools

maxToolCalls

noFailedActions

event

notEvent

eventOrder

maxTokens

maxCost

judge

Test-suite exports

An array export creates stable IDs: file/0000, file/0001, and so on. When you already have stable business keys, you can instead default-export Record<string, EvalDef>. For example, key 15193 in swelancer.eval.ts produces swelancer/15193. A key must be a nonempty single path segment. It cannot be . or .., and cannot contain /, \\, or control characters. Discovery order is fixed by lexical key order.

Context material and official usage measurements

defineMaterialMatch<C, T>({ name, read, match, capture? }) is exported from niceeval and niceeval/expect. An Adapter’s create provides the application context. NiceEval injects the current read-only context into the reader when check or closeQA is registered. The reader returns { state: "complete", items: [{ id, value }] }, or a partial/unavailable result with a reason. IDs must be unique within a collection and no longer than 128 UTF-8 bytes. The default capture limits are 256 items and 48 KiB; you can explicitly raise them to 16,384 items and 4 MiB. check(materialMatch) returns a Boolean handle and defaults to checking that at least one item matches. closeQA(materialMatch, question, options?) returns a measurement handle. Agent readers read managed toolCalls or eventOccurrences from AgentMatchContext<Scope> and use the corresponding ToolMatch or EventMatch. Agent’s additional closeQA(question, options?) reads the current scope’s complete recorded history. closeQA(domainMatch, question, options?) selects all matches with an EventMatch or ToolMatch; use a ToolMatch for tool input and output material. The question must be nonempty and no longer than 8 KiB. options uses JudgePresetOptions. All matched items go to one Judge in their original order; a complete empty collection scores 0 without a model call. Partial, unknown, or oversized evidence is unavailable. defineContextMatch<C, T>({ name, read, match, capture? }) reads one small fact. The reader returns { state: "available", value } or { state: "unavailable", reason }. Pass it to t.check(contextMatch): a BooleanMatch produces a Boolean handle; a ScoreMatch produces a measurement handle. Unknown facts do not invoke the scorer. t.usage reads a frozen value from the official ledger, retaining failed calls, retries, cache buckets, and cost provenance. An Adapter’s basis is recorded-calls; an Agent’s is reported-sends, with Turn, Session, and Attempt scopes preserved. Fields such as usage.totalTokens are NumericMaterial. Unknown values are not filled with zero. t.maxTokens(max) limits the complete token total. t.maxCost(usd) accepts a finite non-negative number or canonical non-negative decimal string and precisely limits effective USD cost. Reported cost, including zero, takes precedence over explicit pricing estimates. A known lower bound above the limit fails; incomplete cost evidence cannot pass. t.elapsedMs is wall-clock time from the runtime Attempt start to the point of reading. Pass it to t.check(..., atMost(limit)) as needed. The application’s own first-completion time is evaluated by an application Match.