Skip to main content
NiceEval provides two eval kinds. defineEval answers whether one run meets requirements, and the default Report reads pass rate. defineScoreEval answers how much work was completed, and the default Report reads accumulated score. Both kinds register an Assertion when you call it; the handle only configures that same entry. Keep each Experiment and Eval Group homogeneous: all pass/fail Evals or all scored Evals. niceeval check, dry runs, and normal runs reject a mixed selection before an Agent or Sandbox starts. Split the selection into two Experiment or Group definitions when a project needs both readings.

Use defineEval for requirements that must hold

Boolean matched enters the Verdict. mismatched makes the final Verdict failed, but does not prevent other Assertions from continuing to register and settle. When later code depends on this result, use await handle.orStop().

Use a continuous quality threshold

A Match such as similarity(...) produces a measurement in [0, 1]. It calculates character-level Levenshtein edit similarity, not semantic similarity. In a pass/fail Eval, call .gate(minimum) on the registered handle to set the quality threshold:
Below the measurement threshold, that gated requirement fails. Without .gate(minimum), the measurement does not affect the Verdict. When you only want an explanation that does not affect the Verdict, use t.diagnostic(...); do not register a measurement that nobody consumes.

Award points for each completed step

Step-by-step tasks fit defineScoreEval. Assertions save only evaluation by default; .score(n) is what makes one contribute points. A Boolean match contributes n, a mismatch contributes 0, and a measurement m contributes m * n.
t.score(n) directly registers a contribution. n must be finite and nonnegative, and its returned handle can configure only key and label. When test returns normally, NiceEval closes the score automatically. No scoring entries is still a valid result with an official score: 0: evaluation successfully formed a zero score; it is not an execution failure or insufficient evidence.

Results for the two eval kinds

scored can be a zero score. Zero means evaluation successfully formed a score; it is not execution failure or insufficient evidence. Every Attempt still records a terminal Verdict, but the Report uses scored or errored as the primary result for a Score Eval; it does not convert that Verdict into a pass rate.

Judge is also a measurement

A defineJudge call declares a managed Judge Match. Pass the values needed for the decision as fields in the material object. Register the definition directly with t.judge; this does not require a judge option on the Eval. By default, the model comes from the project’s judgeRuntime. Set judge: "judge-model" to use a different model for one Eval.
Register a separate Judge for each requirement that can fail, contribute score, or show its own reason. In a score-style Eval, one Judge can combine .gate(minimum) and .score(points); the measurement multiplied by points is the contribution, and the gate determines the Verdict. Each Judge runs once.

Use real commands to verify coding tasks

When the grading standard itself is a file, use loadText to read hidden tests, a reference implementation, or a test script. A file change reruns the corresponding eval. See Grade Hidden Tests.

Practical advice

  • Register every condition that must hold as a Boolean Assertion so one Attempt gathers complete failure information.
  • When later steps depend on one result, use await handle.orStop().
  • Use defineScoreEval when partial completion is meaningful; a normal test return closes it automatically.
  • Use Judge for open-ended semantics. Prefer deterministic Matches for exact files, commands, and structured output.