defineEval answers whether one run meets requirements, and the default Report reads pass rate. defineScoreEval answers how much work was completed, and the default Report reads accumulated score. Both kinds register an Assertion when you call it; the handle only configures that same entry.
Keep each Experiment and Eval Group homogeneous: all pass/fail Evals or all scored Evals. niceeval check, dry runs, and normal runs reject a mixed selection before an Agent or Sandbox starts. Split the selection into two Experiment or Group definitions when a project needs both readings.
Use defineEval for requirements that must hold
matched enters the Verdict. mismatched makes the final Verdict failed, but does not prevent other Assertions from continuing to register and settle. When later code depends on this result, use await handle.orStop().
Use a continuous quality threshold
A Match such assimilarity(...) produces a measurement in [0, 1]. It calculates character-level Levenshtein edit similarity, not semantic similarity. In a pass/fail Eval, call .gate(minimum) on the registered handle to set the quality threshold:
.gate(minimum), the measurement does not affect the Verdict. When you only want an explanation that does not affect the Verdict, use t.diagnostic(...); do not register a measurement that nobody consumes.
Award points for each completed step
Step-by-step tasks fitdefineScoreEval. Assertions save only evaluation by default; .score(n) is what makes one contribute points. A Boolean match contributes n, a mismatch contributes 0, and a measurement m contributes m * n.
t.score(n) directly registers a contribution. n must be finite and nonnegative, and its returned handle can configure only key and label.
When test returns normally, NiceEval closes the score automatically. No scoring entries is still a valid result with an official score: 0: evaluation successfully formed a zero score; it is not an execution failure or insufficient evidence.
Results for the two eval kinds
scored can be a zero score. Zero means evaluation successfully formed a score; it is not execution failure or insufficient evidence. Every Attempt still records a terminal Verdict, but the Report uses scored or errored as the primary result for a Score Eval; it does not convert that Verdict into a pass rate.
Judge is also a measurement
AdefineJudge call declares a managed Judge Match. Pass the values needed for the decision as fields in the material object. Register the definition directly with t.judge; this does not require a judge option on the Eval. By default, the model comes from the project’s judgeRuntime. Set judge: "judge-model" to use a different model for one Eval.
.gate(minimum) and .score(points); the measurement multiplied by points is the contribution, and the gate determines the Verdict. Each Judge runs once.
Use real commands to verify coding tasks
loadText to read hidden tests, a reference implementation, or a test script. A file change reruns the corresponding eval. See Grade Hidden Tests.
Practical advice
- Register every condition that must hold as a Boolean Assertion so one Attempt gathers complete failure information.
- When later steps depend on one result, use
await handle.orStop(). - Use
defineScoreEvalwhen partial completion is meaningful; a normaltestreturn closes it automatically. - Use Judge for open-ended semantics. Prefer deterministic Matches for exact files, commands, and structured output.