Skip to main content
Judge uses a separate model to assess factual consistency, summary faithfulness, or clarity of explanation. It is a managed evaluator. Ordinary custom ScoreMatch values remain pure; defineJudge, the built-in factories, and the managed defineScoreMatch path all use the same Assertion path.

Declare a scoring definition

Define the criterion at module scope. rubric describes the dimensions to evaluate, and anchors provide calibration points for a continuous quality scale:
When anchors is omitted, NiceEval uses 0 and 1 as the default endpoints. measurement may be any finite value in [0, 1]; the endpoints do not limit the result to two classes or represent model confidence.

Select evaluation material

Pass the values needed for the decision directly to t.judge. The Eval judge option only selects the model configuration for this Eval:
Material accepts strings, finite numbers, booleans, null, arrays, and plain objects. Use field names that express application meaning, such as change and explanation; you do not need to serialize the object first. NiceEval creates a bounded JSON snapshot when the Assertion is registered, so later mutations to the original object do not change the request. When you need to inspect image content, put judgeImage({ body, mediaType }) in the material. It freezes PNG or JPEG bytes without creating another Judge or making an additional request. The image and text are sent in the same Judge call and stored in the same Assertion evidence. See Add image material for the setup. t.judge(material, explainsRisk) and t.check(material, explainsRisk) each register one measurement Assertion. They share the material snapshot, budget, evaluation, audit, sealing, and handle; there is no Match allowlist.

Use built-in Judges directly

Root t, Session, and Turn provide built-in methods that return MeasurementAssertionHandle<Kind> Assertions. For example:
closeQA(selector, question, options?) sends all matched material to one Judge in its original order. The question is an acceptance criterion to answer using only the material, not an expected answer. Agent’s closeQA(question, options?) uses the current Turn, Session, or Attempt’s complete recorded history, including messages and tool facts. It does not select only the final reply or successful operations. An explicit EventMatch selects from the three event views. For questions about tool inputs and outputs, use a ToolMatch or custom MaterialMatch. These forms compose a scoring Match and pass it to the same check. Application assertions such as usedNoTools() encapsulate a fixed criterion and do not accept another Match. The decision is 1 when satisfied, 0 when not satisfied, and unavailable when evidence is insufficient. A complete empty collection yields 0 without a model call. Missing or oversized material is not filled with zero or truncated before scoring. For material from another application, use a custom MaterialMatch. An Adapter’s create provides the application context. The synchronous reader in defineMaterialMatch reads the current read-only context when the Assertion is registered. The selector predicate evaluates one item; and applies its conditions to the same item. Use t.check(selector).gate() for ordinary existence checks and t.closeQA(selector, question).gate(0.8) for a question about the whole matched collection. factuality, faithfulness, instructionFollowing, and pairwisePreference take explicit material, construct a Match, and pass it through the same check receiver.

Configure the Judge Provider

OpenAIProvider, VercelProvider, OpenRouterProvider, and TypesafeProvider are exported from niceeval/judge. Vercel denotes AI Gateway, not Vercel Sandbox or a general AI SDK model object. To change only the model for one Eval, set judge: "another-model"; the selected Provider’s endpoint, credential source, and execution limits remain in effect. An Eval or Experiment can also provide a complete Provider, which replaces the service, default model, and limits. NiceEval uses the selected Provider’s managed protocol to obtain a finite [0, 1] measurement and public rationale. Chat Providers use forced-function requests; TypeSafe uses /systemone. With no model or key, no network request is sent and the Assertion is unavailable. Transport failures and timeouts are unavailable; HTTP 400, protocol incompatibility, and invalid responses are errored. TypeSafe supports score, classify, and batch classify, but not extract. Therefore faithfulness() on TypesafeProvider is unavailable; it is not converted to an overall score or a different denominator.

Thresholds and scores

Pass-style and score-style Evals call .gate(minimum) on the registered handle to form a quality gate. Score-style Evals can also call .score(points); the contribution is measurement multiplied by points. Gate and score may be called in either order, and both execute the Judge only once.

Configuration, preflight, and failures

Static configuration validation does not probe the endpoint. A missing Provider or key leaves the Assertion unavailable; transport failures and timeouts are unavailable; HTTP 400, protocol incompatibility, and invalid responses are errored. The Judge’s reason goes in the general explanation. Sealed results retain the frozen material and the evidence sent to the model, including image evidence.

Reading results

niceeval view, fixed query operations, and failure feedback create summaries from the same sealed Assertion results. Failed and unavailable entries appear first. Configured measurements display their actual value and required threshold. Judge has no special Assertion-result or display branch. Successful requests save the exact material that was sent. A path, URL, alt text, or ordinary base64 string is not visual input; only explicit image material becomes an image part in the model API request.