Skip to main content
Connecting NiceEval is not an all-or-nothing project. Based on where the Adapter connects and what additional observability data it receives, integration has three tiers. Each tier adds one kind of investment to the previous tier and provides a new set of capabilities; every Eval you have already written keeps working unchanged.

Tier 1: connect only send

You do not change a line of application code. Integration is a single Adapter file whose send communicates through the same interface your user-facing frontend already uses. The full assertion set is available at this tier. That includes text assertions and Judge checks, structured-output validation, multi-turn and session isolation, HITL approval flows, tool assertions from event mapping, and usage assertions from the usage returned by send. Every Verdict comes from the Turn returned by send; the next two tiers add no assertions. If the application interface exposes model selection, model comparison Experiments also belong here: the Experiment’s model reaches the Adapter through ctx.model and only needs to be forwarded with the request. Most teams start here, and many need nothing more.

Tier 2: send + OTel

Use the same send and event mapping, but have the application also send its OTel spans to NiceEval. If the application is already instrumented with AI SDK telemetry, LangGraph, OpenLLMetry / OpenInference, or its own gen_ai spans, this requires no application change. Otherwise, add generic OTel initialization. That is observability work, not a change customized for evals. This tier provides observability: the call waterfall in niceeval view. Each model call and tool execution inside the application appears in a per-turn timeline with its latency and tokens. Assertions are unaffected: spans go only into the waterfall, not the event stream or assertions. See OTel Integration for the path.

Tier 3: application changes + Experiment flags

To evaluate a feature A/B test such as “which Prompt, tool set, or feature toggle works better?”, the comparison point is inside the application, where the first two tiers cannot reach it. This tier changes application code to expose variants as externally selectable configuration. An Experiment’s flags pass through ctx.flags to the Adapter; the Adapter sends them to the application with the request, using an HTTP header, request-body field, or environment variable. The application chooses the variant from those parameters. What changes is the application, not the integration surface. The Adapter stays the same and still calls the application only through its external interface.

Non-intrusive is the baseline for the first two tiers

The first two tiers are non-intrusive. You start the application in its normal way, such as pnpm start or a deployment wherever you run it. The eval side does not spawn the application process or open another port. The Adapter only communicates with the same interface your user-facing frontend already uses, such as an HTTP endpoint or SSE stream. If it cannot connect, it reports a clear “start the application first” error. The only difference is observability data: Tier 1 has only the events returned by send, while Tier 2 has an additional copy of spans sent to NiceEval, so the waterfall appears.

How to move up

The three tiers are progressive rather than mutually exclusive. Start at Tier 1 for a baseline and model comparisons. Move to Tier 2 when you need the call waterfall. Move to Tier 3 when you need to compare variants inside the application. Each upgrade only adds something to the Adapter or application. See Experiments for how to organize comparisons on the Experiment side.
  • Connect Your Agent — The integration overview: minimal integration and parameter channels.
  • Adapter — The contract itself: what send receives and returns, and how ctx.model / ctx.telemetry / ctx.flags appear at each tier.
  • OTel Integration — The full Tier 2 path.
  • Experiments — How to declare model / flags comparisons in an Experiment.