Skip to main content
Adapter defines the send contract: it receives TurnInput and AgentContext, and returns a Turn. This tutorial starts by sending one message, then progressively adds the capabilities needed for a complete integration. Each step adds only a small amount of code; existing parts are marked // … omitted. At the end of every step, you will find the assertions the new data supports. After completing any step, write those assertions in an Eval, rerun npx niceeval exp, then inspect multi-turn traces, usage, tool events, or pending requests in niceeval view. Three principles run through the tutorial:
  • Connect the interface your users already use. The Adapter calls the same endpoint and receives the same shape. Do not create an interface just for Evals or import internal application code to call a function directly. See Connect Your Agent for why.
  • Hand-write transport only. The URL, authentication, and request body depend on the application. niceeval/adapter supplies converters from raw responses to the standard event stream. ctx.session supplies the state APIs needed to continue sessions and resume a paused HITL interaction.
  • Send runtime feedback through ctx, not directly to the terminal. Use ctx.progress(...) for long steps. Use ctx.diagnostic(...) for degradation or error context you need to inspect after a run. Throw when execution cannot continue. Do not call console.log/error or write process.stdout/stderr from an Adapter.

Confirm the application interface shape

NiceEval does not define a new application protocol. Existing applications commonly use one of these protocols or a variation of one. Built-in converters support these response shapes: The following steps use a Chat Completions-shaped interface. Steps two and four show the replacements for a Responses-shaped interface and streaming interfaces.

Step 1: send one message and receive a reply

The smallest send does three things: send input.text to the application’s interface, put the reply in one message event, and report this Turn’s status:
When an interface returns a response you can continue processing but whose evidence is incomplete, report a diagnostic instead of printing the full raw response:
progress is not persisted. A diagnostic is written to an Attempt-owned channel and can be reviewed from the selected Run’s details page. Throw directly when an HTTP connection fails or a response cannot be parsed. The Runner records the failure in agent.run and forms an errored Verdict in the niceeval.verdict channel. This step supports direct value checks on t.reply (t.check(t.reply, includes(...)) or pattern(...)), all Judge conversation material, and model comparison on the Experiment side. ctx.model comes from an Experiment’s model. The Runner passes it through unchanged; the Adapter only forwards it. See Tier for the integration tiers. It has two obvious limits: every Turn is a new conversation, so a second t.send cannot continue the first; and it cannot see tool calls at all. The next two steps solve one of those limits each.

Step 2: continue prior messages

The Runner promises one thing for sessions: every send in one session line receives the same ctx.session; a new session line—the Eval’s first Turn or the line after t.newSession()—receives a completely new one. How you continue a session depends on the application interface shape:
  • The client supplies complete history—the server is stateless and receives the full messages list each Turn, as with Chat Completions—→ create an Adapter-private slot with createSessionSlot<TMsg[]>(), then read and write it with ctx.session.get(slot) / set(slot, messages).
  • The server keeps history—the interface accepts a session ID, such as a Responses-shaped previous_response_id or an SDK-native session/thread—→ use ctx.session.id and ctx.session.capture(id).
The main example uses the first form:
Notice that there is no first-Turn branch. On a new session line, ctx.session.get(historySlot) naturally returns undefined, and ?? [] normalizes it to empty history. The first-send form is the natural result of a new line, not a condition you need to check. You do not need to declare any capability on defineAgent, either: connect ctx.session and multi-turn continues; omit it and every Turn starts a new conversation. When the application interface accepts a session ID, history lives on the server and the Adapter records only that ID. Change only these two places inside send:
capture writes only when no ID has been recorded. A backend that repeats an ID—or even changes it due to a fork—cannot overwrite the line currently being continued. This step supports multi-turn conversations and session isolation with t.newSession().

Step 3: record usage

An Agent that answers correctly but burns ten times as many tokens should not receive the same score as a frugal one. Usage is usage, the fourth field on Turn alongside events and status. When the application interface returns usage, report it truthfully; the Runner accumulates it Turn by Turn into the session line and the entire run. A Chat Completions-shaped response includes usage, so copy it over. The rest of send is exactly the same as step two:
The full Usage fields also include optional cacheReadTokens / cacheCreationTokens, reasoning-token count reasoningTokens, request count requests, and costUSD. Fill costUSD only when a Provider or Adapter truthfully returns observed USD cost. Never put upstream estimates, a model-catalog price table, or a local calculation there. Fill each field only when the interface actually reports it. If the interface reports no usage, omit usage entirely; other assertions are unaffected. The Runner always calculates estimatedCostUSD independently from the model, reported token usage, and its Config/runtime price table; observed costUSD does not change that calculation. The Experiment budget and t.maxCost() use the estimate. Fixed Inspection operations read sealed Usage and never consume the Runner estimate. When the Config/runtime price table cannot form a Runner estimate, t.maxCost() is unavailable; missing data must never become 0. When an explicit zero rate truly produces estimatedCostUSD: 0, 0 is a valid estimate and is still evaluated against the limit. This step supports t.maxTokens() / t.maxCost() assertions and usage in reports and niceeval view.

Step 4: parse tools into events

An application’s response contains more than reply text. In a Chat Completions-shaped response, tool_calls records the tools called in this Turn. An Adapter’s most important job is to normalize an interface response into the standard event stream: one object for each thing that happened this Turn, ordered by actual occurrence in Turn.events, with one of the ten types below. For actual field values, the contract page has a complete one-Turn example.
Parsing is a “response field → event” map. Hand-written, it looks like this; the rest of send is exactly the same as step two:
Usually you do not need to write this loop. When a response has a standard shape, one official converter line replaces all of it. The return value includes events, status, and even the manually copied usage from step three, so return it directly:
For a different interface shape, use the matching piece: Built-in converters work by response shape, not by assuming an application protocol. Only a delta stream without a ready-made reducer from the protocol side needs a mapping. That mapping declares only the operation corresponding to each frame; deltaStream handles concatenation, pairing, and landing timing. After normalization, the events you emit determine the assertion families Eval authors can write: A Chat Completions response does not guarantee a complete process record. An application can finish its tool loop server-side and return only a final answer. Therefore, turnFromChatCompletion returns no completeness proof. Positive assertions such as calledTool work, while negative assertions such as notCalledTool report incomplete evidence. The Responses protocol requires the output array to record the complete process, so turnFromResponses carries a completeness proof and negative assertions are trustworthy. The difference in trustworthiness comes from the interface contract. This step supports tool assertions including turn.calledTool(), turn.toolOrder(), turn.maxToolCalls(), and turn.noFailedActions(). For ordering across Turns, use session as the receiver instead.

Step 5: HITL

When an application stops mid-Turn to wait for a person—for tool approval or missing information—both sides of send have obligations:
  • The pausing Turn returns status: "waiting" and emits an input.requested event with a stable id for every pending question. t.check(turn.status, equals("waiting")) and t.requireInputRequest() read them, and answers match by this id.
  • The answer Turn reaches the Adapter from an Eval’s t.respond(...) as another ordinary send on the same session line with the same state. A human decision arrives structurally through input.responses; each entry carries requestId, optionId, or text. See Inputs for the Different Answers for the shapes. The Adapter gives that decision back to the application, then continues fetching the result. For a call a human rejects, set the tool operation.finished event’s status to "rejected", not "failed": rejection is a human decision, not a tool failure, so noFailedActions() does not misfire.
The “scene read halfway through when the Turn paused”—for example, an SSE stream read halfway through—also lives on ctx.session. Create an Adapter-private slot with createSessionSlot<Pending>(), call ctx.session.set(slot, pending) when it pauses, then call ctx.session.take(slot) at the beginning of the answer Turn to recover it. Taking clears it, so it is consumed once. HITL almost always occurs on a streaming interface because the pause happens in the middle of a stream. This example therefore switches to an application that passes native events through SSE. It exercises every preceding step together; see the complete runnable version in the tier1 example:
For an interface that does not need HITL, delete the three pause-scene parts—Pending, its slot, and the opening take branch—and keep the rest. See HITL for the complete mental model of pausing, answering, and resuming. This step supports t.check(turn.status, equals("waiting")), t.requireInputRequest(), and t.respond() / t.respondAll(). You can assert a human-rejected call precisely with calledTool(toolMatch(..., { status: "rejected" })).

Step 6: connect OTel traces

When an application is already instrumented—standard OTel HTTP server instrumentation is enough—the integration has two halves: one is startup configuration, and the other belongs in send. Distinguishing them means distinguishing what never changes from what changes every Turn. The endpoint is startup configuration; it is not passed from send. NiceEval’s OTLP receiver address is the same for every run, so it does not go through ctx. Fix the receiver port in niceeval.config.ts, point the application’s OTel exporter at that fixed URL during startup, and you never need to change it again no matter how many Evals follow:
What send passes is this Turn’s trace context, not the endpoint. ctx.telemetry.headers is a W3C traceparent header the Runner generates freshly for every Turn. Spread it into the request and the spans the application produces for this Turn attach precisely to this Turn’s trace, without misattribution when Evals run concurrently. Back in the main chat-app, send gains exactly one line:
ctx.telemetry appears only when OTel integration is configured. Spreading undefined when it is not configured is safe, so this line can remain permanently. Spans still arrive without this header, but attribution degrades to time windows and that Agent’s Turns fall back to serial execution. Carrying it is what makes attribution accurate under concurrency. This step supports the per-Turn call waterfall in niceeval view, including model calls, tool execution, duration, and tokens. Assertions still read events produced by earlier steps; spans are only for the waterfall. For receiver configuration and span-attribution rules, see OTel Integration.

Step 7: pass through Experiment flags for A/B comparison

Once an application exposes variants as switchable configuration, an Experiment declares flags, and the Runner passes them unchanged through ctx.flags to send every Turn. The Adapter does not interpret their meaning; it forwards them in the request and the application switches variants from the parameter:
Two Experiment files, each declaring its own flags, can run the same Evals separately through npx niceeval exp to form an A/B comparison. This is Tier 3 of the three integration tiers because the application must cooperate by exposing switches. For the cost and payoff, see Tier. See Experiments and Lifecycles for flags, model, attempts, and the other Experiment fields. This step supports score comparison across variants over the same Evals.

Reference: the five t APIs as send sees them

After writing the seven steps, look back at the five Eval-side driving APIs. Reaching send, they are one function receiving different fields; there is no second method to implement:
  • Adapter — send inputs, outputs, three integration tiers, and capability sources.
  • Connect Your Agent — minimal integration, parameter passing, and optional capabilities.
  • HITL — the complete concept of pausing to wait for a person: handshake timing and both sides’ obligations.
  • Drive — how an Eval uses t.send(), t.newSession(), and HITL.
  • Assert — the complete assertion vocabulary driven by the standard event stream.