Skip to main content
This page is the execution path for a Coding Agent that brings NiceEval into a project. It assumes the project already has the niceeval dependency, niceeval init has run, and you reached this page from the packaged INDEX.md. Speak with the user in their language. Treat packaged documentation as the authority for every API, field, and CLI behavior; do not invent details from training memory.

Step 1: Explore the project, then confirm the path with the user

This step determines the rest of the workflow. Inspect the code first, present the findings for the user to confirm, and ask only about facts you could not discover. Do not open with a long questionnaire or make an unexamined assumption. Discover:
  1. What kind of Agent this is. Read the README, package.json dependencies, routes, and Agent-loop code. Identify its stack—AI SDK, LangGraph, OpenAI Agents SDK, Claude Agent SDK, or a custom loop—and its core use case, such as support, SQL, or coding work.
  2. How the frontend and Agent communicate. Is it HTTP, gRPC, or WebSocket? Is the protocol standard or custom: AI SDK UI Message Stream, OpenAI Responses or Chat Completions, an SDK-native event stream passthrough, or custom JSON/SSE frames? This directly determines whether to use a built-in Adapter with zero mapping or hand-write send with event mapping.
  3. Whether the backend already has OTel. Look for OTel SDK initialization, AI SDK telemetry, LangSmith, OpenLLMetry, OpenInference, or related instrumentation. An existing setup makes Tier 2 nearly free.
  4. Whether the user already has A/B tests or feature flags. Existing variation switches are an immediate Tier 3 entry point because an Experiment can pass them through as flags.
  5. Which Judge to use. Semantic evaluation with Judge Matches from niceeval/expect needs a Judge model separate from the Agent under test and an OpenAI-compatible /chat/completions endpoint. Ask which service key the user has and which Judge model they want. A missing key does not silently pass: a Judge Assertion becomes unavailable; when it participates in pass grading, or has .score() / .orStop() in a scored eval, the Attempt’s grading is unavailable. Without a Judge configuration, start with precise Matches only.
  6. Whether the Agent itself needs a Sandbox. A coding-agent CLI, or a Skill, Plugin, Hook, or MCP server written for a coding agent, must edit files and run commands in an isolated Workspace. It must use a Sandbox rather than an HTTP-service-style send path. Recommend dockerSandbox({ source: { type: "image", image: "node:24-slim" } }) by default. Follow a Provider already specified in the task or configuration; ask only when the Provider, Docker availability, or remote-execution need would change the plan. The Provider is declared on the eval or Experiment; there is no CLI flag, project-wide default, or auto-detection. See Choose a Sandbox Provider.
When the task has not already selected a tier, introduce the integration tiers and make a recommendation:
  • Tier 1 (send only): No application change. The full assertion set—text, Judge, multi-turn, tools, and HITL—already works here.
  • Tier 2 (send + OTel): The application also sends OTel spans to NiceEval, where you can review the call waterfall in niceeval view. Existing instrumentation makes this nearly free.
  • Tier 3 (invasive changes + flags): Expose internal variations as flags for feature A/B testing. Existing A/B switches are ready-made entry points.
Recommend Tier 1 first, then Tier 2 by default. Especially when Step 3 finds existing OTel instrumentation, say explicitly: “Tier 2 only sends one more copy of the spans, so its cost is close to zero.” Propose Tier 3 only when the user explicitly wants variation comparison. Choose the right documentation for the discovered shape. Do not start writing an Adapter before reading it:

Step 2: Configure a Judge

Configure the Judge service discovered in Step 1. Judges use an OpenAI-compatible /chat/completions endpoint in niceeval.config.ts:
Write baseUrl explicitly for a compatible gateway. If you set only a key and omit baseUrl, NiceEval calls the official endpoint. The gateway credential goes to OpenAI and returns “Incorrect API key provided.” It looks like an expired key, but the actual problem is the wrong endpoint. Tell the user three things:
  • A misconfiguration does not pretend to pass. If the model or key cannot be resolved, the gateway rejects authentication, or an evaluation request times out, the Judge Assertion is unavailable with a reason. The reading surface does not disguise a missing measurement as a mismatch or zero.
  • Verify it once after configuration. Use defineJudge to declare a rubric and anchors, then call t.judge(material, definition).gate(0.8). After it runs, use niceeval view --run <run-id> to inspect the result.
  • Do not skip a Judge directly when no user is present during autonomous onboarding. First check whether the environment already has a usable key, such as OPENAI_API_KEY or DEEPSEEK_API_KEY. If it does, point apiKeyEnv to it and verify. If not, write precise Matches first and state at handoff that no Judge is configured.
  • Keep the Judge model separate from the Agent under test so a model does not score its own output. Configuration resolves field by field across the eval, Experiment, and project configuration. See Judge for the three factories.

Step 3: Write the three pieces

After reading the documentation selected in Step 1, write these in order:
  1. Adapter (agents/*.ts, or the directory established by the user’s project). Use defineAgent to implement send. Put configuration in factory parameters; do not hard-code it or read process.env. See the Adapter contract, Adapter Reference signature, and Event Reference for event mapping. Two easy mistakes matter: choose an endpoint or mode that exercises the system under test’s core capability, not merely the easiest path to make work—if a platform has both plain LLM chat and a database-connected execution mode, connecting the former evaluates a low-level model proxy rather than the product. Also, declare evidenceCoverage truthfully based on the actual mapping. If you map only final text, do not use completeEvidenceCoverage; the declaration changes assertion completeness, and overstating it is worse than being conservative.
  2. Experiment (experiments/*.ts). Reference the Adapter above and declare model, flags, attempts, and so on. Model comparison uses two Experiment files, each pinned to one model. evals: (eval) => boolean decides which evals each one runs. Paths only provide IDs and batch execution; fixed Inspection operations consume the current published physical results and coverage facts.
  3. Eval (evals/*.eval.ts). First discover what the application does, then write one eval aligned with a real capability. Read its README, routes, tool definitions, or system prompt to find the core use case. A support bot should receive a real support question; a SQL Agent should receive a real query task. Use that as the input and Assertions for the first eval. Two inputs are unacceptable: a placeholder unrelated to the application such as “Hello,” and a meta-question asking the system under test what it is or can do. Those are not tasks users ask it to perform. Start with the smallest shape—one input, t.succeeded(), and one content Assertion about the expected reply—but that shape is only scaffolding. Before handoff, it must also meet both conditions:
    • A content Assertion turns red when the system under test hallucinates. Do not assert a word already in the input—asking “What is X?” and then asserting that the reply contains “X” passes when the system simply repeats the prompt. Assert substantive content unique to the expected answer: concrete facts, structure with hasSections(), a real link with includesUrl(), or a Judge Match registered through check.
    • Include at least one negative case. Give input the system should not be able to answer, such as a nonexistent table or unavailable retrieval topic. Assert that it clearly says it cannot find or do it instead of fabricating a plausible-looking result. For Agents connected to real data or retrieval sources, this is the failure shape most worth testing first. See Authoring Evals for the form, Evaluation Kinds and Assertions for Assertions and kinds, and defineEval Reference for the signature.
For how parameters flow from Experiment to Adapter, and the difference between static factory configuration and dynamic values per turn (ctx), see Connect Your Agent. Two architecture rules are nonnegotiable when writing the Adapter:
  • Do not make an in-process direct call. Even if the Agent runtime and eval live in the same repository, the Adapter must use HTTP or the appropriate transport layer. Do not replace fetch with a direct import of the function under test. See “Why not call it directly?” in Connect Your Agent.
  • Do not manage the system-under-test process from the eval side. Do not spawn the application or open another port. The user starts it normally, such as with pnpm dev. If the Adapter cannot connect, report a direct error such as “start the application first”; do not start a service itself.

Step 4: Run and verify

Use Coding Agent Feedback Loop to complete the cycle of run, inspect, modify, and rerun with niceeval view. When a machine must read the result, use a fixed request discovered by query discover. See View Results for viewer usage. When it does not run, triage three ways: a thrown fetch error means the application is not running or the URL is wrong. A failed t.succeeded() means the application returned a nonsuccess status. If only a content Assertion fails, the integration works; adjust the Assertion or the application. Within the user’s existing authorization and budget, proactively repair an execution-chain error and rerun only the affected scope. Do not modify or rerun for a read-only request. When a real external condition such as the target service, port, credential, database, or model service is missing, report that blocker plainly; do not present a temporary substitute as a run against the real target.

Step 5: Finish by telling the user what changed

Before summarizing, go through this completion checklist. If any item is not true, return to Step 3 and fill it in; do not gloss over it in the summary.
  • The eval input is a core use case of the system under test, not a meta-question about itself or a placeholder greeting.
  • Every content Assertion turns red when the system repeats the prompt or hallucinates; asserted words do not already occur in the input.
  • At least one negative case provides input that should not be answerable and asserts an explicit inability to do it.
  • When a key is available, the Judge is configured and a Judge score has appeared in niceeval view. When no key is available, the summary says so.
  • The Experiment’s declared model / flags are actually consumed by the Adapter. There is no dead configuration and no invented model value the system under test does not support.
After the first run works, summarize before proposing next steps. State the connected system under test, where the Adapter, Experiment, and eval live, how to run niceeval exp <experiment-path> and niceeval view, and what the first result looks like. Do not refactor existing user code or add abstractions outside these files unless asked.

Step 6: Ask whether the user wants a deeper integration

After the summary, present optional next steps. For each, explain what it enables, roughly how much code it changes, and what benefit it buys. Let the user choose; do not continue on your own: Also tell the user what these options have in common: they are incremental additions to the Adapter or application, and the evals already written do not need to change by one line. See Integration Tiers for what the three levels buy and when upgrading is worthwhile. If Step 1 found existing OTel instrumentation, proactively recommend the call-waterfall option because its cost is close to zero.