niceeval dependency, niceeval init has run, and you reached this page from the packaged INDEX.md. Speak with the user in their language. Treat packaged documentation as the authority for every API, field, and CLI behavior; do not invent details from training memory.
Step 1: Explore the project, then confirm the path with the user
This step determines the rest of the workflow. Inspect the code first, present the findings for the user to confirm, and ask only about facts you could not discover. Do not open with a long questionnaire or make an unexamined assumption. Discover:-
What kind of Agent this is. Read the README,
package.jsondependencies, routes, and Agent-loop code. Identify its stack—AI SDK, LangGraph, OpenAI Agents SDK, Claude Agent SDK, or a custom loop—and its core use case, such as support, SQL, or coding work. -
How the frontend and Agent communicate. Is it HTTP, gRPC, or WebSocket? Is the protocol standard or custom: AI SDK UI Message Stream, OpenAI Responses or Chat Completions, an SDK-native event stream passthrough, or custom JSON/SSE frames? This directly determines whether to use a built-in Adapter with zero mapping or hand-write
sendwith event mapping. - Whether the backend already has OTel. Look for OTel SDK initialization, AI SDK telemetry, LangSmith, OpenLLMetry, OpenInference, or related instrumentation. An existing setup makes Tier 2 nearly free.
-
Whether the user already has A/B tests or feature flags. Existing variation switches are an immediate Tier 3 entry point because an Experiment can pass them through as
flags. -
Which Judge to use. Semantic evaluation with Judge Matches from
niceeval/expectneeds a Judge model separate from the Agent under test and an OpenAI-compatible/chat/completionsendpoint. Ask which service key the user has and which Judge model they want. A missing key does not silently pass: a Judge Assertion becomesunavailable; when it participates in pass grading, or has.score()/.orStop()in a scored eval, the Attempt’s grading is unavailable. Without a Judge configuration, start with precise Matches only. -
Whether the Agent itself needs a Sandbox. A coding-agent CLI, or a Skill, Plugin, Hook, or MCP server written for a coding agent, must edit files and run commands in an isolated Workspace. It must use a Sandbox rather than an HTTP-service-style
sendpath. RecommenddockerSandbox({ source: { type: "image", image: "node:24-slim" } })by default. Follow a Provider already specified in the task or configuration; ask only when the Provider, Docker availability, or remote-execution need would change the plan. The Provider is declared on the eval or Experiment; there is no CLI flag, project-wide default, or auto-detection. See Choose a Sandbox Provider.
- Tier 1 (send only): No application change. The full assertion set—text, Judge, multi-turn, tools, and HITL—already works here.
- Tier 2 (send + OTel): The application also sends OTel spans to NiceEval, where you can review the call waterfall in
niceeval view. Existing instrumentation makes this nearly free. - Tier 3 (invasive changes + flags): Expose internal variations as
flagsfor feature A/B testing. Existing A/B switches are ready-made entry points.
Step 2: Configure a Judge
Configure the Judge service discovered in Step 1. Judges use an OpenAI-compatible/chat/completions endpoint in niceeval.config.ts:
baseUrl explicitly for a compatible gateway. If you set only a key and omit baseUrl, NiceEval calls the official endpoint. The gateway credential goes to OpenAI and returns “Incorrect API key provided.” It looks like an expired key, but the actual problem is the wrong endpoint.
Tell the user three things:
- A misconfiguration does not pretend to pass. If the model or key cannot be resolved, the gateway rejects authentication, or an evaluation request times out, the Judge Assertion is
unavailablewith a reason. The reading surface does not disguise a missing measurement as a mismatch or zero. - Verify it once after configuration. Use
defineJudgeto declare a rubric and anchors, then callt.judge(material, definition).gate(0.8). After it runs, useniceeval view --run <run-id>to inspect the result. - Do not skip a Judge directly when no user is present during autonomous onboarding. First check whether the environment already has a usable key, such as
OPENAI_API_KEYorDEEPSEEK_API_KEY. If it does, pointapiKeyEnvto it and verify. If not, write precise Matches first and state at handoff that no Judge is configured. - Keep the Judge model separate from the Agent under test so a model does not score its own output. Configuration resolves field by field across the eval, Experiment, and project configuration. See Judge for the three factories.
Step 3: Write the three pieces
After reading the documentation selected in Step 1, write these in order:- Adapter (
agents/*.ts, or the directory established by the user’s project). UsedefineAgentto implementsend. Put configuration in factory parameters; do not hard-code it or readprocess.env. See the Adapter contract, Adapter Reference signature, and Event Reference for event mapping. Two easy mistakes matter: choose an endpoint or mode that exercises the system under test’s core capability, not merely the easiest path to make work—if a platform has both plain LLM chat and a database-connected execution mode, connecting the former evaluates a low-level model proxy rather than the product. Also, declareevidenceCoveragetruthfully based on the actual mapping. If you map only final text, do not usecompleteEvidenceCoverage; the declaration changes assertion completeness, and overstating it is worse than being conservative. - Experiment (
experiments/*.ts). Reference the Adapter above and declaremodel,flags,attempts, and so on. Model comparison uses two Experiment files, each pinned to onemodel.evals: (eval) => booleandecides which evals each one runs. Paths only provide IDs and batch execution; fixed Inspection operations consume the current published physical results and coverage facts. - Eval (
evals/*.eval.ts). First discover what the application does, then write one eval aligned with a real capability. Read its README, routes, tool definitions, or system prompt to find the core use case. A support bot should receive a real support question; a SQL Agent should receive a real query task. Use that as the input and Assertions for the first eval. Two inputs are unacceptable: a placeholder unrelated to the application such as “Hello,” and a meta-question asking the system under test what it is or can do. Those are not tasks users ask it to perform. Start with the smallest shape—one input,t.succeeded(), and one content Assertion about the expected reply—but that shape is only scaffolding. Before handoff, it must also meet both conditions:- A content Assertion turns red when the system under test hallucinates. Do not assert a word already in the input—asking “What is X?” and then asserting that the reply contains “X” passes when the system simply repeats the prompt. Assert substantive content unique to the expected answer: concrete facts, structure with
hasSections(), a real link withincludesUrl(), or a Judge Match registered throughcheck. - Include at least one negative case. Give input the system should not be able to answer, such as a nonexistent table or unavailable retrieval topic. Assert that it clearly says it cannot find or do it instead of fabricating a plausible-looking result. For Agents connected to real data or retrieval sources, this is the failure shape most worth testing first. See Authoring Evals for the form, Evaluation Kinds and Assertions for Assertions and kinds, and defineEval Reference for the signature.
- A content Assertion turns red when the system under test hallucinates. Do not assert a word already in the input—asking “What is X?” and then asserting that the reply contains “X” passes when the system simply repeats the prompt. Assert substantive content unique to the expected answer: concrete facts, structure with
ctx), see Connect Your Agent.
Two architecture rules are nonnegotiable when writing the Adapter:
- Do not make an in-process direct call. Even if the Agent runtime and eval live in the same repository, the Adapter must use HTTP or the appropriate transport layer. Do not replace
fetchwith a direct import of the function under test. See “Why not call it directly?” in Connect Your Agent. - Do not manage the system-under-test process from the eval side. Do not spawn the application or open another port. The user starts it normally, such as with
pnpm dev. If the Adapter cannot connect, report a direct error such as “start the application first”; do not start a service itself.
Step 4: Run and verify
niceeval view. When a machine must read the result, use a fixed request discovered by query discover. See View Results for viewer usage. When it does not run, triage three ways: a thrown fetch error means the application is not running or the URL is wrong. A failed t.succeeded() means the application returned a nonsuccess status. If only a content Assertion fails, the integration works; adjust the Assertion or the application.
Within the user’s existing authorization and budget, proactively repair an execution-chain error and rerun only the affected scope. Do not modify or rerun for a read-only request. When a real external condition such as the target service, port, credential, database, or model service is missing, report that blocker plainly; do not present a temporary substitute as a run against the real target.
Step 5: Finish by telling the user what changed
Before summarizing, go through this completion checklist. If any item is not true, return to Step 3 and fill it in; do not gloss over it in the summary.- The eval input is a core use case of the system under test, not a meta-question about itself or a placeholder greeting.
- Every content Assertion turns red when the system repeats the prompt or hallucinates; asserted words do not already occur in the input.
- At least one negative case provides input that should not be answerable and asserts an explicit inability to do it.
- When a key is available, the Judge is configured and a Judge score has appeared in
niceeval view. When no key is available, the summary says so. - The Experiment’s declared
model/flagsare actually consumed by the Adapter. There is no dead configuration and no invented model value the system under test does not support.
niceeval exp <experiment-path> and niceeval view, and what the first result looks like. Do not refactor existing user code or add abstractions outside these files unless asked.
Step 6: Ask whether the user wants a deeper integration
After the summary, present optional next steps. For each, explain what it enables, roughly how much code it changes, and what benefit it buys. Let the user choose; do not continue on your own:
Also tell the user what these options have in common: they are incremental additions to the Adapter or application, and the evals already written do not need to change by one line. See Integration Tiers for what the three levels buy and when upgrading is worthwhile. If Step 1 found existing OTel instrumentation, proactively recommend the call-waterfall option because its cost is close to zero.