Skip to main content
A Coding Agent can complete a clear feedback loop from the command line: read the installed documentation, run an Experiment, read the final receipt, use its createdRunIds to inspect results, update the program or eval, and run again.

Start with the installed documentation

NiceEval packages its Chinese documentation in docs-site/zh/ and provides an INDEX.md at the package root for Coding Agents. A Coding Agent should read node_modules/niceeval/INDEX.md first, then use the index to reach pages needed for the current task. Do not rely on training data or online examples from another version. This keeps API and CLI guidance aligned with the installed version. npx niceeval init initializes configuration and writes managed guidance into the project’s AGENTS.md. If a project has only CLAUDE.md, it writes there instead. If neither file exists, it creates AGENTS.md. Run init again after upgrading NiceEval to refresh the managed block. The managed guidance separates writing source from proving what happened in a Run. An Agent can read project source such as evals/ and agents/ while writing code. To prove historical execution, it uses fixed niceeval query operations or niceeval view. Do not scan raw .niceeval/ files or infer historical execution from current source. If the public read surface lacks needed evidence, report a NiceEval presentation gap. You can give an AI this starting task:
That keeps the Agent’s documentation aligned with the version installed in the project.

Run and retain the receipt

Every --json line is feedback from the current process. Progress and diagnostics describe what is happening in this invocation. Exactly one final receipt contains:
The Agent should treat the final receipt as the handoff for this invocation. createdRunIds exhaustively lists successfully created Runs, and publicationCutoff is an opaque boundary of published facts; neither proves that the task is complete. The receipt is not a persisted-results protocol; read business facts from the published Record through its createdRunIds.

Inspect results with Run IDs

After the run ends, inspect the Run’s Report:
query returns fixed closed Inspection results, while view supplies fixed first-party Insight. The Agent should distinguish these states:

Run again after an edit

Have the Agent test one verifiable hypothesis at a time: change the program or eval, run the same scope, then read the new receipt and Run. To confirm that every slot executes again, use:
When an existing Attempt is adopted automatically, --dry explains why. Carry and acceptance reasons are saved with the target Run and visible in results. See Rerun and Carry Results. A completed run hands off this Invocation only; it does not prove that the user’s task is complete. When the task requires completion and authorization already covers the work, repair an execution-chain error and rerun only the affected scope. A read-only request performs attribution only, without edits or reruns. Preserve and report a business failure that satisfies the task; do not rewrite the application or eval merely to turn it green.

Record boundaries

A Record stores published facts only and cannot be modified after publication. To obtain a different result, an Agent changes the program or eval, runs a new Invocation, then reads the new receipt’s createdRunIds. To share run facts with another compatible runtime, copy the canonical Record, not feedback from the current process.