runIds are also stable read boundaries. People use niceeval show in the terminal and niceeval view in the browser. AI, scripts, and CI use niceeval query to read JSON.
Start with the Overview in the terminal
show prints the Overview at the current PublicationCutoff. It expands summaries by Experiment, Eval, and Attempt while retaining pass rate, scores, coverage, missing items, and issues. Start here to find the Experiment, Run, or Attempt locator to inspect next.
The Overview comes from published results in the project’s single .niceeval/record.sqlite. Even when the identity of the current source or installed candidate has changed, show retains the latest result for every logical slot instead of displaying existing history as Observed 0/0. That makes a result viewable; it does not mean the result is eligible for reuse under the current target.
Narrow to an Experiment or Run
--experiment shows one Experiment’s aggregate results and Evals. --run shows the members, Verdicts, scores, coverage, and usage summary for one exact Run. Published Attempts from an active Run appear, while positions not yet published show pending. Both options accept multiple exact IDs, but cannot be mixed with each other or with an Attempt locator.
Inspect an Attempt
--source shows captured source and Assertion locations. --execution shows the bounded conversation, tool, and command outline. To expand one item, copy its stable itemId, toolOccurrenceId, or commandId from the outline:
--timing shows activity order and duration, --usage shows token, request, and cost totals, and --diff shows the captured file-change window. Details read only published Record facts. partial, not-recorded, and unavailable describe the available scope; do not treat them as zero values or failures.
Interpret completion, Verdicts, and scores correctly
First confirm whether the Attempt produced a business result, then interpret its quality:errored: execution stopped and produced no gradeable business conclusion.failed: execution completed, but a business Assertion in a pass/fail eval was not satisfied.passed: the pass/fail eval met its Verdict conditions. Still inspect warnings, coverage, and evidence limits.- Score Eval: after scoring completes, read the Score and individual Assertions. Being complete or passed in the Overview does not mean a perfect score.
unavailable as zero. To explain a specific gain or loss, keep the same Attempt locator and cross-check the overview, --execution, and --source.
Review in the browser
--run selects one or more exact Runs. It also shows published Attempts from an active Run and pending for slots not yet published.
Use --no-open to keep the browser from opening automatically, or --port to choose a loopback port. --json writes lifecycle NDJSON only; it does not write result data.
When new facts are published, project View asks you to confirm a refresh. It moves to the new PublicationCutoff only after confirmation.
Read JSON from automation
query is the machine entry point for AI, scripts, and CI. It writes only structured niceeval.query/v1 JSON and does not provide terminal formatting for people. Get an operation schema from the catalog returned by discover, then construct a request. run writes one niceeval.query/v1 document to stdout. Do not parse View pages or read unpublished Record files.
Fixed operations cover common reads: list Runs, get a Run summary, read Attempt detail, trace, diff, sources, artifacts, or compare two Runs. See the CLI for full details.
Read another set of published facts
show and query accept a copy of the canonical SQLite Record through --record and treat it as hostile input. If the exact current schema, SQLite integrity, or a domain invariant fails validation, the command rejects the entire file. view reads only the current project’s .niceeval/record.sqlite.