Skip to main content
NiceEval is an Agent eval tool that helps teams measure, evaluate, and improve AI in production. With NiceEval, teams can compare models, iterate on Agents, find regressions, and continuously improve AI applications with real user data. NiceEval is local-first: your evals run in your own environment. Published Attempts enter the project’s single Record; teams use fixed query operations for automation and local view to review the same facts. It can evaluate Plugins, Hooks, and Skills for Claude Code / Codex, and it can also evaluate your own AI Agent application directly. Whether your Agent uses AI SDK, LangGraph, Pi, or a custom Agent SDK, you can connect it to the same eval system through an Adapter.

Build an eval

Why use NiceEval when DeepEval, LangFuse, and BrainTrust already exist

NiceEval is an Agent-native eval tool. The Dataset / golden pattern of constructing Input and Expected Output does not fit real Agent evaluation. Modern Agents need checks for fine-grained situations such as multi-turn conversations, multi-Agent collaboration, tool calls, and Skill loading. NiceEval is designed for those situations: assertions operate directly on tool calls, message content, structured output, and usage instead of requiring you to prepare a golden test set by hand. Tools such as LangFuse and BrainTrust focus more on tracing and monitoring. That makes them too heavy for the path of writing evals, running evals, and inspecting results. NiceEval is designed for a friendly developer experience on that path from the beginning. The tools do not conflict: NiceEval can coexist with LangFuse and BrainTrust. Use them for tracing, or upload eval results to them.

What you can evaluate

Coding Agent

Put Claude Code, Codex, or bub in a Sandbox, give it a task, then verify the result with real tests and file assertions.

Your own AI Agent application

Connect directly to your application’s interface and assert replies, tool calls, and structured output. Use Flags to test different Prompt versions.

Two integration modes

This mode suits coding Agents such as Codex and Claude Code that must edit code and run commands in a real file system.

Core concepts at a glance

For the complete terminology map, see Architecture Overview.

Start from your scenario

Connect Your Agent

Connect your Agent and write an Adapter that translates messages, tool calls, and structured output into the standard event stream.

Evaluate Coding Agent Extensions

Measure the effect of Skills, Prompts, and Plugin Benchmarks with a real Workspace and controlled experiments.

Learn to Write Evals and Experiments

Learn to write evals, assertions, Experiment configuration, and reports.
Quickstart walks you through installing NiceEval, building an eval, and reviewing the results published by your first run.