query operations for automation and local view to review the same facts.
It can evaluate Plugins, Hooks, and Skills for Claude Code / Codex, and it can also evaluate your own AI Agent application directly. Whether your Agent uses AI SDK, LangGraph, Pi, or a custom Agent SDK, you can connect it to the same eval system through an Adapter.
Build an eval
Why use NiceEval when DeepEval, LangFuse, and BrainTrust already exist
NiceEval is an Agent-native eval tool. The Dataset / golden pattern of constructing Input and Expected Output does not fit real Agent evaluation. Modern Agents need checks for fine-grained situations such as multi-turn conversations, multi-Agent collaboration, tool calls, and Skill loading. NiceEval is designed for those situations: assertions operate directly on tool calls, message content, structured output, and usage instead of requiring you to prepare a golden test set by hand. Tools such as LangFuse and BrainTrust focus more on tracing and monitoring. That makes them too heavy for the path of writing evals, running evals, and inspecting results. NiceEval is designed for a friendly developer experience on that path from the beginning. The tools do not conflict: NiceEval can coexist with LangFuse and BrainTrust. Use them for tracing, or upload eval results to them.What you can evaluate
Coding Agent
Put Claude Code, Codex, or bub in a Sandbox, give it a task, then verify the result with real tests and file assertions.
Your own AI Agent application
Connect directly to your application’s interface and assert replies, tool calls, and structured output. Use Flags to test different Prompt versions.
Two integration modes
- Sandbox mode
- Direct mode
This mode suits coding Agents such as Codex and Claude Code that must edit code and run commands in a real file system.
Core concepts at a glance
For the complete terminology map, see Architecture Overview.
Start from your scenario
Connect Your Agent
Connect your Agent and write an Adapter that translates messages, tool calls, and structured output into the standard event stream.
Evaluate Coding Agent Extensions
Measure the effect of Skills, Prompts, and Plugin Benchmarks with a real Workspace and controlled experiments.
Learn to Write Evals and Experiments
Learn to write evals, assertions, Experiment configuration, and reports.