Skip to main content
Data-driven testing (dataset fan-out) fits many tests with the same structure and different inputs. Typical examples include SQL generation, intent classification, retrieval QA, and tool selection.

How it works

When there is no external business ID, a .eval.ts file default-exports an array:

Generated IDs

If the file is evals/sql.eval.ts, the generated IDs are:
The numeric suffix is zero-padded, making IDs stable to reference and filter. When the data source already carries a stable case, issue, or benchmark ID, default-export a keyed record instead:
If the file is evals/swelancer.eval.ts and the key is 15193, the ID is swelancer/15193. A key must be a nonempty path segment: it cannot be . or .. and cannot contain /, \\, or control characters. NiceEval discovers keys in lexicographic order, so a changed data-source order does not change execution order.

Load from YAML or JSON

loadYaml and loadJson require a decoder. An unvalidated dynamic value exists only at the decoder boundary. After validation, the returned value is strongly typed data you can use directly.

Filter dataset evals

Datasets versus separate files

Cases have exactly the same structure; the main differences are the input and expected output.
Datasets fit broad horizontal coverage. Separate eval files fit behaviorally complex scenarios.