Terminal-Bench is a collection of harbor-native benchmarks to help agent makers quantify their agents' terminal mastery.
Last run Aug 4, 2026, 07:03 AM · composed from 9 runs
Powered by NiceEval
Last run · Aug 4, 2026, 04:04
Terminal-Bench 是一套 Harbor 原生的终端 agent 基准,用来量化 agent 在真实终端环境里的掌握程度。
最后运行 2026年8月4日 07:03 · 由 9 次运行合成