Who is at the frontier of terminal tasks?
We ran the top 20 models from Terminal-Bench 2.1 on TB-fn, our harder variant of the same 89 tasks. Only OpenAI and Anthropic held the top tier, and every model cost more per task.
We ran the top 20 models from Terminal-Bench 2.1 on TB-fn, our harder variant of the same 89 tasks. Only OpenAI and Anthropic held the top tier, and every model cost more per task.
We rebuilt Terminal Bench 2 into task variants that expose broken verifiers, measure real robustness, and guide concrete improvements to the agent harness.