Who is at the frontier of terminal tasks?
We picked the top 20 models from Artificial Analysis's Terminal-Bench 2.1 leaderboard and ran them on both TB-2.1 and TB-fn. TB-fn is Fidian's variant of Terminal-Bench, built from the same 89 tasks as TB-2.1. Only OpenAI and Anthropic models remained on the TB-fn frontier. GLM-5.3 leads the open-weight families on both benchmarks, but its gap to Sol max grows from 1.0 point on TB-2.1 to 8.6 points on TB-fn. Six models lost 6–11 points on TB-fn. Per-attempt cost increased by 41%.
Newly released models repeatedly appear near GPT-5.6 Sol on Artificial Analysis’s Terminal-Bench 2.1 leaderboard. Are these models actually on par with Sol on terminal tasks? We picked the top 20 models from that leaderboard and benchmarked them on TB-fn to answer that question. TB-fn is Fidian’s variant of Terminal-Bench, built from the same 89 tasks as TB-2.1.
TL;DR
-
TB-fn separates models that look comparable on TB-2.1. The score range widens from ~15 to ~22 points, and the top tier narrows from five labs to OpenAI and Anthropic.
-
Six models lose ~6–11 points. Grok 4.5 loses ~11 points, Kimi K3 ~10, and GLM-5.2 ~9. The other three are Gemini 3.6 Flash (~8), GLM-5.3 (~7), and DeepSeek V4 Pro (~6). These models span a wide spectrum on TB-2.1, from top-tier performers like GLM-5.3 down to lower-ranked options such as Gemini 3.6 Flash.
-
Per-attempt cost rises on TB-fn in every task for every model. Individual increases range from about 12% to 97%. Median turns rise from 16 to 19, with average increases of 45% for input tokens and 32% for output tokens.
-
Claude Opus 5 max effort costs 3x the medium effort with no observed TB-fn gain. All three Opus 5 effort settings score ~82% on TB-fn. The medium effort costs $1 per task, xhigh costs 2x, and max costs 3x.
-
Confirmed security refusals affect Anthropic’s Fable 5 and Opus 5, Google’s Gemini Flash 3.6 and 3.7, and OpenAI’s GPT-5.6 Sol. If Fable 5 were to fall back to Opus 4.8, its performance would improve by approximately 7 points on TB-2.1 and 10 points on TB-fn, reaching roughly 82% on both benchmarks.
-
Open-weight models show reduced performance on TB-fn compared to TB-2.1. While GLM-5.3 maintains its lead within this category, its performance gap versus Sol max expands significantly, and drops out of Tier 1 status.
Scores, rank intervals and tiers across both benchmarks
| | +0.6 | 86.2 85.6 | 10 10 | 0.46M·91% 0.40M·92% | 35.7k 29.6k | 2.2k 1.8k | 6.0 4.0 | $1.52 $1.27 | $1.76 $1.49 |
|---|---|---|---|---|---|---|---|---|---|
| | +1.2 | 82.5 81.3 | 15 13 | 0.89M·94% 0.68M·94% | 64.7k 50.7k | 2.8k 2.3k | 8.1 5.0 | $2.37 $1.84 | $2.87 $2.26 |
| | −2.3 | 82.3 84.6 | 8 8 | 0.41M·92% 0.25M·90% | 22.4k 15.3k | 1.4k 1.2k | 3.9 2.9 | $1.07 $0.73 | $1.30 $0.86 |
| | −0.4 | 81.8 82.2 | 17 16 | 1.06M·94% 0.90M·95% | 92.9k 74.6k | 3.7k 3.1k | 13.2 8.2 | $3.23 $2.59 | $3.95 $3.16 |
| | +0.0 | 81.8 81.8 | 11 10 | 0.64M·96% 0.36M·95% | 22.6k 14.1k | 1.2k 0.9k | 2.2 1.5 | $1.04 $0.65 | $1.27 $0.79 |
| | −4.4 | 78.7 83.1 | 15 13 | 1.07M·94% 0.70M·92% | 43.4k 27.8k | 1.6k 1.2k | 4.0 2.9 | $0.90 $0.60 | $1.15 $0.73 |
| | −7.0 | 77.6 84.6 | 20 14 | 1.25M·96% 0.91M·96% | 67.0k 50.2k | 2.3k 2.1k | 12.7 6.9 | $0.67 $0.50 | $0.86 $0.60 |
| | −1.8 | 75.7 77.5 | 11 11 | 0.91M·95% 0.81M·95% | 80.9k 71.9k | 3.5k 3.4k | 7.5 5.9 | $1.25 $1.12 | $1.66 $1.45 |
| | −0.2 | 75.5 75.7 | 18 17 | 0.93M·92% 0.77M·93% | 116.2k 86.7k | 4.4k 3.5k | 19.5 12.2 | $3.80 $2.85 | $5.04 $3.77 |
| | −2.4 | 74.5 76.9 | 15 14 | 0.72M·87% 0.58M·86% | 80.3k 60.9k | 3.4k 2.8k | 17.1 10.4 | $0.83 $0.65 | $1.11 $0.85 |
| | −10.2 | 72.4 82.6 | 11 9 | 0.94M·90% 0.58M·87% | 26.4k 18.2k | 1.1k 1.0k | 3.0 2.6 | $0.83 $0.57 | $1.14 $0.69 |
| | −4.6 | 72.4 77.0 | 33 30 | 3.54M·86% 2.63M·86% | 43.4k 34.4k | 0.8k 0.8k | 3.5 3.2 | $0.76 $0.58 | $1.05 $0.75 |
| | −3.0 | 72.1 75.1 | 10 10 | 0.36M·91% 0.22M·89% | 27.4k 18.3k | 2.0k 1.4k | 2.7 1.9 | $2.10 $1.41 | $2.92 $1.87 |
| | −3.3 | 71.5 74.8 | 13 12 | 1.70M·97% 1.27M·96% | 70.4k 56.7k | 2.5k 2.2k | 5.1 3.5 | $0.13 $0.11 | $0.19 $0.14 |
| | −2.2 | 69.9 72.1 | 26 19 | 2.49M·96% 1.56M·96% | 174.3k 120.6k | 4.4k 3.9k | 27.4 17.1 | $2.46 $1.67 | $3.53 $2.31 |
| | −4.5 | 69.7 74.2 | 32 27 | 3.80M·88% 2.02M·88% | 34.0k 22.9k | 0.6k 0.5k | 2.6 1.9 | $0.71 $0.40 | $1.02 $0.54 |
| | −6.1 | 68.6 74.7 | 16 14 | 1.92M·84% 1.18M·75% | 97.4k 71.6k | 2.8k 2.6k | 9.6 6.6 | $0.48 $0.37 | $0.70 $0.50 |
| | −8.5 | 67.4 75.9 | 12 12 | 2.08M·82% 1.23M·84% | 65.8k 43.7k | 2.0k 1.8k | 5.8 4.4 | $0.53 $0.28 | $0.79 $0.36 |
| | −10.7 | 66.6 77.3 | 15 12 | 1.54M·95% 0.79M·94% | 27.7k 16.1k | 0.9k 0.7k | 3.2 2.0 | $0.76 $0.42 | $1.13 $0.54 |
| | −5.3 | 66.6 71.9 | 39 19 | 1.13M·35% 0.98M·83% | 53.3k 50.8k | 1.4k 1.9k | 54.7 13.4 | $0.15 $0.07 | $0.22 $0.10 |
| | −5.4 | 66.3 71.7 | 20 18 | 3.01M·61% 2.85M·70% | 129.2k 116.3k | 2.8k 2.7k | 6.3 4.8 | $2.29 $1.87 | $3.45 $2.61 |
| | −4.8 | 65.5 70.3 | 23 20 | 2.06M·90% 1.23M·87% | 130.9k 88.9k | 3.4k 3.0k | 7.2 4.7 | $0.03 $0.02 | $0.05 $0.03 |
| | −7.9 | 64.5 72.4 | 46 40 | 6.25M·87% 4.09M·86% | 51.6k 36.7k | 0.7k 0.6k | 4.0 2.9 | $1.23 $0.84 | $1.91 $1.16 |
GPT-5.6 Sol (max) scores ~86%, while Sol xhigh and all three Opus 5 effort settings score ~82%. OpenAI and Anthropic are the only labs in the tier. Grok 4.6 (~79%), GLM-5.3 (~78%), and Kimi K3 (~72%) fall below that tier; xAI, Z.ai, and Moonshot leave it.
Which models hold and which fall on TB-fn
The cross-model score range widens from ~15 to ~22 points, which naturally leads to more tiers. Each model is evaluated at the same effort setting, through the same serving route, and under the same scoring rule on both benchmarks; only the task set changes.
Six drops exceed four standard errors: Grok 4.5 falls ~11 points, Kimi K3 ~10, GLM-5.2 ~9, Gemini 3.6 Flash ~8, GLM-5.3 ~7, and DeepSeek V4 Pro ~6. These models span ~72% to ~85% on TB-2.1, and the gap between Grok 4.6 and Grok 4.5 roughly doubles from ~6 to ~12 points. Which rewrite causes each fall remains unknown.
TB-fn makes every model work harder
Per-attempt cost rises on TB-fn in every task for every model. This increase occurs because median turns rise from 16 to 19, mean input tokens per attempt rise 45%, and mean output tokens rise 32%. A model whose pass rate holds still moves toward higher cost.
The performance-cost frontier consists of the models for which no cheaper alternative has a higher pass@1. GLM-5.2, Grok 4.5, and GLM-5.3 Flash leave the frontier on TB-fn, while Grok 4.6, Opus 5 medium, and Sol xhigh enter it. Open-weight families occupy part of the low- and mid-cost frontier, while OpenAI and Anthropic occupy its highest-scoring end.
Open-weight families fall further behind on TB-fn
We group GLM, Qwen, Kimi, and DeepSeek as open-weight families. GLM-5.3 is included because the GLM family is open-weight, but its own weights were not public at the time of writing.
GLM-5.3 leads the open-weight families on both benchmarks. On TB-2.1 it scores 84.6%, 1.0 point behind Sol max at 85.6%. On TB-fn it scores 77.6%, 8.6 points behind Sol max at 86.2%, ranks seventh, and is in Tier 2a. No open-weight family reaches Tier 1 on TB-fn. Four of the six models whose scores fall by more than four standard errors are from open-weight families: Kimi K3 (-10.2 points), GLM-5.2 (-8.5), GLM-5.3 (-7.0), and DeepSeek V4 Pro (-6.1).
GLM-5.3 costs $0.67 per attempt on TB-fn versus $1.52 for Sol max, and DeepSeek V4 Flash costs $0.03 per attempt. On the open-weight-family efficiency frontier, TB-2.1 contains DeepSeek V4 Flash, GLM-5.3 Flash, GLM-5.2, and GLM-5.3, while TB-fn contains DeepSeek V4 Flash, GLM-5.3 Flash, DeepSeek V4 Pro, and GLM-5.3.
Claude Opus 5 max costs 3x medium for the same TB-fn score
Opus 5 at medium gives up about two points on a SWE-bench Pro subset for roughly half the cost of default high, according to Anthropic’s current cost guidance. Opus 5 xhigh outperforms max on Anthropic’s official Frontier-Bench v0.1 effort curve.
All three Opus 5 effort settings score ~82% on both benchmarks. The medium setting costs $1.04 per task, xhigh costs $2.37, and max costs $3.23. Medium achieves the same pass rate as max at one-third the cost.
Security refusals by model and effort setting
Security refusals affect Anthropic’s Fable 5 and Opus 5, Google’s Gemini Flash 3.6 and 3.7, and OpenAI’s GPT-5.6 Sol. vulnerable-secret is the one task all five refuse on both benchmarks. The table shows the number of tasks refused at least once and the pass-rate headroom lost to refusal averaged across all 89 tasks.
| | 13 12 | 12.4 12.6 |
|---|---|---|
| | 4 8 | 3.5 7.0 |
| | 3 3 | 3.4 3.4 |
| | 3 3 | 3.4 3.4 |
| | 3 3 | 3.4 3.4 |
| | 2 2 | 1.4 2.2 |
| | 2 1 | 1.1 0.7 |
| | 2 2 | 0.6 1.1 |
To compute the degree of performance degradation, we calculate a hypothetical fallback by routing every refused attempt from Fable and all three Opus 5 effort settings to Opus 4.8, using Opus 4.8’s measured pass rate on the same task. Fable gains 7.4 points on TB-2.1 and 9.7 points on TB-fn, reaching 82.4% and 81.8%. Each Opus 5 effort setting gains 3.4 points on TB-2.1 and 3.2 points on TB-fn, and the settings reach 84.7–85.6% and 84.9–85.7%. Applied simultaneously, Fable ranks ninth on TB-2.1 and sixth on TB-fn, remaining just outside the top tier, while the Opus 5 effort settings and Sol max occupy the top four. Opus 5 max ties Sol max on TB-2.1, and Opus 5 xhigh is 0.4 points behind Sol max on TB-fn. Added fallback cost and latency are not estimated.
TB-fn narrows the top tier from five labs to two
OpenAI and Anthropic are the only labs in the TB-fn top tier. GLM-5.3 leads the open-weight families on both benchmarks, but its gap to Sol max grows from 1.0 point on TB-2.1 to 8.6 points on TB-fn. Six models lose ~6–11 points even though their TB-2.1 scores range from ~72% to ~85%. Per-attempt cost rises by ~12–76% in every matched comparison. Confirmed security refusals affect Anthropic’s Fable 5 and Opus 5, Google’s Gemini Flash 3.6 and 3.7, and OpenAI’s GPT-5.6 Sol. Under a hypothetical fallback to Opus 4.8, Fable 5 reaches ~82% but remains just outside the top tier when Opus 5 is adjusted too, while the Opus 5 effort settings reach ~85–86%, within 1.2 points of Sol max.
Caveats
Statistical resolution. Each benchmark contains about 10k trials. Pass@1 averages 89 task-level pass rates; each cell has at least three valid trials with a median of five. Score intervals resample attempts within tasks without resampling tasks, and rank intervals are the 2.5th to 97.5th rank percentiles. Tiers are descriptive groups, not significance tests. Only Sol max has a rank interval of [1-1] on TB-fn. Bold drops exceed four standard errors.
-
A tmux wedge on large file writes penalizes tasks that require writing big files. A large heredoc write can leave the shell stuck at a continuation prompt, triggering an unrecoverable harness failure that scores as a model failure and most affects file-generation tasks. We run a local mitigation for this issue, which is filed upstream as harbor-framework/harbor#2677.
-
Cost sourcing: Cost is per task attempt. The OpenAI, Anthropic, and Gemini 3.7 rows are quoted as billed. Gemini 3.6 and the OpenRouter-routed rows are repriced from measured token usage at August 2026 provider list prices; GLM-5.2 is quoted as billed on a route charging pre-cut rates and should not be read as a current list price. GLM-5.3 Flash is repriced at Z.ai list rates; a 50% promotion was running at the time of writing, so its billed cost was lower than the figure shown. Among rows with reported costs, only Grok 4.5 omits a pricing tier, specifically the above-200k-prompt-token tier.
-
TB-fn is derived from TB-2.1, not a reproduction of it. The previous post details the construction: we repaired broken verification, closed reward-hacking loopholes, created harder task variants, and made evaluation criteria explicit. Across the benchmark, we rewrote 88 of 89 prompts and all 89 verifiers; the mean number of assertions per task increased from 9.1 to 45.8. Consequently, absolute scores are not comparable to lab-published TB-2.1 scores.
Appendix: harness and reproducibility
Both boards use Terminus 2 (pip install harbor==0.20.0) with two fixes and one settings change.
-
A parse-failure circuit breaker (
max_consecutive_parse_failures) turns the content-filter empty-completion loop into a fast, correctly-labeled failure to prevent timeouts (upstreamed as harbor-framework/harbor#2608). -
A PS2 tmux-wedge recovery clears a shell left stuck at a
>continuation prompt after a large heredoc file-write (filed upstream as harbor-framework/harbor#2677). -
An 1800-second per-request timeout (litellm
timeout) is raised from the 600s default so a >10-minute reasoning turn is not silently killed and retried.
Our harbor changes are public at fidian-ai/harbor, and reproducing the harness requires the published package plus these modifications. TB-fn is withheld to prevent training contamination, meaning third parties cannot reproduce the per-task results until the set is released. We verified our copy of harbor-framework/terminal-bench-2-1 by full diff against upstream commit 7131e437 (2026-08-11): all 89 tasks are byte-identical except one cosmetic fake-secret literal-split in the sanitize-git-repo verifier.
Appendix: how these models are served
We serve models through lab APIs or OpenRouter. We did not control for differences in quantization, parameter translation, tool-call formatting, or system-prompt handling. The serving route is therefore a potential capability confound as well as a determinant of cost and reliability.
| Access route | Models |
|---|---|
| Direct: Anthropic API | Claude Opus 5 (medium/xhigh/max), Opus 4.8, Sonnet 5, Fable 5 |
| Direct: Google Vertex AI | Gemini 3.7 Flash (medium and high), Gemini 3.6 Flash |
| Direct: OpenAI API | GPT-5.6 Sol (max and xhigh), GPT-5.6 Terra, GPT-5.6 Luna |
| Via OpenRouter | Grok 4.5 and Grok 4.6 (xAI), Kimi K3 (Moonshot), DeepSeek V4 Pro 0813 and DeepSeek V4 Flash (DeepSeek), Qwen 3.8 Max (Alibaba), Muse Spark 1.2 (Meta), GLM-5.2, GLM-5.3, and GLM-5.3 Flash (Z.ai) |
Acknowledgements
We thank Hung Le, Ninareh Mehrabi, Reza Pourabolghasem, Ali Parandeh, Ahmad Beirami, Mert Cemri, Melissa Pan, Nino Scherrer, Pooyan Amini, Meisam Raaviyayn, Ziteng Sun, Ananda Theertha Suresh, and Chinnadhurai Sankar for feedback on an earlier version of this blog post.
Citation
If you find this work useful, please cite:
@misc{fidian2026terminalfrontier,
title = {Who is at the frontier of terminal tasks?},
author = {{Fidian} and Chen, Hubert},
year = {2026},
month = {August},
url = {https://fidian.ai/blog/tb-fn-benchmark-results},
}