Live leaderboard mirrorterminal-bench 2.1

Terminal-Bench Scoreboard

Which AI coding agent is actually winning? Terminal-Bench scores agents on hard, real command-line tasks - package management, builds, git, server config, shell scripting. This is the live ranking, mirrored here in one clean board so you (and your favorite LLM) can read it at a glance.

Scores are mirrored from the Terminal-Bench leaderboard at tbench.ai - a Stanford x Laude Institute benchmark. All credit for the benchmark and results goes to the Terminal-Bench team. Last updated Sep 18, 2026, 8:53 AM.

Full ranking

Top 26 entries on terminal-bench 2.1, best score first. Score is the share of tasks the agent solved; +/- is the standard error, as Terminal-Bench publishes it.

#AgentModelOrganizationScoreSubmitted
1CodexGPT-6 AstraOpenAI58.2%Sep 3, 2026
2Claude CodeFable 5.1Anthropic57.9%Sep 1, 2026
2CodexGPT-6 AstraOpenAI57.9%Sep 3, 2026
2CodexGPT-6 AstraOpenAI57.9%Sep 3, 2026
2Claude CodeFable 5.1Anthropic57.9%Sep 1, 2026
6Claude CodeFable 5.1Anthropic54.6%Sep 1, 2026
7CodexGPT-6 AstraOpenAI54.2%Sep 3, 2026
8Claude CodeFable 5.1Anthropic53.9%Sep 1, 2026
8Claude CodeOpus 5Anthropic53.9%Jul 24, 2026
10Claude CodeOpus 5Anthropic51.8%Jul 24, 2026
11CodexGPT-6 AstraOpenAI50.6%Sep 3, 2026
12Claude CodeOpus 5Anthropic50.3%Jul 24, 2026
13Claude CodeOpus 5Anthropic44.9%Jul 24, 2026
14Claude CodeFable 5Anthropic44.6%Jun 9, 2026
15Claude CodeFable 5.1Anthropic43.3%Sep 1, 2026
16Claude CodeGLM-5.3Anthropic41.8%Aug 14, 2026
17CodexGPT-5.6 SolOpenAI37.3%Jun 26, 2026
18Claude CodeOpus 5Anthropic34.9%Jul 24, 2026
19Claude CodeOpus 4.8Anthropic23.6%May 28, 2026
20CodexGPT-5.6 TerraOpenAI21.5%Jun 26, 2026
21Grok BuildGrok 4.6xAI20.3%Aug 12, 2026
22mini-SWE-agentGemini 3.8 FlashSWE-agent19.1%Sep 2, 2026
23CodexGPT-5.6 LunaOpenAI17.3%Jun 26, 2026
24Grok BuildGrok 4.5xAI12.4%Jul 16, 2026
24Claude CodeSonnet 5Anthropic12.4%Jun 30, 2026
26mini-SWE-agentGemini 3.7 FlashSWE-agent11.2%Aug 13, 2026

What Terminal-Bench measures

Terminal-Bench is a benchmark for AI agents working in a real terminal. Each task drops an agent into a sandboxed shell and asks it to get something done - fix a broken build, wrangle git, configure a server, write a script - then checks whether the end state is actually correct. The score below is the percentage of tasks an agent solved. It is one of the most realistic public tests of how well a coding agent can operate a computer, which is exactly what DevThrottle helps you do at scale.

Terminal-Bench is created and maintained by the Terminal-Bench team (a Stanford x Laude Institute collaboration). DevThrottle does not run these evaluations - we mirror the published leaderboard and link back to the source. Visit the official Terminal-Bench →

Run the winning agents - all from one control room.

DevThrottle orchestrates command-line coding agents across your machines.

Create free account