Coding Model Rankings 2026
As of 2026-09-24. Terminal-Bench 4.0: GPT-6 Astra (max) with Codex resolves 58.2%. mini-SWE-agent 2.0.0 on SWE-bench Verified is a second table.

Compiled 2026-09-24. The main table copies the official Terminal-Bench 4.0 board; the board record was updated on 21 September 2026. Each row is one submission: model, reasoning effort and agent are scored together. GPT-6 Astra (max) with Codex resolves 58.2% of tasks; the 95% half-width is 2.8 points. Fable 5.1 (max) with Claude Code resolves 57.9%, half-width 3.8 points, and the two intervals overlap. Coding products: AI coding assistants. Crowd votes: LLM capability. List prices: LLM API pricing.

Terminal-Bench 4.0 official board
These are all 27 rows on that board. Tied ranks stay as printed. Cost is the bill for that evaluation run, not a per-million-token list price. GLM-5.3 (max) with Claude Code is 16th at 41.8%. Grok 4.7 (xhigh) with Grok Build was submitted on 21 September 2026 at 37.6%.
| # | 模型 | 档位 | 代理 | 机构 | 解决率 | 95% 半宽 | 花费 | Token | 提交日 |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | max | Codex | OpenAI | 58.2% | ±2.8 | $3,267 | 1.5B | Sep 3, 2026 |
| 2 | Fable 5.1 | max | Claude Code | Anthropic | 57.9% | ±3.8 | $6,244 | 2.7B | Sep 1, 2026 |
| 2 | GPT-6 Astra | xhigh | Codex | OpenAI | 57.9% | ±2.7 | $2,351 | 1.2B | Sep 3, 2026 |
| 2 | GPT-6 Astra | high | Codex | OpenAI | 57.9% | ±3.0 | $2,269 | 1.2B | Sep 3, 2026 |
| 2 | Fable 5.1 | xhigh | Claude Code | Anthropic | 57.9% | ±3.4 | $4,872 | 2.3B | Sep 1, 2026 |
| 6 | Fable 5.1 | high | Claude Code | Anthropic | 54.5% | ±3.4 | $3,985 | 2.2B | Sep 1, 2026 |
| 7 | GPT-6 Astra | medium | Codex | OpenAI | 54.2% | ±2.7 | $1,915 | 1.1B | Sep 3, 2026 |
| 8 | Fable 5.1 | medium | Claude Code | Anthropic | 53.9% | ±3.4 | $2,833 | 1.6B | Sep 1, 2026 |
| 8 | Opus 5 | xhigh | Claude Code | Anthropic | 53.9% | ±3.2 | $6,086 | 6.9B | Jul 24, 2026 |
| 10 | Opus 5 | max | Claude Code | Anthropic | 51.8% | ±3.4 | $5,969 | 6.5B | Jul 24, 2026 |
| 11 | GPT-6 Astra | low | Codex | OpenAI | 50.6% | ±2.8 | $1,557 | 889.8M | Sep 3, 2026 |
| 12 | Opus 5 | high | Claude Code | Anthropic | 50.3% | ±3.7 | $4,662 | 5.5B | Jul 24, 2026 |
| 13 | Opus 5 | medium | Claude Code | Anthropic | 44.9% | ±3.8 | $3,192 | 3.7B | Jul 24, 2026 |
| 14 | Fable 5 | max | Claude Code | Anthropic | 44.5% | ±3.9 | $7,265 | 3.8B | Jun 9, 2026 |
| 15 | Fable 5.1 | low | Claude Code | Anthropic | 43.3% | ±3.6 | $2,359 | 1.3B | Sep 1, 2026 |
| 16 | GLM-5.3 | max | Claude Code | Z.ai | 41.8% | ±3.2 | $2,728 | 8.7B | Aug 14, 2026 |
| 17 | Grok 4.7 | xhigh | Grok Build | xAI | 37.6% | ±3.5 | $3,683 | 5.5B | Sep 21, 2026 |
| 18 | GPT-5.6 Sol | max | Codex | OpenAI | 37.3% | ±3.8 | $2,542 | 4.4B | Jun 26, 2026 |
| 19 | Opus 5 | low | Claude Code | Anthropic | 34.9% | ±3.9 | $2,394 | 2.7B | Jul 24, 2026 |
| 20 | Opus 4.8 | max | Claude Code | Anthropic | 23.6% | ±3.6 | $6,481 | 6.4B | May 28, 2026 |
| 21 | GPT-5.6 Terra | max | Codex | OpenAI | 21.5% | ±3.2 | $1,734 | 5.7B | Jun 26, 2026 |
| 22 | Grok 4.6 | high | Grok Build | xAI | 20.3% | ±3.1 | $3,592 | 4.0B | Aug 12, 2026 |
| 23 | Gemini 3.8 Flash | high | mini-SWE-agent | 19.1% | ±3.4 | $1,829 | 17.2B | Sep 2, 2026 | |
| 24 | GPT-5.6 Luna | max | Codex | OpenAI | 17.3% | ±2.9 | $347 | 11.6B | Jun 26, 2026 |
| 25 | Grok 4.5 | high | Grok Build | xAI | 12.4% | ±2.6 | $2,094 | 3.4B | Jul 16, 2026 |
| 25 | Sonnet 5 | max | Claude Code | Anthropic | 12.4% | ±3.1 | $9,604 | 21.6B | Jun 30, 2026 |
| 27 | Gemini 3.7 Flash | high | mini-SWE-agent | 11.2% | ±2.5 | $1,262 | 11.1B | Aug 13, 2026 |
SWE-bench Verified, one agent: mini-SWE-agent 2.0.0
The public SWE-bench board stacks different agents. The table below keeps bash-only submissions on mini-SWE-agent 2.0.0 and sorts them by resolve rate. Claude 4.5 Opus (high) is at 76.80%, average cost $0.75, submitted 17 February 2026. MiniMax M2.5 (high) and Gemini 3 Flash (high) are both at 75.80%. Kimi K2.5 (high) is at 70.80%, average $0.15. DeepSeek V3.2 (high) is at 70.00%.
| # | 模型 | 档位 | 解决率 | 平均花费 | 提交日 |
|---|---|---|---|---|---|
| 1 | Claude 4.5 Opus | high | 76.80% | $0.75 | 2026-02-17 |
| 2 | Gemini 3 Flash | high | 75.80% | $0.36 | 2026-02-17 |
| 3 | MiniMax M2.5 | high | 75.80% | $0.07 | 2026-02-17 |
| 4 | Claude 4.6 Opus | — | 75.60% | $0.55 | 2026-02-17 |
| 5 | GLM 5 | high | 72.80% | $0.53 | 2026-02-17 |
| 6 | GPT 5.2 | high | 72.80% | $0.47 | 2026-02-17 |
| 7 | GPT 5.2 Codex | — | 72.80% | $0.45 | 2026-02-19 |
| 8 | Claude 4.5 Sonnet | high | 71.40% | $0.66 | 2026-02-17 |
| 9 | Kimi K2.5 | high | 70.80% | $0.15 | 2026-02-17 |
| 10 | DeepSeek V3.2 | high | 70.00% | $0.45 | 2026-02-17 |
| 11 | Gemini 3 Pro | high | 69.60% | $0.96 | 2026-02-26 |
| 12 | GPT 5 mini | — | 56.20% | $0.05 | 2026-02-17 |
LiveCodeBench default window (through 1 May 2025)
The default board lists 454 problems dated 1 August 2024 through 1 May 2025. o4-mini (high) has Pass@1 80.2, Easy 99.1, Hard 63.5. DeepSeek-R1-0528 is 5th at 73.1. That generation stops at the window’s end date. Astra, Fable 5.1 and Opus 5, released later, are absent.
| # | 模型 | 档位 | Pass@1 | Easy | Medium | Hard |
|---|---|---|---|---|---|---|
| 1 | o4-mini | high | 80.2 | 99.1 | 89.4 | 63.5 |
| 2 | o3 | high | 75.8 | 99.1 | 84.4 | 57.1 |
| 3 | o4-mini | medium | 74.2 | 98.2 | 86.5 | 52.7 |
| 4 | Gemini 2.5 Pro | 06-05 | 73.6 | 99.1 | 87.2 | 50.2 |
| 5 | DeepSeek-R1-0528 | — | 73.1 | 98.7 | 85.2 | 50.7 |
| 6 | Gemini 2.5 Pro | 05-06 | 71.8 | 98.2 | 82.3 | 50.2 |
| 7 | EXAONE 4.0 32B | — | 70 | 98.4 | 82.3 | 46.2 |
| 8 | OpenReasoning-Nemotron-32B | — | 69.8 | 98.3 | 81.4 | 46.3 |
How this list is ranked
Compiled 2026-09-24. The main table is the public board titled Terminal-Bench 4.0 on tbench.ai. The board record was updated on 21 September 2026. All 27 rows are included. Resolve rate, 95% half-width, cost, tokens and submission date are copied from that board’s JSON. Cost is total_cost_usd rounded to the dollar. The second table keeps only SWE-bench Verified rows that use bash-only and mini-SWE-agent 2.0.0, sorted by resolve rate. Those ranks are internal to the 12 rows. The third table is the LiveCodeBench default window, problems dated 1 August 2024 through 1 May 2025 (454 problems). That generation stops in May 2025 and is not folded into the main order.
| Source | Snapshot | What it measures | Weight |
|---|---|---|---|
| Terminal-Bench 4.0 official board | Fetched 2026-09-24; board updated_at 21 Sep 2026 | Terminal-task resolve rate. One row is model + reasoning effort + agent. GPT-6 Astra (max) + Codex 58.2% (±2.8), 330 trials. All 27 rows | Main |
| SWE-bench Verified (bash-only) | Fetched 2026-09-24; row dates 17–26 Feb 2026 | mini-SWE-agent 2.0.0 only. Claude 4.5 Opus (high) 76.80%, average $0.75. Other agents on the public board are left out | Aux |
| LiveCodeBench | Default window 1 Aug 2024–1 May 2025; fetched 2026-09-24 | Contest-style Pass@1. o4-mini (high) 80.2. Models released after the window are absent from this default board | Aux (older window) |
互动体验:拖拽排位《Coding Models 2026》
Coding Models 2026
想打造自己的个性化天梯图并开放给读者嵌入吗?