As of October 2026, the Terminal-Bench 4.0 leaderboard puts Claude Code with Opus 5.5 first at 64.8%, ahead of OpenAI Codex with GPT-6 Astra at 58.2%. Older Terminal-Bench 2.1 is now near saturation (~90% for top agents), so it no longer separates the leaders.
Published agent scores
| Agent | Benchmark | Score | Setup | Date / source |
|---|---|---|---|---|
| Claude Code | Terminal-Bench 4.0 | 64.8% | with Claude Opus 5.5 (max effort) | Oct 2026 · source |
| Claude Code | Terminal-Bench 2.1 | 89.1% | agent-level, Artificial Analysis harness | Aug 2026 · source |
| OpenAI Codex | Terminal-Bench 4.0 | 58.2% | with GPT-6 Astra (max effort) | Oct 2026 · source |
| OpenAI Codex | Terminal-Bench 2.1 | 89.5% | agent-level, Artificial Analysis harness | Aug 2026 · source |
| Google Antigravity | Terminal-Bench 2.1 | 70.7% | agent-level, listed as 'Gemini CLI / Antigravity', Artificial Analysis harness | Aug 2026 · source |
Artificial Analysis runs its own Terminal-Bench 4.0 harness at the model level and ranks differently (Sonnet 5.5 at 63.6%; Opus 5.5 and GPT-6 Astra at 59.6% in October 2026) — a reminder that harness and settings matter. Tools without an agent-level leaderboard entry (e.g. Cursor, Copilot) are not ranked here.
What the major benchmarks measure
SWE-bench (and SWE-bench Verified)
Real GitHub issues from popular Python repositories, graded by the tests that resolved them. The agent must produce a patch that passes. Verified is a human-validated subset that removes ambiguous tasks. Results are mostly reported per model and harness.
Terminal-Bench
Tasks executed entirely in a Linux terminal — installing packages, debugging scripts, configuring systems — testing whether an agent can operate a machine, not just write snippets. Version 4.0 is the current, harder edition; leaderboard entries pair a specific agent with a specific model.
Multi-language, front-end and live boards
Extensions cover more languages, screenshot-judged front-end work and repository-scale tasks. Live leaderboards rerun agents on fresh tasks to fight contamination — a live-board score is worth more than a launch-post screenshot.
How to read a vendor's benchmark claim
✓ Green flags
- Names the exact benchmark and version
- Links to a public leaderboard or reproducible setup
- Reports model, harness, effort and attempt budget
- Shows several benchmarks, not one cherry-pick
✕ Red flags
- “Internal benchmark” with no public definition
- A single percentage with no task count or method
- Competitors run under different conditions
- Scores that never appear on any independent board
What benchmarks don't tell you
How the agent behaves on your codebase, how much supervision it needs, latency, cost per task and whether its diffs match your team's style. Benchmarks measure capability ceilings; your workflow sets the floor.