AI coding agent benchmarks, explained

Vendors love quoting scores. Here is what the major benchmarks actually measure, the latest independent agent-level results we track, and how to read claims like a skeptic.

As of October 2026, the Terminal-Bench 4.0 leaderboard puts Claude Code with Opus 5.5 first at 64.8%, ahead of OpenAI Codex with GPT-6 Astra at 58.2%. Older Terminal-Bench 2.1 is now near saturation (~90% for top agents), so it no longer separates the leaders.

Published agent scores

AgentBenchmarkScoreSetupDate / source
Claude CodeTerminal-Bench 4.064.8%with Claude Opus 5.5 (max effort)Oct 2026 · source
Claude CodeTerminal-Bench 2.189.1%agent-level, Artificial Analysis harnessAug 2026 · source
OpenAI CodexTerminal-Bench 4.058.2%with GPT-6 Astra (max effort)Oct 2026 · source
OpenAI CodexTerminal-Bench 2.189.5%agent-level, Artificial Analysis harnessAug 2026 · source
Google AntigravityTerminal-Bench 2.170.7%agent-level, listed as 'Gemini CLI / Antigravity', Artificial Analysis harnessAug 2026 · source

Artificial Analysis runs its own Terminal-Bench 4.0 harness at the model level and ranks differently (Sonnet 5.5 at 63.6%; Opus 5.5 and GPT-6 Astra at 59.6% in October 2026) — a reminder that harness and settings matter. Tools without an agent-level leaderboard entry (e.g. Cursor, Copilot) are not ranked here.

What the major benchmarks measure

SWE-bench (and SWE-bench Verified)

Real GitHub issues from popular Python repositories, graded by the tests that resolved them. The agent must produce a patch that passes. Verified is a human-validated subset that removes ambiguous tasks. Results are mostly reported per model and harness.

Terminal-Bench

Tasks executed entirely in a Linux terminal — installing packages, debugging scripts, configuring systems — testing whether an agent can operate a machine, not just write snippets. Version 4.0 is the current, harder edition; leaderboard entries pair a specific agent with a specific model.

Multi-language, front-end and live boards

Extensions cover more languages, screenshot-judged front-end work and repository-scale tasks. Live leaderboards rerun agents on fresh tasks to fight contamination — a live-board score is worth more than a launch-post screenshot.

How to read a vendor's benchmark claim

✓ Green flags

  • Names the exact benchmark and version
  • Links to a public leaderboard or reproducible setup
  • Reports model, harness, effort and attempt budget
  • Shows several benchmarks, not one cherry-pick

✕ Red flags

  • “Internal benchmark” with no public definition
  • A single percentage with no task count or method
  • Competitors run under different conditions
  • Scores that never appear on any independent board

What benchmarks don't tell you

How the agent behaves on your codebase, how much supervision it needs, latency, cost per task and whether its diffs match your team's style. Benchmarks measure capability ceilings; your workflow sets the floor.

FAQ

Which AI coding agent scores highest on Terminal-Bench?
On the Terminal-Bench 4.0 leaderboard (tbench.ai) as of October 2026, Claude Code with Claude Opus 5.5 at max effort ranks first at 64.8%, and OpenAI Codex with GPT-6 Astra scores 58.2%.
Is SWE-bench a tool benchmark or a model benchmark?
Mostly a model benchmark: SWE-bench Verified results are usually published per model and harness, not per product, so they don't directly tell you how Cursor or Copilot perform.
Should I choose an agent based on benchmarks?
Use benchmarks to shortlist, not to decide. Run two or three candidates on one real task from your own backlog and compare the diffs, time and cost.