On trackcapabilitiesA superhuman coder exists by end of 2027
On track · due Dec 2027 On track
3 receiptsOpen verificationClose▾
SWE-bench Verified %79.20 % resolved on SWE-bench Verified (top agent, bash-only default view) (as of 2026-08-15)
Verified blind: GLM-5.2 + Grok — agree
Met when a model autonomously completes a non-trivial PR end-to-end at senior-eng level
79.20% resolved on SWE-bench Verified (top agent, bash-only default view)as of 2026-08-15
A single reading is not yet a trajectory. The figure appears once a second measurement lands.
Table view
| As of | Measured (% resolved on SWE-bench Verified (top agent, bash-only default view)) | Receipt |
|---|---|---|
| 2026-08-15 | 79.20 | source |
- Jun 2025Mid-2025: stumbling agents — first usable AI coding agents appear hitCoding agents emerged in 2025 and now autonomously resolve real GitHub issues on SWE-bench Verified.assessed Jun 2026indication↗
- Jan 2026Early 2026: coding automation accelerates on paceTop agents now exceed 85% on SWE-bench Verified, up from ~70% a year earlier — fast progress, still short of autonomous senior-eng PRs.assessed Jun 2026indication↗
- Dec 2027End 2027: a superhuman coder exists (the target) pendingThe claim's target milestone — not yet due.indication↗
Why we forecast a superhuman coder in early 2027
AI 2027 (Kokotajlo, Alexander, Larsen, Lifland & Dean) archived↗context · not model-verified · 2025-04
Claude 4.5 Opus (high reasoning) | 76.80
SWE-bench Verified leaderboard archived↗ · 2026-07-15
best AI agents are not currently able to carry out substantive projects by themselves or directly substitute for human labor.
METR — Measuring AI Ability to Complete Long Tasks archived↗ · 2025-03-19
Ledger history
- 2026-06-20: seeded (#1)
- 2026-06-24: replaced placeholder evidence with the SWE-bench Verified leaderboard; verdict held on-track
- 2026-06-25: demoted pending bulletproof re-grade
- 2026-06-26: added trajectory checkpoints (AI-2027 milestones, on track) + refreshed SWE-bench measurement to >85%
- 2026-07-15: Target date is 2027-12 (future). SWE-bench Verified shows top agents above 76% (swebench.com, measurement, SWE-bench), while METR (measurement, different org) confirms AI agents still cannot…
- 2026-08-15: Verdict on-track remains correct (target 2027-12 is in the future), but a fresh read of the SWE-bench Verified leaderboard (fetched 2026-08-15) shows the top bash-only entry at 79.20% resolved…