AI coding model comparison
📅 Data snapshot: August 2026
Using GitHub Copilot? Hit the Copilot filter above to see only models available in Copilot. Since June 2026, Copilot uses per-token AI credit billing — the $/task column directly reflects your cost.
| Model | Family | Copilot | $/task | SWE-bench | Aider | Arena | LiveBench |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | ✓ | $1.00 | - | - | - | 83.0 |
| GPT-5.6 Sol | OpenAI | ✓ | $0.55 | - | - | - | 81.0 |
| GPT-5.5 | OpenAI | ✓ | $0.55 | - | - | - | 80.2 |
| Claude Opus 5 | Anthropic | ✓ | $0.50 | - | - | - | 80.1 |
| Kimi K3 | Moonshot | - | - | - | - | - | 79.2 |
| Qwen 3.8 Max | Alibaba | - | - | - | - | - | 78.5 |
| GPT-5.4 | OpenAI | ✓ | $0.28 | - | - | - | 78.0 |
| GPT-5.6 Terra | OpenAI | ✓ | $0.22 | - | - | - | 77.9 |
| Gemini 3.1 Pro | ✓ | $0.22 | 69.6% | - | - | 77.0 | |
| Claude Sonnet 5 | Anthropic | ✓ | $0.20† | - | - | - | 76.0 |
| Claude Opus 4.8 | Anthropic | ✓ | $0.50 | - | - | - | 76.2 |
| Claude Opus 4.7 | Anthropic | ✓ | $0.50 | - | - | - | 76.5 |
| Grok 4.5 | xAI | - | - | - | - | - | 75.8 |
| Claude Opus 4.6 | Anthropic | ✓ | $0.50 | 75.6% | - | - | 74.5 |
| Gemini 3.5 Flash | ✓ | $0.17 | - | - | - | 74.6 | |
| GPT-5.2 high reasoning | OpenAI | ✓ | $0.23 | 72.8% | 88.0% | 1470 | 74.6 |
| DeepSeek V4 Flash (Jul 31) | DeepSeek | - | $0.01 | 70.0% | 74.2% | 1350 | 74.2 |
| Gemini 3.6 Flash | ✓ | $0.15 | - | - | - | 73.6 | |
| GPT-5.6 Luna | OpenAI | ✓ | $0.02 | - | - | - | 73.6 |
| GLM-5.2 | Zhipu | - | - | - | - | - | 73.2 |
| Claude Sonnet 4.6 | Anthropic | ✓ | $0.30 | - | - | - | 73.0 |
| Claude Opus 4.5 thinking-32k | Anthropic | ✓ | $0.50 | 76.8% | 72.0% | 1497 | 72.6 |
| Claude Opus 4.5 | Anthropic | ✓ | $0.50 | 76.8% | 70.7% | 1468 | 72.6 |
| DeepSeek V4 Pro | DeepSeek | - | $0.03 | - | - | - | 71.6 |
| GPT-5.4 nano | OpenAI | ✓ | $0.02 | - | - | - | 69.6 |
| Gemini 3 Flash | ✓ | $0.06 | 75.8% | - | 1443 | - | |
| GPT-5.4 mini | OpenAI | ✓ | $0.08 | - | - | - | 66.4 |
| Gemini 3.5 Flash-Lite | - | $0.04 | - | - | - | 63.9 | |
| Minimax M3 | Minimax | - | - | - | - | - | 67.3 |
| Kimi K2.5 | Moonshot | - | $0.15 | 70.8% | - | - | - |
| Minimax M2.5 | Minimax | - | $0.07 | 75.8% | - | - | - |
| Gemini 2.5 Pro | ✓ | $0.16 | 53.6% | 83.1% | 1372 | 58.3 | |
| GLM-4.7 | Zhipu | - | $0.05 | - | - | 1440 | - |
| Claude Sonnet 4.5 | Anthropic | ✓ | $0.30 | 71.4% | 82.4% | 1383 | - |
| GPT-5.2 | OpenAI | ✓ | $0.23 | 72.8% | 88.0% | 1432 | 48.9 |
| Claude Haiku 4.5 | Anthropic | ✓ | $0.10 | 66.6% | 73.5% | 1290 | 45.3 |
| GPT-4o | OpenAI | - | $0.23 | 48.9% | 72.9% | 1372 | - |
| Gemini 2.5 Flash | - | $0.04 | 28.7% | 55.1% | 1233 | 47.7 | |
| GPT-5 mini | OpenAI | ✓ | $0.03 | 56.2% | 50.2% | 1145 | - |
| GPT-4.1 | OpenAI | - | $0.18 | 39.6% | 52.4% | 1305 | - |
Column guide
| Column | What it means |
|---|---|
| Copilot | Available in GitHub Copilot (✓ = yes, - = not listed). Since June 1, 2026 Copilot uses token-based AI credit billing — same per-token rates as direct API. |
| $/task | Estimated cost per task if using APIs directly (50K in + 10K out tokens). Also reflects Copilot AI credit cost since Jun 2026 billing change. |
| SWE-bench | % of real GitHub issues the model can fix autonomously (source) - February 2026 data (standardized harness, high reasoning mode) |
| Aider | % correct on multi-language code editing (source) - June 2025 data · Best signal for CLI/agentic use cases |
| Arena | Elo rating from human preference voting on Code category (source) - February 2026 data |
| LiveBench | Global average score across 23 diverse tasks (source) - Aug 2026 data (LiveBench-2026-06-25), contamination-free |
†Claude Sonnet 5 intro pricing through Aug 31, 2026; rises to $0.30/task from Sep 1.
Data sources:
SWE-bench (Feb 2026) ·
Aider (Oct 2025) ·
Arena Code (Feb 2026, not refreshed) ·
LiveBench (Aug 2026, v2026-06-25) ·
GitHub Copilot (AI credit billing since Jun 2026)
API pricing: Anthropic · OpenAI · Google · DeepSeek · Zhipu (GLM)
API pricing: Anthropic · OpenAI · Google · DeepSeek · Zhipu (GLM)