AI coding model comparison
📅 Data snapshot: October 2026
Using GitHub Copilot? Hit the Copilot filter above to see models listed as available there. Copilot uses per-token AI credit billing; the $/task column estimates a separate direct-API task cost. See GitHub's rate card for Copilot-specific rates.
| Model | Family | Copilot | $/task | SWE-bench | Aider | Arena | LiveBench |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 Max | Anthropic | ✓ | $1.00 | - | - | - | 83.4 |
| Claude Opus 5.5 Thinking Max | Anthropic | ✓ | $0.40 | - | - | - | 83.2 |
| Claude Opus 5 Thinking Max | Anthropic | ✓ | $0.50 | - | - | - | 80.1 |
| Claude Fable 5 Max | Anthropic | ✓ | $1.00 | - | - | - | 83.0 |
| GPT-5.6 Sol Max | OpenAI | ✓ | $0.40 | - | - | - | 81.0 |
| GPT-5.5 Thinking xHigh | OpenAI | ✓ | $0.55 | - | - | - | 80.2 |
| GPT-6 Astra Max | OpenAI | ✓ | $1.00 | - | - | - | 82.2 |
| GPT-6.1 Sol Max | OpenAI | ✓ | $0.20 | - | - | - | 81.6 |
| Muse Spark 1.3 xHigh | Meta | - | - | - | - | - | 81.6 |
| DeepSeek V4.1 Flash Max | DeepSeek | - | $0.03* | - | - | - | 81.1 |
| GPT-6 Sol Max | OpenAI | ✓ | $0.20 | - | - | - | 79.3 |
| GPT-6 Luna Max | OpenAI | ✓ | $0.01 | - | - | - | 72.0 |
| Gemini 3.8 Flash | ✓ | $0.075† | - | - | - | - | |
| Gemini 3.7 Flash | ✓ | $0.075† | - | - | - | 78.8 | |
| Kimi K3 | Moonshot | ✓ | - | - | - | - | 79.2 |
| Grok 4.6 | xAI | ✓ | - | - | - | - | 78.0 |
| Grok 4.7 | xAI | ✓ | - | - | - | - | 77.4 |
| Claude Sonnet 5.5 xHigh | Anthropic | ✓ | $0.20 | - | - | - | 77.8 |
| GLM-5.3 | Zhipu | - | - | - | - | - | 76.1 |
| Qwen 3.8 Flash Next | Alibaba | - | - | - | - | - | 76.2 |
| Qwen 3.8 27B | Alibaba | - | - | - | - | - | 75.3 |
| DeepSeek V4 Pro 0813 | DeepSeek | - | $0.106* | - | - | - | 77.4 |
| Muse Spark 1.2 xHigh | Meta | - | - | - | - | - | 78.0 |
| GPT-5.4 Thinking xHigh | OpenAI | ✓ | $0.28 | - | - | - | 78.0 |
| GPT-5.6 Terra Max | OpenAI | ✓ | $0.22 | - | - | - | 77.9 |
| Gemini 3.1 Pro | - | $0.22 | 69.6% | - | - | 77.0 | |
| Claude Sonnet 5 | Anthropic | ✓ | $0.20 | - | - | - | 76.0 |
| Claude Opus 4.8 | Anthropic | ✓ | $0.50 | - | - | - | 76.2 |
| Claude Opus 4.7 | Anthropic | - | $0.50 | - | - | - | 76.5 |
| Grok 4.5 | xAI | - | - | - | - | - | 75.8 |
| Claude Opus 4.6 | Anthropic | - | $0.50 | 75.6% | - | - | 74.5 |
| Gemini 3.5 Flash | - | $0.17 | - | - | - | 74.6 | |
| GPT-5.2 high reasoning | OpenAI | - | $0.23 | 72.8% | 88.0% | 1470 | 74.6 |
| Gemini 3.6 Flash | - | $0.075† | - | - | - | - | |
| GLM-5.2 | Zhipu | - | - | - | - | - | 73.2 |
| Claude Sonnet 4.6 | Anthropic | ✓ | $0.30 | - | - | - | 73.0 |
| Claude Opus 4.5 thinking-32k | Anthropic | - | $0.50 | 76.8% | 72.0% | 1497 | 72.6 |
| Claude Opus 4.5 | Anthropic | - | $0.50 | 76.8% | 70.7% | 1468 | 72.6 |
| GPT-5.4 nano | OpenAI | ✓ | $0.02 | - | - | - | 69.6 |
| Gemini 3 Flash | - | $0.06 | 75.8% | - | 1443 | - | |
| GPT-5.4 mini | OpenAI | ✓ | $0.08 | - | - | - | 66.4 |
| Gemini 3.5 Flash-Lite | - | $0.04 | - | - | - | 63.9 | |
| Minimax M3 | Minimax | - | - | - | - | - | 67.3 |
| Kimi K2.5 | Moonshot | - | $0.15 | 70.8% | - | - | - |
| Minimax M2.5 | Minimax | - | $0.07 | 75.8% | - | - | - |
| Gemini 2.5 Pro | - | $0.16 | 53.6% | 83.1% | 1372 | 58.3 | |
| GLM-4.7 | Zhipu | - | $0.05 | - | - | 1440 | - |
| Claude Sonnet 4.5 | Anthropic | - | $0.30 | 71.4% | 82.4% | 1383 | - |
| GPT-5.2 | OpenAI | - | $0.23 | 72.8% | 88.0% | 1432 | 48.9 |
| Claude Haiku 4.5 | Anthropic | ✓ | $0.10 | 66.6% | 73.5% | 1290 | 45.3 |
| GPT-4o | OpenAI | - | $0.23 | 48.9% | 72.9% | 1372 | - |
| Gemini 2.5 Flash | - | $0.04 | 28.7% | 55.1% | 1233 | 47.7 | |
| GPT-5 mini | OpenAI | ✓ | $0.03 | 56.2% | 50.2% | 1145 | - |
| GPT-4.1 | OpenAI | - | $0.18 | 39.6% | 52.4% | 1305 | - |
Column guide
| Column | What it means |
|---|---|
| Copilot | Available in GitHub Copilot (✓ = listed, - = not listed). Copilot usage is billed in AI credits based on tokens used. |
| $/task | Estimated direct-API cost for 50K input + 10K output tokens; it is not a Copilot bill estimate. |
| SWE-bench | % of real GitHub issues the model can fix autonomously (source) - February 2026 data (standardized harness, high reasoning mode) |
| Aider | % correct on multi-language code editing (source) - latest listed runs are October 2025; historical comparison only |
| Arena | Elo rating from human preference voting on Code category (source) - February 2026 data |
| LiveBench | Global average score across 23 diverse tasks (source) - current scores shown from release 2026-06-25 |
DeepSeek V4.1 Flash has time-based peak/off-peak rates; this task estimate uses peak pricing. †Gemini 3.7/3.8 Flash promotional API pricing runs through Dec 31, 2026.
Data sources:
SWE-bench (Feb 2026) ·
Aider (latest listed runs: Oct 2025) ·
Arena Code (Feb 2026, not refreshed) ·
LiveBench (latest release: v2026-06-25; leaderboard includes newer models) ·
ProgramBench (leaderboard updated Sep 28, 2026) ·
GitHub Copilot (AI credit billing since Jun 2026)
API pricing: Anthropic · OpenAI · Google · DeepSeek · Zhipu (GLM)
API pricing: Anthropic · OpenAI · Google · DeepSeek · Zhipu (GLM)