📅 Data snapshot: August 2026

Using GitHub Copilot? Hit the Copilot filter above to see only models available in Copilot. Since June 2026, Copilot uses per-token AI credit billing — the $/task column directly reflects your cost.
Model Family Copilot $/task SWE-bench Aider Arena LiveBench
Claude Fable 5Anthropic$1.00---83.0
GPT-5.6 SolOpenAI$0.55---81.0
GPT-5.5OpenAI$0.55---80.2
Claude Opus 5Anthropic$0.50---80.1
Kimi K3Moonshot-----79.2
Qwen 3.8 MaxAlibaba-----78.5
GPT-5.4OpenAI$0.28---78.0
GPT-5.6 TerraOpenAI$0.22---77.9
Gemini 3.1 ProGoogle$0.2269.6%--77.0
Claude Sonnet 5Anthropic$0.20†---76.0
Claude Opus 4.8Anthropic$0.50---76.2
Claude Opus 4.7Anthropic$0.50---76.5
Grok 4.5xAI-----75.8
Claude Opus 4.6Anthropic$0.5075.6%--74.5
Gemini 3.5 FlashGoogle$0.17---74.6
GPT-5.2 high reasoningOpenAI$0.2372.8%88.0%147074.6
DeepSeek V4 Flash (Jul 31)DeepSeek-$0.0170.0%74.2%135074.2
Gemini 3.6 FlashGoogle$0.15---73.6
GPT-5.6 LunaOpenAI$0.02---73.6
GLM-5.2Zhipu-----73.2
Claude Sonnet 4.6Anthropic$0.30---73.0
Claude Opus 4.5 thinking-32kAnthropic$0.5076.8%72.0%149772.6
Claude Opus 4.5Anthropic$0.5076.8%70.7%146872.6
DeepSeek V4 ProDeepSeek-$0.03---71.6
GPT-5.4 nanoOpenAI$0.02---69.6
Gemini 3 FlashGoogle$0.0675.8%-1443-
GPT-5.4 miniOpenAI$0.08---66.4
Gemini 3.5 Flash-LiteGoogle-$0.04---63.9
Minimax M3Minimax-----67.3
Kimi K2.5Moonshot-$0.1570.8%---
Minimax M2.5Minimax-$0.0775.8%---
Gemini 2.5 ProGoogle$0.1653.6%83.1%137258.3
GLM-4.7Zhipu-$0.05--1440-
Claude Sonnet 4.5Anthropic$0.3071.4%82.4%1383-
GPT-5.2OpenAI$0.2372.8%88.0%143248.9
Claude Haiku 4.5Anthropic$0.1066.6%73.5%129045.3
GPT-4oOpenAI-$0.2348.9%72.9%1372-
Gemini 2.5 FlashGoogle-$0.0428.7%55.1%123347.7
GPT-5 miniOpenAI$0.0356.2%50.2%1145-
GPT-4.1OpenAI-$0.1839.6%52.4%1305-

Column guide

Column What it means
Copilot Available in GitHub Copilot (✓ = yes, - = not listed). Since June 1, 2026 Copilot uses token-based AI credit billing — same per-token rates as direct API.
$/task Estimated cost per task if using APIs directly (50K in + 10K out tokens). Also reflects Copilot AI credit cost since Jun 2026 billing change.
SWE-bench % of real GitHub issues the model can fix autonomously (source) - February 2026 data (standardized harness, high reasoning mode)
Aider % correct on multi-language code editing (source) - June 2025 data · Best signal for CLI/agentic use cases
Arena Elo rating from human preference voting on Code category (source) - February 2026 data
LiveBench Global average score across 23 diverse tasks (source) - Aug 2026 data (LiveBench-2026-06-25), contamination-free

†Claude Sonnet 5 intro pricing through Aug 31, 2026; rises to $0.30/task from Sep 1.

Data sources: SWE-bench (Feb 2026) · Aider (Oct 2025) · Arena Code (Feb 2026, not refreshed) · LiveBench (Aug 2026, v2026-06-25) · GitHub Copilot (AI credit billing since Jun 2026)
API pricing: Anthropic · OpenAI · Google · DeepSeek · Zhipu (GLM)