📅 Data snapshot: October 2026

Using GitHub Copilot? Hit the Copilot filter above to see models listed as available there. Copilot uses per-token AI credit billing; the $/task column estimates a separate direct-API task cost. See GitHub's rate card for Copilot-specific rates.
Model Family Copilot $/task SWE-bench Aider Arena LiveBench
Claude Fable 5.1 MaxAnthropic✓$1.00---83.4
Claude Opus 5.5 Thinking MaxAnthropic✓$0.40---83.2
Claude Opus 5 Thinking MaxAnthropic✓$0.50---80.1
Claude Fable 5 MaxAnthropic✓$1.00---83.0
GPT-5.6 Sol MaxOpenAI✓$0.40---81.0
GPT-5.5 Thinking xHighOpenAI✓$0.55---80.2
GPT-6 Astra MaxOpenAI✓$1.00---82.2
GPT-6.1 Sol MaxOpenAI✓$0.20---81.6
Muse Spark 1.3 xHighMeta-----81.6
DeepSeek V4.1 Flash MaxDeepSeek-$0.03*---81.1
GPT-6 Sol MaxOpenAI✓$0.20---79.3
GPT-6 Luna MaxOpenAI✓$0.01---72.0
Gemini 3.8 FlashGoogle✓$0.075†----
Gemini 3.7 FlashGoogle✓$0.075†---78.8
Kimi K3Moonshot✓----79.2
Grok 4.6xAI✓----78.0
Grok 4.7xAI✓----77.4
Claude Sonnet 5.5 xHighAnthropic✓$0.20---77.8
GLM-5.3Zhipu-----76.1
Qwen 3.8 Flash NextAlibaba-----76.2
Qwen 3.8 27BAlibaba-----75.3
DeepSeek V4 Pro 0813DeepSeek-$0.106*---77.4
Muse Spark 1.2 xHighMeta-----78.0
GPT-5.4 Thinking xHighOpenAI✓$0.28---78.0
GPT-5.6 Terra MaxOpenAI✓$0.22---77.9
Gemini 3.1 ProGoogle-$0.2269.6%--77.0
Claude Sonnet 5Anthropic✓$0.20---76.0
Claude Opus 4.8Anthropic✓$0.50---76.2
Claude Opus 4.7Anthropic-$0.50---76.5
Grok 4.5xAI-----75.8
Claude Opus 4.6Anthropic-$0.5075.6%--74.5
Gemini 3.5 FlashGoogle-$0.17---74.6
GPT-5.2 high reasoningOpenAI-$0.2372.8%88.0%147074.6
Gemini 3.6 FlashGoogle-$0.075†----
GLM-5.2Zhipu-----73.2
Claude Sonnet 4.6Anthropic✓$0.30---73.0
Claude Opus 4.5 thinking-32kAnthropic-$0.5076.8%72.0%149772.6
Claude Opus 4.5Anthropic-$0.5076.8%70.7%146872.6
GPT-5.4 nanoOpenAI✓$0.02---69.6
Gemini 3 FlashGoogle-$0.0675.8%-1443-
GPT-5.4 miniOpenAI✓$0.08---66.4
Gemini 3.5 Flash-LiteGoogle-$0.04---63.9
Minimax M3Minimax-----67.3
Kimi K2.5Moonshot-$0.1570.8%---
Minimax M2.5Minimax-$0.0775.8%---
Gemini 2.5 ProGoogle-$0.1653.6%83.1%137258.3
GLM-4.7Zhipu-$0.05--1440-
Claude Sonnet 4.5Anthropic-$0.3071.4%82.4%1383-
GPT-5.2OpenAI-$0.2372.8%88.0%143248.9
Claude Haiku 4.5Anthropic✓$0.1066.6%73.5%129045.3
GPT-4oOpenAI-$0.2348.9%72.9%1372-
Gemini 2.5 FlashGoogle-$0.0428.7%55.1%123347.7
GPT-5 miniOpenAI✓$0.0356.2%50.2%1145-
GPT-4.1OpenAI-$0.1839.6%52.4%1305-

Column guide

Column What it means
Copilot Available in GitHub Copilot (✓ = listed, - = not listed). Copilot usage is billed in AI credits based on tokens used.
$/task Estimated direct-API cost for 50K input + 10K output tokens; it is not a Copilot bill estimate.
SWE-bench % of real GitHub issues the model can fix autonomously (source) - February 2026 data (standardized harness, high reasoning mode)
Aider % correct on multi-language code editing (source) - latest listed runs are October 2025; historical comparison only
Arena Elo rating from human preference voting on Code category (source) - February 2026 data
LiveBench Global average score across 23 diverse tasks (source) - current scores shown from release 2026-06-25

DeepSeek V4.1 Flash has time-based peak/off-peak rates; this task estimate uses peak pricing. †Gemini 3.7/3.8 Flash promotional API pricing runs through Dec 31, 2026.

Data sources: SWE-bench (Feb 2026) · Aider (latest listed runs: Oct 2025) · Arena Code (Feb 2026, not refreshed) · LiveBench (latest release: v2026-06-25; leaderboard includes newer models) · ProgramBench (leaderboard updated Sep 28, 2026) · GitHub Copilot (AI credit billing since Jun 2026)
API pricing: Anthropic · OpenAI · Google · DeepSeek · Zhipu (GLM)