Getting Started: AI Coding Quick Reference
📅 Last updated: October 2026
No-nonsense reference for developers who just want to know which AI model to pick. Bookmark this and stop Googling.
📌 Quick pick - just tell me what to use
These are starting points based on available benchmark results and provider pricing, not controlled tests of latency or code quality for each workflow. There is no public benchmark here specifically for documentation or code review.
Gemini 3.8 Flash
GPT-6 Sol
Gemini 3.7 Flash
DeepSeek V4.1 Flash
GPT-6 Astra
Gemini 3.8 Flash
Gemini 3.8 Flash
GPT-6 Sol
Gemini 3.7 Flash
GPT-5.6 Luna
GPT-6 Astra
GPT-5.6 Sol
A note on versions
You’ll see version numbers everywhere: Sonnet 3.5, Sonnet 4, Sonnet 4.5. Gemini 2.5, Gemini 3. GPT-4o, GPT-5.
Don’t overthink it. The tier matters more than the version. “Sonnet” is the mid-tier Claude. “Opus” is the heavyweight Claude. “Flash” is the fast/cheap Gemini. Your IDE usually offers the latest version of each tier - just pick the tier that fits your task.
When this page says “Sonnet”, it means whatever the current Sonnet is. Same for the others.
The big three model families
Speed key: ⚡⚡⚡ Fast · ⚡⚡ Medium · ⚡ Slow
Anthropic (Claude)
| Model | What it’s for | Speed | Cost |
|---|---|---|---|
| Haiku 4.5 | Fast tasks, scaffolding, CLI | ⚡⚡⚡ | $0.10/task |
| Sonnet 5 / 5.5 | Everyday coding | ⚡⚡ | $0.20/task |
| Opus 5 | Complex reasoning, design | ⚡ | $0.50/task |
| Opus 5.5 | Higher-end reasoning | ⚡ | $0.40/task |
| Fable 5 / 5.1 | Frontier reasoning | ⚡ | $1.00/task |
OpenAI (GPT)
| Model | What it’s for | Speed | Cost |
|---|---|---|---|
| GPT-6 Luna | Lightweight, lowest-cost frontier choice | ⚡⚡⚡ | $0.01/task |
| GPT-6 / GPT-6.1 Sol | General coding and reasoning | ⚡⚡ | $0.20/task |
| GPT-5.6 Luna / Terra / Sol | Budget to frontier variants | ⚡⚡ to ⚡ | $0.02–$0.40/task |
| GPT-6 Astra | Highest-capability tier | ⚡ | $1.00/task |
| GPT-5 mini | Lightweight legacy option | ⚡⚡⚡ | $0.03/task |
- Implementing algorithms (graph traversal, dynamic programming)
- Debugging race conditions or complex state machines
- Mathematical proofs or formal verification
Google (Gemini)
| Model | What it’s for | Speed | Cost |
|---|---|---|---|
| Gemini 3.7 / 3.8 Flash | Current fast general-purpose models | ⚡⚡ | $0.075/task through Dec 31, 2026* |
| Gemini 3.6 Flash | Previous generation | ⚡⚡ | $0.075/task through Dec 31, 2026* |
| Gemini 3.5 Flash-Lite | Lower-cost, high-volume workloads | ⚡⚡⚡ | $0.04/task |
| Gemini 3.1 Pro | Higher-capability tasks | ⚡ | $0.22/task (≤200K context) |
Benchmarks
Want numbers?
- Compare all models - sortable table, filter by Copilot cost
- Benchmark details - methodology, sources, caveats
The TLDR:
- LiveBench’s current leader is Claude Fable 5.1 Max (83.4), followed by Claude Opus 5.5 (83.2) and Fable 5 Max (83.0). LiveBench still labels its question-set release 2026-06-25; these are later leaderboard submissions, not a new benchmark release.
- GPT-6 variants are in the mix: Astra scores 82.2, GPT-6.1 Sol 81.6, Sol 79.3 and Luna 72.0. Direct-API task estimates run from about $0.01 to $1.00 using the same token assumptions.
- The budget frontier moved: DeepSeek V4.1 Flash scores 81.1 on LiveBench; its price varies by peak/off-peak time. Gemini 3.7/3.8 Flash is temporarily discounted through Dec 31, 2026.
- ProgramBench adds a different kind of evidence: its strict full-program reconstruction tasks remain difficult even for the leaders. Don’t compare its pass rates directly with SWE-bench issue resolution.
- SWE-bench Verified, Aider, and Arena have not kept pace with current releases. Their published scores remain useful historical measurements, not evidence about models absent from those runs.
The benchmark page records each source’s access date and methodology; LiveBench’s checked leaderboard date and question-set release are separate.
Google’s Gemini 3.6/3.7/3.8 Flash promotional API rates end Dec 31, 2026. Anthropic’s published Claude Sonnet 5 price is $2/$10 per MTok (the previously announced September increase was cancelled).
Benchmarks are useful for gut-checking, but the real test is running a model on your own work.
Marketing BS decoder
| They say | It means |
|---|---|
| “Most intelligent” | Bigger, slower, pricier |
| “Balanced” | Mid-tier - usually right |
| “Fast” / “efficient” | Smaller, cheaper, simpler |
| “Reasoning” / “thinking” | Extra thinking time - see below |
| “Preview” / “experimental” | Unstable - skip it |
| “200K context” | Can see lots of code - but should it? |
When do “reasoning” models actually help?
Reasoning models (o1, o3, “thinking” variants) work through problems step-by-step before responding.
Worth it for:
- Implementing complex algorithms (A*, red-black trees, constraint solvers)
- Debugging concurrency issues, race conditions, deadlocks
- Untangling deeply nested dependency chains
- Mathematical proofs or formal logic
Overkill for:
- Adding a new API endpoint
- Fixing a null pointer exception
- Writing unit tests
- Refactoring for readability
- Most day-to-day feature work
A standard model with a good prompt is faster and cheaper for 90% of coding tasks.
What about context window size?
Context window (what’s this?) = how much code the model can “see” at once. Bigger sounds better, but:
- More context = more noise. The model gets distracted.
- More context = slower and pricier. You pay per token.
- You rarely need it. Most tasks involve a few files, not hundreds.
Big windows help for: exploring unfamiliar codebases, analysing logs, multi-file refactors. For everyday coding, focused context beats massive context.