THE VAULT / SHIP IT / MARKDOWN
Model routing cheat sheet
Models are employees. The flagship is a $400/hr consultant, the mid-tier is your salaried senior dev, the cheap tier is the intern army, and most vibe coders put the consultant on data entry all day and then wonder where the money went. This is my routing sheet: which model to hire for which job, the July 2026 price list, the break-even math on $20/$100/$200 plans, and the handoff prompt that lets you downgrade mid-project without losing the plot.
LAST VERIFIED 2026-07-08 · FOUNDING MEMBER RESOURCE
NOTE
Use this page as a prompt reference. Every file below is built to be handed straight to an AI. Download it, drop it into Claude, ChatGPT, Cursor, or Lovable, and say "use this as my playbook." The prompts on this page are copy-paste ready.
WATCH OUT
LAST VERIFIED: 2026-07-08. Prices move monthly. Two of the numbers below have expiry dates printed on them (Sonnet 5 intro pricing ends 2026-08-31, GPT-5.6 is preview pricing). Check the vendor page before you commit real budget. All API prices in USD per 1M tokens, input/output.
Stop asking which model is best. Ask which one to hire.
"Which model is best" is the wrong question, the same way "who is the best employee" is the wrong question. Best at what, for how much? Claude Opus 4.8 tops the mid-2026 coding benchmarks and costs $5/$25. Gemini 3.1 Flash-Lite costs $0.25/$1.50 and will happily classify 10,000 support tickets while Opus is still clearing its throat. Neither is "best." They have different salaries and different jobs.
The meta-finding repeated across every serious mid-2026 comparison: no single model wins every lane. Heavy devs who route by task spend $20 to $200 a month. Heavy devs who run the flagship on everything spend $1,000+. Same output, 5 to 10x price difference. Routing is the highest-ROI skill nobody teaches.
| Tier | The employee | Salary (per 1M tok) | Put them on |
|---|---|---|---|
| Consultant | Claude Fable 5, GPT-5.5 Pro | $10/$50 and up | Problems that already defeated the tier below, unattended overnight runs |
| Senior engineer | Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro | $2 to $5 in, $12 to $30 out | Architecture, long agentic builds, hard debugging, big-doc research |
| Mid-level dev | Claude Sonnet 5, GPT-5.4 | $2 to $3 in, $10 to $15 out | Daily coding, features that follow an existing pattern, content |
| Intern army | Haiku 4.5, GPT-5.4 nano, Flash-Lite | $0.20 to $1 in, $1.25 to $5 out | Classification, extraction, renames, subagent grunt work |
The routing table: 9 job types, one hire each
This is the core of the sheet. Cost bands are my estimates from typical session token volumes with caching on; your mileage varies with context size. When two models are listed, the first is the default.
| Job | Hire | Why | Cost band per task (API, est.) |
|---|---|---|---|
| Scaffold a new app | Claude Sonnet 5 | Near-Opus coding at $2/$10 intro pricing (through 2026-08-31). The default in Claude Code and most IDE routings for a reason. | $0.50 to $2 per scaffold session |
| Long agentic build (multi-hour, high stakes) | Claude Opus 4.8 | SWE-bench Verified ~88.6%, SWE-bench Pro ~69.2% vs GPT-5.5 at 58.6% and Gemini 3.1 Pro at 54.2%. Fewest rescue interventions on long runs. 1M context at standard rates. | $10 to $50 per run |
| The hardest problems / overnight runs | Claude Fable 5 | $10/$50, 2x Opus. Only when Opus already failed or the run is unattended and a failure costs more than the tokens. Requires 30-day data retention, no ZDR. | $5 to $50+ per run |
| Quick edits, rename, refactor-in-place | Sonnet 5 (Haiku 4.5 for trivial) | Flagship latency and price are wasted on a 40-line diff. | $0.05 to $0.50 |
| Debugging | Sonnet 5 first, Opus 4.8 after 2 failed attempts, Fable 5 last resort | The escalation ladder beats starting expensive: most bugs die at the cheap tier. | $0.20 to $10 |
| High-volume agentic pipelines, computer use | GPT-5.5 | Faster and cheaper to operate per task at $5/$30, strong multi-tool persistence. Watch the hidden reasoning tokens (see cost traps). | $0.50 to $5 per pipeline task |
| Research, planning, big-doc analysis | Gemini 3.1 Pro | $2/$12 with 1M context and strong multimodal (images, PDF, video). Cheapest way to read a whole codebase or a 300-page PDF. Input doubles to $4 past 200K tokens. | $0.50 to $3 |
| Content: drafts, docs, marketing copy | Sonnet 5 or GPT-5.4 | Writing quality plateaus below flagship prices. $2/$10 (Sonnet intro) or $2.50/$15 (GPT-5.4) is plenty. | $0.10 to $0.50 per piece |
| Cheap bulk: classification, extraction, subagents | Gemini 3.1 Flash-Lite ($0.25/$1.50), GPT-5.4 nano ($0.20/$1.25), Haiku 4.5 ($1/$5) | Batch API cuts all of these another 50%. Flash-Lite batched is $0.125/$0.75. | Pennies per thousand items |
NOTE
Route by task tier, not by loyalty. The one-vendor setup feels tidy and costs you real money. My own split: Claude for anything that touches code, Gemini for reading huge things cheaply, a nano/Flash-Lite tier for bulk. That covers 95% of jobs.
The price sheet: July 2026, all three vendors
Anthropic (per 1M tokens)
| Model | Context | Input | Output |
|---|---|---|---|
| Claude Fable 5 | 1M | $10.00 | $50.00 |
| Claude Opus 4.8 | 1M | $5.00 | $25.00 |
| Claude Opus 4.7 / 4.6 | 1M | $5.00 | $25.00 |
| Claude Sonnet 5 | 1M | $3.00 ($2.00 intro through 2026-08-31) | $15.00 ($10.00 intro) |
| Claude Sonnet 4.6 | 1M | $3.00 | $15.00 |
| Claude Haiku 4.5 | 200K | $1.00 | $5.00 |
Modifiers: prompt-cache reads cost ~0.1x input, cache writes 1.25x (5-min TTL) or 2x (1-hr TTL), Batch API is 50% off. No long-context premium on Opus 4.8: the full 1M window bills at standard rates.
WATCH OUT
Sonnet 5 tokenizer trap: it uses a new tokenizer that produces ~30% more tokens for the same text vs Sonnet 4.6. Effective cost is higher than the sticker suggests, and the intro discount partly exists to soften that. Compare invoices, not price pages.
OpenAI (per 1M tokens)
| Model | Input | Output |
|---|---|---|
| GPT-5.6 "Sol" (limited preview) | $5.00 | $30.00 |
| GPT-5.6 "Terra" (preview) | $2.50 | $15.00 |
| GPT-5.6 "Luna" (preview) | $1.00 | $6.00 |
| GPT-5.5 (flagship) | $5.00 | $30.00 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.4 mini | $0.75 | $4.50 |
| GPT-5.4 nano | $0.20 | $1.25 |
| o3 (post July-2026 price cut) | $2.00 | $8.00 |
| o3-pro | $20.00 | $80.00 |
| o4-mini | $1.10 | $4.40 |
Modifiers: cached input at ~10% of input price, automatic. Batch endpoint 50% off (GPT-5.5 drops to $2.50/$15 batched).
WATCH OUT
o-series trap: reasoning models bill hidden reasoning tokens at output rates. Real cost runs 3 to 10x the visible tokens. And GPT-5.6 Sol/Terra/Luna are preview pricing; expect changes at GA.
Google Gemini (per 1M tokens)
| Model | Input | Output |
|---|---|---|
| Gemini 3.1 Pro | $2.00 (>200K tokens: $4.00) | $12.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 |
Modifiers: Batch mode 50% off all models (up to 24h turnaround), cached input at ~10% of the cache-miss rate on Pro and Flash. The rate-limited free API tier is still generous, which makes Gemini the cheapest place to prototype a bulk pipeline before you pay anyone.
Subscription math: when $20/$100/$200 beats API keys
The flat plans are usage arbitrage. If your usage is steady and heavy, the vendor loses money on you and that is the point. If your usage is spiky or light, you are the one subsidizing the heavy users. Here are the plans, then the break-evens.
| Plan | Price | What you get |
|---|---|---|
| Claude Pro | $17/mo annual, $20 monthly | Claude Code included, plus Cowork, Design, Science |
| Claude Max 5x | $100/mo | 5x Pro usage, higher output limits |
| Claude Max 20x | $200/mo | 20x Pro usage |
| ChatGPT Plus | $20/mo | ~160 msgs/3h instant, 3,000 thinking msgs/wk, 25 deep-research runs/mo |
| ChatGPT Pro | $100/mo | 5x Plus, GPT-5.5 Pro, o3-pro, Codex at 5x (10x promo expired 2026-05-31) |
| ChatGPT Pro | $200/mo | 20x Plus, 250 deep-research runs/mo, Operator agent |
| Google AI Pro | $19.99/mo | Deep Search, Jules agent, $10/mo Cloud credits |
Break-even, real numbers
Working assumption: Sonnet 5 as the daily driver, and a typical 2-hour Claude Code session burning roughly $3 to $8 in API-equivalent tokens (estimated, cache-assisted). From there:
- 1
$20 Pro pays for itself after 3 to 6 sessions a month
One afternoon a week of Claude Code and the sub already wins. This is the easiest yes in the whole sheet: if you code with AI at all, the $20 tier is close to free money.
- 2
$100 Max 5x breaks even around $100/mo of API-equivalent usage
That's daily Sonnet sessions, or a few Opus agent runs a week. Reference point: metered Codex heavy-dev spend runs $100 to $200/mo, so a flat $100 on the Claude side is a genuine discount at the same intensity.
- 3
$200 Max 20x is for Opus-heavy daily agent work
API-metered equivalents run $500 to $1,000+/mo for heavy agentic users. If your agents run while you sleep, the $200 plan is the cheap option. Strange sentence, true in 2026.
- 4
Set the overflow cap on day one
New in 2026: Claude Pro/Max overflow "usage credits" continue at API rates after your plan limit, with a monthly cap you set. Set the cap before you need it, not after the surprise invoice.
When API keys still win
- [ ]Spiky usage: heavy one week, zero the next. Subs charge you for the zero weeks.
- [ ]Bulk jobs: the Batch API's 50% discount has no subscription equivalent.
- [ ]Multi-provider routing: subs lock you into one vendor's models; the routing table above wants at least two vendors.
- [ ]Anything production or user-facing: never serve users off your personal sub. Rate limits, ToS, and your own usage all collide.
PAYOFF
Can't decide? Run one metered week. Put $25 on an API key, work normally, read the invoice. Multiply by 4 and pick the plan the math picks. Ten minutes of arithmetic beats a month of guessing.
The downgrade ladder: start expensive, hand off cheap
The single most expensive routing mistake is symmetric: people either run Opus on everything (burning money) or run the cheap model on everything (burning time on rescue work). The fix is a ladder. The expensive model establishes the pattern once; the cheap model copies it forever.
- 1
Tier 1: Opus 4.8 (or Fable 5) sets the pattern
Architecture, the first vertical slice, CLAUDE.md conventions, the hard first version of every pattern. This is where judgment matters and where the $25/M output tokens buy something real.
- 2
Tier 2: Sonnet 5 grinds the pattern out
Everything that now has an example to copy: more endpoints shaped like the first endpoint, more components shaped like the first component, tests shaped like existing tests. This is 70% of a project's tokens at 40% of the price.
- 3
Tier 3: Haiku 4.5 / Flash-Lite / 5.4 nano do the mechanical bulk
Renames, data cleanup, classification, subagent grunt work. Nothing at this tier should require judgment; if it does, it's mis-routed.
- 4
Escalate back up only after 2 failed attempts
Two strikes at the current tier, then move one tier up with a fresh context. Do not let a cheap model retry itself into a hole; its failed attempts poison the context for attempt three.
The ladder only works if the cheap model inherits the context. It won't have the expensive model's reasoning history, so you make the expensive model write it down before you switch. Run this on the expensive model as the last act of its session:
Save the output as HANDOFF.md, then open the cheap session with exactly: "Read HANDOFF.md, then start task 1." The tripwires section is the part people skip and regret; it's what stops a $1/M model from improvising a rewrite of your auth middleware at 2am.
Cost control: 9 tactics, checkable
- [ ]Cache-friendly prompts: system prompt and stable context at the TOP, variable stuff last. Cache hits need identical prefixes, and reads bill at ~0.1x input.
- [ ]Never edit the top of a long conversation. It invalidates the cache for everything after it, and you re-pay full input price on the whole history.
- [ ]Small contexts: point the agent at 5 relevant files, not the repo. On Gemini 3.1 Pro specifically, stay under 200K input or the rate doubles to $4.
- [ ]Kill loops early: an agent that failed the same test 3 times is not about to succeed on attempt 7. Stop, escalate a tier or fix it yourself.
- [ ]Review the plan before any long autonomous run. 10 minutes of plan review is cheaper than $40 of confidently wrong Opus tokens.
- [ ]Batch anything that can wait 24h: 50% off on Anthropic, OpenAI, and Gemini. The single biggest lever for bulk work.
- [ ]On reasoning models, watch billed tokens, not visible tokens. o-series hidden reasoning runs 3 to 10x what you see on screen.
- [ ]Set the overflow usage-credit cap on Claude Pro/Max before you hit the plan limit.
- [ ]Re-check this sheet monthly. Sonnet 5 intro pricing ends 2026-08-31, GPT-5.6 is preview pricing, and o3 just took a price cut. This market reprices faster than cloud ever did.
Per-tool notes: plan quirks that change the math
The routing logic above is tool-agnostic, but each tool's billing has quirks that shift what "cheap" means inside it. The four I get asked about most:
Claude Code: no standalone SKU, and the Team trap
Claude Code has no standalone price. It rides on Pro ($17 to $20), Max ($100/$200), Team Premium ($100/seat annual, minimum 5 seats), or raw API pay-as-you-go.
WATCH OUT
Team Standard ($20/seat) does NOT include Claude Code. Teams get burned by this constantly: they buy the $20 seats, discover the dev team can't use Claude Code, and end up paying $100/seat Premium anyway. Price the Premium seats from day one.
Default model is Sonnet 5; switch per-task with /model. The downgrade ladder maps directly: /model opus to establish the pattern, /model sonnet to grind it out, Haiku via subagents for the bulk tier.
Cursor: Auto mode is the free lunch
Pro is $20/mo structured as a $20 credit pool, but Auto mode is unlimited and does not burn credits. The play: route routine edits through Auto, spend credits only on named-model requests where you specifically want Opus or GPT-5.5.
Above Pro: Pro+ $60 (3x usage), Ultra $200 (20x, priority features), Teams $40/user with $120 Premium seats at 5x. Annual billing takes 20% off everything.
The Hobby free tier (~2,000 completions, 50 slow premium requests/mo) is fine for a weekend of evaluation, not for work.
Zed: the cleanest BYO-keys setup
Personal plan is free with your own API keys: you pay list price and nothing else, which makes Zed the cleanest home for the API-key routing strategy in this sheet.
Pro is $10/mo plus hosted usage billed at API list +10%, with up to $10 of usage included (call it ~$20/mo typical). At real volume that +10% markup loses to BYO keys, so Pro mostly buys convenience, not savings.
Students: Pro free for a year with $10/mo in token credits, hosted models except Claude Opus.
Windsurf: quotas now, but grandfathering muddies it
Windsurf moved from credits to quotas for subscribers who joined after March 2026: Free (daily quota), Pro $20/mo, Max $200/mo, Teams $40/user/mo. Pre-purchased credits were converted to extra usage, so grandfathered accounts may not match the public pricing page. Check your own plan screen, not the marketing page.
One genuinely nice detail: tab autocomplete is unlimited on every plan, including Free.
Take the sheet with you
The whole cheat sheet as a single markdown file: routing table, price sheet, subscription math, the handoff prompt, and the cost-control checklist. Drop it in your repo or your notes app, and re-verify the prices monthly (the LAST VERIFIED date is at the top of the file for exactly that reason).
PAYOFF
The move this week: pick your top 3 recurring task types, write the model you'll use for each on a sticky note, and stop deciding per-session. Routing decided once beats routing re-litigated daily. My own note reads: code = Sonnet, hard code = Opus, bulk = Flash-Lite.
Get the next drop
Every new Vault system ships to the list first. Twice a week, free.