<!--
ROBOTS ON PAYROLL / THE VAULT
What this is: a cheat sheet for routing tasks to the right AI model tier by job, with API prices as of July 2026.
How to use it: download this file and give it to your AI (Claude, ChatGPT,
Cursor, Lovable chat) as a reference. Say: "Use this as my playbook for
choosing which AI model to use."
Latest version + full guide: https://robotsonpayroll.com/resources/model-routing-cheatsheet
-->

# Model Routing Cheat Sheet

**LAST VERIFIED: 2026-07-08.** Prices move monthly. Check the vendor pricing page before you commit real budget. All API prices in USD per 1M tokens (input/output).

Frame: models are employees. The flagship is a $400/hr consultant, the mid-tier is your salaried senior dev, the cheap tier is the intern army. You would not put the consultant on data entry. Route accordingly.

---

## 1. The Routing Table

| Job | Hire | Why | Cost band per typical task (API, est.) |
|---|---|---|---|
| Scaffold a new app | Claude Sonnet 5 | Near-Opus coding at $2/$10 intro pricing (through 2026-08-31). The default in Claude Code and most IDE routings for a reason. | $0.50 to $2 per scaffold session |
| Long agentic build (multi-hour, high stakes) | Claude Opus 4.8 | SWE-bench Verified ~88.6%, SWE-bench Pro ~69.2% vs GPT-5.5 at 58.6%. Best edge-case handling; fewest rescue interventions on long runs. 1M context at standard rates. | $10 to $50 per run |
| The hardest problems / overnight runs | Claude Fable 5 | $10/$50, 2x Opus. Only when Opus has already failed or the run is unattended and a failure costs more than the tokens. Note: requires 30-day data retention, no ZDR. | $5 to $50+ per run |
| Quick edits, rename/refactor-in-place | Claude Sonnet 5 (Haiku 4.5 for trivial) | Flagship latency and price are wasted on a 40-line diff. | $0.05 to $0.50 |
| Debugging | Sonnet 5 first, escalate to Opus 4.8 after 2 failed attempts, Fable 5 as last resort | Escalation ladder beats starting expensive: most bugs die at the cheap tier. | $0.20 to $10 |
| High-volume agentic pipelines, computer use | GPT-5.5 | Faster and cheaper to operate per task at $5/$30; strong multi-tool persistence. Watch hidden reasoning tokens (see cost traps). | $0.50 to $5 per pipeline task |
| Research and planning, big-doc analysis | Gemini 3.1 Pro | $2/$12 with 1M context and strong multimodal (images/PDF/video). Cheapest way to read a whole codebase or 300-page PDF. Input jumps to $4 past 200K tokens. | $0.50 to $3 |
| Content: drafts, docs, marketing copy | Sonnet 5 or GPT-5.4 | Writing quality plateaus below flagship prices. $2.50/$15 (GPT-5.4) or $2/$10 (Sonnet intro) is plenty. | $0.10 to $0.50 per piece |
| Cheap bulk: classification, extraction, subagents | Gemini 3.1 Flash-Lite ($0.25/$1.50), GPT-5.4 nano ($0.20/$1.25), Claude Haiku 4.5 ($1/$5) | Batch API cuts all of these another 50%. Flash-Lite batched is $0.125/$0.75. | Pennies per thousand items |

Meta-rule (repeated across every mid-2026 comparison): no single model wins every lane. Route by task tier, not by loyalty.

---

## 2. The Price Sheet (July 2026)

### Anthropic (per 1M tokens)

| Model | Context | Input | Output |
|---|---|---|---|
| Claude Fable 5 | 1M | $10.00 | $50.00 |
| Claude Opus 4.8 | 1M | $5.00 | $25.00 |
| Claude Opus 4.7 / 4.6 | 1M | $5.00 | $25.00 |
| Claude Sonnet 5 | 1M | $3.00 ($2.00 intro through 2026-08-31) | $15.00 ($10.00 intro) |
| Claude Sonnet 4.6 | 1M | $3.00 | $15.00 |
| Claude Haiku 4.5 | 200K | $1.00 | $5.00 |

Modifiers: cache reads ~0.1x input; cache writes 1.25x (5-min TTL) or 2x (1-hr TTL); Batch API 50% off. No long-context premium on Opus 4.8.

Traps: Sonnet 5 uses a new tokenizer (~30% more tokens for the same text vs 4.6, so effective cost is higher than the sticker suggests). Fable 5 requires 30-day data retention (no ZDR). Opus 4.8 fast mode (research preview) costs a premium for up to 2.5x output speed.

### OpenAI (per 1M tokens)

| Model | Input | Output |
|---|---|---|
| GPT-5.6 "Sol" (limited preview) | $5.00 | $30.00 |
| GPT-5.6 "Terra" (preview) | $2.50 | $15.00 |
| GPT-5.6 "Luna" (preview) | $1.00 | $6.00 |
| GPT-5.5 (flagship) | $5.00 | $30.00 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.4 mini | $0.75 | $4.50 |
| GPT-5.4 nano | $0.20 | $1.25 |
| o3 (post July-2026 cut) | $2.00 | $8.00 |
| o3-pro | $20.00 | $80.00 |
| o4-mini | $1.10 | $4.40 |

Modifiers: cached input ~10% of input price (automatic); Batch endpoint 50% off.

Trap: o-series reasoning models bill hidden reasoning tokens at output rates. Real cost runs 3 to 10x the visible tokens. GPT-5.6 preview pricing may change at GA.

### Google Gemini (per 1M tokens)

| Model | Input | Output |
|---|---|---|
| Gemini 3.1 Pro | $2.00 (>200K tokens: $4.00) | $12.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 |

Modifiers: Batch mode 50% off all models (up to 24h turnaround); cached input ~10% of cache-miss rate on Pro/Flash. Generous rate-limited free API tier still exists.

---

## 3. Subscription Math

### The plans (July 2026)

| Plan | Price | What you get |
|---|---|---|
| Claude Pro | $17/mo annual, $20 monthly | Claude Code included, plus Cowork/Design/Science |
| Claude Max 5x | $100/mo | 5x Pro usage, higher output limits |
| Claude Max 20x | $200/mo | 20x Pro usage |
| ChatGPT Plus | $20/mo | ~160 msgs/3h instant, 3,000 thinking msgs/wk, 25 deep-research runs/mo |
| ChatGPT Pro | $100/mo | 5x Plus, GPT-5.5 Pro, o3-pro, Codex (5x since promo ended 2026-05-31) |
| ChatGPT Pro | $200/mo | 20x Plus, 250 deep-research runs/mo, Operator |
| Google AI Pro | $19.99/mo | Deep Search, Jules agent, $10/mo Cloud credits |

### Break-even, real numbers

Assume Sonnet 5 as the daily driver. A typical 2-hour Claude Code session burns roughly $3 to $8 in API-equivalent tokens (est., cache-assisted).

- **$20 Pro beats API keys after ~3 to 6 sessions a month.** If you run Claude Code more than one afternoon a week, the sub wins.
- **$100 Max 5x beats keys around $100/mo of API-equivalent usage:** daily Sonnet sessions, or a few Opus agent runs a week. Reference point: metered Codex heavy-dev spend runs $100 to $200/mo, so a $100 flat plan on the Claude side is a real discount for the same intensity.
- **$200 Max 20x is for Opus-heavy daily agent work.** API-metered equivalents run $500 to $1,000+/mo for heavy agentic users. If your agents run while you sleep, this is the cheap option, which is a strange sentence but true in 2026.
- **New in 2026:** Pro/Max overflow "usage credits" continue at API rates after your plan limit, with a monthly cap you set. Set the cap before you need it.

### When API keys still win

- Spiky usage (heavy one week, zero the next). Subs charge you for the zero weeks.
- Bulk jobs: Batch API 50% off has no subscription equivalent.
- Multi-provider routing (subs lock you to one vendor's models).
- Anything production/user-facing. Never serve users off your personal sub.

---

## 4. The Downgrade Ladder

Start expensive, establish the pattern, downgrade the routine work. Ladder:

1. **Opus 4.8 (or Fable 5):** architecture, first vertical slice, CLAUDE.md conventions, the hard first version of the pattern.
2. **Sonnet 5:** everything that now has an example to copy: more endpoints, more components, more tests shaped like existing tests.
3. **Haiku 4.5 / Flash-Lite / 5.4 nano:** mechanical bulk: renames, data cleanup, classification, subagent grunt work.

Escalate back up only after 2 failed attempts at the current tier.

### The handoff prompt (run on the EXPENSIVE model before you downgrade)

```
You are ending a session. A cheaper, less capable model will continue this
work and it will NOT have your reasoning history. Write a handoff brief it
can follow mechanically.

Project: [PROJECT NAME AND ONE-LINE DESCRIPTION]
Work completed this session: [WHAT YOU JUST BUILT]
Remaining work to hand off: [THE ROUTINE TASKS LEFT]

Produce a single markdown file with exactly these sections:

## Pattern to copy
The ONE canonical example of the established pattern, with the file path
and the actual code. If there are multiple patterns, one example each.

## Rules that are not obvious from the code
Every convention I would enforce in review: naming, error handling, where
files go, what to import from where. Bullet list, imperative voice.

## Task list
The remaining tasks as a numbered checklist. Each task: file(s) to touch,
which pattern from above to copy, and a one-line acceptance check.

## Do not touch
Files and behaviors that must not change, with one line on why each is
fragile.

## When to stop and escalate
3 to 5 concrete tripwires (e.g. "any change to the auth middleware",
"test suite fails twice in a row") where the cheaper model must stop and
report instead of improvising.

Constraints: no motivational filler, no restating the project pitch.
Assume the reader executes instructions literally and has zero context
beyond this file. Under 800 words.
```

Save the output as `HANDOFF.md`, start the cheap session with "Read HANDOFF.md, then start task 1."

---

## 5. Cost Control Checklist

- [ ] System prompt and stable context at the TOP of the prompt, variable stuff last (cache hits need identical prefixes; reads are ~0.1x input price)
- [ ] Don't edit the top of a long-running conversation; it invalidates the cache for everything after it
- [ ] Small contexts: point the agent at 5 relevant files, not the repo; on Gemini 3.1 Pro specifically, stay under 200K input or the rate doubles
- [ ] Kill loops early: an agent that has failed the same test 3 times is not about to succeed on attempt 7; stop, escalate a tier or fix it yourself
- [ ] Review the plan before any long autonomous run; 10 minutes of plan review is cheaper than $40 of confidently wrong Opus tokens
- [ ] Batch anything that can wait 24h (50% off on Anthropic, OpenAI, and Gemini)
- [ ] On o-series/reasoning models, watch billed (not visible) tokens; hidden reasoning runs 3 to 10x visible
- [ ] Set the overflow usage-credit cap on Claude Pro/Max before you hit the plan limit
- [ ] Re-check this sheet monthly; intro pricing (Sonnet 5 ends 2026-08-31) and preview pricing (GPT-5.6) both expire

---

## 6. Per-Tool Notes (July 2026)

### Claude Code
- No standalone SKU. It rides on Pro ($17 to $20), Max ($100/$200), Team Premium ($100/seat annual, min 5 seats), or raw API pay-as-you-go.
- Team Standard ($20/seat) does NOT include Claude Code. Teams get burned by this constantly.
- Default model is Sonnet 5; switch per-task with `/model`. The downgrade ladder maps directly: `/model opus` to establish, `/model sonnet` to grind, Haiku via subagents.

### Cursor
- Pro $20/mo is a $20 credit pool, but **Auto mode is unlimited and does not burn credits.** Route routine edits through Auto, spend credits only on named-model requests.
- Pro+ $60 (3x usage), Ultra $200 (20x). Teams $40/user, Premium seats $120. Annual billing is 20% off.
- Hobby free tier: ~2,000 completions and 50 slow premium requests/mo. Fine for evaluation, not for work.

### Zed
- Personal plan is free with BYO API keys: the cleanest setup for API-key routing, you pay list price and nothing else.
- Pro $10/mo + hosted usage at API list +10%, with up to $10 of usage included (~$20/mo typical). At real volume, BYO keys beat the +10% markup.
- Student: Pro free for a year ($10/mo token credits; hosted models except Claude Opus).

### Windsurf
- Moved from credits to quotas for subscribers who joined after March 2026: Free (daily quota) / Pro $20 / Max $200 / Teams $40/user. Grandfathered users keep converted credits, so your plan may differ from the pricing page.
- Tab autocomplete is unlimited on every plan, including Free.

---

*Prices verified 2026-07-08 against vendor pricing pages and independent trackers. Cost bands are my estimates from typical session token volumes; your mileage varies with context size and caching. When in doubt, run one metered week on API keys and let the invoice pick your plan.*
