Always Reach for the Biggest Model
Stop downgrading your model to save money. For real work, reach for the most capable model in whatever stack you're in — Claude Opus, GPT-5.5, Gemini 3 Pro — not the cheaper tier you keep getting nudged toward: Sonnet, Haiku, GPT-5.4 and its mini and nano siblings, Gemini Flash. Step down only when you have a reason. The popular advice is the exact opposite, and I think the popular advice is usually wrong.
You have heard it everywhere. "Don't use Opus for that, Sonnet is fine." "Just route the easy stuff to Flash." "Haiku is 90% as good for a fifth of the price." By now it has hardened into official guidance — the kind of line that lands in a company's AI playbook: "Select the model that fits the task; using Opus as default does not provide additional value — see the GitHub Copilot model comparison." And that comparison delivers exactly what you'd expect: it pens Opus into "deep reasoning and debugging," and tells everyone else, in its own words, to "start with a general-purpose option like GPT-5 mini, then adjust." Pick the fit. Never reach for the ceiling. The default is down.
I think down is the wrong default — and I think "using Opus as default adds no value" is one of the most expensive sentences in that playbook, precisely because the cost it ignores never shows up where anyone is looking.
So I do the opposite. I reach for the biggest model first and make it earn its way out of the job. And I mean first — not just for the gnarly refactor, but for the throwaway question too. Ask "how do I write this Postgres window function" and the popular move is to flick it to Haiku or Flash. My move is Opus on low effort: I want the answer that's right the first time, that I don't have to second-guess, that I don't have to paste back in an hour later because it quietly invented a column. The small model often has to think longer to land the same fact — and then I still have to check it. That instinct holds whoever makes the model: Anthropic, OpenAI, Google. Reach for the smarter brain; turn its reasoning down, not up.
That's a strong claim, so I went looking for the numbers — the ones that back me up and the ones that don't. Both piles are below. I'll tell you where I land, but the last line is yours, not mine.
First, the strongest argument against me
Let me hand my opponents their best weapon: the price tag is real, and it is not subtle.
Per token, the most capable models are genuinely expensive. As of June 2026, Anthropic's own pricing puts Claude Opus 4.8 at $5 / $25 per million input/output tokens, Sonnet 4.6 at $3 / $15, and Haiku 4.5 at $1 / $5. Opus is exactly 5× the price of Haiku and 1.67× Sonnet on the same tokens. It's the same story everywhere: OpenAI runs from GPT-5.5 at $5 / $30 down through GPT-5.4 ($2.50 / $15), its mini ($0.75 / $4.50) and nano ($0.20 / $1.25) — the flagship is 24× the output price of nano — while Google spans Gemini 3 Pro ($2 / $12) all the way down to Gemini 2.5 Flash-Lite at $0.10 / $0.40.
It gets worse for me. Anthropic's own guidance tells you not to do what I do. Their pricing page, under "cost optimization strategies," says plainly: "Choose Haiku for simple tasks, Sonnet for most production workloads, and Opus for the most complex reasoning." Their Opus product page frames the top model as a premium pick "when quality matters and you're shipping every day" — conditional, not a default. The vendor that profits when you spend more is telling you to spend less. That should make you pause.
So: cheaper models are a fraction of the cost, the labs themselves say right-size your model, and Epoch AI finds the capability gap between cheap-and-frontier is only about four months wide. If the case for cheap is this strong, why do I still reach for the biggest model?
Because you don't pay per token. You pay per solved task.
The per-token price is the wrong number to optimize. What you actually care about is the cost of getting the job done correctly — and on that axis the gap collapses, sometimes reverses.
From Anthropic's Opus 4.5 announcement, verbatim:
"Set to a medium effort level, Opus 4.5 matches Sonnet 4.5's best score on SWE-bench Verified, but uses 76% fewer output tokens. At its highest effort level, Opus 4.5 exceeds Sonnet 4.5 performance by 4.3 percentage points—while using 48% fewer tokens."
Read that again. And note the pairing: Sonnet is the exact downgrade the playbook pushes you toward — "Opus is overkill, Sonnet is fine" — the same-line step down, not some cross-vendor swap. To reach the same quality, the "expensive" model burned roughly a quarter of the output tokens. Opus output costs 1.67× Sonnet per token — but at equal quality it emits ~76% fewer of them, so the output-token bill lands somewhere near 40% of Sonnet's for that task. The headline 5× and 1.67× ratios describe a world where both models write the same number of tokens. On hard problems, they don't.
This is a single-vendor benchmark (SWE-bench Verified), and Anthropic notes that Opus 4.7+ uses a new tokenizer that can consume up to 35% more tokens for the same text, which eats into the saving. It is not a free lunch. But the direction is clear and it is the opposite of the cost-dashboard intuition.
There's harder backing than one launch post. A 2025 economic-evaluation paper, Economic Evaluation of LLMs (Zellinger & Thomson), modeled this directly and concluded that "single large LLMs often outperform cascades when the cost of making a mistake is as low as $0.1," and that "practitioners should typically use the most powerful available model, rather than attempt to minimize AI deployment costs, since deployment costs are likely dwarfed by the economic impact of AI errors." The fair caveat: that recommendation is conditional on mistakes costing more than about ten cents, and it was tested on math problems. Below that threshold, cheap routing wins — and they say so. But ask yourself how often a wrong answer in your actual work costs less than ten cents to absorb.
The cheap model has to think harder — and you still pay for it
Here's the part the per-token sticker hides completely. Artificial Analysis runs every model through the same fixed battery — their Intelligence Index — and records how many output tokens each one burns to get through it, because the real bill is tokens × price and reasoning models differ wildly in how much they spew. Plot the score against those tokens:
Read it left to right. The most attractive corner is top-left: high score, few tokens. Who lives there? Capable models at reasonable effort. Now find the small models that tried to compete by thinking harder — Claude Sonnet 4.6 at max effort and GPT-5.4 mini at "xhigh". They're exiled to the right edge at ~195M and ~240M tokens — roughly double what Claude Opus 4.8 spends (110M) — and they still score lower (52 and 49 vs 61). A smaller model wound up to high reasoning isn't a bargain: it burns more tokens than the frontier model and arrives behind it. Claude Haiku 4.5 in reasoning mode tells the same story — 87M tokens, more than double the 36M average, to reach only 37.
And lining up different models here isn't a fudge — it is the real choice, because the cheaper option is usually a smaller and older model at the same time. What a cross-model chart can't isolate is the reasoning dial on its own, so here it is by itself: GPT-5.5 at every effort level — same brain, same questions, only the dial moving — with the cheaper GPT-5.4 mini and nano dropped in for scale:
Walk it from low to xhigh: the score climbs 51 → 57 → 59 → 60, while the cost to run the same evaluation climbs $501 → $1,199 → $2,159 → $3,357 (output tokens: 7M → 22M → 45M → 75M). The first step is the bargain — low to medium buys six points for about $700. Everything after is the trap: three more points for triple the bill. The cheaper siblings make the other half of the case: GPT-5.4 nano tops out at 44 even at xhigh, and GPT-5.4 mini — even cranked to xhigh at $1,354 a run — reaches only 49, below GPT-5.5 at its cheapest setting (51, for $501). You can't out-think your way up from a smaller brain. Anthropic's numbers say the same on Claude — medium effort already matches Sonnet's best score, and the highest setting adds just 4.3 points. So the dial is real, but its returns die fast.
This chart cuts both ways. At low effort, the cheap models are genuinely cheaper — Artificial Analysis priced the full Haiku run at about $583 and Flash even less, against $4,300+ for Opus at maximum effort. On a pure budget, used sensibly, the small model wins. But notice what it can't do: no reasoning setting lets Haiku or Flash reach Opus's 57–61. That ceiling is the model, not the dial. "Same quality, cheaper" only exists below the cheap model's ceiling — and the moment you need quality above it, spending the small model's tokens to chase it is the worst trade on the chart.
And in agents, small gaps explode
If you only ever fired single prompts, I'd have a weaker case. But the work is increasingly agentic — long chains of tool calls, each step depending on the last. That changes the math violently.
A task with n sequential steps succeeds only if every step succeeds — so per-step reliability compounds. Take two models a hand-picked five points apart: at 99% per-step reliability a 100-step task completes ~37% of the time; at 95% it completes 0.6% of the time. Those exact numbers are illustrative — I'm not claiming any particular model sits at 99% or 95% — but the shape is the point: a per-step gap that looks negligible on a leaderboard becomes a factor of sixty end-to-end.
The mechanism is documented, not invented. An ICLR 2026 paper, The Illusion of Diminishing Returns (Sinha et al.), shows that "marginal gains in single-step accuracy can compound into exponential improvements in the length of tasks a model can successfully complete," and that "larger models can correctly execute significantly more turns even when small models have near-perfect single-turn accuracy." METR's long-task measurements point the same way: capability is increasingly measured in how long a task a model can carry without falling over. The caveat: the ICLR headline result used a synthetic task, and error-recovery scaffolding can blunt the compounding — which is why I lean on the paper for the direction, not for a precise multiplier. But the direction is real, and it means the few extra points the frontier model scores are worth far more than the benchmark delta suggests — they are the difference between an agent that finishes and one that derails at step 40.
The hidden bill follows from this. When the cheap model derails, you pay the difference — in review time, in debugging its plausible-but-wrong output, in the retry. A token is fractions of a cent. A developer-minute spent untangling a confidently wrong refactor is not. That's the "cost of a mistake" the economics paper was talking about, and for knowledge work it is almost never under ten cents.
Reasoning effort: more thinking is not always better
"Use the biggest model" is not the same as "crank reasoning to maximum." This is where I break with the naive version of my own claim.
More reasoning effort genuinely helps on hard problems — that's how Opus 4.5 buys its accuracy. But it costs tokens, latency, and money, and past a point it can make answers worse. Anthropic researchers documented exactly this in Inverse Scaling in Test-Time Compute (Gema et al.): they built tasks where "extending the reasoning length of Large Reasoning Models deteriorates performance." They catalogued five failure modes — Claude models get distracted by irrelevant detail, OpenAI's o-series overfits to the framing, models drift from sensible priors to spurious correlations, all of them lose focus on long deductive chains, and extended reasoning can even amplify concerning behaviors. Longer thinking is a dial, not a virtue. Turn it up for genuinely hard reasoning; leave it low for the routine, where it just burns money and invites the model to overthink a simple thing into a wrong thing.
You can watch this in the wild — and it shows the model, not the dial, is what moves the needle most. SkateBench, an independent benchmark that has models define 390 technical skateboarding tricks, lines them up like this:
Read it as two within-line comparisons, because that's the choice you actually face — you pick a tier inside one provider's lineup, not across providers. In OpenAI's line: nudging the dial up (gpt-5 default → high) adds two points (75% → 77%), but dropping to the small sibling, gpt-5-mini, erases forty (38%). In Anthropic's line the same downgrade is even harsher — the cheaper tier, claude-4.6-sonnet, craters to 15% against the biggest model's 64%. Same lesson on both ladders: the dial is a small lever, the tier is the big one. Cost and latency confirm the trade rather than rescue the cheap picks: gpt-5-high is the most accurate but also the priciest ($9.57/run) and slowest (53.7s), while the cheap options are fast, cheap, and much worse. One note: this is a niche, single-author benchmark, so read the exact figures as directional. And I'm deliberately comparing each provider's biggest model to its own cheaper tiers — not OpenAI against Anthropic. That gpt-5-high outscores Opus here is beside the point; "reach for the biggest" means the biggest in the lineup you're already in, not a cross-vendor leaderboard.
There's a quieter point hiding in a comparison like that, and it's the reason mixing model generations is fair rather than sloppy: in real life, cheaper and older travel together. The discount tier is where last-generation models go to live out their days, and the model picker rarely makes that obvious. It tells you an option is "faster" or "efficient"; what it's actually handing you — which model, which generation, drawing from which budget — sits in a one-line description most people never read. So when you reach for the cheaper choice, you're often taking a quiet step backwards in time as well, and you never decided to. The downgrade you didn't notice you were making is exactly the one that costs you.
So where's the sweet spot? My answer: medium or high — mostly high — and almost never the ceiling. The top notch goes by different names — xhigh, max effort, "ultrathink" in Claude Code — and the GPT-5.5 curve already showed what it buys: a few points of intelligence for triple the cost. Worse, the inverse-scaling results say that past some point the extra thinking doesn't only cost more — it can drag accuracy down. High effort is the honest trade-off: most of the accuracy, a fraction of the runaway cost. The ceiling is for the rare problem that genuinely deserves it, not a default — and the dial behaves the same on a frontier model from any lab.
So the recipe I actually argue for is narrow and specific: the smartest model, on high — not the cheapest model on max. Capability comes from the brain you pick; you spend the reasoning budget to finish, not to compensate for a brain that was too small to begin with.
The counter-case
Here is the strongest case against everything I just argued — because all of it is true:
- The gap is narrow and shrinking. Anthropic quotes a partner saying Haiku 4.5 reaches ~90% of Sonnet 4.5 on agentic coding (a testimonial, not a clean benchmark — but still). Epoch puts the open-vs-frontier gap at about four months. For a lot of tasks, 90% is plenty.
- Routing works. RouteLLM and similar systems demonstrably cut cost while preserving most quality by sending only the hard queries to the big model. That is right-sizing done well — and it's exactly the regime where the economics paper says cheap wins.
- The frontier is getting more expensive to run, not less. A 2026 MIT FutureTech paper, The Price of Progress, finds that while price-for-fixed-performance falls ~5–10× a year, the cost of running the actual frontier is rising 3–18× a year, and warns that benchmarks "present a warped picture of progress in practical capabilities per dollar."
- At real-time, high-volume scale, latency and rate limits decide for you. If you're classifying ten thousand support tickets, Anthropic's own example runs them on Haiku for ~$37. Putting Opus on that job would be silly, and slow.
All of that is real. None of it is about complex, high-stakes, long-horizon work — which is most of what I actually do.
But you're probably not paying per token anyway
Now step back from the API price list, because most of us never see it. We reach these models through flat-rate subscriptions — Claude Code on a Pro or Max plan, Cursor, GitHub Copilot — paying a fixed monthly fee and spending against an allowance, not a per-token meter. That $5-versus-$1 gap the entire "use the cheap model" case is built on? Invisible at your desk. Choosing Opus over Haiku for that Postgres question costs you, in dollars, precisely nothing extra this month.
The catch — and it's a fair one — is that the platforms still ration the frontier; they've only swapped the unit from dollars to quota. The per-task economics from the charts above don't vanish here, they just change denomination: GitHub Copilot metered Opus 4.8 at a 15× "premium request" multiplier (before June's move to token-based credits), Cursor drops you to a cost-efficient "Auto" router once your credits are spent, and Claude Code defaults to Sonnet and only escalates to Opus past ~20% of your usage. So the same ratio that made the frontier look expensive per token reappears as a quota multiplier — except the provider now absorbs part of it through your flat fee, which if anything makes the frontier a better deal at your desk than the raw price list implies. The tools still nudge you down, because those multipliers optimize the vendor's bill, not your result. And you can override every one of them.
Once the dollar is off the table, the cheap model's real price tag comes into focus — and it was never tokens. It's trust. It's the answer you re-read because you're not sure it saw the whole file. The second prompt to make it reconsider the edge case it skipped. The half-right refactor that looked fine until it wasn't. None of that lands on a pricing page; it lands in your afternoon, and in how much you can lean on what the thing just told you. The bigger model's pitch is not that it's cheaper per token. It's that you believe it the first time.
So where's the line?
Here's the shape of the decision, and then I'll stop talking. The whole thing reduces to one question the cost dashboard never asks: what does it cost when the model is wrong?
So here's the rule, and it keeps the default where I started it — at the top. Reach for the biggest model first; step down only when a wrong answer is both cheap to catch and cheap to fix, and something underneath will catch it anyway — a one-shot classification, a draft you'll skim regardless, a high-volume pipeline with a human net. That's a real category, and there the cheap model isn't just acceptable, it's correct. But it's the exception you justify, not the default you assume. Everywhere else — the answer that feeds an agent forty steps deep, ships to production, or won't look wrong until it's expensive — the few points of reliability you'd "save" by stepping down were the only points that mattered. And watch where the throwaway question actually lands: the moment you re-read an answer because you're not sure, or paste it back an hour later, your own time has already blown past the pennies you saved. For interactive work that bar almost always favors staying at the top — which is why I default there even for the small stuff.
My opinion is not subtle, and I won't pretend the framing above is neutral. But I told you at the top: the last line is yours. You've now got both piles of facts. Where does your line sit?
One last thing — because this is where the argument really wants to go. Every chart here leaned on public benchmarks that test models in isolation. The comparison that would settle it runs inside the harnesses we actually work in: the same task, the same repo, handed to each model the tool puts in front of you — default versus the cheap fallback — in Claude Code, Copilot, Codex, and the rest. Same prompt, same harness, only the model swapped. That's the test that proves this instead of arguing it. It doesn't exist yet, so I'm going to build it. That's the next post.
All prices, model versions, and figures are accurate as of June 2026 (Opus 4.5–4.8, Sonnet 4.5/4.6, Haiku 4.5, GPT-5.4/5.5, Gemini 2.5) and will date quickly — every claim links to its primary source so you can check whether it still holds when you read this.