groundy
Agents & Frameworks

When Claude Code Needs Opus 5.5, and When It Just Burns Tokens

Vendor data and independent studies show Opus 5.5 gains apply to long tasks, while prompt fixes and harness changes often yield larger coding-agent improvements than model up.

Published 9 references
A translucent green resin dinosaur examines a yellow connector above a long segmented assembly on an ivory surface, with hard shadows and visible casting bubbles.
On this page10 sections

If your Claude Code bill is climbing, the model selector is probably not your biggest problem, and upgrading everything to Opus 5.5 is probably not your fix. Anthropic’s own usage guidance for Opus 5.5, published alongside the model’s September 22, 2026 launch, quietly makes that case: the vendor’s claimed gains concentrate on long, multi-step autonomous work, while some of the ways teams spend more on Opus 5.5 add zero capability. Meanwhile, independent research says the variables that actually move coding-agent outcomes are often the prompt, the harness, and whether anyone rechecks the agent’s intermediate conclusions, none of which requires the higher tier to work.

The practical rule that falls out of the evidence: route Opus 5.5 to long-horizon tasks with a written finish line and checkpoints that survive context summarization, keep cheaper tiers for interactive and well-specified work, and fix your prompts and CLAUDE.md before you touch the model dropdown. What follows is the reasoning, including the parts Anthropic’s guide does not cover.

What Anthropic actually claims Opus 5.5 changes in Claude Code

Start with what the vendor says, clearly labeled as vendor-sourced. Anthropic’s launch page claims Opus 5.5 “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5”. A customer testimonial Anthropic published, from a company that tests models on real engineering and trading-desk work, reports that on their agentic coding tasks the new model “matched Opus 5’s quality in about half the turns, time and output tokens, cutting the cost of that workload by 40 to 50%” (launch page). Both figures are Anthropic-published statements, not audited spend data, and I have not seen an independent replication as of this writing.

The more useful vendor statement is about where the gains sit. According to the official guide, “Compared to prior Opus models, its biggest gains are on multi-step work, like carrying a change through a large repository until the tests pass. Early testers had it run long coding tasks for hours with little oversight.” The hours-long autonomy claim is an early-tester anecdote the vendor chose to publish; treat it as a hypothesis about task shape, not a measured result.

That framing is still informative, because it tells you where Anthropic itself thinks the model earns its keep: tasks that chain many steps, hold state across a long run, and end at a verifiable stopping condition. The guide’s concrete recommendations follow that shape:

  • Whole-task handoff with an explicit finish line. Describe the end state (tests pass, migration complete) rather than narrating each step.
  • Subagent splits for audits, migrations, and reviews across a large codebase, with each subagent’s result checked.
  • A file-based task list (the guide suggests something like TASKS.md) because long runs fill the context window, at which point Claude Code summarizes older turns; a list on disk survives that summarization.
  • Keep-going rules paired with human gates. The guide warns: “A rule to keep going means fewer stops, so keep your own check before anything risky or hard to undo,” and recommends keeping permission prompts on for destructive commands. That parallels the permission controls on Anthropic’s product pages, which note Claude “asks before permanently deleting files” by default.

There is also one attention-grabbing anecdote worth flagging precisely because it is weak evidence: a single early tester cited in the guide claimed Opus 5.5 at its lowest effort caught more bugs than Opus 5 at high effort, with fewer false alarms. One testimonial is not a routing policy.

Anthropic’s own numbers come with a noise floor the company deserves credit for disclosing. On its Terminal-Bench-Science 0.1 eval, the launch page states a standard error of ±3.5–5 points per model, and notes its reproduction of a public leaderboard score for Opus 5 (29.0% versus the reported 30.0%) is “within noise.” Practical consequence: a one-to-three-point gap between two Claude tiers on a public benchmark is not routing signal, and no one should be re-tiering their team over it.

Where the frontier model earns its cost, and where it does not

Combining the vendor’s task-shape framing with the independent evidence below, the routing heuristic I would put in front of an engineering lead is:

  • Route up when the task is long-horizon and multi-step (a change carried through a large repository until tests pass, a migration, a broad audit), when you can write down a finish line, and when you have a way to checkpoint progress outside the context window.
  • Stay on a cheaper tier for interactive back-and-forth, well-specified single-file work, and anything where a human is steering each turn anyway. The autonomy advantage the vendor claims does nothing for you if you are the autonomy.
  • Do not route up to fix a bad prompt or a missing instruction. The evidence for that ordering is strong, and it is the next section’s subject.

This is a heuristic, not a measured crossover point. No independent Opus 5.5 benchmark appears in this article’s sources; the independent studies tested Opus 5, 4.8, and 4.6. Any team making a real routing decision should run its own A/B on a sample of its own tasks and measure cost per completed task, because that is the metric the next sections argue actually matters.

Token burn by default: fast mode, thinking boilerplate, reasoning dumps

The most quotable thing in Anthropic’s guide is the part where the vendor tells you to stop paying for things. Three spend items are called out, explicitly or by structure, as non-capability costs:

Fast mode. The guide describes it plainly: “Fast mode is available for Opus 5.5 at launch as a research preview. You get the same model, and the text arrives sooner. It needs extra usage turned on, and it costs more per token than standard mode.” Same model, higher price, lower latency. The launch page lists standard mode at $4 per million input and $20 per million output, with cache reads at $0.20 per million. Fast mode is exactly double, $8 and $40 per million (same page). That is a purchase of human wait time, not quality. It is worth buying when a developer is sitting idle waiting on the agent and their loaded hourly cost dwarfs the token premium; it is waste in a CI pipeline, an overnight batch, or any unattended run where nobody is watching the cursor. Framing fast mode as “smarter” would be wrong, and the vendor does not claim it.

Thinking boilerplate. Per the same guide: “Delete “think carefully” lines. Opus 5.5 already thinks before every reply.” Prompt rituals carried over from older models, the “think hard” and “think step by step” lines pasted into CLAUDE.md files and saved instructions, are now redundant on this model and cost tokens every time they are sent. One caveat the guide implies but teams should state explicitly: this applies to Opus 5.5. If older models are still in your rotation, stripping those lines from shared instructions could degrade them.

Reasoning dumps. Requests to reproduce the model’s internal reasoning in the reply can be declined and fall into one of the flag categories, according to the guide, which recommends asking instead for a short explanation of the approach. If your team has prompts demanding full chain-of-thought output, they are generating refusals or spend without better code.

None of these three require a model decision. They are hygiene, and they are cheap to fix this week.

What the guidance omits: unverified intermediate conclusions

Here is where the vendor checklist runs out and the independent evidence starts. Anthropic’s guide covers stop rules, task files, and checking subagent results. It does not address the failure mode that a five-week account of running Claude Opus 5 through Claude Code documents in detail: wrong intermediate conclusions.

In that author-reported project, the authors gave a coding agent (Opus 5 with shell access) hard search problems in coding theory. The project produced real results, including a new lower bound, but the authors describe a recurring pattern: “Each time, an intermediate result was written down, never rechecked, and treated as a fact that ruled out further search. Checking final outputs, as our protocol required, does not catch such errors.” One wrong intermediate conclusion hid a genuine improvement for three weeks. And the decisive result, per the paper, “came from a second agent session with no shared context that audited our repository.”

This is one project, one domain, and a prior model, so do not over-extrapolate. But the mechanism generalizes in a way that matters for the routing decision: the failure was not a capability gap that a bigger model would have fixed. The agent was already the frontier tier. What was missing was a process layer, a re-verification step for the conclusions the agent writes down and then builds on, plus a fresh-session audit that cannot inherit the first session’s assumptions. A model upgrade does not add either of those. You add them.

This is also the expensive version of the problem. A long unattended Opus run that builds on a wrong intermediate conclusion does not just produce a wrong answer; it produces hours of confidently wrong work at frontier prices. The routing rule “send long tasks to Opus 5.5” is only safe if it travels with “and re-verify what it concluded along the way, not just what it shipped.”

When a better prompt or harness beats a bigger model

The strongest counter-evidence to route-up-by-default comes from controlled work showing how much non-model variables move outcomes.

A study of SOC 2 compliance in generated code across Claude Fable 5, Opus 4.8, and Opus 5 ran 24 generations across four use cases and found that “The single SOC 2 sentence moved every case to 86–100%, worth 23 to 50 points, and removed every insecure construction” (same study). One sentence naming the compliance standard. That effect, 23 to 50 rubric points within one frontier Claude family, is larger than any model-tier delta cited anywhere in this research, and it cost nothing. Important scope limits: the authors themselves restrict the model-comparison claim to a single vendor’s family, say it “needs narrowing even there,” and list cross-vendor replication as the main limit on generalization. So the finding is not “prompts beat models universally.” It is “within one family, one prompt line swamped the tier difference on this task type.”

A second study complicates the prompt-fix story in a useful way. In a case study on Claude Opus 4.6 deobfuscating JavaScript with poisoned identifier names, the authors report that explicit verification instructions, “verify each name matches the math,” failed to stop wrong names from propagating in 12 out of 12 runs. What changed behavior was reframing the task as fresh generation rather than deobfuscation. The lesson is narrower than “prompt harder”: sometimes prompting harder at the same task framing does nothing, and the fix is to change what the agent is doing, not how sternly you ask. That study ran partly in Claude Code CLI under a Max subscription, covers two code archetypes, and its authors disclaim generality.

Zoom out one more level and a 314-page monograph synthesizing 164 scholarly works and 100 practitioner records on coding-agent reliability concludes: “Across this evidence, many apparent model failures originate elsewhere in the system, while improvements at one layer often fail to propagate to end-to-end outcomes.” Harness, execution state, retrieval, memory, permissions, review interfaces: the authors argue the system around the model is where much of the reliability lives, and the records underlying that synthesis vary in strength, so treat it as a map of where to look rather than a set of measured effects.

The harness layer also has concrete, measured wins. AtomicCommitBench, which tests whether agents can reconstruct clean commit histories, found agents typically return one squashed patch mixing features, fixes, and tests. Structured evidence, what the authors call DACE, improved exactly the lower-scoring setups (MiniMax +0.075 ARI, Kimi +0.049), according to the paper, by making structural cues explicit. A harness change lifted the weaker configurations without any model swap.

The pattern across these: before you pay for a tier, pay attention to the instruction and the scaffolding. They are cheaper levers, and in this evidence they are often the binding ones.

Cost per finished task, not per token

The accounting mistake that makes model routing feel urgent is watching per-token prices. A position paper compiling listed API prices across major providers as of early 2026 found posted prices “span over an order of magnitude,” without quality normalization, which is precisely why per-token price is a weak proxy for what a task costs. A cheaper model that needs three attempts and a human cleanup pass can cost more per finished task than a frontier model that lands it in one. Conversely, the Anthropic-published testimonial claiming 40 to 50% workload cost cuts, on the Opus launch page, is a claim about turns, time, and output tokens per completed task, which is the right unit. If you audit anything, audit that: pick twenty representative tasks, run them on both tiers, and compare cost per accepted result, not the rate card.

The decision table, and where the evidence runs out

SituationRouteWhy
Multi-step change across a large repo, finish line is written down, checkpoints in a task fileOpus 5.5Vendor’s claimed gains concentrate here; failure risk is manageable if intermediate conclusions get rechecked
Interactive pairing, human steering each turnCheaper tierThe autonomy advantage has no room to act
Well-specified small task (single file, clear spec)Cheaper tierLong-horizon capability is irrelevant to task shape
Agent is idle-waiting a developerFast mode is a latency purchase, not qualitySame model at a higher per-token price, per the vendor
Unattended or CI runsStandard mode, no fast modeNobody benefits from lower latency
Output quality is poor on any tierFix prompt, CLAUDE.md, task framing, or harness firstOne sentence moved compliance 23–50 points; reframing fixed what verification prompting could not
Choosing between tiers on a 1–3 point benchmark gapDo not route on itOn Terminal-Bench-Science 0.1, the vendor’s disclosed standard error is ±3.5–5 points

The limits deserve equal weight with the table. Every Opus 5.5-specific number in this article is vendor-sourced, from the Opus launch page and its published testimonials: the 40%-cheaper-to-run claim, the 40–50% testimonial figures, the hours-long autonomy story, the low-effort bug-catching anecdote. Among the sources cited here, no independent Opus 5.5 benchmark or measured spend comparison appears. The independent studies cover earlier models in the same family, and the single strongest prompt-over-tier result is confined to one vendor’s family by its own authors.

What the evidence does support, firmly, is the order of operations. Write the finish line. Put the task list in a file. Delete the thinking boilerplate from shared instructions. Re-verify intermediate conclusions on anything long-running, ideally with a fresh session that cannot inherit the first one’s assumptions. Then, and only then, decide whether the remaining gap is a model-tier gap. For a lot of teams, that sequence will shrink the Opus 5.5 bill more than any routing policy would, because it removes the spend that was never buying capability in the first place.

Frequently Asked Questions

What should teams do to avoid paying for redundant thinking instructions?

Per the same guide: “Delete “think carefully” lines. Opus 5.5 already thinks before every reply.” Prompt rituals carried over from older models, the “think hard” and “think step by step” lines pasted into CLAUDE.md files and saved instructions, are now redundant on this model and cost tokens every time they are sent.

Why is per-token price a weak proxy for task cost?

A position paper compiling listed API prices across major providers as of early 2026 found posted prices “span over an order of magnitude,” without quality normalization, which is precisely why per-token price is a weak proxy for what a task costs. A cheaper model that needs three attempts and a human cleanup pass can cost more per finished task than a frontier model that lands it in one.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Getting the most out of Opus 5.5claude.devAccessed
  2. Anthropic Claude Opus Launch Pageanthropic.comAccessed
  3. Running Claude Opus 5 through Claude Codearxiv.orgAccessed
  4. SOC 2 Compliance in Generated Codearxiv.orgAccessed
  5. Claude Opus 4.6 Deobfuscation Case Studyarxiv.orgAccessed
  6. Coding-Agent Reliability Monographarxiv.orgAccessed
  7. AtomicCommitBencharxiv.orgAccessed
  8. API Price Position Paperarxiv.orgAccessed
  9. Anthropic Product Overviewclaude.comAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy