On February 7, 2026, Anthropic shipped a research preview that charged six times the normal rate for faster Claude Opus output: $30 per million input tokens and $150 per million output against the standard $5/$25, promising up to 2.5x faster responses for interactive coding (Simon Willison). At that multiplier, the speedup had to justify a 500% premium on every token. Today you cannot buy the 6x tier at all. Fast mode for Opus 4.7 was deprecated on June 25, 2026 and removed on July 24, 2026 (Claude Code docs), and Opus 4.6 requests asking for fast speed now run at standard speed and standard rates (Anthropic’s API docs). What remains is fast mode on Opus 5.5, Opus 5, and Opus 4.8, each at a 2x multiplier. Neither model that carried the 6x rate supports fast mode anymore, so the premium you can actually buy is two-thirds smaller.
The Price of Speed
At launch the multiplier was a flat 6x: $30/$150 per million tokens against Opus 4.6’s $5/$25, with a 50% promotional discount through February 16, 2026 that brought the effective rate to 3x for the first ten days (Willison). Willison also flagged the worst case: past 200k input tokens, the 1M-context surcharges stacked on the fast rate to reach $60 input and $225 output per million. That surcharge is gone. Current documentation prices fast mode flat across the full 1M token context window, including requests over 200k input tokens (pricing page, fast mode docs).
The current rate card, per Anthropic’s pricing page:
| Model | Standard (per MTok, in/out) | Fast (per MTok, in/out) | Premium |
|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | $8 / $40 | 2x |
| Claude Opus 5 | $5 / $25 | $10 / $50 | 2x |
| Claude Opus 4.8 | $5 / $25 | $10 / $50 | 2x |
| Claude Opus 4.7 | $5 / $25 | returns an error | not available |
| Claude Opus 4.6 | $5 / $25 | standard speed, standard rates | none |
The two retired rows behave differently. Opus 4.7 rejects fast requests outright, with no fallback to standard speed (fast mode docs). Opus 4.6 accepts them, silently answers at standard speed, and reports usage.speed: "standard" in the response, so API users can verify what they were billed for (fast mode docs).
Two modifiers stack on top of these rates: prompt caching multipliers and the 1.1x data residency multiplier for US-only inference (pricing page).
One billing detail trips up subscription users: on Pro, Max, Team, and Enterprise plans, fast mode draws from usage credits and is not included in plan limits (Claude Code docs). Until credits are turned on, /fast reports “Fast mode requires usage credits.” Pro and Max users enable them under Settings > Usage on claude.ai; on Team and Enterprise, a member with billing access enables them for the organization. Console organizations pay per token with their other API usage, but must have fast mode access provisioned while it stays in research preview (Claude Code docs).
What You’re Actually Getting
Fast mode is not a different model. The Claude Code documentation describes it as “a different API configuration that prioritizes speed over cost efficiency,” with “identical quality and capabilities.” The documented gain is “up to 2.5x higher output tokens per second,” a ceiling rather than a guarantee, and it applies to output tokens per second (OTPS), not time to first token (fast mode docs). Willison reported at launch that Anthropic’s teams had “been building with a 2.5x-faster version of Claude Opus 4.6” internally (Willison). Since long code responses spend most of their wall-clock time generating output rather than waiting for the first token, output throughput is the number interactive users actually feel.
The model behind the toggle has changed with Claude Code releases. Fast mode launched on Opus 4.6 at the 6x rate (Willison); in Claude Code it then defaulted to Opus 4.7 from v2.1.142 through v2.1.153, to Opus 4.8 from v2.1.154 through v2.1.218, to Opus 5 from v2.1.219, and to Opus 5.5, the current default, from v2.1.280 (Claude Code docs). The pricing trail is only partly dated. Vercel’s AI Gateway added fast mode for Opus 4.7 on May 12, 2026 at $30/$150, still the 6x multiplier (changelog), and Anthropic deprecated that model’s fast mode on June 25, 2026 (Claude Code docs). The current rate card prices every supported model at 2x (pricing page); when the multiplier changed is not dated in the documentation, so pinning it to a specific Claude Code release would be guesswork.
Two cost details sit underneath the multiplier. First, the 4.6-to-4.7 transition changed tokenization: Opus 4.7 and later models use a tokenizer that maps the same text to approximately 30% more tokens (pricing page), which Anthropic’s launch announcement put at roughly 1.0 to 1.35x depending on content (Opus 4.7 announcement). Per-session bills moved between those models even where list prices did not. Second, Opus 5.5’s standard rate is $4/$20, below the $5/$25 of Opus 5 and 4.8, which makes its $8/$40 fast rate the cheapest fast option, 20% below the $10/$50 charged on the other two (pricing page). Among today’s supported models the tokenizer is uniform, so the remaining cost lever is which Opus you pin.
When you disable fast mode with /fast, you stay on Opus; /model is the command for switching models. Switching to Opus 4.7 now turns fast mode off, because that model no longer supports it (Claude Code docs).
When Fast Mode Makes Sense
Anthropic’s documentation names three cases where fast mode earns its keep (Claude Code docs):
Rapid iteration on code changes: when you’re trying approaches back to back, wait time fragments attention. Output arriving up to 2.5x faster keeps the loop tight.
Live debugging sessions: latency matters most when you’re watching output as it streams. On a long response, a 2.5x output rate can turn a ten-second wait into roughly four, which is the difference between staying inside the problem and context-switching away from it.
Time-sensitive work with tight deadlines: if response latency is your bottleneck, cutting it by more than half compresses interactive work. At 2x per token, this case is far easier to justify than it was at launch, when the same speedup cost six times the price.
When to Stick With Standard Mode
Long autonomous tasks: you’re not watching the responses, so latency you don’t experience is a premium paid for nothing (Claude Code docs).
Batch processing or CI/CD pipelines: automated workflows gain nothing from latency, and the Batch API does not offer fast mode at all (fast mode docs). The async path points the other way entirely: batch requests carry a 50% discount on both input and output (pricing page).
Cost-sensitive workloads: standard mode is the same model with the same outputs; the only thing the premium buys is shorter waits.
Infrastructure on third-party clouds: fast mode is not available on Amazon Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS (fast mode docs, Claude Code docs). Gateway routes that existed were model-specific: Vercel AI Gateway added fast mode for Opus 4.6 on April 7, 2026 (changelog) and for Opus 4.7 on May 12, 2026 (changelog), both at the 6x rate. Anthropic has since made fast mode unavailable on both of those models, so gateway availability is worth re-checking per model rather than treating as a standing feature. For comparison on what premium Opus access costs elsewhere, GitHub Copilot carries Opus 4.7 at a 15x premium request multiplier (changelog).
Claude Managed Agents: speed is part of the agent’s model configuration, set at creation by passing model as an object such as {"id": "claude-opus-5", "speed": "fast"} (agent setup docs). The same three supported models apply (fast mode docs).
Priority Tier: fast mode is not available with a Priority Tier commitment (fast mode docs). Those commitments are no longer available for purchase, and organizations holding one keep it through their contract end date (service tiers).
Team or Enterprise without owner opt-in: fast mode is disabled by default. An Owner must enable it, at Admin Settings > Claude Code for Claude AI organizations, or in Claude Code preferences plus provisioned access for Console organizations. The error is unambiguous: “Fast mode has been disabled by your organization.” (Claude Code docs).
The Hidden Cost: Mid-Conversation Switching
The first time you enable fast mode in a conversation, you pay the full fast uncached input price for the entire existing context (Claude Code docs). At 50,000 tokens of accumulated context, that one-time re-bill is $0.50 at the $10 per million fast input rate on Opus 5 and 4.8 (pricing page), or $0.40 at Opus 5.5’s $8 rate, on top of the ongoing premium for new tokens.
The re-bill applies once per conversation, not on every toggle: “toggling fast mode off and on again later does not repeat it” (Claude Code docs). What a toggle does cost you is cached-prefix pricing. Requests at different speeds do not share cached prefixes, so every switch forfeits the prompt cache until it rebuilds (fast mode docs).
The documentation’s advice stands: “For the best cost efficiency, enable fast mode at the start of a session rather than switching mid-conversation” (Claude Code docs).
Optimization Strategies
-
Toggle at session boundaries.
/fastin the CLI, or the Toggle fast mode command in the VS Code extension, turns it on. Because the context re-bill is one-time but the cache split happens on every switch, the cheapest toggle is the one made before context accumulates (Claude Code docs). -
Stack with lower effort where quality allows. Anthropic notes fast mode can combine with lower effort levels for “maximum speed on straightforward tasks” (Claude Code docs); the tradeoff is less thinking time and potentially lower quality on complex work.
-
Know the rate limit pool. Fast mode has a dedicated limit separate from standard Opus, and all supported Opus models share one pool. When you hit it, Claude Code falls back to standard speed and pricing, the ↯ icon grays out, and fast mode re-enables when the cooldown expires (Claude Code docs). At the API, exceeded limits return a 429 with a
retry-afterheader (fast mode docs). -
Watch usage credits, not just limits. If credits run out mid-session, Claude Code retries at standard speed and pricing and turns fast mode off for the rest of the session; run
/fastagain after topping up (Claude Code docs). -
Per-session opt-in for organizations.
fastModePerSessionOptIn: truein managed settings makes every session start with fast mode off, which keeps multiple concurrent sessions from drawing usage credits by default (Claude Code docs).
ROI Calculation Framework
Compare your hourly rate to the premium. The break-even test is unchanged in shape: if the dollar value of the time saved exceeds the extra token spend, fast mode pays. What changed is the size of the spend, 2x per token rather than 6x, on a default model whose fast rate is $8/$40 (pricing page).
Measure your bottleneck. Time a few standard-mode sessions and estimate what fraction of your cycle is spent waiting on responses versus thinking, typing, and reviewing. If waiting accounts for under a fifth of your time, faster tokens won’t convert into finished work.
The Bigger Picture
The most instructive part of this pricing experiment is its ending. The 6x rate is documented from launch on February 7, 2026 through Vercel’s May 12, 2026 snapshot, which still listed Opus 4.7 fast mode at $30/$150 (Vercel changelog); fast mode for Opus 4.7 was deprecated on June 25 and removed on July 24, and every model on the current rate card sits at 2x. The dated record shows that the multiplier changed, not when. Latency is now a priced dimension of the API in the way throughput tiers are priced in cloud compute, but at a multiplier close enough to the speedup (2x price for up to 2.5x output rate) that the trade is arguable per session rather than absurd on its face. The operational lesson matters as much: the economics of /fast are version-dependent. Anthropic has moved the model behind the toggle across releases, and the premium itself did not stay at 6x. Budget against the model you pin, not the toggle.
The Verdict
Is 6x pricing worth it? The question has expired: no model you can select today carries that multiplier. Fast mode for Opus 4.6 silently downgrades to standard speed, and fast mode for Opus 4.7 was removed on July 24, 2026 (fast mode docs, Claude Code docs). The live premium is 2x on Opus 5.5, Opus 5, and Opus 4.8, with Opus 5.5 the Claude Code default and the cheapest of the three (pricing page).
At 2x, the case for interactive work is straightforward: same model, identical output, up to 2.5x faster generation, double the per-token price. For debugging and rapid iteration where you watch output arrive, that is a defensible trade at a professional billing rate. For automated workflows, batch jobs, or anything you’re not actively watching, it remains money spent on latency nobody experiences, and the Batch API’s 50% discount points the other way.
Willison called the launch tier “Anthropic’s fastest best model” (Willison). The $60/$225 extended-context rate he flagged is gone, and so is the 6x multiplier that prompted the question. The pragmatic approach survives unchanged: enable fast mode at the start of your most latency-bound sessions, confirm which model sits behind it, and measure whether the wait you saved was worth the premium you paid.
Join the discussion
Share a useful perspective or ask a question about this article.