
UpTrajectory Review
Anthropic has released Claude Haiku 5.5, its newest small model, at roughly 25% of what Haiku 4.5 cost to run — a 75% price cut in a single generation. The company also halved the cache-read price on Claude Sonnet 5.5, which matters because cached context is where a lot of real-world API spend concentrates. The launch comes just two weeks after Opus 5.5 debuted on Sept. 22, completing Anthropic's 5.5 family refresh in rapid succession. For operators who have been treating frontier-model pricing as a fixed cost of doing business, this is a signal that the floor on inference pricing is still falling fast — and that the 'small model' tier is now where the most aggressive economics live.
If you run a small business that touches AI in any way — customer support triage, drafting, classification, extraction, internal search — Haiku-tier pricing is the line item to re-examine this quarter. A 75% cut means workloads you priced out six months ago as marginal may now be comfortably profitable. The Sonnet cache-read halving is arguably just as significant for anyone running agents or long-context workflows: if you're repeatedly sending the same system prompts, document sets, or conversation history, cache reads are often the dominant cost. Anthropic is effectively repricing the 'memory' of your AI stack, not just the compute.
What stands out here is the speed and sequencing. Anthropic didn't just launch a cheaper model — it cut prices on an existing model (Sonnet 5.5) that is barely a generation old, which suggests competitive pressure from OpenAI's GPT-5 family and Google's Gemini lineup is forcing faster-than-usual price compression. We're somewhat skeptical of the framing that this is purely generosity or efficiency gains; it's also market-share defense. But the practical effect is real regardless of motive. The under-reported angle is what this does to smaller model providers and open-source alternatives, which just lost a chunk of their price advantage.
Second-order effects cut in a few directions. Developers who built around Haiku 4.5 should see immediate margin relief, but anyone who locked in long-term contracts or committed spend based on older pricing may want to renegotiate. The bigger shift is architectural: as small models get this cheap, the calculus changes from 'one big model for everything' to 'route cheap tasks to Haiku, reserve Opus for hard problems.' That routing discipline is where the savings actually materialize — and where sloppy implementations leave money on the table. Expect your AI vendors and consultants to start quoting new baselines.
Watch whether OpenAI and Google match these cuts within the next few weeks, because the last time Anthropic moved this aggressively on pricing, the rest of the market followed within a month. Also watch for benchmarks: the real question is whether Haiku 5.5 at a quarter of the price matches Haiku 4.5 on the tasks you actually run, or whether the savings come with capability trade-offs that push work back up to Sonnet. If you're an operator, the actionable move this week is simple: pull your last 90 days of API invoices, identify your Haiku and cache-read spend, and model what the new pricing does to your unit economics before your next billing cycle closes.
“Anthropic PBC today released Claude Haiku 5.5, pricing its newest small model at roughly a quarter of what Haiku 4.5 costs to run.” — SiliconAngle
Takeaway: Audit your last 90 days of AI API spend now — Haiku 5.5's 75% price cut and Sonnet's halved cache-read costs could materially change your unit economics this quarter.
Excerpt from the original — SiliconAngle
Anthropic PBC today released Claude Haiku 5.5, pricing its newest small model at roughly a quarter of what Haiku 4.5 costs to run. Claude Sonnet 5.5 is getting cheaper as well, since the company is halving what it charges for cache reads on that model. Two weeks after Opus 5.5 launched Sept. 22, Haiku 5.5 […]
The post Anthropic releases Claude Haiku 5.5 small model and halves Sonnet 5.5 cache read prices appeared first on SiliconANGLE.