Enterprise

Anthropic cuts Claude Haiku to $0.10 per million tokens, resetting the floor for small-team AI tools

Claude Haiku 5.5 lists 90% below its predecessor for short requests, bringing live support agents and high-volume lead work into reach for lean sales and marketing stacks.

Photo: Wikimedia Commons / Anthropic, CC0 1.0 Public Domain — The Claude AI symbol, the logo used by Anthropic for its Claude model family

Anthropic released Claude Haiku 5.5 on October 7, 2026, pricing short requests at $0.10 per million input tokens and $0.50 per million output tokens, a roughly 90% reduction from Haiku 4.5’s $1.00 and $5.00. Cache reads in the lower tier drop to $0.01 per million tokens, down from $0.10. The new price exactly matches OpenAI’s GPT-6 Luna, released weeks earlier at identical rates.

The pricing has a hard edge. Below 100,000 tokens per request, those headline rates apply; above it, input jumps to $0.50 and output to $2.50, a 5x step and a 50% cut from Haiku 4.5 at that tier. Anthropic says roughly 90% of Haiku traffic sits under the threshold, so most workloads see the full cut. Factoring in request sizes and tokenizer changes, the company estimates a 75% workload cost reduction versus Haiku 4.5. Anthropic halved Sonnet 5.5 cache reads to $0.10 per million tokens the same day, trimming an estimated 20% off most agentic Sonnet workloads.

The customer evidence Anthropic published alongside the release is unusually specific. Ze’ev Klapow, a distinguished software engineer at HubSpot, reported that Haiku 5.5 “achieved the best score HubSpot had seen on its simulated CRM task suite, at 92.8% averaged over three runs, and was the fastest model tested on a CRM audit task, with the highest hit rate and the lowest false positive rate.” Box’s VP of AI Products, Yashodha Bhavnani, reported an 11-point improvement over Haiku 4.5 with approximately half the latency. AlphaSense’s Daniel Campos said Haiku 5.5 scored 0.84 versus 0.76 for Haiku 4.5 across 400 production-style queries for Ask in Document, a workload handling about 8 million calls a week. Asana’s Aaron Vinh reported “over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn” in its AI Teammates evaluation suite.

Benchmarks published by Anthropic place Haiku 5.5 at 1,620 on GDPval-AA v2.1 against Sonnet 5.5’s 1,840, and at 72.4% on the OSWorld 2.1 offline subset versus Sonnet’s 83.9%. Terminal-Bench reaches about 39.2% at maximum effort.

The structural point is the floor. When OpenAI priced GPT-6 Luna at these rates in September, it was a unilateral move. Anthropic matching it within weeks, during the run-up to the IPO signaled by its $100B run-rate, confirms a new equilibrium. Live support agents, outbound lead classification, and CRM enrichment at scale are no longer gated by inference economics. They’re gated by whether anyone has built the product yet.

Sources