The 90% price cut is real. So is the fine print that eats it. Anthropic released Claude Haiku 5.5 this week, and on the surface it is the cleanest "prices are falling" story in the AI economy: up to 90 percent cheaper on the requests that make up about 90 percent of real traffic, a benchmark leap in computer use from 15.7 to 72.4 percent, and a first-in-class feature set (adjustable reasoning levels, a one-million-token context). But the two asterisks that follow the price table matter more than the price table. One: Haiku 5.5 ships a new tokenizer that consumes more tokens per task, the same quiet inflation Anthropic did with its Opus 4.x models. Two: at maximum effort, the model burns roughly 162,000 output tokens per task on an intelligence benchmark — about three times what GPT-6 Luna needs for a comparable result, which puts GPT-6 Luna ahead on the performance-per-dollar curve even though Haiku 5.5 wins the headline score. Read it that way, Haiku 5.5 is not the end of the pricing war. It's the most sophisticated round of it yet. Here's the full scoreboard.
The Headline Numbers
What Anthropic actually shipped, before the asterisks:
- Price: about 75 percent cheaper on average than Haiku 4.5, and up to 90 percent cheaper for prompts under 100,000 tokens — which the company says is roughly 90 percent of previous Haiku requests. Long prompts cost five times more. The input price for short prompts: $0.10 per million tokens versus $1.00 before. Output: $0.50 versus $5.00.
- Context: up from 200,000 tokens to one million.
- Reasoning levels: the first Haiku-class model with adjustable effort, so you can trade cost against quality on a per-task basis. Anthropic's own guidance: great for compaction, summarization, sub-agent work; for complex agentic coding, Sonnet 5.5 and Opus 5.5 remain the better picks.
- Benchmarks: OSWorld computer use jumps to 72.4 percent (from 15.7), Terminal-Bench agentic coding hits 39.2 percent (from zero), HLE with tools reaches 57.4 percent, and it beats OpenAI's budget model GPT-6 Luna in every head-to-head category Anthropic tested.
- The side deals: Sonnet 5.5 cache reads cut 50 percent (a move the company says saves most agentic workloads about 20 percent), and monthly API credits for subscribers — $100 to $500 per month depending on tier. The Sonnet cut is almost certainly a direct reaction to OpenAI's GPT-6.1 series.
Artificial Analysis, the independent benchmark platform, ranks Haiku 5.5 as the leading small-class model at 43 on its Intelligence Index, just ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41), and GPT-6 Luna (38) — with a 13-point gap to Claude Sonnet 5.5. That's the top of the small-model class, and it's a class that now includes Chinese and Korean players. The "small model" category is no longer a backwater; it's where the volume of the AI economy actually runs.
Asterisk One: The Tokenizer
Per-token prices are not per-task prices, and the gap between the two is the tokenizer. Haiku 5.5 uses an updated tokenizer that consumes slightly more tokens per task than Haiku 4.5, and Anthropic is explicit about it. It's not the first time: the Opus 4.x models had token usage jump about 30 percent from a tokenizer change alone, even with flat per-token pricing. The practical effect is that the "90 percent cheaper" headline is a per-token number, and your actual invoice is a per-task number. When the unit of billing changes in your favor but the quantity goes up, the net savings land somewhere in the middle. For a cost-sensitive deployment, this is the single most important line in the release — because the entire pitch of Haiku 5.5 is high-volume, cost-sensitive work. If the volume is the product, the quantity is the price.
Asterisk Two: The Burn Rate
The second asterisk is where the pricing war gets interesting. Artificial Analysis measured token consumption at the highest effort level: Haiku 5.5 uses about 162,000 output tokens per task, versus roughly 50,000 for GPT-6 Luna — about three times more. And the gap doesn't close at comparable performance: at the "high" effort setting, Haiku 5.5 scores 38 with about 55,000 tokens, while GPT-6 Luna reaches the same score with around 50,000. So on the performance-per-dollar curve — the thing that actually decides what you spend — OpenAI's budget model comes out ahead, despite losing the headline benchmark. The "effort" dial is a feature, and it's also a lever: turn it up and you buy intelligence with output tokens, and Haiku 5.5's top of that dial is expensive in quantity even if it's cheap in price. The model is a bargain. The setting is not.
What It Means
1. The pricing war moved from "cheaper tokens" to "cheaper tasks," and the new battleground is burn rate. Haiku 5.5 and GPT-6 Luna have the same conversation: the one with the lower per-token price loses the invoice, because the one with the higher effort dial burns three times the output. The next twelve months of AI pricing are going to be priced per task, not per token, and the number that matters is not the sticker but the quantity. That's a fundamentally different way to sell intelligence, and it's the reason the "90 percent off" headline and the "Luna is more efficient" result can both be true in the same release. The sticker is a marketing number. The burn is the P&L.
2. The small-model class is now a genuine three-region race, and the benchmark leader is on a Chinese open model's doorstep. GLM-5.3 Flash at 42, one point behind Haiku 5.5's 43, means the top of the small class is no longer "Anthropic versus OpenAI." It's Anthropic, OpenAI, Google, and a Chinese lab all within a four-point band. That's the same shape as the frontier race, compressed into the tier where the volume actually is. The strategic read: the tier that runs summarization, classification, support, and sub-agents is now as contested as the frontier, and the winner of that tier is going to own the cost structure of the entire agent economy. Haiku 5.5 leading by one point is not a moat. It's a race position.
3. The "admit I don't know" property is the quiet differentiator, and it's the one that will matter for agents. Artificial Analysis flags that Haiku 5.5 hallucinates in 40 percent of cases versus 77 percent for GPT-6 Luna, and that its weaker factual knowledge is partly offset by a stronger tendency to say it doesn't know. For a chatbot that's a nice-to-have. For an agent that is going to call tools, open databases, and act on its own confidence, "I don't know, here's what I do know" is the safety property. The agent era is going to price models on calibration, not just accuracy, because the cost of a confident wrong answer in an automated workflow is the cost of the action it triggered. That's a new line item that Haiku 5.5 is quietly setting.
4. The API credits and the Sonnet cut are the competitive tells, not the product. Anthropic cutting Sonnet cache reads in half the same week it shipped Haiku 5.5, and handing out $100 to $500 of monthly API credits to subscribers, is not a pricing event. It's a retention and a volume event. The company is buying two things with the same move: it's making the expensive model cheaper to run (lock-in) and it's giving its largest subscribers a budget to build agents on its stack (migration). The credits are a down payment on the agent economy. The Sonnet cut is the price of keeping the big workloads home while the GPT-6.1 series is circling. Read the two together and the release is not a product launch. It's a market-defense play with a small model attached.
🔥 Hot Takes
1. "Up to 90 percent cheaper" is the most overpriced marketing line in the release, and the tokenizer is the real price. Every cost-sensitive deployment should stop looking at the per-token sticker and start measuring per-task burn, because that's the number that lands on the invoice. Haiku 5.5's "90 percent" is a per-token claim in a world that bills per task, and the new tokenizer is the silent line item that moves the task number. The companies that price this model correctly are going to get a real discount. The ones that price it from the press release are going to find out in their next quarterly forecast. The sticker is a number. The tokenizer is the price.
2. The "effort" dial is the most important new feature in the release, and it's also the most dangerous one. A per-task quality-cost slider is the first time a model provider has put a dial directly on your bill. Turn it up and you buy intelligence. Turn it down and you buy speed. The feature is a genuine advance for cost control. The danger is that the dial makes the model's cost a variable the user sets, and when the user sets it, the user is also setting the error rate. The agent that's running at "max" effort at 162,000 tokens per task is not the same agent as the one at "low." The provider has just handed the cost of a mistake to the person who clicked the slider. That's not a feature. That's a liability transfer, and the next "why did our agent spend $40,000 in a week" incident is going to be a story about the effort dial.
3. GLM-5.3 Flash at one point behind is the line the whole press release should be focused on, and it's the line nobody's going to make the headline. A Chinese open model within four points of the top of the small-model class is not a "competition" story. It's a "the tier is global" story. The entire agent economy — the summarizers, the classifiers, the support bots, the sub-agents — is going to be built on the model that owns the small tier, and that tier now has four serious players in a four-point band. Haiku 5.5 leading by one point against a Chinese open model is not a win. It's a position in a race where the gap is the noise. The next release cycle is going to decide the tier, and the tier is going to decide the cost of the next decade of AI. The small model is where the volume is. The volume is where the economics are. The economics are where the war is.
The Bottom Line
Claude Haiku 5.5 is the cleanest "prices are falling" launch in months and the most complicated one at the same time. The 90 percent cut is real on the sticker. The tokenizer inflates the quantity. The effort dial moves the burn from 50,000 to 162,000 tokens per task. And one point behind it, a Chinese open model is sitting in the same band. The pricing war did not end with this release. It got its first honest per-task scoreboard. The next round is going to be priced in burn, not in tokens, and the winner of the small tier — the tier that runs the volume of the agent economy — is still up for grabs. Haiku 5.5 is at the top of the board right now. One point is not a moat. It's a starting line.