Claude Haiku 5.5 matches GPT-6 Luna's price, but hides a cost increase in the tokenizer
Anthropic launched Claude Haiku 5.5 on October 7, charging the same price as OpenAI's GPT-6 Luna up to 100,000 tokens. But a less generous tokenizer and a price jump above that limit change the math for anyone running agents in production.
Anthropic launched Claude Haiku 5.5 on October 7, 2026, the new generation of its fast and cheap model, according to a report by developer Simon Willison on his blog. The announcement was already expected: Willison had commented days earlier that the Haiku family was outdated, and Haiku 4.5 (launched almost a year earlier, at $1 per million input tokens and $5 per million output tokens) had become too expensive compared to more recent competition.
The direct comparison is with GPT-6 Luna, OpenAI's model launched the previous month. Up to 100,000 tokens, Haiku 5.5 charges exactly the same price as Luna: $0.10 per million input tokens and $0.50 per million output tokens. That's a 10x drop compared to Haiku 4.5, and according to Willison, the new model also reports higher benchmarks in that usage range.
The detail that changes the math above 100,000 tokens
The price parity doesn't last long. Past the 100,000-token limit, Haiku 5.5's price jumps 5x, to $0.50/$2.50 per million tokens. Luna also has an adjustment above a certain volume, but the threshold is higher (272,000 tokens) and the increase is much smaller, to $0.20/$0.75.
In practice, this separates two well-defined usage profiles for those building with AI:
| Model | Price up to the limit (input/output per million) | Token limit | Price above the limit |
|---|---|---|---|
| Claude Haiku 4.5 | $1 / $5 | not applicable | flat price |
| Claude Haiku 5.5 | $0.10 / $0.50 | 100,000 tokens | $0.50 / $2.50 |
| GPT-6 Luna | $0.10 / $0.50 | 272,000 tokens | $0.20 / $0.75 |
If the agent's context fits within 100,000 tokens, Haiku 5.5 ties Luna on price and still has the benchmark edge. Above that, especially in agents with long conversation history or RAG with heavy retrieved context, Luna comes out cheaper by a good margin.
The tokenizer that shrinks the budget behind the scenes
There's a second catch that Willison points out and that doesn't show up in Anthropic's pricing table: Haiku 5.5 uses a new, less generous tokenizer. Testing the same long prompt with his own token-counting tool (the Claude Token Counter), he measured that Haiku 5.5 uses about 1.25 times more tokens than Haiku 4.5 for the same text.
That's a hidden cost increase: the price per token dropped 10x, but every real request now uses more tokens to represent the same content. For anyone migrating from Haiku 4.5 to Haiku 5.5 and projecting savings based solely on the pricing table, it's worth running your own token count before finalizing the budget.
Reasoning as a cost dial, tested with the pelican on a bicycle
Haiku 5.5 ships with configurable reasoning effort levels: low, medium, high, xhigh, and max. Unlike previous versions, the new model doesn't allow reasoning to be turned off entirely, and the default is medium. This matters because every reasoning token is billed as an output token, so the effort level is, in practice, a direct control over cost and latency.
Willison tested this with his informal benchmark, the pelican riding a bicycle in SVG, running the model via the llm-anthropic plugin:
llm install -U llm-anthropic
llm anthropic refresh
llm -m claude-haiku-5.5 "Generate an SVG of a pelican riding a bicycle" -o thinking_effort lowThe results show the trade-off in practice: the pelican generated with low effort came out bad, but cost 0.0936 cents and took 7 seconds. From low effort upward the result improved and the bicycle frame came out correct. At the opposite extreme, max effort took 5 minutes and 9 seconds and cost 3.3826 cents, still a fraction of what Haiku 4.5 used to cost: the same test, a year earlier, cost 0.7583 cents on a model that didn't even support reasoning levels and still drew bad pelicans.

API credits that pay for their own subscription
In the same announcement, Anthropic cut Sonnet 5.5's cache read price in half and created an API credit scheme for subscribers to the Max and Team plans. The official text, quoted by Willison, describes the benefit this way:
Second, this week, we'll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users.
Anthropic, official statement
The credit amount matches exactly the cost of the subscription, which in practice lets you use the API without paying anything beyond the plan you already have. To activate it, just go to Settings -> Billing and choose which API organization should receive the credit. You can also turn off automatic top-up, so API calls simply stop when the balance hits zero, instead of generating a surprise bill.
A detail that catches out those who don't read the fine print: the credits don't roll over from one month to the next. Whatever isn't used is lost, so teams paying for Max or Team and not using the API regularly are leaving money on the table.
What changes for those running agents in production
The practical takeaway for developers is calculator in hand, not automatic adoption. Worth considering:
- Agents with short context (below 100,000 tokens per call): Haiku 5.5 ties Luna on price and, according to the benchmarks cited by Willison, delivers more quality. It's a natural candidate to replace Haiku 4.5.
- Agents with long context (heavy RAG, extensive history, windows above 100,000 tokens): Luna remains cheaper, because Haiku 5.5's price jump is more aggressive and kicks in sooner.
- Projects that already have prompts measured in tokens: remeasuring with the new tokenizer before comparing cost is mandatory, since the same text uses about 25% more tokens on Haiku 5.5.
- Teams with an Anthropic Max or Team subscription: it's worth setting up the API credits and using automatic top-up disabled as a budget safety lock.
Willison also notes, in passing, that OpenAI still allows using the Codex subscription for personal API use, which remains a better deal for anyone consuming heavy amounts of API. Anthropic's credit scheme narrows that gap but doesn't close it: anyone choosing between the two platforms based on cost for heavy usage still has math to do beyond the per-token price announced in the headline.
Translated from the Brazilian Portuguese original · Read the original
GPT-6's Intelligent UI makes ChatGPT generate interface, not just text
OpenAI launched Intelligent UI alongside GPT-6 on October 7: the model now decides, with each response, whether to draw a chart, a button, or a calculator. For those building products with AI, what matters isn't the visuals, it's the engine behind them, and what it still doesn't open up.