NEWS

Claude Haiku 5.5 arrives 90% cheaper and with adjustable effort

Claude Haiku 5.5 was launched by Anthropic on October 8, nearly a year after the previous version. The company's most affordable line had been stalled

Claude Haiku 5.5 arrives 90% cheaper and with adjustable effort
Image: Redação iMasters

Claude Haiku 5.5 was launched by Anthropic on October 8, nearly a year after the previous version. The company's most affordable line had been stalled since October 2025. However, the new model targets high-volume requests and lower-cost tasks.

That gap drew attention in the industry. Meanwhile, competitors kept updating their budget options, such as GPT Luna and Gemini Flash Lite.

The novelty that changes everyday use

This is the first model in its class with an adjustable effort level. Therefore, it's possible to reduce reasoning when the priority is saving tokens.

When the task demands more, you simply raise that level. This way, the same model handles different scenarios without switching endpoints.

Notice how it fits into an agent architecture. Small models execute simple tasks, while larger options orchestrate activities and set objectives.

Claude Haiku 5.5 outperforms GPT 6 Luna in published tests

The launch benchmarks compared the new model to its direct competitor. According to the company, it wins in every highlighted test.

The internal comparison also came out positive. In fact, the model outperforms Haiku 4.5 in the evaluations presented.

Still, there's a clear limit. It loses to Sonnet 5.5 across all synthetic evaluations.

This positioning makes sense within the portfolio. The line prioritizes speed and low cost, while Sonnet and Opus take on complex tasks.

The pricing table now has two tiers

Here's the point that requires attention in your calculations. Pricing changes once usage goes past 100,000 tokens.

Up to that limit, input costs US$0.10 per million tokens. Above it, the price rises to US$0.50.

On output, the numbers are US$0.50 and US$2.50 respectively. Cache reads cost US$0.01 and US$0.05.

Cache writes, in turn, cost US$0.125 and US$0.625.

The comparison with the previous generation is striking. The price is 90% lower for requests of up to 100,000 tokens and 50% lower above that.

According to the company, 90% of Haiku 4.5 requests fell into the first tier. Additionally, the new version consumes fewer tokens to get the job done.

Claude Sonnet 5.5 also got cheaper

The announcement brought a parallel reduction. The price of cache reads for Sonnet 5.5 was cut in half.

The price went from US$0.20 to US$0.10 per million tokens.

The effect shows up where it matters most. According to the company, the model tends to run up to 20% cheaper in agentic work.

There's also another commercial update. The company will start distributing a new amount of monthly credits for API use on the Max and Team plans.

Where to use it and what to measure before migrating to Claude

The model is already available on compatible platforms. The list includes Amazon Web Services, Google Cloud and Microsoft Azure.

On the Claude Platform, the identifier is claude haiku 5 5.

First, measure how many of your calls fall below 100,000 tokens. That number defines the real price you'll end up paying.

Second, test the effort levels in your workflow. Plenty of tasks perform well at the lowest setting.

Third, review your caching strategy. With reads this cheap, reusing context becomes a concrete advantage.

Finally, think about routing. Simple tasks on Haiku and open-ended work on Sonnet or Opus tends to perform better than relying on a single model.

Follow our profile on Instagram!

Translated from the Brazilian Portuguese original · Read the original

More from Redação iMasters
View profile →