Tutti gli articoli

IA7 min di letturaDi JH Akash

Claude Haiku 5.5 vs GPT-6 Luna: The $0.10 Small-Model Showdown

Claude Haiku 5.5 costs $0.10 per 1M input tokens, matches GPT-6 Luna's price, and beats it on benchmarks. Pricing, context, speed, and which to pick, compared.

Claude Haiku 5.5 vs GPT-6 Luna: benchmarks, price and which to pick compared

Anthropic launched Claude Haiku 5.5 on October 7, 2026, and the small-model market suddenly got interesting. The company calls it the cheapest, fastest, and most capable small model it has ever shipped, with a 1M-token context window, adaptive thinking, and a price cut of about 90% versus Haiku 4.5.

The catch for buyers: OpenAI's GPT-6 Luna carries the exact same list price. So the fight is no longer about the sticker, it is about what you actually get per dollar. We compare the two on price, benchmarks, speed, and the fine print.

All benchmark figures below are vendor-reported by Anthropic in its October 7 announcement unless marked independent.

Quick verdict

Claude Haiku 5.5GPT-6 Luna
List price (up to 100K tokens)$0.10 in / $0.50 out per 1M$0.10 in / $0.50 out per 1M
Long prompts$0.50 / $2.50 above 100K$0.20 / $0.75 above 272K
Context window1M tokensSmaller long-prompt tier is cheaper
Knowledge work (GDPval-AA v2.1)1620 Elo1437 Elo
Computer use (OSWorld 2.1)72.4%48.9%
Raw output speed237.5 tokens/sec136.9 tokens/sec
Real cost per task~$0.21~$0.07

Haiku 5.5 wins on capability-per-dollar-listed. Luna wins on the bill you actually pay per task, because it writes far fewer tokens. Read on for why.

Pricing: same sticker, different bill

The headline prices are identical. For prompts up to 100,000 tokens, both models charge $0.10 per million input tokens and $0.50 per million output tokens. Haiku 5.5 is a 90% cut from Haiku 4.5 ($1.00 / $5.00). Prompt caching on Haiku 5.5 saves up to 90% (cache reads at $0.01 per million), and batch processing saves another 50%.

The tiers differ where it matters:

ModelPrompt lengthInput / 1MOutput / 1M
Claude Haiku 5.5Up to 100K tokens$0.10$0.50
Claude Haiku 5.5Over 100K tokens$0.50$2.50
GPT-6 LunaShort tier$0.10$0.50
GPT-6 LunaOver ~272K tokens$0.20$0.75
Claude Haiku 4.5Flat$1.00$5.00

Two gotchas. First, Haiku 5.5's long-prompt tier is expensive: above 100K tokens you pay 5x the input price, while Luna's long tier is gentler. For a 150K-token prompt, Luna is cheaper on list price. Second, Haiku 5.5 uses a newer tokenizer that counts about 30% more tokens for the same text, which eats into the headline cut. Anthropic's "75% cheaper on average than Haiku 4.5" claim already bakes this in, but you should still rerun your own cost estimates on real traffic.

Benchmarks: Haiku 5.5 leads on the shared tests

These are Anthropic's published numbers. Haiku 5.5 was tested at max effort.

BenchmarkHaiku 5.5Haiku 4.5GPT-6 Luna
GDPval-AA v2.1 (Elo)16207351437
AA-Briefcase v1.1 (Elo)15786141336
OSWorld 2.1, computer use72.4%15.7%48.9%
Terminal-Bench 4.0, agentic coding39.2%0.0%16.4%
FrontierCode 1.1 Main46.4%n/a42.4%
Chartography, no tools46.4%6.4%29.1%
Humanity's Last Exam, no tools45.9%10.2%n/a
Humanity's Last Exam, with tools57.4%18.7%n/a
SWE-bench Multilingual83.7%67.4%n/a

The OSWorld result is the headline. Haiku 4.5 barely registered on computer use (15.7%), while Haiku 5.5 scores 72.4% against Luna's 48.9%. A Haiku-class model that can operate real software changes what you can build at this price. On Terminal-Bench, Haiku 4.5 scored literally 0.0%; Haiku 5.5 now completes 39.2% of agentic coding tasks.

Read the table with three caveats. Scores were run at max effort, and Anthropic's own system card shows GDPval-AA dropping from 1620 to 1277 at medium effort (though medium uses about a tenth of the output tokens). Artificial Analysis measured Terminal-Bench at 33% in its own harness versus 39.2% in Anthropic's run. And Anthropic did not publish SWE-bench Verified, GPQA Diamond, or tau-bench for this model, so treat any article quoting those as Anthropic figures with skepticism.

Speed: fast, but verbosity decides the bill

Haiku 5.5 is genuinely fast. Artificial Analysis measured 237.5 tokens per second of output against Luna's 136.9. Anthropic pitches it as a high-volume workhorse and as a subagent for its larger Opus 5.5 and Sonnet 5.5 models, which plan the work while Haiku executes small tasks in parallel.

But identical list prices can still produce very different bills. On Artificial Analysis' measure of what a typical benchmark task actually bills, Haiku 5.5 lands at $0.21 against Luna's $0.07. The cause is verbosity: Haiku 5.5 writes about 4.4x the median output tokens on the same test suite, while Luna runs near typical.

The lever Haiku 5.5 gives you is new: it is the first Haiku-class model with adjustable effort controls (Low to Max), with adaptive thinking on by default. The model decides how much reasoning a task needs, and you can dial it down to trade intelligence for a smaller bill. Note that at high effort, Artificial Analysis measured a 26-second time to first token, so the speed advantage depends heavily on the effort setting.

What's new vs Haiku 4.5 (and what's breaking)

Beyond price: the context window jumps from 200K to 1M tokens, max output from 64K to 128K, and safety classifiers can now return stop_reason: "refusal" with no server-side fallback.

Migration is not a drop-in swap. The Claude Platform docs list breaking changes: manual extended thinking with budget_tokens now errors, non-default temperature/topp/topk values error, assistant prefill is rejected, and computer use requires the new computer_toolset_20260801 version. Audit your code before switching the model ID to claude-haiku-5-5.

On launch day Anthropic also halved cache-read prices for Sonnet 5.5 (it says this makes most agentic work about 20% cheaper) and added a monthly API credit for Max and Team subscribers. The model shipped everywhere at once: Claude.ai, the Claude Platform, Claude Code, Amazon Bedrock, Google Cloud, Microsoft Foundry, and GitHub Copilot.

Early customer numbers

Anthropic's launch materials include results from enterprise customers. These are vendor-published, but several carry specific numbers:

  • AlphaSense: ran 400 queries for its "Ask in Document" feature (about 8M calls a week in production). Haiku 5.5 scored 0.84 versus 0.76 for Haiku 4.5, a statistically significant gain.
  • Asana: over 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn.
  • HubSpot: 92.8% on an internal CRM evaluation, the best score it had seen, with the fastest completion and the lowest false positive rate.
  • Box: an 11-point improvement at about half the latency versus Haiku 4.5.

Which to pick

Pick Claude Haiku 5.5 if you need agentic capability at the lowest list price. It leads Luna on every shared benchmark, pushes 237 tokens per second, offers a 1M-token window, and is the natural subagent inside the Claude stack. Teams already running Sonnet or Opus 5.5 get the cheapest fast executor for classification, summarization, extraction, context compaction, and browser use. If you build AI agents into client workflows, this is the model that makes high-volume agentic work affordable. That is exactly what we do at CodeMyPixel: our AI agent development services put models like this inside real business pipelines, and agencies looking to automate delivery should talk to a hire-an-AI-automation-agency style partner that knows which tier actually fits.

Pick GPT-6 Luna if your workload is high-volume and simple. Identical list prices aside, Luna's leaner outputs make the real bill roughly 3x cheaper per task in independent testing, and its long-prompt tier is cheaper too. For summarization, classification, and extraction at scale, Luna is the cheaper runtime.

Stick with Sonnet 5.5 for hard agentic coding. Terminal-Bench stays wide open: 39.2% for Haiku 5.5 versus 70.6% for Sonnet 5.5 at $2/$10. Use the small models where they win, not where they struggle.

One honest read: this is a two-model race with no single winner. Haiku 5.5 is the capability buy, Luna is the economy buy, and your prompt mix decides which math wins. Our comparison of the bigger models, GPT-6.1 Sol vs Claude Sonnet 5.5, covers the flagship tier, and Claude Opus 5.5 sits above both for the heaviest work.

FAQ

When was Claude Haiku 5.5 released?

Anthropic released Claude Haiku 5.5 on October 7, 2026. It launched the same day on Claude.ai, the Claude Platform, Claude Code, Amazon Bedrock, Google Cloud, Microsoft Foundry, and GitHub Copilot. The API model ID is claude-haiku-5-5.

What does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens: $0.10 per million input tokens and $0.50 per million output tokens. Above 100K tokens: $0.50 input and $2.50 output. Cache reads cost $0.01 per million, five-minute cache writes $0.125, and batch processing takes another 50% off.

What is the context window of Claude Haiku 5.5?

1M tokens of context, up from 200K on Haiku 4.5, with up to 128K output tokens (up from 64K). Note the price jumps fivefold for prompts over 100K tokens, so budget the higher tier if you plan to fill the window.

Is Claude Haiku 5.5 a drop-in replacement for Haiku 4.5?

No. Manual budget_tokens thinking, non-default sampling parameters, and assistant prefill now return errors, and computer use requires a new toolset version. The new tokenizer also counts about 30% more tokens for the same text, so rerun cost estimates on real traffic.

Which is cheaper: Claude Haiku 5.5 or GPT-6 Luna?

Same list price ($0.10/$0.50 per million), but Luna is usually cheaper in practice. It writes far fewer tokens per task, costing about $0.07 versus $0.21 per typical benchmark task in independent testing, and its long-prompt tier ($0.20/$0.75 above ~272K tokens) undercuts Haiku 5.5's $0.50/$2.50 tier above 100K.

Which is better for AI agents: Haiku 5.5 or GPT-6 Luna?

Haiku 5.5, on the current numbers. It leads Luna on OSWorld 2.1 (72.4% vs 48.9%) and Terminal-Bench 4.0 (39.2% vs 16.4%), and Anthropic positions it explicitly as a subagent for Opus 5.5 and Sonnet 5.5. Luna still wins on raw cost per task for simple high-volume work.

  • Claude Haiku 5.5
  • GPT-6 Luna
  • AI pricing
  • AI benchmarks
  • Anthropic
  • OpenAI