Tous les articles

IA6 min de lecturePar Sabbir Ahmed Minhaz

Claude Opus 5.5: The New Most Powerful AI Model, Beating GPT-6 Astra and Fable 5.1

Claude Opus 5.5 tops the Artificial Analysis Intelligence Index, beats GPT-6 Astra and Fable 5.1 on coding and real work, and costs 40% less than Opus 5. Charts inside.

Bar chart showing Claude Opus 5.5 leading the Artificial Analysis Intelligence Index with 58, ahead of Claude Fable 5.1 and GPT-6 Astra at 53

Anthropic released Claude Opus 5.5 on September 22, 2026, and the numbers are hard to ignore. Only two months after Opus 5, the new model now sits at the top of the Artificial Analysis Intelligence Index, beats Anthropic's own larger Claude Fable 5.1 on most benchmarks, and pulls clearly ahead of OpenAI's GPT-6 Astra on coding and real professional work.

And it does all of that while costing less to run than the model it replaces.

At CodeMyPixel we use Claude every day for engineering, content and client work, so we went through the launch data to see how much of the hype holds up. Here is what we found.

The short version

  • #1 on the Artificial Analysis Intelligence Index with a score of 58, five points ahead of both Claude Fable 5.1 and GPT-6 Astra (53 each).
  • 66.4% on Terminal-Bench 4.0, compared with 57.9% for GPT-6 Astra and 55.8% for Fable 5.1.
  • 1846 Elo on GDPval-AA, a benchmark of real professional tasks. That is 304 points above GPT-6 Astra.
  • About 40% cheaper to run than Opus 5 on typical workloads, with lower prices on every token type.
  • Around 30% faster output generation than Opus 5.

Smartest model on the leaderboard

The Artificial Analysis Intelligence Index combines a wide set of reasoning, knowledge, math and coding evaluations into one score. It is one of the most useful independent snapshots we have for comparing frontier models side by side.

Horizontal bar chart of the Artificial Analysis Intelligence Index. Claude Opus 5.5 leads with 58, followed by Claude Fable 5.1 and GPT-6 Astra at 53, Muse Spark 1.3 at 48, Grok 4.7 and MiMo-V2.6-Pro at 46, GLM-5.3 at 45, Gemini 3.8 Flash at 41, DeepSeek V4.1 Flash at 39 and GPT-5.6 Luna at 37.

Opus 5.5 scores 58. The next best models, Claude Fable 5.1 and GPT-6 Astra, are tied at 53. A five point lead at the very top of this index is a big gap. For context, the distance between GPT-6 Astra and the fourth place model, Muse Spark 1.3, is also five points.

The interesting part is that Opus 5.5 beats Fable 5.1, which is Anthropic's larger and more expensive model. A mid tier flagship outscoring the premium tier is not something we see often.

Coding: where the gap is widest

For a studio like ours, coding performance is what matters most. This is where Opus 5.5 pulls furthest ahead.

Grouped bar chart of agentic coding benchmarks. On Terminal-Bench 4.0, Claude Opus 5.5 scores 66.4%, Claude Fable 5.1 55.8% and GPT-6 Astra 57.9%. On FrontierCode v1.1, Opus 5.5 scores 54.4%, Fable 5.1 50.3% and GPT-6 Astra 53.3%.

On Terminal-Bench 4.0, which tests an agent working in a real terminal to finish multi step engineering tasks, Opus 5.5 solves 66.4% of tasks. GPT-6 Astra solves 57.9% and Fable 5.1 solves 55.8%. That is an 8.5 point lead over the best OpenAI model.

On FrontierCode v1.1 the race is closer. Opus 5.5 still leads with 54.4%, with GPT-6 Astra just behind at 53.3%.

Anthropic also shared results on CursorBench 4.0 (57.8% versus 51.8% for Fable 5.1) and OSWorld 2.0 for computer use (81.8% versus 80.7%). One early tester reportedly finished a 680,000 line code migration in less than a day, work that would normally take an engineering team weeks.

Real professional work, not just puzzles

Benchmarks can feel abstract. GDPval-AA tries to fix that by scoring models on the kind of tasks people actually get paid for: analysis, reports, planning and business documents. Results are ranked with an Elo rating, like chess.

Horizontal bar chart of GDPval-AA v2.1 Elo ratings. Claude Opus 5.5 scores 1846, Claude Fable 5.1 1735, Claude Opus 5 1708 and GPT-6 Astra 1542.

Opus 5.5 reaches 1846 Elo. Fable 5.1 sits at 1735 and GPT-6 Astra at 1542. In Elo terms, a 300 point gap means the higher rated model wins the large majority of head to head comparisons. For knowledge work, this is the clearest win in the whole launch.

Anthropic also says the model communicates better: less jargon, shorter answers, and the most important information placed first. In an earnings research test, 16 out of 18 reports passed quality checks.

Cheaper than the model it replaces

Usually a smarter model means a bigger bill. Opus 5.5 goes the other way.

Grouped bar chart comparing API prices per million tokens. Claude Opus 5 costs $5 input, $25 output, $6.25 cache write and $0.50 cache read. Claude Opus 5.5 costs $4 input, $20 output, $5 cache write and $0.20 cache read.

  • Input: $5 down to $4 per million tokens
  • Output: $25 down to $20 per million tokens
  • Cache writes: $6.25 down to $5 per million tokens
  • Cache reads: $0.50 down to $0.20 per million tokens

Anthropic says Opus 5.5 costs around 40% less than Opus 5 on typical workloads. A fast mode is also available at $8 input and $40 output per million tokens, with up to 2.5x the speed.

What it costs per task, honestly

Price per token is only half the story. Artificial Analysis also measures the average cost to complete one task from its Intelligence Index.

Horizontal bar chart of cost per Intelligence Index task in USD. Claude Fable 5.1 costs $7.63, Claude Opus 5.5 $5.98, Grok 4.7 $3.74, GPT-6 Astra $3.26, GLM-5.3 $2.01, Muse Spark 1.3 $1.60 and Gemini 3.8 Flash $1.24.

Here we want to be fair. Opus 5.5 costs $5.98 per task, which is 22% cheaper than Fable 5.1 ($7.63) while being smarter. But GPT-6 Astra is cheaper still at $3.26 per task. So the real question is what you are paying for. If your work is coding agents or complex business tasks, the extra accuracy of Opus 5.5 usually saves more time than the price difference. For simple, high volume jobs, a cheaper model can still be the better fit.

Speed is also not its headline. In the Artificial Analysis output speed ranking, Gemini 3.8 Flash leads by a wide margin. Opus 5.5 is about 30% faster than Opus 5, which is a nice upgrade, but it is built to be the smartest model, not the fastest one.

See it in action

Want to see how it feels in a real coding session? This hands on session from BridgeMind puts Opus 5.5 to work on a real build.

YouTube
Vibe Coding With Claude Opus 5.5, by BridgeMind

Safety came along for the ride

Anthropic did not trade safety for speed here. Opus 5.5 was evaluated before release by outside groups including METR, and Anthropic reports its best score yet on internal behavioral audits. Resistance to prompt injection matches or beats Opus 5 in every setting tested, which matters a lot if you run agents that read emails, websites or documents.

This is also Anthropic's first release since CEO Dario Amodei's call to "pace the frontier", so the company is clearly trying to show that it can improve capability and safety together.

Opus 5.5 vs GPT-6 Astra vs Fable 5.1 at a glance

  • Intelligence Index: Opus 5.5 58, Fable 5.1 53, GPT-6 Astra 53. Winner: Opus 5.5
  • Terminal-Bench 4.0: Opus 5.5 66.4%, Fable 5.1 55.8%, GPT-6 Astra 57.9%. Winner: Opus 5.5
  • FrontierCode v1.1: Opus 5.5 54.4%, Fable 5.1 50.3%, GPT-6 Astra 53.3%. Winner: Opus 5.5
  • GDPval-AA Elo: Opus 5.5 1846, Fable 5.1 1735, GPT-6 Astra 1542. Winner: Opus 5.5
  • Cost per index task: Opus 5.5 $5.98, Fable 5.1 $7.63, GPT-6 Astra $3.26. Winner: GPT-6 Astra

Where you can use it

Opus 5.5 is available today in the Claude apps on the Pro, Max, Team and Enterprise plans (with higher five hour usage limits), on the Claude API as claude-opus-5-5, and through AWS, Google Cloud and Microsoft Azure. Anthropic says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks.

Our take

Claude Opus 5.5 is the most capable model you can use right now. It leads the independent intelligence ranking, it has a clear lead over GPT-6 Astra on coding and real professional work, and it beats Anthropic's own bigger Fable 5.1 while costing less. GPT-6 Astra still wins on raw cost per task, and speed focused models are faster, but for serious engineering and knowledge work Opus 5.5 is our new default.

At CodeMyPixel we build with Claude every day, and Opus 5.5 is an easy upgrade for our AI assisted development and content workflows. If you want help building AI agents or AI powered products on top of it, get in touch with our team.

Sources: Anthropic, Introducing Claude Opus 5.5, TechCrunch, and Artificial Analysis benchmark highlights.

  • Claude Opus 5.5
  • Anthropic
  • Claude
  • GPT-6 Astra
  • Claude Fable 5.1
  • AI benchmarks
  • AI coding
  • LLM comparison
  • AI news