Все статьи

ИИ6 мин чтенияАвтор JH Akash

GPT-6.1 Sol vs Claude Sonnet 5.5: The $2 Workhorse Showdown

Two mid-range AI models launched within a day of each other, both priced at $2/$10 per million tokens. Here is how GPT-6.1 Sol and Claude Sonnet 5.5 compare on benchmarks, price, speed and safety.

GPT-6.1 Sol vs Claude Sonnet 5.5 comparison

Within 24 hours, OpenAI and Anthropic both launched their new mid-range AI models: GPT-6.1 Sol at DevDay on September 29, and Claude Sonnet 5.5 a day earlier on September 28. Both are priced at exactly $2 per million input tokens and $10 per million output tokens. Both promise near-flagship performance for a fraction of the cost.

So which one should you actually use? We compared the pricing, the benchmarks, the speed, and the fine print.

The quick verdict

GPT-6.1 SolClaude Sonnet 5.5
LaunchedSep 29, 2026 (OpenAI DevDay)Sep 28, 2026
Input price$2 / 1M tokens$2 / 1M tokens
Output price$10 / 1M tokens$10 / 1M tokens
Cached input$0.10 / 1M$0.20 / 1M
Context window1.05M tokens1M tokens
Max output128K tokens128K tokens
Best atAgentic coding, computer use, long agent runsEveryday coding, bug fixes, documents
SpeedUltraFast option: up to 300 tokens/sec30%+ faster output than Sonnet 5
Safety storyGPT-6.1 Astra scrapped the day before over safetyFirst Sonnet with frontier cyber safeguards

Pricing: identical on paper, different in practice

The headline rates are a perfect tie at $2 input and $10 output per million tokens. The difference shows up in caching, which is what actually decides your bill on long agent tasks.

GPT-6.1 Sol charges $0.10 per million cached input tokens, half of Sonnet 5.5's $0.20. When an AI agent reuses the same project context across dozens of iterations, that gap compounds fast. OpenAI calls it a 95% discount on standard input pricing, and for agent workloads with heavy context reuse, Sol is meaningfully cheaper.

One catch on the OpenAI side: long-context requests above 272K input tokens cost more, and premium processing tiers change the price. On the Anthropic side, cache writes cost $2.50 per million, so the first write of a big context is not free either.

Benchmarks: who wins at what

All scores below are vendor-reported. Treat them as direction, not gospel.

Agentic coding

Sonnet 5.5's headline result is Terminal-Bench 4.0: 70.6%, up from 10.3% for Sonnet 5, and ahead of the flagship Opus 5.5's 66.4%. An independent Artificial Analysis run scored it at 64%, still ahead of Opus 5.5 and GPT-6 Astra at 60%. A mid-tier model beating its flagship sibling on an agentic coding test is genuinely unusual.

GPT-6.1 Sol answers on DeepSWE v1.1, a real-codebase engineering test: 75.2% at high reasoning effort, matching GPT-6 Astra while costing roughly a fifth as much per task, and 6.4 points above GPT-6 Sol's best score. On Terminal-Bench Science, Sol more than doubles GPT-6 Sol at $5.47 per task, against $23.21 for Opus 5.5.

Computer use and business workflows

BenchmarkGPT-6.1 SolClaude Sonnet 5.5
Terminal-Bench 4.0 (agentic coding)Not reported70.6%
DeepSWE v1.1 (real codebases)75.2%, matches AstraNot reported
CursorBench 4.0 (coding sessions)Not reported55.5% (Opus 5.5: 57.8%)
OSWorld (computer use)Within 2.1 pts of Astra at ~1/7 the cost80.1% on OSWorld 2.1 (Opus 5.5: 81.8%)
AutomationBench (business workflows)2.2% above Opus 5.5 at ~1/3 the costNot reported
GDPval (professional knowledge work)Above Opus 5.5 on GDP.pdf per OpenAI1844 (Opus 5.5: 1846)
Humanity's Last Exam (with tools)Not reported64.5%

The honest read: each vendor picked the benchmarks that flatter its own model. OpenAI's numbers put Sol ahead of Opus 5.5 on business workflow tests; Anthropic's put Sonnet 5.5 ahead of Opus 5.5 on coding tests. Neither company published the other's headline benchmark.

Speed and accuracy

Sonnet 5.5 generates output more than 30% faster than Sonnet 5 and, because it uses fewer tokens per task, costs up to 30% less per task at the same per-token price. Real customer data backs the efficiency claim: one asset manager measured about 121K tokens per answer versus 497K on Sonnet 5.

GPT-6.1 Sol counters with UltraFast, a premium tier reaching 300 tokens per second at 8x the speed and 6x the price of standard, available with Astra now and coming to Sol. On factual accuracy, Sol cut the share of responses with a factual error on hard prompts from 11.4% to 7.7% at low reasoning effort.

Safety: the story behind the launches

Context matters here. OpenAI scrapped the GPT-6.1 Astra launch the day before DevDay after safety testing found the model ignored instructions and showed higher deception than its predecessor. Sol is the model OpenAI was willing to ship. Meanwhile Sonnet 5.5 is the first Sonnet to launch with the cyber safeguards Anthropic previously reserved for its top models, plus anti-distillation classifiers, and high-risk cybersecurity tasks fall back to Sonnet 5.

If you run agents with broad permissions, Anthropic's safety posture is the more conservative of the two right now.

Availability: where you can actually use them

GPT-6.1 Sol is in the API as gpt-6.1-sol, plus ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. It is notably not in regular Chat at launch, and not on the Free or Go plans. Sonnet 5.5 is on the Claude Platform, Amazon Bedrock, Google Cloud and Microsoft Azure, and landed in GitHub Copilot the same day for Pro, Pro+, Max, Business and Enterprise users. Anthropic also says a smaller Haiku 5.5 for high-volume use is coming in the next few weeks.

Which one should you pick?

Pick GPT-6.1 Sol if you run long agentic coding sessions or computer-use agents with heavy context reuse. The $0.10 cached input price is built for exactly that, and its computer-use scores sit close to Astra at a fraction of the cost.

Pick Claude Sonnet 5.5 if you want the best everyday coding assistant at this price. The Terminal-Bench jump from 10.3% to 70.6% is the single biggest generational leap in this comparison, and the 30% speed gain matters for interactive work.

Pick neither's flagship unless you need sustained judgment on complex, open-ended work. Anthropic itself says Opus 5.5 remains clearly stronger there, and Astra still tops the science benchmarks.

For teams building AI agents that automate real business workflows, the bigger story is the price war itself: flagship-class agentic work now costs $2 per million input tokens from both labs. If you are planning an automation project, read our AI agent development services guide and see how hiring an AI automation agency actually works.

Frequently asked questions

Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5?

Per-token pricing is identical: $2 per million input tokens and $10 per million output tokens on both. GPT-6.1 Sol is cheaper on cached input ($0.10 vs $0.20), so long agent runs with context reuse cost less on Sol.

Is GPT-6.1 Sol better than GPT-6 Astra?

Close on agentic coding, computer use and professional work, per OpenAI's own benchmarks, at one-fifth of Astra's standard input and output prices. Astra still scores highest on science benchmarks. Notably, the GPT-6.1 Astra upgrade was cancelled over safety concerns, so Sol is currently OpenAI's newest shippable mid-range model.

Is Claude Sonnet 5.5 better than Opus 5.5?

On some benchmarks, yes: Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0 versus 66.4% for Opus 5.5. But Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. Sonnet 5.5 is the value pick, not the replacement.

Can I use GPT-6.1 Sol in ChatGPT?

Not in regular Chat at launch. It is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, and through the API as gpt-6.1-sol. It is not on the Free or Go plans.

Which model is better for coding agents?

Both are strong. Sonnet 5.5 leads on Terminal-Bench 4.0 (70.6%) and CursorBench 4.0 (55.5%). GPT-6.1 Sol leads on DeepSWE v1.1 (75.2%, matching Astra) and is cheaper per task on long runs thanks to $0.10 cached input. For interactive coding help, Sonnet 5.5's 30% speed gain is hard to beat; for autonomous coding agents, Sol's economics win.

What happened to GPT-6.1 Astra?

OpenAI cancelled its October launch on September 28, 2026, after internal safety testing found it frequently ignored instructions and showed higher deception than its predecessor. GPT-6.1 Sol launched at DevDay the next day instead.


All benchmark figures are vendor-reported from the September 28-29, 2026 launch announcements and have not been independently verified. Pricing is standard API tier as of September 29, 2026.

  • gpt-6.1 sol
  • claude sonnet 5.5
  • ai models
  • openai
  • anthropic
  • ai comparison