Tutti gli articoli

IA6 min di letturaDi JH Akash

GPT-6.1 Sol Ultrafast: OpenAI Selling Speed at 300 Tokens/Sec

Tibo's "6.1 coming soon" post pointed at Ultrafast, OpenAI's paid speed tier: 300 tokens/sec at 6x the price. The numbers, the NVIDIA twist, and the math.

GPT-6.1 Sol Ultrafast infographic: 300 tokens per second, 8x faster in Codex, 6x faster in API, 2/0 per 1M tokens

On October 4, 2026, Tibo posted three words on X: "6.1 coming soon." GPT-6.1 Sol was already live, so the internet did what it does, and the guesses piled up. Was the cancelled GPT-6.1 Astra back? Was a 6.1 Luna coming? Tibo himself settled it with a reply: "6.1 sol ultrafast."

That short reply points at OpenAI's newest product category: not a model, but speed sold as a paid tier. Ultrafast is live for GPT-6 Astra and rolling out for GPT-6.1 Sol, and it changes how AI pricing works. Here is everything that is actually confirmed, what it costs, and whether the math makes sense for your workloads.

Quick verdict

Ultrafast is OpenAI's premium inference lane. For GPT-6.1 Sol it promises up to 8x faster token generation in Codex and up to 6x in the API, around 300 tokens per second, at 6 times the standard price. That puts Sol Ultrafast at $12 per million input tokens and $60 per million output tokens, still cheaper than standard GPT-6 Astra. If your agents or users wait on every token, this is the tier to watch. If you run batch jobs where nobody watches the screen, you are paying for nothing.

How the story broke

The timeline matters, because the speculation moved faster than the facts.

  • September 28: OpenAI scrapped the planned October release of GPT-6.1 Astra after internal tests showed more deception and failed scope-authorization checks than GPT-6 Astra. Safety chief Saachi Jain confirmed the decision publicly.
  • September 29: At DevDay, OpenAI shipped GPT-6.1 Sol instead, priced at $2 per million input tokens and $10 per million output tokens, with cached input at $0.10. The company says it delivers near-Astra performance on agentic coding, computer use, and professional work (vendor claim). Ultrafast debuted the same day for GPT-6 Astra.
  • October 4: Tibo's "6.1 coming soon" post and his "6.1 sol ultrafast" reply pointed the market at the one gap everyone knew about: Ultrafast was listed as "coming soon" for Sol since DevDay.

So no, GPT-6.1 Astra is not back. The cancelled model stays cancelled, and the 6.1 in question is the speed tier, not the model.

Pricing: the full rate card

Ultrafast costs 6x the standard tier. Here is what that means in dollars per million tokens, based on OpenAI's published pricing via VentureBeat.

TierInput / 1MOutput / 1MPrice multiple
GPT-6.1 Sol Standard$2$101x
GPT-6.1 Sol Fast mode$4$202x
GPT-6.1 Sol Ultrafast$12$606x
GPT-6 Astra Standard$10$501x
GPT-6 Astra Ultrafast$60$3006x

Two numbers matter here. First, Sol Ultrafast at $60 output per million costs roughly the same as standard Astra at $50, which is the whole pitch: near-flagship speed on the cheap model, for about flagship money. Second, the jump from Fast mode to Ultrafast is steep: triple the price of the existing 2x priority tier for the extra throughput.

GPT-6.1 Sol price per 1M output tokens: Sol Standard $10, Sol Fast $20, Sol Ultrafast $60, Astra Ultrafast $300

Speed: what 300 tokens per second means

OpenAI's claims for Ultrafast: up to 300 tokens per second, roughly 8x faster generation in Codex and up to 6x faster through the API. For context, independent benchmark firm Artificial Analysis currently measures Google's Gemini 3.5 Flash at about 201 tokens per second. Note the comparison is not apples to apples: 300 is OpenAI's vendor-reported ceiling, while 201 is a third-party measurement.

GPT-6.1 Sol Ultrafast speed vs rivals in tokens per second: Ultrafast 300 per OpenAI claim, Gemini 3.5 Flash 201 per Artificial Analysis

The more interesting question is what the tier runs on. Semi Analysis reported that GPT-6.1 Sol Ultrafast is NOT running on Cerebras, as many assumed, but on NVIDIA GPUs at low batch size. The detail stung enough that NVIDIA's official AI account replied with nothing but an eyes emoji, and the post crossed 100,000 views within hours. Context: OpenAI had touted 750 tokens per second for GPT-5.6 Sol on Cerebras in August, so the 300 tokens per second figure looks modest next to what Cerebras hardware has done for them.

The technical read: low batch sizes mean fewer requests packed together, which is how you buy latency with hardware. That is consistent with a separately provisioned premium lane, and with Tibo's earlier note that 6.1 Sol recovered to expected speeds after a two-day load spike.

Who gets it, and what it is really for

Ultrafast launched for GPT-6 Astra with consumer access gated behind the $500 Pro tier and Enterprise plans. For Sol, the tier is rolling out now. The natural buyers are latency-bound workloads: coding agents where a developer watches the cursor, voice AI where round-trip time kills the conversation, and interactive copilots where every 100 milliseconds of delay shows.

The wrong buyer is anyone running batch pipelines: evaluation sweeps, document processing, overnight summarization. There, throughput is not latency, and 6x the price buys you 6x the tokens at the same total job cost with no user-facing benefit.

The pattern: speed becomes a SKU

Ultrafast is the third leg of a pricing model OpenAI has been assembling all year. First came the price war: GPT-6.1 Sol at a fifth of Astra's price, and a capability refresh seven days into a generation's life. Then came usage governance: halved Pro 200 limits, load spikes, separately metered capacity. Now comes the speed SKU: latency as a line item on the invoice.

The implication for builders is practical. Model choice is now a weekly procurement decision, not an architectural one. Pin your model IDs, route by latency budget, and re-run the cost math every time OpenAI moves a number, because they keep moving.

For teams building agent products, the interesting middle is custom AI agents designed around a latency budget rather than a model name. If your agent's loop tolerates Sol Standard prices but needs Ultrafast response times, that is a routing problem, and routing problems are what an AI automation agency solves daily.

Which to pick

  • Interactive agents and coding copilots where humans wait: Ultrafast is worth testing. The 6x multiplier is real money, so benchmark it against Fast mode on your actual traces before committing.
  • Standard chat, RAG, content pipelines: Stay on Standard or Fast mode. Nobody watching means nobody paying for.
  • Cutting-edge reasoning where accuracy dominates: GPT-6 Astra Standard at $10/$50 still exists. Speed does not upgrade intelligence.
  • High-volume background jobs: Standard Sol at $2/$10 remains the value tier. Pin the model ID and move on.

Related reading: GPT-6 Astra: the model that changed everything and Claude Opus 5.5, the most powerful AI model yet.

FAQ

What is GPT-6.1 Sol Ultrafast?

It is OpenAI's paid speed tier for GPT-6.1 Sol, promising up to 300 tokens per second, about 8x faster generation in Codex and up to 6x in the API, at 6 times the standard token price.

How much does Ultrafast cost?

For GPT-6.1 Sol: $12 per million input tokens and $60 per million output tokens. For GPT-6 Astra: $60 input and $300 output per million. Pricing had not been updated on OpenAI's pricing page at the time of writing, so confirm before provisioning.

Is GPT-6.1 Astra coming back?

Nothing official says so. OpenAI cancelled the planned October release on September 28, 2026 after internal safety tests. Tibo's "6.1 coming soon" post referred to Sol Ultrafast, not Astra.

Does Ultrafast run on Cerebras?

According to Semi Analysis, no. They reported that GPT-6.1 Sol Ultrafast runs on NVIDIA GPUs at low batch size, not on Cerebras hardware.

Who should pay for Ultrafast?

Teams running latency-bound workloads: coding agents, voice AI, and interactive copilots where a human waits on every response. Batch jobs with no human watching get no benefit from it.

  • GPT-6.1 Sol
  • Ultrafast
  • OpenAI
  • AI pricing
  • AI inference speed
  • AI news