📑 Table of contents

GPT-6.1 Sol at 1/5 the price of Astra while Gemini 4 disappoints in real-world use: the market rationalizes

LLM & Modèles 🟢 Beginner ⏱️ 15 min read 📅 2026-10-03

GPT-6.1 Sol at 1/5 the price of Astra while Gemini 4 disappoints in real-world use: the market is rationalizing

🔎 Two announcements, one signal: the benchmark party is over

Two news stories broke just days apart, and they tell the same story. OpenAI launches GPT-6.1 Sol, a model pitched as nearly on par with its flagship GPT-6 Astra, at one-fifth the price. In the same window, Bloomberg reveals that Google employees doubt the real-world performance of Gemini 4 Argon, despite it being excellent on paper.

While the tech press — AI Weekly leading the charge — churns out announcement after announcement, one underlying figure stands out: only 2.2% of consumers were paying for AI in May 2026. The consumer economy isn't taking off. It's plateauing, and everyone knows it by now.

The market is rationalizing. No more benchmark records as a selling point — what matters now is cost per task, margins, and real-world usage. And for once, this rationalization works in users' favor — provided you know how to read the numbers.


The Essentials

  • OpenAI launches GPT-6.1 Sol: nearly the level of GPT-6 Astra in coding, computer use, and professional tasks, at one-fifth of Astra's price.
  • Cost per task down about 30% compared to GPT-6 Sol, which itself launched with a 50% price cut (analyst Hesamation).
  • Gemini 4 Argon matches GPT-6 Astra at 60% of the cost per task according to Artificial Analysis, but Google employees cited by Bloomberg consider it disappointing at front-end coding and multi-step tasks.
  • 2.2% of consumers were paying for AI in May 2026, average spend of $31/month (PNC study, relayed by Andreessen Horowitz's State of Markets report).
  • Cost per task is becoming the market's reference metric — benchmarks are being demoted to a marketing argument.
  • On the developer side: a well-designed agent architecture cuts the bill by up to 77% (feedback from r/AI_Agents), and a small open-source model costs about 9 cents per monthly active user (Inworld AI data, July 2026).

Model Main use Price (October 2026) Ideal for
GPT-6.1 Sol (OpenAI) Coding, computer use, everyday work ≈ 1/5 of Astra's price (check openai.com) The workhorse: agents, dev, high volumes
GPT-6 Astra (OpenAI) The most complex tasks Premium reference (check openai.com) Critical steps where failure is costly
Gemini 4 Argon (Google) General-purpose, long context ≈ 60% of Astra's cost per task (Artificial Analysis) Tight budgets — test it on your real use cases
Claude Opus 5.5 (Anthropic) Demanding coding, reasoning Premium pricing (see 2026 comparison) Maximum quality with no compromise

GPT-6.1 Sol: the "quasi-Astra" that slashes prices

Direct answer: GPT-6.1 Sol is a commercial offensive disguised as a technical launch. OpenAI presents its new model as nearly on par with GPT-6 Astra in coding, computer use, and professional tasks — for one-fifth the price of its own flagship.

That's the positioning openly embraced in the announcement, notably picked up by BNB: Sol is OpenAI's "daily workhorse." Astra stays at the top for the hardest tasks. Sol takes everything else — that is, the overwhelming majority of real request volume.

Analyst Hesamation offers the nuance that matters (source): the cost per task of GPT-6.1 Sol is about 30% lower than that of GPT-6 Sol. And above all, GPT-6 Sol's cost reduction didn't come from an efficiency leap, but from a simple 50% price cut. Translation: OpenAI is buying market share with its margins. There is, for now, no technical revolution behind it.

This isn't an isolated move either. The Sol line has followed a clear trajectory since the GPT-5.6 Sol preview, launched right at the start of the price war. OpenAI understood that the market would no longer pay the premium by default, and each iteration widens the price gap with the flagship.

Early usage feedback is pretty good: on r/AI_Agents, developers find Sol close to Astra for complex work. If that holds at scale, Astra's premium will only be justified for a minority of cases. That's exactly the goal.

Why this launch is an admission

A vendor doesn't slash its flagship's price without a reason. By launching a quasi-Astra at one-fifth the price, OpenAI implicitly admits that the premium no longer sells itself — and that growth has to come from volume, not unit price. It's the same conclusion as the consumer data: the only way to grow is to make mass usage affordable.

My take: this is the most important move of the month. Not because Sol is brilliant, but because it forces the entire market to reposition. When the number one sells nearly its best model at one-fifth the price, the others no longer have the option of playing it safe.


Gemini 4 Argon: when benchmarks and real-world performance diverge

Direct answer: Gemini 4 Argon is probably a good model whose marketing runs ahead of reality. Independent tests from Artificial Analysis place it on par with GPT-6 Astra at 60% of the cost per task, with a significantly lower hallucination rate according to early tests.

The problem? According to Bloomberg, Google employees themselves doubt its real-world performance. The two cited weak points: front-end coding and multi-step tasks. In other words, precisely the use cases for which you'd choose a premium model in 2026.

It's yet another example of a phenomenon we've been observing for two years now: benchmarks get optimized, real-world performance lags behind. A model can shine on standardized evaluation suites and stumble on a real front-end codebase, with its dependencies, its UI states, and its technical debt.

The positive point, however, is real: the drop in hallucinations reported in early tests. That's often where professional users' trust is decided — a model that makes things up less is a model you can leave more on its own in production. What remains to be confirmed is whether it holds up outside the test bench.

For coding, the lesson is simple: never choose a model based on a benchmark alone. Our selection of the best LLMs for coding is built on real-world usage feedback, not lab scores — and that's the only method that holds up.

For those hesitating between the two ecosystems, our ChatGPT vs Gemini comparison details what actually changes day to day, beyond Artificial Analysis's numbers.

An important nuance, and it applies to all of this week's news: Bloomberg's claims rest on internal sources, not public tests. On paper, Gemini 4 Argon remains Google's most aggressive offering in terms of cost per task. The doubt concerns the gap between the demo and production — not the model's existence or its claimed qualities.


Cost per task, the market's new compass

Direct answer: the metric replacing benchmark rankings is the cost per task actually accomplished. Hesamation sums it up well: the market benchmark has become a cost per task comparison, not academic scores.

Why now? Because the dominant use of AI in 2026 is no longer the occasional chat — it's the agent. An agent that chains together dozens of calls to accomplish a single task turns a difference in unit price into a massive difference in the bill. A model that's five times cheaper but nearly as good isn't a compromise: it's an obvious financial decision.

Concretely, how do you measure it? Take a typical task from your daily routine — fixing a bug, writing up a report, running a search — and count what it costs end to end, calls included. Tedious the first time, then invaluable for all your future decisions.

This is exactly the calculation that teams building agentic systems are making. Our guide to the best LLMs for AI agents starts from this principle: the best model for an agent is almost never the most expensive one.

Market consequence: labs will no longer compete on "who has the best score," but on "who completes the most tasks per dollar spent." That's a far better war for users. And a far tougher one for provider margins — which brings us to the uncomfortable number.


2.2%: The Wall Facing the Consumer AI Economy

Direct answer: the consumer AI economy doesn't hold up, and the 2.2% figure proves it. According to Andreessen Horowitz's semiannual State of Markets report, cited by TechCrunch on September 30, 2026, only 2.2% of consumers were paying for an AI service in May 2026 — at an average spend of $31 per month.

These figures come from a PNC study conducted this summer. The most concerning part isn't the level, it's the curve: growth is linear, with no jump, even as the models improve. The leap between GPT-5.2 and Astra is barely visible on the spending curves. Performance gains aren't converting free users into paying ones.

The Netflix Calibration, or the Scale of the Chasm

TechCrunch's exercise puts the problem in perspective: even matching Netflix's 325 million subscribers at $34 per customer, you get roughly $11 billion in annual revenue. That's less than a third of OpenAI's operating costs. The consumer market, even fully saturated, cannot finance the infrastructure.

WebProNews adds its "brutal math," backed by Inworld AI data from July 2026: a basic text conversation on a small open source model costs about 9 cents per monthly active user. The same load on OpenAI's flagship speech-to-speech model climbs to $18.24. A single paying user has to subsidize more than 30 free users, and a single high-end voice interaction can wipe out the margin of an entire subscription.

What This Means for You

Concretely: free tiers are going to tighten up. If you live off free tiers, enjoy it while it lasts — our roundup of the best free LLMs lists what's still genuinely usable without a credit card. But don't build any serious workflow on top of them.

And for businesses wondering where the AI money is going: it will go to APIs and professional licenses. The model economy is shifting from consumers, who don't pay, to businesses, who have no choice. It's from this shift that aggressive price cuts like Sol's come: the servers have to be filled, and individual consumers will never manage that alone.


The developers' counterattack: agent architecture and local models

Direct answer: while the labs battle over prices, developers are already cutting their bills themselves — by up to 77%. A popular thread on r/AI_Agents describes the approach: never let the model code directly.

The 77% principle

The LLM plans and decides, but execution stays deterministic. Code is generated once, reviewed, then executed by standard scripts. Simple tasks are routed to cheap models, redundant calls are cached. The claimed result on this type of workload: a 77% reduction in real cost, with GPT-6.1 Sol as the budget brain.

It's not magic, it's architecture. Every premium model call you replace with deterministic code is a call you no longer pay for. The lesson holds for any model: the chattier the agent, the more model choice and routing matter.

The local lever

Second lever: local models. The Inworld AI figure cited earlier says it: 9 cents per month per active user for basic text on a small open source model. With self-hosting, the API bill drops to zero — all that remains is hardware and electricity.

If you've never taken the plunge, our guide to installing a local LLM with Ollama or LM Studio takes an hour to follow, and our roundup of the best LLMs to run locally will tell you what to install based on your machine.

To host the stack — Ollama, your agent scripts, a small vector database — a VPS is more than enough. Hostinger offers plans suited to this kind of budget (prices as of October 2026, check their site): it's the cheapest entry point for a personal agentic infrastructure.

My take: this is the real lesson of the week. The labs' price war is spectacular, but the cost war users are waging is even more so. Those who optimize their architecture today will never look at pricing grids the same way tomorrow.


Who Should Choose What Now (October 2026)

Direct answer: GPT-6.1 Sol becomes the rational default choice, Astra the exception, Gemini 4 Argon a bet you have to validate yourself.

Developers and agent teams: Sol is made for you. Nearly Astra-level at coding and computer use, at a fifth of the price — the math is done. Keep Astra for the steps where an error costs more than the price difference.

Budget-conscious users: Gemini 4 Argon shows the best performance-to-price ratio on paper. But given the doubts reported by Bloomberg about front-end coding and multi-step tasks, test it on your real workloads before migrating anything.

Those who want premium with no compromise: Claude Opus 5.5 remains the quality benchmark. It was, in fact, its launch that triggered the GPT-6 Sol and Luna price war, unleashed at half price 90 minutes later. You pay more, but you're not playing benchmark roulette.

Small budget, big needs? The combination that keeps coming up this week: GPT-6.1 Sol for volume, a local model for trivial tasks, Astra as a rare exception. It's the stack documented by the teams active on Reddit, and it costs a fraction of an all-premium setup.

And if you're lost in the middle of all this, that's normal: three price cuts and one launch in a few weeks — nobody can keep up in real time. Our monthly comparison of the best LLMs exists precisely for that, and the 2026 Claude, GPT, Gemini, Llama comparison gives you the big picture.


❌ Common Mistakes

Mistake 1: Choosing a model based on its benchmark

An excellent Artificial Analysis score tells you nothing about your use case. Gemini 4 Argon matches Astra on the tests, but would reportedly fall behind on front-end coding, according to Google employees cited by Bloomberg. The solution: put together 10 to 20 tasks representative of your actual work, test the candidates on them, and compare the cost per completed task — not the score.

Mistake 2: Confusing price cuts with efficiency gains

GPT-6 Sol's cost reduction came from a 50% price cut, not from technical progress (Hesamation). A commercial price cut can be reversed overnight; an architectural efficiency gain cannot. Always check what's driving a price drop before building a cost strategy on top of it.

Mistake 3: Letting the agent do everything, model included

An agent that asks the LLM to write and then execute every piece of code pays the premium price at every step. The solution documented on r/AI_Agents: LLM for planning, deterministic execution in plain code, routing simple tasks to cheap models, systematic caching. Up to 77% in savings for an afternoon of refactoring.

Mistake 4: Building a product on consumer subscriptions

With 2.2% of consumers paying $31/month and a market unable to cover a third of OpenAI's costs even at Netflix's scale (TechCrunch calibration), AI B2C is a bottomless pit. If you're selling AI, sell it to businesses or via API. That's where the margins are — not with the average consumer.


❓ Frequently Asked Questions

Does GPT-6.1 Sol replace GPT-6 Astra?

No, the two are complementary. Sol is the everyday workhorse: nearly Astra's level in coding, computer use, and professional tasks, at one-fifth the price. Astra remains the flagship for critical steps. OpenAI is deliberately segmenting: premium for the use cases that justify it, the rest in volume at low prices.

Is Gemini 4 Argon a bad model?

No. Independent tests from Artificial Analysis place it at the level of GPT-6 Astra at 60% of the cost per task, with fewer hallucinations according to early tests. The doubt reported by Bloomberg concerns real-world usage — front-end coding and multi-step tasks — according to Google employees. Test it on your own use cases before any migration.

What exactly does the 2.2% figure mean?

According to the PNC study relayed by Andreessen Horowitz's State of Markets report (via TechCrunch, September 2026), 2.2% of consumers were paying for AI in May 2026, with an average spend of $31/month. Growth is linear, with no acceleration despite model progress. The consumer market will never fund the infrastructure on its own.

Can you really save 77% with an agent architecture?

That's a field report from r/AI_Agents, not a universal guarantee. The principle: never let the model code directly, run deterministic code, route simple tasks to cheap models, and cache redundant calls. The gain depends on your workload, but the direction is right in almost every case.

Should you switch to local models now?

For basic text tasks, yes: a small open source model costs about 9 cents per month per active user via API (Inworld AI data, July 2026), and zero when self-hosted. For flagship capabilities — speech-to-speech, complex agents — remote models remain necessary. The ideal is a mix: local for volume, API for peaks.


✅ Conclusion

By slashing prices with GPT-6.1 Sol while Gemini 4 Argon struggles to convince in the field, the AI market has just shifted from benchmarks to cost per task — and that's great news for your bills. To get a clear picture every month, follow our comparison of the best LLMs.