Skip to content
Artificial intelligence

Best AI models of September 2026: GPT-6 Astra, Claude Fable 5.1, Gemini, Grok, Muse, DeepSeek and MiMo compared

Seven frontier AI models shipped in three weeks of September, and none of them wins at everything. GPT-6 Astra and Claude Fable 5.1 share the top score on the independent index, Astra is the best at operating a computer, Muse Spark 1.3 posts the top score on long coding tasks by Meta’s own numbers, MiMo V2.6 Pro is the strongest open model, and DeepSeek-V4.1-Flash is the cheapest off-peak. We compared them on price, architecture and ten benchmarks.

Cover of our comparison of 7 AI models from September 2026: released in three weeks, GPT-6 Astra and Claude Fable 5.1 score 53 on the AA index, flagships cost 46 times more than MiMo

01Who wins where

CategoryWinnerKey number
General intelligenceGPT-6 Astra and Claude Fable 5.1tied at 53 on the AA Intelligence Index v4.3
Computer useGPT-6 Astra72.6% on OSWorld 2.0
Math and scienceGPT-6 Astra97.6% on FrontierMath Tier 4
Long coding tasksMuse Spark 1.375.4% on DeepSWE v1.1 (Meta’s figure)
Terminal agentGPT-6 Astra57.9% on Terminal-Bench 4.0
Open modelMiMo V2.6 Pro46 on the AA Intelligence Index v4.3
Lowest priceDeepSeek-V4.1-Flash off-peak$0.15 / $0.60
Value among closed modelsGemini 3.8 Flash73.7% on DeepSWE at $0.75 / $3.75
Prices are per million input / output tokens. Figures come from vendors and independent testing; sources are listed in each section. At peak hours DeepSeek costs $0.30 / $1.20, and MiMo V2.6 Pro becomes the cheapest.

We wrote a separate guide to each model, and you’ll find the links in the overview table below. Here we put them side by side.

02Overview: 7 AI models in one table

Four US companies, Meta and two players from China shipped models within three weeks, and their prices differ by as much as 80 times. Only DeepSeek and MiMo publish open weights under the MIT license.

ModelMakerReleasedWeightsContextInput / output (USD per 1M)
GPT-6 AstraOpenAISep 3, 2026closed1.05M10 / 50
Claude Fable 5.1AnthropicSep 1, 2026closed1M10 / 50
Gemini 3.8 FlashGoogleSep 2, 2026closed1M0.75 / 3.75 (through 2026)
Grok 4.7xAI (SpaceXAI)Sep 21, 2026closed500K2 / 6
Muse Spark 1.3MetaSep 2, 2026closed1.05M1.25 / 4.25
DeepSeek-V4.1-FlashDeepSeekSep 10, 2026open (MIT)1M0.30 / 1.20 (peak)
MiMo V2.6 ProXiaomiSep 22, 2026open (MIT)1M0.435 / 0.87
Sources: OpenAI, Anthropic, Agentpedia, xAI, Meta, Hugging Face, Xiaomi. GPT-6 Astra launched as a limited preview on Sep 3 and opened to everyone on Sep 4.

Outside the comparison: Claude Opus 5.5 also came out on September 22. It costs $4 / $20 and tops the entire Artificial Analysis v4.3 leaderboard with 58 points; Anthropic says it works at the level of Fable 5.1 for a lower price. The same day, OpenAI added GPT-6 Sol and GPT-6 Luna. This comparison sticks to the seven models from early and mid-September, and we show Opus 5.5 for context in the intelligence chart. (Anthropic)

03Timeline: September 2026, day by day

September 2026 timeline: Sep 1 Claude Fable 5.1, Sep 2 Gemini 3.8 Flash and Muse Spark 1.3, Sep 3 GPT-6 Astra, Sep 10 DeepSeek-V4.1-Flash with open weights, Sep 21 Grok 4.7 and Sep 22 MiMo V2.6 Pro with open weights
Release dates per vendor announcements. Open-weight models in red.

Four models landed in the first three days of September. Anthropic went first on September 1, Google and Meta followed a day later, and OpenAI shipped on the 3rd. GPT-6 Astra started as a limited preview and opened to all users on September 4. The open models from China came later: DeepSeek on September 10 and Xiaomi not until September 22, a day after Grok.

04Architecture and context window

OpenAI, Anthropic, Google and Meta don’t disclose parameter counts. The Chinese open models show everything: both use a mixture-of-experts (MoE) design, where only a small slice of the network works on each token.

ModelTotal parametersActive per tokenArchitectureMax outputInputs
GPT-6 Astranot disclosednot disclosednot disclosed128Knot stated
Claude Fable 5.1not disclosednot disclosedadaptive thinking128Knot stated
Gemini 3.8 Flashnot disclosednot disclosednot disclosed65,536text, images, audio, video, documents
Grok 4.7about 2.1T (per Elon Musk)not disclosednew, larger base modelno fixed limittext, images
Muse Spark 1.3not disclosednot disclosednatively multimodalnot statedtext, images, video, PDF, limited audio
DeepSeek-V4.1-Flash552B + 196B Engram8B (input), 16B (generation)MoE, causal encoder-decoder384Ktext, images
MiMo V2.6 Pro1.02T42BMoE with a frozen router128Ktext, images, video, audio
Sources: OpenAI, Anthropic, Agentpedia, xAI, Meta, Hugging Face, DeepSeek API, Xiaomi. Grok’s parameter count comes from Elon Musk; xAI has not officially confirmed it.
Context window (thousands of tokens)
  • GPT-6 Astra1,050
  • Muse Spark 1.31,049
  • Claude Fable 5.11,000
  • Gemini 3.8 Flash1,000
  • DeepSeek-V4.1-Flash1,000
  • MiMo V2.6 Pro1,000
  • Grok 4.7500

Source: vendor documentation. Muse Spark 1.3 has exactly 1,048,576 tokens. Grok 4.7 is the only one with half the window, and it charges double from 200K tokens.

A million tokens holds a mid-sized codebase or several hundred pages of documents. What differs is what that long context costs you. Muse Spark 1.3 adds no surcharge for the full window, while Grok 4.7 charges double from 200K tokens. Thanks to compression, DeepSeek needs only 890 bytes of KV cache memory per token, which keeps it cheap even for agents with huge inputs.

05API pricing: which AI model is cheapest

The open models from China are the cheapest. MiMo V2.6 Pro costs $1.31 for a million input tokens plus a million output tokens, while GPT-6 Astra and Claude Fable 5.1 cost $60, roughly 46 times more. DeepSeek-V4.1-Flash costs $1.50 at peak and half that off-peak.

Price of 1M input + 1M output tokens
  • GPT-6 Astra$60.00
  • Claude Fable 5.1$60.00
  • Grok 4.7$8.00
  • Muse Spark 1.3$5.50
  • Gemini 3.8 Flash$4.50
  • DeepSeek-V4.1-Flash (peak)$1.50
  • MiMo V2.6 Pro$1.31
  • DeepSeek-V4.1-Flash (off-peak)$0.75

Per vendor price lists. Gemini at its intro price through the end of 2026, Grok for prompts under 200K tokens, Muse on the Standard tier. DeepSeek’s peak hours are weekdays 1:00 to 4:00 and 6:00 to 10:00 UTC.

The budget five: input and output price per 1M tokens
  • Grok 4.7Input: $2.00Output: $6.00
  • Muse Spark 1.3Input: $1.25Output: $4.25
  • Gemini 3.8 FlashInput: $0.75Output: $3.75
  • DeepSeek-V4.1-Flash (peak)Input: $0.30Output: $1.20
  • MiMo V2.6 ProInput: $0.43Output: $0.87

GPT-6 Astra and Claude Fable 5.1 ($10 / $50) would not fit the scale. MiMo V2.6 Pro’s input price is exactly $0.435.

Full price list

ModelInputCached inputOutputNotes
GPT-6 Astra$10$1$50Fast mode costs double
Claude Fable 5.1$10$0.25$50cheapest cache among the flagships
Grok 4.7$2$0.50$6double from 200K tokens; OpenRouter $1.60 / $4.80
Muse Spark 1.3$1.25$0.15$4.25Contributor tier $0.10 / $0.20, but Meta trains on your data
Gemini 3.8 Flash$0.75not stated$3.75$1.50 / $7.50 from Jan 1, 2027
DeepSeek-V4.1-Flash$0.30$0.006$1.20half price off-peak
MiMo V2.6 Pro$0.435$0.0036$0.87Flash version $0.14 / $0.28
Prices per 1M tokens. Sources: OpenAI, Anthropic, xAI, OpenRouter, Meta, Agentpedia, DeepSeek API, Xiaomi.

Price per token isn’t price per task. GPT-6 Astra and Muse Spark 1.3 need fewer tokens than their predecessors for the same work, and cheap cache reads make Claude Fable 5.1 25% to 45% cheaper than Fable 5 for agent workloads. The real gap on your bill is often smaller than the chart suggests. If you’re budgeting for 2027, plan for the higher Gemini 3.8 Flash price.

06Overall intelligence: which AI model is smartest

On the independent Artificial Analysis Intelligence Index v4.3, GPT-6 Astra and Claude Fable 5.1 share first place among the seven models with 53 points. The best open model is MiMo V2.6 Pro with 46 points. The leaderboard as a whole, though, is led by Claude Opus 5.5, which came out on September 22.

Artificial Analysis Intelligence Index v4.3 (points)
  • Claude Opus 5.5 (outside the comparison)58
  • GPT-6 Astra53
  • Claude Fable 5.153
  • Muse Spark 1.3 (max)48
  • Grok 4.746
  • MiMo V2.6 Pro46
  • Gemini 3.8 Flash41
  • DeepSeek-V4.1-Flash39

Source: Artificial Analysis, as of Sep 27, 2026. Before rounding, Grok 4.7 is just ahead of MiMo (46.45 vs 46.32). Muse Spark 1.3 at the xhigh level scores 45.

Watch out for the older version of the index. Until early September, version v4.1 used a different scale, and many articles still cite it: Claude Fable 5.1 scored 65.7 there and GPT-6 Astra 61.2. Version v4.3 grades harder, and points can’t be converted between versions or mixed in one comparison. This article uses only v4.3, just like our guides.

The gap between 53 and 46 points is clear, but you won’t notice it on every task in everyday work. Results in the area you actually need tell you more, which is why we break them down in the sections that follow.

07Coding: the best AI model for code

On long coding tasks (DeepSWE), the models are nearly tied and the budget ones have caught up with the flagships. The gap only shows up on the harder terminal agent test. On Terminal-Bench 4.0, GPT-6 Astra and Claude Fable 5.1 succeed 1.5 to 3 times as often as the cheaper models.

DeepSWE v1.1: long agentic software tasks
  • Muse Spark 1.375.4%
  • DeepSeek-V4.1-Flash74.2%
  • GPT-6 Astra74.1%
  • Gemini 3.8 Flash73.7%
  • MiMo V2.6 Pro71.9%
  • Grok 4.771.0%
  • Claude Fable 5.167.4%

Vendor figures, each measured in its own setup, so treat the comparison as a rough guide. Claude Fable 5.1 per OpenAI’s table; xAI reports 70.0%, and Anthropic has not published a number.

Terminal-Bench 4.0: general terminal agent
  • GPT-6 Astra57.9%
  • Claude Fable 5.155.8%
  • Grok 4.737.6%
  • MiMo V2.6 Pro34.9%
  • DeepSeek-V4.1-Flash31.2%
  • Gemini 3.8 Flash19.1%

Vendor figures: OpenAI, Anthropic, xAI, Xiaomi, Hugging Face, Agentpedia. Muse Spark 1.3 has not published a score.

Terminal-Bench 2.1: older version of the test
  • DeepSeek-V4.1-Flash90.6%
  • MiMo V2.6 Pro89.9%
  • Gemini 3.8 Flash89.4%
  • Muse Spark 1.388.8%

Vendor figures. GPT-6 Astra, Claude Fable 5.1 and Grok 4.7 have not published a score on this version.

  • DeepSWE barely separates them anymore. Six of the seven models land between 71% and 75.4%, and Claude Fable 5.1 scores 67.4% per OpenAI’s table. Muse Spark 1.3’s top score is also Meta’s own measurement and isn’t on the public DeepSWE leaderboard yet.
  • Terminal-Bench 4.0 splits them into tiers. GPT-6 Astra (57.9%) and Claude Fable 5.1 (55.8%) sit where the cheaper models can’t reach yet. Grok 4.7 scores 37.6%, Gemini 3.8 Flash 19.1%.
  • Terminal-Bench 2.1 is saturated. The budget models cluster around 90%, and the test can’t tell them apart.
  • Independent numbers: CursorBench 3.2 gives Claude Fable 5.1 73.4% and Gemini 3.8 Flash 69.2% (BenchLM), and the AA Coding Agent Index gives Fable 5.1 70 points and GPT-6 Astra 67 (MindStudio).

For routine bug fixes and small code changes, a budget model will do. For hours of unattended work in a terminal, the flagships still earn their price.

08Computer use and AI agents

On computer use, GPT-6 Astra leads OpenAI’s comparison table with 72.6% on OSWorld 2.0. Anthropic reports an even higher number for Claude Fable 5.1, but on a different task set. On the AutomationBench business workflow test, the budget models’ own measurements put them ahead of both GPT-6 Astra and Claude Fable 5.1.

OSWorld 2.0: computer use (vendor figures)
  • Claude Fable 5.1 (partial, Anthropic’s task set)77.9%
  • GPT-6 Astra72.6%
  • Muse Spark 1.3 max (partial)66.9%
  • Gemini 3.8 Flash59.0%

Each vendor used a different task set or grading, so the numbers are not directly comparable. Anthropic used the August 2026 task set (on the strict variant, Fable 5.1 scores 41.7%). Grok, DeepSeek and MiMo have not published a score.

AutomationBench: business workflows
  • DeepSeek-V4.1-Flash54.8%
  • MiMo V2.6 Pro53.1%
  • Muse Spark 1.349.6%
  • GPT-6 Astra41.4%
  • Claude Fable 5.131.4%

Vendor figures: Hugging Face, Xiaomi, Meta, OpenAI, Anthropic. The test version may differ between vendors.

  • OSWorld isn’t one test. Anthropic used the August task set and reports the partial variant. OpenAI used its own setup, and Meta reported its max level. In OpenAI’s table, GPT-6 Astra scores 72.6% and Claude Opus 5 scores 70.2%, which is why we name Astra the winner on computer use.
  • Astra is the fastest. It finishes an OSWorld 2.0 task in about 40 minutes, versus 75 for GPT-5.6 Sol.
  • Budget models’ AutomationBench scores need checking. DeepSeek (54.8%), MiMo (53.1%) and Muse (49.6%) all report their own numbers. For Claude Fable 5.1 (31.4%), OpenAI’s and Anthropic’s figures agree, and GPT-6 Astra (41.4%) is OpenAI’s figure. On a newer version of the test, Xiaomi reports 52.0% for Astra.
  • Ready-made computer use products come mainly from OpenAI (ChatGPT Work, Codex) and Anthropic (Claude Code, Claude Cowork). With the open models, you usually build the agent setup yourself.

09Science, math and reasoning

On science questions the models are close. GPT-6 Astra clearly leads on the hardest math, and Claude Fable 5.1 leads on the multidisciplinary Humanity’s Last Exam.

GPQA Diamond: PhD-level science questions
  • GPT-6 Astra96.0%
  • Gemini 3.8 Flash95.3%
  • Claude Fable 5.193.7%
  • DeepSeek-V4.1-Flash90.9%

Vendor figures; Gemini 3.8 Flash per Artificial Analysis (Google does not report this number).

Humanity’s Last Exam with tools
  • Claude Fable 5.165.0%
  • DeepSeek-V4.1-Flash63.9%
  • GPT-6 Astra57.2%

Vendor figures. The other models have not published this variant of the test.

Only OpenAI and Anthropic published the hardest tests

BenchmarkGPT-6 AstraClaude Fable 5.1What it measures
FrontierMath Tier 4 (v2)97.6%87.8%research-level math
Terminal-Bench Science 0.164.6%52.6%scientific research in a terminal
ARC-AGI-295.0%90.0%abstract reasoning
ARC-AGI-399.9% in a modified harness, 62.7% in the standard testnot statedabstract reasoning
Sources: OpenAI, MindStudio. More detail in our GPT-6 Astra guide and Claude Fable 5.1 guide.

Alongside Astra, OpenAI published new results on gaps between primes and is itself asking mathematicians to check them; Astra officially solved two of 68 open Erdős problems. Claude Fable 5.1 more than doubled Fable 5’s score on Terminal-Bench Science. Xiaomi shows a theorem formalized in Lean 4 with more than 6,000 lines of verified code. None of the models can yet do reliable research on its own. (OpenAI, Anthropic, Xiaomi)

10Cybersecurity and safeguards

The open models report the strongest results at finding vulnerabilities, but they come with no central safeguards. The flagships, on the other hand, are the most tightly guarded. GPT-6 Astra is the first OpenAI model to reach the “Critical” level of cybersecurity capability, which is why the public version refuses to write exploits.

CyberGym: finding vulnerabilities
  • MiMo V2.6 Pro94.0%
  • DeepSeek-V4.1-Flash88.1%
  • Gemini 3.8 Flash Cyber86.2%

Vendor figures: Xiaomi, Hugging Face, Agentpedia. The Gemini entry is the specialized Cyber version. OpenAI, Anthropic, xAI and Meta have not published a score.

Unsafe actions during computer use (lower is better)
  • GPT-5.6 Sol (older)22.0%
  • Claude Opus 5 (older)11.5%
  • Claude Fable 5.19.5%
  • GPT-6 Astra2.4%

OpenAI internal test without production safeguards, so this is one competitor measuring the others.

How each vendor handles safety

ModelApproachStronger version for vetted users
GPT-6 Astrarefuses exploits; classifiers stop unauthorized activityDaybreak program (planned)
Claude Fable 5.1hands sensitive cyber requests to a weaker Opus model; vulnerability finding is allowedClaude Mythos 5.1, so far only in a life sciences program; the cyber program offers Opus and Sonnet
Gemini 3.8 Flashstandard model not tested on cyber benchmarksGemini 3.8 Flash Cyber in the Fairwind program
Grok 4.7new safeguard stack; let through 3.3% of risky requests in an internal testred-team capabilities by invitation
Muse Spark 1.3per Meta, more resistant to prompt injection and careful with irreversible actionsnot stated
DeepSeek-V4.1-Flashopen weights; no central safeguards when self-hostednot needed
MiMo V2.6 Proopen weights; little independent testing so farnot needed
Sources: OpenAI, Anthropic, Agentpedia, xAI, TrueFoundry, MindStudio, Xiaomi.

With open models, the usage rules are your responsibility. With closed models, expect legitimate security work such as penetration testing or exploit development to hit refusals until you get access to a specialized version.

11Open vs closed models: where your data goes

If your data can’t leave your own infrastructure, DeepSeek-V4.1-Flash and MiMo V2.6 Pro are the only options. Both have open weights under the MIT license and can run on your own servers. With every other model, prompts always go to the vendor or one of its cloud partners.

ModelWeightsWhere it runsData protection
GPT-6 AstraclosedOpenAI API, Microsoft Azure, Amazon BedrockZero Data Retention for selected customers
Claude Fable 5.1closedClaude API, Amazon Bedrock, Google Cloud, Microsoft FoundryZero Data Retention for eligible customers
Gemini 3.8 FlashclosedGemini API, Google AI Studio, Gemini appper Google’s terms
Grok 4.7closedxAI API, OpenRouter, Vercel, CloudflareUS regional endpoint for a 10% premium
Muse Spark 1.3closedMeta Model API, OpenRouter, LLM Gatewayon the Contributor tier Meta trains on your data; check your settings in Muse Code
DeepSeek-V4.1-Flashopen (MIT)your own servers, third parties (e.g. NVIDIA), direct APIthe direct API is run by a company based in China
MiMo V2.6 Proopen (MIT)your own servers, OpenRouter, Xiaomi APIXiaomi’s API runs on servers in China
Sources: OpenAI, Anthropic, BenchLM, xAI, TrueFoundry, MindStudio, Xiaomi.

Self-hosting isn’t free. DeepSeek-V4.1-Flash has 552 billion parameters, and even a quantized version needs hundreds of gigabytes of memory. One community build runs on two RTX PRO 6000 cards with 96 GB each. MiMo V2.6 Pro, at 1.02 trillion parameters, is heavier still; the smaller MiMo V2.6 Flash with 309 billion is the more practical choice. (MindStudio)

For EU companies, a model hosted on your own infrastructure in Europe settles the question of where data goes. For any vendor’s API, check the specific data processing terms with your lawyer.

12Performance per dollar: the best-value AI model

For coding, you get nearly the same result for a fraction of the price. DeepSeek-V4.1-Flash matches GPT-6 Astra on DeepSWE and, even at peak, costs one fortieth as much.

Scatter plot of price vs DeepSWE v1.1 score: MiMo $1.31 and 71.9%, DeepSeek $1.50 and 74.2%, Gemini $4.50 and 73.7%, Muse $5.50 and 75.4%, Grok $8 and 71.0%, GPT-6 Astra $60 and 74.1%, Claude Fable 5.1 $60 and 67.4%
The horizontal axis is logarithmic; top left means the most performance per dollar. DeepSWE scores are vendor-reported, DeepSeek at its peak price.

The chart has two limits. DeepSWE only covers long software tasks; on the harder Terminal-Bench 4.0, GPT-6 Astra and Claude Fable 5.1 lead by a wide margin (57.9% and 55.8% vs 31.2% for DeepSeek). And a cheap model that needs three tries can end up costing more than an expensive one that gets it right the first time. So measure the cost per acceptable answer on your own tasks.

13Which AI model should you pick

There is no single best model for everything. It comes down to what you need, and three questions will help:

How to pick an AI model in three questions: data must stay on your servers, then DeepSeek or MiMo; hours of unattended work or computer use, then GPT-6 Astra or Claude Fable 5.1; price matters most, then Gemini 3.8 Flash or Grok 4.7; otherwise a model by field
A simplified guide based on the state of play in September 2026.

Recommendations by use case

Use casePickWhy
Automating work in apps with no APIGPT-6 Astra72.6% on OSWorld 2.0, ready-made ChatGPT Work environment
Large code changes, overnight agents, researchClaude Fable 5.153 points on the AA index, strong in the terminal, cheap cache
Math and scientific computingGPT-6 Astra97.6% on FrontierMath Tier 4
Coding agents at scale on a budgetGemini 3.8 Flash or DeepSeek-V4.1-Flashover 73% on DeepSWE at a fraction of the price
Analyzing a huge codebase or document setMuse Spark 1.398.1% on MRCR up to 1M tokens, no long-context surcharge
Legal and engineering agent workGrok 4.7leads Harvey Legal, second to GPT-6 Astra on EEBench
Finance and biologyGemini 3.8 Flashleads both Vals Finance and LABBench2
Data must stay on your own serversDeepSeek-V4.1-Flash or MiMo V2.6 Proopen weights under the MIT license
Text, images, video and audio in one open modelMiMo V2.6 Pronative support for every input type
Finding vulnerabilities in codeMiMo V2.6 Pro or DeepSeek-V4.1-Flash94% and 88.1% on CyberGym (vendor figures)

Before you commit, run two or three candidates on your own tasks and compare the cost per acceptable answer. If you want to bring AI into your business processes and don’t know where to start, we can help with AI implementation.

14FAQ

What is the best AI model in September 2026?

Of the seven models compared, GPT-6 Astra and Claude Fable 5.1 score highest on the independent Artificial Analysis Intelligence Index v4.3, tied at 53. Astra leads on computer use and math, Fable 5.1 on Humanity’s Last Exam. The leaderboard as a whole is led by Claude Opus 5.5 (58 points), which came out on September 22.

Is GPT-6 Astra or Claude Fable 5.1 better?

It depends on the task. Both score 53 on the Artificial Analysis Index v4.3. Astra leads on computer use (72.6% on OSWorld 2.0), math and Terminal-Bench 4.0, Fable 5.1 on Humanity’s Last Exam. API pricing is the same for both: $10 per million input tokens and $50 per million output tokens.

Which AI model is the cheapest?

Among the models compared, DeepSeek-V4.1-Flash off-peak ($0.15 / $0.60 per million tokens). At peak it costs $0.30 / $1.20, and MiMo V2.6 Pro ($0.435 / $0.87) is cheaper. The smaller MiMo V2.6 Flash is cheaper still at $0.14 / $0.28.

What is the best open source AI model?

Xiaomi’s MiMo V2.6 Pro. It scores 46 on the Artificial Analysis Intelligence Index v4.3, the most of any open-weight model. DeepSeek-V4.1-Flash scores 39 but does better on DeepSWE and costs less off-peak.

Which AI model is best for coding?

On long software tasks the models are nearly tied: six of the seven score 71% to 75.4% on DeepSWE, and Claude Fable 5.1 scores 67.4%. For autonomous work in a terminal, GPT-6 Astra (57.9%) and Claude Fable 5.1 (55.8%) lead on Terminal-Bench 4.0.

Which model has the largest context window?

GPT-6 Astra (1.05 million tokens) and Muse Spark 1.3 (1,048,576 tokens). The rest offer 1 million, except Grok 4.7 at 500,000.

15Verdict: a market in three tiers

At the top are GPT-6 Astra and Claude Fable 5.1. Of these seven, they are the only ones you can trust with several hours of unattended work on a computer or in a terminal. But they cost 46 times more than MiMo V2.6 Pro. If you want similar intelligence for less, Claude Opus 5.5 has been available at $4 / $20 since September 22.

In the middle sit Gemini 3.8 Flash, Grok 4.7 and Muse Spark 1.3. They match the flagships on coding for a seventh to a thirteenth of the price, but they fall short as general agents.

At the bottom on price, not quality, are the open DeepSeek-V4.1-Flash and MiMo V2.6 Pro. They keep pace with the leaders on DeepSWE, cost a fraction, and you can run them on your own servers.

For most companies the sensible setup is a mix: a budget model for volume and a flagship for the hardest jobs. The market shifts week by week, so it pays not to lock yourself into one vendor.

Who wins where: intelligence GPT-6 Astra and Claude Fable 5.1 with 53 points; computer use, math and terminal GPT-6 Astra; long coding Muse Spark 1.3; open model MiMo V2.6 Pro; lowest price DeepSeek-V4.1-Flash; value Gemini 3.8 Flash
The summary in one image, free to share.

16Sources

LISTIFY teamWebsites, apps and marketing from Prague since 2008

More articles

All articles →
Artificial intelligenceSeptember 27, 2026 · 13 min read

The AI Act Hasn’t Been Postponed. What Your Business Needs to Do Now

Artificial intelligenceSeptember 26, 2026 · 14 min read

GPT-6 Astra (ChatGPT 6): the technical guide to benchmarks, API, pricing and access

Artificial intelligenceSeptember 26, 2026 · 15 min read

Claude Fable 5.1: the technical guide to benchmarks, API, pricing and how it compares

Share this page

By email

Got an idea? In 15 minutes, you'll know how to make it happen.

A short call, no sales pitch. We'll tell you what makes sense, what it will cost and how fast we can deliver it.

+420 771 166 199Mon to Fri, 8:30 a.m. to 4:00 p.m. (Prague time) · info@listify.cool

When should we call you?

Pick a day and a time window. We'll call you, and it takes about 15 minutes.

Day