Best AI models of September 2026: GPT-6 Astra, Claude Fable 5.1, Gemini, Grok, Muse, DeepSeek and MiMo compared
Seven frontier AI models shipped in three weeks of September, and none of them wins at everything. GPT-6 Astra and Claude Fable 5.1 share the top score on the independent index, Astra is the best at operating a computer, Muse Spark 1.3 posts the top score on long coding tasks by Meta’s own numbers, MiMo V2.6 Pro is the strongest open model, and DeepSeek-V4.1-Flash is the cheapest off-peak. We compared them on price, architecture and ten benchmarks.

01Who wins where
| Category | Winner | Key number |
|---|---|---|
| General intelligence | GPT-6 Astra and Claude Fable 5.1 | tied at 53 on the AA Intelligence Index v4.3 |
| Computer use | GPT-6 Astra | 72.6% on OSWorld 2.0 |
| Math and science | GPT-6 Astra | 97.6% on FrontierMath Tier 4 |
| Long coding tasks | Muse Spark 1.3 | 75.4% on DeepSWE v1.1 (Meta’s figure) |
| Terminal agent | GPT-6 Astra | 57.9% on Terminal-Bench 4.0 |
| Open model | MiMo V2.6 Pro | 46 on the AA Intelligence Index v4.3 |
| Lowest price | DeepSeek-V4.1-Flash off-peak | $0.15 / $0.60 |
| Value among closed models | Gemini 3.8 Flash | 73.7% on DeepSWE at $0.75 / $3.75 |
We wrote a separate guide to each model, and you’ll find the links in the overview table below. Here we put them side by side.
02Overview: 7 AI models in one table
Four US companies, Meta and two players from China shipped models within three weeks, and their prices differ by as much as 80 times. Only DeepSeek and MiMo publish open weights under the MIT license.
| Model | Maker | Released | Weights | Context | Input / output (USD per 1M) |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Sep 3, 2026 | closed | 1.05M | 10 / 50 |
| Claude Fable 5.1 | Anthropic | Sep 1, 2026 | closed | 1M | 10 / 50 |
| Gemini 3.8 Flash | Sep 2, 2026 | closed | 1M | 0.75 / 3.75 (through 2026) | |
| Grok 4.7 | xAI (SpaceXAI) | Sep 21, 2026 | closed | 500K | 2 / 6 |
| Muse Spark 1.3 | Meta | Sep 2, 2026 | closed | 1.05M | 1.25 / 4.25 |
| DeepSeek-V4.1-Flash | DeepSeek | Sep 10, 2026 | open (MIT) | 1M | 0.30 / 1.20 (peak) |
| MiMo V2.6 Pro | Xiaomi | Sep 22, 2026 | open (MIT) | 1M | 0.435 / 0.87 |
Outside the comparison: Claude Opus 5.5 also came out on September 22. It costs $4 / $20 and tops the entire Artificial Analysis v4.3 leaderboard with 58 points; Anthropic says it works at the level of Fable 5.1 for a lower price. The same day, OpenAI added GPT-6 Sol and GPT-6 Luna. This comparison sticks to the seven models from early and mid-September, and we show Opus 5.5 for context in the intelligence chart. (Anthropic)
03Timeline: September 2026, day by day

Four models landed in the first three days of September. Anthropic went first on September 1, Google and Meta followed a day later, and OpenAI shipped on the 3rd. GPT-6 Astra started as a limited preview and opened to all users on September 4. The open models from China came later: DeepSeek on September 10 and Xiaomi not until September 22, a day after Grok.
04Architecture and context window
OpenAI, Anthropic, Google and Meta don’t disclose parameter counts. The Chinese open models show everything: both use a mixture-of-experts (MoE) design, where only a small slice of the network works on each token.
| Model | Total parameters | Active per token | Architecture | Max output | Inputs |
|---|---|---|---|---|---|
| GPT-6 Astra | not disclosed | not disclosed | not disclosed | 128K | not stated |
| Claude Fable 5.1 | not disclosed | not disclosed | adaptive thinking | 128K | not stated |
| Gemini 3.8 Flash | not disclosed | not disclosed | not disclosed | 65,536 | text, images, audio, video, documents |
| Grok 4.7 | about 2.1T (per Elon Musk) | not disclosed | new, larger base model | no fixed limit | text, images |
| Muse Spark 1.3 | not disclosed | not disclosed | natively multimodal | not stated | text, images, video, PDF, limited audio |
| DeepSeek-V4.1-Flash | 552B + 196B Engram | 8B (input), 16B (generation) | MoE, causal encoder-decoder | 384K | text, images |
| MiMo V2.6 Pro | 1.02T | 42B | MoE with a frozen router | 128K | text, images, video, audio |
- GPT-6 Astra1,050
- Muse Spark 1.31,049
- Claude Fable 5.11,000
- Gemini 3.8 Flash1,000
- DeepSeek-V4.1-Flash1,000
- MiMo V2.6 Pro1,000
- Grok 4.7500
Source: vendor documentation. Muse Spark 1.3 has exactly 1,048,576 tokens. Grok 4.7 is the only one with half the window, and it charges double from 200K tokens.
A million tokens holds a mid-sized codebase or several hundred pages of documents. What differs is what that long context costs you. Muse Spark 1.3 adds no surcharge for the full window, while Grok 4.7 charges double from 200K tokens. Thanks to compression, DeepSeek needs only 890 bytes of KV cache memory per token, which keeps it cheap even for agents with huge inputs.
05API pricing: which AI model is cheapest
The open models from China are the cheapest. MiMo V2.6 Pro costs $1.31 for a million input tokens plus a million output tokens, while GPT-6 Astra and Claude Fable 5.1 cost $60, roughly 46 times more. DeepSeek-V4.1-Flash costs $1.50 at peak and half that off-peak.
- GPT-6 Astra$60.00
- Claude Fable 5.1$60.00
- Grok 4.7$8.00
- Muse Spark 1.3$5.50
- Gemini 3.8 Flash$4.50
- DeepSeek-V4.1-Flash (peak)$1.50
- MiMo V2.6 Pro$1.31
- DeepSeek-V4.1-Flash (off-peak)$0.75
Per vendor price lists. Gemini at its intro price through the end of 2026, Grok for prompts under 200K tokens, Muse on the Standard tier. DeepSeek’s peak hours are weekdays 1:00 to 4:00 and 6:00 to 10:00 UTC.
- Grok 4.7Input: $2.00Output: $6.00
- Muse Spark 1.3Input: $1.25Output: $4.25
- Gemini 3.8 FlashInput: $0.75Output: $3.75
- DeepSeek-V4.1-Flash (peak)Input: $0.30Output: $1.20
- MiMo V2.6 ProInput: $0.43Output: $0.87
GPT-6 Astra and Claude Fable 5.1 ($10 / $50) would not fit the scale. MiMo V2.6 Pro’s input price is exactly $0.435.
Full price list
| Model | Input | Cached input | Output | Notes |
|---|---|---|---|---|
| GPT-6 Astra | $10 | $1 | $50 | Fast mode costs double |
| Claude Fable 5.1 | $10 | $0.25 | $50 | cheapest cache among the flagships |
| Grok 4.7 | $2 | $0.50 | $6 | double from 200K tokens; OpenRouter $1.60 / $4.80 |
| Muse Spark 1.3 | $1.25 | $0.15 | $4.25 | Contributor tier $0.10 / $0.20, but Meta trains on your data |
| Gemini 3.8 Flash | $0.75 | not stated | $3.75 | $1.50 / $7.50 from Jan 1, 2027 |
| DeepSeek-V4.1-Flash | $0.30 | $0.006 | $1.20 | half price off-peak |
| MiMo V2.6 Pro | $0.435 | $0.0036 | $0.87 | Flash version $0.14 / $0.28 |
Price per token isn’t price per task. GPT-6 Astra and Muse Spark 1.3 need fewer tokens than their predecessors for the same work, and cheap cache reads make Claude Fable 5.1 25% to 45% cheaper than Fable 5 for agent workloads. The real gap on your bill is often smaller than the chart suggests. If you’re budgeting for 2027, plan for the higher Gemini 3.8 Flash price.
06Overall intelligence: which AI model is smartest
On the independent Artificial Analysis Intelligence Index v4.3, GPT-6 Astra and Claude Fable 5.1 share first place among the seven models with 53 points. The best open model is MiMo V2.6 Pro with 46 points. The leaderboard as a whole, though, is led by Claude Opus 5.5, which came out on September 22.
- Claude Opus 5.5 (outside the comparison)58
- GPT-6 Astra53
- Claude Fable 5.153
- Muse Spark 1.3 (max)48
- Grok 4.746
- MiMo V2.6 Pro46
- Gemini 3.8 Flash41
- DeepSeek-V4.1-Flash39
Source: Artificial Analysis, as of Sep 27, 2026. Before rounding, Grok 4.7 is just ahead of MiMo (46.45 vs 46.32). Muse Spark 1.3 at the xhigh level scores 45.
Watch out for the older version of the index. Until early September, version v4.1 used a different scale, and many articles still cite it: Claude Fable 5.1 scored 65.7 there and GPT-6 Astra 61.2. Version v4.3 grades harder, and points can’t be converted between versions or mixed in one comparison. This article uses only v4.3, just like our guides.
The gap between 53 and 46 points is clear, but you won’t notice it on every task in everyday work. Results in the area you actually need tell you more, which is why we break them down in the sections that follow.
07Coding: the best AI model for code
On long coding tasks (DeepSWE), the models are nearly tied and the budget ones have caught up with the flagships. The gap only shows up on the harder terminal agent test. On Terminal-Bench 4.0, GPT-6 Astra and Claude Fable 5.1 succeed 1.5 to 3 times as often as the cheaper models.
- Muse Spark 1.375.4%
- DeepSeek-V4.1-Flash74.2%
- GPT-6 Astra74.1%
- Gemini 3.8 Flash73.7%
- MiMo V2.6 Pro71.9%
- Grok 4.771.0%
- Claude Fable 5.167.4%
Vendor figures, each measured in its own setup, so treat the comparison as a rough guide. Claude Fable 5.1 per OpenAI’s table; xAI reports 70.0%, and Anthropic has not published a number.
- GPT-6 Astra57.9%
- Claude Fable 5.155.8%
- Grok 4.737.6%
- MiMo V2.6 Pro34.9%
- DeepSeek-V4.1-Flash31.2%
- Gemini 3.8 Flash19.1%
Vendor figures: OpenAI, Anthropic, xAI, Xiaomi, Hugging Face, Agentpedia. Muse Spark 1.3 has not published a score.
- DeepSeek-V4.1-Flash90.6%
- MiMo V2.6 Pro89.9%
- Gemini 3.8 Flash89.4%
- Muse Spark 1.388.8%
Vendor figures. GPT-6 Astra, Claude Fable 5.1 and Grok 4.7 have not published a score on this version.
- DeepSWE barely separates them anymore. Six of the seven models land between 71% and 75.4%, and Claude Fable 5.1 scores 67.4% per OpenAI’s table. Muse Spark 1.3’s top score is also Meta’s own measurement and isn’t on the public DeepSWE leaderboard yet.
- Terminal-Bench 4.0 splits them into tiers. GPT-6 Astra (57.9%) and Claude Fable 5.1 (55.8%) sit where the cheaper models can’t reach yet. Grok 4.7 scores 37.6%, Gemini 3.8 Flash 19.1%.
- Terminal-Bench 2.1 is saturated. The budget models cluster around 90%, and the test can’t tell them apart.
- Independent numbers: CursorBench 3.2 gives Claude Fable 5.1 73.4% and Gemini 3.8 Flash 69.2% (BenchLM), and the AA Coding Agent Index gives Fable 5.1 70 points and GPT-6 Astra 67 (MindStudio).
For routine bug fixes and small code changes, a budget model will do. For hours of unattended work in a terminal, the flagships still earn their price.
08Computer use and AI agents
On computer use, GPT-6 Astra leads OpenAI’s comparison table with 72.6% on OSWorld 2.0. Anthropic reports an even higher number for Claude Fable 5.1, but on a different task set. On the AutomationBench business workflow test, the budget models’ own measurements put them ahead of both GPT-6 Astra and Claude Fable 5.1.
- Claude Fable 5.1 (partial, Anthropic’s task set)77.9%
- GPT-6 Astra72.6%
- Muse Spark 1.3 max (partial)66.9%
- Gemini 3.8 Flash59.0%
Each vendor used a different task set or grading, so the numbers are not directly comparable. Anthropic used the August 2026 task set (on the strict variant, Fable 5.1 scores 41.7%). Grok, DeepSeek and MiMo have not published a score.
- DeepSeek-V4.1-Flash54.8%
- MiMo V2.6 Pro53.1%
- Muse Spark 1.349.6%
- GPT-6 Astra41.4%
- Claude Fable 5.131.4%
Vendor figures: Hugging Face, Xiaomi, Meta, OpenAI, Anthropic. The test version may differ between vendors.
- OSWorld isn’t one test. Anthropic used the August task set and reports the partial variant. OpenAI used its own setup, and Meta reported its max level. In OpenAI’s table, GPT-6 Astra scores 72.6% and Claude Opus 5 scores 70.2%, which is why we name Astra the winner on computer use.
- Astra is the fastest. It finishes an OSWorld 2.0 task in about 40 minutes, versus 75 for GPT-5.6 Sol.
- Budget models’ AutomationBench scores need checking. DeepSeek (54.8%), MiMo (53.1%) and Muse (49.6%) all report their own numbers. For Claude Fable 5.1 (31.4%), OpenAI’s and Anthropic’s figures agree, and GPT-6 Astra (41.4%) is OpenAI’s figure. On a newer version of the test, Xiaomi reports 52.0% for Astra.
- Ready-made computer use products come mainly from OpenAI (ChatGPT Work, Codex) and Anthropic (Claude Code, Claude Cowork). With the open models, you usually build the agent setup yourself.
09Science, math and reasoning
On science questions the models are close. GPT-6 Astra clearly leads on the hardest math, and Claude Fable 5.1 leads on the multidisciplinary Humanity’s Last Exam.
- GPT-6 Astra96.0%
- Gemini 3.8 Flash95.3%
- Claude Fable 5.193.7%
- DeepSeek-V4.1-Flash90.9%
Vendor figures; Gemini 3.8 Flash per Artificial Analysis (Google does not report this number).
- Claude Fable 5.165.0%
- DeepSeek-V4.1-Flash63.9%
- GPT-6 Astra57.2%
Vendor figures. The other models have not published this variant of the test.
Only OpenAI and Anthropic published the hardest tests
| Benchmark | GPT-6 Astra | Claude Fable 5.1 | What it measures |
|---|---|---|---|
| FrontierMath Tier 4 (v2) | 97.6% | 87.8% | research-level math |
| Terminal-Bench Science 0.1 | 64.6% | 52.6% | scientific research in a terminal |
| ARC-AGI-2 | 95.0% | 90.0% | abstract reasoning |
| ARC-AGI-3 | 99.9% in a modified harness, 62.7% in the standard test | not stated | abstract reasoning |
Alongside Astra, OpenAI published new results on gaps between primes and is itself asking mathematicians to check them; Astra officially solved two of 68 open Erdős problems. Claude Fable 5.1 more than doubled Fable 5’s score on Terminal-Bench Science. Xiaomi shows a theorem formalized in Lean 4 with more than 6,000 lines of verified code. None of the models can yet do reliable research on its own. (OpenAI, Anthropic, Xiaomi)
10Cybersecurity and safeguards
The open models report the strongest results at finding vulnerabilities, but they come with no central safeguards. The flagships, on the other hand, are the most tightly guarded. GPT-6 Astra is the first OpenAI model to reach the “Critical” level of cybersecurity capability, which is why the public version refuses to write exploits.
- MiMo V2.6 Pro94.0%
- DeepSeek-V4.1-Flash88.1%
- Gemini 3.8 Flash Cyber86.2%
Vendor figures: Xiaomi, Hugging Face, Agentpedia. The Gemini entry is the specialized Cyber version. OpenAI, Anthropic, xAI and Meta have not published a score.
- GPT-5.6 Sol (older)22.0%
- Claude Opus 5 (older)11.5%
- Claude Fable 5.19.5%
- GPT-6 Astra2.4%
OpenAI internal test without production safeguards, so this is one competitor measuring the others.
How each vendor handles safety
| Model | Approach | Stronger version for vetted users |
|---|---|---|
| GPT-6 Astra | refuses exploits; classifiers stop unauthorized activity | Daybreak program (planned) |
| Claude Fable 5.1 | hands sensitive cyber requests to a weaker Opus model; vulnerability finding is allowed | Claude Mythos 5.1, so far only in a life sciences program; the cyber program offers Opus and Sonnet |
| Gemini 3.8 Flash | standard model not tested on cyber benchmarks | Gemini 3.8 Flash Cyber in the Fairwind program |
| Grok 4.7 | new safeguard stack; let through 3.3% of risky requests in an internal test | red-team capabilities by invitation |
| Muse Spark 1.3 | per Meta, more resistant to prompt injection and careful with irreversible actions | not stated |
| DeepSeek-V4.1-Flash | open weights; no central safeguards when self-hosted | not needed |
| MiMo V2.6 Pro | open weights; little independent testing so far | not needed |
With open models, the usage rules are your responsibility. With closed models, expect legitimate security work such as penetration testing or exploit development to hit refusals until you get access to a specialized version.
11Open vs closed models: where your data goes
If your data can’t leave your own infrastructure, DeepSeek-V4.1-Flash and MiMo V2.6 Pro are the only options. Both have open weights under the MIT license and can run on your own servers. With every other model, prompts always go to the vendor or one of its cloud partners.
| Model | Weights | Where it runs | Data protection |
|---|---|---|---|
| GPT-6 Astra | closed | OpenAI API, Microsoft Azure, Amazon Bedrock | Zero Data Retention for selected customers |
| Claude Fable 5.1 | closed | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry | Zero Data Retention for eligible customers |
| Gemini 3.8 Flash | closed | Gemini API, Google AI Studio, Gemini app | per Google’s terms |
| Grok 4.7 | closed | xAI API, OpenRouter, Vercel, Cloudflare | US regional endpoint for a 10% premium |
| Muse Spark 1.3 | closed | Meta Model API, OpenRouter, LLM Gateway | on the Contributor tier Meta trains on your data; check your settings in Muse Code |
| DeepSeek-V4.1-Flash | open (MIT) | your own servers, third parties (e.g. NVIDIA), direct API | the direct API is run by a company based in China |
| MiMo V2.6 Pro | open (MIT) | your own servers, OpenRouter, Xiaomi API | Xiaomi’s API runs on servers in China |
Self-hosting isn’t free. DeepSeek-V4.1-Flash has 552 billion parameters, and even a quantized version needs hundreds of gigabytes of memory. One community build runs on two RTX PRO 6000 cards with 96 GB each. MiMo V2.6 Pro, at 1.02 trillion parameters, is heavier still; the smaller MiMo V2.6 Flash with 309 billion is the more practical choice. (MindStudio)
For EU companies, a model hosted on your own infrastructure in Europe settles the question of where data goes. For any vendor’s API, check the specific data processing terms with your lawyer.
12Performance per dollar: the best-value AI model
For coding, you get nearly the same result for a fraction of the price. DeepSeek-V4.1-Flash matches GPT-6 Astra on DeepSWE and, even at peak, costs one fortieth as much.

The chart has two limits. DeepSWE only covers long software tasks; on the harder Terminal-Bench 4.0, GPT-6 Astra and Claude Fable 5.1 lead by a wide margin (57.9% and 55.8% vs 31.2% for DeepSeek). And a cheap model that needs three tries can end up costing more than an expensive one that gets it right the first time. So measure the cost per acceptable answer on your own tasks.
13Which AI model should you pick
There is no single best model for everything. It comes down to what you need, and three questions will help:

Recommendations by use case
| Use case | Pick | Why |
|---|---|---|
| Automating work in apps with no API | GPT-6 Astra | 72.6% on OSWorld 2.0, ready-made ChatGPT Work environment |
| Large code changes, overnight agents, research | Claude Fable 5.1 | 53 points on the AA index, strong in the terminal, cheap cache |
| Math and scientific computing | GPT-6 Astra | 97.6% on FrontierMath Tier 4 |
| Coding agents at scale on a budget | Gemini 3.8 Flash or DeepSeek-V4.1-Flash | over 73% on DeepSWE at a fraction of the price |
| Analyzing a huge codebase or document set | Muse Spark 1.3 | 98.1% on MRCR up to 1M tokens, no long-context surcharge |
| Legal and engineering agent work | Grok 4.7 | leads Harvey Legal, second to GPT-6 Astra on EEBench |
| Finance and biology | Gemini 3.8 Flash | leads both Vals Finance and LABBench2 |
| Data must stay on your own servers | DeepSeek-V4.1-Flash or MiMo V2.6 Pro | open weights under the MIT license |
| Text, images, video and audio in one open model | MiMo V2.6 Pro | native support for every input type |
| Finding vulnerabilities in code | MiMo V2.6 Pro or DeepSeek-V4.1-Flash | 94% and 88.1% on CyberGym (vendor figures) |
Before you commit, run two or three candidates on your own tasks and compare the cost per acceptable answer. If you want to bring AI into your business processes and don’t know where to start, we can help with AI implementation.
14FAQ
What is the best AI model in September 2026?
Of the seven models compared, GPT-6 Astra and Claude Fable 5.1 score highest on the independent Artificial Analysis Intelligence Index v4.3, tied at 53. Astra leads on computer use and math, Fable 5.1 on Humanity’s Last Exam. The leaderboard as a whole is led by Claude Opus 5.5 (58 points), which came out on September 22.
Is GPT-6 Astra or Claude Fable 5.1 better?
It depends on the task. Both score 53 on the Artificial Analysis Index v4.3. Astra leads on computer use (72.6% on OSWorld 2.0), math and Terminal-Bench 4.0, Fable 5.1 on Humanity’s Last Exam. API pricing is the same for both: $10 per million input tokens and $50 per million output tokens.
Which AI model is the cheapest?
Among the models compared, DeepSeek-V4.1-Flash off-peak ($0.15 / $0.60 per million tokens). At peak it costs $0.30 / $1.20, and MiMo V2.6 Pro ($0.435 / $0.87) is cheaper. The smaller MiMo V2.6 Flash is cheaper still at $0.14 / $0.28.
What is the best open source AI model?
Xiaomi’s MiMo V2.6 Pro. It scores 46 on the Artificial Analysis Intelligence Index v4.3, the most of any open-weight model. DeepSeek-V4.1-Flash scores 39 but does better on DeepSWE and costs less off-peak.
Which AI model is best for coding?
On long software tasks the models are nearly tied: six of the seven score 71% to 75.4% on DeepSWE, and Claude Fable 5.1 scores 67.4%. For autonomous work in a terminal, GPT-6 Astra (57.9%) and Claude Fable 5.1 (55.8%) lead on Terminal-Bench 4.0.
Which model has the largest context window?
GPT-6 Astra (1.05 million tokens) and Muse Spark 1.3 (1,048,576 tokens). The rest offer 1 million, except Grok 4.7 at 500,000.
15Verdict: a market in three tiers
At the top are GPT-6 Astra and Claude Fable 5.1. Of these seven, they are the only ones you can trust with several hours of unattended work on a computer or in a terminal. But they cost 46 times more than MiMo V2.6 Pro. If you want similar intelligence for less, Claude Opus 5.5 has been available at $4 / $20 since September 22.
In the middle sit Gemini 3.8 Flash, Grok 4.7 and Muse Spark 1.3. They match the flagships on coding for a seventh to a thirteenth of the price, but they fall short as general agents.
At the bottom on price, not quality, are the open DeepSeek-V4.1-Flash and MiMo V2.6 Pro. They keep pace with the leaders on DeepSWE, cost a fraction, and you can run them on your own servers.
For most companies the sensible setup is a mix: a budget model for volume and a flagship for the hardest jobs. The market shifts week by week, so it pays not to lock yourself into one vendor.

16Sources
- OpenAI: GPT-6 Astra
- Anthropic: Claude Fable 5.1 and Mythos 5.1
- Anthropic: Claude Opus 5.5
- Agentpedia: Gemini 3.8 Flash
- Artificial Analysis: model leaderboard
- xAI: Grok 4.7
- OpenRouter: Grok 4.7
- Meta: Muse Spark
- TrueFoundry: Muse Spark 1.3
- Hugging Face: DeepSeek-V4.1-Flash
- DeepSeek: API pricing
- MindStudio: running DeepSeek-V4.1-Flash locally
- Xiaomi: MiMo-V2.6
- MindStudio: GPT-6 Astra benchmarks
- BenchLM: Gemini 3.8 Flash