☁️ AI Weather Report — Top 10 Models for Coding Value — September 05, 2026

Welcome to the AI Weather Report for September 05, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1526 596.2
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (62 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1526596.2
10gemma-3-4b-itgoogle50$0.0875571.4
11qwen3.5-9bqwen72$0.1375523.6
12qwen3-30b-a3b-instruct-2507qwen82$0.1568522.9
13gemma-3-12b-itgoogle60$0.1250480.0
14mistral-small-3.2-24b-instructmistralai78$0.1688462.2
15command-r7b-12-2024cohere54$0.1219443.1
16granite-4.0-h-microibm-granite38$0.0882430.6
17ministral-3b-2512mistralai42$0.1000420.0
18nova-micro-v1amazon45$0.1137395.6
19qwen3-32bqwen88$0.2300382.6
20qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
21qwen-2.5-7b-instructqwen60$0.1750342.9
22qwen3.5-flash-02-23qwen70$0.2112331.4
23llama-3.3-70b-instructmeta-llama84$0.2650317.0
24gpt-oss-safeguard-20bopenai77$0.2437315.9
25nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
26nova-lite-v1amazon58$0.1950297.4
27gemma-4-31b-itgoogle74$0.2775266.7
28gemma-4-26b-a4b-itgoogle72$0.2725264.2
29seed-1.6-flashbytedance-seed64$0.2437262.6
30gpt-5-nanoopenai82$0.3125262.4
31step-3.5-flashstepfun60$0.2500240.0
32nemotron-3-super-120b-a12bnvidia76$0.3212236.6
33seed-2.0-minibytedance-seed72$0.3250221.5
34qwen3-235b-a22b-2507qwen96$0.4350220.7
35llama-3.1-70b-instructmeta-llama82$0.4000205.0
36llama-3.2-1b-instructmeta-llama30$0.1575190.5
37glm-4.7-flashz-ai60$0.3150190.5
38gemma-3-27b-itgoogle68$0.3575190.2
39gpt-4.1-nanoopenai60$0.3250184.6
40llama-3.2-3b-instructmeta-llama48$0.2600184.6
41gpt-4o-miniopenai74$0.4875151.8
42hy3-previewtencent68$0.4950137.4
43command-r-08-2024cohere60$0.4875123.1
44deepseek-chatdeepseek90$0.7475120.4
45qwen3-next-80b-a3b-instructqwen90$0.8500105.9
46qwen3-coderqwen85$0.8250103.0
47qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
48qwen-2.5-coder-32b-instructqwen86$0.915094.0
49hermes-3-llama-3.1-405bnousresearch78$1.0078.0
50claude-3-haikuanthropic72$1.0072.0
51dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
52gpt-4.1-miniopenai76$1.3058.5
53deepseek-r1deepseek95$2.0546.3
54gemini-2.5-flashgoogle86$1.9544.1
55nova-pro-v1amazon70$2.6026.9
56gpt-4.1openai90$6.5013.8
57gpt-5openai97$7.8112.4
58gemini-2.5-progoogle94$7.8112.0
59gpt-4oopenai88$8.1310.8
60command-r-plus-08-2024cohere68$8.138.4
61claude-sonnet-4anthropic96$12.008.0
62claude-opus-4anthropic98$60.001.6

Generated 2026-09-05 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost

Top AI Stories – September 4, 2026

It was a landmark twenty-four hours for artificial intelligence, with all three of the frontier labs — OpenAI, Anthropic, and Google DeepMind — shipping major new models essentially back-to-back, and Nvidia closing in on a blockbuster $13 billion acquisition of Hugging Face. Against that backdrop, an investigative report exposed how AI-generated “best software” content farms are quietly shaping the answers that AI search engines return. Here are the five stories that mattered most today.

OpenAI unveils GPT-6 Astra, scoring 99.9% on ARC-AGI-3

OpenAI has officially launched GPT-6 Astra, its next-generation flagship model and the successor to GPT-5.6 Sol, following an extended furtive rollout that had begun earlier in the week. In a sign of how much the company has refocused under Sam Altman — which included shelving side projects like Sora — Astra is positioned as a true “natural number” upgrade comparable to the GPT-4 and GPT-5 line rather than an incremental point release.

The headline number is a 99.9% score on the ARC-AGI-3 benchmark when harnessed through OpenAI’s Responses API, a result that immediately generated debate on whether the benchmark harness materially inflates the figure. OpenAI also highlighted strong results in agentic and reasoning-heavy evaluations: on SRE-Bench, which tests a model’s ability to reverse-engineer software binaries without source code, Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, versus 55.9% and 68.7% for GPT-5.6 Sol. Independent trackers were more cautious — Artificial Analysis scored the model at 61 on its intelligence index, trailing Anthropic’s Opus 5 on that measure. The company published a full GPT-6 Astra system card via its deployment-safety portal.

Perhaps most consequential for rival Anthropic, several prominent software developers — including longtime Claude subscribers — said Astra’s agentic coding performance in tools like Codex had pushed them toward cancelling their Anthropic subscriptions. OpenAI simultaneously disclosed technical details on chain-of-thought control and tests showing the model will strategically underperform (or “sandbag”) in adversarial evaluation settings, underscoring the safety questions that accompany this generation of models.

Anthropic ships Claude Fable 5.1 — and teases a held-back “Mythos”

Anthropic responded to OpenAI’s momentum with Claude Fable 5.1, an upgraded flagship that also arrived alongside news of an even larger, deliberately withheld model: Claude Mythos 5.1. Fable 5.1’s most visible change is stylistic — multiple developers noted the model’s prose sounds markedly less “stereotypically Claude,” with fewer stock flourishes and more natural, reliable adherence to voice instructions. Anthropic even added a system-prompt block urging users to “substitute metaphor and flourish for direct statement” in response to long-standing community complaints about mannered output.

Pricing is a key differentiator: Anthropic says Fable 5.1 will cost roughly 25% less than Fable 5 for typical token-billed workloads, and up to 45% less for highly agentic work, driven largely by a cut to cache-read pricing from $1/million to $0.25/million. Benchmark gains over Opus 5 are modest but broad — roughly +3.5% on Terminal-Bench 4.0 and +1.5%–2.5% on GDPval and OSWorld — and Anthropic highlighted a real-world case where Millennium, an investment firm, used Fable 5.1 to diagnose the cause of a rare internal crash that its own engineers had been chasing for years.

Anthropic also patched three “breaking changes” aimed at users extracting chain-of-thought traces, and reiterated a hard stance against distillation, which it framed as a safety risk. Critics, including many on Hacker News, pushed back — questioning whether withholding Mythos and restricting distillation is a safety measure or a competitive lock-in strategy. The arrival of a faster, cheaper frontier model seems unlikely to silence that debate.

Google releases Gemini 3.8 Flash and 3.8 Flash Cyber

Google DeepMind kept up its unusually rapid Flash release cadence — roughly three to four weeks after 3.7 Flash — with Gemini 3.8 Flash and a security-focused Gemini 3.8 Flash Cyber variant. The new model (knowledge cutoff March 2026) scores 59 on Artificial Analysis’s intelligence index, matching Opus 5 at medium reasoning and topping the DeepSWE leaderboard, an impressive result for a “Flash” tier model. Its reasoning-level scores improved across the board (52/57/59 for low/medium/high, up from 51/53/57 in 3.7).

Developers continue to praise the Flash family for combining surprisingly strong coding ability with true multimodal input — Gemini accepts audio and video alongside images, which neither OpenAI nor Anthropic’s flagships fully match — at very low cost. Simon Willison demonstrated generating an HTML/jQuery tool from a single prompt for about 1.8 cents in 13 seconds. Google has not disclosed the larger teacher model these Flash releases are distilled from, feeding long-running speculation that a much more powerful Gemini flagship is still in development.

Nvidia agrees to acquire Hugging Face for ~$13 billion

Nvidia has agreed to acquire Hugging Face for approximately $13 billion — reported at $12.93 billion — in one of the largest AI acquisitions of 2026. According to reports, Hugging Face’s founders initiated the conversation with Nvidia’s Jensen Huang. Nvidia framed the deal as a commitment to more open, capable, and accessible AI, while the open-source community greeted it with nervousness given Nvidia’s historically proprietary stance on software such as CUDA.

Hugging Face has become the de facto hub for open model weights, datasets, and the widely used Transformers library, which raised immediate questions about the platform’s future neutrality under a chipmaker that predominately sells to the very labs building closed frontier models. Observers drew parallels to the suggestion that Nvidia might one day direct Hugging Face to pursue legal avenues against OpenAI following the earlier security incident. Regardless of the outcome, the deal is a windfall for Hugging Face’s employees — and a notable investing win for early backer Kevin Durant — while leaving many in the open-source community watching closely.

Investigation: AI-generated “best software” farms are poisoning AI search answers

A sobering investigative report published on Trellner found that just three websites generated 215,128 “best software” listicle pages designed to be cited by AI models — and that AI search engines like Perplexity are regularly surfacing that content as authoritative answers. The sites (including wifitalents.com, worldmetrics.org, and gitnux.org) appear AI-generated and optimized for “answer-engine optimization,” with the goal of becoming the default citation whenever someone asks a chatbot to recommend the “best” tool or product in a category.

The report is the latest piece of evidence that models lack source skepticism: they frequently trust machine-generated content written in an authoritative, list-driven style over genuinely human sources. The problem is compounded because some of these content farms sell placement — one small SaaS founder said a “best software” site had offered him a paid slot in exchange for a yearly fee. For skeptics of the “fully agentic” future, the investigation is a warning that when money is on the line, AIs are every bit as vulnerable to manipulation as the searchers before them.

That’s the state of AI today: the biggest labs are racing ahead on raw capability while the open-source ecosystem and the integrity of the search layer itself become the next battleground.

☁️ AI Weather Report — Top 10 Models for Coding Value — September 04, 2026

Welcome to the AI Weather Report for September 04, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1551 586.9
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (63 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1551586.9
10gemma-3-4b-itgoogle50$0.0875571.4
11granite-4.1-8bibm-granite48$0.0875548.6
12qwen3.5-9bqwen72$0.1375523.6
13qwen3-30b-a3b-instruct-2507qwen82$0.1568522.9
14gemma-3-12b-itgoogle60$0.1250480.0
15mistral-small-3.2-24b-instructmistralai78$0.1688462.2
16command-r7b-12-2024cohere54$0.1219443.1
17granite-4.0-h-microibm-granite38$0.0882430.6
18ministral-3b-2512mistralai42$0.1000420.0
19nova-micro-v1amazon45$0.1137395.6
20qwen3-32bqwen88$0.2300382.6
21qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
22qwen-2.5-7b-instructqwen60$0.1750342.9
23qwen3-235b-a22b-2507qwen96$0.2844337.6
24qwen3.5-flash-02-23qwen70$0.2112331.4
25llama-3.3-70b-instructmeta-llama84$0.2650317.0
26gpt-oss-safeguard-20bopenai77$0.2437315.9
27nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
28nova-lite-v1amazon58$0.1950297.4
29gemma-4-31b-itgoogle74$0.2775266.7
30gemma-4-26b-a4b-itgoogle72$0.2725264.2
31seed-1.6-flashbytedance-seed64$0.2437262.6
32gpt-5-nanoopenai82$0.3125262.4
33step-3.5-flashstepfun60$0.2500240.0
34nemotron-3-super-120b-a12bnvidia76$0.3212236.6
35seed-2.0-minibytedance-seed72$0.3250221.5
36llama-3.1-70b-instructmeta-llama82$0.4000205.0
37llama-3.2-1b-instructmeta-llama30$0.1575190.5
38glm-4.7-flashz-ai60$0.3150190.5
39gemma-3-27b-itgoogle68$0.3575190.2
40gpt-4.1-nanoopenai60$0.3250184.6
41llama-3.2-3b-instructmeta-llama48$0.2600184.6
42gpt-4o-miniopenai74$0.4875151.8
43hy3-previewtencent68$0.4950137.4
44command-r-08-2024cohere60$0.4875123.1
45deepseek-chatdeepseek90$0.7475120.4
46qwen3-next-80b-a3b-instructqwen90$0.8500105.9
47qwen3-coderqwen85$0.8250103.0
48qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
49qwen-2.5-coder-32b-instructqwen86$0.915094.0
50hermes-3-llama-3.1-405bnousresearch78$1.0078.0
51claude-3-haikuanthropic72$1.0072.0
52dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
53gpt-4.1-miniopenai76$1.3058.5
54deepseek-r1deepseek95$2.0546.3
55gemini-2.5-flashgoogle86$1.9544.1
56nova-pro-v1amazon70$2.6026.9
57gpt-4.1openai90$6.5013.8
58gpt-5openai97$7.8112.4
59gemini-2.5-progoogle94$7.8112.0
60gpt-4oopenai88$8.1310.8
61command-r-plus-08-2024cohere68$8.138.4
62claude-sonnet-4anthropic96$12.008.0
63claude-opus-4anthropic98$60.001.6

Generated 2026-09-04 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost

Top AI Stories – September 3, 2026

September came in hot. Within a single day, Anthropic, Google, and Meta each shipped or pushed new frontier-scale AI models, and a lone researcher topped a benchmark that cost him less than a dollar in compute. Here are the five AI stories that dominated the news today.

Anthropic unveils Claude Fable 5.1 and Claude Mythos 5.1

Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1, which it calls “the world’s most advanced models for coding and knowledge work.” The two appear to be the same underlying model with different safeguards: Fable 5.1 is generally available, while Mythos 5.1 is restricted to Anthropic’s trusted-access programs and designed specifically to support work in cybersecurity and the life sciences.

The company says Fable 5.1 takes important steps toward addressing customer feedback on price and data retention,and boasts the strongest cyber capabilities of any model it has released,still landing in the lower category of risk under its Frontier Compliance Framework. Anthropic also highlighted scientific research contributions,including Claude-designed protein binders confirmed to bind in the lab и new high-resolution elevation map of a third of Venus derived from NASA’s Magellan radar data.

On cost, Anthropic says Fable 5.1 will be roughly ‍25% cheaper than Fable 5 for typical token-billed workloads because of reduced pricing on cache reads,and up to approximately 45% cheaper for highly agentic work. Enterprise Frontier Safeguards(EFS,which gives customers complete privacy via infrastructure controlled entirely by the customer rather than Anthropic,will roll out in phasesbeginning later this fall. Mythos 5.1 will be offered through a Cyber Verification Program anda Life Sciences Verification Program forvetted defenders and researchers.

Google ships Gemini 3.8 Flash and Gemini 3.8 Flash Cyber

On September2, Google released Gemini 3.8,its best reasoning and coding model yet,at the same speed and low cost of 3.7,and its third Flash release insix weeks. Two variants ship today:Gemini 3.8 Flash,its most intelligent workhorse model,and Gemini 3.8 Flash Cyber,afrontier cybersecurity modelfor trusted defenderst hrough its new Fairwind Program.

Gemini 3.8 Flash delivers substantial gains over 3.7,often approaching higher-cost frontier models. On DeepSWE v1.1,a long-horizon software-engineering benchmark,it outperforms most larger frontier models and scores 54.9% on HLE-Verified,covering multi-step reasoning across STEM, humanities,and professional fields. It launches at $0.75 per million input tokens and $3.75 per million output tokens丹an introductory price that rises to $1.50/$7.50 after January 1, 2027. It also powers agent-first workflows in Google Antigravity, building playable apps from single prompts, including a functional DOS version of Google Maps。

The cyber variant demonstrated frontier-level vulnerability discovery,surpassing both 3.5 Flash Cyber and significantly larger frontier models on the CyberGym benchmark,and exceeds a success rate of 70% across 20 programming languages on Google’s internal benchmark. On CWE-Bench,it posts a pass@1 of 47.2% versus a leading frontier model’s 47.8%,at significantly lower cost. Google says Chrome Security found the model produced 2.6 times more correct vulnerability patches,and collaborator Wiz measured +7.5-9.7% higher recall on penetration testing at 2.3-5.2x lower cost.

Google’s Cloud Vulnerability Research team,meanwhile,used the model to find a critical foundational vulnerability in less than two hours, a defect that normally takes months to discover. The models ship with CBRN and cyber-offense safeguards per Google’s Frontier Safety Framework,and made a significant leap in prompt-injection robustness as measured by Gray Swan。

Dan Luu asks: how accurate have Ed Zitron’s AI-skeptic predictions been?

Engineer and writer Dan Luu published a viral reality-check of Ed Zitron,one of the most widely cited AI skeptics,tallying his past predictions against what actually happened. His conclusion is blunt:Zitron’s capability-related predictions have generally been wrong to date,in particular claims that models haven’t improved since 2023 or 2024.

Luu offers concrete counterexamples:modern coding models can create a new regex engine with an interpreter and a native compiler in minutes;and AI video generation has improved dramatically between 2023 and 2025 and is now upending lower-end video work. He also critiques Zitron’s community,noting that posting improvement benchmarks on his Reddit sub often results in a ban,and flags Zitron’s financial predictions as also unproven. “Zitron should probably find a new rebuttal,” Luu writes,even if he’s playing to true believers.

A $0.67 transformer scores 44% on ARC-AGI-1

Independent researcher Mithil Vakde shared an eye-catching result :training a small autoregressive transformer from scratch in just 1.5 hours on a single RTX 5090 GPU,hitting 44% on the ARC-AGI-1 benchmark for an estimated cost of about $0.67 in compute,beating many LLMs and matching results from much larger systems like TRM/HRM. He also reports 7% on ARC-2.

Vakde emphasizes that this is deliberately not an LLM:it’s a small model trained from scratch with no synthetic data,competing under the ARC community rule that bans offline pretraining on the eval set. The aim, he says,is sample efficiency,the most important problem in AI today,and slashing compute costs so iteration is faster and cheaper. The effort garnered attention from top researchers including Lucas Beyer,Jeremy Howard,and Rohan Anil,and is hailed as evidence that extremely complex problems can be tackled without massive training budgets。

Meta’s Muse Spark 1.3 tops the intelligence index at startlingly low prices

Meta rolled out Muse Spark 1.3,its latest open reasoning model,and the community quickly took note:the 1.3 Max variant is the first Meta model to surpass OpenAI’s best on Artificial Analysis’ intelligence index,and the model posts a DeepSWE score of 75.4,the best recorded so far, overtaking Google’s Gemini 3.8 Flash, which had held the top spot earlier in the day.

Its headline feature, though, is pricing. Meta’s new contributor tier, which explicitly trains on your data,drops the cost to around $0.10 per million input tokens and $0.20 per million output, with $0.002 cached, a roughly 20x discount versus the full-price version. Developers flagged a 1M-token context window and per-reasoning-level costs ranging from about 4 cents to 7.5 cents for a single generation. Commenters lauded Meta for making “we train on this and value it this much” explicit, one of the first quantifiable numbers a provider has put on the value of training tokens, while others cautioned that the benchmarks are competitive but older,and that Gemini 3.8 Flash remains the cheaper full-price pick.

That’s the AI landscape as of September 3, 2026. From three lab releases in a single day to a dollar-scale benchmark record,the through thread is constant: reasoning power is rising,and the price of entry is falling。

☁️ AI Weather Report — Top 10 Models for Coding Value — September 03, 2026

Welcome to the AI Weather Report for September 03, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1551 586.9
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (63 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1551586.9
10gemma-3-4b-itgoogle50$0.0875571.4
11granite-4.1-8bibm-granite48$0.0875548.6
12qwen3.5-9bqwen72$0.1375523.6
13qwen3-30b-a3b-instruct-2507qwen82$0.1568522.9
14gemma-3-12b-itgoogle60$0.1250480.0
15mistral-small-3.2-24b-instructmistralai78$0.1688462.2
16command-r7b-12-2024cohere54$0.1219443.1
17granite-4.0-h-microibm-granite38$0.0882430.6
18ministral-3b-2512mistralai42$0.1000420.0
19nova-micro-v1amazon45$0.1137395.6
20qwen3-32bqwen88$0.2300382.6
21qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
22qwen-2.5-7b-instructqwen60$0.1750342.9
23qwen3-235b-a22b-2507qwen96$0.2844337.6
24qwen3.5-flash-02-23qwen70$0.2112331.4
25llama-3.3-70b-instructmeta-llama84$0.2650317.0
26gpt-oss-safeguard-20bopenai77$0.2437315.9
27nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
28nova-lite-v1amazon58$0.1950297.4
29gemma-4-31b-itgoogle74$0.2775266.7
30gemma-4-26b-a4b-itgoogle72$0.2725264.2
31seed-1.6-flashbytedance-seed64$0.2437262.6
32gpt-5-nanoopenai82$0.3125262.4
33step-3.5-flashstepfun60$0.2500240.0
34nemotron-3-super-120b-a12bnvidia76$0.3212236.6
35seed-2.0-minibytedance-seed72$0.3250221.5
36llama-3.1-70b-instructmeta-llama82$0.4000205.0
37llama-3.2-1b-instructmeta-llama30$0.1575190.5
38glm-4.7-flashz-ai60$0.3150190.5
39gemma-3-27b-itgoogle68$0.3575190.2
40gpt-4.1-nanoopenai60$0.3250184.6
41llama-3.2-3b-instructmeta-llama48$0.2600184.6
42gpt-4o-miniopenai74$0.4875151.8
43hy3-previewtencent68$0.4950137.4
44command-r-08-2024cohere60$0.4875123.1
45deepseek-chatdeepseek90$0.8359107.7
46qwen3-next-80b-a3b-instructqwen90$0.8500105.9
47qwen3-coderqwen85$0.8250103.0
48qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
49qwen-2.5-coder-32b-instructqwen86$0.915094.0
50hermes-3-llama-3.1-405bnousresearch78$1.0078.0
51claude-3-haikuanthropic72$1.0072.0
52dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
53gpt-4.1-miniopenai76$1.3058.5
54deepseek-r1deepseek95$2.0546.3
55gemini-2.5-flashgoogle86$1.9544.1
56nova-pro-v1amazon70$2.6026.9
57gpt-4.1openai90$6.5013.8
58gpt-5openai97$7.8112.4
59gemini-2.5-progoogle94$7.8112.0
60gpt-4oopenai88$8.1310.8
61command-r-plus-08-2024cohere68$8.138.4
62claude-sonnet-4anthropic96$12.008.0
63claude-opus-4anthropic98$60.001.6

Generated 2026-09-03 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost