There is now a supermarket for machine intelligence, and like any supermarket its shelves are wildly uneven. The same unit of "genius in a data center" can cost you pennies or a fortune depending only on which box you reach for.
We all need LLMs
We are now entering an era when LLMs are no longer seen as a strange gimmick but as a genuinely useful assistant and co-worker. Our use of them is no longer a one-off; it is steady and expanding. In a couple of years we have gone from debating whether to pay $20 a month for an exotic thing to learning how to budget our LLM spend in tokens.
In a recent blog post, Andrew Ng, co-founder of Google Brain and Coursera and one of the most influential teachers and researchers in modern AI, argued that once Anthropic and the US government showed that access to the best models can be revoked on short notice, the incentive for everyone else to invest in open-source alternatives grows sharply, though he cautioned that training frontier models is hard and how far those alternatives get remains to be seen.
Time for aggregators
In these new times one cannot afford to stay ignorant about OpenRouter and the other model aggregators, or about how jagged the current pricing of these models really is.
OpenRouter is the best known aggregator, a single API that reaches over 300 models from OpenAI, Anthropic, Google, Meta, Mistral and dozens of other providers, charging the model's own price plus a small fee on top. But it is no longer alone. WaveSpeedAI offers close to 300 language models through one interface, and adds a far larger multimodal catalog beside them. AI/ML API reaches past 400 models and spans text, image, video and audio on the same pay-as-you-go billing. Eden AI takes a different angle, built in France with European data residency by default, a detail that matters more by the month. SiliconFlow competes on raw speed, promising faster inference and lower latency than its rivals. And LLM Gateway arrives as the open-source answer to OpenRouter, a unified API you can run yourself. The shelves of this supermarket, in other words, are filling fast.
20 supermodels
Ethan Mollick, a Wharton professor and AI influencer, recently tweeted his appreciation of the AA-Briefcase benchmark, and I will show you it’s list along with details on models and their pricing. The list is sorted from highest rating to lowest, and the difference between neighbouring places is not dramatic. Each model name carries its country in parentheses; cost per million tokens appears as ¢50 / $2.5, meaning 50 cents for input and $2.50 for output; context appears as 1M/256k, meaning a 1-million-token window and a 256k maximum output.
Claude Fable 5 (USA): access was switched off.
Claude Opus 4.8 (USA): $5 / $25, context 1M/128k.
GLM-5.2 (China): $1.20 / $4.10, context 1M/131k.
GPT-5.5 (USA): $5 / $30, context 1M/128k.
MiniMax-M3 (China): ¢30 / $1.20, context 1M/512k.
Claude Sonnet 4.6 (USA): $3 / $15, context 1M/128k.
DeepSeek V4 Pro (China): ¢43.5 / ¢87, context 1M/384k.
Qwen3.7 Max (China): $1.25 / $3.75, context 1M/65k.
Gemini 3.5 Flash (USA): $1.50 / $9, context 1M/65k.
MiMo-V2.5-Pro (China): ¢43.5 / ¢87, context 1M/131k.
DeepSeek V4 Flash (China): ¢9 / ¢18, context 1M/65k.
Kimi K2.6 (China): ¢67 / $3.50, context 262k/262k.
Grok 4.3 (USA): $1.25 / $2.50, context 1M, no fixed output cap.
GPT-5.4 mini (USA): ¢75 / $4.50, context 400k/128k.
Muse Spark (USA): 0/0, context 262k.
Claude 4.5 Haiku (USA): $1 / $5, context 200k/64k.
Mistral Medium 3.5 (France): $1.50 / $7.50, context 262k.
Gemini 3.1 Pro Preview (USA): $2 / $12, context 1M/65k.
Gemma 4 31B (USA): 0/0.
AA-Briefcase, by Artificial Analysis, tests models on realistic long-horizon knowledge work rather than disconnected prompts. Built over months by experts from Google, McKinsey and BCG, its scenarios run as multi-week projects with linked tasks and thousands of fragmented inputs: company documents, meeting transcripts, data exports, and tens of thousands of Slack messages and emails, often messy and contradictory. Models must produce real deliverables such as financial models, board presentations, and design mock-ups. Grading combines binary rubric checks for correctness with pairwise scoring of analytical and presentation quality, exposing outputs that look polished but are wrong.
Shopping the jagged frontier
So treat this like any supermarket: do not reach for the priciest box out of habit. The dearest model is rarely 10 times smarter than one costing a tenth as much, and for most tasks the cheaper shelf is more than good enough. Match the model to the job, keep an eye on the price-to-intelligence ratio, and let the jagged frontier work for you rather than against you. Being smart about intelligence is now a skill of its own.