Who this is for
This is for people who buy models and agents the way they buy any other tool: by whether the job actually gets done. Not by a token price on a pricing page, and not by a screenshot of a leaderboard.
If you are staffing daily knowledge work, you need a default that can finish the inbox at a price you can keep paying month after month. If you are staffing the hard tail, the leftover problems that only a few models can touch, you need a different default. Those are not the same purchase, even when the same company is selling both.
Pareto is a set, not a crown
People like to say a model is Pareto optimal. What they usually mean is that it is the one they already prefer. That sounds rigorous. It is not a buying rule.
An unconstrained front is a set of points, not a throne. A cheap model that only does search can sit on it. An expensive model that only does novel research can sit on it too. Neither is the winner. The front only says you cannot beat one point on every axis at once. That is how a cheap incomplete model gets praised. It looks efficient until you notice the last slice of the job never shipped.
Finish the list before you price the time
Start by writing down the jobs you actually need done. Not a soft bench, and not an average across every public exam. Just the work you would pay someone to finish this week.
A model only gets a score if it finishes all of those jobs. Miss the last slice and it does not count, no matter how cheap the tokens were. If only one model can finish the whole list, that model is the only point that matters.
Among the models that do clear the bar, pick the one that costs the least once you include real dollars and real waiting time. Count retries. Count tools. An average price per completed task sounds similar, but it rewards partial coverage. The number you want is the price of finishing the full set. Once every survivor has done that, the ranking is simple: dollars times seconds, among the models that actually did the work.

Four corners, not a ranking
Gemini, Grok 4.6, GPT-5.6 Sol, and Claude Fable 5 can all sit on that open front because they win different corners. Gemini is fast and cheap on search and retrieval. The set it finishes is thinner, and it finishes that set quickly. Grok 4.6 is the daily knowledge-work corner, at about $2 in and $6 out per million tokens. GPT-5.6 Sol is long-horizon coding and agents, where extra effort can buy accuracy, at about $4 to $5 in and $20 to $30 out. Claude Fable 5 is the hardest reasoning and the research tail, at $10 in and $50 out.
Those prices are not a ranking of intelligence. They are the cost of different corners. Name the job list and one corner becomes the buy. Leave the list unnamed and all four stay optimal, which is true and also useless when you are trying to staff a team.

What the benches actually separate
The big aggregate indexes are close. Artificial Analysis has put Grok 4.6 and GPT-5.6 Sol around 61, and Fable around 62. That is not a hierarchy. The split is the suite you care about.
Knowledge-work suites such as GDPVal-AA and AA-Briefcase are where a workhorse earns the name. Grok 4.6 has led or tied those rows at a fraction of Sol and Fable price. That is the inbox most people actually have. Hard software is a different row. Fable still leads SWE-Bench Pro class work, and Sol leads several terminal and effort-scaled agent rows. Grok is good there. It is not the tip. Frontier science is another row again: novel reasoning, long first-principles loops, the jobs Anthropic describes Fable and Mythos doing with its own scientists. That is not the inbox either.
A fair reading is simple. Grok is the default for almost all daily knowledge-worker jobs. Fable and Sol are the default when the job is the leftover hard tail. Gemini is the default when the job is search and snappy retrieval. An unfair reading is that Grok does everything, or that Fable is only for science. Fable also wins brutal software.
On the overlap of work that both Grok and the expensive models can finish, Grok is the cost-time winner. That is the workhorse claim. The long-haul claim is different. It is about the stack under that corner, not the model card alone.
Two different buyers
The labor market and the lab are not the same buyer, even when they spend the same watt. Payroll is a volume problem. It is headcount times hours times ordinary tasks. Route the day to the model that finishes that list at the lowest dollars times seconds, and send only the leftover tail upstairs.
The labs are running a different objective. OpenAI and Anthropic put a large share of gigawatts into research. Claude has authored most of Anthropic merged code, and the engineer merge rate inside that lab has moved by large multiples versus 2024. That is recursive self-improvement in the form they actually have: use the model to write the next model, and to do work only the next model can do.
Their claim is not that expensive tokens beat cheap tokens on the same memo. Their claim is that a smarter model expands the job list. New tasks appear that cheaper models cannot enter, so revenue per watt can stay high even if the token is expensive. Both claims can be true. The labor market wants completed ordinary work per dollar. The lab wants the growth rate of the task set.
The dangerous mix is sending Fable the mail because it is smarter. The other dangerous mix is assuming the workhorse is enough after the valuable work has moved into the tail.
Tokens are a manufactured good
A token is easy to talk about as if it were a slogan. It is not. Someone has to turn electricity into an answer, and the bill is made of familiar parts: the energy in each token, the price of delivered power, the cost of the chips as they wear down, and the cost of running the cluster around them. Power is a local problem. Chips are scarce. Advanced packaging sits in a queue. Inference margins at the model layer have been thin, while NVIDIA still takes most of the profit on merchant accelerators.
A workhorse list price only holds if those inputs sit inside the same P and L. Rent them and the volume price gets competed away. Raise price to cover the rent and you fall off the cost axis. Nobody owns the whole stack from dirt to token. The honest picture as of September 2026 looks like this.
Tesla plus SpaceX owns the tokens and the cluster: Grok, X, Colossus. Electrons are largely owned on site, with gas turbines, solar, Megapacks, and their own substations. Dirt is only partial. There is Texas lithium refining, cathode work, and domestic LFP cells, but not the mines. Silicon is still rented from NVIDIA. Terafab is a plan, not a fleet.
Google owns tokens, cluster, and silicon. TPUs are already in production at scale. Electrons are partial: Intersect Power, nuclear PPAs, not Colossus-speed turbines sitting behind the meter. Dirt is empty at material scale.
OpenAI owns the tokens. The cluster is mostly rented. Silicon is now partial. Jalapeño, designed with Broadcom and taped at TSMC, is an inference chip. Lab samples and August 2026 benches exist. A small deploy is targeted for late 2026, with volume in 2027. Training is still NVIDIA. Electrons and dirt are rented.
Anthropic owns the tokens. The cluster is rented across AWS Trainium, Google TPUs, and NVIDIA. Silicon is empty. More than a million Trainium2 chips sit on Project Rainier. A chip team exists, and Samsung 2nm talks are exploratory, but no Anthropic part is in the fleet.
Google still has the only silicon cell that is actually in production at scale. OpenAI just entered partial. Anthropic has not. Tesla plus SpaceX is empty on silicon until something ships.

Why the column matters
Over time, the column that fills downward is what keeps a workhorse corner on that solid. You want a model that can finish ordinary jobs, stay one edge from Fable and Sol on the tail, and do it at a token cost set by owned electrons and amortization instead of someone else's markup.
Empty cells are rent paid to the chip vendor and to the grid. Those rents are how a volume model falls a generation behind. Custom silicon is the next rent to kill. Google built TPUs to stop paying NVIDIA. Jalapeño is OpenAI version of the same move, not yet the fleet. Tesla plus SpaceX talk Terafab. Same idea.
Speed of standing up power is also product. Colossus came up on turbines and Megapacks while others sat in interconnection queues. Time to a gigawatt is part of the cost of finished work at company scale, not just model scale.
That is the long-haul claim. Grok is already the workhorse on everyday knowledge work. The provider that keeps that corner on the front, at a cost set by its own stack, is the one that can still be the default in five years.
What to do with this
- Write down the jobs you actually buy. Daily knowledge work is not SWE-Pro, and SWE-Pro is not a research exam.
- Drop any model that cannot finish all of them. Then rank dollars times seconds.
- Default the volume to the workhorse on that list. Escalate the tail.
- Do not treat lab research spend as proof that the expensive model should price the inbox.
- If you are choosing a provider for years, look down the column. Tokens are the easy row.
The model card is a corner. The company is a column. The job is the list. Use all three. Do not crown a single Pareto winner.
Sources
xAI Grok 4.6 pricing, August 2026. Artificial Analysis and vendor benches for Grok 4.6, GPT-5.6 Sol, and Claude Fable 5. OpenAI, Jalapeño with Broadcom, 24 June 2026 and first results 25 August 2026. Anthropic and Amazon Trainium expansion, 20 April 2026. Tesla lithium refinery ramp, 2026. SpaceX and xAI Colossus power reporting, 2025-2026. NVIDIA merchant accelerator share estimates, 2026.
