nextbig.dev
Vancouver, B.C. · Intelligence on AI and the machines that run it
nextbig.dev
The Briefing · Friday, July 24, 2026

Anthropic replaced its flagship model today and did not change the price, while Stripe moved to buy the switch that chooses between models for close to $10bn, the engine tier keeps improving at a flat number and the routing layer just repriced roughly eightfold in two months

Claude Opus 5 takes over the Opus tier with stronger agentic judgement and noticeably less run-to-run variance, at exactly what Opus 4.8 cost: $5 and $25 per million tokens. Hours later, Stripe was reported in talks to acquire OpenRouter near $10bn, against $1.3bn in May, for a company that makes no models and takes about five percent of each call. Reliability turns a model into a part. Parts get sourced through a switch.

12 min read
The Rundown No. 155 · Audio Edition · 5 min All episodesRSSMP3
0:00 / 5:09
VTT
The Big Story
Anthropic replaced its flagship model today and did not change the price, while Stripe moved to buy the switch that chooses between models for close to $10bn — the engine tier keeps improving at a flat number and the routing layer just repriced roughly eightfold in two months

Anthropic replaced its flagship model today and did not change the price, while Stripe moved to buy the switch that chooses between models for close to $10bn — the engine tier keeps improving at a flat number and the routing layer just repriced roughly eightfold in two months

Anthropic replaced the top of its lineup today. Claude Opus 5 takes over from Opus 4.8 as the flagship Opus-tier model, with better agentic judgement, more efficient tool calling, and 84 percent on the Online-Mind2Web computer-use benchmark. It costs five dollars per million input tokens and twenty-five per million output, which is exactly what Opus 4.8 cost. The capability moved. The price did not.

On the same day, The Information's reporting that Stripe is in talks to acquire OpenRouter for close to ten billion dollars spread across every wire that covers this industry. OpenRouter raised in May at $1.3bn. It aggregates more than three hundred models from over sixty providers, lets an application compare prices, switch models mid-flight and fall back automatically when one degrades, and takes roughly five percent of each call. Annualised revenue was around fifty million dollars in March, up from nineteen million at the end of 2025, across more than 1.5 million monthly active developers and tens of billions of inference requests. Roughly eight times the valuation in about two months, for the company that does not make a single model.

Put the two next to each other and the shape of the market is legible. The engine tier improved and held its price, which is what an improving commodity does — quality rises, the number stays, and the buyer's cost of changing suppliers drops to a line in a config file. The switch in front of the engines got repriced eightfold. One detail from the launch explains the mechanism better than any valuation does: Lovable's co-founder measured Opus 5 as twenty-two percent better on their hardest agentic coding tasks and, more importantly, far less variable run to run. Reliability is what turns a model from a personality into a part, and parts get sourced through a switch.

This desk has argued the same thing in the other direction. Our essay two weeks ago made the case for building so that any model can be fired — routing, evaluations and abstractions kept deliberate so no single provider becomes load-bearing. Today a payments company put roughly ten billion dollars on the machinery that does the firing, and it is worth being precise about why the buyer is a payments company. Stripe already processes OpenRouter's transactions. As software starts paying for its own inference, per call, at machine frequency, the thing that meters model consumption and the thing that meters money are converging into one product, and Stripe would rather own that junction than sit beneath it.

The counter-argument is strong enough that it should be stated plainly rather than saved for later. Five percent of somebody else's inference is the most attackable margin in this stack. Cursor has shipped its own routing, Ramp built comparable functionality internally, Databricks introduced routing and reportedly held talks to buy OpenRouter itself. Routing is not hard; what is hard is the integration surface and the billing relationship, and every large platform already owns both. The labs, for their part, would prefer the switch not exist. Ten billion dollars is a claim that a toll booth on commodity traffic holds its position for years. The traffic is certain. The toll is the part being priced, and it is the part with the most people aiming at it.

@anthropicai Read source
The Bill Reaches the Meter

The build-out's electricity bill lands on people who never ordered any compute

Fortune laid out today how the AI boom's power costs are reaching ordinary customers, and the mechanism is not subtle. PJM Interconnection, the largest US grid operator, serving roughly 67 million people from Illinois to Virginia, cleared its 2028-29 capacity auction at $325 per megawatt-day, the maximum its price cap allows, while still coming up about 6.8 gigawatts short of what the grid needs to stay reliable. That is a market saying it cannot buy enough at any permitted price. The regulatory machinery has started moving in response: the Federal Energy Regulatory Commission issued show-cause orders to the six largest grid operators, directing them to defend or rewrite how they handle gigawatt-scale load requests, co-located generation and upgrade-cost allocation, while Virginia began levying a consumption tax of 1.1 cents per kilowatt-hour on datacentre electricity from July 1. Yesterday the White House gathered utility and developer commitments meant to keep the bill with the technology companies. All of it is downstream of the same arithmetic: the load arrived faster than the generation, and somebody pays the difference.

Memory stays scarce into the next decade, by the supplier's own account

Datacentres are on course to absorb roughly seventy percent of global memory output through the rest of 2026 and into 2027, as Samsung, SK Hynix and Micron — together more than ninety-five percent of DRAM — divert wafers toward the high-bandwidth memory that AI accelerators require. The supply response is not coming quickly: IDC puts 2026 DRAM supply growth at sixteen percent and NAND at seventeen, both below the twenty to thirty percent that has historically been normal, and SK Hynix's chief executive Kwak Noh-Jung has said the crunch will probably persist beyond 2030. This is the constraint underneath everything else in this edition. It is why a graphics card sits finished in a warehouse, why Apple restructured how it sells phones, and why a company that raised $8.6bn yesterday to build Chinese memory capacity was able to do so at that size. The token price keeps falling. The silicon the tokens run on does the opposite, on a schedule its own suppliers now describe in decades.

A New Lab Picks a Different Target

Reid Hoffman's new lab is aiming at routine work rather than code

Prentis, a new AI lab co-founded by Reid Hoffman and Mark Pincus, is in talks to raise $100m at a $1bn valuation to build computer-use models, on the premise that automating ordinary routine work is a larger opportunity than automating programming. It is a contrarian position at a moment when nearly every frontier release, including the one Anthropic shipped this morning, leads with coding benchmarks. The reasoning is defensible: coding is where the capability is easiest to measure and where the buyers are most sophisticated, which is exactly why it is the most crowded and least defensible market. The work that happens in a browser at a desk is harder to benchmark, far larger in aggregate, and almost entirely untouched. Whether a new lab can compete on computer use against labs already posting 84 percent on Online-Mind2Web is the open question, and a billion-dollar valuation before the first model is the price of finding out.

Cognition buys Poke, and pays for a personality

Cognition acquired Poke, an AI assistant that lives inside messaging apps, in a deal valuing Poke's parent in the low nine figures. Cognition builds coding agents, an area where capability differences between competitors are narrowing by the month and where every entrant now claims similar benchmarks. What it bought was not a model or a distribution channel of consequential size but a distinctive voice and interaction style, in a market where the underlying intelligence is increasingly interchangeable. That is the commoditisation thesis showing up as an acquisition strategy: when the engines converge, buyers spend on the things that do not — the interface, the tone, the routing, the billing. Same logic as a payments company paying ten billion for a switch, applied to the surface a person actually touches.

Quick Hits
The Takeaway

Anthropic put a better model at the top of its lineup this morning and left the price exactly where it was, at $5 and $25 per million tokens, with the most consequential improvement being less variance run to run rather than a higher benchmark. Within hours, Stripe was reported to be paying close to $10bn for OpenRouter, roughly eight times its May valuation, for the layer that chooses between models and takes about five percent of each call. Read together, that is a market telling you where it thinks the durable margin sits: not in the engine, which improves on schedule at a flat number, but in the switch, the meter and the billing relationship in front of it. We argued the builder's half of this two weeks ago, that you should construct systems so any model can be fired. The other half is that whoever owns the firing mechanism gets paid, which is why the buyer here is a payments company rather than a lab. Underneath all of it the physical bill kept climbing in public: a grid auction clearing at its cap and still short 6.8 gigawatts, regulators rewriting interconnection rules, Virginia taxing datacentre electricity by the kilowatt-hour, and memory suppliers describing a shortage in units of decades. Cheap on top, scarce underneath, and this week the money finally said out loud which layer it wants to own.

The Call C-20260724

Routing gets given away. By June 30, 2027, at least one major cloud or developer platform — AWS, Google Cloud, Azure, Databricks, Vercel, Cloudflare or a company of comparable scale — ships multi-provider model routing with automatic failover at no incremental take rate, bundled into existing platform pricing rather than billed as a percentage of routed inference.

The case

A five percent toll on other people's inference is the most exposed margin in this stack, and the people best placed to attack it already own what makes it valuable. Cursor shipped its own routing, Ramp built the equivalent internally, and Databricks both introduced routing capabilities and reportedly explored buying OpenRouter outright. The hard part was never the technical work; it is the integration surface and the billing relationship, and every large platform holds both already. When a capability's price is a toll rather than a cost, the incumbent with a bundling motive gives it away to protect the surface underneath — that is what happened to CDNs, to CI minutes and to object-storage egress tooling. Stripe paying near $10bn is a claim that this particular toll holds for years. The cheapest possible answer from a cloud is to make routing a free feature.

What proves us wrong

If June 30, 2027 arrives and no major cloud or developer platform has shipped multi-provider routing with automatic failover at no incremental take rate — every one of them still charging a margin on routed third-party inference, or not offering the capability at all — the call is wrong.

Settles by June 30, 2027
The Tape T-20260724
▲ Long TSM TSMC medium conviction

We hold the TSMC long opened yesterday on reported price increases of around ten percent into demand it cannot fully serve. Today's grid and memory news sharpens why the position sits where it does in the chain: power is constrained by regulators and queues, memory is constrained by a supply response its own producers describe in decades, and leading-edge logic is constrained by a single company that has now started collecting for it. Being the bottleneck is only valuable if you can price it, and this is the bottleneck that just did.

Sole leading-edge supply into demand that exceeds it, with price increases being collected rather than negotiated, sitting upstream of every accelerator vendor. The offsets are unchanged: geographic concentration and extreme capital intensity.

Wrong if The reported increases failing to materialise, or advanced-node utilisation slipping as customers defer accelerator orders. Settles 12 months
▲ Long AMD AMD medium conviction

We hold the AMD long from Wednesday. Nothing today touched the thesis and one thing quietly reinforced it: Anthropic shipped a new flagship model two days after committing to up to two gigawatts of AMD silicon, which is the sequence that matters for a second supplier. Adoption follows deployment, and deployment follows a customer willing to go first. ROCm maturity is still the constraint and the first gigawatt is still not until 2027.

A frontier lab shipping on schedule while committing multi-gigawatt capacity to AMD strengthens the case that the second source is operational rather than aspirational. The binding constraint remains software maturity, which the partnership explicitly commits to closing.

Wrong if A slipped or reduced first-gigawatt deployment, or two more quarters without a comparable frontier-lab commitment to Instinct, argues the anchor was bought rather than earned. Settles 12 months
▲ Long MU Micron medium conviction

We hold the Micron long, and today the supplier side made the argument for us. Datacentres are set to take roughly seventy percent of global memory output through 2026 and into 2027, supply growth is running below its historical range on both DRAM and NAND, and SK Hynix's chief executive has put the end of the crunch beyond 2030. A model tier that improves at a flat price does nothing to reduce how much memory the deployment consumes; if anything, cheaper capable models increase the number of them running. The standing risk remains that memory over-corrects on a lag, and CXMT's $8.6bn raise yesterday is when that clock started.

Producers themselves now describe the shortage in multi-year terms while supply growth runs below historical norms, and memory demand is indifferent to which model or accelerator wins. The offset is newly funded Chinese capacity that will eventually arrive.

Wrong if DRAM and NAND contract pricing rolling over before Q4, or Micron's next report showing AI demand failing to offset consumer softness. Settles 6 months
◆ Watch GOOGL Alphabet medium conviction

We hold the Alphabet watch on yesterday's disclosure that it will lease third-party datacentre capacity as a bridge while free cash flow sits negative. Today's grid news is the reason that admission should be read as structural rather than tactical: PJM's capacity auction cleared at its price cap and still fell 6.8 gigawatts short, and FERC has ordered the six largest US grid operators to rewrite how they handle gigawatt-scale load requests. Alphabet's constraint is not its balance sheet, it is the queue, and the queue is now a regulatory proceeding.

The physical constraint behind Alphabet's leased-capacity bridge is being confirmed at the grid level, where auctions clear at caps while falling short and interconnection rules are under federal review. The $514bn backlog remains the offsetting argument that the spending is against contracted demand.

Wrong if Free cash flow returning positive within two quarters while leased capacity winds down retires the concern; another capex raise with cash flow still negative and cloud margins compressing confirms it. Settles 9 months
◆ Watch Private Stripe low conviction

We open a watch on Stripe, private and therefore unbuyable, because the reported price is the cleanest statement anyone has made about where AI margin settles. Paying close to $10bn for OpenRouter, against $1.3bn in May, values a roughly five percent toll on other people's inference as a durable position rather than a temporary one. The strategic case is real: software is starting to purchase its own compute per call, and the meter for model consumption and the meter for money are converging. The risk is equally real and better understood by the buyers than the sellers, since Cursor, Ramp and Databricks have all built or bought routing already. We watch because the outcome here tells us whether the middle of the AI stack is a business or a feature.

The reported multiple prices routing as a durable toll on commodity inference, at a moment when several well-capitalised platforms are building the same capability in-house and would rather bundle it than pay for it. Stripe's existing payment relationship with OpenRouter is the strongest argument that the junction is genuinely one product.

Wrong if The deal closing and routed volume continuing to grow at an intact take rate through 2027 argues the position holds; the deal collapsing, or a major platform bundling routing at no incremental margin, argues it was a feature that got priced as a company. Settles 12 months
Desk signals from the day's verified wire — falsifiable, dated, settled in public. Analysis, not individualized investment advice.

Get this briefing in your inbox

What changed in AI and compute, what it costs, and what to build. One email per week. No spam, unsubscribe anytime.