nextbig.dev
Vancouver, B.C. · Intelligence on AI and the machines that run it
nextbig.dev
The Briefing · Monday, July 27, 2026

Kimi K3 puts 2.8 trillion parameters and a million-token window into open weights, but activates only 104bn at a time, making the frontier portable before it is cheap

Moonshot AI's Kimi K3 activates 16 of 896 experts, uses 104bn of its 2.8tn parameters per token and claims 2.5 times the scaling efficiency of K2. Its published results reach 88.3 on Terminal-Bench 2.1 and 94.5 on MCPMark, close enough to closed leaders that the deployment harness now matters as much as access to weights.

7 min read
The Rundown No. 158 · Audio Edition · 2 min All episodesRSSMP3
0:00 / 1:50
VTT
The Big Story
Kimi K3 puts 2.8 trillion parameters and a million-token window into open weights, but activates only 104bn at a time, making the frontier portable before it is cheap

The number that makes Kimi K3 possible is not 2.8tn. It is 104bn, the share of parameters active for each token. Moonshot AI built the open-weight model with 896 experts and selects 16 at a time, pairing that sparse mixture with Kimi Delta Attention and Attention Residuals. The company says the design improves overall scaling efficiency by about 2.5 times over Kimi K2 while supporting a one-million-token context window and native image input. Sparsity is what lets a model grow faster than the compute used on each step.

Its own benchmark table puts K3 inside the frontier group rather than beneath it. The model scores 88.3 on Terminal-Bench 2.1, beside 88.8 for GPT-5.6 Sol and 88.0 for Claude Fable 5. On MCPMark it posts 94.5, ahead of the closed models listed, while FrontierSWE lands at 81.2 against Fable's 86.6. Vendor tables always deserve replication, especially when harnesses differ, but these are narrow gaps across work that requires tools rather than polished single-turn answers. K3 also reaches 91.2 on BrowseComp and 84.8 on OSWorld-Verified in Moonshot's runs.

The weights change who can inspect and adapt the engine. They do not make a 2.8tn-parameter system fit under a desk. K3 uses MXFP4 weights and MXFP8 activations, and Moonshot recommends vLLM, SGLang or TokenSpeed for serving. Even at low precision the full model implies an infrastructure job involving many accelerators, fast interconnect and disciplined routing. Open here means deployable without the vendor's API, not inexpensive or operationally simple. The repository is a transfer of control; the cluster remains a substantial bill. Operators still have to place 93 layers across real machines.

That limitation is also the commercial opening. Cloud and inference providers can turn a difficult open deployment into a catalog item, then compete on latency, regional control, observability and price. Teams gain credible negotiating power with model vendors without having to own a cluster. Moonshot gains distribution beyond its own endpoint. The providers lose some engine-level differentiation, but they keep the integration and operations margin that managed open-source software has supported for two decades. A hard model to serve can produce a very good managed service.

A builder evaluating K3 should separate three questions that launch coverage tends to merge: are the weights available for the intended use under the Kimi K3 License, does the model clear an internal task set using the incumbent's harness and permissions, and can a provider meet the memory, latency and residency requirement at acceptable cost? The first answer arrived today. The second takes an evaluation week. The third is why the likely winner from an open 3T-class model is the platform that can reliably make 2.8tn parameters feel like one ordinary API route. That provider owns the mundane work between downloadable weights and a dependable endpoint.

@Kimi_Moonshot Read source
The Gateway Becomes the Product

Satya Nadella tells companies to keep their context separate from any one model

Microsoft chief executive Satya Nadella warned that companies relying on one AI system for everything may not survive, and argued for retaining prompts, context and metadata independently from the model that processes them. The advice is strategically self-serving because Azure wants to host every provider, yet it is also sound systems design. A model gateway can preserve policy, evaluations and audit data while engines change underneath it. The important boundary is the harness: tool permissions, retrieval and memory should belong to the application, with the model treated as a replaceable compute dependency. Kimi K3 makes that architecture useful rather than theoretical, because open weights provide an exit route alongside commercial APIs.

Windows gives JavaScript and TypeScript first-class paths into native AI work

Microsoft laid out a set of Windows improvements for JavaScript and TypeScript developers, broadening the routes from familiar web tooling into native applications and local AI features. The strategic value is distribution. JavaScript teams already own the interfaces where agent workflows appear, but native packaging, device APIs and model access have historically pushed them into a second toolchain. Reducing that translation cost gives Microsoft another way to make Windows the harness even when the model comes from elsewhere. It also tests Nadella's portability claim: first-class support matters only if developers can move context and tool logic between local, Azure and third-party engines without rewriting the application boundary.

Defense Moves Into the Harness

Nvidia's open security alliance defines the agent as more than its weights

Nvidia launched the Open Secure AI Alliance around an explicit claim: an agent is a stack of models, harnesses and guardrails, so security must cover identity, permissions, isolation, logs and evaluation rather than rely on closed weights. Contributions include Nvidia's NOOA research framework, HPE's work around SPIFFE and SPIRE identity, Hugging Face's Safetensors, Red Hat and IBM's signed-patch work, and Microsoft's multi-model MDASH scanner. That breadth is the point. Open models give defenders inspectable engines, while open controls give them a way to test what the engine can reach. The alliance will matter if those pieces become interoperable defaults rather than a list of member projects.

Microsoft uses several models to argue security quality comes from routing

Microsoft introduced AI security tools built around specialized models that discover, debate and verify vulnerabilities instead of asking one general model to perform every step. The design is a security version of the gateway Nadella described: route each subtask to the system best suited to it, preserve evidence and require a proof before a finding becomes action. It also creates a harder evaluation problem. A strong aggregate score can hide one brittle handoff between agents, and a convincing debate can still share the same blind spot across models trained on similar code. Teams should inspect routing policy, reproducible evidence and false-positive cost before accepting a leaderboard gain as an operational control.

Quick Hits
The Takeaway

Kimi K3 turns model portability from procurement language into a cluster-sized engineering option. The weights are available, the active path is 104bn parameters and the vendor's results sit close to closed leaders, but deployment still requires memory, networking and a harness that can reproduce the work. Nadella's instruction to retain context outside any one model and Nvidia's decision to organize agent security around identity, permissions and logs both point to the layer worth owning: the control plane around the engine. That control plane must do more than change an endpoint. It needs comparable evaluations, scoped tools, evidence that survives routing, and costs that can be measured per task. An open model supplies negotiating power only when the rest of the application is portable enough to use it. The immediate action is to run K3 through the same gateway, permissions and task set as the incumbent, then price the infrastructure honestly. A 2.8tn-parameter release lowers dependence on a lab before it lowers the operations bill.

The Call C-20260727

By October 31, 2026, at least one of AWS, Azure, Google Cloud, Databricks or Cloudflare will list Kimi K3 as a fully managed production model, including hosted weights and a supported endpoint rather than a customer-operated cluster recipe.

The case

K3's published scores make it useful enough to attract enterprise evaluations, while its 2.8tn parameters make self-operation difficult enough to preserve a managed margin. Cloud platforms can offer model portability without surrendering the gateway, observability or infrastructure relationship. Moonshot gains distribution and a reference deployment, and customers gain negotiating power against closed API suppliers. The incentives align around a catalog listing faster than around widespread private clusters.

What proves us wrong

If October 31 arrives without a managed Kimi K3 listing from any of the five named platforms, and they offer only do-it-yourself deployment instructions or third-party marketplace images, the call is wrong.

Settles by October 31, 2026
The Tape T-20260727
◆ Watch MSFT Microsoft medium conviction

We sharpen the Microsoft watch opened July 13. Nadella is describing the model-independent architecture Azure is positioned to sell: customer context and policy stay in a gateway while engines rotate underneath. Windows support for JavaScript and TypeScript broadens the application surface, and multi-model security supplies a concrete routed workload. The risk is that Microsoft's economic interest in OpenAI keeps product behavior less neutral than the architecture sounds.

Azure can collect infrastructure, identity and observability revenue across competing models, making portability economically attractive. Microsoft must prove that context, evaluations and tool permissions actually move cleanly between its own and third-party engines.

Wrong if Material Azure features remaining exclusive to one model family, or customers needing to rebuild memory and policy to switch engines, would reduce the gateway thesis to positioning. Settles 6 months
◆ Watch Private Moonshot AI medium conviction

We open a private-company watch on Moonshot AI. Kimi K3 puts an open 3T-class model within narrow benchmark distance of closed leaders and gives providers a substantial managed-serving opportunity. Distribution is the hinge. If major clouds list it, Moonshot can become an engine supplier without owning every customer relationship. If serving remains specialist work, the weights will be influential while the commercial endpoint stays peripheral.

The model combines a 104bn active path, one-million-token context and competitive tool benchmarks with infrastructure requirements that reward managed deployment. Open weights widen adoption, but the Kimi license and independent benchmark replication still determine enterprise comfort.

Wrong if No major managed listing by October, or independent evaluations showing large regressions outside Moonshot's chosen harnesses, would weaken the case for broad commercial distribution. Settles 9 months
Desk signals from the day's verified wire — falsifiable, dated, settled in public. Analysis, not individualized investment advice.

Get this briefing in your inbox

What changed in AI and compute, what it costs, and what to build. One email per week. No spam, unsubscribe anytime.