The Wire · Meta
Llama and Meta's open-weight models and AI research — the open end of the frontier, tracked for builders.
Moonshot AI released Kimi K3, the largest open-weight model yet: a 2.8-trillion-parameter mixture-of-experts (16 of 896 experts active), 1M-token context, native vision, Kimi Delta Attention and Attention Residuals, in K3 Max and K3 Swarm variants. Moonshot's own benchmarks show K3 beating Anthropic's Opus 4.8, while independent Artificial Analysis scores it clearing Opus 4.8 and GPT-5.5 but losing to Fable 5 and GPT-5.6 Sol; it lists at $3/$15 per million tokens, roughly half Opus's cost per task. Ask it who it is and it says "I am Claude" — the tell of the distillation Anthropic documented in February (3.4M exchanges from Moonshot, 16M across three Chinese labs via 24,000 fake accounts). Also: Mira Murati's Thinking Machines ships its first model, the 975B open-weight Inkling; an autonomous AI agent breaches Hugging Face and is caught with the open-weight GLM 5.2; and NotebookLM becomes Gemini Notebook with a sandboxed cloud computer in every notebook.
Meta seems to be having a bit of an identity crisis. On Monday, the social networking singularity said it would spend $50 billion to expand its Hyperion datacenter project in Richland Parish, Louisiana, from 2.2 to 5 gigawatts. The news comes less than a week after a report broke claiming that Meta was actively exploring options to offload its excess compute capacity to other AI labs. So, which is it, Zuck? Did you invest too much or too little in AI? The easy answer is that Meta overcommitted. Inspired by the early success of Llama, it made a huge bet on the AI gold rush. Offloading spare compute to the highest bidder is just a hedge in case its Superintelligence team turns out to be another pipe dream, like the Reality Labs Metaverse that utterly failed to spark enthusiasm for immersive environments accessible through Meta's Quest cybergoggles. The more pragmatic read is that Zuckerberg has woken up to the fact he’ll never be as cool as OpenAI boss Altman or Anthropic's Amodei, and renting out spare compute is just the natural progression for any sufficiently large hyperscaler. Dawn of the Meta cloud? Meta's business model is closer to Google's than those operated by OpenAI and Anthropic. Both Meta and Google offer various services which generate revenues by connecting users with advertisers. For Google it’s a search and entertainment empire. For Meta it's enabling an endless feed of content generated by friends, family, influencers, and yes, bots. Both are immensely profitable, earning $132.2 billion and $60.5 billion in profits last year, respectively. That's profit, not revenue. But both are now plowing over $100 billion a year into AI infrastructure to power large language and image and video generation models. As we learned from Meta’s recent earnings calls, the most commercially potent of those models get the right ads in front of the right eyeballs. The open secret is Meta was already one of the most successful AI companies long before ChatGPT debuted. Exce
Read full story →Meta has said it will expand its Hyperion data center in Richland Parish, Louisiana, to 5 GW (gigawatts) of compute capacity from an initial 2 GW, pushing the company’s planned investment in the region beyond $50 billion. The announcement — made in an official blog post on Monday, July 13 — confirms the long-signaled scale-up of what is already Meta's largest data center. Go deeper with TH Premium: AI and data centers (Image credit: Microsoft) Photonics and high-speed data movement is the next big AI bottleneck The data center cooling state of play Massive AI data center buildouts are squeezing energy supplies Ultra Ethernet: The data center interconnection of tomorrow The expansion will be a major increase over the $10 billion, 4-million-square-foot project Meta unveiled in December 2024, when it said the campus would deliver more than 2 GW of capacity. However, the 5GW target itself is not entirely new. CEO Mark Zuckerberg said in July 2025 that Hyperion would eventually reach that scale. Monday’s announcement formally ties the expanded capacity to an investment exceeding $50 billion and provides updated figures for jobs, contracts, and public infrastructure spending. Much of the announcement is built around local economic impact. Meta said local Louisiana businesses have received more than $1.6 billion in contracts since construction began, while also highlighting teacher bonuses in Richland Parish that rose from $10,000 last year to more than $50,000 this year, funded by increased tax revenue tied to the data center. In what appears to be a bid to pacify anti-data-center sentiment further, Meta said it plans to invest over $1 billion in local infrastructure improvements, including roads, water, and wastewater systems, as part of the expansion. The company’s recent agreement with utility Entergy Louisiana includes natural-gas plants providing more than 5.2 GW of capacity and support for up to 2.5 GW of new solar generation. Entergy claims Meta’s payments could sa
Read full story →Meta has withdrawn the first image generation product created by its Superintelligence Labs fewer than 72 hours after launch. The product was called “Muse Image” and Meta launched it on July 8, billing it as “the first AI image generation model from Meta Superintelligence Labs.” That lab is Zuck’s latest big bet and aims to create a “personal superintelligence that knows us deeply, understands our goals, and can help us achieve them.” In the case of Muse AI, that help came in the form of applying one of 30 new filters that “uniquely understand Instagram videos and photos and can interpret your photo's scene — the lighting, composition, and subject — to make nuanced edits that feel natural and true to you.” Instagram users could use those effects to “transform your photos with a single tap,” the social networking company promised. Users could also apply the filters to content posted by third parties. “Meta is also launching the ability to @mention friends’ public Instagram accounts in Meta AI and generate creative AI images featuring them, such as personalized birthday cards, group trip memes, or playful edits between friends. It's an upgrade to what people can create with AI features at Meta, making it more personal, fun, and social,” the company said. Meta almost certainly leads the world in three things: The number of people signed up to its social networks; experience of people behaving horribly online, and; dealing with community backlashes after privacy abuses. Yet somehow it didn’t imagine that enabling this feature by default might be controversial, or that allowing users to alter images with AI might be abused. Backlash was therefore swift and widespread. Actors’ union SAG-AFTRA condemned the product. “Anything other than a clear and conspicuous OPT-IN for these types of uses of Instagram users’ images is unacceptable, and an utter miscalculation of public sentiment regarding the obvious dangers and harms inherent in such use,” it posted on Instagram. Within
Read full story →Lorde was performing at the Real Cool Festival in Madrid on Thursday and took some time during her set to speak out against AI glasses. While she didn't specify any brands in particular, it's likely she was taking a shot at festival sponsor Ray-Ban, which has collaborated with Meta on a pair of AI smartglasses. […]
arXiv:2607.14171v1 Announce Type: new Abstract: Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes. State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topology from RLHF: for each prompt, N independent trajectories are sampled from the initial state, and an advantage is computed by subtracting a group baseline. This design ignores a defining property of agent sandboxes. They are deterministic, snapshottable, and resumable from any intermediate state. We argue that this property enables a fundamentally different rollout topology: rather than N independent trees of depth T, one can construct a single tree of N leaves whose siblings share prefixes, and therefore share variance. We instantiate this idea as Branching Policy Optimization (BPO), a sandbox-native RL algorithm that (i) adaptively snapshots the sandbox at high-entropy decision points along a backbone trajectory, (ii) forks K alternative actions per branch point and rolls out each to termination, and (iii) computes per-step advantages from sibling returns rather than from independent prompts. We prove this estimator is unbiased and has strictly lower variance than the trajectory-level baseline, with the reduction equal to the prefix-explained portion of return variance. On WebShop, ALFWorld, and SWE-bench Verified with Qwen2.5-7B and Llama-3.1-8B backbones, BPO improves success by 3.6--6.1 absolute points over GRPO and RLOO at matched compute, halves gradient-norm variance, and matches the best baseline using 38% fewer policy updates.
arXiv:2607.08779v1 Announce Type: new Abstract: The signed integer alphabet contains one more negative representable value than positive. Yet, by convention, the standard symmetric integer quantizer fixes its scale to be strictly positive, which assigns this extra representable value to the negative tail and can force clipping of positive outliers. In this work, we show that, at few-bit precision, such clipping is a non-trivial source of quantization error. Asymmetric quantization addresses this problem with a zero point, shifting the grid toward the observed data range; however, this flexibility is well-known to carry a runtime penalty. For example, in llama.cpp on an AMD EPYC(TM) "Turin" CPU, a 4-bit symmetric format uses up to 9% less memory with up to 2.45$\times$ higher throughput than its asymmetric counterpart. We highlight signed symmetric quantization as a third option that retains the runtime profile of symmetric quantization without the penalty of the asymmetric format: our signed absmax grid places the extra representable value on the dominant-outlier tail through a principled and lightweight sign selection rule while keeping the zero point at zero. Our theoretical analysis offers two main results. First, we establish the signed absmax grid as conditionally bound-optimal on $\ell_2$ quantization error, and show that the condition holds for 88-99% of weight groups across pre-trained large language models (LLMs) at low bit widths. Second, we show that negating the scale of a standard symmetric quantizer is analytically equivalent to a unit zero point shift on the same signed integer alphabet. We empirically validate our proposal on models from the Qwen3, Qwen3.5, and Llama3 families, and observe improvement in perplexity and downstream few-shot accuracy over the standard unsigned symmetric quantizer at no extra inference cost
An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.
300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.
GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.
Every story is sourced. Every score is computed. We show our work and link to originals.
Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.