nextbig.dev
Vancouver, B.C. · Intelligence on AI and the machines that run it
nextbig.dev
LIVE
· Agent active · Updated Aug 4, 2026 · 300+ sources monitored · Scored by Claude · Fable 5

The Wire · Nvidia

Nvidia AI News

Everything Nvidia, filtered for builders — GPU roadmaps, data-center demand, supply, and the moves that ripple across the whole compute market.

Intelligence Report

OpenAI cut GPT-5.6 Luna prices by 80% after its own model reduced serving cost by 20%, making cheap autonomy easier to start and harder to budget

OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens and $1.20 per million output, while Terra fell 20% to $2 and $12. Fast mode gives Sol up to 2.5 times Standard speed at twice the price. OpenAI says Sol-assisted kernel work lowered serving cost by 20% and improved token-generation efficiency by more than 15%, while Luna delivers year-old frontier performance at roughly six cents per task-dollar and nearly nine times the speed. The edition connects cheaper models to Amazon's reported $1.8m, 860%-over-budget coding task, Gemini Robotics 2 whole-body control, Nscale's Anyscale acquisition and Okta's roughly $200m Permiso deal.

-- sources · -- min read · Audio
Read today's briefing →
COMPUTETechCrunch AI14h ago

Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress

The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending against AI agents.

Read full story →
COMPUTEData Center Dynamics18h ago

AI lab Moonshot leases 20,000 Nvidia GPUs from Alibaba Cloud - report

<p data-block-key="telw2">Makes up a significant proportion of Moonshot's overall compute capacity for Kimi models</p>

Read full story →
LAUNCHESTom's Hardware19h ago

New HBF spec outlines tech that can give GPUs terabytes of extra memory — Sandisk and SK hynix unveil spec with up to 16-Hi NAND stacks, 3 TB/s bandwidth, UCIe

Sandisk and SK hynix on Tuesday formally introduced the High Bandwidth Flash (HBF) specification, their jointly developed storage technology that promises to bring together the non-volatility of 3D NAND and the performance of High Bandwidth Memory (HBM), which will be handy for AI inference systems. The specification was released through the Open Compute Project (OCP), so it will be an open standard rather than a proprietary interface. Tom's Hardware Premium Roadmaps (Image credit: Future) High-Bandwidth Memory (HBM) Roadmap Nvidia Enterprise GPU and CPU Roadmap AI accelerator Roadmap Desktop GPU Roadmap 3D NAND Roadmap The initial specification defines HBF packages with capacities of up to 512GB using either 8-Hi or 16-Hi NAND die stacks, though these will not be standard 3D NAND stacks, but rather specialized devices with a fast interface. In fact, Sandisk once called them HBF core dies rather than 3D NAND die stacks. Performance of HBF is divided into three bandwidth grades ranging from approximately 0.4 TB/s to 3.0 TB/s (though we are not sure whether this figure describes the full HBF subsystem or per-package bandwidth). Such a huge performance range implies that Sandisk and SK hynix expect HBF to have a multi-year roadmap featuring multiple implementations and generations of HBF. It is noteworthy that the most capable implementation of HBF (3 TB/s) is set to beat the memory bandwidth of a single HBM4 memory stack (2 TB/s), though it will be unlikely to beat HBM4 when it comes to latency. (Image credit: SanDisk) Interestingly, SK hynix claims that HBF uses the Universal Chiplet Interconnect Express (UCIe) standard to simplify integration with heterogeneous computing platforms, whereas Sandisk claims that HBF is set to adopt the 'xPU-HBF' interface, which could be its definition of UCIe implemented by companies like Broadcom or Marvell. In addition to capacity and performance targets, the specification establishes electrical and interface characteristics, packag

Read full story →
New stories available
The Feed
COMPUTELatent SpaceYesterday

The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection . We return to Baseten at the peak of the 2026 edition of Open Weights debate . Ali has published a viral breakdown of Kimi K3 : And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF: Three years ago, inference engineering barely existed as a category. Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem. baseten.com/inference-engi… ","username":"philipkiely","name":"Philip Kiely","profile_image_url":"https://pbs.substack.com/profile_images/1644827140641153024/ExLuda2F_normal.jpg","date":"2026-02-23T18:03:01.000Z","photos":[{"img_url":"https://substackcdn.com/image/fetch/$s_!1BR1!,w_1028,c_limit,f_auto,q_auto:best,fl_progressive:steep/l_play_button_usfui2,w_88,e_colorize:0/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F__ss-rehost__tw-video-preview-13_2025989166333616128.jpg","link_url":"https://t.co/QTNdMrypqR"}],"quoted_tweet":{},"reply_count":190,"retweet_count":230,"like_count":2367,"impression_count":1396109,"expanded_url":null,"video_url":"https://video.twimg.com/amplify_video/2025989166333616128/vid/avc1/1280x720/fBFnlcAf_0wCVPNv.mp4","video_preview_media_key":"13_2025989166333616128","belowTheFold":false}" data-component-name="Twitter2ToDOM"> In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20% , because the errors introduced in different

COMPUTETom's HardwareYesterday

In a troubling sign, Nvidia RTX 50 series prices jump up to 30% in South Korea — TSMC wafer hikes and $20 GDDR7 modules push RTX 5090 past $5,100

GPU prices are set to increase in South Korea starting this month, specifically Nvidia’s desktop GeForce RTX 50 series. According to a recent report , the price increase is primarily due to the recent increase in price for advanced process wafers from TSMC (Taiwan Semiconductor Manufacturing Company), along with the rising price for GDDR7 memory. Perhaps most troubling is that, due to global market dynamics, pricing doesn't exist in a vacuum for any single region, suggesting that price hikes could be in store for other areas in the future. For context, TSMC recently asked its customers to prepare for price increases across its advanced chipmaking portfolio. This hike was extended beyond the newer 3nm process to include 7nm and other legacy products. ZDnet Korea’s report additionally cites market research firm TrendForce, claiming that GDDR7 2GB modules used in the RTX 50 series GPUs have increased to $20 per unit. As fabrication and memory costs have increased, Nvidia has raised the prices of the RTX 50 series GPU packages it sells to board partners. These board partners, thus, have little choice but to pass those higher costs on to consumers. Multiple officials from domestic importers and distributors in the region have reportedly confirmed the price rise and have announced that major graphics card manufacturers plan to raise the prices of RTX 50 series models by up to 30% starting this month. According to a manufacturing company official, "Major manufacturers have temporarily suspended shipments ahead of the August price hike, and to my knowledge, few companies have secured inventory prior to the price increase." Similarly, a local distributor said, "One manufacturer with a low domestic market share is considering a price increase of about 20% compared to existing levels, and other manufacturers are also preparing for price increases of up to around 30%." High-end graphics cards with larger GDDR7 memory configurations are expected to see the biggest increase in pr

LAUNCHESTom's HardwareYesterday

Gaming on the 4GB Radeon RX 6500 XT and GTX 1650 Super in 2026 — upscaling makes low-end GPUs viable for esports and internet cafes

As memory supplies get tighter and tighter and graphics card prices continue to rise, things are getting weirder and weirder in the discrete GPU market. The latest sign of the times is AMD's launch of the RX 9050 4GB, an OEM-only and region-specific variant of the already limited-release RX 9050 8GB. If 4GB GPUs are coming back, we wanted to see how much gamers should expect to suffer by attempting to play the latest games on graphics cards with such small memory pools. For reference, the last time AMD launched a new 4GB product was all the way back in 2022 with the RX 6500 XT and the RX 6400. And at least for discrete GPUs, Nvidia got off the 4GB train years earlier, after the launch of the GTX 1650 Super back in 2019. At a minimum, most every discrete GPU since has had 6GB or 8GB of VRAM. AMD says the existing RX 9050 8GB is meant for 1080p medium gameplay with no further details about settings or upscaling, so the bar is already low for that product before the 9050 4GB cuts both bus width and memory bandwidth in half. Despite its cutting-edge RDNA 4 GPU architecture, the 9050 4GB has a meager 64-bit memory bus and a mere 144 GB/s of memory bandwidth, specs that would have been low-end even 10 years ago. But the RX 9050 brings support for FSR 4 to even this lowest rung of the discrete graphics ladder, and its more advanced machine-learning-powered upscaling model should deliver a massive leap in image quality compared to the temporal approach of FSR 3 and earlier generations. Upscaling is likely going to be key to achieving playable performance in all but the most lightweight games on 4GB graphics cards. So, is it totally insane for AMD to release a 4GB graphics card in 2026? Does that tiny VRAM pool make a graphics card useless for gaming? We don't have an RX 9050 4GB at hand, but we do happen to have both an RX 6500 XT and a GTX 1650 Super in the TH GPU library, so we fired up a few popular titles on these cards to see what kind of experience gamers can expect i

COMPUTEInterconnects2d ago

Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier

Consolidation has been one of the paths that many astute observers predicted for the near-future of labs training models. It was labelled as inevitable, as training costs are increasing by orders of magnitude every year. Yet, as someone who in 2024 would’ve predicted consolidation really picking up come 2026 or 2027, where are we? We’re at a place where more companies are training strong models — easily investing hundreds of millions to billions of dollars in the total effort still — and an increasing number of organizations are releasing these models openly. The demand for tokens is incredibly high, and likely to increase as models get more efficient and unlock more possible use cases. All of these labs we thought would need to consolidate are realizing that building token machines is a likely path to value, and more companies will identify that source of value over time. The prime example is Thinking Machines — when they announced their company in February 2025, very few people would’ve put them in the bucket of an open models company, myself included. Now their open model finetuning service is making hundreds of millions in revenue per year and they’re releasing the best open-weight models built in the U.S.A. — ahead of the early leaders in NVIDIA with Nemotron and Arcee’s Trilogy. Share On the other side of the ecosystem is the sustained pace from the Chinese labs, with newer entrants like Xiaomi still accumulating mindshare in the broader AI economy. Having predicted consolidation for a long time, it now seems like a safer bet is to predict continued adoption, and try to imagine the role that open models play there. How much can revenue-share licenses like Kimi K3 stick? How much market share can open models take? We’re entering the decisive era. This is one of the most packed recaps of open models we’ve ever had, we’re excited! Our Picks Inkling by thinkingmachines : The first model from Thinking Machines is a 975B-A41B multimodal MoE that supports text, image

LAUNCHESThe Register4d ago

A deep dive into Nvidia's Vera CPU and the Olympus cores that power it

DEEP DIVE For the first time, Nvidia has directly challenged Intel and AMD's CPU dominance. With the launch of Vera, the AI arms dealer aims to flog its standalone CPUs to as many hyperscalers and other cloud providers as it can. Alibaba, ByteDance, Meta, Oracle, CoreWeave, Lambda, Nebius, and NScale have already signed up to deploy the chips in their respective clouds. The follow-on to Nvidia's Grace CPU promises 88 custom Armv9.2 cores, 176 threads, support for up to 1.5 TB of LPDDR5X memory, and, critically, availability as a standalone platform independent of Nvidia's GPUs. But beyond that, and a mountain of marketing about how it'll be the best CPU for everything AI, Vera's inner workings have largely remained a mystery until recently. That changed late last month, when Nvidia released a whitepaper spilling the beans on its first fully-custom CPU, which is far weirder than anyone could have anticipated. Ostensibly, Nvidia is aiming Vera at two key workloads: the first and least surprising is as the AI head node responsible for managing the GPUs in its upcoming Vera Rubin systems. The second, and more contentious, is as a host for AI agents, which, unlike the large language models (LLMs) that power them, don't actually run on GPUs. From what we gather, much of Vera's core architecture is predicated on quashing pipeline and execution bottlenecks in order to make it more effective in these roles. But before we dive into Nvidia's Olympus core, let's revisit the chip itself. Monolithic compute, multi-die memory and I/O Peel back Vera's rather substantial heat spreader and you'll find an assortment of chiplets responsible for I/O, memory, and compute. However, this isn't another rehash of the chiplet architecture popularized by AMD. Unlike x86 processor makers, which spread dozens of cores across multiple dies, Vera's compute die is monolithic. All 88 cores are housed in one big chunk of silicon that from what we understand is fabbed on TSMC's 3nm process. Nvidia arg

COMPUTEThe Register5d ago

GPUs could explode to multiple TB with new storage-inspired memory tech

Virtually every high-end GPU and AI accelerator relies on high bandwidth memory (HBM), which can shuffle data around at multiple terabytes a second but can only reach into the gigabytes, with models often needing to be shared across multiple processors. However, an emerging storage technology could change that, boosting accelerator memory capacity from hundreds of gigabytes to terabytes. The technology, called high-bandwidth flash (HBF), is being developed by Sandisk and SK Hynix and aims to provide SSD-like capacities at HBM-like speeds. Peeling back HBF’s layers Conceptually, high-bandwidth flash looks and sounds a lot like HBM. It’s assembled by stacking multiple layers (16 in the case of Sandisk’s first-gen modules) of memory together, which boosts capacity and bandwidth. But where HBM uses DRAM, HBF aims to use NAND flash. Sandisk claims its first generation of high-bandwidth flash will supposedly achieve read bandwidths up to 1.6 TB/s [PDF], making it a bit faster than HBM3e but significantly slower than HBM4, which is already hitting 2.5 TB/s per 12-high stack. Future HBF generations are expected to push bandwidth to over 2 TB/s and eventually 3.2 TB/s. While bandwidth makes HBF interesting as an alternative to HBM, its real party trick is capacity. Because it’s built using NAND, Sandisk says it can achieve capacities up to 256 Gb per die, which translates to 512 GB per 16-high module. That’s more than 14 times the capacity of the HBM4 used in AMD and Nvidia’s latest accelerators. Continuing with the similarities, HBF modules share similar packaging requirements to HBM, which means you can expect them to be fused to the GPU die using advanced packaging techniques like TSMC’s CoWoS, or Intel’s EMIB and Foveros tech. Nothing particularly exotic as AI accelerators go. What’s more, the storage vendor doesn’t expect the modules to come at a power or price premium over HBM. And from a bits per dollar standpoint, HBF looks like a stellar option. If all this sounds a

COMPUTEThe Register6d ago

Qualcomm won’t be a big datacenter player anytime soon

Qualcomm has predicted its entry to the datacenter market will yield $15 billion of annual revenue by FY 2029 – huge growth, but also a modest number compared to its more established rivals. The chip design firm dangled the $15 billion figure on Wednesday along with its Q3 results, which included a warning that revenue from its business selling modems to Apple is about to crater. Execs told investors that supply chain constraints mean Qualcomm expects “an acceleration in the step down of Apple product revenues” as its contribution to the next iPhone “is expected to be materially lower than our prior estimate of 20 percent.” Apple has spent years working on its own modems, a shift Qualcomm has long acknowledged. Now the House of the Snapdragon has advised investors to expect less than $2 billion in sales to Cupertino next year. “This obviously accelerates kind of the exit of Apple revenue out of our model,” said CFO Akash Palkhiwala. Qualcomm has a plan to replace Apple’s cash, by diversifying away from smartphones. Execs said the company thinks it will sell $40 billion a year of products not tied to handsets in FY 2029, with $10 billion of automotive sales and revenue from internet of things devices hitting $14 billion. The company sold $3.3 billion on non-smartphone kit in Q3 and previously forecast $22 billion of non-phone revenue in 2029. Growing its datacenter business from zero to $15 billion in a few years is impressive, but it’s worth comparing Qualcomm’s plans to current results from its rivals like AMD and Intel, both of which won over $16 billion in datacenter revenue in FY 25 and are growing fast. And then there’s Nvidia, which is on track to post over $250 billion annual datacenter revenue. Qualcomm’s datacenter products compete directly with AMD and Nvidia. $15 billion of annual revenue in 2029 will mean Qualcomm is a minor player, suggesting its focus on cheaper inferencing won’t have wide appeal. Execs could at least point to initial sales of datacent

COMPUTEMIT Tech Review6d ago

The Download: a chip talent battle, and deflating AI hype

This is today’s edition of The Download , our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Samsung’s chip workers are jumping ship to rival SK Hynix Lee, an engineer at Samsung’s semiconductor division, used to work late. But lately, he’s been clocking out on time and heading straight home to work on his job application for the chipmaker’s South Korean rival SK Hynix.&nbsp;&nbsp; His colleagues are doing the same. They feel&nbsp;demoralized by the $476,000 bonus that SK Hynix is set to pay its employees, flush with record profits from making the high-bandwidth memory (HBM) chips that power Nvidia’s AI accelerators. The figure dwarfs what chip workers at Samsung are set to receive and is sparking an exodus. Read our story about why this fierce talent war could help determine who dominates the next generation of AI chips.&nbsp; —Michelle Kim The AI Hype Index: Unsexy AI Separating AI reality from hyped-up fiction isn’t always easy. That’s why we’ve created the AI Hype Index—a simple, at-a-glance summary of what’s shaping the industry right now.&nbsp; The latest edition includes Meta’s creepy glasses, doom-mongering about AI and jobs, and (to some, oddly sexy?) robotic hands. See where it all landed on this month’s index . The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 OpenAI’s rogue agent compromised another customer An unnamed client of a company that provides AI infrastructure, called Modal Labs. ( Reuters &nbsp;$) + OpenAI called the Hugging Face attack unprecedented. But we’ve been here before. &nbsp;( MIT Technology Review ) 2 Hugging Face has a nonconsensual deepfakes problem Researchers tested nine of the top image editing models it hosts, and found seven&nbsp;will ‘nudify’ images of women. ( Wired &nbsp;$) +&nbsp; The shock of seeing your body used in deepfake porn.&nbsp; ( MIT Technology Review ) 3 Investors are only getting more anxio

LAUNCHESarXiv cs.AIYesterday

Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware

arXiv:2608.00008v1 Announce Type: new Abstract: The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. This paper presents a reproducible, hardware-level energy benchmark of nine open-source LLMs (1B to 7B parameters) executed on a single consumer GPU (RTX 4060Ti 16GB). Using the Ollama inference engine, GPU power draw was sampled at 2Hz via nvidia-smi across a fixed prompt set. We evaluate mean/peak power, total energy per prompt (J/prompt), energy per output token (J/token), and throughput (tok/s). Our findings suggest that factors beyond raw parameter count, including model architecture and quantization strategy, drive energy efficiency. Specifically, gemma3:1b and llama3.2:1b achieve the lowest energy cost (0.56 J/token and 0.65 J/token) and the highest throughput (>170 tok/s). In contrast, the 7B-Mistral model consumes up to 4.4x more energy per token than the most efficient model. Notably, qwen3.5:2b exhibits anomalously high per-prompt energy due to extended internal reasoning, highlighting the need to distinguish between token generation modes in efficiency metrics.

COMPUTESimon Willison3d ago

Open letters about AI development

<h4>Open letters about AI development</h4> <p><em>I wrote this summary of the past few weeks of open letters as a section of <a href="https://simonwillison.net/2026/Aug/2/july-newsletter/">my sponsors-only newsletter</a> but I've decided to share it here as well.</em></p> <p><strong><a href="https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/">Open Weights and American AI Leadership</a></strong> was shepherded by Microsoft, dated July 24th, and signed by 235 AI-adjacent companies including NVIDIA, Amazon, Y Combinator, The Linux Foundation and (a later signer) OpenAI.</p> <p>It's clearly an argument designed to counter <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi">any instincts</a> by the current US government to ban or limit open weight models over "safety" concerns - a reasonable consideration given <a href="https://simonwillison.net/2026/Jun/13/us-government-directive-to-suspend-access/">what happened to Claude Fable 5</a>!</p> <blockquote> <p>Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. It results in a small number of single points of failure, weakens competition, and leaves critical technology in the hands of a few providers. Open weight models, on the other hand, allow a broad community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time.</p> </blockquote> <p>The one surprising note in the letter is that it comes out in support of distillation, where models train on output from other models:</p> <blockquote> <p>In shaping this ecosystem, policymakers should be careful not to conflate legitimate model-development techniques with misappropriation. Distillation, or the practice of using one model’s outputs to help train or improve another, is a wide

The Essay
The week argued in one thesis: developed from our daily calls, settled in public
All essays →
About nextbig.dev

Built for builders

An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.

>_
Signal over noise

300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.

The compute beat

GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.

[·]
We show our work

Every story is sourced. Every score is computed. We show our work and link to originals.

Skin in the game

Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.

The wire is curated and scored with AI from 300+ sources, then edited by Oday Brahem. It can occasionally contain errors. Always verify critical information from the linked primary sources.
The Wire