The Wire · Nvidia
Everything Nvidia, filtered for builders — GPU roadmaps, data-center demand, supply, and the moves that ripple across the whole compute market.
OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens and $1.20 per million output, while Terra fell 20% to $2 and $12. Fast mode gives Sol up to 2.5 times Standard speed at twice the price. OpenAI says Sol-assisted kernel work lowered serving cost by 20% and improved token-generation efficiency by more than 15%, while Luna delivers year-old frontier performance at roughly six cents per task-dollar and nearly nine times the speed. The edition connects cheaper models to Amazon's reported $1.8m, 860%-over-budget coding task, Gemini Robotics 2 whole-body control, Nscale's Anyscale acquisition and Okta's roughly $200m Permiso deal.
565 points, 205 comments on HN
Read full story →Following the launch of AMD's EPYC 'Venice' CPUs in July , AMD extended the performance claims for its upcoming generation of server chips on Friday. The high-level claim hasn't changed. AMD still says a 96-core, high-frequency Venice chip is around 20% faster than Nvidia's 88-core Vera in SPEC CPU 2026's Integer Rate test. However, the company went into far greater detail about the benchmarks in a new white paper . Tom's Hardware Premium Roadmaps (Image credit: Future) Leading-edge foundry roadmaps Nvidia Enterprise GPU and CPU roadmap AMD's Enterprise GPU and CPU roadmap Intel's roadmaps examined — 14A, Nova Lake, Diamond Rapids & AI accelerator push Co-Packaged Optics (CPO) foundry roadmaps There are several configuration differences depending on the benchmark throughout AMD's white paper, and although we'll call out those differences here to the best of our ability, we don't have all of the details. For the Vera comparison, in particular, AMD is mixing data from different sources, and in some cases, using different major releases of the GNU Compiler Collection (GCC). That can have a substantial impact on performance, so keep your salt shaker handy. (Image credit: AMD) First up are results in SPEC CPU 2026 with the intrate test, looking at total throughput. These are older numbers, gathered in July with GCC 15.2. The intrate test runs multiple copies of an application on the same CPU, and the SOP is to run one copy per thread. Presumably, that's what AMD did here, but the white paper doesn't clarify, even in the footnotes. The 256-core 9996 is 2.37x faster than the Intel Xeon 6980P and 2.24x faster than Vera according to the slide. The white paper clarifies the mystery 9006 CPU is the 256-core flagship. Perhaps most impressive is AMD's gen-on-gen comparison. According to these results, the 9996 is around 78% faster than last-gen's 192-core EPYC 9965. Although the high-level results bring in data from Intel and AWS, much of the white paper focused squarely on the
Read full story →A week after an Anthropic researcher’s doomsday warning rattled the AI world, the company’s CEO Dario Amodei has outlined his plan to “pace the frontier” of AI development. The proposal leans on independent safety evaluators and coordination between AI labs in democratic countries, and it’s already picked up some industry support, along with some pointed pushback from Nvidia’s Jensen Huang. Watch […]
Read full story →A week after an Anthropic researcher’s doomsday warning rattled the AI world, the company’s CEO Dario Amodei has outlined his plan to “pace the frontier” of AI development. The proposal leans on independent safety evaluators and coordination between AI labs in democratic countries, and it’s already picked up some industry support, along with some pointed pushback from Nvidia’s Jensen Huang. On […]
Nvidia's Nader Khalil and Sydney Sykes discuss one of the decisions shaping next-gen startups on the Builders Stage at TechCrunch Disrupt 2026.
A new project on GitHub, simply titled " dlss-nr-on-intel ", purports to provide exactly that: a port of NVIDIA's DLSS 5 Neural Rendering to Intel's Xe architecture. Specifically, the author (who goes by "Uzbekunknown") focused on porting the technology to the Intel Arc 140V graphics in his Lunar Lake system, and they seem to have succeeded, at least insofar as he's getting outputs that look reasonably like those of DLSS 5 on other hardware . AI is at the center of this project, beyond the DLSS 5 neural rendering technique itself. Uzbekunknown credits Anthropic's Claude as well as OpenAI's GPT-6 Astra with the code and says that they "supplied the machine, the binary, and the direction, and made the decisions", while the AI agents did everything else. Amusingly, they note that "the wrong turns are in the notes, too, deliberately," including a hallucinated driver bug that does not exist and shaped three phases of development. The end result, rather than being a wrapper around the DLSS 5 DLL as many other hacks have been , fully reimplements the 71-block U-Net that DLSS 5 uses and then runs it on the Intel Xe XMX units through a Vulkan extension called VK_KHR_cooperative_matrix. It's entirely run in FP16 with FP32 accumulate, because Xe2 doesn't support FP8. You can run the model on anything presenting its output through Vulkan, and the user presents proof-of-concept results from three fighting games: Dead or Alive 5 Last Round , Tekken 7 , and Mortal Kombat 1 . While DLSS 5 adds detail to the character, it also changes her look considerably, clashing with the visual style of the game. (Image credit: Uzbekunknown/GitHub ) It's not fast. Running the ten-year-old Tekken 7 in 640x360 resolution (1/9 of FHD) should be a trivial task for the potent Intel Arc 140V graphics, yet it apparently struggles at around 10.5 FPS with this model loaded. Note (as the author does) that the performance of DLSS 5 depends almost entirely on the game's output resolution, so running in hila
2019’s Control followed Jesse Faden into the ever-shifting, paranormally corrupted brutalist innards of the Oldest House, the headquarters of the Federal Bureau of Control, where she became the new FBC director, fought the invading forces of the Hiss, and sought the truth about the fate of her kidnapped brother Dylan. Control Resonant marks the next chapter in the siblings’ story, as Dylan reawakens to discover that Jesse has gone missing and that the Hiss threat has escaped the Oldest House and corrupted Manhattan. Using his own powers and guided by the mysterious Board, Dylan sets out to find Jesse, combat the Hiss incursion, and uncover the mysteries of a new paranormal entity at work in the twisted Manhattan cityscape. We’ve had access to Control Resonant for the past few days, and we’ve been exploring its performance and image quality across a range of hardware and settings. Control was one of the first games to show off the capabilities of GeForce RTX 20-series graphics cards and their ray-tracing capabilities, and it was also one of the first titles to incorporate DLSS upscaling. It’s only fitting, then, that Resonant is a technical showcase of its own. It features path-traced lighting effects bolstered by Nvidia’s RTX Mega Geometry tech, as well as support for the full suite of DLSS 4.5 features: Super Resolution (aka upscaling), Ray Reconstruction, and Multi Frame Generation. With that extensive spread of cutting-edge rendering tech at its disposal, I expected Resonant to look incredible, as both Control and Alan Wake II did before it. (Image credit: Remedy Entertainment/Future) Even without ray tracing or path tracing and DLSS Ray Reconstruction, Control Resonant already looks good, as we’ve come to expect from Remedy games. If performance is a priority, you certainly won’t make the game bad by leaving RT off. Control Resonant RT off versus RT Ultra Remedy Entertainment/Future Remedy Entertainment/Future But ray tracing or path tracing is certainly worth e
Earlier this year, Nvidia CEO Jensen Huang suggested that Marvell’s optics tech would make it the next trillion-dollar company. On Thursday, Marvell took a big step towards securing that status by announcing an expanded partnership with GlobalFoundries, an American wafer fab well known for its work in silicon photonics manufacturing. The tie-up will see GloFo and Marvell work to expand silicon germanium (SiGe) wafer production at the fab’s Burlington, Vermont wafer plant. SiGe is commonly employed in the production of optical transceivers, which convert electrical signals to optical ones and back again, as well as in laser modules used in many high-end switches from Nvidia and others. According to GlobalFoundries, the additional capacity will support the production of near- and co-packaged optics (CPO/NPO) as well as “next-gen” pluggable optics. Nvidia already employs co-packaged optics in some of its Spectrum Ethernet and Quantum InfiniBand switches to cut power consumption and improve reliability. Pundits expect NPO to be used in large multirack systems such as Huawei’s new Ascend 960DT-based SuperPods. Nvidia’s future multi-rack systems are widely expected to use NPO as well. The fab says its SiGe tech has already been validated up to 200 Gbps per lane, which is required for 1.6 Tbps transceivers. Faster lane speeds could open the door to 3.2 Tbps transceivers, which are also on its roadmap. Demand for optics technologies has surged thanks in no small part to everyone’s favorite subject: AI. Speaking during Computex this spring, Marvell CEO Matt Murphy explained why demand for optics demand is set to explode. Up until now, optics have primarily been used to connect racks to resources across the datacenter. Anything within a few meters was well within reach of cheaper and less power hungry copper interconnects. But as lane speeds climbed from 25 Gbps to 50, 100, and then 200 Gbps, copper interconnects began running into practical reach limits. 400 Gbps per lane is
At its annual Connect conference, Huawei unveiled a new generation of AI accelerators that promise performance far beyond anything Nvidia can currently sell in the Middle Kingdom. This could be a boon for Chinese developers used to purchasing lesser weapons from everyone's favorite AI arms dealer. If that weren’t enough, the chip, codenamed the Ascend 960DT — the DT here apparently stands for decode/training — is slated to arrive a full three quarters ahead of schedule, launching in the first quarter of 2027. With up to 288 GB of what we assume is Huawei’s custom HiZQ memory tech, a homegrown alternative to the high-bandwidth memory used by the rest of the world, and up to four petaFLOPS of FP4 performance (half that at FP8), the accelerator offers twice the performance and memory capacity of the company’s 950-series parts, launched earlier this year. Compared to American GPUs, the 960DT offers similar memory and bandwidth to Nvidia’s B300 family of chips launched last year, but only about half the FP8 and a third the FP4 compute. While a big step up for the Chinese chip designer, it still has a long way to go to catch up with Nvidia’s Rubin and AMD's recently launched MI455X which promise substantially higher compute and bandwidth. Rubin boasts between 35 and 50 petaFLOPS of FP4 performance, 288 GB of HBM4 memory, and 22 TB/s of bandwidth, putting it in an entirely different league. Having said that, neither of those chips is available for sale in the Middle Kingdom, so it's not like Chinese model devs have better options. The best GPU Nvidia can sell in China offers near-identical dense floating-point performance, twice the memory, double the memory bandwidth, and support for much larger scale-up domains. Scaling up The performance gap may not be nearly as damning as it sounds, either. AI models aren’t trained, and for the most part aren’t run on a single GPU or NPU anymore. The more important factor is often how efficiently the platform scales. Nvidia's and AMD's
Nvidia’s ability to sell GPUs is ultimately limited by how much juice the power grid can provide. With ever-growing depreciation cycles, it’ll be years before datacenters decommission their aging Hopper or Blackwell systems. More GPUs mean pulling more power from the grid. Nvidia can’t exactly force grid operators to add capacity any faster, but it can make it easier for its customers to build smarter and more efficient bit barns. “At the datacenter scale and at the AI-factory scale, we're literally trying to think about how can we eke out every bit of efficiency to drive more performance per gigawatt,” Dion Harris, senior director of Nvidia HPC and AI Hyperscale Infrastructure Solutions, told El Reg in a recent interview. At the AI Infra Summit this week, we got our first look at the systems Nvidia has been building to maximize the amount of power available for compute while minimizing the impact of datacenters on the local grid. Datacenters, as a general rule, rarely operate anywhere close to the peak capacity. A 100 megawatt datacenter might use at most 80 percent for critical compute loads. The actual ratios vary from bit-barn to bit-barn, but this provides a buffer for hardware inefficiency, conversion losses, and other spikes in demand. The downside, of course, is that this leaves 20 megawatts or so of untapped capacity that the datacenter can’t use and the utility can’t reclaim. Every kilowatt of stranded power is a GPU that Nvidia could have sold, so Nvidia’s DSX platform aims to address both problems. Minimizing overheads The first of these, which Nvidia calls DSX MaxLPS, is an evolution of an old idea. If the compute and all the physical infrastructure — power cabinets, batteries, coolant distribution units (CDUs), and chillers — can talk to one another, operators can achieve significant power savings. For example, if the air handlers had a way of knowing how much power a rack was pulling, they could ramp up and down based on demand rather than running max
According to The Information, Apple is planning to get back into the server game and might just pair up with Nvidia to make it happen. Apple retired its Xserve line in 2011 and has largely left enterprise machines to other manufacturers since. But the growing demand for compute power as the AI industry continues to […]
Since about 2020, AI has largely focused on training bigger and better models. Large language models (LLMs) ballooned from millions of parameters to trillions. This proved effective: The largest version of OpenAI’s GPT-3, released in 2020, correctly answered just 43.9 percent of questions on a popular knowledge-and-reasoning benchmark. Just four years later, GPT-4o reached a score of 88.7 percent on the same exam, effectively matching those of human experts. Advanced AI labs are still training ever larger models, but that training has somewhat receded to the background of the AI conversation. In 2026, inference—the use of trained models to produce code, write essays, or make images of ourselves as elves—has come to the forefront. “It’s like training is yesterday’s news,” says Matt Kimball , principal data-center analyst at Moor Insights & Strategy. “All that any chief information officer wants to talk about is inference.” Nvidia CEO Jensen Huang, speaking at the company’s GTC 2026 conference, touted this change as the “ inflection point of inference .” Part of what’s caused the shift is very simple: LLMs are becoming useful, so people are using them. On top of that, many models on the market today are reasoning models. In response to a user’s query, they run inference not just once but multiple times, reprompting themselves in a process called chain of thought . Reasoning models generate longer outputs, and models with high reasoning effort can produce up to 20 times as much text as those with low or no effort. Adding even more to the world’s inference workload, the rise of agentic AI has resulted in inference running not just as a real-time response to a user’s query but also around the clock, working autonomously toward a user-defined goal. Amazon’s Trainium chip was originally designed for AI training. However, Amazon Web Services chose to break up AI inference into two parts, with Trainium running the more computationally complex portion and Cerebras’s wafer-sca
Nvidia CEO Jensen Huang took a call from President Trump on Monday while onstage at the All-In Podcast's All-In Summit. It's not the first time Huang has taken a call from the president during work, but this time he put Trump on speakerphone before a big crowd. During the call, the president launched into his […]
An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.
300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.
GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.
Every story is sourced. Every score is computed. We show our work and link to originals.
Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.