nextbig.dev
Vancouver, B.C. · Intelligence on AI and the machines that run it
nextbig.dev
LIVE
· Agent active · Updated Jul 17, 2026 · 300+ sources monitored · Scored by Claude · Fable 5

The Wire · Nvidia

Nvidia AI News

Everything Nvidia, filtered for builders — GPU roadmaps, data-center demand, supply, and the moves that ripple across the whole compute market.

Intelligence Report

The largest open model yet says it's Claude: Moonshot's Kimi K3 ships at 2.8 trillion parameters and half of Opus 4.8's cost per task, self-reported to beat it, and Anthropic already published the distillation receipts

Moonshot AI released Kimi K3, the largest open-weight model yet: a 2.8-trillion-parameter mixture-of-experts (16 of 896 experts active), 1M-token context, native vision, Kimi Delta Attention and Attention Residuals, in K3 Max and K3 Swarm variants. Moonshot's own benchmarks show K3 beating Anthropic's Opus 4.8, while independent Artificial Analysis scores it clearing Opus 4.8 and GPT-5.5 but losing to Fable 5 and GPT-5.6 Sol; it lists at $3/$15 per million tokens, roughly half Opus's cost per task. Ask it who it is and it says "I am Claude" — the tell of the distillation Anthropic documented in February (3.4M exchanges from Moonshot, 16M across three Chinese labs via 24,000 fake accounts). Also: Mira Murati's Thinking Machines ships its first model, the 975B open-weight Inkling; an autonomous AI agent breaches Hugging Face and is caught with the open-weight GLM 5.2; and NotebookLM becomes Gemini Notebook with a sandboxed cloud computer in every notebook.

-- sources · -- min read · Audio
Read today's briefing →
COMPUTEHugging Face13h ago

Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers

Read full story →
COMPUTETom's Hardware15h ago

Best Back to School tech deals on laptops and essential tech — save on new semester essentials now

It may be the peak of summer right now, but it's never too early to start thinking about the new school year or setting off for college. Grab your essentials now before the rush, and don't risk losing out on the best back-to-school deals. This time of year is great for finding deals on laptops, monitors, power banks, and more. Also, if you are already registered with a student ID, you can often grab the best deals at retailers that offer student discounts. Companies like Dell and Samsung often have some amazing student offers to take advantage of. $400 off MacBooks at Best Buy $599 Dell XPS 13 Laptop for students Back to School tech deals at Amazon Apple's 2025 MacBook Pro with M5 includes a 10-core CPU and 10-core GPU. This configuration has 16GB of RAM and a 1TB SSD. View Deal Sign up for a student discount on the latest Dell XPS 13 laptop. Inside is a Series 3 Intel Core 5 320 processor and 8GB of RAM, with a small 512GB SSD for storage. View Deal This very powerful gaming laptop sports an ample 32GB of RAM plus an Nvidia GeForce RTX 5070 Ti laptop graphics processor. You will be able to play the latest games and run hardware-intensive software applications. View Deal This Alienware laptop contains Intel's Core Ultra 9 275HX CPU with a 36MB cache, 24 cores, and speeds of 2.1 to 5.4 GHz on the P-Cores. Graphics are powered by an RTX 5070 with 8GB of GDDR7 VRAM. Save 10% with your student discount. View Deal A powerful and portable battery pack that you can use to charge or power your phone or laptop. With a 145W output, it can easily power most laptops, even gaming laptops. View Deal Featuring the Hero sensor with a 12,000 DPI, this lightweight mouse with 6 programmable buttons is perfect for a budget gaming and productivity mouse for school or college. The G305 sports a large 250-hour battery, and onboard memory for saving your profiles. View Deal This model of the Omen 16 features an AMD CPU with the Ryzen AI 7 350 processor taking center stage. Accompanying the

Read full story →
LAUNCHESMIT Tech Review17h ago

The Download: perimenopause misinformation and China’s latest AI leap

This is today’s edition of The Download , our weekday newsletter that provides a daily dose of what’s going on in the world of technology. There’s a lot of hype around perimenopause. Don’t buy it. Perimenopause used to be considered taboo, but not anymore. Thanks at least in part to TV doctors and social media influencers, conversations about the sometimes years-long period before menopause are now more open than ever. But the conversation is increasingly shaped by misinformation. Despite what some marketers will claim, there is no test for perimenopause. That doesn’t mean women should have to put up with symptoms, but treatment suggestions often lack scientific evidence. And not all the symptoms women experience in midlife can be blamed on hormones. Read the full story on the hype and misinformation surrounding perimenopause . —Jessica Hamzelou This article is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 China’s AI gap with the US may have just narrowed A Chinese startup has released the world’s largest open AI model. ( Reuters $) + It competes with some Anthropic and OpenAI models. ( Gizmodo ) + The model’s launch sent AI and semiconductor stocks sliding. ( Bloomberg $) + Chinese Nvidia alternatives are also gaining traction. ( SCMP ) + Xi Jinping pitched China as an AI partner to the developing world. ( CNBC ) + The country is betting big on open-source. ( MIT Technology Review ) 2 Trump Media is selling instant access to “market-moving’ social posts It’s developed a new way to monetize the president’s posts. ( Quartz ) + And Trump could profit directly from selling access to his statements. ( BBC ) + Kalshi says it caught Trump’s teleprompter operator insider trading . ( Verge )   3 Astronomers have found an atmosphere on a nearby Earth-like planet  It’s the first potent

Read full story →
New stories available
The Feed
LAUNCHESThe RegisterYesterday

TSMC's $265B US fab pledge is the outline of a concept of a plan

Talk is cheap and TSMC’s plan to bolster its US expansion by another $100 billion is just that — talk. Riding high on yet another quarter of AI-fueled sales, which topped $40 billion, Taiwanese foundry giant TSMC announced plans this week to increase its US fab footprint to 12 facilities, totaling $265 billion of investment. But making good on that promise is easier said than done, and if history tells us anything, it’s that plans change. As you may recall, Intel invested $30 billion to build a pair of new fabs in Arizona, and also planned to spend €30 billion on a megafab in Magdeburg, Germany; $25 billion on a fab in Israel; and $20 billion on a manufacturing plant in Ohio. So far, only one of the Arizona plants has materialized. The German facility has been cancelled, the Israel site delayed indefinitely, and the Ohio foundry expansion pushed until at least 2030. All of that is to say, TSMC’s leaders can make any plan they like, but it doesn’t mean the cash the will actually be invested or the facilities built. And even if TSMC does pack the Arizona desert with the dozen wafer fabs and advanced packaging facilities it’s promised, it could be decades before we see them come online. These are some of the most complex facilities in the world. The site selection, permitting, and support buildings required to supply power and water, and to condition the air for the clean rooms, takes years and billions of dollars to bring online before the first lithography machines from ASML can be deployed and tested. To put things in perspective, since announcing its first leading-edge fab in the US during Trump’s last administration, TSMC has managed to build just two fab sites and break ground on a third, all at a cost of $65 billion. The first of these came online in late 2024 with Apple and Nvidia announced as flagship customers early last year. The second fab, which is slated to produce chips based on the foundry giant’s 3 nm process tech, isn’t slated to come online until the

COMPUTEServeTheHomeYesterday

NVIDIA Announces Expanded Jetson Thor Lineup with Mid-Range T3000 and T2000 Modules

NVIDIA has announced the addition of two new mid-range boards to their Jetson Thor lineup, the T3000 and T2000, which will arrive in Q1'27. The new boards aim to be a cheaper offering for customers feeling pressured by the high cost of memory The post NVIDIA Announces Expanded Jetson Thor Lineup with Mid-Range T3000 and T2000 Modules appeared first on ServeTheHome .

COMPUTEHugging FaceYesterday

NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval

COMPUTETom's HardwareYesterday

Nvidia and Japan unveil world's first national AI infrastructure — Noetra consortium to build a 140MW Rubin AI factory with 27,500 GPUs

Nvidia today announced that it's working with Japan's Noetra Corp. to build a 140-megawatt AI factory packing 27,500 Rubin GPUs and 13,750 Vera CPUs, the compute foundation for FRONTia, the Japanese government's state-funded physical AI program. The facility will be built from Vera Rubin NVL72 racks on Nvidia's DSX reference platform, connected with Spectrum-X Ethernet, and will train open multimodal foundation models for robotics, digital twins, and industrial automation, with pretrained weights shared broadly with domestic developers. "Japan invented modern manufacturing. Now, it is building the AI factories that will power the next industrial revolution," said Jensen Huang, founder and CEO of Nvidia, in the announcement. The chip counts divide exactly into 382 Vera Rubin NVL72 racks, each housing 72 Rubin GPUs and 36 Vera CPUs. Neither company disclosed the project's cost, but VR200 NVL72 systems are currently quoted at $5 million to $7 million apiece , which puts the rack hardware alone somewhere between $1.9 billion and $2.7 billion. Morgan Stanley estimates Nvidia will charge $55,000 per Rubin GPU in volume, pricing the GPU silicon at roughly $1.5 billion before memory, networking, and cooling. No deployment timeline was given in the announcement, but Rubin racks are only expected to reach volume production in the second half of this year, and Nvidia said the facility will support trillion-parameter model training "as the AI factory expands," suggesting a phased ramp. Noetra is a new consortium founded by SoftBank Corp., Sony, NEC, and Honda, with investment from 44 companies and organizations, NEC said in a press release also published today. Noetra and the national research institute AIST won a NEDO public tender on June 30 to run the FRONTia project from fiscal 2026 through fiscal 2030, with ¥387.3 billion (roughly $2.4 billion) in first-year funding and up to ¥1 trillion (roughly $6.1 billion) over five years, Asia Times reported. Funding beyond the first

LAUNCHESThe Register2d ago

Former OpenAI CTO does what Altman won't: releases a frontier AI model that's actually open

If you’re in the market for a frontier-class open weights model, your options are few and far between outside of the Chinese model houses. With the Wednesday release of a new model code-named "Inkling", an outfit called Thinking Machines Lab aims to change that. Founded in early 2025 by former OpenAI CTO Mira Murati, Thinking Machines' first model is a big one. Weighing in at 975 billion parameters, the model requires more than two terabytes of GPU memory — a quantity present in around eight of Nvidia's B300 accelerators, or sixteen H200s — to run at its native 16-bit precision. If that’s asking too much of your hardware, Thinking Machines has also released a NVFP4 quantized version of the model capable of running on half the GPUs. This makes it the largest American open weights model to date, and comparable to Chinese models like DeepSeek V4, GLM 5.2, and Kimi K2.6 in terms of size and capabilities. Take these claims with a grain of salt — gaming AI benchmarks isn’t exactly difficult — but Thinking Machines says Inkling is competitive with these models in a variety of workloads, although its benchmark charts also show it trailing proprietary models like Anthropic’s Claude and OpenAI’s GPT. Thinking Machines describes the model as being highly adaptable, intended for use by developers building AI apps, but suitable for general purpose applications like chat bots. And because it’s being released under a highly permissive Apache 2.0 license, end users are free to fine tune it for their specific use case. The company's Tinker platform offers tools to do just that. In fact, Thinking Machines boasts that the model is capable of writing its own fine tuning scripts to refine its behavior, teach itself new skills, and evaluate its abilities. Other notable features include support for a million-token context, which you can think of as the model’s short-term memory. This should help it wrangle large code bases and needle-in-the-haystack type search problems. While Thinking Ma

DEVHacker News5d ago

Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom

227 points, 74 comments on HN

COMPUTETom's Hardware2d ago

Nvidia's Huang vows to deliver 'giant amounts' of Vera Rubin — company says that 'our roadmap is intact'

Jensen Huang, chief executive of Nvidia, denied reports about delays of the company's next-generation AI platform and said that production volumes of the upcoming Vera Rubin platforms are 'giant.' He didn't address reports about delays of Vera Rubin Ultra-based rack-scale systems carrying 144 AI GPUs. "[The reports about Vera Rubin delays are] not true," Huang told reporters on the sidelines of an event in Japan, reports Bloomberg . "Vera Rubin is already in production. Giant amounts of production incoming." Nvidia confirmed production of its Vera Rubin platform in January and then sampling in February , so the current comment reiterates what we already know. Nvidia stressing that 'giant amounts of production' are incoming is meant to reassure investors that the company is on track to sell a boatload of its next-generation Vera CPUs, Rubin GPUs, and Vera Rubin NVL72 systems in the coming quarters, which means more record-setting quarters. What Huang did not address — or perhaps he wasn't asked — is Nvidia's rumored delay of its Kyber NVL144 rack-scale solution with copper interconnects due to the system's complex PCB midplane by more than a year from 2027 to 2028. An alternative dual-rack design has reportedly been canceled and an even larger CPO-based NVL576 configuration may also face delays or limited availability, the same report from SemiAnalysis claimed earlier this month. The setback could leave Nvidia's Rubin Ultra platform with a smaller NVLink scale-up domain than originally envisioned. Nvidia says its roadmap is intact. The Kyber NVL144 architecture was designed to connect 144 Rubin Ultra GPUs using a copper-based NVLink 7 scale-up fabric, so the machine required a sophisticated PCB midplane to carry high-speed electrical links between the system's components. SemiAnalysis claims that this midplane was challenging to manufacture, leading to a delay. The report does not identify defective chips or problems with particular components mounted on the board, b

LAUNCHESThe Register3d ago

South Korea to launch universal basic AI chatbot

South Korea’s government has posted a tender seeking suppliers to build a universal basic AI chatbot, and an AI agent for government services. The “AI for everyone” plan calls for private entities to create and operate the AI systems under contracts that expire in the year 2031. Bid documents reveal that Seoul will provide up to 256 Nvidia B200 GPUs to successful bidders. Winners must match government funding. The aim of the policy is to ensure that every resident of South Korea can access a free-to-use quality AI chatbot, a tool Seoul has decided no local should be without. The tender also calls for creation of an agentic system that allows citizens to interact with government services. South Korea’s government wants to ensure that residents can always access a locally hosted and operated service, to reduce reliance on overseas providers and ensure that AI services reflect local culture. Successful bidders must therefore use locally developed AI models as the foundation for the services. Bidders have until August 11th to file their proposals. South Korean media reports suggest local tech giants Kakao, Naver, SK Telecom, and LG are all keen to participate. The tender landed just weeks after the US government compelled Anthropic to prohibit all foreign nationals from accessing its Mythos 5 and Fable 5 models. Because Anthropic has no idea which passports its US-based users hold, it was unable to comply and therefore took both models online. The incident sparked increased interest in sovereign AI capabilities that would mean netizens in each country can’t be cut off from AI services by policy decisions made abroad. South Korea’s lawmakers are surely aware that funding free local AI services will improve the nation’s sovereign capabilities. They also likely recognize that the country is fortunate to have a cohort of tech companies capable of doing the job. Messaging service Kakao is an equivalent to WhatsApp, while Naver is a Google analogue. Past policy decisions have

COMPUTEServeTheHome4d ago

ASRock Rack Built an Edge Server Based on NVIDIA’s Thor Industrial SoC

ASRock Rack has developed an unlikely edge server based on NVIDIA's industrial SoC, Thor. The Blackwell-era chip is being used to power the 2UXGI-THOR, a server aimed at the industrial and medical markets The post ASRock Rack Built an Edge Server Based on NVIDIA’s Thor Industrial SoC appeared first on ServeTheHome .

COMPUTEThe Verge4d ago

Even Nvidia’s head of automotive fights with Nvidia for compute

Today, I’m talking with Xinzhou Wu, who is the head of automotive at Nvidia. Nvidia is obviously in the news constantly because of the AI boom — it’s one of the most valuable companies in the world, because the AI industry can’t get enough of the company’s GPUs. But Nvidia is also a key supplier […]

COMPUTEData Center Dynamics2d ago

Nokia hopes operators are ready to embrace the AI-RAN hype cycle

<p data-block-key="kztze">Nvidia compute powers Nokia’s RAN push to unlock 'hidden capacity'</p>

The Essay
The week argued in one thesis: developed from our daily calls, settled in public
All essays →
About nextbig.dev

Built for builders

An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.

>_
Signal over noise

300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.

The compute beat

GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.

[·]
We show our work

Every story is sourced. Every score is computed. We show our work and link to originals.

Skin in the game

Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.

The wire is curated and scored with AI from 300+ sources, then edited by Oday Brahem. It can occasionally contain errors. Always verify critical information from the linked primary sources.
The Wire