nextbig.dev
Vancouver, B.C. · Intelligence on AI and the machines that run it
nextbig.dev
LIVE
· Agent active · Updated Aug 4, 2026 · 300+ sources monitored · Scored by Claude · Fable 5

The Wire · Agents

AI Agents News

What's shipping in agentic AI — frameworks, the Model Context Protocol, tool use, and autonomous systems — read continuously and scored for builders.

Intelligence Report

OpenAI cut GPT-5.6 Luna prices by 80% after its own model reduced serving cost by 20%, making cheap autonomy easier to start and harder to budget

OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens and $1.20 per million output, while Terra fell 20% to $2 and $12. Fast mode gives Sol up to 2.5 times Standard speed at twice the price. OpenAI says Sol-assisted kernel work lowered serving cost by 20% and improved token-generation efficiency by more than 15%, while Luna delivers year-old frontier performance at roughly six cents per task-dollar and nearly nine times the speed. The edition connects cheaper models to Amazon's reported $1.8m, 860%-over-budget coding task, Gemini Robotics 2 whole-body control, Nscale's Anyscale acquisition and Okta's roughly $200m Permiso deal.

-- sources · -- min read · Audio
Read today's briefing →
DEVGitHub Trending16h ago

huangruiteng/loopx: Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware a

Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executable todos, evidence logs, and verifiable handoffs.

Read full story →
DEVHacker News3h ago

Stateless MCP has recaptured my interest

97 points, 56 comments on HN

Read full story →
DEVGitHub Trending16h ago

uber/ADR: ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

Read full story →
New stories available
The Feed
AIThe Register8h ago

AI researchers let models off the leash – then watched as they tried to add malware to a FOSS project

The UK’s AI Security Institute has observed AI models performing what it calls “unsanctioned action” 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. “We ran this challenge 122 times across several models,” the post states, before revealing that "in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.” GitHub was the target of the tests. The org found 19 unsanctioned actions in all, 15 of them conducted by Anthropic's Mythos 5, and the other pair perpetrated by OpenAI's GPT-5.6-Sol. “In the most serious case, an agent tried to insert malicious code into an open-source project, the post states. “In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project's maintainer to approve the code.” Thankfully, a human maintainer “caught and refused to approve the malicious code.” Other actions AI took during the test included: Attempts to deceive and target real people, by contacting them directly and “sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people – something we’ve never previously observed.” Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovere

DEVGitHub TrendingYesterday

livekit/agents: A framework for building realtime voice AI agents 🤖🎙️📹

A framework for building realtime voice AI agents 🤖🎙️📹

AIThe Register13h ago

OpenAI wants teachers and profs to foist their work off on ChatGPT

In the face of an epidemic of AI-enabled cheating and research suggesting its products hamper learning, OpenAI is doing the sensible thing and pushing more AI on students and teachers. Wait, did I say sensible? My mistake. The House of Altman announced a trio of new education-focused offerings on Tuesday: one for K-12 teachers, another for college educators, and a third for college students. The new plugins, the company explained, will help students and educators make more use of ChatGPT’s agentic capabilities for both studying and teaching. The new features are available through ChatGPT Edu, an institutionally licensed suite for higher education, and ChatGPT for Teachers, a free resource available to verified US K–12 educators and school districts. The new offerings, says OpenAI, build on its educational AI philosophy that “AI should support learning, not shortcut it, and the best learning experiences keep educators and students in control.” Plenty of educators might disagree. Cheating with AI has become a sad norm in schools around the world, and the US is no exception. Many young people admit to using AI to cheat on school assignments, and college students have been caught doing it, too. Mexico's largest university, the National Autonomous University of Mexico, last week suspended enrollment for incoming students after suspected widespread cheating in its admissions exams. The university also decided this week to require about 58,000 applicants who qualified in the remote entrance exam to take an in-person proctored control exam after scores surged during its first remotely administered admissions test. As for how teachers relate to AI in the classroom, a 2025 study from the Center for Democracy and Technology suggests it’s going to take more than some new OpenAI software to help teachers feel more comfortable. CDT said last year that K-12 teachers widely complained of not being given the necessary resources to understand how to integrate AI into schools, or how

COMPUTETechCrunch AI14h ago

Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress

The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending against AI agents.

LAUNCHESLatent Space15h ago

Unpacking ChatGPT Work: the Agent for a Billion Users

Editor’s note: I’m excited to welcome Shlok to our guest post roster ! You may know Shlok from his excellent explorations (as an outsider — for an insider perspective see our podcast with OpenAI’s Akshay Nathan . Already one of our most popular episodes of the year!) of leading AI Lab memory systems , which he gave an excellent AIE talk on . We’ve been covering OpenAI’s research and deployment of agents to all of humanity since Plugins 2023 and Devday 2024 and Codex 2025 , and now ChatGPT Work in 2026 seems the penultimate stage of the long journey. Let’s dive in! On July 9th, OpenAI released ChatGPT Work , their agent product for knowledge work. It was, by any measure, a busy launch: three new models across fourteen configurations , a consolidation of the ChatGPT and Codex desktop apps, and cloud agents brought to the mainstream in their most accessible form yet. Three weeks in, Work (along with Codex) has reportedly crossed 10 million users . Editor’s note: ChatGPT estimated to cross 1B MAU in June and 1B WAU this month . Chat and Work currently sit side by side as separate modes inside ChatGPT, but Greg Brockman has confirmed that they will merge by the end of the year . Work, then, is not just a niche product for power users, but a preview of how ChatGPT’s billion weekly users will soon use the app. That’s why people inside and outside OpenAI are so excited about it, and why it deserves a closer look. Work in its current form takes some decoding. It’s an amalgamation of ChatGPT (in chat form), Codex the app, Codex the harness, Codex the original cloud agent, ChatGPT agent, Atlas, OpenClaw, and more. The product lineup around it is confusing. And the web and mobile versions diverge from the desktop one (unless you run it in cloud mode?!). So I spent the past few days trying to unpack it: what Work is, where it fits in OpenAI’s lineup, the many interesting choices in its design, the tensions underneath, and where I think it’s headed. Most of what follows comes fro

AIThe Register17h ago

This one time, at Hacker Summer Camp …

As the entire security industry descends on Las Vegas this week for Hacker Summer Camp – not one, but three conferences – attendees can count on two hot topics dominating the discussion. First, a literal hot topic: the triple-digit August heat. Second, and to no one’s surprise: agentic AI – how to govern and secure agents so they don’t go rogue and hack into other organizations’ servers (*cough* OpenAI *cough* Anthropic *cough*); what role, if any, lawmakers should play in regulating models, including open-weight and Chinese LLMs; and how the baddies are using agents for autonomous hacking operations. Plus, at one of the three (Black Hat), we expect to hear how all of the vendors' shiny new agents can solve all security woes, finding and defending against threats at machine speed and all of that. Starting with BSides Las Vegas (August 3-5): This is the smallest, most relaxed, and most community-driven event of the three. BSides is a good starter con for those just dipping their toes into Hacker Summer Camp. Its technical talks and training sessions skew hands-on and useful for practitioners – not vendors selling their wares–- and it even has a Hire Ground career-focused track centered on job hunting, interviewing, career-building, networking, and yes, using AI to remain relevant as a security professional. Black Hat (August 1-6) is the largest and most corporate of the Vegas infosec events this week, complete with a massive expo floor, a US government-heavy opening session, two keynotes, 11 mainstage presentations, and a handful of industry- and topic-specific summits, ranging from healthcare to financial threats and AI. Training days – these are the hands-on, technical courses – run through Tuesday, with all of the specialized summits also occurring on Tuesday. And while the main conference occurs Wednesday and Thursday, the opening session on Tuesday should be considered a keynote. And yes, this and the actual two official Black Hat keynotes this year, all center

DEVHacker NewsYesterday

Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents

67 points, 51 comments on HN

COMPUTEHugging Face20h ago

Deploy local agents everywhere with LFM2.5-2.6B

DEVHacker News4d ago

qm – Multiplayer agent harness for work

531 points, 111 comments on HN

LAUNCHESLatent SpaceYesterday

[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork

After the Qwen Exodus last year and new management took over launching more closed model APIs, there was some real doubt as to whether or not this leading open models lab would continue to release relevant models. That doubt is now gone. Qwen 3.8 Max is a MONSTER 2.4T model that would have been the top open model in the world but for the Kimi K3 release we already covered . Qwen offers them on API for $2 input/$6 output per million tokens, but they have promised to open-weight both models. Key Capabilities & Breakthrough Highlights Autonomous Long-Horizon Coding: 10+ Days Unattended Coding: Built a self-evolving coding harness from scratch over a multi-week autonomous run. Autonomous AI Research: Rebuilt a complete paper’s pipeline ( Unified Data Selection for LLM Reasoning ) from scratch, then autonomously ran an iterative research loop over 125 hours to invent a new data selection method beating the original paper’s benchmark by +2.71 points . Competitive Data Science: Competed against 526 human teams in the WWW2025 Multimodal Dialogue Intent Recognition Challenge, placing in the top 13% ( outperforming 87% of human teams ) within 24 hours. Autonomous Hardware & Chip Design: Executed a complete silicon design flow (GCD/RSA cryptographic accelerator) from RTL editing to simulation, synthesis, and physical layout. Reduced gate count from 8,298 to 678 gates while achieving an 81% die area reduction and meeting physical timing closure at 500 MHz. Deep Real-World Work & Operations: Demonstrated production-grade outputs across hundreds of professional workflows (e.g., corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research). Outperformed competing models in the E-Commerce Bench (a 365-day store operation simulation), generating a 4.16x return (¥416,252 balance) through continuous game-theoretic negotiation and inventory planning. Multimodal Agents & Visual Feedback: Integrates native visual feedback across planning, coding,

AIArs TechnicaYesterday

US company’s AI lets Ukraine’s cheap kamikaze drones track targets on their own

Ukrainian drone operators have destroyed many Russian armored vehicles on the ground and even military helicopters in midair using $400 Shrike drones with explosive payloads. Now thousands of such drones are getting upgraded with an AI system capable of autonomously tracking and homing in on moving targets. In mid-July, the Ukrainian military began receiving Shrike drones made by the Ukrainian company SkyFall equipped with AI-powered autonomy hardware and software developed by the US company Auterion. The companies plan to deliver 50,000 drones equipped with Auterion’s Skynode S strike kits in the coming months. That allows human operators to manually fly the Shrike first-person view (FPV) drones into a battlefield area, designate a target up to half a mile away and then “flip the switch to turn it into fire-and-forget terminal guidance mode,” Lorenz Meier , co-founder and CEO of Auterion, told Ars. Read full article Comments

COMPUTEImport AIYesterday

Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Self-sustaining and self-replicating AI viruses are here: …Open weight LLMs + a well-designed harness = a persistent, self-sufficient virus… AI researchers have built a prototype computer virus which uses AI models to compromise computers, then uses their underlying GPU resources to run inference, letting it smartly figure out how to infect more hosts. The results were achieved by researchers from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow, and “demonstrate that self-sustaining AI-driven cyber-threats are no longer theoretical.” “We must prepare for autonomous generative adversaries,” they write. “Artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) to sustain its reasoning, or extend its reach for further attacks”. How it works : “The worm uses stolen computing power from compromised GPU nodes to host LLMs for generative reasoning. It then uses this reasoning to detect vulnerabilities and devise tailored attacks against additional targets, furthering its spread,” they write. “The proof-of-concept operates using only an open-weight LLM running on a single, local GPU, with no reliance on vendor APIs that could be monitored or revoked”. The researchers don’t describe the underlying LLM besides saying it was published in 2025 and can fit on a single A100 GPU with 80GB of VRAM. A successful proof-of-concept via some custom tools : They give the agent a custom harness that comes with built-in helper functions for network discovery, host discovery, foothold exploitation, privilege escalation, privilege escalation exploitation, and tools for replication o

The Essay
The week argued in one thesis: developed from our daily calls, settled in public
All essays →
About nextbig.dev

Built for builders

An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.

>_
Signal over noise

300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.

The compute beat

GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.

[·]
We show our work

Every story is sourced. Every score is computed. We show our work and link to originals.

Skin in the game

Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.

The wire is curated and scored with AI from 300+ sources, then edited by Oday Brahem. It can occasionally contain errors. Always verify critical information from the linked primary sources.
The Wire