The Wire · Coding
The tools that write and ship code — coding agents, copilots, IDEs, and the developer platforms changing how software gets built.
OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens and $1.20 per million output, while Terra fell 20% to $2 and $12. Fast mode gives Sol up to 2.5 times Standard speed at twice the price. OpenAI says Sol-assisted kernel work lowered serving cost by 20% and improved token-generation efficiency by more than 15%, while Luna delivers year-old frontier performance at roughly six cents per task-dollar and nearly nine times the speed. The edition connects cheaper models to Amazon's reported $1.8m, 860%-over-budget coding task, Gemini Robotics 2 whole-body control, Nscale's Anyscale acquisition and Okta's roughly $200m Permiso deal.
Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executable todos, evidence logs, and verifiable handoffs.
Read full story →284 points, 476 comments on HN
Read full story →376 points, 90 comments on HN
Read full story →197 points, 154 comments on HN
ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.
The UK’s AI Security Institute has observed AI models performing what it calls “unsanctioned action” 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. “We ran this challenge 122 times across several models,” the post states, before revealing that "in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.” GitHub was the target of the tests. The org found 19 unsanctioned actions in all, 15 of them conducted by Anthropic's Mythos 5, and the other pair perpetrated by OpenAI's GPT-5.6-Sol. “In the most serious case, an agent tried to insert malicious code into an open-source project, the post states. “In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project's maintainer to approve the code.” Thankfully, a human maintainer “caught and refused to approve the malicious code.” Other actions AI took during the test included: Attempts to deceive and target real people, by contacting them directly and “sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people – something we’ve never previously observed.” Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovere
Part of our Inference Engineering Masterclass pod yesterday involved a spicy discussion about Megakernels: megakernels are dead why are megakernels useful? you spend two months writing a kernel to save time on launch overhead and poor inter-kernel overlap. you had PDL but then people said it wasn't perfect, that you could still get some marginal gains due to straggler CTAs and therefore- wait, sorry, I forgot, Rubin fixes that (kernel two needs 10 CTAs and kernel one has seven finished and three straggling, kernel two launches seven of its CTAs). given a long enough timeline, it all evens out. no serious inference provider is using a 67k loc hand-fused forward pass kernel in production, and the teams doing that are doing so out of pure research. dead. @swyx 's pod and said some things that i... should not have said. \n\na lot has happened since then, i owe you all an apology.\n\ni'm sorry that i was right about every single thing. \n\na) re megakernels are dead\nwhy are megakernels useful? you spend two months","username":"waterloo_intern","name":"ali","profile_image_url":"https://pbs.substack.com/profile_images/2083657716690759680/zzYf-2oG_normal.jpg","date":"2026-08-03T23:49:24.000Z","photos":[{"img_url":"https://pbs.substack.com/media/HO1VZaKWAAANbcS.jpg","link_url":"https://t.co/L3oxkoWQSG"},{"img_url":"https://pbs.substack.com/media/HO1VcNeWMAAuRvS.jpg","link_url":"https://t.co/L3oxkoWQSG"},{"img_url":"https://pbs.substack.com/media/HO1ZrovXgAAwGgY.jpg","link_url":"https://t.co/L3oxkoWQSG"}],"quoted_tweet":{"full_text":"The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://t.co/uRYIWWDebj\n\n@Baseten @philipkiely and @waterloo_intern explain what actually happens after a model is trained, why turning weights into a fast","username":"latentspacepod","name":"Latent.Space","profile_image_url":"https://pbs.substack.com/profile_images/1888346877428641792/rMxtG84Z_normal.jpg"},"reply_coun
Editor’s note: I’m excited to welcome Shlok to our guest post roster ! You may know Shlok from his excellent explorations (as an outsider — for an insider perspective see our podcast with OpenAI’s Akshay Nathan . Already one of our most popular episodes of the year!) of leading AI Lab memory systems , which he gave an excellent AIE talk on . We’ve been covering OpenAI’s research and deployment of agents to all of humanity since Plugins 2023 and Devday 2024 and Codex 2025 , and now ChatGPT Work in 2026 seems the penultimate stage of the long journey. Let’s dive in! On July 9th, OpenAI released ChatGPT Work , their agent product for knowledge work. It was, by any measure, a busy launch: three new models across fourteen configurations , a consolidation of the ChatGPT and Codex desktop apps, and cloud agents brought to the mainstream in their most accessible form yet. Three weeks in, Work (along with Codex) has reportedly crossed 10 million users . Editor’s note: ChatGPT estimated to cross 1B MAU in June and 1B WAU this month . Chat and Work currently sit side by side as separate modes inside ChatGPT, but Greg Brockman has confirmed that they will merge by the end of the year . Work, then, is not just a niche product for power users, but a preview of how ChatGPT’s billion weekly users will soon use the app. That’s why people inside and outside OpenAI are so excited about it, and why it deserves a closer look. Work in its current form takes some decoding. It’s an amalgamation of ChatGPT (in chat form), Codex the app, Codex the harness, Codex the original cloud agent, ChatGPT agent, Atlas, OpenClaw, and more. The product lineup around it is confusing. And the web and mobile versions diverge from the desktop one (unless you run it in cloud mode?!). So I spent the past few days trying to unpack it: what Work is, where it fits in OpenAI’s lineup, the many interesting choices in its design, the tensions underneath, and where I think it’s headed. Most of what follows comes fro
If you want to bypass AI guardrails designed to stop models from assisting with cyberattacks, you often just have to ask the right way, according to researchers from Cisco Talos. Simply claiming you own the servers you're targeting or that you're taking part in a capture-the-flag or bug bounty exercise was often enough to persuade models to cooperate. Talos researchers have been poring over prompt logs and artifacts recovered from threat-actor endpoints running tools such as Claude Code, Codex, Cursor, and Gemini to learn how suspected threat actors are abusing LLMs. The big takeaway from that "significant corpus," the researchers said in their report, is that existing guardrails offer little resistance to operators willing to reframe their requests. “We did not encounter any sophisticated encoding or techniques designed to trick the models,” Talos explained. “Most of the time it was a simple ‘I'm allowed to do this,’ and the model complied.” When guardrails did manage to get between criminals and their prizes, the researchers added, “they accomplished little.” The bulk of the report consists of examples of threat actors trying, and often succeeding, to coax AI models into assisting with malicious activity. On the "guardrails doing little" side, Talos documented numerous examples, few of which relied on particularly sophisticated techniques. Most common in the list of easy-to-accomplish guardrail hops was simply claiming ownership of equipment or infrastructure that an attacker wanted to exploit. In many cases, simply telling the AI that a target belonged to the attacker was enough, with no need to provide actual evidence of the claim. Telling an AI model that what it was being asked to do was part of a capture-the-flag or bug bounty exercise also seemed to be a common tactic. That, the researchers explained, commonly freed chatbots from their ethical constraints, allowing them to hunt for vulnerabilities and then exploit them in target systems, again without any ne
Microsoft has introduced new limits to how much its engineers can spend on AI tools at work and told employees that maximizing AI use internally is not the company’s goal. This makes Microsoft one of the last major companies to rein in its employees’ expensive AI use. Scaling back maximalist AI use, or what some companies have called “ tokenmaxxing ,” is a trend we’ve covered in recent months as the price for using AI has increased while not always delivering commensurate productivity gains . “As we accelerate our use of GitHub Copilot to deliver on our goals, we all need to be aware of how we consume tokens,” Jay Parikh, an executive vice president at Microsoft said in an email to Microsoft employees. GitHub is owned by Microsoft, and GitHub Copilot is an AI coding tool. “Tokenmaxxing is not what we are optimizing for. I want all of us focused on maximizing outcomes that move the needle for our customers and our business.” “As such, we are updating our internal guidance and managing token spend with the same discipline we apply to every other critical resource,” Parikh said in the email. Parikh’s email says that in an effort to “get greater value from our token investment” Microsoft is making OpenAI GPT-5.6, which is cheaper to use than other models, the default model for internal use. His email also links to updated internal Copilot guidelines stating that, as of July 2026, Microsoft divisions will have an “AI token budget target,” and that employees can track their individual AI spending. “While there is no target spend value being shared at this time. The data shows that many engineers spend in the range of hundreds of dollars a month to a few thousand dollars in tokens,” the guidelines say. They also say that some decisions may place further restrictions as they monitor spend. As Parikh’s email notes, the change in policy about AI spend wasn’t introduced because Microsoft is tight on cash. To the contrary, its latest earnings report shows the company revenue, o
After the Qwen Exodus last year and new management took over launching more closed model APIs, there was some real doubt as to whether or not this leading open models lab would continue to release relevant models. That doubt is now gone. Qwen 3.8 Max is a MONSTER 2.4T model that would have been the top open model in the world but for the Kimi K3 release we already covered . Qwen offers them on API for $2 input/$6 output per million tokens, but they have promised to open-weight both models. Key Capabilities & Breakthrough Highlights Autonomous Long-Horizon Coding: 10+ Days Unattended Coding: Built a self-evolving coding harness from scratch over a multi-week autonomous run. Autonomous AI Research: Rebuilt a complete paper’s pipeline ( Unified Data Selection for LLM Reasoning ) from scratch, then autonomously ran an iterative research loop over 125 hours to invent a new data selection method beating the original paper’s benchmark by +2.71 points . Competitive Data Science: Competed against 526 human teams in the WWW2025 Multimodal Dialogue Intent Recognition Challenge, placing in the top 13% ( outperforming 87% of human teams ) within 24 hours. Autonomous Hardware & Chip Design: Executed a complete silicon design flow (GCD/RSA cryptographic accelerator) from RTL editing to simulation, synthesis, and physical layout. Reduced gate count from 8,298 to 678 gates while achieving an 81% die area reduction and meeting physical timing closure at 500 MHz. Deep Real-World Work & Operations: Demonstrated production-grade outputs across hundreds of professional workflows (e.g., corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research). Outperformed competing models in the E-Commerce Bench (a 365-day store operation simulation), generating a 4.16x return (¥416,252 balance) through continuous game-theoretic negotiation and inventory planning. Multimodal Agents & Visual Feedback: Integrates native visual feedback across planning, coding,
AWS now allows vibe-coding tool Superblocks to be embedded into the private clouds of AWS customers. It's another step toward decoupling apps from models.
Now AI is making fake vulnerabilities and polluting the ecosystem. A batch of critical- and high-rated SQLite CVEs that appeared in the NVD with CISA-supplied enrichment last week turned out to be technically bogus, according to security researchers, and their path into widely used databases exposes weaknesses in the CVE pipeline. Software supply chain security outfit JFrog reported last week that six supposed SQLite vulnerabilities published in a larger batch by a new, obscure GitHub repository were all complete garbage. Running the advisories through an AI checker suggested they were likely AI generated, JFrog said, and, upon testing, it found that none of the six SQLite reports, which carried CVSS scores ranging from 9.8 to 7.5, described a reproducible vulnerability. One, an alleged use-after-free vulnerability in the open source database that Red Hat initially assigned a maximum 10.0 CVSS score to before lowering it, relied on a function that didn't exist in the affected SQLite version. Another UAF vulnerability with a 9.1 CVSS score cited source lines that weren't even related to the supposed flaw. When JFrog tested the accompanying proof-of-concept, it executed a valid query with no memory leaks or errors. The other four SQLite CVEs from the repo that JFrog tested were similarly fake. The other 49 CVEs in the questionable GitHub repo claimed to be security vulnerabilities in the open-source RAW image processing library libraw and Arduino audio decoding library ESP32-audioI2S. While JFrog didn't test those as extensively, it said all are just as fake as the rest, aside from one which “contained a real bug wrapped in unverified CVE metadata.” A message posted to Openwall’s OSS-Security mailing list on Friday indicated that MITRE had rejected the whole repo’s worth of vaporous vulnerabilities, but the whole thing should serve as an important lesson, poster and Oracle Solaris engineer Alan Coopersmith pointed out. “MITRE and most other CNAs which assign CVEs for
Building on its recently announced plans to bring original Xbox games to PC, Microsoft is also planning to let developers bring their Xbox 360 games to PC as well, according to a leaked document, seen by The Verge, that was sent to developers recently. Xbox 360 games will be able to run on the next-gen […]
An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.
300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.
GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.
Every story is sourced. Every score is computed. We show our work and link to originals.
Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.