The Wire · Google
Gemini, DeepMind, and Google's AI across cloud and devices — the releases and research that matter to builders.
OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens and $1.20 per million output, while Terra fell 20% to $2 and $12. Fast mode gives Sol up to 2.5 times Standard speed at twice the price. OpenAI says Sol-assisted kernel work lowered serving cost by 20% and improved token-generation efficiency by more than 15%, while Luna delivers year-old frontier performance at roughly six cents per task-dollar and nearly nine times the speed. The edition connects cheaper models to Amazon's reported $1.8m, 860%-over-budget coding task, Gemini Robotics 2 whole-body control, Nscale's Anyscale acquisition and Okta's roughly $200m Permiso deal.
280 points, 259 comments on HN
Read full story →377 points, 245 comments on HN
Read full story →758 points, 342 comments on HN
Read full story →Google is refocusing its CC AI agent on household coordination, letting families share emails, schedules, and tasks so the AI can manage calendars, fill out forms, make shopping lists, plan meals, and more.
if you see this, it’s beacuse you’re a real fan. AI News for 9/16/2026-9/17/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies! AI Twitter Recap Agent Runtimes, Long-Horizon Workflows, and the Rise of Coordinator UIs Claude Code Projects pushes “one conversation, many cloud threads” into product : Anthropic rolled out Projects in Claude Code , where a single conversation can spawn parallel cloud sessions, pass context between threads, and continue running after the user leaves. Follow-up posts clarify availability and that threads currently run in the cloud, with local workflows coming . Internally, Anthropic staff describe it as a higher-level coordinator abstraction with evolving long-lived memory and aggregated status updates via a single controlling Claude ( Cat Wu , MikeyK ). This is one of the clearer productizations yet of multi-session orchestration instead of just “chat + tools.” Google and others are standardizing agent infrastructure around managed harnesses, files, and secrets : Google updated Gemini managed agents with a new Antigravity-based harness plus two notably practical APIs: a Credentials API that keeps secrets out of model context via placeholders and trusted-domain egress proxying, and a Files API for artifact movement and persistent sandboxes. The same release claims up to 30% lower costs and 22% higher cache hits . Meanwhile, Perplexity’s Computer , Base44’s phone-calling Superagent , Google Labs’ family-oriented CC agent , and Meta’s desktop Muse for Mac all point in the same direction: persistent agents with scoped permissions, user-specific context, and asynchronous execution as the default UX rather than an add-on. Jev and “System One” Classification Models as a New Agent Primitive TypeSafe’s Jev dominated discussion as a fast, cheap constrained-output primitive : The cleares
Amazon Web Services does a lot of things right, but the cloud giant’s user console is arguably not one of them. Which may be why it’s just created an easier alternative for new users. In a Wednesday post, Senior Solutions Architect Micah Walter said that when AWS launched its earliest storage, compute, and queuing services, “anyone with an idea could start building.” Over the years, AWS added more services, and more options, making its console and overall UI quite complex. “That combination of global reach, breadth, and depth remains essential for those customers,” Walter wrote, “but if you are at the start of a new idea, every configuration option is effort standing in the way of shipping your dream product fast.” “We’ve heard from builders that they do not want to spend their first hours configuring an AWS environment,” he added. The cloud giant’s response is a new “getting started experience” aimed at “builders who are working at the pace of AI.” The Register understands that AWS has been working on this for almost a year, after realizing new users find its existing console intimidating. Walter said the new UI allows users to establish an AWS account with credentials from Google, GitHub, and Apple. When new users sign up, AWS will automatically create a project. “Instead of having to complete configuration tasks before you can work on your project, you start with sensible defaults and simple administration,” Walter wrote. Once logged in, users “get a prompt to paste into your coding agent that configures it to work with your new AWS environment. From there, your agent can deploy resources, run workloads, and iterate on your application following best practices for working with AWS.” The cloud colossus will apply what Walter described as “sensible defaults” to resources, and those who sign up with the tool can also apply a spending limit starting from $20 a month. AWS will “suggest a spend limit based on your usage, and you can accept that recommendation or set a c
The new institute aims to surface differing views between Google, Google DeepMind, and the broader global research community around AGI. "They will not always agree, and they will likely change their minds, as more data and information comes to light at the fast-moving frontier."
A zero-click vulnerability that allows remote code execution affects all of the major AI coding agents - Anthropic’s Claude Code, OpenAI’s Codex, Google's Gemini CLI, Microsoft’s Copilot, and Microsoft-owned GitHub Copilot - and could give attackers full access to every asset and piece of data that the agent can reach, researchers say. The exploit, dubbed “Plugin4Shell,” is a “first-of-its-kind AI supply-chain attack,” according to threat hunters at Air, a security startup focused on protecting enterprise AI agents. Instead of targeting the model or agent, Plugin4Shell attacks trusted marketplaces that host plugins for major coding agents. Such attacks could therefore reach millions of users and machines, the researchers said. Almost 90 percent of Fortune 500 companies use Copilot, according to Microsoft, which also happens to be one of the two that didn’t ship a patch for the flaw. “The fix has to ship in the agent, and updating is the only complete mitigation where one exists,” Air researchers Or Nevo, Dor Granat, and Niv Hoffman said in a Thursday report. The Air team reported the security issue to all four vendors in June, and both Anthropic and OpenAI patched it in Claude Code 2.1.179 and Codex 0.146.0, respectively. Google has deprecated the Gemini CLI, and therefore told Air it will not patch, so every install remains vulnerable. Google does, however, suggest users migrate to its newer Antigravity agentic development environment, which is protected from this attack. Microsoft didn’t fix the flaw in Copilot. However, a GitHub spokesperson told us the Plugin4Shell attacks do not affect GitHub. “To prevent abuse of SHAs, GitHub does not allow users to create branch or tag names that resemble commit SHAs,” the spokesperson said. “This mitigation ensures the reported vulnerability cannot be exploited on GitHub.” The Air researchers said that the GitHub mitigation isn’t sufficient to defeat Plugin4Shell attacks. This is “because marketplaces can also be hosted in o
Google's AI models may not be in the lead by most measures right now, but the company does have one key advantage: your data. If you're deep in the Google ecosystem, Gemini models have a lot of context on you already. Google's latest Google Labs experiment, known as CC , aims to expand that kind of customization to the whole family. CC is an evolution of something that Google announced in 2025. The original CC eventually became Gemini's Daily Brief, which churns through the data in your Google account to offer daily action items and suggestions. The new CC has a similar goal, but it's designed to be a shared resource for up to six users in a family. Google says that CC has its own Google account, allowing each family member to interact with it (or not) as they choose. For example, CC only sees emails from its connected users if they are explicitly shared. You can do that by designating certain addresses as always available to the agent—something like school scheduling emails. You can also send content to CC via email or Google Chat. The agent can even monitor a shared Google Drive folder, into which you can dump invitations, documents, and other content. Read full article Comments
The shift comes after a UNICEF test found leading AI models struggled to accurately retrieve global development statistics.
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text , an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be “cloudy,” the key might change it to “overcast.” Anyone who knows the key can determine if it was generated by the platform using it. New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow. The threat can become greater in the face of an adversarial prompt, in which an attacker attempts to cause a model to carry out a harmful action, such as revealing a password or other sensitive information. Instructions that normally wouldn’t be followed will, in some cases, be performed once the watermarking is deployed. The finding underscores the need for developers to thoroughly test how their LLMs and agents behave when watermarking is in place. Changing safety behavior “As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they’re powering an agent,” Andrea Siposova, an AI security researcher at Lasso Security, told Ars. “Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it’s going to show up somewhere.” Read full article Comments
Less than two weeks after Google released a mapping of the complete brain and central nervous system of an adult male fruit fly, we've seen enthusiasts put the structure to work everywhere from turning a fruit fly into a day trader to teaching it parallel parking . Now, one Balatro fan says they trained the structure with an algorithm to play the game, with the win rate currently sitting at a cozy 20%. The famous Fruit Fly has beaten Balatro from r/balatro The player shared a sped-up video of the model apparently playing the game. Based on the video, the player chose the lowest difficulty (White Stake) and the default Red Deck. We've already seen OpenAI's GPT-6 'Astra' model beating the game with the Black Deck on Gold Stack difficulty, which is generally considered the hardest combination in the game. ActualAerie1011, the Reddit user who shared the video, says they trained the model using a trainer algorithm they developed to discover useful Balatro seeds. Like other roguelike games, Balatro is randomized, so algorithms like this can discover seeds that are unique and can potentially lead to very high scores (including the game's scoring limit). In order to train the brain, both the brain apparatus (a connectome alongside the actual model) and the algorithm play a seed. Then, the results are compared, and the model on the brain is rewarded or punished based on its choices. Currently, the user says that the brain has a 20% success rate on a random seed, presumably at that same White Stack/Red Deck difficulty. The user says the model doesn't know anything about the seed outside of what's immediately visible on-screen, and that training is ongoing. "The fruit fly will return, strong and smarter," they wrote in a comment on their original post. It's an impressive feat, though some commenters have cast doubt on the project. The player didn't share many details about how they trained the model outside of what's above, nor any repo for the project or references to other o
Watermarks that European law requires be added to AI-generated content to establish provenance may come at a cost. According to Lasso Security, AI model watermarking changes how AI agents handle tools and safety refusals. The altered behavior isn't necessarily worse but can be, particularly under adversarial prompt injection. With the implementation of the EU AI Act, providers of AI models must mark the output of their software with machine-readable code. Google DeepMind's SynthID-Text is one method for doing so, and has been adopted by Anthropic and by OpenAI. The benefit of this sort of digital labeling is that manipulative or deceptive AI-generated content can be more easily detected, even if it does have the potential to stigmatize the usage of AI. Anthropic's explanation of how it applies watermarks to Claude output involves intervening in the prediction that results in specific words. For example, if Claude were emitting the sentence "The weather today was cold and…" then it might favor one statistically likely candidate (e.g. "overcast") over an alternative (e.g "gray"). It may be possible to detect those additions. "Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token," Lasso explained in a blog post provided to The Register. "At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection." "Watermarking uses low-stakes choices like these – which occur many times over a piece of generated text – to leave a pattern in Claude’s responses," Lasso Security added. "That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it." While a reader might not notice the word choice bias, AI agents can be subtly sensitive to vocabulary differences. Lasso found that this sort of digital content tagging can affect tool calling and refusal behavior. Watermarking, the co
<p data-block-key="p6vdq">Search giant looking at first facility in the Land of Enchantment</p>
Snap is introducing "Specs Intelligence," a new AI assistant that can connect other digital accounts to help you with things like work tasks and keeping track of travel information. It seems similar to AI assistants like Meta's Muse and Gemini's Spark, though Snap is pitching Specs Intelligence as an "anticipatory AI service" that "helps you […]
An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.
300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.
GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.
Every story is sourced. Every score is computed. We show our work and link to originals.
Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.