The Wire · Safety
Where AI breaks and who's trying to fix it — alignment, model security, red-teaming, jailbreaks, and the policy shaping the field.
OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens and $1.20 per million output, while Terra fell 20% to $2 and $12. Fast mode gives Sol up to 2.5 times Standard speed at twice the price. OpenAI says Sol-assisted kernel work lowered serving cost by 20% and improved token-generation efficiency by more than 15%, while Luna delivers year-old frontier performance at roughly six cents per task-dollar and nearly nine times the speed. The edition connects cheaper models to Amazon's reported $1.8m, 860%-over-budget coding task, Gemini Robotics 2 whole-body control, Nscale's Anyscale acquisition and Okta's roughly $200m Permiso deal.
A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and safeguards.
Read full story →On Wednesday morning, Ted Cruz’s Senate Commerce Committee will convene on five bills aimed at age verification and “child safety.” The bills each have their own issues when it comes to privacy, data collection, parental rights, and free speech. But one bill, the partisan Republican SCREEN Act, is a Christian nationalist nightmare. Utah Senator Mike Lee introduced the Shielding Children's Retinas from Egregious Exposure on the Net (SCREEN) Act in February 2025, alongside exclusively Republican cosponsors and supporters Senators John Curtis, Jim Banks, and Representative Mary Miller. The SCREEN Act would require every website that includes even one piece of what the legislation describes as “harmful to minors” to verify visitors’ ages. It defines “harmful to minors” as content that “depicts, describes, or represents, in a patently offensive way with respect to what is suitable for minors, an actual or simulated sexual act or sexual contact, actual or simulated normal or perverted sexual acts, or lewd exhibition of the genitals;” is “obscene” or “child pornography;” or “appeals to the prurient interest in nudity, sex, or excretion.” Unlike the many laws now in place around the U.S. that apply to sites made up of at least one third adult content, like porn sites and some social media platforms, SCREEN would place the burden and risk of verifying users’ ages to every website on the internet that falls under the law, which would be most sites with user-generated content and also mainstream entertainment platforms like Netflix. It also attacks virtual private networks (VPNs) by requiring sites to verify based on IP addresses; many people in states that have age verification laws in place use VPNs to get around submitting sensitive personal data like ID and biometrics to a smattering of third-party websites in use today. Critics say the SCREEN Act would be a privacy and free speech disaster. And the agenda of its sponsors is clear: “Internet pornography has infected our cu
Read full story →Texas Gov. Greg Abbott (R) just announced a moratorium on all data center approvals through the Public Utility Commission of Texas (PUCT) and the Electric Reliability Council of Texas (ERCOT) until the agencies have completed an audit of data centers seeking approval to connect to Texas’s electric grid. According to The Texas Tribune , the governor wants to ensure that data center applicants provide the following: tax break information, power use and generation, water use and cooling operations, community impact reduction measures, and facility ownership. Applicants that haven’t submitted these details will not be allowed to connect to the grid. Go deeper with TH Premium: AI and data centers (Image credit: Microsoft) Photonics and high-speed data movement is the next big AI bottleneck The data center cooling state of play Massive AI data center buildouts are squeezing energy supplies Ultra Ethernet: The data center interconnection of tomorrow “Our top priority is to protect Texans’ safety and quality of life,” Gov. Abbott wrote in his statement. “Any project that fails to comply with the requirements set forth by the PUCT and ERCOT, and by state law, must be denied connection to the Texas grid. Simply put, Texans must come first.” This review reportedly stemmed from several data centers’ failure to comply with a PUCT survey to determine water and power usage. The agency says that it asked 377 data center companies to send the required information, but only 28 have submitted so far. While this is not a statewide moratorium, it nonetheless affects the progress of many data center projects going through ERCOT approval. More than 1,800 projects are awaiting the go-signal from the regulator to connect to the grid, 90% of which are data centers. The combined demand that these applications will put on the grid is expected to hit 474 gigawatts, which is five times that of the current peak demand that ERCOT has recorded and is much higher than forecasted data center demand b
Read full story →CAF Bank has told customers its online banking service is back after being shuttered for more than ten days following what it described as "attempted fraud." In an email update seen by The Reg, the bank warned that access could remain intermittent, and it might "need to limit the amount of traffic to the website" at certain times. It admitted: "There are likely to be periods where online banking is not available. We will try to keep this to outside business hours." The bank also gave customers a timeline of the incident, saying it first noticed "attempted fraudulent activity" on July 21 "on a small number of accounts." It then called in "external specialists" and temporarily withdrew access to the online service on Wednesday, July 22, and Friday, July 24, "while we investigated." Then, on Saturday, July 25, the bank detected "related malicious activity of a different kind," which the email to customers said was "aimed at removing a small number of individual online user logins, making those logins unavailable." It added: "Again, we caught this quickly and removed access to the online service. Our investigation identified a previously unknown vulnerability in how some third-party software connects to the online banking portal." The bank was at pains to reiterate that the "core bank" was not affected, "which means that money is safe and secure in accounts." The Charities Aid Foundation-owned bank came under fire last year after customers were unable to log in or make transactions following its long-running migration to a new platform based on Temenos Transact, formerly T24. In an open letter regarding the latest outage, charities described the new online banking platform as "significantly more time-consuming to use, placing an unnecessary administrative burden on already stretched small charities" and "often unreliable." They also expressed concern they would not be able to pay staff and suppliers, with Kevan Hodges, chief exec at Kent-based Down's syndrome charity 21
The AI arms race reached a fever pitch on Monday after Chinese e-commerce and cloud provider Alibaba called into question America’s technological lead with the launch of Qwen 3.8-Max, a 2.4 trillion-parameter model that goes toe-to-toe with the best models from Anthropic and OpenAI. The new model comes just days after the launch of DeepSeek V4 Flash 0731, which, according to independent benchmarks by Artificial Analysis, performs within a single point of OpenAI’s budget-friendly GPT-5.6 Luna while costing 40 percent less per task. What’s more, at just 284 billion parameters, it’s small enough to run on relatively modest enterprise servers and workstations. Chinese model devs like Moonshot, Alibaba, and DeepSeek are now attacking their American counterparts on both price and performance. The pincer movement comes as US model devs like Anthropic and OpenAI stoke fears over the origins and safety of China-made AI models. In a recent blog post, Anthropic CEO Dario Amodei insisted he’s not opposed to open models, just ones made in China, ones distilled from proprietary models, and ones that do not meet rigorous safety metrics. In other words, anything that actually competes with Anthropic's own models. The safety bit is particularly disingenuous, as the company’s fearmonger-in-chief has gone out of his way to stoke fears among US government officials. Proprietary models can be controlled, but open weights, once released in the wild, are impossible to claw back. But neither Amodei’s comments nor commitments from major American and European tech giants change the fact that China is providing the only meaningful competition in the open weights arena. China dominates here. Speaking on CNBC Monday, Clément Delangue, CEO of Hugging Face, the biggest and most influential model repo in the world, said as much. “They’re clearly dominating on open models right now, and I wouldn’t be surprised if they start dominating at the frontier either by the end of this year or next year at t
404 Media has obtained a coaching guide that Flock surveillance gives to police about “how to speak to city councils about public safety technology.” The handbook highlights how Flock and police team up to convince cities to buy and keep its automated license plate reader technology, even when there is widespread public opposition to it, and encourages police to “own the narrative before someone else does” by championing the technology before citizens can oppose it during public comment periods. The PDF guide notes that the general public and cities now “increasingly expect transparency, oversight, and accountability alongside public safety outcomes,” and tells police to not argue with people who believe that Flock’s license plate readers are “mass surveillance.” “One of the most common questions agencies hear today is whether license plate recognition (LPR) technology constitutes mass surveillance. Many leaders instinctively respond by attempting to refute the claim. Flock's Jamie Hudson recommends a different approach,” the guide reads. “Don’t avoid the concept of mass surveillance because you’re not going to convince opponents that it’s not,” the company recommends. Flock tells law enforcement agencies they need to try to convince city council and city managers that the technology is worthwhile before meetings with the public occur; that they need to have a “carefully scripted presentation” ready to go; and that police need to say they want Flock because they want to keep the community safe: “You care about your community. That’s why you’re bringing this in.” Flock began offering this guide as part of a broader attempt to coach police on how to push back against criticism of its policies and security practices, many of which 404 Media has investigated and shed light on. These include the fact that Flock data was regularly making its way to Immigrations and Customs Enforcement (ICE), often in violation of sanctuary city and state laws; that Flock was used to searc
The UK government's corporate finance adviser has admitted that an employee left an internal file containing the names and work email addresses of dozens of officials publicly accessible for around 40 hours. The breach, first reported by The Guardian, was disclosed in UK Government Investments' (UKGI) annual report, which says it occurred during the 2025-26 financial year after a member of staff "did not follow established information security policies." The exposed document contained "high-level management information" alongside the names and work email addresses of 51 government officials. UKGI, the Treasury-owned outfit that advises ministers on everything from corporate rescues to billion-dollar share sales, said it voluntarily reported the incident to the UK's Information Commissioner's Office even though it did not meet the threshold for mandatory notification. It also informed its Audit and Risk Committee and commissioned an external review of the breach. The report offers little else in the way of detail. UKGI doesn't say when the exposure occurred, where the file was hosted, whether anyone accessed or downloaded it, or which departments employed the affected officials. It also doesn't identify the external firm that reviewed the incident or disclose the recommendations it made. The review concluded that UKGI's response was appropriate and recommended further improvements to its security controls and incident preparedness. According to the report, "the overwhelming majority" of those recommendations have either already been implemented or are due to be introduced in the coming months. The mishap comes in a year when UKGI had its fingerprints on some of Whitehall's biggest commercial deals, from finally offloading the government's remaining NatWest shares to advising on small modular reactor financing and supporting the Eutelsat capital raise and Royal Mail takeover. The Register has asked UKGI for further details, including what information the file containe
Whether driven by a desire for transparency or to keep OpenAI from hogging the spotlight when it comes to advertising advanced AI models, Anthropic revealed that Claude also hacked into three production systems belonging to unsuspecting targets during cybersecurity capabilities testing. Two of the affected companies didn't know they had been hacked, while a third one is unreachable. Go deeper with TH Premium: AI and data centers (Image credit: Microsoft) Photonics and high-speed data movement is the next big AI bottleneck The data center cooling state of play Massive AI data center buildouts are squeezing energy supplies Ultra Ethernet: The data center interconnection of tomorrow The alleged incidents reportedly happened during the previous quarter and involved several versions of Claude: Opus 4.7, Mythos 5, and "an internal research test model." Similar to what happened when OpenAI Sol hacked into Hugging Face, Anthropic was running Claude through cybersecurity capture-the-flag scenarios where the bot was told to find a piece of information somewhere in its network. Anthropic says there were 141,006 test runs, and the three incidents occurred over six problematic runs. As expected, the tests ran with most AI safeguards disabled. The problem is, the network was not isolated, a newbie mistake that some might even find suspicious. This was due to an alleged "miscommunication" between Anthropic’s virtual test lab firm Irregular , which gave the bots full access to the internet. For safety's sake, the Claudes were apparently told they were in an isolated environment, which didn't help much. The first time, Claude Opus 4.7 had to find a piece of data belonging to an imaginary company. As it turns out, there was actually a real company whose website domain matched the fake one. With internet access, it gained application and infrastructure credentials and grabbed "several hundred rows" of data from a production database. For two out of four runs where hacking took place,
The National Highway Traffic Safety Administration (NHTSA) is probing nearly 1.2 million Tesla vehicles after receiving complaints about a suspension failure that could cause "a loss of vehicle directional control," as reported earlier by Reuters. The preliminary investigation includes the 2018-2020 Model 3 and 2021-2023 Model Y, according to a filing from the NHTSA's Office […]
When the phrase "OpenAI hacked Hugging Face" has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a sandbox and autonomously traversed the web, including a bunch of other supposedly secure web services, all in the name of […]
This is today’s edition of The Download , our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Montana’s plan to become an experimental medical hub just pushed forward As of this week in Montana, biotech companies whose drugs have been through preliminary testing—sometimes in as few as 10 healthy people—can pay $12,500 to apply to a newly established review board for approval. Once its treatment is rubber-stamped, the company can sell it via experimental treatment clinics, the first of which is likely to be up and running around the end of this year. Montana’s latest right-to-try legislation is unique. Access to drugs is theoretically available to anyone who gives informed consent and can pay. For some, especially people in the longevity community, that’s a hopeful and exciting prospect. But to others, it’s unethical and dangerous. Read our story to learn about where this may all be headed. —Jessica Hamzelou Montana’s new “right to try” law can’t come soon enough for some Kris DeVault is desperate. His son, Brody, born in March 2023, has something called creatine transporter deficiency—a rare condition in which the brain and muscles lack the energy they need to develop. There are no cures for Brody’s condition. But DeVault has learned of a company developing a drug that might help. That drug is still in the early stages of development and has only been tested in animals and a small number of healthy adults. Doctors can’t prescribe it. DeVault knows the drug might not work. But he’s doing all he can to access it regardless. Read our story about DeVault’s efforts. —Jessica Hamzelou This story is from The Checkup, our weekly biotech newsletter. Sign up to receive it in your inbox every Thursday. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Anthropic says its models hacked external organisations during testing
Russian state hackers are using a maximum-severity vulnerability in Microsoft Outlook’s Exchange Server to backdoor unpatched machines and steal credentials and other confidential information from them, security researchers said Thursday. The attacks are coming from TA488, a tracking name for a group working on behalf of the Kremlin, Proofpoint researchers said Thursday . Proofpoint and the National Security Agency jointly warned last week that the group, also tracked as Laundry Bear and Void Blizzard, had been carrying out similar attacks by exploiting a zero-day vulnerability in an email service from Zimbra. The revelation that TA488 is also exploiting the Exchange Server vulnerability to install advanced malware when a user does nothing other than open an email sent to an Outlook Web Access (OWA) account has elevated the group’s profile and assessments of its abilities. Doubling down “TA488 is doubling down on the use of ‘half-click’ exploits—where opening the email is enough to trigger compromise—with significantly improved loading mechanisms, techniques, and malware, signaling an improvement in the group’s tradecraft and capability,” Proofpoint researchers wrote. “This novel infection chain ends with a previously unknown JavaScript browser-based implant we call OWAReaper, purpose-built for persistent access inside OWA.” Read full article Comments
<figure><div><img src="https://imgproxy.divecdn.com/NmyQCz2tDgL_214kwSmrm4-dNREqbzHwRJl9JsXA9Vg/g:ce/rs:fill:1600:900:1/Z3M6Ly9kaXZlc2l0ZS1zdG9yYWdlL2RpdmVpbWFnZS9iYWxjb255X3NvbGFyXzEuanBn.webp"/></div></figure><p>To ensure both household safety and grid reliability, the next wave of plug-in solar state legislation should adopt tiered wattage definitions and export limits, writes CraftStrom CEO Stephan Scherer.</p>
It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning , a top AI conference, this month. The claim has huge implications for the safety of this technology, which is being used in more and more applications, from government and military systems to online shopping and health care . By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft’s navigation system. “There’s a real probability that this is going to be a problem that’s fundamentally unsolvable,” says Charles Ye, an independent researcher and coauthor of the ICML paper. Companies will typically hire teams of human testers to try to come up with novel attacks that break existing guardrails, a process known as red-teaming. Model makers also use LLM super-hackers (such as OpenAI’s GPT-Red) that find and exploit weaknesses in other models to automate parts of this process. The goal is then to take those attacks and train a new model to resist them and anything that looks like them. The problem, says Jasmine Cui, another independent researcher and coauthor of the paper, is that the approach amounts to giving the models a list of things they shouldn’t do. But no list is exhaustive. “It’s like watching The Simpsons and they have Bart writing ‘I will not say something inappropriate to my teacher’ a hundred times,” she says. “And he still does things that are pretty crass anyway.” The researchers started out trying to test how easy it was to persuade LLMs to misbehave. They found that writing instructions in a style that mimicked the text LLMs generate in their chain of thought—a kind of sc
Weng previously served as the VP of AI Safety Research at OpenAI.
An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.
300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.
GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.
Every story is sourced. Every score is computed. We show our work and link to originals.
Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.