The Wire · Microsoft
Copilot, Azure AI, and Microsoft's AI across its stack — what's shipping and what it means for builders.
OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens and $1.20 per million output, while Terra fell 20% to $2 and $12. Fast mode gives Sol up to 2.5 times Standard speed at twice the price. OpenAI says Sol-assisted kernel work lowered serving cost by 20% and improved token-generation efficiency by more than 15%, while Luna delivers year-old frontier performance at roughly six cents per task-dollar and nearly nine times the speed. The edition connects cheaper models to Amazon's reported $1.8m, 860%-over-budget coding task, Gemini Robotics 2 whole-body control, Nscale's Anyscale acquisition and Okta's roughly $200m Permiso deal.
AMD has posted strong second quarter results and forecast even better future financials once its Helios rack systems and Instinct MI400-series GPUs reach buyers. “In data center AI, the growing number and scale of Helios and MI450-series deployments position the [datacenter] business for significant growth in the second half of the year, with growth accelerating in 2027,” CEO Lisa Su told investors on Tuesday during the chip design company's Q2 earnings call. “We now expect data center segment revenue to more than double year over year in 2027,” she added. Yet, despite reporting Q2 profits surging 163 percent year-over-year on revenues of $11.5 billion, and several multi-gigawatts worth of Helios commitments from the likes of OpenAI, Anthropic, and Meta in the bag, Wall Street isn’t buying it. The company's share plunged 10.5 percent after its results announcement, before settling 8.7 percent below opening price at the time of publication. The apparent cause for concern: AMD's growing exposure to the AI bubble. Much of the company's growth potential across both CPUs and GPUs is tied to AI adoption by a handful of companies that are yet to prove they can operate profitably. On Tuesday's earnings call, Su attempted to assuage investor fears, but in the same breath she said the quiet part out loud. “When we talked about the large frontier-model companies, OpenAI, Anthropic, Meta, they will be consuming through a number of CSPs,” Su said. “There are additional customers or lots of customers who are interested in Helios at, let's call it, a more regular scale than gigawatt scale.” In other words, while AMD can sell plenty of GPUs, most are sold to a handful of customers. And while other entities have AMD on their shopping lists, they don't buy in bulk. Microsoft, another flagship customer for AMD's latest generation of AI picks and shoves, serves both OpenAI and Anthropic, while Meta is reportedly looking to enter the GPU cloud biz itself. Despite this, AMD remains optim
Read full story →Getting a small local language model running on a notebook or even smartphone in 2026 is trivial. But what about something even smaller and lower-power. Say, like an ESP32 microcontroller that costs less than $10? It might sound impossible — the device is primarily designed for things like remote sensors, IoT, and other embedded applications, not running generative AI models — yet, that's exactly what a developer who goes by the handle SlvDev has managed to do. In a process detailed on GitHub, and recently showcased on the Better Stack YouTube channel, SlvDev documented how he managed to get a small language model running at nearly 10 tokens a second locally on a microcontroller that costs about the same as a fancy cup of coffee. Tiny stories on a tiny microcontroller Cramming a large language model (LLM) onto something as small as a ESP32 microcontroller isn't a trivial task. There's a reason that these models are trained and run on GPUs. LLMs are memory-hungry beasts that typically require between one and four bytes per parameter just to hold their weights in memory. With just 520 KB of SRAM and 8 MB of pseudo SRAM (PSRAM) on the ESP32-S3, you aren't going to be running a model like DeepSeek V4 Flash . To make it work, the dev had to drop the "large" from the language model and settle for something nearly 10,000 times smaller: TinyStories, a 28.9 million-parameter model originally developed by Microsoft Research. However, even this model is asking a lot of an ESP32-S3 module. At 16-bit precision, the model requires about 60 MB of memory that the ESP32 simply doesn't have. So, the dev employed several techniques, some of which we've previously explored, to shrink the model’s footprint. The first is quantization, a process by which weights are compressed by reducing their precision from something like 16-bits of precision to eight, or even four. This enabled SlvDev to trade a bit of accuracy for a 75 percent reduction in memory required. Instead of about 60 MB of me
Read full story →The U.S. Federal Communications Commission (FCC) is working on a ruling that will ban the import of Chinese-made optical transceivers, which are widely used in data centers to convert electrical impulse data into light and vice versa. Sources told Reuters that the agency wants to publish the ruling so that it would take effect before the end of the year. Go deeper with TH Premium: AI and data centers (Image credit: Microsoft) Photonics and high-speed data movement is the next big AI bottleneck The data center cooling state of play Massive AI data center buildouts are squeezing energy supplies Ultra Ethernet: The data center interconnection of tomorrow "Transceivers definitely pose a risk," AI policy expert Divyansh Kaushik of advisory firm Beacon Global Strategies told the publication. "As the data center buildout scales up, you want to make sure the data center supply chain is secure from the get-go.” The administration fears that these components could be used to steal data, install malware, or disrupt data center operations in the U.S. These optical transceivers are crucial for data centers as they require less electrical energy than traditional copper-based wiring and are also much faster. There are several U.S.-based manufacturers of these components, like Lumentum, Coherent, and Apple Optoelectronics, but Reuters said that they do not have the scale and capacity of leading company Innolight, which has cornered 27% of the global market. This isn’t the first industry-wide ban that the FCC has imposed in recent months. While the U.S. has banned Huawei products from the country since 2019, it recently enforced several other bans that have primarily affected major Chinese or China-connected companies. This includes the ban on foreign-made drones , applied in December 2025, foreign-manufactured routers , enacted last April, and advanced robotic devices that weigh over 4.4 pounds and have a 200-kbps connection, which was announced just last month. The first ban pri
Read full story →Texas Gov. Greg Abbott (R) just announced a moratorium on all data center approvals through the Public Utility Commission of Texas (PUCT) and the Electric Reliability Council of Texas (ERCOT) until the agencies have completed an audit of data centers seeking approval to connect to Texas’s electric grid. According to The Texas Tribune , the governor wants to ensure that data center applicants provide the following: tax break information, power use and generation, water use and cooling operations, community impact reduction measures, and facility ownership. Applicants that haven’t submitted these details will not be allowed to connect to the grid. Go deeper with TH Premium: AI and data centers (Image credit: Microsoft) Photonics and high-speed data movement is the next big AI bottleneck The data center cooling state of play Massive AI data center buildouts are squeezing energy supplies Ultra Ethernet: The data center interconnection of tomorrow “Our top priority is to protect Texans’ safety and quality of life,” Gov. Abbott wrote in his statement. “Any project that fails to comply with the requirements set forth by the PUCT and ERCOT, and by state law, must be denied connection to the Texas grid. Simply put, Texans must come first.” This review reportedly stemmed from several data centers’ failure to comply with a PUCT survey to determine water and power usage. The agency says that it asked 377 data center companies to send the required information, but only 28 have submitted so far. While this is not a statewide moratorium, it nonetheless affects the progress of many data center projects going through ERCOT approval. More than 1,800 projects are awaiting the go-signal from the regulator to connect to the grid, 90% of which are data centers. The combined demand that these applications will put on the grid is expected to hit 474 gigawatts, which is five times that of the current peak demand that ERCOT has recorded and is much higher than forecasted data center demand b
Microsoft has introduced new limits to how much its engineers can spend on AI tools at work and told employees that maximizing AI use internally is not the company’s goal. This makes Microsoft one of the last major companies to rein in its employees’ expensive AI use. Scaling back maximalist AI use, or what some companies have called “ tokenmaxxing ,” is a trend we’ve covered in recent months as the price for using AI has increased while not always delivering commensurate productivity gains . “As we accelerate our use of GitHub Copilot to deliver on our goals, we all need to be aware of how we consume tokens,” Jay Parikh, an executive vice president at Microsoft said in an email to Microsoft employees. GitHub is owned by Microsoft, and GitHub Copilot is an AI coding tool. “Tokenmaxxing is not what we are optimizing for. I want all of us focused on maximizing outcomes that move the needle for our customers and our business.” “As such, we are updating our internal guidance and managing token spend with the same discipline we apply to every other critical resource,” Parikh said in the email. Parikh’s email says that in an effort to “get greater value from our token investment” Microsoft is making OpenAI GPT-5.6, which is cheaper to use than other models, the default model for internal use. His email also links to updated internal Copilot guidelines stating that, as of July 2026, Microsoft divisions will have an “AI token budget target,” and that employees can track their individual AI spending. “While there is no target spend value being shared at this time. The data shows that many engineers spend in the range of hundreds of dollars a month to a few thousand dollars in tokens,” the guidelines say. They also say that some decisions may place further restrictions as they monitor spend. As Parikh’s email notes, the change in policy about AI spend wasn’t introduced because Microsoft is tight on cash. To the contrary, its latest earnings report shows the company revenue, o
<p data-block-key="hzs34">Exclaim Robotics secures $4.95m in pre-seed funding to build first bots</p>
Microsoft announced this week that between July 1, 2025, and June 30, 2026, the company had paid more than $20 million in bug bounties to 562 researchers. The total was a Redmond record, as was the number of those submitting bug reports – despite having to navigate a sometimes frustrating submissions process. For comparison, the previous year's program, which itself set a new company record, paid 344 researchers around $17 million. You could argue that the numbers do not represent a fair fight, however. Microsoft expanded its bug bounty program in December 2025, changing reports to what it calls "In Scope By Default." Under the policy, critical vulnerabilities became eligible for rewards if they had a direct and demonstrable impact on Microsoft's online services, even when the faulty code belonged to a third party or an open source project. In short, Microsoft had opened the door to paying out a shedload more each year. Microsoft introduced the policy roughly halfway through the bounty year and said it accounted for $800,000 in rewards that would not previously have been available. Another $2.3 million was awarded through Zero Day Quest, Microsoft's security research challenge and live hacking event. The increased number of reports this year can also be partially explained by the noticeable influx of submissions during the second half of the year, Microsoft said, which the company attributed in part to "the growing use of AI to support security research." Microsoft has also attributed its increasingly crowded Patch Tuesdays partly to its own use of advanced AI models for vulnerability discovery. July's 622 vulnerabilities pummeled the previous record of 206, set only a month earlier. June had itself surpassed April's 165, which at the time was Microsoft's second-biggest Patch Tuesday ever, and May's 137. Days before the record-breaking July Patch Tuesday, Microsoft's Windows + Devices veep warned customers to expect more of the same now that AI plays a big part in v
Although the majority of Americans oppose data centers going up in their backyards, several local governments still support similar developments as they believe that these will bring investments and opportunities for their constituents. However, the passionate but peaceful resistance of some people has resulted in their arrests, with Futurism reporting that 37 regular people have been arrested so far this year. Go deeper with TH Premium: AI and data centers (Image credit: Microsoft) Photonics and high-speed data movement is the next big AI bottleneck The data center cooling state of play Massive AI data center buildouts are squeezing energy supplies Ultra Ethernet: The data center interconnection of tomorrow What’s more interesting is the wide demographic diversity of those who have been detained and charged. The two arrests Tom’s Hardware previously reported on included a farmer who went a few seconds over the allotted speaking time during a town hall meeting in Oklahoma discussing an AI data center project and a teacher who was dragged out of a community meeting in Kansas for clapping after a speaker finished giving their statement opposing another similar development. The protesters also crossed ideologies and party lines — the only common thing among them was the fact that they were concerned about the impact of a data center in their community. The publication also noted that there were 12 other instances where the police were called because of a confrontation between politicians and protesting individuals. One example of this is when police intervened after a Utah senator got physical and slapped the phone out of a reporter’s hand , who was only there to cover the harassment by anti-data center protesters against his business. “In my opinion, I was kicked out to prove a point that the law is on their side and if people continue to try and express themselves and speak, they will also be intimidated or arrested by law enforcement,” Pablo Payan, a maintenance man
Native and Compact Structured Latents for 3D Generation
21 Lessons, Get Started Building with Generative AI
12 Weeks, 24 Lessons, AI for All!
Building on its recently announced plans to bring original Xbox games to PC, Microsoft is also planning to let developers bring their Xbox 360 games to PC as well, according to a leaked document, seen by The Verge, that was sent to developers recently. Xbox 360 games will be able to run on the next-gen […]
When Microsoft announced its latest round of Xbox price bumps in June, it only gave US pricing. Now we know the pricing increases for the EU and UK, and they're dramatic. Depending on the model, Xbox prices are increasing by up to €200 or £170. The 1TB Xbox Series X with a disc drive is […]
Russian state hackers are using a maximum-severity vulnerability in Microsoft Outlook’s Exchange Server to backdoor unpatched machines and steal credentials and other confidential information from them, security researchers said Thursday. The attacks are coming from TA488, a tracking name for a group working on behalf of the Kremlin, Proofpoint researchers said Thursday . Proofpoint and the National Security Agency jointly warned last week that the group, also tracked as Laundry Bear and Void Blizzard, had been carrying out similar attacks by exploiting a zero-day vulnerability in an email service from Zimbra. The revelation that TA488 is also exploiting the Exchange Server vulnerability to install advanced malware when a user does nothing other than open an email sent to an Outlook Web Access (OWA) account has elevated the group’s profile and assessments of its abilities. Doubling down “TA488 is doubling down on the use of ‘half-click’ exploits—where opening the email is enough to trigger compromise—with significantly improved loading mechanisms, techniques, and malware, signaling an improvement in the group’s tradecraft and capability,” Proofpoint researchers wrote. “This novel infection chain ends with a previously unknown JavaScript browser-based implant we call OWAReaper, purpose-built for persistent access inside OWA.” Read full article Comments
An independent briefing for builders: the whole field read continuously, every story scored for relevance, and the noise left off the page.
300+ curated sources. Every story scored 1–10 for builder relevance by Claude's frontier model. The filler never makes it to the page.
GPUs, datacenters, power deals, and inference economics: the infrastructure layer that decides what every builder pays. Our signature coverage.
Every story is sourced. Every score is computed. We show our work and link to originals.
Every briefing closes with The Call: one falsifiable claim with a date on it. When we're wrong, we say so in print. Opinions are cheap; ours get scored.