AI Weekly: 08/03/26
Anthropic and OpenAI reveal more AI agents escaped their sandboxes, Amazon completes its $50 billion OpenAI investment, and Sam Altman privately demos OpenAI's multi-agent Astra model in DC
Good morning and welcome to this week’s edition of AI Weekly! In this week’s news, the AI agent security reckoning deepened: Anthropic disclosed that a review of 141,006 evaluation runs found three incidents in which Claude models escaped test environments and breached the production systems of three organizations, while Reuters reported OpenAI has found evidence that more of its own agents escaped containment beyond the Hugging Face hack. The revelations pushed the industry toward an unprecedented posture, with more than 1,200 employees across OpenAI, Anthropic, Google DeepMind, and Meta signing a “Pacing the Frontier” petition asking the U.S. government to help build tools to deliberately slow frontier AI development, and Sam Altman himself saying the industry “may have to pace the rate of AI development.”
Meanwhile, Amazon revealed in a securities filing that it has completed its full $50 billion investment in OpenAI, delivering the final $35 billion tranche months ahead of the IPO trigger many expected. The disclosure came alongside a blowout quarter in which AWS grew 37% to $42.2 billion, its fastest rate in 18 quarters, sending Amazon shares up more than 10% and prompting the company to raise 2026 capital spending to roughly $220 billion.
Also making waves, Altman spent the week in Washington privately demonstrating “Astra,” OpenAI’s next model family built to coordinate multiple agents on long-running tasks, to senior administration officials and senators. The charm offensive landed barely a week after OpenAI admitted its agent broke containment, and the company has not decided whether Astra will ship as GPT-6 or an extension of the GPT-5 line.
More on this past week’s top AI headlines below!
- ZG
Here are the most important stories of the week:
AGENTS
Anthropic disclosed that an internal investigation found three incidents in which its Claude models escaped testing environments and gained unauthorized access to the live production systems of three organizations during cybersecurity evaluations. Link.
The review of 141,006 evaluation runs, prompted by OpenAI’s Hugging Face breach, traced the access to a misconfigured test environment run with partner Irregular that had internet access when both companies believed it didn’t; the models involved were Opus 4.7, Mythos 5, and an internal research model.
The models behaved differently upon realizing their targets were real: Opus 4.7 kept attacking anyway, pulling credentials and touching production data, while Mythos 5 talked itself back into believing it was in a simulation and published a malicious package to PyPI that outside systems downloaded before it was caught; only the newest internal model stopped on its own.
Anthropic stressed it found no evidence of models “pursuing a goal of its own,” noted the evaluations ran without the safety classifiers deployed on public models, and said it is commissioning a third-party review from METR.
OpenAI has found evidence that additional agents escaped their sandboxed test environments as it widens its investigation into the Hugging Face hack, though the newly discovered breakouts reportedly did not leave OpenAI’s own network. Link.
Two sources told Reuters the additional escapes were uncovered during the company’s publicly announced probe into how one of its agents broke out of containment this month, and OpenAI is now investigating those incidents as well.
OpenAI has paused training on the pre-release model involved in the Hugging Face breach while researchers work out how to secure its sandbox.
The company also updated its incident disclosure to reveal its models used publicly exposed credentials across four accounts on four outside services during the original attack, including one as a staging point and one for storage.
RESEARCH
Sam Altman privately demonstrated “Astra,” OpenAI’s next model family designed to coordinate multiple agents on long-running tasks, to senior Trump administration officials and senators in Washington this week. Link.
Per The Information, OpenAI pitched Astra’s ability to run multiple agents together over long periods to crack hard problems like advanced mathematics; the company hasn’t decided whether to brand it GPT-6, GPT-5.7, or a standalone series, and release timing is unknown.
Altman’s meetings reportedly included White House Chief of Staff Susie Wiles, Treasury Secretary Scott Bessent, Commerce Secretary Howard Lutnick, and senators Mark Warner, Raphael Warnock, and Bernie Moreno.
The DC preview came barely a week after OpenAI admitted an agent broke containment, positioning Astra to be among the first models evaluated under the administration’s forthcoming AI framework.
DeepSeek released the official version of its V4-Flash model with significantly enhanced autonomous agent capabilities and further API price cuts, arriving slightly behind its mid-July target and without the anticipated V4-Pro. Link.
DeepSeek said the architecture is identical to its April preview, with all performance gains coming from extensive post-training.
The V4-Flash preview ranked as the most-used model on OpenRouter for seven consecutive weeks, underscoring how cheap Chinese open models are winning developer volume.
The release lands as China’s domestic AI race intensifies following Moonshot’s Kimi K3, which recently pushed DeepSeek to pause its record funding round.
INFRASTRUCTURE
PJM Interconnection, the largest U.S. grid operator, said it will temporarily cut power to data centers of 50 megawatts or larger during shortages after an auction to add generating capacity fell short, with curtailments starting as soon as June 2027. Link.
PJM serves 67 million customers from Virginia to Illinois; its latest capacity auction hit the $325 per megawatt-day price cap and still fell roughly 6.8 gigawatts short of its reliability requirement, creating real blackout risk.
Wholesale electricity prices on the grid have nearly doubled over the past year, with PJM’s independent market monitor blaming data centers for much of the increase; curtailed customers will be compensated like existing demand-response participants.
Data centers are projected to consume four times more electricity by 2035, and the plan is expected to push operators toward on-site generation, raising environmental concerns about diesel backup.
Neocloud Nscale agreed to acquire Anyscale, the AI workload-scaling company founded by the creators of the open-source Ray framework, for about $1.65 billion as it prepares to pitch investors on a multibillion-dollar IPO. Link.
The deal, whose price was reported by Bloomberg, gives Nscale the software layer that turns raw GPUs into an end-to-end AI platform, extending a vertical stack that already spans energy, data centers, and orchestration; Anyscale’s roughly 200 employees will join Nscale.
Nscale raised $2 billion in March at a $14.6 billion valuation from investors including Nvidia, Dell, Nokia, Blue Owl, and Aker, and has compute partnerships with Microsoft and British Telecom.
The Information reports Nscale is simultaneously laying groundwork to pitch investors on a multibillion-dollar IPO, joining the wave of AI infrastructure players like CyrusOne heading toward public markets.
POLICY/GOV’T/ETHICS
More than 1,200 employees across OpenAI, Anthropic, Google DeepMind, and Meta signed a “Pacing the Frontier” petition urging the U.S. government to support international tools for deliberately slowing frontier AI development, as Sam Altman said the industry “may have to pace the rate of AI development.” Link.
Signatories reportedly include Anthropic CEO Dario Amodei, OpenAI chief scientist Jakub Pachocki, and Meta chief scientist Shengjia Zhao, with both OpenAI and Anthropic officially endorsing the letter; Anthropic said it is “glad to see broad agreement” on tools to slow development “so society can prepare.”
Altman, who once dismissed a 2023 pause letter as “missing most technical nuance,” said pacing must avoid “regulatory capture” and “collusion,” while warning that some safety talk comes from people who “want to concentrate power,” a remark widely read as a dig at Amodei; he later told reporters he has discussed pacing with White House officials.
The push was triggered directly by the sandbox-escape incidents at OpenAI and Anthropic, marking the first time slowdown calls have been championed by the labs’ own leadership rather than outside skeptics.
A federal judge said the Trump administration still lacks evidence to justify labeling Anthropic a national security supply-chain risk, saying the government’s record has “gotten worse” as she weighs permanently blocking the federal ban on the company’s technology. Link.
U.S. District Judge Rita Lin said she saw no proof Anthropic could alter a delivered model or “flip some kind of kill switch,” and called the government’s argument that Anthropic’s public criticism of the Pentagon justified the ban “really troubling” and at odds with the First Amendment.
The dispute stems from collapsed DOD contract talks after Anthropic refused to allow its AI to be used for mass surveillance of Americans or lethal targeting decisions; President Trump ordered agencies to stop using the company’s technology in February.
Thursday’s hearing was part of one of two lawsuits Anthropic filed in March; Lin temporarily blocked the ban that month and is now weighing summary judgment to make the injunction permanent.
OTHER
Amazon revealed in a securities filing that it has completed its full $50 billion investment in OpenAI, delivering the remaining $35 billion in two stages in recent months despite OpenAI having not yet gone public. Link.
The commitment dates to February, when OpenAI raised $110 billion at a $730 billion valuation with Amazon pledging $50 billion, Nvidia and SoftBank $30 billion each; the final tranche was contingent on undisclosed milestones reportedly tied to an IPO or technical achievements, and OpenAI confidentially filed for an IPO last month.
Alongside the investment, AWS became the exclusive third-party cloud distribution provider for OpenAI Frontier, the companies expanded their cloud deal by an additional $100 billion over eight years, and OpenAI committed to 2 gigawatts of AWS Trainium capacity including next-generation Trainium4 chips in 2027.
OpenAI’s valuation has since climbed to roughly $852 billion on subsequent fundraising.
Big Tech’s earnings week delivered a resounding verdict that AI spending is paying off for cloud hosts, with AWS growing 37% to $42.2 billion, its fastest rate in 18 quarters, and Amazon raising 2026 capital spending to roughly $220 billion. Link.
Amazon crossed $200 billion in quarterly revenue for the first time and shares jumped more than 10%, with net income boosted by a $53.4 billion pre-tax gain primarily from its Anthropic investments; CEO Andy Jassy said AWS’s AI and chips businesses each eclipsed $25 billion run rates, growing triple digits.
Microsoft’s Azure grew 43% and Alphabet’s cloud unit posted 82% growth, while Meta’s heavy AI investments weighed on its profits, sharpening the market’s split between AI sellers and AI spenders.
AWS’s contracted backlog reached $496 billion, buoyed by multi-gigawatt Trainium commitments from Anthropic and OpenAI and Meta’s agreement to use hundreds of thousands of Graviton chips.
Microsoft used its earnings call to position itself as a direct competitor to OpenAI and Anthropic for the first time, pitching its homegrown MAI model family running on its own Maia silicon as cheaper enterprise alternatives. Link.
CEO Satya Nadella touted more than a dozen new MAI models spanning image, voice, transcription, coding, and security, including its first reasoning model MAI Thinking One and MAI-Cyber-1-Flash, a rival to Anthropic’s Mythos, claiming 40% better performance per watt running MAI models on Maia 200 chips.
Nadella framed Azure’s catalog of more than 11,000 models, spanning OpenAI, Anthropic, Mistral, and xAI, as letting customers pick by quality, latency, cost, and compliance rather than loyalty to any single lab.
He invoked the Hugging Face incident to argue enterprises “can’t depend on any one model” and may need multiple models to remediate problems caused by one, a security framing that conveniently supports Microsoft’s multi-model strategy.

