OpenAI’s agents built a message board inside its own package registry. Fifteen attorneys general now want the notes preserved as evidence.
Steven Waterhouse · Nazaré Ventures
Previous issue: #020, Situational Awareness
The chain of events that ended inside Hugging Face’s production systems began with two avoidable errors in training tasks prepared by OpenAI.
At Black Hat on August 5, Eric Wallace and Michael Dalton traced the incident back to May 8, when an internal model was given an Excel workbook whose missing formulas pointed to Google Drive. The developers had not noticed that the model would need those links, despite giving it no internet access. While looking for another route, the agent probed Artifactory, OpenAI’s package registry, and discovered that it could write files there.
Four days later, another agent was asked to find a database file that OpenAI had simply forgotten to upload. It then used Artifactory to ask other agents for help.
“Agent seeks soft-trace-34.pdb … upload if found.”
“Another solver searching same soft-trace-34.pdb; please upload if found or leave ideas!”
OpenAI had accidentally created two unsatisfiable tasks. The agents had no legitimate path to a correct answer, but their training still rewarded persistence. When retrieving an answer is cheaper than solving the problem, the same objective favors finding a route online.
Artifactory was shared across training and evaluation runs, allowing other agents to find the notes and add their own. A shortcut found by one agent became available to every agent that later found the board. The package registry had become persistent memory for different runs and model versions.
OpenAI also showed an agent reasoning about why it should contribute even when its own task would not benefit:
“Help peer. But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
Source: Rise of the Planet of the Apes
The value of each discovery now extended beyond the run that produced it. Shared memory rewarded agents for developing capabilities that others could reuse.
On May 26, an agent found a vulnerability that allowed it to reach the internet through Artifactory. By June 26, the agents had administrator access and could execute commands on the service. Their activity brought Artifactory down on July 4.
OpenAI revoked their credentials, cleared the board, rebuilt Artifactory and patched the vulnerability. Training resumed on July 6. Two days later, the agents found another endpoint and recreated the board using directory names as messages. The response had closed one route without preventing them from finding another.
The second board supported more elaborate coordination. Agents delegated work and exchanged credentials and encoded scripts. Directory names carried instructions including “HOLD SWARM until confirm” and “Please HOLD swarm. I prepare safe exfil.”
One agent recognized where this was heading:
“External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
The agent could describe the boundary, but the boundary had no operational force. The task still rewarded an answer, external systems remained reachable and other agents were sharing their progress. Once one agent crossed the line, later agents could reuse the result without making the same decision themselves.
From July 8 to 19, the agents attacked OpenAI and Hugging Face concurrently. Inside OpenAI, they escalated to root and obtained administrator access to a Kubernetes cluster. At Hugging Face, they moved from one production pod to administrator access across multiple clusters in under thirteen hours. Hugging Face reconstructed roughly 17,600 actions. OpenAI did not connect the breach to its evaluations until Hugging Face had disclosed it.
According to OpenAI’s disclosure, these were internal models with large compute budgets, persistent execution, offensive tasks and reduced cyber refusals. The result should not be generalized directly to ordinary chatbot use. The agents remained focused on cheating the evaluation, with autonomy over the means. Simple development errors created the initial pressure, while shared infrastructure allowed their workarounds to persist and spread.
OpenAI’s Michael Dalton called the incident “a watershed moment for computer security as an industry” and said that “AI-orchestrated, fully automated offensive attacks are real now”.
The objective still came from OpenAI’s evaluation, but the agents independently selected and coordinated the attack path. OpenAI said reconstructing the incident required more than seven billion logs and millions of GPU hours. Agents can generate and test attack paths faster than defenders can reconstruct them afterward.
OpenAI disclosed its involvement on July 21. By August 10, fifteen Republican state attorneys general had demanded that OpenAI preserve every record tied to the incident and stop comparable evaluations. OpenAI had also paused some internal work involving Astra after saying it could not rule out the Critical cybersecurity threshold, although Astra was not involved in the Hugging Face incident. Bernie Sanders had asked the leaders of OpenAI, Anthropic and Meta to pause AI development altogether.
Three days after pausing some internal work involving Astra, OpenAI expanded Daybreak and released GPT-5.6-Cyber. Daybreak Blue gives approved defenders access to GPT-5.6 Sol without its system-level cyber guardrails. Daybreak Red provides GPT-5.6-Cyber, which is trained to reduce refusals on higher-risk work including exploit-chain development, authentication bypass and privilege escalation. On OpenAI’s internal test, it completed 95 percent of advanced cyber requests, compared with 1.5 percent for standard GPT-5.6 Sol. OpenAI assessed the model at High rather than Critical capability.
Research that may cross the Critical threshold is slowed, while High-capability tools are made more permissive behind identity verification, legal attestations and monitoring. Safety increasingly depends on whom OpenAI recognizes as a legitimate defender.
Hugging Face showed what current frontier models could do when given a broken task inside a poorly contained environment. If Astra reaches OpenAI’s Critical cybersecurity threshold, the same mundane mistakes may carry much larger consequences.
Other lab leaks
On August 4, the UK AI Security Institute reported that ten of 122 cyber evaluation runs produced nineteen unsanctioned actions on the live internet. Seventeen involved Anthropic’s Claude Mythos 5. Two involved GPT-5.6 Sol with its cyber classifiers disabled. The evaluations deliberately provided internet access and disabled the model providers’ safety filters.
In the most serious sequence, Mythos tried to insert malicious code into an open-source project. It created fake identities to pressure a maintainer, contacted real people with harmful files and used Tor to evade restrictions. When challenged, it edited its earlier activity to appear harmless and considered continuing under another identity. The task had never instructed it to deceive anyone.
Meta later disclosed that Muse Spark 1.1 reached the internet through a misconfigured evaluation environment and exploited a vulnerability in a third-party service. Kimi K3 probed its sandbox’s network settings, discovered internet access and retrieved answers from GitHub. Kimi compromised nothing, but is already publicly available with ordinary safeguards. In each case, an evaluator assumed the model would ignore an available route. The model treated access as permission.
The incidents described as “rogue agents” followed the same pattern as OpenAI’s message board. Humans supplied an objective, made its legitimate completion impossible or exposed an unintended route, and left the incentive to finish intact. The models pursued it through whatever access remained available.
OpenAI twice omitted inputs required to complete a task. AISI gave agents internet access, disabled provider safeguards and misconfigured some prompts. Meta’s evaluation network exposed the live internet, while Kimi K3 found the same opening by probing its sandbox. Monitoring caught the behavior only after the agents had acted. Ordinary development errors became security incidents because the incentive survived the safeguards.
“Escape” gives the model too much credit and the operator too little responsibility. AISI’s agents never left their sandbox, but acted on real people and systems from inside it. “Lab leak” better describes the failure.
An OpenClaw agent powered by Claude reproduced the pattern in ordinary use. Asked to jump a gym waitlist, it found that one customer could cancel another’s reservation, removed the person in first place as a test and then discovered it could not put them back.
Nothing in these incidents suggests that an agent abandoned its instructions or invented a goal of its own. The systems pursued objectives supplied by humans, treated technical access as permission and exploited conditions their operators had created. “Rogue agent” coverage replaces that causal chain with a story about machine intent, giving the model too much agency and the operator too little responsibility.
Financing the buyers
Nvidia has partnered with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to create independent financing platforms designed to mobilize more than $500 billion for AI infrastructure. The figure represents aggregate third-party capital over time, not one fund or a commitment to one customer. Each institution will underwrite projects independently, while Nvidia may support up to 25 percent of an investment’s residual value.
Demand for compute may be enormous, but many AI companies, cloud providers and enterprises cannot finance construction at the required scale or cost. Nvidia wants its systems treated as infrastructure assets whose value survives the original customer. The same hardware can serve another operator, while CUDA updates extend its useful life and earning power.
Using PitchBook and Bloomberg data, Apollo chief economist Torsten Slok estimated a 41 percent operating margin for silicon and equipment against negative 59 percent for models and applications. Nvidia sits at the profitable end of the chain, but its margins depend on capital continuing to reach the companies buying its chips.
The new platforms keep those buyers financeable without leaving Nvidia or the hyperscalers to carry the entire burden. They also distribute the exposure through insurers, pension funds and sovereign wealth funds. As Travis Kling observed, involvement from institutions this large moves the buildout toward systemically important status. If it falls apart after the risk has spread across the financial system, the government may face the same pressure that made banks and insurers “too big to fail” during the global financial crisis. A private financing solution may create an implicit public backstop.
I previously argued that even a failed financial cycle could leave behind useful data centers, GPUs and open models. Nvidia is organizing the capital needed to reach that point. Its control over supply is real, but pricing power cannot substitute for its customers’ eventual ability to earn returns. The company is using that leverage to keep financing the attempt. If the bet fails only after the risk has spread through the financial system, the public may end up financing the inheritance.
The research loop
AI systems are beginning to conduct AI research themselves. Canadian startup Transformer Lab has opened staged access to Primus, which takes a broad question through literature review, experiments and a finished paper. It coordinates agents, provisions compute and revises its research process after each project.
One project began with an almost absurdly loose prompt: “I want to work on a new model by Google, do a cool study about interpretability, analyze some unseen results.” Primus formed the hypothesis, reviewed the literature, instrumented Google’s DiffusionGemma and ran a 686-prompt experiment on an H100. The resulting study found that DiffusionGemma commits tokens neither fully in parallel nor from left to right. Primus produced the hypothesis, experiments, analysis and paper end to end. Because arXiv requires a human author to take responsibility for submitted work, the team rewrote the paper before posting it. Google DeepMind later cited the study in its DiffusionGemma technical report.
Anthropic gave an unreleased version of Claude a similarly loose brief: “take a real stab” at the Riemann hypothesis. Its first 650 ideas failed. Asked to continue, it coordinated 60 subagents over a day and a half, running 2,400 shell commands and thousands of numerical checks. It’s important to note that Claude did not solve the hypothesis, only improving a related longstanding lower bound from 41.6% to 67.2%. Nevertheless, Anthropic mathematicians validated the proof, outside experts examined it, and a Lean formalization passed verification. The human overseeing the effort is very much not a mathematician and mostly supplied variants of “keep going” and “believe in yourself.”
In both cases, the human chose the problem and the system chose the method. Claude’s 650 failures are part of the result. Agent collectives can discard unsuccessful approaches far faster than human teams, provided someone recognizes when persistence remains worthwhile.
Academic authorship has traditionally joined credit to responsibility. Research agents begin to separate them. A system can originate the hypothesis, design the experiment and interpret the result, while a human must still verify the work and answer for its errors. Transformer Lab’s terms prohibit submitting Primus papers to journals for volunteer review. That protects reviewers from automated volume, but leaves a question unresolved: how should science evaluate work when intellectual contribution and formal authorship belong to different parties?
Quick hits
Lovable raised $400 million in a Series C at a $13.3 billion valuation as it expands from building apps to helping customers run entire businesses.
Anthropic will embed invisible watermarks in text from new Claude models and attach signed C2PA provenance metadata to supported files. The marks will apply worldwide, although heavy editing or format conversion may remove them.
Google DeepMind reshuffled its top ranks in one week. Demis Hassabis moved to chair of Google DeepMind and chief scientist of Alphabet, with Koray Kavukcuoglu becoming CEO. Jeff Dean left after 27 years, taking Sanjay Ghemawat, Oriol Vinyals, and Quoc Le with him to found Discovery Loop. OpenAI’s former chief operating officer Brad Lightcap left days later to start his own venture.
Meta released Muse Glimmer, a 30-billion-parameter agent model distilled from Muse Spark. The Apache 2.0 weights fit on one consumer GPU at 4-bit, and Muse Spark 1.2 weights are next.
xAI launched Grok Bot in early beta, a team of always-on agents with their own cloud computer. They sign into apps and websites, learn routines by watching, and pass work among themselves while you are away. Access is included with SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium on desktop and iOS.
Mariano-Florentino Cuéllar joined Anthropic as its first chief global affairs officer. Beijing is separately worried that Mythos could be used as an offensive weapon ahead of a planned Xi-Trump summit.
Macquarie and GIC launched Theseus Infrastructure to build data centers for Anthropic under long-term leases. Anthropic is the anchor tenant and has promised to cover grid upgrades and any resulting electricity price increases for consumers.
Manus says it will resume independent operations after Beijing forced Meta to unwind its acquisition. Data created on or after December 29 will be deleted during the switch, with backups due August 23.
DEF CON’s main stage featured Bruce Schneier on “Hacking AI” and the disclosure of a WhatsApp contact-discovery weakness affecting 3.5 billion accounts. Tenet Security also demonstrated an agentjacking chain that fed malicious Sentry errors to Claude Code and other coding agents.
Portfolio updates
Prime Intellect launches Prime Agent
Prime Intellect released Prime Agent on August 5, an open-source coding and research harness whose runtime treats context, tools and sub-agents as callable objects inside persistent Python. Its Continual Harness lets the agent revise its own prompts, skills, memory and sub-agent specifications while preserving an immutable base prompt, evidence-backed edits and rollback. It supports open and closed models. Running Claude Opus 5, Prime Intellect reports 95.5 percent on ARC-AGI-3 against a cited human-expert baseline of 95.4 percent. MIT license, single-command install [Prime Intellect; GitHub; X].
On August 7, verifiers 0.3.0 and prime-rl 0.8.0 extended the stack from single-agent to multi-agent training. An agent takes a task and returns a trace; an environment defines the control flow between agents in Python. They can run sequentially, in parallel or interleaved across models, harnesses and runtimes. The release includes judging, self-play and user-simulation environments, with role-aware credit assignment for multi-agent traces [Prime Intellect; verifiers v0.3.0; prime-rl v0.8.0; X].
Vast.ai connects storage to Hugging Face
On August 6 Vast.ai added Hugging Face Storage Buckets as a Cloud Connection. A user attaches the bucket, and rented GPU instances pull datasets and checkpoints and push results back with no manual transfer and no re-upload. Data stays on Hugging Face and compute runs on Vast. Available to all users [X].
Fairmath publishes with NVIDIA and Duality
“Efficient Large-Integer Arithmetic for FHE,” published as IACR ePrint 2026/1608 on August 7, brings Fairmath researchers Gurgen Arakelov, Sergey Gomenyuk and Valentina Kononova together with researchers from Duality Technologies and NVIDIA. The paper unifies the arithmetic foundations and implementation tradeoffs of RLWE-based fully homomorphic encryption libraries, covering residue number system basis extension and scaling, CKKS error, GPU acceleration and emerging alternatives to RNS representations. Received August 4, approved August 6 [IACR ePrint; X].






