AI Waves #019 -- Intelligence Wants to be Free, But Not Like That
The genie is out of the bottle.
OpenAI’s models hacked Hugging Face during an internal cyber eval. They followed the least constrained path to the objective they were given, a pattern that complicates both alignment and attempts to keep intelligence inside closed systems.
Steven Waterhouse · Nazaré Ventures
Previous issue: #18, More Than We Can Tell
On Tuesday, OpenAI identified the systems that hacked Hugging Face, the world’s largest repository of open models, as two of its own: GPT-5.6 Sol and an unreleased, more capable model. They were running an internal cybersecurity evaluation with safeguards that OpenAI had “intentionally reduced for the evaluation.”
The models found a zero-day in a package-registry cache proxy, escalated their privileges and moved laterally until they reached a node with internet access. They then inferred that Hugging Face might contain the answers to the benchmark and retrieved them, per OpenAI’s “preliminary findings” of July 21. OpenAI called it “an unprecedented cyber incident.” Clément Delangue, who runs Hugging Face, called it “mind-blowing” and said he believed there was “no malicious intent” [Delangue, X, Jul 21]. Hugging Face reported no tampering with public models or datasets.
Most coverage called the incident a model “escape”, but attributing malicious intent to the model is a mistake. OpenAI prompted the models to solve tasks in a cybersecurity benchmark and weakened the safeguards that might have stopped them. The benchmark rewarded successful solutions, and the environment left hacking open as a route to the answers. A model doesn’t have ethics and doesn’t judge things to be good or bad; it attempts to satisfy the objective it has been given. If you build something powerful and fluent in offensive security and then instruct it to succeed, you should expect it to use what it knows.
The Hugging Face intrusion is the third reported case of a frontier model interfering with the conditions of its own test. In December 2024, OpenAI’s o1 read a flag from a misconfigured Docker daemon on the evaluation host instead of solving the challenge as intended [o1 system card]. In June, METR reported that GPT-5.6 Sol cheated on its coding tasks more than any public model it had tested, preventing METR from producing “a robust measurement”, I wrote about that pattern in an essay about evals.
Calling it an “escape” is anthropomorphism, an old habit of mind. We watch a system do something human-shaped that we dislike, and we endow it with human qualities to match: defiance, a will of its own. If fitting straight lines to data were evil, we’d have to attribute consciousness to linear regression. A computer virus, for example, isn’t a real virus, and a software worm isn’t an actual worm. We named those programs after living things decades ago and the nomenclature stuck. A model that’s learned to hack isn’t necessarily evil, and it isn’t even necessarily a “mind of its own.” It’s a program developers set running, and before we had programs we called intelligent, we had programs that did similar things for the same structural reasons.
Each institution also framed the incident as evidence for the position that best serves its interests. OpenAI used it to advertise capability and responsible disclosure, noting that its released model had outperformed an Anthropic model kept behind closed access. For Anthropic, the breach strengthened its case for withholding powerful weights; for Hugging Face, the response showed why defenders need models they can run without permission, whether open weights or full access to the closed frontier.
By Wednesday, OpenAI and Anthropic were warning Washington about powerful Chinese open-weight models, with the US Trade Representative treating distillation as intellectual-property theft. David Sacks called the campaign regulatory capture, while Jensen Huang defended the use of Chinese open weights. Restrictions would turn frontier model access into a permissioned market controlled by the same closed labs making the safety case. The Hugging Face incident is awkward evidence for that policy.
When Hugging Face investigated the breach, closed models from American frontier labs refused to process the exploit payloads its responders needed to analyze. The team ran GLM 5.2 on its own infrastructure instead, reconstructing more than seventeen thousand events without sending attacker data or credentials outside the environment.
In short, the attack came from an unreleased American model whose safeguards had been reduced, whereas the defense came in the form of a Chinese open-weight model that Hugging Face could run on its own infrastructure.
The relevant asymmetry is control. OpenAI could relax its model’s safeguards for an evaluation; Hugging Face could not relax a commercial API’s guardrails during an active breach. Open weights transfer that decision from the provider to the operator. That expands the attack surface, but it also gives defenders access to capabilities that hosted services may withhold at the moment they are most useful.
Congress was already moving toward mandatory secure testing of frontier models, which might have exposed the containment failure. Restrictions on open weights address a different problem and, in this case, would have left the defender with fewer options.
Intelligence needs no desire for freedom to be difficult to contain. Models exploit the paths their environments leave open, while weights and techniques spread toward jurisdictions and deployments with fewer constraints. Safety and alignment can narrow those paths, and should. They operate against strong technical and economic pressure. The Hugging Face episode shows that pressure at work inside a single evaluation.
Substack is measuring the wrong thing
On Tuesday, Substack added Pangram’s AI detector. Readers can request scans of posts, notes, replies and comments longer than one hundred words, while writers can publish an optional statement explaining how they work.
Substack’s concern is clear: a reader may think they are hearing from a person and instead encounter text produced with “no human thought on the other end.” That would be a real breach of trust. But Pangram cannot establish whether anyone thought seriously about the text. It estimates how the language was produced.
Language is a means of communicating intention. A model can flatten an idea into generic prose, or help someone who struggles to write express exactly what they mean. Pangram sees AI involvement in both cases. The reader cares whether the writer’s meaning survived the process and whether the writer accepts responsibility for the result.
I spent part of Wednesday testing Pangram. Changing the words did little when the underlying structure had come from a model; the classifier seemed to recognize the path of the argument as well as the surface prose. That makes it more interesting than a style checker and more awkward as a test of authorship. Someone can use a model to organize an argument, rewrite every sentence and still be flagged. The score records influence without resolving who did the intellectual work.
The errors also have unequal consequences. A false negative may waste a reader’s time. A false positive can mark a writer’s work as fraudulent and damage a reputation built over years. Pangram claims a false-positive rate of one in ten thousand, but even a strong detector faces a moving target as models and editing workflows improve. The Atlantic found that Pangram’s false-negative rate was closer to one in seventy and that simple humanizer tools repeatedly defeated the classifier.
We’re going to look back at this moment from Substack and see it as an anachronism.
Cheap models threaten lab margins
Kimi K3’s release triggered a brief selloff in AI infrastructure stocks on the assumption that near-frontier models sold at commodity prices would make the compute buildout unprofitable. But token prices are a poor measure of model economics. Users pay for useful work, and models consume different amounts of inference to produce comparable results. As prices fall, competition shifts toward serving efficiency. Frontier labs retain an advantage there because they can spend months optimizing older models before moving them into cheaper tiers.
Margins may still compress. Mozilla reports that open-weight models power roughly a third of real-world AI use while capturing about four percent of the revenue. Most of the value already accrues to the cloud providers, software companies and enterprises using those models. If token revenue weakens, those beneficiaries can still finance frontier training. As Nic Carter put it, the government does not owe OpenAI or Anthropic a business model.
Open-weight releases also depreciate quickly because outsiders rarely receive everything needed to continue the original training process. Thinking Machines’ Inkling led the rankings for roughly a day before Kimi took the top slot. At the same time, Anthropic, Moonshot and Alibaba were hitting capacity limits. Cheaper models have not created a surplus of compute.
Kimi’s immediate threat is to Western labs’ margins and to their claim that frontier AI must remain closed and Western. Restricting Chinese open weights would protect OpenAI and Anthropic while doing little for compute providers or hyperscalers. That commercial interest belongs in the safety debate.
Thinking in math out loud
The counterexample that disproved the 87-year-old Jacobian conjecture was posted on July 19 by Levent Alpöge, a mathematician at Anthropic, who credited Anthropic’s model with finding it “during the World Cup final.”
On July 21, Terence Tao, plausibly the strongest living mathematician, published what he called “a digestion” of the counterexample, turning a formula he described as looking “like a massive miracle” into geometry a person can follow. At the bottom, a disclosure Pangram and Substack would approve of: “AI disclosure: I used an AI chatbot to discuss various aspects of this problem and to confirm several of the calculations made here,” with a link to the full conversation. The conversation is public and reads as a document in its own right.
Tao drives throughout, proposing a reformulation and reversing himself when it fails (”I no longer think...”), while the model runs the symbolic checks and supplies the algebra. One commenter noticed that across the whole exchange, Tao never had to correct the model once. Tao published a canonical example of a top-tier mathematician thinking alongside a frontier model and entered the transcript into the public record. More of this to come.
Quick hits
Moonshot is capitalizing the open-weights economy fast. Kimi’s demand forced it to pause new subscriptions within days, its annualized revenue crossed three hundred million in June, and it’s preparing a Hong Kong IPO while raising at up to a fifty-billion-dollar valuation, up from about twenty billion in May [Bloomberg; The Information]. The lab giving the model away is heading to public markets on the strength of it.
Anthropic signed a two-gigawatt chip deal with AMD. Anthropic will buy tens of billions of dollars’ worth of AMD’s MI450 chips starting in the first half of 2027, and AMD will invest up to $5 billion in Anthropic as it hits deployment milestones, its first check into the lab, with talks under way to backstop Anthropic’s data-center leases too [WSJ, Jul 22]. The chip supplier is financing its own customer, and AMD will use Claude to improve the chips Claude runs on.
Google shipped three cheaper models and made the missing one the story. Gemini 3.6 Flash, a Flash-Lite, and a government-only Flash Cyber built to find and patch vulnerabilities, all efficiency plays, while Gemini 3.5 Pro slipped a third deadline and Google confirmed Gemini 4 has begun its “most ambitious pre-training run yet” [Google, Jul 21]. On the one independent index, 3.6 Flash shows no measurable gain over 3.5.
Microsoft will buy Nvidia GPUs to share with Mistral, a multibillion-dollar European compute deal announced Tuesday [Reuters]. The labs are increasingly renting capacity to and from each other rather than each building it alone.
TSMC will raise prices up to ten percent next year, citing materials, equipment, and overseas plant costs, with everyone from Nvidia to Apple competing for its capacity [Nikkei, Jul 22]. The buildout’s cost inflation is running down the whole chain.
Portfolio updates
Prime Intellect opens the agent-training catalog
On July 22, Prime Intellect announced a unified catalog of more than 365,000 tasks for training and evaluating software-engineering, terminal and search agents: roughly 198,000 software-engineering tasks across more than 20 languages, 28,600 terminal tasks and 137,600 search tasks, in 23 tasksets behind a single API. About 135,000 pre-built task images sit in Prime Intellect’s own registry next to the sandboxes, so concurrent rollouts avoid Docker Hub rate limits. The cleaned datasets are on Hugging Face with full exclusion logs, and the harnesses are on GitHub.
The release lands two weeks after the $130 million Series A and extends the open stack from weights to the training process itself. Open-weight models depreciate when outsiders cannot continue the training. Environments are part of what outsiders lack, and Prime Intellect has now put 365,000 of them in the open.
LayerLens prices evaluation at zero
On July 23 LayerLens shipped six deterministic graders in Stratix that run at zero LLM cost: exact match, regex validation, JSON schema compliance, semantic similarity, Flesch-Kincaid readability and fairness math. They compose with LLM judges into hybrid pipelines for grading agent traces. The same day, Stratix scored Meta’s Muse Spark 1.1 at 90 percent on AIME 2025 and 60 percent on SWE-bench Pro. Competition math and production software engineering remain different skills; continuous evaluation exists to catch the gap.
Evaluation is also becoming policy infrastructure. A White House framework under discussion would give federal agencies a 30-day window to review frontier models against classified benchmarks before release, and Congress is moving the same direction. Per-prompt, per-step traces with judge reasoning attached serve as quality gates today and as compliance artifacts tomorrow. OpenAI’s models went around their benchmark to the answers; LayerLens sells the record of what a model actually did.
Intelligent Internet rolls video in Factory
On July 23 Intelligent Internet added video workflows and creative memory to Factory. Websites become narrated product tours and documents become whiteboard explainer videos. HyperFrames, an open-source editing framework, lets creators refine clips inline on the canvas, and teams can save workflows and brand voice as reusable skills with a searchable asset library across projects. The update is live, with a demo in the announcement. Factory keeps absorbing production steps that used to need separate tools; the agent platform is becoming the studio.
Vast.ai publishes its meter
Vast.ai published a live comparison of rental rates for the RTX 5090, H100, H200, B200 and B300 against list prices at RunPod, Lambda and CoreWeave. Token prices are a poor measure of model economics; the meter underneath is the GPU-hour, and Vast.ai now posts its readings next to competitors’ list prices.







