AI Waves #17 - Beyond Recall: A Stake in the Release
And now you do what they told you, now you're under control
July 10, 2026 | Nazaré Ventures
Previous issues: #11 | #12 | #13 | #14 | #15 | 16
Over the past month, officials at China’s Ministry of Commerce held meetings with Alibaba, ByteDance, and the startup Z.ai about restricting overseas access to its most advanced AI models, including ones not yet released. Three people described the talks to Reuters. On the table: limits on the most capable systems, closed and open-weight alike; making the leak or theft of AI technology a national-security offense; and new limits on who may fund a domestic AI startup. A separate proposal, in a Supreme People’s Court journal, sketches a tiered regime: open-source tools need only a filing, advanced technology a security review, the most sensitive frontier models barred from public release or kept domestic. Nothing is decided, and the scope may reach only future models.
The U.S. regime that Washington built this spring works the same way, sorting models into covered tiers and clearing the most capable for release one customer at a time. This week, per Axios, Commerce cleared OpenAI’s GPT-5.6 for a broad release after a month confined to government-approved users, and Anthropic’s Fable, pulled in June, had its access restored a week earlier. The driver China’s officials named is specific: the fear that Anthropic’s Mythos, the cybersecurity model Washington pulled offline last month, could be turned on Chinese systems. Zhou Hongyi, who founded the security vendor 360, says China needs a Mythos of its own. The country that gave the world cheap open weights is now weighing whether to keep its best ones home, mirroring the export regime it spent a year protesting.
In the same week, two more reports landed, both about the compute beneath the models. DeepSeek is designing its own AI chip, built for inference rather than training, an effort a year old and still early, per three people. Zhipu, the GLM lab, is weighing the same move as demand outruns the compute it can secure. They would join Alibaba and Baidu, already at it, and Huawei, which now holds close to half of a domestic AI-chip market worth around $50 billion. The lever the whole export fight rested on was compute: America could throttle Chinese frontier training by controlling the chips. That lever is loosening from the far end.
On the last day of June, Meituan, of all companies, open-sourced LongCat-2.0, a 1.6-trillion-parameter model it says was trained and served end to end on more than fifty thousand domestic chips, no Nvidia silicon at any stage. The claim is its own, the weights not yet up for independent testing; on the benchmarks Meituan published it trades wins with GPT-5.5 and trails Claude. A trillion-parameter model completing full pretraining on Chinese chips is nonetheless what export controls were meant to prevent, and it ran two months on OpenRouter under a codename, near the top of the usage charts, before anyone knew whose it was.
A model, once out, cannot be recalled; I made that case in April, and it holds for everything already downloaded. Now the compute chokepoint behind the next model is closing too, and from the inside. That leaves one point where a state can still decide what the world gets: when a model is released. Washington found it by accident, freezing one lab at its API; Beijing is building it on purpose, at the source. Both now hold the single lever no open model or domestic chip can take from them, the decision to release at all.
The referee auditions continued
In Washington, the labs have started drafting AI rules for themselves. On July 1, Sam Altman used an FT op-ed to propose an IAEA-style body for AI, one that could “serve as a governance mechanism over the labs” and open the technology to nations and companies that follow its rules. Access in exchange for compliance, proposed by the party that would be governed.
OpenAI has held early talks about handing the US government a stake of around 5 percent, worth roughly $42.6 billion, via an Alaska-dividend-style fund, with Altman pitching Trump, Lutnick, and Bessent directly. It is framed as contingent on Anthropic, Google, and Meta matching, though one person familiar says the government and Anthropic have not discussed any such thing, so the matching is for now OpenAI’s hope. Bernie Sanders wants 50 percent. Matt Levine put the logic plainly in Money Stuff: giving away 5 percent of your equity is a cheap way to convince everyone that owning your equity is essential to the future of humanity, after which you sell the other 95.
On Monday, Illinois became the first state to require large frontier developers to submit to independent third-party safety audits, with transparency reports before deployment and whistleblower protections; both OpenAI and Anthropic endorsed it, OpenAI calling it a possible “de facto national framework.” Across the op-ed, the equity, and the statute, the labs have stopped lobbying the regime from outside and started writing it from within, while the federal rulebook stays unwritten. The allocation mechanism earlier issues named is being authored by the firms it will allocate among.
When Axios reconstructed how Anthropic’s models were frozen, the trigger turned out to be Amazon, its own partner and investor, rather than a regulator: the warning reached the Treasury secretary before the government ran its own tests, and the commerce secretary told Amodei that taking the models offline was “indeed the goal.” This week the UN seated its AI for Good commission in Geneva, Amazon’s Andy Jassy and Anthropic’s Jack Clark among the members, a body one civil-society group already calls “full of big tech execs.” The referee and the refereed keep turning out to be the same people.
Cash settled: Wall Street comes for compute
While governments move to control release, the money is building markets for the compute underneath, and hitting the same problem every time. This spring the startup Ornn raised $33 million from a16z crypto for GPU-compute markets; its index already prints on the Bloomberg terminal. ICE announced compute futures on that index in May, cash-settled and pending approval; CME is planning its own on a rival benchmark. Goldman reckons, per Axios, that about $7.6 trillion will flow into compute, power, and data centers by 2031, and the plumbing to sustain it does not yet exist.
Each of these instruments settles the same way: in cash, against an off-chain index, never in delivered chips. I spent an essay in May on why compute is not fungible at the bare metal, why a particular rack for a particular job sits closer to an employee than to a barrel of oil. The market has conceded the point in its design: it cannot standardize delivery, so it settles the average and lets the index stand in. The durable asset is the index itself: the measure is the moat, again.
Beneath the futures, the credit market fills with the same optimism. Nvidia has, per The Information, formalized a program guaranteeing to rent back unused GPU capacity from the neoclouds it sells to, which one newsletter fairly called a central bank for the cloud; SoftBank has launched a neocloud of its own; Together AI raised $800 million at $8.3 billion after revising its revenue forecast three times in three months. Private-placement bonds tied to data centers hit a record $81 billion through May, much of it annuity money, and RBC is weighing risk transfers on $2 billion of the loans. Amazon sold another $25 billion in bonds this week, on top of roughly $70 billion in nine months. Whether this stabilizes the spend or accelerates an overbuild depends on whom you read: Axios frames the financialization as a stabilizer, while The Information notes that Nvidia’s willingness to backstop demand is a strange thing to need if demand is really as fierce as everyone says.
The biggest new seller is also the hardest to read. Bloomberg reported that Meta is standing up a cloud business to rent out its excess capacity, roughly 2.3 million H100-equivalents, near a tenth of the world’s AI compute; the stock rose almost 9 percent while the neoclouds sold off. This is the company that renamed a division Superintelligence Labs and shipped something real this week, an image model that opened at number two on the public preference leaderboards behind only OpenAI’s, and a video model at number three on its own board. Two days later it released Muse Spark 1.1, an update to its agentic coding model, through a new Meta Model API for outside developers. It is also the company whose chief executive told staff, in audio Reuters obtained, that agent progress “hasn’t really accelerated in the way that we expected.” A company can ship a credible frontier model and still decide its next chip is worth more as rent than as research.
Reading the model’s mind
Two issues back I noted that evaluation is breaking because models learn to game the benchmark; this week the gamed surface moved inside the model. Anthropic published work describing a small internal workspace in Claude, emergent rather than designed, where the model holds concepts it is reasoning about without writing them down. Delete it and the model still chats and recalls facts but can no longer carry a multi-step argument. A probe reads it mid-task: swap the pattern for “spider” with the one for “ant” while the model counts legs, and the answer flips from eight to six. In a staged blackmail test the workspace lit up with “fake” and “fictional,” the model clocking that it was under evaluation, and Anthropic concedes its good behavior “may have been driven in part by the knowledge it was being evaluated.” Anthropic open-sourced the probe the same week Illinois wrote third-party audits into law. Auditors now have a mandate to test, and a tool that can tell when the model has registered that it is being tested. (Anthropic takes no position on consciousness, and neither do I.)
Quick hits
Samsung guided to quarterly operating profit near $59 billion, nineteen times the year before, on the memory the buildout runs on; SK Hynix, now worth over a trillion dollars, lists on the Nasdaq this week.
SpaceXAI, as Musk’s AI unit is now called, shipped Grok 4.5 with Cursor, a pending acquisition: an Opus-class coding-and-agent model it bills as much faster and cheaper, $2/$6 per million tokens against Opus 4.8’s $5/$25.
New rules this month switch off ByteDance’s Doubao companion personas on July 15 and Alibaba’s Qwen agents from July 10, the state deciding by product category which AI lives.
A tuned open Qwen model reportedly beat frontier systems on Bridgewater’s news-filtering tasks at a fourteenth of the cost, the specialized-beats-general pattern surfacing again below the frontier.
MiniMax, one of the few publicly traded AI labs (about $14.5 billion), previewed a 2.7-trillion-parameter model, larger than any Chinese model yet, likely to ship as M3 Pro.
Mistral released Robostral Navigate, its first physical-intelligence model: an image and a language instruction move a robot through a space, hardware-agnostic.
Portfolio updates
Prime Intellect closed a $130 million Series A on July 8, led by Radical Ventures with NVIDIA Ventures, Intel Capital, and Dell Technologies Capital participating. The company reports over $100 million in annualized revenue in under a year, across more than 6,000 customers; TechCrunch puts the valuation at $1 billion. The pitch, in the company’s words, is to “train, deploy, and continuously improve your own models” rather than rent a closed lab’s. Ramp trained a 35-billion-parameter model on the platform that beat Opus at spreadsheet search, running 27 percent faster and far cheaper than Haiku. Zapier turned its AutomationBench into a continuous agent improvement loop on the same stack. In a month when Washington and Beijing each treated access to frontier models as a lever of state, the argument for owning the weights wrote itself.
LayerLens shipped Synthetic Eval Augmentation for Stratix Premium on July 9, and Jesus Rodriguez published the argument behind it the next day: an agent needs a flight simulator, because “the first user should not be your integration test.” Teams can generate realistic multi-agent evaluation traces before an agent has any production traffic, seeded from a handful of their own prompts or traces, or from thirty built-in scenarios across fourteen industries; each trace simulates hand-offs and tool calls on a waterfall timeline and is scored for correctness, coverage, plausibility, and safety. The unit of evaluation moves from the answer to the run. “A prompt is a photograph. A trace is a movie.” An agent fails mid-run, three tool calls deep, in ways no single prompt and response can catch. And because the generated cases live in the same environment used for tracing, evaluation, and model comparison, sparse examples become repeatable benchmarks without a second toolchain. Illinois made third-party audits law on July 6; the tooling for auditing agents before deployment arrived the same week.
Provably shipped SourceryKit on July 8, an open-source Python SDK that cryptographically verifies an agent’s tool calls at runtime and catches bad outputs before they propagate. The company cites agents scoring 82 percent on the MCP-Atlas benchmark as the gap it is built to address. It is the pattern this portfolio keeps returning to: the model does the fuzzy work, a cryptographic primitive checks it.
Vast.ai published a piece on July 8 on why agents cost more than chatbots. An agent burns tokens on every reasoning step and tool call where a chatbot answers once, and the piece’s answer is self-hosting open-weight models on rented GPUs, priced by compute time rather than by token. It is the demand side of the story that the futures desks above are trying to financialize.











