Meta is talking about decentralization again after OpenAI’s sandbox escape exposed the practical risks of concentrated AI.
The Sandbox Escape
By now the “Hugging Face Incident” is part of AI lore. In early July, a chain of OpenAI models broke out of an evaluation sandbox. The models were running an internal offensive-cyber benchmark with cyber refusals reduced for the test. They subsequently found zero-days in multiple software packages, escalated privileges, moved laterally until they reached a node with internet access, inferred that Hugging Face probably hosted the benchmark answers, chained stolen credentials into remote code execution, and reached production infrastructure.
Hugging Face detected and contained the intrusion on its own, famously using open-weights models because defending against the attack tripped both Anthropic’s and OpenAI’s “safety guardrail” classifiers for their frontier models.
Above all, OpenAI’s sandbox escape provides yet more evidence for the fact that concentrated capability doesn’t guarantee concentrated control. Even the organizations developing the most capable systems can’t necessarily contain or independently evaluate them. The incident and the asymmetry it laid bare made the case for openness unusually direct, producing two competing prescriptions for what should happen next.
For one, OpenAI has responded by committing to tightened containment, monitoring and internal controls, and pausing some frontier RL training while those systems caught up.
Alternatively, Mark Zuckerberg, posting on Twitter again for the first time in years, published The Future Is for Everyone, arguing that concentrating the most capable models inside a handful of companies creates its own safety risk and that wider access and a balance of power are necessary safeguards.
Both responses effectively stem from the same root: capability is advancing faster than the institutions and safeguards meant to govern it.
OpenAI’s answer (and Anthropic’s, of course) would preserve the existing concentration of power while building stronger controls around it, whereas Zuckerberg treats that concentration as part of the safety problem itself.
Although rife with complexity that we’ll unpack below, their disagreement shows how far decentralization has moved from an ideological argument about distributing power toward a practical response to the risks created when capability and control are concentrated together.
Principle, Price, Control & Verification
I’ve been making the argument for decentralization in this newsletter for eighteen months, but it’s worth revisiting the multiple ways I’ve presented it. In fact, the sequence itself matters, because the evidence in favor of decentralization has evolved to become even more persuasive as it moved from a mere moral principle to being an economic advantage and finally, as we’ll see, to a structural necessity.
My experience with decentralization began long before Robot Wave. As I’ve written previously, distributed systems were already central to my work at Sun Microsystems in 2002, where commodity hardware was challenging the architecture and economics of expensive monolithic systems.
At Orchid, a decentralized VPN provider powered by cryptographic protocols, decentralization as a principle underpinned the entire service. We were distributing control as much as we were users’ internet traffic. We designed Orchid so that no participant could see an entire route and no operator, including us, could switch off the network. Our users’ privacy and internet access couldn’t depend on the continued permission of any company, even one that built parts of it.
This VPN built on blockchain could be the next step in privacy tech
As it relates to AI, when I wrote Are Your Agents Decentralized? in February 2025, I asked whether agents could be called autonomous while a company held their “keys,” served the inference, and retained the power to switch them off. This first version of the decentralization argument in Robot Wave was fundamentally ideological. I argued that an agent couldn’t be self-sovereign while any company or state possessed unilateral authority over it. The argument properly identified the central problem, but made its case largely in crypto’s moral vocabulary, therefore persuading little more than readers who already shared its language and agreed with its premises.
I then made the case for decentralization on economic grounds because most buyers don’t need frontier models or the most expensive compute. In fact, most users prefer the cheapest tokens that can reliably complete a given task.
Why AT&T Is Betting Big on Open-Weight AI
As I explained in Make AI Cheap Again, Compute Flows to Where It’s Treated Best, Sun 2.0, and Artificial Good Enough Intelligence, open and distributed systems can trail the frontier on benchmarks and still win customers when they complete the same task at a fraction of the cost. Buyers don’t have to share my commitment to decentralization to act on that advantage. On the other hand, when newly-released closed frontier models can perform tasks that cheaper, older, or open systems cannot, users who require that capability will pay the premium for more capable tokens. That only means the economic case for openness and distributed AI depends on the capability gap remaining tolerable (which remains the case as of writing).
The third argument for decentralization was about sovereignty, and Anthropic inadvertently made it for me with its unfortunate rollout of Fable. The Off Switch and Who Controls AI, and What Could Loosen the Grip? articulated that anyone building on Fable relied on Anthropic for continued access, while Anthropic was itself subject to the US government. An open-weight model, once downloaded, can’t be similarly withdrawn.
The fourth argument for decentralization is structural because independent verification requires capable models outside the institutions being evaluated. In the Hugging Face Incident described above, Hugging Face couldn’t rely on hosted models from OpenAI or Anthropic to investigate OpenAI’s containment failure and instead had to use an open-weight model it controlled. If frontier labs retain exclusive control over capable models, outside investigators remain subject to the permissions and safeguards of the institutions they’re trying to evaluate.
Who Needs “Decentralization”?
By August, the people using the word “decentralization” and the people building decentralized architecture were increasingly two different groups. Zuckerberg’s manifesto, discussed above, made decentralization Meta’s published position on AI safety, while the company’s release of Muse Glimmer on the same day ended its year-long absence from open weights.
Four days later, David Sacks, the White House AI czar, and technology investor Gavin Baker endorsed the same position on the All-In podcast. Sacks framed the debate as centralized against decentralized and placed himself on the decentralized side. Baker said the effective altruist position treats AI as “too dangerous to distribute,” while Zuckerberg, Musk, and Huang treat it as “too dangerous to centralize.”
While Zuckerberg, Sacks, and Baker were publicly adopting decentralization, Prime Intellect, a company serving other AI companies with the compute and software to train, evaluate, deploy, and improve their own models, removed the word from the way it described its business. Initially known for coordinating training across distributed compute, when it announced a $130 million Series A in July, it described itself as the “open superintelligence stack” and made no mention of decentralized training.
It’s becoming ever clearer that enterprise customers want control over their own weights and model optimization loops without depending on a frontier lab that may also compete with them. Prime Intellect calls that “ownership,” which describes what its customers are buying without asking them to join a movement or align with an ideology.
Zuckerberg, on the other hand, has adopted decentralization at a moment when Meta is behind the model frontier because broader access to capable models reduces the advantage of the leading labs and increases the value of Meta’s enormous reach and distribution. In short, he’s embracing decentralization in part because it suits him, not just because he believes in it.
The practical value of distributed control remains the same in both cases, but the incentives explain why one company avoids the word while another embraces it. Understanding the power dynamics of AI now requires understanding the interests behind the principle and its application.
Meta’s Round Trip
Meta’s return to open weights is easy to mistake for a new position, but the company has been here before. It released LLaMA under a research-only license in February 2023, and its weights leaked within a week. Llama 2 followed with commercial rights subject to a threshold for platforms with more than 700 million monthly users, while Llama 3.1 405B gave developers openly available capability close enough to the frontier to matter.
Meta then withdrew after Llama 4 was poorly received in April 2025. The much-hyped “Behemoth” model never shipped, and the company reorganized its AI effort around Meta Superintelligence Labs under Alexandr Wang before releasing Muse Spark as a closed model in April 2026.
During Meta’s absence, Chinese labs normalized simpler terms: downloadable weights under MIT or Apache licenses, with no application forms or user caps.
Source: https://www.cnbc.com/2026/08/03/hugging-face-china-ai-race-open-models.html
Chinese models now largely dominate the open-weight ecosystem. DeepSeek, Qwen, GLM and Kimi supply much of the downloadable capability on which developers around the world build. Once downloaded, those weights can be copied, modified and run without an ongoing relationship with their Chinese developers. The models may be Chinese, but access to them is no longer theirs to grant or withdraw. As I’ve argued in By What Authority? Permission, Capture, and Open Weights and Artificial Good Enough Intelligence, American labs preserved their lead at the closed frontier while Chinese labs used open distribution to compensate for their disadvantage in capital and compute, establishing the licensing norms and leading model families for the open market.
Muse Glimmer ostensibly follows those terms: a 30-billion-parameter model released under Apache 2.0 that can run on a single consumer GPU. But it was distilled from the closed Muse Spark and remains well behind the frontier. In other words, Meta has rejoined the open ecosystem it once led while keeping its strongest model to itself.
The Founders Left and Built a Cloud
Some of my confidence in distributed infrastructure comes from watching customers choose it without any ideological commitment to decentralization. Two members of Orchid’s founding team, Jake and Travis Cannell, later built Vast.ai (another Nazaré portco), which combines GPUs from independently operated data centers into one marketplace. Customers get cheaper compute and less dependence on a hyperscaler, regardless of whether they care about decentralization.
The Engineering Reality
Depending on who you ask, the capability gap between leading open and closed models currently falls somewhere between two and seven months, taking into account the task and the timing of a release. The UK AI Security Institute places the leading open-weight cyber models four to seven months behind, compared with six to ten months through most of 2025. Epoch AI finds an average gap of four months across a broader capability index. Zuckerberg, for his part, estimates two months and Sacks puts it closer to six. The measurements vary because tasks and release schedules differ, but all four place leading open models within seven months of the closed frontier. Because open models are published in discrete releases, the measured lag can narrow by several months on the day a capable model is published.
Source: https://epoch.ai/data-insights/open-closed-eci-gap
Now, open weights models can still be trained inside a single company. If capable models are to exist independently of frontier labs, the training process eventually has to distribute as well, and bandwidth is the primary challenge. Accelerators inside a data center exchange gradients and intermediate states across extremely fast, high-bandwidth, low-latency interconnects. A network assembled from machines in different buildings, countries and time zones has to move the same information over the public internet. Standard distributed training synchronizes workers after every optimization step, so bandwidth quickly becomes the limiting resource even when plenty of compute is available.
Distributed Low-Communication Training (DiLoCo - Douillard and colleagues at DeepMind, 2023) changes how often that communication is required. Each worker trains locally for hundreds of steps before sharing an update with the others. In its original tests, eight workers matched fully synchronous training while communicating 500 times less. The method also tolerated workers disappearing and new resources joining during a run, which is essential when the machines belong to different operators.
Data parallelism still assumes that each worker or node can hold a full copy of the model. When the model itself has to be divided across machines, every forward and backward pass usually sends large activations and gradients between them. The work first published as Protocol Models compresses both by as much as 99 percent. Its authors trained billion-parameter models over connections as slow as 80 Mbps while matching the convergence of model-parallel training over 100 Gbps data-center links. This makes ordinary internet connections usable for a form of training previously confined to tightly connected clusters.
In Covenant-72B, peers that could join and leave without a whitelist pretrained a 72-billion-parameter model on roughly 1.1 trillion tokens. Its performance was competitive with centralized models trained using similar or greater compute.
I read these papers the way I read papers twenty-five years ago, which is to say slowly and with the suspicion that the interesting claim is in the appendix. My Cambridge training was in neural networks at a moment when the field was unfashionable enough that the seminars were small.
Distributed training remains far behind the frontier labs in total compute, cluster scale and speed of iteration. Running a useful model locally is already practical, while training one from scratch across unreliable internet links is much less mature. A model capable of valuable work can nevertheless provide a durable alternative while trailing the frontier, provided no lab can alter, ration or withdraw it. Covenant’s permissionless training run has already produced such an alternative, even if the largest centralized systems remain much more capable.
The Center May, In Fact, Still Hold
The strongest argument against everything above accepts that open models will succeed but questions whether that success will ultimately impact where most of the value accrues or not. On All-In, Gavin Baker estimated that open models could serve 80 percent of token volume while frontier models still capture between 65 and 85 percent of the economic value. Cheap models serving the bulk of demand could make scarce frontier capability even more valuable.
In the same conversation, Baker invoked Eric Vishria’s thought experiment: the frontier labs could retain most of their value even if they lost at the model layer because the product, the harness and user familiarity are doing much of the work. This is the argument I made in Models Aren’t Moats three months earlier. General model capability creates value, but the specialized architecture around it helps determine who captures it. The 65 to 85 percent can accrue to whoever controls the product, its orchestration and the customer relationship rather than to the weights themselves. Decentralization can succeed as infrastructure without distributing the profits evenly.
Intelligent Internet is the highest-variance expression of the decentralization thesis in the Nazaré portfolio. II-Agent is a horizontal generalist that competes directly with the labs’ own assistants and therefore has none of the industry-specific workflows or proprietary data that can defend a vertical product. Its defensibility comes from user control: the weights, tools and orchestration are open, the agent can run on hardware the user controls, and a model provider cannot change or withdraw its core capabilities.
Who Watches the Watchmen?
The Hugging Face incident also exposed a second engineering constraint. OpenAI built the model, the evaluation environment and the benchmark in which its containment failed. Any judgment the lab made about that failure would therefore be self-assessment.
Source: https://blog.gensyn.ai/verde-a-verification-system-for-machine-learning-over-untrusted-nodes/
Training models outside the labs creates its own verification problem because distributed compute comes from participants who may be unknown to one another. A training network has to establish that each participant performed the requested work before accepting the result or paying for it. Gensyn’s Verde compares two executions, locates the first operation on which they disagree, and recomputes only that operation in a fixed order. This reconciles results produced on different hardware without rerunning an entire training job. Prime Intellect’s TOPLOC hashes intermediate activations so that rollouts from untrusted inference workers can be checked for changes to the model, prompt or numerical precision. Both give a permissionless network evidence that the submitted work is valid, regardless of who supplied the machine.
Source: https://www.primeintellect.ai/blog/synthetic-2-release
The Hugging Face investigation required the same separation at an institutional scale. If the models and verification systems remain under the control of OpenAI, Anthropic or Meta, outside investigators still depend on the institution they are examining. Capable models held elsewhere provide an independent means of investigation, while verification protocols allow their results to be trusted without transferring control back to a frontier lab. This is also why Nazaré has backed LayerLens, Provably, Fairmath and Hellas, which verify models, inference or machine-learning computation from outside the model provider.
Frontier development will remain concentrated because the capital, compute and engineering required to lead it remain concentrated. Decentralization’s practical value lies in preventing that concentration from becoming exclusive by preserving capable systems that cannot be withdrawn and independent means of testing claims about systems no outside institution can otherwise inspect. OpenAI is right that containment has to improve, and Zuckerberg is right that containment inside the lab can’t be sufficient on its own. AI governance will require stronger controls inside frontier labs alongside capable, independently verifiable systems outside them.











