Extinction warnings fuel calls to restrict AI and the frontier labs agree to “pace” the frontier. Meanwhile, personal agents make access increasingly tempting, and mathematics is in existential-crisis mode, too.
Steven Waterhouse · Nazaré Ventures
Following AI Waves #023: Mutually Assured Dependence.
Over the weekend, Dario Amodei called for “pacing the frontier” of AI. Sam Altman endorsed the proposal and promised comparable access for outside evaluators, while Elon Musk publicly agreed.
The overlords appear to agree on something after all? After months of competing over capability and squabbling in public, the industry’s leaders are discussing how to constrain its development.
Dario’s post was published on the heels of Jacob Coxon’s announcement of his resignation from Anthropic on September 8, accusing both Anthropic and OpenAI of behaving irresponsibly. By the following day, the Wall Street Journal was covering his departure, and he appeared on Fox News’s Special Report.
It all reeks a bit of an organized effort to spread fear, uncertainty, and doubt (FUD for my crypto-native readers). It’s now being reported that Coxon will team up with Bernie Sanders and Steve Bannon for an event this coming Tuesday pushing for “human-controlled AI.”
People I know and respect are deeply concerned about AI swarms, and I share their concern. But I’m tired of arguing over whether this ends in 2027 or 2030. If something can change the outcome, that is worth discussing. An endless succession of extinction forecasts becomes engagement farming, however earnestly delivered.
Dario’s immediate commitment gives us something more concrete to examine, but other than his colleagues leading other frontier labs chiming in to agree, his proposal has been roundly panned as yet another attempt at regulatory capture.
Ben Thompson on pacing the frontier, Stratechery, September 2026
Anthropic proposes embedded external evaluators with access to internal work and the right to publish findings, subject to specified redactions. Industry-wide limits and international coordination would follow. He explicitly says pacing does not mean halting training or technical progress.
Anthropic also wants competitors held to the safety requirements it advocates. Dario explicitly argues that coordination would allow caution without sacrificing commercial advantage. If adopted, the proposals would also give Anthropic considerable influence over the conditions under which frontier development proceeds.
As I have argued previously, the frontier labs have a direct commercial interest in how the rules are written. A regime that restricts competing models or makes deployment contingent on a costly approval process would protect some of their pricing power. For Anthropic, competitive pressure ahead of a public listing could make that prospect particularly attractive. It is difficult to disassociate the substance of these warnings from the commercial and political interests surrounding them, including the possibility of regulatory capture.
Some go further, insinuating that “pacing” the frontier is just a way to excuse poor financial results and boost margins as they prepare for IPOs.
Perhaps the most important reaction, however, came from David Sacks, the White House AI czar. In a long tweet posted on Sunday, Sacks called the labs’ bluff, reading the requests for restraint as attempts to secure the leaders’ commercial position. He writes:
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead. (…)
But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. (…)
Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want. (…)
This saga is far from over, but we would do well to stop fomenting fear and get back to trying to educate folks on how better to understand this powerful new technology. Reductive takes on doom and regulation just don’t do anyone any good.
Muse and Instinct: What’s privacy got to do with it?
Meta’s Muse brings persistent memory and background execution into a personal assistant that can handle email, arrange travel, and make purchases. Its usefulness depends partly on how much of a person’s digital life they connect to it.
Instinct makes the appeal concrete. In The Atlantic’s test, the assistant selected a book for an editor based on his writing and arranged bookstore pickup within a $25 budget. The reporter used a temporary phone number and a disposable virtual credit card, then deleted the account afterward. Useful delegation and caution about access can coexist, although few people will maintain that discipline once an assistant becomes part of their routine.
As I wrote last July about AI as a relational technology, personal context improves a service, and the better service encourages people to share more. The accumulated understanding becomes part of what they value. Starting again with another system means rebuilding some of that relationship.
Personal agents extend this process into permission to act. An assistant that resolves an irritating errand gives its user a practical reason to connect another account. Each successful task can make broader access feel reasonable, while the consequences of accumulated access remain harder to assess. People can understand the privacy tradeoff and still decide the help is worth the exposure.
TechCrunch documented early testers’ concerns about unsolicited actions and data retention, including a deletion problem the company subsequently addressed. Instinct’s current terms explicitly distinguish disconnecting a service from deleting its indexed data. The company may continue using that data unless the user separately requests deletion. Revoking access stops one part of the relationship without necessarily removing what the assistant has already learned.
Meta has engineered substantial controls around Muse’s access. Its security architecture isolates the agent, withholds real credentials, and relies on a separate system called Sentinel to authorize connector actions and outgoing network requests. These protections assume the agent may be compromised and constrain what it can do afterward.
Meta’s own access follows different rules. Staff access is restricted through operational policies. Sanitized conversations and tool activity are used for training by default, with an opt-out. Meta excludes conversations and VM data from advertising systems, though activity on outside services may indirectly affect targeting. A Confidential VM intended to prevent Meta from accessing data inside it remains in testing.
Remembering a user’s preferences helps their assistant complete a task. Using those interactions to improve the provider’s general model serves a broader purpose. The immediate usefulness of the first does not settle the terms of the second.
The Information reports capacity constraints, a possible $1 billion raise, and founder Noah Shinn’s preference against charging users. That leaves a commercial question behind the enthusiasm. If advertising or transaction commissions eventually finance these assistants, the provider would have an incentive to influence where delegated spending goes. An agent choosing a hotel or making a purchase could act on that incentive without presenting anything recognizable as an advertisement. This is a possible business model, not an announced Instinct plan.
Portability would give users more room to respond if those terms change. Muse already lets them inspect, edit, and download memory and files, a concrete improvement on the problem I described last year. Whether another assistant can use that context and reconstruct the connected workflows will determine how much freedom the export actually provides. The more useful these relationships become, the more weight falls on their terms of departure.
Why would you ruin your career?
OpenAI’s proposed Navier-Stokes solution has become a dispute about credit, private data, and what mathematical progress is supposed to accomplish.
On September 8, OpenAI published a 166-page proposed solution and a Lean formalization to one of the seven Millennium Prize Problems. Its argument constructs a smooth external force that drives a three-dimensional fluid from rest to infinite velocity in finite time. The official Clay formulation explicitly permits this route.
OpenAI says the project used around 10,000 agents over 88 hours, followed by another 17 hours of formalization using GPT-6 Astra. It began after the company heard rumors of progress by outside researchers.
NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge had been pursuing related work, building on ideas from Diego Córdoba and Luis Martínez-Zoroa. They were preparing their results for publication when OpenAI’s effort overtook them.
In his public statement, Buckmaster says OpenAI researcher Sébastien Bubeck proposed either closely timed announcements or an arrangement in which Buckmaster alone would present OpenAI’s proof. He alleges that Bubeck pushed to exclude Alpöge because of his Anthropic employment.
When Buckmaster threatened to describe the negotiations publicly, he says Bubeck replied:
“Why would you ruin your career?”
Bubeck disputes Buckmaster’s account, saying he did not seek to deny Alpöge credit for work Alpöge had performed. His objection concerned an Anthropic employee co-authoring a paper presenting a proof produced inside OpenAI. The disagreement remains a contested account of how publication and authorship were negotiated.
Buckmaster also questioned whether the unpublished work he had shared with Codex had contributed to OpenAI’s result. The company’s response has changed since the initial announcement.
In a September 10 update, OpenAI says its investigation established that Buckmaster’s Codex prompts during the preceding two months could not have influenced the system, including through training. That is OpenAI’s finding, rather than an independent audit, but it supersedes the uncertainty in its original statement. The dispute over credit and publication conduct remains separate.
The broader question about private research still deserves clear contractual and technical answers. An unpublished mathematical argument retains its value after its author’s name has been removed. Researchers need to understand how their work can be used before entrusting it to a service. This case should no longer be presented as established evidence that customer research contributed to OpenAI’s proof.
The mathematical community’s response has also widened. A declaration signed by 25 Fields medalists argues that solving famous problems has historically served as a proxy for developing mathematical understanding. The work continues through explanation, simplification, discussion, and teaching until other mathematicians can use the new ideas.
Four days before OpenAI’s announcement, Anthropic had shown that Claude could formalize Fermat’s Last Theorem in 11 days. Mathematical production and verification are accelerating together. The human work of understanding what those results make possible runs on a different schedule.
The declaration’s objection is more substantial than resentment at being beaten to a theorem. If AI makes results abundant while understanding them remains demanding, the profession needs ways to recognize and support that work. Otherwise, the activity rewarded most visibly can accelerate while the processes that make its output useful struggle to keep up.
The Clay Mathematics Institute welcomed the apparent settlement of Navier-Stokes while emphasizing the human understanding that should follow. Its process for evaluating the achievement and assigning credit will remain deliberately unhurried. A proof can be checked before a community has established what it has learned.
Quick hits
Mistral’s industrial backing
Mistral raised €3 billion in a Series D led by Samsung Electronics at a post-money valuation above €21 billion.
Samsung adds a major industrial partner to Europe’s effort to develop frontier models and infrastructure outside the American labs. Open weights are becoming part of industrial policy as well as a distribution strategy. The commercial test is whether that backing produces services European businesses choose for their capabilities, alongside their preference for an alternative supplier.
The cheaper model absorbs the premium tier
DeepSeek says V4.1 Flash outperforms V4 Pro and will temporarily serve Pro API requests at Flash prices. The company’s cheaper tier has absorbed the more capable one, making its previous product hierarchy difficult to sustain.
OpenRouter’s analysis of the GPT-5.6 Luna promotion supplies a demand-side counterpart. Daily usage increased 13.8 times, its traffic share rose from 0.7 percent to 7.8 percent, and 32 percent of customers continued using it after the promotion.
Once performance clears a useful threshold, price changes both which model developers select and which tasks become economical to automate. A benchmark gain and a price cut can create very different kinds of demand.
Voice becomes a separately priced component
GPT-Live-1 provides full-duplex speech for $0.05 per minute, including listening while speaking and handling interruptions. Deeper reasoning and tool use pass to a separately selected backend, whose costs are additional.
Developers can improve the conversational interface without replacing the system that performs the work. The product is increasingly assembled from components with different performance requirements and economics.
Astra’s benchmark result depends heavily on its operating setup
ARC Prize verified Astra at 62.7 percent using its standard interface and 99.9 percent with OpenAI’s provider adapter. The best runs used different reasoning settings. Even at the same maximum setting, the scores were 62.7 percent and 98.6 percent.
The adapter preserves private reasoning between requests and compacts longer interactions. The standard interface leaves the model responsible for choosing what to carry forward in visible notes. A model name alone tells a buyer surprisingly little about the capability of the system they will actually deploy.
Nvidia buys distribution while its earlier deal draws scrutiny
Nvidia has signed an agreement to acquire Hugging Face for $12.93 billion, with closing expected in the first half of 2027, subject to regulatory approval. It says the platform will continue supporting competing hardware, clouds, and inference providers without requiring Nvidia compute.
That promise gives the open-model ecosystem a concrete standard against which to assess the acquisition.
Meanwhile, the Justice Department is investigating Nvidia’s licensing arrangement with Groq. The agreement transferred access to technology and brought senior people into Nvidia while Groq remained independent. The inquiry concerns whether that structure avoided scrutiny an acquisition would otherwise receive, rather than establishing that wrongdoing occurred.
Image generation becomes an ongoing production process
ChatGPT Images 2.5 adds local editing, comments placed directly on images, templates, and a Sketch feature. OpenAI also reports better consistency across revisions.
The practical improvement is continuity. A usable image can remain the basis of further work through annotation and revision. That makes generation easier to incorporate into an actual production process, where preserving previous decisions matters as much as producing the first impressive result.
DeepMind precomputes nine billion DNA changes
Google DeepMind’s AlphaGenome Atlas contains predictions for all nine billion possible single-letter changes in the human genome. The one-petabyte collection includes scores for ranking predicted effects and explanations of the molecular features involved.
It is free for academic research and has not been validated or approved for clinical use. Precomputing the prediction space turns a model researchers query individually into shared scientific infrastructure. Researchers still have to establish which predictions survive contact with biology.
Portfolio updates
LayerLens: the score went up and the product got worse
LayerLens sells independent testing for AI models and agents. On September 9 it added run-to-run comparison to Stratix, its testing platform. Two test runs can now be read question by question instead of as one score each.
A single number hides more than it reports. Two runs of the same 198-question test scored 86.9 percent and 88.9 percent, which looks like a small improvement. Underneath, 13 questions went from wrong to right and 9 went from right to wrong. Twenty-two answers flipped and the summary reported two points of progress.
A follow-up on September 10 explains a second problem. Many of these tests are marked by a second AI model, called a judge. Providers update the model behind a judge without changing its name. LayerLens put a number on that on September 11. One judge agreed with human marking 87 percent of the time in one week and 79 percent the next. The test did not change and the prompt did not change. Fifteen labeled examples, re-marked before each comparison, catch the drift without extra model calls.
Dario Amodei wants pacing enforced by evaluators placed inside the labs. Those evaluators will be reading scores. A score can rise while the work behind it gets worse, and the marker producing that score can change without notice.
Memco: the lesson the customer keeps
On September 11 Memco opened the loop behind its Fenmoor benchmark as an SDK, so a company can run continual learning inside its own product instead of only inside a coding agent. Scope boundaries run at the personal, team and organization level, separating a user’s preferences and working context from the information a team’s agents share.
The company’s approach centers on memory that carries across models, tools, and teams. Every lesson carries provenance naming the run, the reviewer and the scope, and passes a human approval gate before reuse. Retirement is a first-class action, so a lesson that stops being true can be withdrawn.
Memco states on its Spark product page that it never trains on customer memory. Muse uses sanitized private interactions for training by default. Same data, opposite defaults.
Fairmath: the bottleneck gets a GPU
On September 14 Fairmath opened two challenges on its FHERMA platform with NVIDIA’s HPC developer group, adding a GPU acceleration layer to fully homomorphic encryption through NVIDIA’s cuPQC library. The first asks for high-performance polynomial multiplication on the cuPQC bigint backend, and is aimed at CUDA programmers rather than only at cryptographers. The second asks for GPU-native key switching in CKKS, one of the heaviest costs in the scheme. Winning work is meant to land in open-source infrastructure that other projects reuse.
Meta answers the question about private data with operational policy and a confidential virtual machine still in testing. Encrypted computation would make the policy unnecessary. Cost has kept it out of production, and the cost is arithmetic.
Prime Intellect: the memory runs out before the reasoning does
Prime Intellect sells the stack companies use to train and run their own agents. With vLLM and Red Hat AI it shipped Hybrid HiSparse for GLM 5.3. A model holds the conversation so far in the fastest memory on the chip. Long agent sessions outgrow that space and the request stalls. HiSparse keeps the request running. On one node of eight H200s at a one-million-token context, a thirteen-turn agent workload held 19 to 25 requests at once, against 5 to 6 before. The work is planned for vLLM version 0.30.
Agent sessions are long by construction, and every turn adds to the store. Serving three to four times as many of them on the same hardware cuts the cost of running agents, which is the business Prime Intellect is in. Prime Agent passed 20,000 GitHub stars on September 8.
Dimensional: what shipped and what did not
Dimensional builds DimOS, an operating system for robots. On September 10 it added support for Habitat, a library of laser-scanned real buildings. A customer can now drive a simulated robot through a scanned real building and watch what its navigation software does, without owning the robot or the building.
Testing robot software usually needs a robot and a room, and both are expensive. A scanned building can be run again at no cost.











