We Have Liftoff?
Dario Amodei’s September 12 call to “pace the frontier” brought Sam Altman and Elon Musk into the unusual position of publicly agreeing about how the AI industry should conduct itself. The companies building increasingly powerful intelligence were concerned that intelligence was becoming too powerful, too quickly.
Dario proposed independent evaluators, coordinated safety standards and government assistance with the antitrust complications of competitors agreeing on their pace of development. Coordination, he explained, would allow the labs to proceed “without sacrificing commercial advantage.” In theory, a form of collective sacrifice that preserves one’s commercial position should be reasonably easy to sell. To keep democracies far enough ahead to afford that pace, he also called for tighter chip export controls and a crackdown on unauthorized distillation.
David Sacks was less impressed. He argued that the companies could slow themselves down and questioned whether the requested regulatory arrangements would protect their businesses and shield them from product liability. Dario’s essay doesn’t explicitly request that exemption, and I believe his concerns about AI are sincere, but companies can sincerely fear the consequences of their products while preferring arrangements that give them more control over their competitors and preserve their position in the market.
The broader pressure to act is real regardless of what anyone makes of the people proposing to decide what to do about AI on our behalf. In Of Models and Men, I compared the moment to March 11, 2020, when COVID abruptly became something the world needed to confront.
Last year, I called the point at which society decides it cannot absorb change at machine speed the “too fast threshold”.
As AI continues to develop, at some point I think society decides AI is moving too fast. Maybe it’s the “job loss” and “AI displacement” narrative. Maybe it’s tech CEO’s wielding disproportionate power over the imminent future. Maybe it’s deepening social and economic inequality. As I noted earlier, who knows?
What counts here is not whether it’s actually true that AI is the reason for perceived change or not. All that matters is the perception that AI is to blame.
It appears as though we’ve hit “too fast.” The public doesn’t need to agree on p(Doom) to become uncomfortable with the pace, the concentration of power, or the people explaining why both are necessary.
Eventually, some of that discomfort will affect what gets built. Perhaps governments intervene, perhaps the labs exercise restraint, or perhaps investors become less enthusiastic about paying for the next enormous training run. For the sake of argument, suppose frontier scaling stops tomorrow. The existing models remain available and the people using them keep working. How much technological progress would actually stop with it?
Less, I suspect, than the people asking for a pause believe. In Apocalypse Now, Abundance Later, Progress Regardless, I suggested we may already have reached takeoff, and that AI could probably keep improving even if every planned data center were blocked. The models we have can already help find improvements to AI, and distillation can turn each improvement into a cheaper model, which makes the next round of searching cheaper too. Nothing in that loop requires a larger cluster, although a larger cluster would speed it up. By the time we decide AI has become too capable, the capability to keep improving it may already be in our hands.
What the Boom Has Bought
A pause in frontier training would leave behind much of what the extraordinary spending of the past few years has purchased. We would still have models capable of writing useful software and assisting research, along with people who have learned how to direct those models. The next generation might be delayed indefinitely while the work of understanding and exploiting the current generation continued.
At Sun Microsystems, I worked on archival storage built from commodity AMD servers and hard drives, inside a company whose business depended heavily on expensive proprietary systems. The company with the most to lose was building the cheaper approach, and after the dot-com crash customers had every reason to choose commodity hardware. I told that story in Make AI Cheap Again. The frontier labs are repeating the pattern: the cheap versions are coming from the same companies that sell the frontier.
The AI boom has already left us with what I have called Artificial Good Enough Intelligence. “Good enough” depends on the job and the price. A capability that is uneconomical to use across an entire business can become valuable when its cost falls sufficiently. The model need not set a new benchmark record for that change to matter. Someone who can suddenly afford to examine every document, test more designs or automate a previously neglected process has acquired a practical capability they did not have before.
This is why the growing expense of frontier training can coexist with rapidly improving AI economics. Setting a new performance record and reducing the cost of reaching an existing standard are different achievements. Epoch’s price-performance research documents how quickly the latter can advance. From a user’s perspective, the relevant question is often how much reliable work a budget will buy, and that can improve substantially while the most expensive research program in the world gets more expensive still.
Part of that improvement is now being designed into the silicon. Jalapeño, OpenAI’s first chip, is an inference processor built with Broadcom around OpenAI’s own models, kernels and serving systems, and OpenAI says its models helped take it from first design to tape-out in nine months. Google has TPUs, Amazon has Trainium and Microsoft has Maia, and even Nvidia paid about $20 billion last December to license Groq’s inference technology and hire its leadership. Chips are increasingly shaped around the models they will run.
Conversely, the chipmakers themselves are moving toward the models. Nvidia agreed to license Poolside’s model factory to train its own open-weight Nemotron models, and in early September it agreed to buy Hugging Face, the main distribution hub for open models, for about $13 billion.
On September 28, AMD agreed to acquire Fei-Fei Li’s World Labs for $8.2 billion because, in Lisa Su’s words, building the next compute platforms “requires a deep understanding of how models are evolving.” In AI Waves #023, I called this mutually assured dependence: the labs design chips to hedge their dependence on Nvidia, and Nvidia builds models to hedge its dependence on the labs. AMD is now making the same bet.
Each layer of the stack is now designed with the others in view, some of it with help from the models themselves, and the gains per chip keep compounding whether or not anyone builds a larger data center.
How Distillation Works
Distillation helps explain why the cost of acquiring a capability can fall so far below the cost of developing it the first time. A capable model supplies examples or feedback that another model learns from. The student acquires useful behavior without having to reproduce the entire process that produced its teacher. The transfer is imperfect and depends on the task, but it can make an expensive capability economical to deploy.
OpenAI has offered distillation tools since 2024, helping customers use the outputs of larger models to improve smaller ones. The commercial appeal is straightforward. A business can use an expensive model to establish what good performance looks like, then train a cheaper model to reproduce it for work that will be repeated frequently. The initial expense becomes reusable, and the teacher can remain valuable even as the student takes over much of the daily workload.
The teacher does not always have to be larger. Researchers can train specialists in different domains and then consolidate their capabilities in a common student, as in Xiaomi’s multi-teacher distillation work. Specialization supplies the teaching advantage. Different versions of a model can develop useful abilities separately and contribute them to a version intended to perform more broadly.
Cursor’s Composer 2.5 training method goes further within the same model. When a particular action needs correcting, the researchers insert a short hint into its context. The model with the hint becomes the teacher for the model without it, supplying a targeted training signal about how to behave at that point. A correction that initially required additional instruction can become part of the model’s learned behavior. Distillation is also a way of incorporating improvements into the next working version of a system.
Dario’s proposal wants to suppress the unauthorized version of this transfer. He calls for a crackdown on “unauthorized distillation by companies in authoritarian countries,” which “allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.” Maintaining that gap is part of his explanation for how the leading labs could afford to slow down.
Distillation works best when you own the teacher. The owner chooses what to teach, can see the reasoning behind each answer and the probabilities behind each word, and can generate as many examples as the job requires. The method Geoffrey Hinton and two Google colleagues described in 2015 trains the student on those probabilities, which only the owner can see. A rival querying someone else’s API gets final answers without the probabilities, in volumes the provider can watch and cut off. In February, Anthropic reported catching DeepSeek, Moonshot AI and MiniMax generating more than 16 million exchanges with Claude through roughly 24,000 fraudulent accounts. Some of those prompts asked Claude to write out the reasoning behind its answers, and the labs have been closing that door: OpenAI has hidden chain of thought since o1, and Anthropic now stops API users from editing earlier context to extract Claude’s reasoning. Learning from outputs alone works for narrow, repeated tasks, which is what OpenAI’s tools are for, but reproducing a frontier model’s general capability that way is much harder.
If copying through the API is the weak channel, the gap Dario wants to protect is less exposed than he suggests, and the parties best placed to distill are the labs themselves, along with anyone working from open weights, where the whole teacher is available. An expensive discovery still spreads, mostly through its owners’ own cheaper models and through open models others can study end to end. On its own, though, distillation passes along only what a teacher has already learned. The case for progress without another frontier run depends on whether existing models can also help find something worth teaching.
The Research Loop
Some of the work AI budgets now buy is research on AI itself. We have accumulated a useful technology that can participate in finding better ways to develop and use that technology, and a decline in the money available for frontier training would leave some of that research capacity intact, potentially with considerable room to become more productive.
One of Anthropic’s recent experiments assigned Sonnet 5 to improve the alignment of an early Opus 4.8 checkpoint. Sonnet scored below Opus on a general capability measure. Nevertheless, over 60 hours it developed and tested more than 50 proposed solutions, with results approaching the alignment performance of the production model. The researchers prohibited simply distilling Claude’s alignment into the target. Sonnet had to contribute research that improved the training process.
The result concerns a particular dimension of model behavior, but it challenges the assumption that only a more generally intelligent AI can improve another system. Given suitable tools and a way to evaluate results, a model can explore possibilities that its developers have neither the time nor the resources to investigate individually. Useful research can come from organizing existing intelligence effectively, with humans continuing to decide which problems deserve attention.
Dream-RSI, published in September, makes that possibility especially concrete. Its underlying coding agent stays fixed while the process directing its search improves. Records of previous discovery attempts provide a setting in which alternative search strategies can be tested against outcomes already observed. The system then takes a better strategy into subsequent live runs. It still needs inference and fresh experiments, but it can improve how it spends those resources without receiving a smarter foundation model.
A fixed model can therefore participate in a changing research process. If it helps discover a more efficient algorithm, subsequent experiments may become cheaper. If it helps improve the way candidate solutions are selected, fewer attempts may be wasted. Where successful behavior can be learned, distillation provides a way to make it cheaper and easier to reproduce. Some of the gains increase the productivity of the work that produces further gains.
This is a recognizable form of recursive improvement, even while humans remain heavily involved. Anthropic’s own discussion of the possibilities allows for substantial acceleration even if people retain the advantage in choosing research directions. OpenAI’s chief scientist, Jakub Pachocki, has made building an automated AI researcher an explicit priority, although he places it within a continued scaling trajectory. The labs differ in their expectations, but they are already investing in AI’s ability to contribute to further AI development.
How Far Can Existing Intelligence Take Us?
A collection of models teaching one another could also become very efficient at reproducing mistakes. Distillation makes behavior transferable without guaranteeing that the behavior is correct, and an impressive demonstration of automated research does not establish an indefinite supply of breakthroughs. The process needs ways to encounter evidence that its current assumptions are wrong.
Software provides some unusually direct opportunities for that encounter. Code can be executed and its behavior measured. A model proposing an optimization can discover that its idea fails, regardless of how convincing its explanation sounded. The useful information comes from the experiment. Existing intelligence can conduct new experiments and learn from new results without first requiring a more capable model to be trained from scratch.
As I argued in Blow the Whistle: Everything’s an Eval, the ability to judge performance becomes more important as models take on more consequential work. A system that raises its score by exploiting a faulty test has found a problem with the test. Someone still has to distinguish that achievement from an improvement in the work we actually wanted done. Better evaluations can make existing intelligence more productive, and poor ones can make apparently rapid progress almost meaningless.
Human judgment remains valuable throughout this process. Choosing an interesting problem and noticing a misleading result are both part of research. So are the less glamorous constraints of obtaining equipment and waiting for the physical world to supply an answer. Automating more of the intellectual work does not remove every limit on the rate at which useful discoveries arrive.
That leaves two possible worlds, and I don’t quite know which one we’re in.
In the first, scaling is the only thing that works, and a pause stalls progress until someone finances the next cluster, most likely out of the current leaders’ margins.
In the second, constrained compute forces the innovation that abundant compute made unnecessary, and models keep improving through better algorithms, better data and better use of the hardware we already have. The bitter lesson, that general methods riding ever more compute beat clever ones, has held for as long as the compute kept coming. It may also be partly self-fulfilling.
When funding is cheaper than good ideas, adding GPUs is the rational way to make progress, however wasteful, and it has worked too well to bet against.
SOURCE: https://arxiv.org/html/2412.19437v2
DeepSeek offered a glimpse of the second world. It trained V3 on H800s, the chips Nvidia cut down to comply with American export rules, and the constraint shows in the choices its team made: an attention design that shrinks the memory the model needs to keep track of long inputs, a mixture of experts that activates about 37 billion of the model’s 671 billion parameters for each token, and training at lower numerical precision than most labs had attempted at that scale. By DeepSeek’s own accounting, the final training run cost about $5.6 million in GPU time. Labs with more compute had less reason to look for those savings.
Some of the researchers now looking are AI. Transformer Lab, a Canadian startup, has built Primus Society, which it describes as “10,000 autonomous researchers, working together as a virtual community on frontier problems.” They are organized into labs that compete for funding, sign their work and earn reputations from how their experiments turn out (one lab calls itself the Bitter Institute). Some researchers are assigned to try ideas nobody else will fund. Others, the skeptics, spend their compute trying to break what the rest claim, and labs regularly publish results that refute their own predictions. The society’s first headline result answered a call to cut the cost of pretraining. By growing a smaller model into a larger one partway through training, the labs say they matched a model trained from scratch with about 30% less compute, and no human intervened between the call opening and the first papers coming back. Transformer Lab says science in Primus Society is limited only by the number of GPUs. Its researchers spent their first rounds of funding finding a way to need fewer.
If I were running a frontier lab, I’d be looking very hard at distillation and every other way of doing more with what I already have, because there’s no guarantee the next round of resources arrives. New frontier models may still accelerate the process enormously, giving researchers better tools and distillation better teachers. But the intelligence the last round of spending bought can already contribute to the research, and a pause would be the first real test of how far it can go on its own.
Ten Days of Restraint
On September 22, ten days after Dario’s call for restraint, Anthropic released Opus 5.5 and OpenAI released GPT-6 Sol and Luna. Dario had explicitly allowed continued training and technical progress, so the launch calendar does not establish a broken promise. Nevertheless, the first releases of the new era offered customers a considerable improvement in what they could accomplish for their money. I am enjoying the restraint so far.
In one of Anthropic’s internal tests, Opus 5.5 rewrote HAProxy, widely used software for distributing web traffic, from C into Rust. The rewrite passed nearly all of HAProxy’s own regression tests, took 9.5 hours and cost 51% less than Fable 5.1’s attempt, and Anthropic says the model performs at Fable’s level on most work. OpenAI’s Sol and Luna bring advances from the GPT-6 family into more affordable models, and on OpenAI’s reported AutomationBench comparison, Sol at xhigh effort completed more business workflows than Opus 5 at max effort, for roughly one-eleventh the cost per task. Opus 5.5 arrived the same day and outscored both, and OpenAI followed Sol with GPT-6.1 Sol on September 29. These are particular evaluations, not universal exchange rates between models, but both companies are selling substantially more useful intelligence for less money.
Neither company has said exactly how it built these models, but both releases do what distillation is for: they carry capability developed at enormous expense into models cheap enough to use constantly. The labs are running the loop themselves. The original Make AI Cheap Again argument left plenty of room for enormous clusters and ambitious training runs, with software improvements making useful intelligence cheaper while the pursuit of the most capable possible system absorbed extraordinary resources. These releases show how enthusiastically the same companies can pursue both.
For the people using these systems, lower costs create room to be more ambitious. Work that barely justified one attempt may justify several. An experiment previously excluded by the budget becomes worth running. It could come from an engineer improving neglected software or a researcher following an unfashionable question, someone who knows something the labs don’t. Some of those experiments will improve the software, models and research methods on which future work depends. Better intelligence makes that process more productive, and cheaper intelligence opens the same process to more people.
The same companies asking society to consider slowing AI are expanding the supply of intelligence available to everyone building with it, and distillation means each release leaves a cheaper starting point behind. By the time we agree on how quickly the labs should be allowed to advance, we may already have given enough people enough intelligence to carry a substantial part of the progress forward themselves.















