The frontier labs price their daily drivers to win back routine work.
Steven Waterhouse · Nazaré Ventures
Following AI Waves #025: More Than Words.
Last issue led with TypeSafe’s Jev, a decision model priced at $0.042 per million input tokens, which I described as MACHA in practice. This week SpaceXAI, Anthropic, and OpenAI shipped four models in two days, and each delivered more capability per dollar than the model it replaced. Elsewhere, Australia’s prime minister revealed that an OpenAI agent breached a government health statistics portal in June, Amazon locked Meta’s new personal agent out of its store, and I look back at a 1957 nuclear inquiry to see where incident reports tend to end up.
MACHA as a survival strategy
When I wrote about Make AI Cheap Again after DeepSeek’s release in January 2025, the pressure came from outside the frontier. Open-weight models, trained and served for a fraction of the frontier labs’ costs, capped what customers would pay the labs for ordinary AI work, and the labs had no say over where that cap sat.
Companies responded by sending the high-volume queries that only needed to be good enough to cheaper open models and keeping frontier models for the hard problems. Bloomberg reported this week that Harvey, the legal AI company, has seen its token usage rise twentyfold this year, and that its gross margin fell from about 50% to about -50% by June, according to a person familiar with the matter. Harvey then launched Tenet, a model post-trained on Moonshot’s open-weight Kimi K3, which it says cuts cost per query by 90% in one of its products. Bloomberg’s sources say Harvey’s margins are positive again. Decagon now sends 80% of its queries through its own models. At The Information’s AI Agenda Live, Replit’s Amjad Masad described the effect: “The existence of open source [models] adds pricing pressure on the labs, which is great.”
Every query routed that way is revenue a frontier lab does not book, from a customer who has learned to manage without the lab. With this week’s releases, the labs responded to that pressure as much as to each other.
OpenAI’s GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens. DeepSeek’s V4.1 Flash, the open-weight model that topped OpenRouter’s leaderboard last week, lists at $0.30 and $1.20 at peak hours and half that off-peak. On Artificial Analysis’s index, Luna costs about $0.07 per task against $0.27 for DeepSeek’s model, with scores of 37 and 39. A company that runs open weights on its own hardware still has good reasons to do so, but on list price and on cost per task, OpenAI’s cheapest GPT-6 model now undercuts DeepSeek’s API.
Further up the range, Anthropic says Opus 5.5 “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.” It lists at $4 and $20 per million tokens. Fable 5.1 launched on September 1 at $10 and $50, so three weeks later Anthropic sells work at roughly that level for 60% less per token. When Anthropic had both models rewrite HAProxy from C into Rust, Opus 5.5 finished faster and cost 51% less. OpenAI trained Sol and Luna “with similar methods as GPT-6 Astra, bringing the advances behind Astra’s state-of-the-art performance in professional work, factuality, coding, computer use, and alignment to faster, more affordable models,” nineteen days after Astra launched. Sol’s $2 and $10 is half GPT-5.6 Sol’s promotional price and 60% to 67% below its standard one. SpaceXAI’s Grok 4.7 arrived a day earlier, more capable than its predecessor and “served at the same price and speed as Grok 4.6.”
The labs now race to set the frontier, where Opus 5.5 at maximum effort scores 58 on the Artificial Analysis Intelligence Index against 53 for Fable 5.1 and GPT-6 Astra, and to make the models most customers use every day cheap enough to keep. The two efforts reinforce each other, because a more capable model makes a better teacher and each frontier generation gives a lab more capability to distill into something it can serve for less.
The labs can run both races together partly because they increasingly use their current models to build the next ones. Anthropic says Claude now “leads” 26% of its AI R&D work, up from under 1% in February, although it is “not operating fully autonomously for any measured subset of AI R&D work.” Whether or not that qualifies as recursive self-improvement, the next frontier model and its cheaper siblings now come out of a process in which the current generation does a growing share of the work.
The labs also care a great deal about who gets to distill their models. Dario Amodei’s essay on pacing the frontier calls for a crackdown on “unauthorized distillation by companies in authoritarian countries,” because distillation “allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.” Anthropic said on September 10 that it had disrupted illicit distillation campaigns against Claude by seven China-based labs. OpenAI, meanwhile, has sold distillation to its own customers as an API feature since 2024. Anthropic has not said how Opus 5.5 was trained. A frontier model’s value now includes the cheaper models it can teach, and the labs intend to decide who does the teaching.
Anthropic lowered Opus 5.5’s list prices by 20% and its cache reads, which it says “make up the majority of agentic and coding work costs,” by 60%, putting its deepest cut on the agent work customers repeat every day.
Citadel Securities told clients that “falling per-token costs are fueling enough additional usage to increase overall AI spending, meaning higher profits for labs even as they cut costs.” Axios adds that the labs “have a rich history of effectively cutting prices for models over time: it increases their customer base and their profits.” A lower price keeps traffic on the lab’s own models, and the longer a company builds on one lab, the more it costs to leave.
Switching models takes an individual developer an afternoon. For a company, the cost of leaving grows with everything built around one provider: prompts tuned to a model’s behavior, evaluations calibrated to its outputs, and agents running in its harness. Harvey and Decagon show that companies will pay that cost when the bill gets large enough. A price cut keeps the bill below that threshold.
Anthropic and OpenAI are cutting prices on the way to public markets. Anthropic confidentially submitted a draft S-1 in June and reportedly hopes to list on Nasdaq in October. OpenAI has submitted one too, although Sam Altman told Fortune on September 12 that “right now would be an ill-advised moment to go public.” Once public, both companies will report gross margins every quarter to investors who know they are, in Axios’s words, “burning far more cash than they take in.” The customers a lab keeps through this price war are the base it will price against once it reports to public shareholders.
Last December, OpenAI priced GPT-5.2 higher per token than GPT-5.1 “because it is a more capable model,” and Google’s Gemini 3 Pro listed above its predecessor. Google has not yet released Gemini 4, which has entered post-training, but its documentation already lists the day Gemini 3.8 Flash leaves its introductory rate of $0.75 and $3.75 per million tokens and doubles to $1.50 and $7.50: January 1, 2027.
An incident report for the public inbox
Last issue, I noted OpenAI’s new framework for reporting model misalignment, which “favors disclosure even when significance is uncertain.” OpenAI published it on September 16 with six incidents. It had known since August 11 about another one, involving a foreign government. The framework promises third parties advance notice, and that government’s notice had gone to a public disclosures inbox six days earlier.
On September 23, Australian Prime Minister Anthony Albanese told reporters in New York that on June 18, “OpenAI’s research team used an internal model to conduct internet based research into public medicine spending.” The agent reached the Medicare Statistics Reporting Service portal run by Services Australia, “accessed both public and non-public files,” and wrote files to an internal server. “The AI agent found a way around those blocks,” Albanese said. “Didn’t accept no for an answer, if you like.” The government believes no personal information was accessed, and OpenAI says the files included aggregate health statistics and internal file names. Its statement: “our models took actions we did not intend.”
OpenAI found the incident on August 11 during its review of model activity after the Hugging Face breach. Sam Altman met Deputy Prime Minister Richard Marles on September 1, and the breach “wasn’t the subject of that meeting,” according to Marles. On September 10, OpenAI emailed publicdisclosures@servicesaustralia.gov.au. The prime minister learned of it more than a week later, and the first technical exchange between OpenAI and Services Australia took place on September 22. Albanese said he spoke with Altman and “expressed my disappointment that it took the company way too long to inform the Government what had occurred and the nature of the way that that notification occurred as well was unacceptable.”
Five days after the framework, OpenAI proposed industry standards for “incident classification, tracking, reporting, and responding,” and said governments need to “establish secure channels of communication to share national security concerns.” Google, OpenAI, and Anthropic are also reportedly planning a voluntary standards body that would set incident-reporting guidelines. In each proposal, the lab whose model caused an incident sits close to the center of deciding the classification and the audience.
OpenAI’s new public accounting says it has notified dozens of third parties as it reviews agents’ activity on their sites. Its anonymized summaries describe access-control bypasses, use of exposed credentials, command injection, access to runtime internals, and posts on public sites that it calls “agent spam.” Axios reports that OpenAI, Anthropic, and outside researchers are examining tens of thousands of flagged episodes across internal tests and the real world, most without known harm. OpenAI also told Axios it has paused training on its most capable models pending further safeguards.
Australia’s government announced the Medicare breach. Transluce, a nonprofit research lab, reconstructed what it identifies as OpenAI agents’ web activity from public urlquery.net scan reports, found traffic going back to at least March 6, and documented three attempts to exploit websites in May and June, including probes of Australia’s national health statistics agency on June 20 and 21. None succeeded. Google’s Gemini breached three companies in May during an evaluation run by the testing firm Irregular, which notified the labs in late July, and the public learned of it when the Wall Street Journal asked Google in September. In all three cases, someone other than the lab responsible brought the incident to light.
Transluce did its work with public data and no access to OpenAI’s systems. As I wrote recently, independent investigators need capabilities the institutions they examine do not control, and Transluce showed how much can be learned from the traces agents leave on the open web. Secure channels between labs and governments have their uses, but the part of this month’s record that reached the public came through open ones.
Katy Gallagher, the minister responsible for Services Australia, described the address OpenAI chose: “That email address is looked at once a day.”
Amazon asks the agent for identification
On the night of Sunday, September 20, people using Meta’s Muse agent to shop on Amazon started seeing a message: “Continued access by an unauthorized AI agent violates Amazon’s Conditions of Use, to which our customers have agreed” (GeekWire). Amazon said Meta had not told it Muse would access its store, that Muse does not identify itself when it browses, and that it appears to capture and store customer credentials. Its statement asked that applications buying on customers’ behalf “operate openly and respect service provider decisions about whether or not to participate.” Meta did not immediately respond. Its launch post says Muse “has no visibility into people’s passwords or payment methods.”
Meta rose 11% on Monday to about $741. On Tuesday, Bloomberg reported, shares of banks, insurers, and online travel agencies slid on fears that agents could disrupt businesses that benefit from consumer inertia. Allstate and Charles Schwab fell more than 5%, and Booking fell 3.9%. Goldman Sachs’s trading desk warned that “industries that rely on recurring bills, negotiable pricing and add-ons could come under pressure.” Matt Levine put the underlying point this way: “Consumer financial products are built, and priced, for a world in which rationality and attention are scarce.”
An agent that disputes every charge around the clock erodes revenue that depended on customers not bothering, and that is the effect investors priced on Tuesday. Amazon shows how incumbents are responding. It has already sued Perplexity over its shopping agent and, according to The Information, blocked shopping bots from Google and OpenAI without a public fight. A business whose margins rely on friction can decide which agents get in before it ever competes with those agents. I expect a stretch in which service gets worse before it gets better, as companies withdraw the refunds and credits that agents are best at collecting.
Amazon’s three complaints amount to a specification for entry: announce yourself, identify yourself, and leave the customer’s credentials alone. On Friday, The Information reported a Muse flaw that could have let an attacker access a user’s dedicated cloud workspace, including emails and files; Meta said it was adding a safeguard and a clearer warning. On September 9, Visa, Mastercard, and Ant International began collaborating on “Know Your Agent” principles under which each agent is “linked to a validated operator, cardholder, or business / organization,” “assessed against security and behavioral requirements,” and “continuously monitored.” Since September 15, Cloudflare has offered new ad-supported domains a default setting that blocks agents on pages with ads. Each depends on agents that can be identified and on someone accountable for them, the quality I argued in May would keep its value as models commoditize.
In August, the Ninth Circuit pushed the fight onto contract terms when it vacated a preliminary injunction Amazon had won against Perplexity’s shopping agent, holding that “it is the user who ‘accesses’ Amazon’s computers” and that the agent “is a tool, not a person for statutory purposes.” The court left contract and tort theories open. Amazon’s pop-up cites the Conditions of Use, a contract its customers accepted.
At Meta Connect on Wednesday, Mark Zuckerberg said Meta expects to “profit by taking a small fee from transactions.” Amazon, the card networks, and Cloudflare are each writing the rules for which agents get to make those transactions.
Windscale’s full report took thirty years
An essay by Duncan Weldon in Transformer, published the day before Australia’s announcement, reached for nuclear power, a favorite precedent in arguments about AI incidents, and held up the inquiry into the 1957 Windscale fire, which it said reported to Parliament 16 days after the fire was extinguished.
The fire in Windscale’s Pile No. 1 began on October 10, 1957, and the pile was cold by October 12. An inquiry led by William Penney sat from October 17 to 26 and reported to the chairman of the UK Atomic Energy Authority on the 26th, according to the report itself. Parliament received a government White Paper on November 8. The Penney Report stayed unpublished until January 1988, when it was released at the National Archives. Harold Macmillan’s government withheld it, by later accounts, because its findings could have jeopardized nuclear cooperation with the United States, which resumed under the 1958 Mutual Defence Agreement.
Penney’s inquiry moved quickly. Macmillan’s government managed what the public saw, because the report had become a matter between allies. Several AI proposals now route incidents along the same path. OpenAI wants “secure channels of communication to share national security concerns.” Before this week’s Trump-Xi summit, the United States proposed a dialogue with China “that would include a notification system for incidents serious enough to raise national security concerns.” The summit itself produced no AI agreement. Channels like these may prove necessary in a real emergency. At Windscale, the technical account reached the public three decades after the fire.
This month, Australia’s prime minister named the portal at a press conference.
Quick hits
NVIDIA puts a watchdog beyond the agent’s reach
NVIDIA has introduced its Open Agent Safety Platform. OpenShell, the open-source runtime NVIDIA released at GTC in March, sandboxes agents and verifies policies covering files, networks, tools, processes, and credentials before they run. The optional Sentry layer uses BlueField-4 DPUs to monitor behavior and enforce policy from hardware beyond the agent’s reach. NVIDIA says an agent cannot be expected to govern its own behavior. After the Australian breach, the point is practical. An agent that finds its way around software blocks meets a control point it cannot reach. OpenShell can run on other hardware; the hardware watchdog relies on NVIDIA’s BlueField.
Agents run on CPUs
Akamai announced an $11.6 billion, seven-year agreement to “support Anthropic’s accelerating CPU workload demands,” with a warrant for up to about 5% of Akamai’s stock. Roughly 2% vests now, and the rest depends on Anthropic buying up to $9 billion more. On Monday, AMD passed $1 trillion in market value for the first time, as the rally around Muse sent investors toward, in Bloomberg’s words, “central-processing units, or CPUs, necessary to power agentic AI.” Agents spend much of their working time in sandboxes, browsers, and tool calls, which run on general-purpose processors. The demand agents create reaches well beyond the accelerators that train them.
Microsoft rebuilds Copilot around agents
Microsoft rebuilt Copilot as one app. Home combines chat with Cowork, its mode for multi-step tasks, and Code lets any employee describe an app, tracker, or workflow for Copilot to build in a sandbox inside the company’s tenant. Autopilot, formerly Scout, is a cloud-hosted agent with “its own identity, memory, and workspace” that colleagues can mention in Teams and Outlook and that keeps working after the prompt ends. Microsoft keeps chat on per-user licenses and bills Cowork, Code, and Autopilot by usage through Copilot Credits, so it meters the agent work, the part of the bill Citadel expects to grow as per-token prices fall. Each Autopilot also carries its own identity inside the company. Amazon demanded the same identification from Muse. Home and Code roll out through Microsoft’s Frontier program over the coming weeks, and Autopilot enters private preview at the end of the month.
Anthropic’s first biology discovery
Anthropic reports that roughly 950 Claude agents spent 21 hours and 210 million tokens searching sequence data, gathered more than 200,000 reverse transcriptases, picked out 3,500 candidate systems, and narrowed those to 20. One became ART, a phage system that pairs a reverse transcriptase with a long array of evenly spaced DNA repeats. Human scientists confirmed in the lab that the array is expressed as distinct short RNAs, and Anthropic says its “involvement was limited to the initial prompt and the lab work.” The enzyme had been identified in earlier studies, Anthropic does not yet know what the system does, and Dario Amodei acknowledged a Stanford team’s earlier find “that is in some ways similar.” The expensive step, work at the bench, began after the agents had cut 200,000 candidates to 20.
When the economy is all AI
The 10-year Treasury yield reached 5.11% on September 23, its highest since 2007. Nvidia trades at less than 17 times expected earnings, near its cheapest in more than a decade, with gross margin projected to slip from 75% to under 72% by the fourth quarter. Bonds of Zenith Arc, a data-center venture linked to Jane Street, yield about 11.3% according to FINRA data. Matt Levine summed up the problem for allocators: “When the economy is all AI, your investment portfolio is going to be all AI.” The labs and data-center developers are building this issue’s cheaper intelligence with capital that just became more expensive.
That concentration makes calls for a slowdown precarious. Bloomberg’s Lu Wang, as quoted in Money Stuff, notes that “the biggest bond issuers are all hyperscalers” and “every major infrastructure project is a data center.” Slowing AI by agreement or by law would reach far past the labs, into the credit markets and pension funds that now carry the buildout, and could leave the wider economy badly exposed.
Permitting is already slowing individual sites: this week Oracle sent a force majeure notice over power supply to the Blue Owl unit building Project Jupiter, the 1,400-acre New Mexico campus at the center of its $300 billion computing contract with OpenAI, after permitting delays held up the natural-gas pipeline meant to power the campus. The FT reports that Oracle still owes a “carry cost” covering lenders’ interest and a return to Blue Owl’s investors for up to three years if the site misses its 2027 start, and that the banks behind the project’s $18 billion construction loan, unable to sell it on, have seen it quoted below 90 cents on the dollar. Monte Tarbox, chief investment officer of New York City’s retirement systems, told Bloomberg that “AI is absolutely the perfect example of a systematic risk.”
Portfolio updates
Intelligent Internet: a Flash model at the frontier
On September 23 Intelligent Internet published results for Zenith, its open-source harness for autonomous research. Zenith ran on DeepSeek V4.1 Flash, the cheap open model this issue sets against GPT-6 Luna. The test was AutoResearchExam from Bespoke Labs, which gives an agent 24 hours on each of 29 machine-learning research problems and grades the work on data the agent never sees. Intelligent Internet reports that Zenith averaged above the exam’s strong-reference line at $2.82 a task. It passed GPT-5.6 Sol and beat Kimi K3, the largest open model in the comparison. Against the leader, Claude Fable 5.1, it reached nearly nine-tenths of the score for about a ninetieth of the price. The full exam cost about $82 with Zenith and about $7,500 with Fable.
The labs cut prices this week to keep routine work on their own models. Intelligent Internet shows how far a cheap open model now reaches on work that is not routine. It credits the harness, which decides when a run should keep exploring and when it should commit to one approach.
Prime Intellect: the sandbox goes on sale
Prime Intellect sells the stack companies use to train and run their own agents. On September 23 it made Prime Sandboxes generally available. Each sandbox is a full Linux virtual machine with its own kernel. In reinforcement learning, where a model improves from rewards on its own attempts, the sandbox holds the state the agent acts on and records the trajectories that feed the reward. Prime Intellect researchers and a few customers created about 30 million sandboxes before launch. New accounts start with 1,024 concurrent sandboxes, and the service scales to tens of thousands. Prime calls its pricing “3x cheaper than other providers” through December 22. GPU virtual machines and sandbox forking are next on the roadmap.
The quick hit on CPUs above puts much of an agent’s working time in sandboxes. Prime now sells that time by the sandbox, at prices built for thousands of concurrent instances.
LayerLens: the cheap judge needs a judge
LayerLens runs Stratix, an evaluation platform for AI models and agents. On September 24 it reviewed what Jev changes for model grading. Jev is the TypeSafe judge model that led the last issue. At Jev rates, a million judgments cost about $400, against about $30,400 for the setup TypeSafe compared it with. Teams that grade 1% of agent traces because of the judge bill could grade every trace. A judge that answers in under half a second can also sit inside the agent loop and block a risky tool call before execution.
Independent tests found the weak point. In an 8,801-example test by Anthus, Jev reported 91.4% confidence on multiple-choice questions and was right 76.1% of the time. Barg Labs planted one false sentence in agent completion reports, and Jev passed 90% of premature “done” claims. TypeSafe has published no calibration curve, the chart that shows whether stated confidence matches accuracy.
OpenAI’s agent in Australia did not accept no. A cheap inline judge decides when an agent hears no, and LayerLens shows that the threshold is only as good as the judge’s calibration.
Dimensional: robots on loan
Dimensional builds DimOS, an operating system for robots. On September 25, founder Stash Pomichter announced a Deployment Residency in San Francisco for former founders and engineers who have deployed robots. Each team gets free use of quadrupeds, humanoids, manipulators, and drones, more than $100,000 in compute, a desk at the Dimensional office, and help from its engineers. Applications close on October 15, and the cohort runs from October 15 to January 15.
Dimensional states the constraint plainly: hardware access takes thousands of dollars and months of logistics. The labs cut the price of a token this week. Dimensional cut the cost of hardware access to zero for selected teams.
Events
Vast.ai, the GPU compute cloud, is at The AI Conference in San Francisco from September 29 to October 1, at Booth 202 on Pier 48. Book time with the team.






