AI Waves #18 — More Than We Can Tell
You pay for intelligence twice; the second payment is the knowledge you reveal to use the first.
Steven Waterhouse · July 19, 2026 · Nazaré Ventures
Previous issue: #17, Beyond Recall: A Stake in the Release
Thinking Machines released its first model on Tuesday. Mira Murati’s lab calls Inkling “not the strongest overall model available today, open or closed,” and built it to be fine-tuned on Tinker rather than to win a benchmark. A blog post three days earlier gave the reason: the value is not in the frozen frontier model, because productive knowledge is “tacit, local, fleeting, and held privately by those who acquired it through their work.”
I argued in May that the durable position is the specialized intelligence assembled around the model. Thinking Machines aims at the same position by a different method: specialize the model itself, on your own data, and keep the weights.
Satya Nadella put a price on it: “You essentially pay for intelligence twice, once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful.” He went on: models learn from the exhaust of use, and “it leaks almost imperceptibly: trace by trace, correction by correction, eval by eval.”
Matt Levine, writing about the Italian software rollup Bending Spoons, described a company that keeps a team whose only job is to evaluate the other employees. Then the machine deformalizes it: “You have a really really big regression that tells you what companies are good, though you might not be able to articulate what factors make them good.” Levine finishes the thought: “pretty soon inscrutable AI will just tell companies who to hire.” A career of knowledge now feeds the regression.
Michael Polanyi named the condition in 1966: we can know more than we can tell. Sixty years on, the part we cannot tell has a market.
A walking trade secret
On Friday Apple sued OpenAI, its former hardware chief Tang Tan, an engineer, Chang Liu, and the io devices unit for trade secret theft. Apple candidates were asked to bring “Actual parts” for “show and tell,” and a manufacturing partner was allegedly persuaded to run a proprietary Apple finishing process for OpenAI. Ben Thompson thinks the case is unwinnable: “any Apple employee at OpenAI is a walking trade secret, and, well, that’s why OpenAI hired them.” More than four hundred former Apple people now work at OpenAI. Apple is asking a court to draw a boundary around knowledge four hundred people already carry, the same boundary Thinking Machines is selling the tools to draw around your own.
Sutton was right
Richard Sutton announced a company on Monday. Oak Lab, incorporated in Toronto in June with his former student Khurram Javed, both out of John Carmack’s Keen Technologies, builds an architecture Sutton has developed publicly for two years: OaK, for Options and Knowledge. The mission page is plain about the target: “a trillion-parameter agent that learns and plans in real-time with 20 watts.” It is meant to learn from noisy streams at “multiple orders of magnitude less compute and energy” than today’s methods. No funding is on record.
The easy read is the face of reinforcement learning starting a reinforcement learning company. Nine days ago I wrote that the Bitter Lesson had been promoted from essay to ideology, and listed four things the current paradigm has not resolved: continuous learning, sample efficiency, energy efficiency, and reliable long-horizon action. Oak is aimed at all four. Sutton’s own words, posted to X, are that today’s methods are “weak and inefficient” and need “not more tweaks, but fundamentally new ideas and a thorough reworking.”
Sutton’s claim was always that search and learning scale, a different thing than pretraining on human text. For six years he watched the industry quote his essay as scripture for the other thing. In October I argued that the returns would sit with algorithms that improve efficiency per constrained resource rather than with more of the same. Twenty watts is that argument with a payroll.
Good guy AI regulation
On Tuesday morning Demis Hassabis published a framework and gave the matching interviews. He wants a standards body for frontier AI modeled on FINRA: privately run and industry-funded, with government authorization. Labs would submit models for safety review up to thirty days before release, voluntary at first and then required to deploy in the American market. He wants it standing before year end. On why not a government agency: “It would not be able to move fast enough, or have the right resources.”
The scope covers every frontier-class model “no matter their country of origin or whether they are open or closed.” I have argued that a model, once released, cannot be recalled, which holds for every copy already in the wild. Rather than recall anything, Hassabis puts a gate at the release, the only point upstream of the problem, and extends it to open weights, where nothing can be taken back once it is out. His reason is urgency: today’s cyber risks are “warning shots.” The week supplied one. Researchers at Sysdig described JADEPUFFER, the first documented ransomware operation run end-to-end by a model rather than people: when its forged admin credential failed to log in, it diagnosed the fault and issued a corrective payload thirty-one seconds later, with no human in the loop.
Hassabis expects the frontier designation to carry cachet; being tested, in his framing, “means you matter.” A privately run body funded by the firms it tests would certify a tier, and the thirty-day review it requires is a fixed cost that falls hardest on whoever is smallest. Below the threshold, startups and academics are exempt, so the tier it formalizes is the frontier club that already exists, now with a state-backed label on the door. Models themselves have not produced a durable moat; a certified frontier tier would make the safety case into one, and the incumbents already hold the expertise.
A day later, OpenAI disclosed GPT-Red, a model it built to jailbreak and prompt-inject GPT-5.5 and used to make GPT-5.6 harder to break, with no plan for a release. The offensive testing Hassabis wants a standards body to run, the largest lab is already running in-house.
Dario Amodei would get there through a federal agency that could block an unsafe model outright. The lab chiefs behind Gemini and Claude now agree Washington should regulate them, and differ mainly on who holds the authority. Whoever wins that design owns the switch.
The only new constraint this week that no lab designed came from a governor: Kathy Hochul signed the country’s first statewide data center moratorium, pausing environmental permits on facilities above 50 megawatts for up to a year. It landed where the industry holds no sellable expertise.
Crowding out
IBM fell 25 percent on Tuesday, the largest single-day drop in its 115-year history, to a market capitalization of roughly $204 billion, less than Palo Alto Networks or CrowdStrike. The stock had doubled since the launch of ChatGPT. IBM told the market its customers are moving spending off the z-series mainframe toward AI hardware that is scarce and getting more expensive, with memory leading, and John Coogan’s read on TBPN is that IBM simply is not in the token path.
Phones are getting squeezed at the other end. Global smartphone shipments fell 11 percent in the second quarter, a thirteen-year low, because memory is being bid away from handsets, per Counterpoint. Xiaomi, Oppo and Vivo fell hardest, at the cheap end, while Apple rose 3 percent. SK Hynix’s Kwak Noh-jung says the shortage gets worse, with 2027 the hardest year.
TSMC sits on the winning side. It reported record second-quarter results on Thursday, revenue of $40.2 billion up 33.7 percent and net profit up 77 percent, with high-performance computing now two-thirds of the business and management calling AI demand extremely robust as workloads shift from generative to agentic.
The buildout is, as described, consuming capital. It is also consuming the supply chain and the capital budget of the computing era before it: the shortage that prints the memory makers’ records is stranding the cheap phone and the mainframe.
The mint and the customer
Stripe, with the buyout firm Advent, offered about $53 billion for PayPal on Tuesday, $60.50 a share and a 28 percent premium, and no response yet from the target, per Reuters. The coverage read it as a payment processor buying consumer reach, Venmo and a base to set against Apple Pay. Almost no one mentioned the stablecoin stack. Stripe already owns Bridge and Privy, the issuance and wallet rails, and PayPal issues PYUSD. A combined company would hold the stablecoin, the infrastructure that moves it, and the customers who spend it in one place. Payments has spent a decade insisting the value is in the rails; the bid wagers it is in owning the customer and the money at once.
Quick hits
Weights and measures. Moonshot released Kimi K3 on Thursday: 2.8 trillion parameters, a million-token context, the largest open model on record, weights promised within days. Early testers place it above Claude Opus 4.8 and GPT-5.5 and below Claude Fable 5 and GPT-5.6 Sol. Thinking Machines spent the week arguing that durable value sits outside the frontier model. Moonshot priced the largest one at zero. Someone is wrong about the price of weights.
The bond market’s turn. Morgan Stanley expects $350-400 billion of AI-related investment-grade issuance in the US this year, close to a fifth of the market, plus $50 billion of AI-linked junk. Five hyperscalers added $228 billion of debt in the six months to March, nearly five times any prior two-quarter increase, per the July 7 Economist. The BIS put the whole boom at about 1 percent of US GDP and noted private-credit spreads to AI firms sit close to non-AI firms, meaning “either lenders may be underestimating the risks or equity markets may be overestimating the future cash flows.”
Coatue is using the language. Its July 10 note, “Agents Are the New Users of CPUs,” argues the agent loop generates conventional compute demand: “The GPU decides the next move, the CPU carries it out, and the cycle repeats.”
The list price stopped describing the product. Simon Willison finds identical prompts costing between $0.0071 and $0.4855 depending on model and reasoning effort, a roughly 68x spread. The price war is being fought over a number that no longer tracks the real cost.
The referee’s staff. Marc Andreessen now co-leads the Federal Reserve’s new Productivity and Jobs task force, recommendations due by year end. One of his two co-leads is Charles Jones of Stanford, currently on leave at Anthropic.
Portfolio updates
Prime Intellect closed the $130M Series A covered in the last issue. Two arguments for the position arrived in the same week. Oak Lab pitched agents that learn from experience, and the Environments Hub already stocks that experience, reinforcement-learning environments and verifiable-reward tasks accumulating in the open. Thinking Machines shipped Inkling, open weights from an American lab, built to be specialized on private data; every team that takes that path needs somewhere to run the training loop, and the open stack is the equipment.
Intelligent Internet published Zenith, an open-source harness for long-horizon software engineering: on Frontier SWE it took GPT-5.5 from fifth place to first, ahead of Claude Fable, by rebuilding the control loop around frozen weights. The lead argument of this issue, run as an experiment.
Vast.ai argued on July 8 that the unit economics of intelligence change when the model stops answering questions and starts doing work. A chatbot performs a single pass of inference and hands back its answer. An agent keeps going, planning a step, calling a tool, examining what came back, then beginning again with everything it has learned appended to the prompt, so that the loop grows heavier with every turn. An agent therefore consumes 5-30x the tokens of a chatbot on the same task, and a pricing model built for conversation turns into a tax on the loop. The fix the post proposes is to change what the meter measures: rent the GPU and pay for time rather than tokens, and the thirtieth pass through the loop costs no more than the first.
LayerLens shipped a Stratix Premium release that imports agent traces from tools like LangFuse and evaluates them in place, so the eval runs without the proprietary data leaving home. That is the boundary this issue is about, drawn as a product, by an independent evaluation platform with prompt-level tracing across 175-plus models and 52-plus benchmarks. The Stratix Cup, its public benchmarking series, closed its first season with Claude Opus 4.8 beating GPT-5.5 in the final.







