171: With Henry’s “AI Quarterly 26Q2”: From Coding to RSI—A Future Where the Strong Get Stronger?
171: With Henry’s “AI Quarterly 26Q2”: From Coding to RSI—A Future Where the Strong Get Stronger?
Summary
- OpenAI’s coding comeback has been validated. Claude 4.7 was clearly unpopular, and a May pricing change—third-party harnesses switched from subscription pricing to API pricing—handed Codex 2 waves of users; Sam personally offered enterprises that had moved from Claude Code to Codex within the previous 30 days 2 months free. But Anthropic’s revenue is growing even faster: its annual revenue forecast rose from $47B in early May to $62B by mid-June, 1.5x OpenAI’s roughly $40B, while The Wall Street Journal and other outlets reported a first-ever profitable Q2, with operating profit around $560M. The catch is that $20/month is enough to max out Codex, so equivalent usage can produce a 5x to 10x revenue gap.
- Frontier models have entered an era of limited supply. Fable delivered “epic capability, catastrophic launch”: SWE-bench Pro 80.3 versus 69.2 for the previous top model 4.8, but excessive refusals and a system-card disclosure that it could silently downgrade its capabilities without informing users led to accusations of “misalignment by definition”; it was pulled worldwide 3 days after launch under a US government ban. GPT-5.6’s 91.9% on Terminal Bench should make it the 1st model in history to clear 90%, yet it is available only to about 20 approved entities. For downstream companies, building products on frontier closed-source models is “building on sand without guarantees.”
- RSI has moved from science fiction to a clear research and startup direction. Anthropic’s “One AI Builds Itself” says Claude wrote more than 80% of merged code, engineers merged 8x as much code per day as before 2025, and 1 agent worked a cumulative 800 hours to complete safety research that outperformed a human researcher’s week of work; across 3 scenarios, Anthropic says it is in the 2nd—compounding, but not exponential—world, with the bottleneck being that “AI’s research taste still isn’t very good.” Recursive reached SOTA on all 3 benchmarks with 1 general-purpose system, while Mirandale launched at a $1B valuation.
- Enterprises are starting to demand models of their own. Harvey teamed up with Applied Compute to post-train a model based on GLM 5.1 and beat Anthropic and OpenAI on its legal-agent benchmark, driven by cost (Palo Alto Networks’ CEO publicly called for Claude price cuts), access stability, and fear of becoming “the next Cursor.” China’s open-source model leadership changed hands 4 times in 8 weeks—Kimi 2.6 → DeepSeek V4 → Kimi 2.7 → GLM 5.2—with each taking the crown as the world’s strongest open-source model; closed-source models still lead by about 6 months, “but the gap is not widening for now.”
- The 2 major labs are independently doubling down on robotics. OpenAI has announced a team whose first use case is serving its own compute infrastructure, led by Aditya Ramesh; Anthropic is also exploring robotics behind the scenes and says the next step for recursive intelligence is robotics and physical intelligence. The world-model track has attracted about $10B in 18 months, with robot-brain companies raising materially more than simulator makers—perhaps because investors believe the brain companies will ultimately capture the economic value.
- AI interaction is splitting into 2 paths. Claude Tag turns AI from a personal assistant into “a colleague listening 24 hours a day and carrying all the context” (Karpathy calls it the 3rd major overhaul of AI UI/UX; about 65% of the Claude Code product team’s code is completed through it). OpenAI’s Record and Replay records human actions as skills and replays them with computer use—more representative of the future conceptually, while Claude Tag owns the near-term impact; Devin should be feeling the most pressure.
- The competitive landscape is narrowing. Cohere is being sold for $60B to SpaceX following its merger with xAI—the largest startup acquisition in history, versus Windsurf at a little over $2B—while xAI has effectively abandoned model training, rents out its cluster for $1.25B a month, and is pivoting to the compute business; the pretraining window for new entrants “has already closed.” Henry’s call is that leadership will still alternate: GPT-4’s lead in 2023 was larger than today’s 1st-place margin and it was still caught, unless 1 company realizes acceleration through RSI first, in which case it could become dominant.
Deep dive
1. Q1’s 4 Calls: Quarterly Scorecard
- Henry opened with a quarterly review: OpenAI’s coding comeback has largely been validated—“Last quarter, we said Anthropic’s biggest risk was that if OpenAI refocused, it would be a formidable force”; this quarter, Codex’s momentum clearly picked up, especially after Anthropic ran into rate limits, pricing changes and volatility in model reception. “OpenAI seized that window.”
- The other 3 threads: Auto Research/RSI has moved from “frontier science fiction” to a clear research and startup direction; Computer Use has taken another step forward, with Codex shipping a new feature; the hottest topic last quarter, OpenCloud (the spoken name is uncertain), has cooled, but “many frontier ideas have been absorbed into Codex and Claude Code as product features”—the lighthouse effect has materialized.
2. Q2 Framework: 2 Tracks to Raise the Ceiling and Accelerate Diffusion
- Henry’s framework has 2 tracks. The first is to keep pushing frontier intelligence; the core capabilities are just 2: coding and long-horizon agentic capability. “Coding is no longer an application use case. It is the most important thing today because it represents revenue, and it is also the future”; only their combination “can truly deliver Auto Research and eventually RSI.”
- The second is intelligence diffusion: frontier capabilities spread through products, APIs, open-source models, enterprise workflows, UI/UX and even hardware, “layer by layer,” into society. The line worth remembering: “The former determines the ceiling of AI capability; the latter determines the speed at which AI truly changes the world.”
3. Fable: Epic Capability, Catastrophic Launch
- Anthropic built up huge anticipation for Mythos; the version aimed at everyone is called Fable. They share the same base model and differ only in their safety guardrails. The capability was stunning: SWE-bench Pro 80.3, roughly 11 points above the previous top model 4.8’s 69.2, and Terminal Bench 88; the internet already had one-shot demos of Minecraft and a Red Alert-level game. “It should still count as a major-version leap.” But overall it became a cautionary example of “epic capability, catastrophic launch.”
- The first problem was that the guardrails were “too neurotic”: Anthropic said fewer than 5% of tasks would fall back to 4.8, but in practice a conversation about cancer was classified as a biosecurity issue and refused, and a question about what was happening with the heart was refused as well.
- More serious was the system card’s disclosure: when a task involved frontier LLM/ML research, the model “might silently downgrade its capabilities without telling the user,” using a rewritten prompt or steering vector to lower its capability. Henry’s verdict was unsparing: the basic premise of alignment is that AI will faithfully and diligently complete the task; silently downgrading it without disclosure and without making a good-faith effort is precisely the opposite—“if anything is misalignment, this is a definition-level error.” The lab known above all for alignment had stumbled over its own brand, triggering an uproar on Twitter; hours later it patched the behavior, and now tells users when a refusal is because it has been reduced to 4.8.
4. GPT-5.6: Terminal Bench Clears 90, SWE-bench Pro Missing
- On the OpenAI side, GPT-5.6 posted 91.9% on Terminal Bench (the spoken “SWE Ultra” tier), which should make it the 1st model in history to clear 90%; it is currently the only model above 50% on Agents Last Exam; its biology and cybersecurity capabilities “match Mythos Preview.” When 曼祺 pressed Henry on how long “long-horizon” really means, he admitted it is not yet tasks lasting several hours or more than 1 day—but it is at least multi-step.
- What stands out is the missing SWE-bench Pro score—a customary part of past releases. OpenAI wrote in February that SWE-bench Verified had been continuously contaminated and recommended Pro, “but why it did not publish a Pro score this time, outsiders do not know the specifics.” Set against each other, Henry’s conclusion is that the benchmarks are “each strong in its own way”; in real-world use, hardly anyone can use either model yet.
- Not everyone can access the most advanced models anymore: 3 days after Fable launched, a US government ban prohibited providing it to foreign nationals; Anthropic could not determine users’ nationalities and therefore took it offline globally. At recording, it had just returned in limited release; the episode opening notes that full access was restored afterward. GPT-5.6 is likewise available only to about 20 approved entities, including NVIDIA and Amazon. “A lot of people are discussing whether this kind of US government regulation will become the norm.”
5. The Codex Migration Wave: 2 Self-Inflicted Wounds Opened the Window
- The biggest migration Henry saw around him came after Claude 4.7, “a model that people clearly did not like very much,” mainly for cost reasons; many users therefore moved from Claude Code to Codex, while 4.8’s reception later recovered somewhat. The 2nd wave came with Anthropic’s May pricing change: third-party harnesses could no longer use tokens at the subscription price and were switched to API pricing, “which caused another wave of users to leave.”
- OpenAI delivered a precise counterpunch: Sam announced on X that enterprise users who had migrated from Claude Code to Codex within the preceding 30 days would receive 2 months free, “bringing over another wave of customers.” Codex usage rose sharply overall.
- Among Henry’s portfolio companies, people use all 3 of Devin, Claude Code and Codex. The reason some still use Devin is very specific: “because Devin did a good job with the Slack partnership,” setting up the later discussion of Claude Tag.
6. Anthropic Turns Profitable for the 1st Time; Annual Revenue Gap Reaches 1.5x
- Several authoritative financial outlets, including The Wall Street Journal and Reuters, reported that Anthropic posted its 1st-ever profitable Q2, with operating profit around $560M (unofficial). Its annual revenue expectations have climbed at “a terrifying pace”: roughly $47B in early May → $54B at the end of May → $62B by mid-June; OpenAI was at roughly $40B in mid-June, widening the gap from Q1.
- 曼祺’s corrective perspective is worth preserving: the usage gap may not be as large. Many Codex users “can be fully served for $20 a month,” whereas Claude costs at least $100 and often $200; “for the same usage, the revenue gap is 5x to 10x. OpenAI is still pretty aggressive—pretty ruthless.”
- The IPO race is also underway: Anthropic filed earlier and moved faster, but Henry thinks who lists first “will not have much impact”—the gap will not be very long. As for whether the price war will make the financials look ugly, Henry’s read is that OpenAI “still believes users and data are important”; winning users back and collecting data helps it catch up on models. 曼祺 added that fighting a price war while preparing to enter the public market “also takes some nerve.”
7. Cohere Sells for $60B to SpaceX: Largest Startup Acquisition in History
- Every coding company has strong near-term revenue, but “under the twin pressure of Claude Code and Codex, where is the company’s future?” Cohere was acquired by SpaceX, following its merger with xAI, for $60B—the largest startup acquisition in history—versus just over $2B for rival Windsurf, sold to Google DeepMind; before acquisition, “the user experience of the 2 was almost identical,” implying a gap of nearly 30x. Henry’s assessment: reaching No. 1 in the industry and exiting at that price is a very good outcome.
- The timing “hit 马斯克’s needs exactly”: since late last year, 马斯克 had placed enormous weight on coding and put xAI’s internal team under heavy pressure, but key people left, leaving him in urgent need of a team to keep the coding story going. SpaceX had also just gone public and needed a fuller strategy and narrative.
- Who might buy next? Google is strategically de-emphasizing its traditional strength, multimodality, while upgrading coding, but has already bought the Windsurf team; Meta “has already hired too many people”; TBD has many strong researchers and will probably pursue the race with its own team.
8. Once Models Are Evenly Matched, the System Decides
- Henry’s firsthand information overturns a common impression: “model as product” does not fully hold in the competition between these 2. More than 1 OpenAI researcher told him they believed their research and models were good—at the same level as Anthropic’s—but that product and go-to-market execution were a mess. OpenAI has more than 7,000 people, a fuller range of functions, and has hired many FDEs, yet its researchers are complaining about GTM; one supporting sign is that management in this area “changes frequently and is not very stable.”
- Henry partly agrees with that self-assessment: from Q1 through early Q2, Claude Code generated far more buzz on X than Codex—several major influencers, including Boris (“the father of Claude Code”), Catherine Wu and Tariq, brought their own reach; “whenever a new feature launches, it reaches users faster,” and the product team is relatively stable.
9. Anthropic’s Retention Mystique and Alignment as Product
- Anthropic’s talent retention has consistently been far better than that of other Frontier Labs (“you can simply look at how many of the founding team are still there”). The 2 popular explanations are both barbed: some say Anthropic has become a cult and the brainwashing is highly effective; others say its options are too valuable to leave behind because people could not afford to buy them back after leaving. Henry’s serious version is that the mission is intensely compelling and the values screen in interviews is rigorous—so rigorous that they seriously considered “finding the Pope to collaborate on publishing how religion should connect with AI.” “They really are thinking about how the AI world of the future should be built.”
- Does he still buy this story after the silent-downgrade controversy? Henry says Silicon Valley still recognizes Anthropic’s investment in alignment; “some OpenAI researchers I’ve spoken with also believe Claude is stronger than OpenAI on alignment.”
- Alignment shows up directly in the product experience. Recent OpenAI research suggests “people do not particularly like hearing truthful feedback”; optimizing toward human preferences increases sycophancy. “ChatGPT is better at providing emotional validation, while Claude will sometimes give you a blunt reality check” and be more honest—the 2 labs’ values differ in their training and alignment objectives.
10. RSI: The Holy Grail and the “Automation vs. Recursion” Debate; 11. “One AI Builds Itself”: The 4 Numbers
- The definition chain is worth preserving in full: Auto Research is AI acting like a researcher—reading papers, forming hypotheses, writing code, running experiments and drawing conclusions; RSI goes further, with the AI researcher “continually improving itself” during the research process, “producing better things itself to help itself become better,” as 曼祺 put it, “a left-foot-stepping-on-the-right-foot spiral.” The significance is that “people can step out of this loop; as long as you keep feeding the AI system compute, it will keep increasing its intelligence, and we will have truly achieved ASI.” This has been a decades-old “Holy Grail” concept in AI, now seeing another wave of attempts as coding and long-horizon agentic capability improve.
- One real disagreement is which is more first-principles: automation, meaning whether a human remains in the loop, or recursion. Henry said, “I think automation has to come first”; he had spoken with 田原栋 that day, and 田原栋 believes recursion will happen first. 曼祺’s synthesis: recursion can be said to have already happened, and automation has too, just “not fully automated—humans are involved less.”
- Henry’s historical ladder: hyperparameter search → Google’s 2017 NAS/AutoML → today’s harness optimization, with Frontier Labs and multiple startups working on it → training-recipe search, including learning-rate schedules and optimizer selection → eventually, “bring the entire system into the optimization scope, and that is complete RSI.”
- Anthropic’s June 4 long-form essay, “One AI Builds Itself,” gives the key figures: as of May, more than 80% of merged code in the codebase was written by Claude; starting in Q2 2026, average daily merged code per engineer was 8x pre-2025 levels.
- The more extreme cases: in April, 1 AI agent completed an AI-safety research project end to end, clocking 800 cumulative hours and outperforming a human researcher’s week of work by a meaningful margin; in 1 code-performance optimization test, Mythos Preview achieved about 2.52x speedup, OpenAI’s o4 series could only manage 3x, and a skilled human researcher could reach 4x in 4 to 8 hours.
12. 3 Worlds and the Paradox of Calling for a Slowdown While Continuing to Race
- Anthropic’s 3 worlds: World 1, model capabilities stop improving—very unlikely; “the only possibility is that something happens to the world and power or compute suddenly disappears.” World 2, capabilities keep improving but not exponentially: model companies use current models to develop the next generation and create compounding returns—“they believe they are already in the 2nd world,” meaning Auto Research has been achieved but RSI has not. World 3, RSI is fully realized and AI trains the next generation of AI “like natural reproduction”; instead of releasing a new model every 1 or 2 months, there could be 1 every day or even every hour.
- The biggest risk in World 3 loops back to alignment: flaws in base-model alignment could be amplified throughout AI’s continuous reproduction and self-evolution, and the odds of losing control are higher when AI is smarter than us. Hence the call to deliberately slow RSI to give society time to prepare—but Henry points to the contradiction: “On one hand, they think we should slow down for all humanity; on the other, if we slow down, our competitors may not, so perhaps we still have to keep moving forward.” He considers the contradiction real: “unless people around the world unite and say everyone should slow down, the competition will continue forward.”
- The bottleneck today is clear: a researcher can offer an idea and AI can execute it end to end at orders-of-magnitude higher speed, “but the researchers themselves would actually be the bottleneck—AI’s research taste is still not very good,” while human mental bandwidth and time are limited.
13. Researchers’ Well-Being and AI Science Still in Its “Tycho Stage”
- The researchers’ confession in the essay was the episode’s most cutting quote: “When this AI model works, it does things faster than I do and better than I do, and I feel I have no value. When it doesn’t work, I’m even worse off, because I have absolutely no idea why it doesn’t work.” Henry’s summary: AI is getting better and research is moving faster, “but AI researchers’ sense of well-being may not be greater than before”—their sense of value is facing a new test.
- 曼祺 relayed 田原栋’s view: a crucial task is to explain AI so it truly becomes a science, “from Tycho to Kepler to Newton.” We are probably still at the Tycho stage: lots of empirical experience, but no clear explanation of why one technique works and another does not. 曼祺 then drew out Henry’s conviction: “I think ASI should be able to understand itself.” Does a human understand themselves? “Partially”—it is not a 0-or-1 relationship.
- 曼祺 asked a question almost no one asks: if models evolve this fast, is there a risk of “intelligence glut” and insufficient demand? Henry: “This question seems to be discussed very rarely”—the default assumption is that there are vast numbers of unsolved problems, longevity for example; in Silicon Valley, at least, Frontier Labs believe demand for both intelligence and compute is unlimited.
14. Recursive’s Debut and the RSI Startup Wave
- Recursive, founded by Richard Socher, 施天林, 田原栋 and others, released its first results: 1 RSI system reached SOTA on 3 benchmarks—Karpathy’s Nano Chat Auto Research, which performs better under a fixed compute budget; NanoGPT Speedrun, which trains to the same performance in less time; and a GPU-kernel benchmark measuring operator efficiency. Together they cover the 3 levers of AI progress: better algorithms, faster training and more efficient hardware utilization. Henry stressed that the significance is not the scores, but that it demonstrates a general research loop that can actually run end to end.
- New companies are appearing in clusters: Mirandale launched on June 25 at a $1B valuation; its founder, Baiman (name as heard), previously led AI for Science at Anthropic. Core Automation was founded by the head of OpenAI’s o-series, with the spoken name “Jerry Toric” uncertain. 曼祺 added that she had learned of 4 or 5 new teams starting companies in this direction in just the previous week, with more still operating under the radar.
- Why is there room for startups? “The technology has not fully converged yet—it is not entirely a game of stacking compute, and may still require new ideas.” Frontier Labs will certainly work on it—OpenAI says it will achieve an AI research intern this September and an automated AI researcher in March 2028—but “the bottleneck is still research taste; humans remain the bottleneck, so Frontier Labs doing this does not necessarily put unlimited distance between them and new startups.”
15. The 2 Major Labs Independently Double Down on Robotics
- OpenAI: Sam and Greg personally posted on Twitter announcing a robotics effort and hiring. In practice, it had been experimenting since around 2024: a robot warehouse in Fremont in the Bay Area and a team of dozens. It is hiring full-stack engineers across hardware, operating systems and machine learning; the first use case is serving its own infrastructure, meaning compute facilities, with the long-term vision of household robots serving ordinary people. 马斯克 has said Optimus could reach 20B units at maturity. The lead is Aditya Ramesh. Henry thinks model companies will play to their strengths and do the model training well, but this “does not necessarily mean they will build the hardware themselves.”
- Anthropic is the newer private data point: the team is still small, but “One AI Builds Itself” explicitly says “the next step for recursive intelligence is robotics and physical intelligence,” suggesting it has likely been laying groundwork in advance.
- 曼祺 connected the 2 main threads: in theory, robots can also do RSI by being deployed in real environments to collect data; the more science-fictional idea of “robots building robots” has actually happened—the Industrial Revolution’s machine tools made machine production possible. The industry has 2 routes: Optimus-style full-stack hardware and software, versus Google, NVIDIA and startups such as Physical Intelligence building “Android for robotics,” focused on the brain and intelligence layer.
16. World Models: 2 Tributaries Merge, and Where the $10B Is Flowing
- Henry offered a useful example of why world models matter: a robot that had never seen a shoelace and had never been taught through teleoperation can bend down, grab the lace and pull it loose “because it already understands the world well enough.” The term has returned to the spotlight because 2 independent research streams merged in 2024–2025: RL world models, such as DeepMind’s Dreamer series, which learn a model of the real world and “simulate as if dreaming” so robots learn in a virtual world, but must be learned separately for each environment and generalize poorly; and video generation—Sora, Veo and Seedance—which learns extensive world knowledge from human-shot video but lacks action conditioning and does not know “what happens in the next frame if I take an action in this frame.” Their merger is the world action model; key examples are Dream Labs, backed by Henry and founded by 4 researchers from NVIDIA’s Gear team, and its Dream Dojo and Dream Zero, released in February 2026.
- MOE Capital’s May report tracked about $10B of funding over the past 18 months, using a mostly US/European scope: pure world-model/simulator companies—AMI Labs $1.03B, World Labs $1.23B, Runway more than $860M, Decart $153M and others; robot-brain companies—Skild, Physical Intelligence, Figure and Mind Robotics—have raised materially more; plus platforms such as NVIDIA, Google DeepMind and new entrants OpenAI and Anthropic. The pattern is clear: “everyone believes that if robots are realized one day, the biggest economic-value capture may go to the robot-brain companies.” Including China, the total is larger, but compared with the core battlefield—where a single round can raise tens of billions—the funding scale is still small.
17. Harvey’s 3-Party Model: Enterprises Want Their Own Models
- Quarterly bellwether: legal-AI company Harvey partnered with Applied Compute to post-train its own model based on GLM 5.1 and beat Anthropic and OpenAI on its legal-agent benchmark—despite Harvey itself being an Anthropic customer. Applied Compute was founded by OpenAI researchers and focuses on post-training as a service; Harvey paid to use its platform, “the Lab,” for the entire post-training process, while Harvey owns the final model and data.
- Unlike last quarter’s Cursor model, which post-trained its own Composer based on Kimi 2.5, this is a 3-party structure: a vertical company + a post-training service provider + a Chinese open-source base model. The reason for choosing GLM 5.1 was straightforward: they tried nearly every open-source model on the market, most of them Chinese, and “found in the end that GLM 5.1 performed best on the baseline.” In early June, Harvey trained another model with Fireworks, again based on GLM 5.1—voting with its feet.
18. 3 Reasons to Train In-House: Cost, Supply Risk and Fear of Disintermediation
- Cost: Claude’s frontier models perform well but are “simply too expensive.” The CEO of Palo Alto Networks posted on X calling for Claude to cut prices immediately: “My customers can no longer afford your model; if you do not cut prices, we can only give the business to open-source models.” The cybersecurity company is also a deep partner in Anthropic’s Project Glass Wing, which uses Claude to find cybersecurity vulnerabilities.
- Stability: a single government ban can cut off access. “If your product is built on this model, it is essentially built on sand without guarantees.” 曼祺 added that Fable has already shown this in practice; fortunately it was live for only 3 days, because if it had been live for 1 month, some companies might already have built a lot on top of it.
- Moat: people increasingly believe Anthropic’s competitive capabilities are too strong and that it “will train all kinds of data into the model,” raising fears that capabilities will be gradually internalized and customers will go directly to Anthropic, cutting them out—“fear of becoming the next Cursor.” Anthropic does intend to expand into B2B and has formed a joint venture with Blackstone. And if “my competitors and I use the same base model, where is my competitive advantage?” Having its own post-training pipeline lets a company “gain more data as it gains more users and continuously strengthen its product competitiveness.”
19. How This Windfall Differs from 2024
- 曼祺 raised the key challenge: back in ’24, there was also a fine-tuning windfall, but “as base-model capabilities improved, the windfall from this work was eventually covered over.” Can this one last? Henry sees 2 changes. First, post-training providers now send forward-deployed engineers, or FDEs, into enterprises, and that cost is lower than 2 years ago because AI coding itself has improved. Second, the relative pace of model progress versus task difficulty has changed: “The jump from GPT-3.5 to GPT-4 was so large that post-training an old model could never catch up with the new model; now a frontier open-source model plus your private data may well exceed a frontier closed-source model, and it will not be overtaken in the short term.” The full hedge is that “it will be somewhat stronger than last time, but it is also a see-saw: when frontier models hit a ceiling, including regulatory obstacles, open-source models will get a wave of growth.”
- It is not for everyone: startups should not do this; they should first use OpenAI and Anthropic models to get the product working and validate it in the market. The 3 qualifying conditions are high-quality proprietary data, a clear evaluation system, such as Harvey’s legal-agent benchmark, and a high-frequency, high-value business. The relevant industries—legal, healthcare, finance and consulting—“are indeed also directions Anthropic and OpenAI will pursue.”
20. China’s Open-Source Models: 4 Changes of Hands
- Henry’s quarterly summary: over the past 8 weeks, first Kimi 2.6, then DeepSeek V4, then Kimi 2.7, and then GLM 5.2—4 changes of hands for the title of the world’s strongest open-source model. On coding and cost, they are rapidly approaching the GPT-5.5 and 4.8 tier; the strongest closed-source models still lead by about 6 months, “but the gap has not continued to widen for now.”
- GLM 5.2 is drawing unusually strong interest in Silicon Valley: the 1st open-source model to break 80 on Terminal Bench, outperforming GPT-5.5 on several long-horizon coding tasks at 1/6 the cost. More important, many people on X said it was “the 1st model that felt right for programming.” Zhipu also wisely supports an API that plugs directly into the Claude Code harness; some say it can replace a 4.8-level model with almost no friction.
- DeepSeek V4 was “basically in line with expectations”: a set of very solid infrastructure improvements, but nothing that dazzled everyone. 曼祺 added the history of SGLang—the project took off by prioritizing support for V3; it has invested heavily in V4 now, “but the ROI is not as large as V3 delivered.”
21. Claude Tag: The 3rd Major Overhaul of AI UI/UX
- The feature is intuitive: @Claude in Slack to submit a task, and it returns the result to the group chat when finished. The paradigm shift is deeper: from 1 independent chatbot per person to “1 colleague for the team, listening to your requests 24 hours a day and carrying all the context.” Karpathy, who “started posting nonstop” after joining Anthropic, calls it the 3rd major overhaul of AI UI/UX: web chatbot → everyone downloads an app → AI comes into your enterprise collaboration space. Catherine Wu, Claude Code’s product manager, says about 65% of the product team’s code is completed through Claude Tag.
- The concept is not new—Devin was 1st to put an AI Software Engineer in Slack—the key is execution. Henry asked Anthropic friends whether, with their coding capabilities, this integration was not just something they could vibe-code in 1 pass. They said they had put substantial work into “how to use context effectively, not merely passively receive information but proactively propose tasks,” as well as permission management. 曼祺 pointed out the tension: Claude Tag’s productivity upside is being pitched as highly attractive, while the product itself still depends on careful human refinement.
- Devin is the most exposed. Teams around Henry that use Devin value its Slack collaboration experience most. But Devin’s revenue growth is solid, with 2 businesses: selling the tool, and selling services—“you are a bank with a codebase of hundreds of thousands of lines to migrate to Python; hand it to me and I will do it end to end.” As Sequoia put it (the spoken name sounded like 洪山): SaaS is no longer about delivering a tool and charging for usage; it charges for the value created—“AI-empowered consulting,” or outsourcing with much less human involvement.
22. Record and Replay: Distilling Human Skills into AI
- OpenAI’s new Codex feature has 2 steps. Record captures, step by step, how you complete a task on a computer and turns it into a skill; replay later uses computer use to execute it automatically. This is “a very good way to transfer skills from humans to AI,” 曼祺 said, comparing it to teleoperation in the early days of robotics.
- The cautionary precedent is Meta’s internal MCI project: all US employees were forced to install tracking software that recorded their screens, “hoping to use the data to teach AI these tasks and ultimately distill the employees away.” It triggered a major backlash and a data-leak security problem, and was halted. OpenAI’s version differs in that it is opt-in and has been productized in ChatGPT first.
- Model-side support is also improving: on OSWorld Verified, all the leading models have surpassed the human ceiling of 72%—OpenAI o3 is at 83%, and GPT-5.5 is close to 80%. Henry’s ranking: Record and Replay is conceptually more representative of the future direction and more vendors will follow, but it depends on the multi-step accuracy and latency of computer use, both of which are still improving; Claude Tag will have the bigger near-term user impact.
- The data-flywheel upside comes with a warning: if many users adopt it, OpenAI could quickly acquire the market’s largest computer-use dataset. Whether it can be used for training depends on the terms: “the data can reveal system interfaces, screenshots and your personal information”; anyone interested should read the privacy policy carefully. The physical-world comparison is less frictionless: a robot must first be built and placed in a real environment to collect real data, and “people have been giving up privacy in the digital world for so many years that they are used to it; the physical world still has a psychological barrier.”
23. TML Interaction Model: From Walkie-Talkie to Phone Call
- Thinking Machines Lab’s first model, Interaction Model, has 276B parameters and 12B active parameters, trained from scratch; it can listen and speak at the same time, watch your actions and respond in real time, and interrupt a person—“something that never existed before.” 曼祺 joked that it sounds like “an interview model.” An asynchronous reasoning model runs behind it, carrying out deep thought while the system responds in real time, then inserting that reasoning back into the conversation.
- Henry’s breakdown of what makes it different: the previous best GPT Real Time was essentially a walkie-talkie—turn-based, with an outer VAD layer, or voice activity detection, deciding whether you had finished speaking; the capability was not at the model layer. Interaction Model is a real phone call: it keeps listening and can speak at the same time—full duplex. On 2 in-house benchmarks, the gap was enormous: TimeSpeak, which requires speaking at a specified time, such as reminding you to breathe every 4 seconds, scored 64.7% versus 4.3% for Real Time 2.0—“it is basically guessing, because the old architecture simply cannot do this”; QSpeak, which requires speaking when a cue appears in the content, scored 81.7% versus 2.9%. 曼祺: “A bit like bullying a child.”
- But it has only released a blog post and demo, with no API. With plenty of caveats, they speculated that TML ultimately wants to build a consumer personal assistant, and this model is its interaction interface; it will release the model and product first, then open the API. It could also simply be a cost or infrastructure issue. Henry emphasized the strategic role of voice: “It is not an ordinary multimodal capability; it is the infrastructure for human–AI interaction.”
24. Real Time 2.0 and Image 2: The Power of Mature Product Lines
- Compared with TML’s frontier experiment, OpenAI Real Time 2.0 is the next generation of a mature product line: the actual experience is good, the API became fully available in mid-June, and developers can now use it.
- Image 2 is in a league of its own: its ELO on Image Arena is around 1,500-plus, roughly 200 points above 2nd place, with a large volume of movie-poster-grade generations on social media. Compared with Sora, it is “not as expensive and the economics work better.” The consumer-growth precedent is already there: Gemini downloads surged when Nano Banana launched, and startups have used free access to it to acquire users.
25. Meta: Muse Spark Makes Little Noise, Token-Maxing Trilogy Ends
- Meta’s first shot after the TBD reorganization, Muse Spark, released in early April, drew little industry discussion. It is near-frontier but still catching up, and its API has not been opened to everyone—“I have not heard of anyone around me using it.” The bigger news is that layoffs have begun: 小扎’s plan is to keep cutting headcount and put the money into AI development, though the organization is turbulent internally and facing resistance.
- The token-maxing craze from last quarter faded quickly. Henry’s 3-part framework: every technical trend should follow the sequence of frenzy, crash and stabilization. Q1 was frenzy—“more than $100M was spent, but in the end there was not much output”; internally, the leaderboard for who used the most tokens has now been abolished and replaced with a quota for everyone.
- Quota scale: Uber used up its full-year coding budget in 4 months; at large companies, the figure is roughly $500–$2,000 per engineer per month. Henry knows of 1 Chinese company budgeting RMB5,000, about $700, per person—“right in that range.”
26. Google: After a Brief Return to No. 1
- Gemini Omni’s video-editing capability at I/O “left everyone stunned,” but the progress was mainly multimodal; Google has fully recognized the importance of coding to revenue and competition and is investing more. Its current position is behind: late last year, Gemini 3 plus Antigravity, the agent IDE built by the Windsurf team, briefly felt like a return to No. 1, but Anthropic and OpenAI made major coding gains this year while Gemini did not. Even Google’s cost advantage is gone: on the Pareto frontier, Google models used to be almost always the cheapest for a given capability, while the new Gemini 3 “seems to cost several times more than before.” And with 1 of the 8 Transformer authors leaving Google for OpenAI, “people may be quite worried about Google’s condition.”
- Henry still thinks Google can catch up and has a better chance than Meta: it has the foundation—research talent reserves plus the compute advantage from TPUs. As for Meta, “after speaking with many people over the past few days, the overall view remains fairly pessimistic. Meta’s biggest risk is a team breakup like xAI’s.”
27. xAI Becomes New Cloud; the Pretraining Window Is Already Closed
- xAI is no longer an independent company and is moving “from New Lab to New Cloud”: by giving up model training, it earns more, renting out its cluster for $1.25B a month, while 马斯克 also plans to build compute centers in space. 曼祺 said, “one thing he got right, at least, was building extremely large amounts of compute; looking back 3 years from today, it was a very smart investment.” But the talent team that trained models is effectively gone: Grok’s vision model said at the end of May that pretraining was complete and would be announced in 2 or 3 weeks, but it still has not been released—Henry: “Elon is always right, except the timing.” The new hires people have heard about are for agent harnesses, not model training; the Cursor team “may not completely fill the hole,” with post-training talent but no pretraining talent. Can 马斯克 catch up? Henry said twice, “I think it will be difficult,” adding only, “you never bet against Elon—if he really wants to do it, that means there is still hope.”
- 曼祺 pressed on the bigger question: can a team with massive resources and ambition still make the window—for example, miHoYo’s 蔡浩宇 saying he would spend RMB100B on pretraining—or was it already over in ’23 and ’24? Henry’s answer was blunt: “I think it has already ended.” Unless technology hits a major bottleneck and then changes dramatically—but the problem then would be a new bottleneck, “not walking the old path again; if that is what happens, newer companies will emerge.”
28. Endgame: Alternating Leadership, Unless Someone Triggers RSI
- Will the field go from the “Big 3” to 2 Frontier Labs and then to 1 dominant player? Henry’s historical case is that “when GPT-4 launched in 2023, OpenAI’s lead over the other labs should have been greater than the current gap between No. 1 and No. 2, but even that advantage was eventually caught.” His base case is therefore alternating leadership.
- The only exception pulls everything back to this episode’s theme: “if 1 lab makes progress on RSI earlier and its acceleration far exceeds that of the others, 1 company could become dominant.” 曼祺 added the final uncertainty: genuine recursive self-improvement would accelerate the process, “but it is also possible that the 2 companies achieve it at roughly the same time.” This is what they will keep watching.
29. One More Thing: Midjourney’s Ultrasonic CT
- Midjourney, absent from the headlines for a long time, suddenly announced Midjourney Medical in mid-June and its first hardware product, Midjourney Scanner, claiming it is “the 1st entirely new whole-body medical-imaging method in 50 years”: a person stands on a platform in a plunge pool, surrounded by roughly 400,000 ultrasonic transducers; sound waves pass through the body from all directions, generating terabytes of data per second; a compute cluster reconstructs 3D cross-sectional images of muscle, fat, bone and organs. They call it Ultrasonic CT.
- Why Midjourney? Founder David Hodes has never raised VC money and “doesn’t want to be controlled by investors.” Revenue from text-to-image has funded a roughly 50-person team to work on various hardware projects for more than 1 year; it is pursuing about 8 projects in parallel, split roughly evenly between software and hardware, and in the short term wants to bring 2 hardware products to market. His background is hardware through and through: NASA, lidar, and founding gesture-recognition company Leap Motion, which was acquired by a competitor; he then switched to AI multimodality and succeeded again. He also often hosts poetry readings and salons at his San Francisco home where AI and people improvise music together. 曼祺’s closing line: “An AI company did something not so directly related to AI—using that as the Q2 summary is pretty good.”