Pioneers Insight Method Research Author
Gemini 3’s Comeback, Agent Models, and the RL Startup Opportunity
Back to Episodes

Gemini 3’s Comeback, Agent Models, and the RL Startup Opportunity

Summary

  • The significance of Gemini 3 Pro is not that it leads on any single benchmark, but that Google has finally combined long context, cost efficiency, multimodality, generation quality, and developer access into a product moment that reaches SOTA across different dimensions. The guest attributes the comeback to long-term co-design: vertical integration across TPU, model infrastructure, and applications, with Workspace and multiple hardware surfaces continuously feeding data back into the system; Google Labs then used editors, writers, and people skilled at creating viral posts and works to refine the product’s distribution. More importantly, the team says pre-training still has “no war in sight”—this is not a one-off catch-up.

  • GPT-5.2’s benchmark surge only makes sense on the cost curve: GDPval rose from 38.8% for GPT-5.1 to 77.9%, while ARC-AGI-2 climbed from 17.6% to 52.9%, even as the actual cost of equivalent tasks fell. The guest sees this as higher “intelligence per token,” not merely brute-force test-time scaling; but GDPval’s task mix and document entry points may still be toy-like, and the model can only answer how many r’s are in Garlic probabilistically—“AGI still has a long way to go.”

  • Foundation models will continue to internalize general-purpose Agent scaffolding, but will struggle to absorb the heavy tail created by proprietary enterprise workflows. Model vendors can train models on mock tool interactions to develop “their own shortcuts,” but enterprise-specific, non-public data will leave a persistent gap. Agents such as Cursor, with access to real scenarios and users, can in turn push models to improve coding, debugging, and multi-step capabilities. “Models and Agents make each other better”; the opportunity worth backing is the net-new market that continues expanding even after models absorb legacy functionality.

  • Precur is betting on a trainable tool layer: tools are no longer static APIs, but systems that remember “where they themselves have failed,” turn failure trajectories into enterprise assets, and improve autonomously. Its code-driven tool aims to reduce context pollution in long tasks; in early customers’ blind evaluations, it delivered roughly a 12% improvement over incumbent tools while reducing latency and expanding coverage, without requiring customers to replace their existing Agents. The end market is neither To C nor To B, but “To Agent”—a market served by the millions of Agents created by people through vibe coding.

  • The RL startup opportunity is shifting from “squeezing more out of data” to competing for interaction, reward, and high-fidelity environments, across three layers: RL environments, RL as a Service, and high-value vertical applications. The environment layer will become the “practice ground and exam center” for future Agents, but replicating Jira is no longer a moat; cybersecurity offense and defense, robotics, and autonomous-driving simulation require domain expertise. The service layer serves enterprises that cannot build RL in-house, while the applications layer targets drugs, trading, and science discovery. Moe Capital has invested in Preference Model, while Periodic Labs represents the third category.

  • An Agent’s model choice is first a question of cloud, discounts, and ecosystem, and only then a question of leaderboards: migrating enterprise data to another cloud often takes 1–2 years, and existing Azure and GCP relationships can outweigh short-term model leadership. Claude’s advantage is not just coding, but the closed loop formed by MCP, programmatic tool calling, Claude Code, and its developer paradigm. Real-time Agents may instead prefer low-latency models such as Gemini Flash and OpenAI mini. Qwen and other Chinese open models offer size choice and room for ablation, and teams will also post-train on open models; policy restrictions and value concerns among US enterprises still leave room for an “American DeepSeek.”

  • The next Agent capability race will center on native long-horizon reasoning, multimodality, and open models that can be post-trained—not merely another fragile layer of scaffolding. Developers need open foundations that have seen abundant Agent traces during training, can directly understand dirty PDFs, images, and mixed modalities, and can plan, act, and correct themselves across tasks spanning dozens of steps. Even if models internalize some context engineering, framework usage data will continue feeding back into model training; the guest’s view is that the serviceable market will expand faster than the portion absorbed by foundation models.

Deep dive

1. Gemini 3 Ignites Google’s Morale, Stock, and Developer On-Ramp

  • Henry and Naomi held their celebration at a San Francisco gallery, with guests mainly consisting of founders of recently hot Silicon Valley startups and Gemini 3 users. Team morale surged after the launch, and the market reaction was strong; the event celebrated both Gemini 3’s release and its user adoption.

  • AI Studio was described as undergoing “explosive growth.” Google’s new AI Studio team, formed only about a week earlier, is led by Logan Kilpatrick, who previously worked on developer experience at OpenAI. He later conducted frequent interviews with core figures such as Demis Hassabis and Jeff Dean and is now effectively Google DeepMind’s external spokesperson, pushed to the front line of Google’s developer portal.

  • 程曼祺 asked whether GPT-5.2, released the same day, was responding to Gemini 3. The explanation the guest heard was that the new version fixed several issues from GPT-5.1’s pre-training phase; OpenAI’s internal “Code Red” was also seen as a response.

2. GPT-5.2’s Real Gains Are Hidden in the Cost Curve, Not a Single Score

  • GPT-5.2 rose from 38.8% for GPT-5.1 to 77.9% on GDPval, nearly doubling the probability of matching or exceeding top human experts across 44 categories of knowledge work; ARC-AGI-2 also climbed from 17.6% to 52.9%. On the surface, this looks like an exceptionally rapid generational leap.

  • 戴韩俊 cautioned that every benchmark number embeds token and cost dimensions. If a model is allowed to think and search for longer, scores on tests such as AIME can even be pushed to 100%; the more important question is whether the score rises at the same cost, or whether achieving the same performance falls from tens of dollars to something much lower.

  • Although GPT-5.2’s list price is nominally higher, its overall cost per completed task may be lower, implying more efficient use of tokens. By contrast, the thinking traces of some open models continually “talk to themselves” and take detours, exposing substantial room to improve intelligence per token.

  • Databricks’ OfficeQA offers a reality check: GDPval’s task distribution may not match the problems customers actually care about, and enterprise workers often do not even know which document to start searching. GPT-5.2 still improves under harder, more production-like settings, but 77.9% cannot be read directly as a real-world automation rate.

3. GDPval Requantifies the AGI Vision, While Garlic Shows Capability Remains Uneven

  • Naomi linked GDPval back to OpenAI’s 2018 charter, which defined AGI as AI capable of highly automating and outperforming humans at most economically valuable work. The dataset covers important industries contributing more than 5% of GDP, samples 44 occupations across manufacturing, real estate, and other sectors, and includes 1320 tasks—an initial direct measurement of that definition.

  • No score eliminates the capability gaps: GPT-5.2’s internal codename is Garlic, but when asked how many letters r are in Garlic, it may still answer 0, 1, or 2. “AGI still has a long way to go”; real-world capability does not rise smoothly with the aggregate score.

  • Naomi also noted that OpenAI introduced GDPval, PaperBench, and SWE-Lancer this year, and that the best models reported in the initial tests for those benchmarks were Anthropic models. She respects OpenAI’s researchers for separating marketing from research: the first results on self-created standards did not favor OpenAI’s own models, which made the benchmarks more credible.

4. Foundation Models Will Absorb General Scaffolding, but the Private Enterprise Tail Remains Agent Territory

  • 戴韩俊 acknowledged that model vendors are organizing training sets by industry coverage and using mock interactions to expose foundation models to a range of tool workflows. With enough experience, models will discover “their own shortcuts,” call tools directly, and eliminate much of the Agent scaffolding previously built in external code.

  • The other side of the coin is that enterprise workflows follow a “long-tail and heavy-tail” distribution: processes, permissions, and data are highly private, and companies are unwilling to publish them. Foundation models therefore struggle to train comprehensively; expanding model coverage does not close the gap, but merely moves its boundary.

  • Cursor shows how models and applications can make each other better. Once a Coding Agent has access to users and task contexts, model vendors can specifically strengthen coding, debugging, and multi-step capabilities for it. Once ChatGPT or an API connects to search, file reading, and other tools, it is already approaching an Agent, making the boundary between models and Agents increasingly blurred.

5. Gemini 3’s Distribution Advantage Comes from Product Orchestration, Not Just Model Scores

  • In the month or two before launch, Silicon Valley kept raising expectations, with some OpenAI and Anthropic employees reportedly feeling nervous. Gemini 3 still met those expectations at an unusually high bar, becoming the moment when Google moved from “having highlights across different dimensions” to achieving SOTA across the board.

  • Google had previously showcased Gemini 1.5’s one-million-token long context, cost efficiency on the Pareto frontier, reasoning with budgets, multimodal understanding, and Nano Banana. Gemini 3’s key achievement was to “put these components together,” rather than release another champion in a single category.

  • Gemini’s native generation of interactive webpages was seen as more polished and usable than some dedicated products, making it naturally suited to viral distribution. GPT-5.2’s improvements for office and white-collar work may be highly valuable, but they are less attention-grabbing than an immediately playable interactive application; this is not a fair comparison on pure performance.

  • Bethany explained that Google Labs brings editors, chief editors, and writers into NotebookLM’s product experience and demos, while also recruiting people skilled at creating viral online posts and works. The name Nano Banana itself came when a PM working until 2 or 3 a.m. looked at the two bananas on their manicure and said, “Why not we just call it Nano Banana?”

6. Gemini 3 Suggests Pre-Training Still Has Runway, While TPUs Turn Scale Training into a System Advantage

  • Ilya Sutskever once predicted that “pre-training as we know it will end,” but after Gemini 3’s release, its co-author and Google researcher Oriol Vinyals stressed that this round of pre-training still produced breakthroughs and that there was “no war in sight”—the team is not down to its last usable round of improvements.

  • Henry therefore sees Gemini 3 as the result of a platform with further runway, rather than a one-time exhaustion of the final data dividend. At least internally, Google’s progress in pre-training has not been described as the end of an era.

  • Once training scales far enough, the bottleneck shifts from compute-bound to network-bound. TPUs were designed from the outset around a mesh architecture, while GPUs began as standalone cards and only later evolved network interconnects; this gives TPUs a potential architectural advantage in large-scale training.

7. Google’s Moat Is a Three-Layer Co-Design Loop Linking Data and Compute

  • 戴韩俊 believes Google is slow to make decisions but excels at long-term positioning. The first layer of co-design runs from TPU kernels and low-level libraries through JAX to large-model infrastructure, allowing teams to optimize jointly. OpenAI and other companies designing their own chips are pursuing the same control over the hardware-model stack.

  • The second layer connects models and applications. Google Workspace has accumulated years of Calendar, email, Docs, and other workplace data, while usage patterns from enterprise products such as Agentspace can feed back into model training. Google does not merely own workplace distribution; it owns a “train–deploy–retrain” loop.

  • The third layer connects data to hardware surfaces: cars, Assistant, Home, Project Astra glasses, and Pixel phones can collect different forms of data. 戴韩俊’s summary was that this ecosystem may simply not have reached its breakout point before—“the breakout point just happens to be now.”

8. TPUs Are Becoming Transparent and Usable, but CUDA Still Holds Deployment and Ecosystem Inertia

  • Google has long made TPUs available externally, and some of the latest versions were even reserved by outside customers first, with the remaining capacity used for internal training. Through its partnership with Google, Precur participated in Tunix, the first RL on TPU framework, and received a strategic research TPU grant.

  • For developers, both PyTorch and JAX can run on TPUs through XLA, while inference frameworks such as vLLM and SGLang are reducing the perceived migration cost. JAX is more like a numerical-computing abstraction spanning GPUs and TPUs, with Flax, Orbax, and RLax on top; Pallas is closer to CUDA’s lower-level role.

  • One guest’s reservation was that TPUs still face insufficient deployment across major CSPs, data-center retrofits, a smaller JAX user base, front-end and back-end packaging capacity, and the need to balance relationships with Broadcom and MediaTek. For now, “CUDA is still king,” especially when customers prioritize ease of use and fast GTM.

  • But if model update velocity slows and large teams become willing to work down at the instruction-set level, CUDA’s ease-of-use premium could decline. Large customers may also use TPUs as a bargaining chip against NVIDIA. The show’s joke about “buy Broadcom and TSMC” came with an explicit disclaimer that it was not investment advice, while 程曼祺 relayed a market rumor that NVIDIA was using GPU discounts to persuade Google to slow TPU promotion.

9. After the Brain–DeepMind Merger, Google Traded Research Redundancy for Stronger Delivery Discipline

  • Before the merger, Google Brain and DeepMind each had model-training teams and infrastructure, creating duplicated roadmaps and inevitable politics. The name Gemini was explained as a “twin” born from the merger of the two organizations; once the technical direction became clear, the talent density and division of labor at a large company began to compound.

  • 戴韩俊 experienced both cultures: DeepMind was more top-down, with projects such as AlphaFold and StarCraft requiring organizational discipline and even daily stand-ups; Brain had a broader spectrum and more bottom-up freedom, including some self-directed research that in hindsight was “not very useful.”

  • Language models happen to require both cultures: clear outcomes and short-term delivery, alongside research unknowns where nobody can guarantee the work will succeed. 戴韩俊’s exact phrasing was that “research needs redundancy, it needs some waste,” but as competition intensified, publication policies tightened and many results that once could have been published are now delayed or kept internal.

  • For researchers, satisfaction has also shifted from public papers toward internal visibility, compensation, and product impact. If work becomes part of Gemini and is used by people worldwide, its impact may far exceed that of being noticed by peers in a narrow field—another form of compensation for “anonymous engineering.”

10. Startups Must Either Go Deep in Verticals or Fill Immature Horizontal Layers

  • Vertical opportunities require “π-shaped talent”: people who understand biology, drugs, and clinical-trial design, also understand computer science, and can mobilize pharmaceutical companies, industry, and academic labs such as Stanford. Foundation models can supply life-science data, while vertical teams handle the last mile; the two are not zero-sum.

  • Horizontal opportunities are driven by intense enterprise anxiety. 戴韩俊 said top companies fear most of all falling from first place in the world to second. That ambition and formality of demand creates three opportunities: orchestrators that use RL to optimize Agent scheduling, data connectors represented by Composio, and the tool layer where Precur operates.

  • Explaining his departure from a large company, 戴韩俊 said startups cannot begin by assuming AGI already exists. Today’s models are still fundamentally “as much data, as much intelligence.” Even against native multimodal models, a voice startup can differentiate through lower latency and higher single-task quality; continual learning, memory, and self-evolution all remain far from solved.

11. Precur Turns Tools from Static Interfaces into Stateful Systems That Remember Failure

  • Bethany defines tools as the core way an Agent interacts with the real world. Precur’s vision is not to build another set of static APIs, but to embed intelligence in tools so that after a failed call they can learn, iterate, and become “better the more they are used” in coverage or model quality.

  • A single tool is merely one action within a stateful platform; the end state is “tools making tools,” with the same tool evolving into different states for different enterprises based on their failure trajectories. Those trajectories are no longer disposable logs, but core assets that enterprises accumulate over time.

  • The team defines its larger user base as Agents rather than people: “There used to be To C and To B; now there is To A, To Agent.” The millions of Agents created through methods such as vibe coding will be the primary callers, so the product’s interface, reliability, and optimization targets all prioritize agent-friendliness.

12. Code-Driven Tools Trade Less Context for Greater Reliability on Longer Tasks

  • Precur chose code-driven tools, aligning with the 2024 CodeAct research and Anthropic’s later programmatic tool calling (PTC): Agents take action by generating and executing code rather than relying only on static tool calls.

  • The value of this approach is not only a more flexible calling format, but also greater stability and lower context usage. In a long trajectory, every input and output accumulates; the team sees context pollution as one of the main causes of long-horizon Agent failure.

  • Early customer feedback has focused on coverage, latency, and accuracy: the tool can handle large-data tasks that incumbent systems struggle to cover, materially reduces latency in some scenarios, and outperformed existing tools by roughly 12% in blind evaluations designed by customers without Precur’s prior knowledge.

  • Because the product sits at the tool layer, customers do not need to replace their existing Agents and can plug it directly into platforms such as Microsoft Copilot. Google, Microsoft, and other unnamed Agent platforms are viewed more as ecosystem partners than as customer interfaces that must be displaced head-on.

13. Self-Evolution Is Far More Than Online RL, and It Is Not an External Memory Layer

  • The ideal reference point for continual learning is advertising and recommendation systems: every video watched and every swipe-away feeds back into the online system and prompts an update. Language models still cannot reliably achieve this kind of “better the more they are used” behavior after deployment.

  • Cursor says it is pursuing online RL by feeding whether users accept an autocomplete suggestion back into the training system as immediate feedback. 戴韩俊 believes the direction is sound, but outsiders do not know how far the implementation has actually progressed. Whether the feedback is reliable and whether it is interpreted correctly are both deployment challenges.

  • RL is not a stable algorithm, and online environments make rewards difficult to verify, requiring substantial engineering and algorithmic work. Theoretical feasibility does not imply current reliability.

  • RAG and its variants, Agent memory, and ChatGPT memory are early attempts at lifelong learning, but external memory does not change the model’s parameters. These approaches work in the short term but may not be the final answer; “self-evolution should be far more than this.”

14. Moe Capital Breaks the RL Opportunity into Environments, Services, and Specialized Applications

  • Naomi and Henry announced the launch of Moe Capital with Professor 王梦迪, co-director of Princeton’s AI acceleration and innovation center, focusing on early-stage AI infrastructure, applications, and AI for science. “Moe” refers both to Mixture of Experts and to an expert community of young researchers from OpenAI, Anthropic, xAI, and Google DeepMind who participate in investment decisions and portfolio support.

  • Naomi used “squeezing juice” to describe model growth: the juice in traditional data has mostly been extracted, so the new objects of value are the interaction and reward required by RL. Whoever can produce interactive environments that can be trained repeatedly and scored precisely may control the next generation of data infrastructure.

  • The first category, RL environments, resembles a practice ground and examination center; Moe Capital has invested in Preference Model. The second, RL as a Service, gives training capabilities to enterprises that cannot build RL in-house, using private data to develop AI specialists for sales, customer service, compliance, and other functions.

  • The third category consists of “special forces” applications in drug development, financial trading, and science discovery, with Periodic Labs cited as an example. Thinking Machines Lab’s Tinker is closer to infrastructure as a Service: it provides low-level primitives such as forward, backward, and optimization, but still requires users to understand RL themselves.

15. The Moat for RL Environments Has Shifted from Software Replication to Domain Fidelity

  • Henry recalled that an early team once used a Coding Agent over a weekend to replicate the complex Atlassian Jira and train an AI to understand engineering teams’ task-management workflows. That was the first wave of RL environments, but as more participants enter, building a commercial-software replica is no longer enough to create an advantage.

  • Teams such as Preference Model have begun shifting toward cybersecurity offense-and-defense environments. To make an environment realistic while providing researchers with suitable APIs and rewards, a team must possess both cybersecurity expertise and RL capabilities; domain fidelity is the new moat.

  • The physical world also needs environments. Autonomous driving and robotics already involve extensive simulation work, and one autonomous-driving virtual-environment company was acquired by Waymo, whose internal team was called “Sim City.” Simulators can generate rare edge cases that are scarce in the real world, especially low-frequency, high-value data such as crashes.

  • 戴韩俊 believes the boundary of “world models” is blurry and that simulation itself is not new. He previously worked on a retail generative model that used SKUs as tokens to simulate users purchasing items concurrently and their relationship to time and context. It was then called a generative model for retail; world models now have stronger rendering and expressive capabilities, but the underlying simulation idea is continuous.

16. For Enterprises Choosing Models, Cloud Lock-In and Credits Usually Come Before Performance

  • The first hard constraint Precur observed after leaving Google was the cloud: OpenAI models cannot be used directly in GCP, and Gemini is not available in Azure. Enterprise data often cannot leave its existing cloud, so “where you sit determines what you think” narrows the candidate set before leaderboards do.

  • Migrating years of data at a large enterprise typically takes 1–2 years or longer, making it a high-risk operation for a CTO. Azure’s trust and stickiness among mid-sized and large enterprises are important distribution advantages for OpenAI, while Anthropic’s presence across all three major clouds gives it a more neutral position.

  • Cloud credits, discounts, and arrangements for purchasing third-party services through the cloud also change the real cost. When companies buy services with credits rather than cash, a model only needs to be “not too bad” to become the default. Precur will try three models across different clouds rather than insist on a single foundation.

17. Claude’s Agent Advantage Comes from Coordination Across the Model, Tool Paradigm, and Developer Ecosystem

  • Claude’s long-standing popularity is not just a function of high coding benchmark scores; coding ability transfers into the ability of an Agent “brain” to act. Code is both an output and a medium for calling tools, organizing workflows, and executing multi-step tasks.

  • Bethany observed that Anthropic began visibly increasing its investment in the Agent ecosystem in October, with its blog addressing Agent developers directly. PTC lets models take action through code, while MCP sparked a wave of tool connectivity; even with MCP’s flaws, it forced other model vendors to add similar training data.

  • Claude Code has also gone well beyond “writing code for programmers”; people use it to generate n8n workflows, webpages, and PPTs. When 程曼祺 asked whether Google and OpenAI catching up in coding would leave Claude’s advantage intact, the answer was that beyond the model’s capabilities, the real differentiation includes who better understands how Agents are actually written and used.

18. Real-Time Agents Need Low-Latency Decisions More Than Deep Thought at Every Step

  • Extended thinking can demonstrate a model’s mathematical and reasoning ceiling, but it may not suit Agents that need to call tools continuously. Some developers turn off a model’s thinking mode and use prompts to preserve a small amount of reasoning, manually balancing quality against latency.

  • As a result, Gemini Flash, OpenAI mini, and smaller models may be the best choices within a given response-time budget even if their absolute scores are not the highest. The key metrics are task-level cost and intelligence per token, not the price of a single call or the maximum thinking budget.

  • Precur will test different models across the major clouds and has no strong attachment beyond customer cloud constraints. Real-time Agent selection therefore cannot focus only on the highest capability of a single universal foundation; it must also consider performance within the available response time.

19. Chinese Open Models Offer Trainable Choice, but Policy Boundaries Still Apply

  • The team praised the Qwen series: its range of sizes lets startups pick and choose and run ablations, while its reasoning capabilities may have a natural advantage in some Agent scenarios. The reservation is that benchmark gains do not necessarily transfer to real tasks, and thinking traces sometimes consume unnecessary tokens.

  • US open-model supply is temporarily weaker. The guest noted that Llama is no longer advancing as consistently as it once did and named Reflection AI, which claims to be building an “American DeepSeek.” Part of the opportunity for such companies comes from real constraints: some enterprises explicitly prohibit Chinese models, rather than rejecting them because of insufficient capability.

  • The guest does not want to see the market separated along ideological lines, but acknowledged that human alignment under different value systems will become a procurement criterion. Investment and product decisions must therefore treat policy availability as an independent variable alongside model performance.

  • On Google’s Gemma, 戴韩俊 said it showed “insufficient sincerity”: a smaller model with multimodality to support has difficulty balancing the tradeoffs, and real-world use can produce occasional low-level errors such as garbled text. Google’s open-model strategy is currently more about earning developer trust and avoiding the “CloseAI” label than challenging closed frontier models with Gemma.

20. DeepSeek-V3.02 Defined a More Sincere Form of Openness by Publishing Training Details

  • 戴韩俊 recalled that during NeurIPS, someone observed that roughly one-third of the passengers on a flight were reading the formal DeepSeek-V3.02 paper. Without Wi-Fi, the paper became material researchers were happy to use to pass the time.

  • The paper disclosed architectural improvements, failures encountered in RL, and methods for making many off-policy behaviors train more stably. Some frontier-model labs may already know these techniques, but the willingness to publish the details for everyone to discuss was itself worthy of respect.

  • This led to his distinction between degrees of openness: “Some open source is fairly fake open source, and some is more sincere open source.” Open weights can improve reputation, but publishing reusable methods and failure experience comes closer to advancing the ecosystem as a whole.

21. Good Evaluations Must Test Long-Range Interaction While Preserving Unhackable “Weird Questions”

  • 戴韩俊 said bluntly, “How high a benchmark is scored depends entirely on one’s conscience.” The longer a public test has existed, the more likely it is to suffer information leakage or hill climbing. Frontier labs still retain some restraint because if scores diverge from real behavior, users will quickly punish them.

  • Precur focuses more on multi-step, long-horizon tasks such as SWE-bench Pro and Mind2Web. Across a dozen or more actions, the context keeps growing, and the model must understand the preceding trajectory, plan the next step, and take action. Needle in a haystack tests only a single, simple task and cannot represent this kind of reasoning.

  • Sierra’s τ²-bench adds a user simulator: users change their replies based on what the Agent says, naturally creating long, multi-turn interactions. The Agent no longer answers a static question, but must keep adjusting while lacking full knowledge of the user’s and environment’s state.

  • Private “weird questions” function like interview problems. Sebastian Bubeck used code to draw a unicorn, while 戴韩俊 used a new dictionary to construct a “circuit board” and test whether a model approximates a universal Turing machine. The Generative Burrito Test asks for a half-eaten burrito with the correct filling. These tasks may not be practical, but they reduce artificially high scores produced by targeted training.

22. Multimodality and Native Long-Horizon Capabilities Will Expand the Agent Market, Not Merely Compress the Application Layer

  • What Precur most wants to see is an open model trained on abundant Agent traces and designed to support post-training. Teams must be able to change the foundation model’s behavior through post-training; if the model natively understands Agent scenarios, developers can move beyond patching limitations with prompts and scaffolding.

  • Multimodality is another clear direction. Enterprise data often consists of scanned PDFs and dirty documents, and Databricks’ use of a parser to improve OfficeQA also shows that native understanding remains inadequate. Feeding in a blank exam paper, reasoning through the questions, and outputting an image of the completed paper tests understanding, reasoning, and generation simultaneously.

  • 戴韩俊 believes humans are inherently multimodal, and Gemini has pursued this direction from Day One. Video and mixed modalities will materially increase model capacity and resource requirements. Elon Musk once said online that he wanted Grok to compete against professional League of Legends teams, which also highlights that visual understanding, action, and latency must all work together.

  • Long-horizon training is also moving from single-step RLHF, where a reward is given after one output, toward agentic RL, advance planning, and parallel thinking. Bethany’s final answer was that as more people use frameworks such as ADK, memory and tool-call data will increasingly feed back into model training. Models will internalize some downstream capabilities, but the delta from enterprise-market expansion may still exceed the portion they absorb.