Pioneers Insight Method Research Author
New Year Livestream 1: AI in 2025 and 2026, Consensus and Non-Consensus in Tech
Back to Episodes

New Year Livestream 1: AI in 2025 and 2026, Consensus and Non-Consensus in Tech

Summary

  • Enterprise AI crossed the “whether to do it” decision threshold in 2025, but budgets did not automatically flow to the most expensive, most capable general-purpose models. The boardroom question has shifted from yes or no to “How and how much?” Vertical small models, localized fine-tuning and “cocktail” stacks have gained consensus on cost, privacy and regulatory grounds. 张璐 notes that financial services, healthcare, insurance and other service industries covering over 50% of U.S. GDP offer extensive opportunities for vertical AI Agents.

  • DeepSeek rewrote the model race from a “four or five big labs” oligopoly into one where open-source models and new labs can still change the outcome. 徐皞 says Chinese models such as Kimi can produce strong results at relatively low cost, and a growing number of U.S. companies have begun using them; Thinking Machines, SSI and other New Labs show that core research can happen outside the major incumbents. “DeepSeek is only the beginning,” but it is still “too early” to say whether 2026 will deliver a GPT-3- or GPT-4-level huge moment.

  • Scaling Law is not dead; the debate has shifted to who can afford it and how to realize it through systems engineering. GPT-5’s underwhelming performance briefly led the market to conclude that pretraining had hit a wall, while Gemini later restored attention to pretraining; 徐皞 believes data cleaning, allocation, domain knowledge, cluster interconnects and fault tolerance all remain far from exhausted, making “10x room” in pretraining entirely imaginable in hindsight. 张璐’s qualification is that Scaling Law still holds, but is no longer the only path and is increasingly something only a handful of players can afford.

  • Meta’s investment thesis is stuck on a real fork: should it fill the gap in foundation models, or first turn model capability into user experience? The host said Meta reportedly acquired Manus for roughly $2B-$3B, with talks lasting just over ten days. 张璐 sees Manus as a strong team on execution, product and data, but questions why Meta would buy an application company when it needs model capability more; 徐皞 argues Meta can simply call OpenAI or Anthropic APIs and still fill the model gap in 2027.

  • If OpenAI IPOs in 2026, its billion-scale user base and user inertia give it a very large story and substantial potential, but costs, margins and retention will determine whether the valuation holds. It has advanced its for-profit structure and arrangements with Microsoft, but 张璐 says Gemini’s training costs may be under 30% of OpenAI’s, while the latest YC cohort is also relying more heavily on the Gemini API; the path to profitability remains unclear. 徐皞’s view is that OpenAI is “mainly a story”: no single business is nailed down, yet “every one of them makes me feel there is huge potential.”

  • Anthropic is the more focused winner in enterprise APIs and AI coding, but the guests explicitly rejected the idea that it has already built a moat. Very few Fortune 500 or Fortune 1000 companies use only OpenAI, 徐皞 says; even if Anthropic has not fully overtaken it, it is on the way. 张璐 believes Claude Code and Anthropic’s safety and To B positioning could keep attracting regulated-industry budgets. The problem is that switching costs remain low: “In today’s AI era, no company truly has a moat,” and an IPO must answer both the high-growth and cost-optimization questions.

  • The more monetizable application opportunity in 2026 may be vertical Agents; the bottleneck is not the demo but industry accuracy approaching “all correct.” 张璐 says companies in insurance, supply chain and other sectors have grown rapidly from early stage to tens of millions, and in some cases hundreds of millions, in revenue; she is also bullish on healthcare and space tech. 徐皞 warns that industry customers do not accept systems that are 90% or 99% correct—they want them to be “all correct.” That forces startups to build extensively around proprietary data, fine-tuning, platforms and workflows, leaving room in niches that general-purpose models cannot directly cover.

Deep dive

1. Enterprise AI’s decision has shifted from whether to do it to how much to invest

  • Looking back at 2025, 张璐 says the biggest surprise was how quickly enterprise consensus shifted. She had spent the past decade building around enterprise AI and vertical small models, but did not expect the market to acknowledge so quickly that Scaling Law is not a universal key to every problem.

  • In high-privacy, highly regulated industries, companies do not need the most expensive model for every task. They can combine small language models, localized fine-tuning and multiple deployment modes into a “cocktail” tailored to real workflows.

  • From the J.P. Morgan Healthcare Conference to Davos and then corporate budgets in Q3 and Q4, boards are no longer asking whether to apply AI—the answer is clearly yes—but “How and how much?”

  • Another change that exceeded expectations was the sharp easing of regulation. What 张璐 sees is not a single AI narrative, but overlapping acceleration across AI-native companies, healthcare, finance, space and defense technology, leaving investors “busy and exhausted” while also “extremely enthusiastic and energized.”

2. DeepSeek showed that frontier-model research is not sealed off by five giants

  • 徐皞 says that by the end of 2024, the market had largely concluded that the open-source war was over and Llama would “rule the market.” The 2025 DeepSeek moment showed that model capability need not come only from OpenAI, Anthropic, xAI, Google and Meta.

  • The deeper signal is not DeepSeek’s success alone. Open-source Chinese models such as Kimi can also achieve good results at relatively low cost, and a meaningful number of U.S. companies have begun using them. “DeepSeek is only the beginning.”

  • New research vehicles are also emerging outside the major labs: OpenAI’s former CTO Mira Murati founded Thinking Machines, Ilya Sutskever founded SSI, and a founding member of xAI has started a new venture. 徐皞 believes these teams each have their own thesis and angle, and could potentially do something similar to what OpenAI did around 2018 and 2019.

  • Expectations are not the same as a near-term catalyst. 徐皞 thinks 2026 may bring results, but cannot judge whether a GPT-3- or GPT-4-level “huge moment” will appear; 张璐 stresses that these New Labs are designed around long-term architecture and safety research, not short-term commercial breakout.

3. New Labs’ value may be ruling out wrong paths, not immediately rewriting the industry

  • A research company may go a long time without publishing a paper or may produce only a few open-source models, 张璐 says, without implying stagnation. It may be preparing for world models, 3D worlds or next-generation infrastructure over the next 3-5 years rather than replicating today’s LLMs.

  • 张璐 cites and likes one definition of AGI: “AI can do more than 90% of the work, and do it better than more than 90% of people.” By that standard, the industry remains far from AGI, while the current debate over its definition is highly confused.

  • A significant amount of research funding may indeed be “wasted” if the conclusion is simply that a particular path does not work. But 张璐 argues that failed validation still creates value for the industry and should not be judged by the standards applied to fast-growing companies in traditional VC.

  • The bar for impressing the market is also rising. “It is now hard to say that an architecture has appeared that is 10x better than what exists; 10x may simply be a quantitative parameter.” She says training costs per 1M tokens have fallen by hundreds of times in just over a year, making a pure multiple improvement increasingly unlikely to qualify as a true breakthrough.

4. Scaling Law remains powerful, but the unit of growth has expanded from parameters to the full system

  • 徐皞 is in the more optimistic camp on Scaling Law. He speculates that after GPT-5 underwhelmed, OpenAI may have shifted substantial effort away from pretraining; Gemini’s performance suggests that conclusion may have been premature, and OpenAI has reportedly refocused on pretraining.

  • Saying that the internet has been “scraped clean” does not mean the data problem is solved. Data cleaning, allocation ratios, quality assessment and the addition of safety and other domain knowledge still leave countless combinations unexplored; 徐皞 attributes part of the Gemini 2.5 and 3 breakthroughs to better data curation and use.

  • Compute is not simply a matter of wiring more cards together. Elon Musk is building one of the world’s largest data centers and may eventually operate 1M cards, but cluster topology, bandwidth, fault tolerance, how long continuous training can run before failure and why it fails remain unresolved computer-architecture problems. That is why 徐皞 considers 10x room in pretraining entirely plausible in hindsight.

  • 张璐 agrees that Scaling Law will still hold in 2026, but adds two constraints: it is no longer the only growth path and is becoming affordable only to a small number of players. Effective scaling will require coordinated optimization across compute, data and the full system, not “blindly piling on data.” Google’s real-world user feedback can also create a product feedback loop.

5. Systems companies offer another path beyond model competition

  • 张璐 divides AI companies into two groups: OpenAI and Anthropic are model-centric, while Google is system-centric, able to scale data, multiple product lines, real-user feedback, DeepMind’s new architectures and infrastructure both horizontally and vertically.

  • xAI is also pursuing a systems route. Musk is not merely training models; he is trying to build a complete hardware-to-software ecosystem and advance it with capital and execution. The value of this approach is that cost, feedback and deployment capabilities can iterate together.

  • Apple has been heavily criticized in the foundation-model race, but still controls the hardware, data gateway, application gateway and real-world AI interface. It can negotiate with OpenAI and then switch to Google without immediately owning the strongest foundation model.

  • 张璐 puts the shift in market evaluation standards plainly: the question is no longer only whose AI is “the smartest,” but also who is cheaper and who can actually deploy it. Systems advantages have therefore become an important variable beyond model rankings.

6. Meta’s Manus acquisition magnified the conflict between models and applications

  • The transaction context cited by the host is that Meta reportedly acquired Manus for $2B-$3B after only a little over ten days of talks. 张璐 calls it a good exit for Manus and praises its execution, product capabilities and data, but still asks: “What does Meta need more right now—application capability or model capability?”

  • Her disappointment with Meta stems from Llama 4: Meta shifted to the product side too early, weakening reasoning and model progress, and the result fell far short of expectations. It then went through the Scale AI deal, internal reshuffling and 杨立昆’s departure, visibly falling out of the first tier. A large company that cannot remain in the top 3 over an extended period will find that “very painful,” she says.

  • 徐皞 takes the opposite view: “Meta doesn’t need to go looking for trouble by building a large model. It should just build applications.” There is no problem with calling OpenAI or Anthropic APIs; the real issue is that users do not see a meaningful GenAI experience improvement inside Meta’s products.

  • 张璐 counters with the analogy that AI is like electricity. If Meta does not control its own electricity, it will face long-term risks to independence and survival, especially given its persistent dependence on its delicate relationship with Apple. 徐皞 agrees that long-term technical control is necessary, but stresses that a model’s “shelf life is too short”: even if it is the world’s best today, six months of sleep could send it into decline, and waiting until 2027 to fill the gap would not be too late.

  • 张璐 also acknowledges Meta’s substantial advantages. Zuckerberg has execution and gut, while the company has the ability to pivot, strong control and solid cash flow, giving it room to experiment. The key question is whether its talent bench and internal structure can stabilize more quickly.

7. OpenAI’s IPO story is large enough; the financial answers remain too soft

  • 张璐 views the progress OpenAI has made on its for-profit structure, particularly its arrangements with its largest investor Microsoft, as an important milestone. The question is whether current revenue, growth and the latest valuation can command the same multiple in the secondary market.

  • She sees pressure from several directions: GPT-5 and Gemini have changed user preferences, the latest YC cohort is relying more heavily on the Gemini API, and Gemini’s training costs may be under 30% of OpenAI’s. OpenAI still has difficulty showing a path that is “truly profitable.”

  • Given its massive costs and unclear margins, 张璐 believes OpenAI “definitely needs an IPO.” Whether the price can hold after listing will challenge both the company and the market. Anthropic has also expressed interest in going public, seeking capital-market support at a more appropriate revenue-to-valuation ratio.

  • 徐皞 calls OpenAI “mainly a story,” without intending either praise or criticism. ChatGPT may already be approaching 1B users, but user stickiness and its status as the default entry point are weakening; even he may switch from ChatGPT to Gemini this year. OpenAI’s strategy is far broader than Anthropic’s: “I haven’t seen a single thing that is nailed down, but every single thing makes me feel there is huge potential.”

8. Anthropic is closing in on the top spot in enterprise APIs through focus, but still has no moat

  • 徐皞 traces Anthropic’s enterprise breakthrough to the weekend when OpenAI’s CEO was removed and then reinstated. The turmoil made enterprises realize that OpenAI could become unavailable overnight, prompting them to seek a Plan B. Roughly 2 years and 2 months later, few Fortune 500 or Fortune 1000 companies use only OpenAI.

  • Even though Anthropic has not fully overtaken OpenAI in enterprise API call volume, 徐皞 believes it is on the way to doing so, while stopping short of claiming that the lead has already changed hands.

  • AI coding is also still in its first phase. GitHub once appeared to have won the market, then Cursor emerged, and Claude Code subsequently attracted developers away from Cursor. 张璐 believes the combination of native code models and coding agents has improved the development experience, but the market is far from settled.

  • 张璐 values Anthropic’s long-term investment in safety and To B markets. Finance, healthcare and insurance require regulation, compliance, privacy and localized deployment first, rather than an absolutely strongest model. Feedback on its financial products has been mixed, but the strategy is clear; it is also not tied to a single cloud, allowing it to take money from Google and Amazon while working with Microsoft. Before an IPO, it still needs to show that cost optimization will not come at the expense of growth.

9. Whether 2026 is the “Year of Applications” depends on whether you look at depth or breadth

  • 张璐 believes applications will continue accelerating in 2026 because the technology sector accounts for only about 10% of U.S. GDP, while finance, healthcare, insurance and other services account for more than 50%. Every industry offers opportunities to embed vertical AI Agents into workflows.

  • 徐皞’s objection is that if “Year 1” only means the appearance of interesting applications, then 2025 qualifies; if it means a mass breakout, it “definitely does not.” On breadth, the number of real applications is even lower than he expected a year ago.

  • His strongest examples are AI coding and vibe coding, including Cursor, Claude Code and Lovable for non-programmers. Enterprise customer service, such as Sierra, comes next, followed by AI browsers. Browsers are now viewed in Silicon Valley as a meaningful battleground, but remain far from entering millions of households.

  • Slow application growth is not merely a cyclical issue. ChatGPT has left a deep mark on user behavior, so a niche product must be much better to justify switching; improvements in foundation models may also replace part of an application’s functionality. Even when the model version is unchanged, its underlying capabilities and characteristics may shift, leaving developers in “a different world every day.”

10. Vertical Agents are growing revenue quickly, but the accuracy bar is far higher than for general chat

  • 张璐 cites Newfront Insurance, an insurtech company that uses AI to automate brokerage. It may not offer a full platform, but reached several hundred million dollars in revenue within 5-6 years and exited for roughly $1.5B-$2B, showing that a single vertical workflow can reach substantial scale.

  • Supply chain has yet to produce a comparable leader, but some companies have reached tens of millions of dollars in revenue. 张璐 is especially focused on healthcare and space. She believes healthcare privacy and data-silo problems have been addressed relatively well through federated learning, while the space ecosystem naturally combines data, AI and robotics.

  • JPMorgan Chase is a strong example of enterprise demand. Its CEO said the bank’s AI budget exceeds that of the other top 10 banks combined. The 3 roughly Series A companies 张璐 invests in all work with JPMorgan, and the fastest signed a large order within a few months. Those orders provide not only commercial and market validation, but also real-world feedback that helps products close the loop quickly.

  • 徐皞 offers a cooler constraint: vertical industries do not accept the error tolerance of general chat. “It is not 90% correct, and not even 99% correct—they want you to be completely correct.” Nothing in the real world is completely correct, so companies must invest heavily in domain-specific data, fine-tuning, Agents, platforms and process controls to reach higher accuracy.