100: How Silicon Valley Sees DeepSeek, with Fusion Fund’s Zhang Lu
Summary
- Zhang Lu sees DeepSeek R1’s breakout as a victory for the open-source ecosystem, not as a Chinese company independently catching up with OpenAI. From V2 and V3 through R1, DeepSeek has consistently published technical reports and intermediate experiments, allowing teams worldwide to share architectural exploration. That openness is especially scarce as OpenAI is mocked as “CloseAI” and Mistral and AlphaFold 3 have tightened access to details. For investors, if open source catches up with closed source, the most direct beneficiaries are startups that can call, distill and modify models at low cost.
- R1’s most important industry implication is not any single benchmark, but that unsupervised reinforcement learning can produce chain of thought without process labels. Zhang Lu does not see this as overturning the scaling law; rather, it suggests that after reducing the need for expensive labeling, more compute could still lift capabilities by “another order of magnitude.” Lower costs are only the result of this technical path; the real upside is that model training, inference and self-exploration have been reopened as areas for innovation.
- DeepSeek’s cost cuts triggered near-term concern about Nvidia, but Zhang Lu’s long-term view is the opposite: cheaper models will pull AI into every industry sooner and ultimately lift GPU demand. The comparison raised by the hosts was that V3 used 2,048 H800s and beat Llama 3.1 on some benchmarks, while the latter was trained on 16,000 H100s. Zhang Lu’s point is that lower per-task costs will bring AI to more use cases that previously could not afford it; this is not the end of compute demand.
- Meta will face brand pressure over “why a small company did it better,” but it may also become a long-term beneficiary of the open-source boom. DeepSeek’s disclosed V3 training cost of roughly $5.57M covers GPU hours only, excluding upfront investment, R&D and personnel, so it cannot be directly compared with Meta’s full spending. Counting only the “raw ingredients for cooking” is not the same as counting the kitchen, pots and pans. DeepSeek’s new architectural direction could help Llama iterate, while Meta can continue differentiating itself from closed-source players such as Google through the open-source ecosystem.
- The closed-source camp has not stopped, and competitive advantage is moving beyond model scale toward proprietary data, internal tools and commercial data flywheels. Anthropic is acquiring more high-quality B2B data through enterprise orders; xAI may tap 2D/3D industrial data from Tesla factories and supply chains, as well as SpaceX and Starlink. Gemini, Claude and ChatGPT each have advantages in accuracy and writing style. Zhang Lu heard from a founding member of an unnamed company that 70%—80% of its internal code is already written by proprietary AI tools; undisclosed applications are being used first to accelerate the company’s own iteration.
- The Agent thesis is right, but today’s products look more like “very junior interns” than analysts. When Zhang Lu tested Operator, its search speed was “like an old lady,” and it still fabricated information. Enterprise customers are therefore more likely to buy narrow, specialized and accurate industry Agents than a generic assistant that costs $200 per month but tries to do everything. Compared with traditional RPA, the breakthrough is lowering the user barrier to “can you chat?”; deployment still requires workflow integration, inference optimization and cost infrastructure.
- Outside AI, Tech Bio and space tech are the two areas to watch most closely, though both are being accelerated by AI as a “catalyst.” Longevity has shifted from whether people can live to 150 or 200 toward whether they can reach 100 or 120 with a clear mind and healthy body. In space, there is a possibility that over the next 3—5 years the cost of sending one person into space could fall to $50K—$100K, while launching a single satellite could fall to $10K—$20K. Zhang Lu also warned that in 2025, “uncertainty will become the only certainty”; early-stage companies will derive more of their edge from rapid iteration, industry data, commercial channels and resistance to the giants’ control of resources.
Deep dive
1. DeepSeek Went from Tech-World Dark Horse to Global Talking Point
程曼祺 described DeepSeek’s breakout as a case of a domestic bloom drawing attention abroad: it was neither a Chinese tech giant nor one of the so-called “six little tigers” of large-model startups, yet it was the first to attract intense attention in the U.S. tech community.
At Davos, 张璐 found that interest in Chinese and U.S. models and DeepSeek was no longer limited to tech professionals. Industry figures including Scale AI founder Alex discussed it publicly, while leaders outside technology began asking about it on their own.
The attention did not begin with R1. Before year-end, friends at OpenAI and Anthropic were already tracking DeepSeek’s architectural direction at NeurIPS, but they did not expect it to outperform. Interest accelerated toward the end of the year, and R1 delivered “a very big surprise.”
2. R1’s First-Order Significance Is That Open Source Can Catch Closed Source
张璐’s core assessment was: “The company has done an excellent job, but the most important thing is that this is a victory for the open-source ecosystem.” DeepSeek’s achievement reflects its own innovations, but also builds on paths that open- and closed-source teams around the world had already explored, validated or disproved.
Starting with V2 and V3, DeepSeek released more than just its models; it also published detailed technical reports and intermediate experimental results. 程曼祺 said researchers admired not only the performance, but its willingness to hand the community “how it got there.”
As OpenAI is mocked as “CloseAI,” Mistral stops open-sourcing, and DeepMind does not fully release AlphaFold 3, DeepSeek’s full disclosure of new methods has earned greater respect from U.S. peers for both its operating style and its contribution to the ecosystem.
Meta’s Yann LeCun made the same point as 张璐: this should be viewed as an open-source model surpassing some closed-source systems. 程曼祺 cited the Perplexity CEO’s view: “Once open source catches up with or even surpasses closed-source software, developers will move to open source.”
3. Unsupervised Reinforcement Learning Strengthens the Scaling Law Rather Than Invalidating It
张璐 said R1’s biggest surprise for the industry was unsupervised reinforcement learning: without relying on large volumes of process-data labels, the model can explore, think, reflect and form chain of thought on its own.
The result challenged a previous industry view, including among DeepMind personnel, that process labels were necessary for spontaneous reasoning. R1 showed that this may be possible without labeling process data.
She rejected the idea that this means the scaling law has failed: “This is actually a major boost to and validation of the scaling law.” Once the labeling bottleneck is reduced, applying more compute could still improve capabilities by another order of magnitude.
4. Lower Costs Open the Door to Enterprise ROI and Vertical Small Models
Among the B2B customers 张璐 has met, the first question is usually not whether a model is the strongest, but “ROI”: how much compute and electricity it requires, and whether total investment can generate a proportional business return. Cost remains the main obstacle to large-scale industrial adoption.
DeepSeek lowered both training and reasoning costs while allowing developers to distill smaller models. That is especially important for startups with limited capital and compute. On the day of the episode, its app had risen from fourth to first in the U.S. App Store’s free overall ranking, surpassing ChatGPT.
The vertical AI companies backed by Fusion Fund generally do not train foundation models from scratch. They call open-source models or APIs, then add fine-tuning and an industry-specific data library “like a cocktail,” continuously tuning the result into a sellable small model.
张璐 has never offered a precise timetable for open source catching up with closed source, but remains bullish on the long-term direction: “The open-source ecosystem will always be most beneficial to startups, while closed source will always be most beneficial to large enterprises.”
5. DeepSeek Rewrites the Old View That China Executes and the U.S. Innovates from Zero to One
张璐 cautiously said DeepSeek “may be the first company” to show U.S. model companies and startups at scale that Chinese teams are also innovating at the foundational architecture level, rather than simply calling new U.S. APIs, copying the Llama architecture and building applications on top.
Unlike many Chinese model companies, DeepSeek may not be focused primarily on commercialization, putting more effort into foundational structure and engineering innovation. That difference could change the perceptions of U.S. peers more than any single benchmark.
AMD announced that DeepSeek was using what the program called “three hundred X” for large-model inference. 张璐 said that while many models remain tied to CUDA and AMD is still at a disadvantage, the partnership would also bring DeepSeek attention from large companies.
6. Cheaper Models Do Not Mean Nvidia Demand Has Peaked
Asked why DeepSeek would short Nvidia, 张璐 pushed back directly: “That is still a long way from shorting Nvidia.” The next phase of AI is moving from technology into every traditional industry with large volumes of high-quality data, and cost is the main gatekeeper.
She cited a technical elasticity example: one portfolio company reduced model size by at least 4x and improved efficiency by 2.5x, together cutting costs by roughly 10x. More architectures of this kind could pull forward the timing of digital transformation across the economy.
程曼祺 retained the market’s most direct arithmetic: V3 used 2,048 H800s and beat Llama 3.1 on some benchmarks, while the latter was trained on 16,000 H100s. 张璐’s response was that the lower the unit cost, the more use cases can afford AI—and total GPU demand may therefore rise.
7. Meta Faces Perception Pressure but May Capture the Technical Dividend
On the employee jokes about Meta circulating online, 张璐 did not confirm their authenticity. She first clarified the cost basis: DeepSeek’s roughly $5.57M covers GPU hours only, not personnel, upfront investment or R&D.
Her analogy was cooking: counting only ingredients and seasoning is obviously different from including the kitchen, pots and pans. The $5.57M therefore cannot be directly compared with the full organizational cost of a Meta division.
Meta is under real pressure because Llama has long been one of the most widely used mature architectures in the open-source ecosystem, and the company wants Llama 4 to match or surpass closed-source models. A small team taking the lead creates a brand and perception challenge: why did the larger company not do it first?
At the product level, however, DeepSeek has effectively explored a new architecture for Llama, which is a long-term positive. Meta differentiates itself from closed-source giants such as Google through open source and can continue benefiting as more developers adopt open architectures.
8. Closed-Source Models Extend Their Edge into Proprietary Data and Internal Efficiency
程曼祺 asked whether the market was overestimating how close open source is to closed source and underestimating the capabilities OpenAI, Anthropic and Gemini have not yet disclosed. 张璐 said those companies are iterating at “an astonishing speed,” with absolute advantages in resources, compute and talent.
OpenAI remains the industry benchmark, but Anthropic is building a data flywheel through enterprise orders: more customers bring more diverse, higher-quality B2B data. As public consumer data approaches saturation and the problems and challenges of synthetic data receive more attention, this resource becomes increasingly important.
xAI’s differentiation is not limited to talent. It may gain access to data from Tesla vehicles, 3D factories, production scheduling, supply chains and automation, as well as rocket, satellite and 2D/3D industrial data from SpaceX and Starlink—information ordinary web crawlers cannot easily obtain.
Gemini, Claude and ChatGPT are also showing distinct strengths. 张璐’s partner believes ChatGPT is more accurate on some tasks, while Claude writes better. One unnamed company has already assigned 70%—80% of its internal code to proprietary AI tools, using undisclosed applications first to accelerate itself.
9. A Model’s “Personality” May Come from Its Architecture or Its Language Corpus
程曼祺 observed that DeepSeek sometimes produces unexpectedly poetic, science-fiction-like language, such as: “Every time you enter text, my brain… instantly lights up a vast data-nebula space.” Models now differ not only in performance, but also in their conversational temperament.
张璐 linked this style to unsupervised reinforcement learning: when a model explores, reflects and organizes answers on its own, its final mode of expression may develop an architectural “personality.”
She stressed that this was only “a very immature idea.” Whether joint training in Chinese and English brings the thought patterns and cultural context behind each language into the model will require validation across models trained on more varied corpora.
Her example was the American Revolutionary War: even when American and British textbooks are both written in English, they use different narratives for the same facts. Switching between language families could further change an answer’s concision, poetic quality and use of imagery and reference.
10. The 2025 Model Opportunity Is Smaller, Closer to Devices and Built on New Architectures
张璐 is most excited about vertical small models and AI on edge devices. Beyond phones, microphones, earbuds, desk lamps and factory sensors could all become AI interfaces, provided the models are “small and useful” while operating within extremely limited compute and battery budgets.
Qualcomm, Broadcom and HP are all exploring edge deployment. The smallest model from one Fusion Fund portfolio company was, in the program’s words, “below one billion tokens,” ran on a Raspberry Pi and was said to perform comparably to GPT-4 in the second half of last year.
Local processing addresses three issues at once: lower round-trip latency to the cloud, less power consumed by data transfer and computation, and stronger privacy if sensitive data can be processed locally. Finance, insurance and healthcare may therefore be more willing to adopt edge AI.
On architecture, “we are only just getting started.” DeepSeek modified the Attention architecture, which has been developed for years; another unnamed team made its model run more efficiently on CPUs than GPUs. 张璐 emphasized that this “is not a simple replacement,” but an expansion of industrial diversity.
11. Operator Shows the Direction of Agents and Exposes Their Current Limits
After testing Operator, 张璐’s immediate assessment was that it had moved beyond answering questions to “helping you do things”: it could search for flights, call webpages and execute workflows, but its movement speed on simple queries was “like an old lady,” far slower than doing the task manually.
More seriously, it still fabricated data. Some errors users can identify, while others they cannot. She nevertheless sees the outlook as “very promising,” because OpenAI, Salesforce and others all view Agents as the industry’s next phase.
The current capability boundary is closer to a “very junior intern” at a financial institution and cannot yet reliably perform an analyst’s job. Application companies need to solve how to embed this capability into existing workflows, rather than assume it can already own a role independently.
The longer-term vision is more aggressive: a manager who once oversaw 10 analysts might eventually need only 1 analyst, with that person calling multiple Agents to cover the remaining workload. 张璐 did not present this target as an achieved reality.
12. Enterprise Agents Will Compete on Specialization Before Generality
张璐 estimated that U.S. financial-firm interns earn $50K—$100K a year, while Operator was then available only through the $200-per-month Pro subscription. 程曼祺 mentioned Sam Altman’s statement that it would be moved to the $20-per-month Plus plan; 张璐 cautioned that “his timeline may need to be extended a little further.”
Enterprise customers are relatively unlikely to pay $200 a month for a generic Agent because accuracy and specialization remain inadequate. They are more likely to choose a specialized product that handles one clearly defined use case with sufficient precision, then compare it with Salesforce, Microsoft or an industry startup.
The fundamental difference between Agents and traditional RPA is the interaction barrier. ChatGPT spread because the question was “can you chat?” Operator likewise resembles giving verbal instructions to a personal assistant. Traditional process automation often requires integration, technical operation and a long implementation cycle.
Many vertical-Agent experiments have emerged over the past year, and the opportunity extends beyond the application layer. Fusion Fund has invested in Agent inference and infrastructure companies that help developers reduce inference costs, simplify industry integration and build usable products faster.
13. Healthcare, Finance, Insurance and Space Have the Right Data Conditions for Agents
Healthcare ranks first for 张璐. The figure she cited was that roughly 30% of data relevant to human society is related to healthcare, while perhaps only 5% is being used. The data is abundant and high quality, supporting both personalized medicine and better quality of life.
On the idea of an “AI doctor,” she remained cautious: “It is hard to say about making an AI doctor.” But Agent-assisted doctors are “absolutely foreseeable” in the near term, with clinical diagnosis, treatment, medical imaging enhancement and medical coding having accumulated years of deployment.
Finance needs simpler workflows and more effective data use. Insurance is large, under-automated and more standardized than healthcare, with fewer variables across business processes, making it suitable for more uniform vertical Agents.
Space depends on high-quality satellite data. 张璐 summarized the priority industries as needing three things at once: vast amounts of data; data with high value and quality; and a sufficiently large industry with diverse application scenarios.
14. As Privacy Risks Rise, Young Users May Be More Willing to Confide in AI
At the JPMorgan Healthcare Conference, 张璐 heard an unexpected observation: when companies offer both online therapists and AI mental-health support, most people are more willing to share sensitive information with AI, with the tendency especially pronounced among younger users.
That does not mean privacy concerns disappear. Operator can access identity, login and payment information when buying a flight, while healthcare Agents are even more sensitive. But user behavior suggests that people may sometimes trust AI more than human therapists and find it easier to open up.
张璐 explained this as a generational shift in interaction habits. Smartphones and the internet are a natural environment for one generation; AI and Agents may be the same for the next. AI pets and AI toys mean some children may begin talking with models from a very young age.
15. America’s Modular AI Ecosystem Lets Six People Win Major Enterprise Orders
程曼祺 summarized the U.S. model as a combination of open-source foundations, third-party infrastructure and vertical applications, contrasting it with domestic investors’ concern that any single layer is “too light,” lacks a moat and will be squeezed by giants. 张璐 agreed that this is a genuine difference between the two markets.
U.S. giants focus mainly on foundation models and are willing to let startups build out applications on their platforms. In highly regulated sectors such as pharmaceuticals, finance and insurance, startups can also offer on-prem deployment, local data control and more manageable partnerships; their trust level is not necessarily lower than that of giants such as Google.
She cited two examples of extreme efficiency. A six-person financial AI company had signed 3 Fortune 500 customers with orders in the tens of millions of dollars. Another company with more than 20 employees took annual revenue from zero to $20M, initially spending perhaps only a few million dollars, and had used less than half of its financing.
Mature enterprise sales systems and AI coding tools amplify this efficiency. Work that once required 10 engineers may now take 3 people with GitHub Copilot. Large enterprises outside technology are also increasing their AI investment.
16. Silicon Valley’s New Normal Is One DeepSeek Every Week, with Excitement and Pressure Rising Together
张璐 described the local venture scene as “continuously accelerating, continuously getting hyped”: “This week it is DeepSeek; next week it may be something else.” New papers, architectures and products appear every week. The excitement is real, as is the pressure of falling behind.
In 2018, Fusion Fund built an AI analyst named Ada Lovelace to compile weekly updates on papers and research institutions. Today, when traveling, 张璐 does not even need to wait for the report; partners are already sending papers, differences and key points continuously in the group chat.
She pushed back on the idea that the center of gravity remains on Sand Hill Road. The team is on University Avenue by Stanford, with Google, Facebook and Nvidia all roughly 30 minutes away. NeurIPS, the JPMorgan Healthcare Conference and GTC in March continue to increase information density.
Silicon Valley remains central, but it is no longer the only center. Roughly 60% of Fusion Fund’s companies are in Silicon Valley and 40% outside it. Boston, New York, Austin and Los Angeles are also rising. In local cafés, people are “either discussing startups or discussing investing,” and the topic almost always comes back to AI.
17. Tech Bio Recasts Longevity Around Quality of Life Rather Than Lifespan
Tech Bio refers to the combination of technology and life sciences. 张璐’s observation at the JPMorgan Healthcare Conference was that longevity is no longer mainly about whether people can live to 150 or 200, but whether they can reach 100 with a healthy body and a clear mind.
The target could extend to 120, but only with better early diagnosis, personalized treatment and continued advances in targeted therapies, immunotherapy and mRNA. The metric is shifting from lifespan alone to quality of life.
AI is more of a “catalyst” here: it accelerates both the digitalization of traditional industries and medical technology itself. Computational biology, digital biology, digital diagnostics and digital therapeutics are different points along the same direction.
18. Space Tech Enters a Three-to-Five-Year Payoff Window, While 2025’s Only Certainty Is Uncertainty
张璐 believes the relevant horizon is the next 3—5 years, not 10: the cost of sending one person into space could fall to $50K—$100K, while launching a single satellite could fall to $10K—$20K. If achieved, both launch volume and frequency would rise significantly.
One Fusion Fund portfolio company uses AI for space-traffic management and satellite-data trading and already generates more than $10M in annual revenue. Los Angeles has developed hundreds of companies around SpaceX, many founded by former employees and many operating as suppliers or strategic partners. Wildfires badly hit the city, but industrial areas were farther from the fires, and many small teams worked remotely, limiting the impact.
U.S. technology trends are still led by private companies and private capital. The government mainly provides non-dilutive funding, research grants, military orders and permitting support. The logic behind speeding up power-plant approvals was also industry-led: companies first said, “We are developing AI, and there is not enough electricity,” and policy responded.
张璐’s conclusion on 2025 was that “uncertainty will become the only certainty.” Political shifts, regional economies, financial black swans and weekly technology changes all require startups to iterate quickly. Fusion Fund uses a network of CTOs from more than 40 Fortune 500 companies to help portfolio companies validate, sell and commercialize directly, while preventing technology and resources from being fully controlled by giants.