Pioneers Insight Method Research Author
105: 尤洋 on DeepSeek Cloud Costs and Why 潞晨 Quit MaaS
Back to Episodes

105: 尤洋 on DeepSeek Cloud Costs and Why 潞晨 Quit MaaS

Summary

  • DeepSeek’s disclosed 24-hour measured data pushed the dispute over third-party deployments to a roughly 46x gap in output throughput, but the two datasets were not measured on a like-for-like basis. The run processed 608B input and 168B output tokens, using an average of 226.75 eight-GPU H800 nodes—about 1,814 GPUs in total—to support 20M–30M DAUs; at R1 pricing, the profit-to-cost ratio was 545%, equivalent to about an 84.5% gross margin. But V3 calls and overnight discounts mean this is only a theoretical ceiling, and the interview did not complete a like-for-like recalculation for both sides.

  • 尤洋 insists his conclusion targets small and mid-sized cloud vendors reselling DeepSeek API and overstating acceleration, not DeepSeek itself. When 潞晨 tested four eight-H800 machines, each produced roughly 300 token/s, versus DeepSeek’s disclosed 14.8K token/s; based on an earlier estimate, third-party MaaS could lose as much as RMB400M per month. He even said small and mid-sized cloud vendors could recover only RMB1 for every RMB100 invested. Faced with the official data, he declined to comment on DeepSeek’s technical details, stressing only that online mixed workloads differ from offline benchmarks.

  • 尤洋 sees MaaS’s structural deadlock as a billing mismatch: upstream GPUs are billed by time, downstream applications pay only for actual tokens, and the platform in the middle absorbs all idle-capacity risk. He estimates that supporting 100B output tokens per day may require 500B tokens of capacity. The host suggested that enough users, cross-time-zone traffic and overnight offline workloads could reduce redundancy to 2x; he still called instantaneous demand “unpredictable.” The 5x figure is only his personal estimate, but it is the key—and disputed—capacity assumption behind his loss calculation.

  • Open-source inference frameworks can raise efficiency across the industry, but may leave third-party MaaS with no chargeable technical moat. 尤洋 believes SGLang, vLLM and TensorRT are already strong enough that 99% of small and mid-sized cloud vendors will struggle to build a meaningful lead. Even further optimization may simply level the playing field again, leaving competition to price. “All losses come from machine idleness,” not from the outcome of any single impressive throughput test.

  • 潞晨 briefly launched DeepSeek API and then shut it down, redirecting resources toward private models, Post-Training and dedicated serving instances. The company announced a partnership with Huawei Ascend to provide DeepSeek API on Feb 4, then announced on Mar 1 that the service would be paused a week later. 尤洋 acknowledged that MaaS can generate revenue, but “the price is a little too high”; startups should “niche down, niche down again; focus, focus again.” A dedicated instance or appliance lets the customer take an entire machine and directly control capacity and resource costs.

  • In 尤洋’s commercial framework, MaaS is better suited to companies with exclusive model pricing power, or large cloud vendors with super apps and enormous traffic pools. He cited a report that ChatGPT-4.5 was many times better than GPT-4 while costing 500x more, calling the excess profit “a reward for innovation.” The host then reframed the comparison as GPT-4.5 versus V3 at roughly 500x the price and asked whether the market would accept it; the program did not clarify the two comparison sets. The investment implication is not whether monopoly pricing should exist, but that API resellers without model differentiation have virtually no bargaining power.

  • 潞晨’s alternative route is an “AI version of Databricks”: private deployment for large domestic enterprises and low-cost GPU cloud services for smaller overseas customers. Its target contract value is roughly RMB100K–RMB1M. Overseas, it offers H200 and even B200 capacity to customers in Southeast Asia and the Middle East; 尤洋 says GPU prices can be 70%–80% below AWS and Microsoft cloud. Open-source Colossal-AI drives customer acquisition, while the enterprise edition, compute and delivery generate revenue. The company is also testing a video model: if the results are good, it becomes a cash-flow product; if not, it can serve as infrastructure and be open-sourced.

Deep dive

1. The Dispute Began with Two Cost Ledgers, but No Like-for-Like Reconciliation Was Completed

  • The host first laid out DeepSeek’s 24-hour data: 608B input tokens, 168B output tokens, and an average of 226.75 nodes, each with eight H800s. That works out to roughly 1,814 GPUs, described as supporting 20M–30M DAUs.

  • DeepSeek calculated a 545% profit-to-cost ratio using prevailing GPU prices and API pricing, equivalent under the standard convention to an 84.5% gross margin. The host retained two qualifications: the traffic included lower-priced V3, and overnight calls received discounts, so the actual margin could be below this theoretical figure.

  • 尤洋 repeatedly drew the distinction: “I don’t really want to put myself in opposition to DeepSeek.” He said he had publicly called DeepSeek China’s best model as early as Jan 2; his actual criticism was aimed at small and mid-sized cloud vendors reselling API access while claiming inference was 10x faster than Nvidia.

  • The host did not avoid the rhetorical contradiction: 尤洋 said he was not fighting DeepSeek, yet defended closing Zhihu comments by saying, “Right now, I’m the only one fighting DeepSeek.” 尤洋 explained that this was his position “in the eyes of keyboard warriors,” not his definition of the object of debate.

2. The 300 vs. 14.8K Gap: 尤洋 Focuses on Load and Utilization

  • 潞晨 previously tested open-source deployment on four eight-H800 servers. 尤洋 put actual output at roughly 300 token/s per machine, while DeepSeek disclosed average output throughput of 14.8K token/s per node and input throughput of about 73.7K token/s. Output throughput differed by roughly 46x, but the two figures came from different test scenarios and the program did not provide a like-for-like recalculation.

  • In the sequence experiments 尤洋 showed, performance was best with 1K input and 1K output. Holding input at 1K while increasing output to 2K, 4K and 8K caused throughput to decline continuously. With 32K input and 8K output, two machines could produce only a little over 100 tokens per second.

  • His online workload profile combines translation, summarization, multi-turn dialogue and paper reading, with highly uneven input and output lengths. “Real MaaS utilization should be extremely volatile.” Even if one solution were 2x faster than vLLM or SGLang, he still believes the utilization problem would remain.

  • The host tried to map DeepSeek’s prefill/decode load balancing onto this sequence-level volatility. 尤洋 responded with “no comments,” then closed with, “Let’s just assume the SGLang or vLLM I used was terrible.” The program preserved the enormous data gap without separating the contributions of framework optimization, sequence length and business-mix differences.

3. MaaS Leaves the Gap Between Fixed GPU Costs and Usage-Based Revenue with the Platform

  • 尤洋’s industry breakdown is straightforward: IaaS providers or data centers charge by machine and time, and a server is billed even at zero utilization. Downstream MaaS customers, however, pay only for generated tokens. Applications neither care whether the machine is idle nor know whether the platform is already fully loaded.

  • He gave the example of IaaS charging RMB80K per machine, with the price unchanged even at 0% utilization. When machines are fully loaded, additional requests can cause the service to crash; when they are idle, no users means no revenue. 尤洋 summarized the model this way: “The core of AI’s middleware or compute layer is reducing idle compute,” and “all losses come from machine idleness.”

  • His capacity estimate is that a platform promising 100B output tokens per day should prepare machines capable of 500B tokens. The premise is that the platform cannot limit calls from upstream applications, because throttling would limit customer growth. 尤洋 explicitly said 5x was a personal estimate, not a precisely measured industry constant.

4. The Host Challenges 5x Redundancy; 尤洋 Still Treats Randomness as Unavoidable

  • The host relayed a rebuttal from another infrastructure startup: with enough users across the US, China and Europe, plus overnight idle capacity redirected to offline workloads, an additional 1x of capacity might be enough. DeepSeek itself also reduces inference nodes at night and shifts resources back to training.

  • 尤洋 rejected the scenario. Even if total hourly users were similar, he argued that without per-application call limits, instantaneous utilization would still be determined by uncontrollable requests. The more applications there are, the larger the actual swings; in his view, 2x redundancy is “extremely unreliable.”

  • The disagreement was not resolved with historical traffic, concurrency distributions or service-level data. The host’s peak-shaving and valley-filling mechanism and 尤洋’s 5x safety buffer remain side by side, meaning the RMB400M monthly-loss conclusion must be read together with its capacity assumption.

5. Inference Optimization Can Spread, but Diffusion Itself Erodes Platform Premiums

  • 尤洋’s benchmark is not proprietary 潞晨 technology, but the open-source stacks vLLM, SGLang and TensorRT. He said he would “bet” that 99% of small and mid-sized cloud vendors could not be materially faster than these systems; some 10x acceleration claims may instead come from parameter settings or corner cases.

  • On DeepSeek’s open-source methods such as PD disaggregation, he spoke only in general principles: “Computer systems have a word called trade-off. If you speed something up 10x, you are definitely sacrificing something.” Ideal results on offline machines do not necessarily transfer fully to mixed sequences and random online traffic.

  • He expects people over the next few months to one or two years to publish better-looking numbers using quantization, pruning and distillation, and others to announce profits from DeepSeek API. But as of early Feb 2025, he still maintained that small and mid-sized cloud vendors would struggle to make money. He also warned that services labeled “full-blooded versions” may already have undergone pruning or quantization.

6. 潞晨 Tried MaaS; the Conclusion Was That Revenue Exists but Is Not Worth Burning Cash

  • After DeepSeek’s popularity surged, 潞晨 revisited MaaS: it announced a partnership with Huawei Ascend to launch DeepSeek API on Feb 4, then announced on Mar 1 that the service would be suspended a week later. 尤洋 acknowledged that it “can indeed generate some revenue,” but the loss rate was high and the cost of obtaining that revenue was excessive.

  • This was not the first time the model had raised doubts. 尤洋 mentioned Fireworks.ai, Lepton AI and Together AI, as well as Anyscale, associated with Berkeley professor Ion Stoica. In his view, Anyscale’s abandonment of MaaS was the earliest trigger for reflection; Together AI’s main revenue also came not from MaaS, but from its compute platform and GPU rental services.

  • DeepSeek’s demand spike made the team “really want to give it a try,” but the result was close to its prior judgment. 尤洋’s resource-allocation principle is: “Start by niching down, niche down again; focus, focus again.” The priority is to expand existing core revenue, then revisit MaaS once pricing structures and GPU utilization become more mature.

  • He later narrowed his earlier statement that monthly losses could reach RMB400M: it referred to smaller platforms, which might invest RMB100 and recover only RMB1. A large company investing RMB1 might recover RMB30–RMB50. Not every provider faces the same loss rate.

7. Dedicated Instances and Appliances Return Capacity Responsibility to Customers

  • 潞晨 has more confidence in serving instances: a fully optimized machine is reserved by one customer—100 users use one machine, 200 use two, 400 use four. The platform no longer aggregates random requests from multiple customers into one resource pool, and the customer can directly see and control its capacity.

  • 尤洋 considers appliances offered by Alibaba Cloud, Volcano Engine and similar providers a more stable format because the machine is fully controlled by the customer. Compared with per-token billing, dedicated instances, appliances and private clusters make the customer responsible for the cost of reserved resources.

  • The host compared this with early cloud computing, when elastic resources faced similar skepticism but eventually became a mature market. 尤洋’s distinction is that most mature cloud customers reserve a block of resources and cannot use the service without machines. If application customers eventually rent model capacity by time period, he said, “That’s great—that’s no longer today’s MaaS business model.”

8. Without a Model Moat, MaaS Ultimately Competes Only on Price

  • 尤洋 believes large cloud providers with super apps and enormous user pools, such as Alibaba Cloud and Volcano Engine, may have lower loss rates than smaller platforms. Model developers such as DeepSeek can also optimize more deeply around their own models. Neither should be evaluated using the same cost structure as a pure third-party reseller.

  • His conclusion is that the only MaaS model that may reliably make money is something like ChatGPT: a model in a league of its own that is not immediately pulled into a homogeneous-model price war. Third parties have no model moat, and their software optimization is unlikely to surpass the leading open-source stacks. All that remains is “vicious price competition.”

  • The host asked why third-party MaaS existed in the first place if official APIs and large clouds could supply the same models. 尤洋’s answer was not mysterious: “Everyone is exploring, after all—we’re all looking for some way to survive.” Experimentation explains market entry, but does not automatically create a durable moat.

9. The “Reward for Innovation” Argument Pushes Model Pricing Power Toward Monopoly

  • 尤洋 called Google the biggest contributor to the current AGI wave, citing its acquisition of the company built around Hinton’s team, support for Google Brain, and work on Transformer, BERT, Switch Transformer and AlphaGo. His causal explanation is that the high margins of search gave the company room to pursue long-term exploration of non-core frontier topics.

  • Setting politics aside and looking at the issue commercially, he cited a report saying ChatGPT-4.5 was many times better than GPT-4 while costing 500x more, then defended OpenAI’s closed-source model and high pricing: “My model is better than everyone else’s, so I should make more money.” R&D is difficult to sustain if it cannot be converted into profit; whether the market accepts the price is a matter of “betting and accepting the result.”

  • The host then reframed the question as GPT-4.5 costing roughly 500x more than V3 and asked whether the market would accept it. A high price becomes profit only if the market accepts it. 尤洋 said the answer depends on whether the model is “substantially better,” while still describing excess profit as “a reward for its innovation.” The program did not clarify the difference between GPT-4 in 尤洋’s cited report and V3 in the host’s question.

  • When the host argued that DeepSeek’s open-source strongest model could expand downstream experimentation, 尤洋 classified that as an externality driven by a “grand ideal,” not a purely commercial act. He argued that AI would need one or two highly profitable giants in the long run, while adding that China should also have its own companies; he was not advocating an OpenAI monopoly.

10. 潞晨 Is Betting on the Full Private-Model Stack, Not Public API Resale

  • 潞晨 positions itself as a PaaS and Post-Training platform: enterprises process private data, then use reinforcement learning, distillation and SFT to build vertical models and deploy them into internal systems or their own applications. 尤洋’s analogy is that GPT-3 is “a very smart child” that still has to learn specialized knowledge.

  • The ideal workflow has enterprises continuously uploading reports, PDFs and other data, with the system automatically updating parameters and replacing the production model. Many industry applications do not need a huge model; 尤洋 says a foundation model with capabilities similar to Llama 7B, combined with “valuable private data,” may be sufficient.

  • 尤洋 has a Berkeley PhD and previously taught at the National University of Singapore, where he built a research group. He believes the team is naturally stronger in infrastructure. Large domestic customers buy the enterprise edition and deploy it on their own clusters; 潞晨 must send staff to deliver the system, but needs to control labor costs. He cited one Fortune 500 customer that bought roughly 1,000 A100 and A800 GPUs. Overseas, the company mainly uses an automated cloud platform to serve smaller customers.

  • The product boundary is three steps: process private data, complete RL, distillation and SFT, then deploy inference. The first step can rely on mature external tools; 潞晨 integrates the latter two. 尤洋 believes a sufficiently focused product with clear functionality is the right fit for a startup with limited resources.

11. Open-Source Acquisition, Low-Cost GPUs and Video Models Make Up the Next-Stage Bets

  • 潞晨 wants to replicate Databricks’ open-source-to-enterprise path: Colossal-AI stars, forks and dependents demonstrate capability and drive leads, while sales convert users into paying customers. 尤洋 also acknowledges that for large enterprises, open-source influence and proactive sales “are both important.”

  • He uses “commodity trading” to explain why website traffic is not high: an average customer pays about RMB100K per year, so 1,000 customers generate RMB100M. Customers are generally in the RMB100K–RMB1M range; the company is not competing with Microsoft, AWS and Alibaba for projects worth tens of millions of RMB.

  • The overseas platform is operated by a Singapore company and offers H200 and even B200 capacity to smaller customers in Southeast Asia and the Middle East. 尤洋 says optimization and regional cost advantages allow GPU prices to come in 70%–80% below AWS and Microsoft cloud, while Colossal-AI’s functionality and ease of use improve margins. He explicitly limits this logic to smaller customers.

  • 潞晨 is also developing a video model. 尤洋 believes that in AI infrastructure, large language models have largely exhausted what they can do, and the next major breakthrough may come from multimodality. Video can also be delivered directly as a product: “A one-minute video costs RMB10K in the market.” If the model performs well, it becomes a cash-flow product; if not, it can be used as infrastructure and open-sourced to bring more users onto that infrastructure.

  • Looking back to the company’s founding in 2021, he said GPT-3 was already hugely popular within the AI community, but had not reached the mainstream. The company chose infrastructure from the outset, where it had an edge, rather than directly building a foundation model. Newly joined National University of Singapore PhD students are adding algorithm and data capabilities. On the controversy, his final attitude was: “Whatever—this is just my view, isn’t it?” He also said he would stop responding to personal attacks and continue discussing only technical facts and business models.