Pioneers Insight Method Research Author
[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model
Back to Episodes

[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model

Summary

  • Max calls Kimi K3 a clear third-best model, while Dylan says it may be second-best for his practical use. Max says benchmark composites show a clear, stable top three including Soul 5.6 and Kimi K3, ahead of Google, Meta, and SpaceX; he says Google should feel embarrassed. K3 is still slower and officially trails Fable 5 and GPT 5.6 full, but Dylan finds it less frustrating than Opus when access is rejected or limited.
  • K3’s 2.8 trillion parameters make inference capacity an immediate bottleneck. Dylan says it cannot fit in a B200 and needs B300, GB300, or MI355X hardware for service on a single 8-way HGX system; multi-node pipeline parallelism would hurt performance. Max speculates that the roughly 10-day weights delay gives serving teams time to prepare and may let Moonshot arrange licensing and GB300 capacity with Together, Fireworks, Nebius, and CoreWeave.
  • K3’s pricing highlights attractive frontier-model economics but a constrained buyer base. Pricing rose from $0.95/$4 for the previous 2.7 Code model to $3/$15 per million input/output tokens. Max argues that if K3 is probably not operating at negative margins, similarly sized Fable charging $10/$50 would imply extraordinary margins, potentially better than SaaS. Cost-conscious users may choose GLM-5.2 or MiniMax M3, while SemiAnalysis-like users continue using GPT-5.6 and Opus; Max doubts large enterprises will adopt K3 seriously.
  • Dylan attributes the narrowed open-versus-closed gap squarely to US restrictions that limit access to Anthropic’s best models. He points to Mythos versus Fable, says he cannot use Opus and can use Sonnet only sometimes, and thinks Soul 5.6 may not be OpenAI’s largest model. He expects policy changes or another frontier step near the end of summer, while also arguing that open models may eventually reach true frontier parity.
  • Western open source has a sovereign-computing opportunity beyond cost savings. Max sees strong demand for a Western model that does not suck and expects many US enterprises to reject Chinese weights even in air-gapped systems; he also says a Chinese-model ban may be coming. Dylan says Western models must beat both bulk Chinese open models and tier-2/3 frontier-lab models, with government as the larger market. K3’s quantization during SFT—MXFP4 weights and MXFP8 activations—supports broad hardware compatibility and China’s domestic-accelerator push.
  • Harnesses, routing, and untapped adoption still support frontier-lab revenue. Max says model quality is increasingly difficult to distinguish across maximum, high, and medium effort, making features in tools such as OpenCode, Hermes, and Pi decisive. Dylan suggests outcome-based pricing could produce 95%-plus margins, and Max agrees that routers and adjustable thinking effort could do so. Max sees substantial room for new use cases across coding, audio, video, deep research, robotics, and world models; Dylan and Max believe new adopters will overwhelm any K3 substitution and prevent ARR-growth deceleration.

Deep dive

1. Kimi K3 has entered the frontier conversation

  • Max gives K3 a “clear yes” as the world’s third-best model. He says benchmark composites remain directionally useful and show a clear, stable top three including Soul 5.6 and Kimi K3, ahead of other open models as well as Google, Meta, and SpaceX. He calls Kimi’s achievement remarkable and says Google should feel embarrassed.

  • Max explicitly qualifies the ranking: K3 is still worse overall than Fable and Soul 5.6. The release blog says K3 has a noticeable user-experience gap versus Fable 5 and GPT 5.6 full. Max speculates that the self-deprecation could reflect Chinese humility or a desire to avoid US scrutiny.

  • Dylan says K3 is slow but capable enough that he has not found much complicated work it cannot handle. He calls it his practical No. 2 because Opus access can be frustrating: he is sometimes rejected or limited in the web console, deep research, and coding plan. With a paid API key, however, he says he is not being rejected in the same way.

2. A 2.8T model makes hardware part of the launch strategy

  • Dylan describes K3 as a new base model and architecture, roughly twice the size of the previous models, incorporating Kimi Delta Attention, attention residuals, and Stable Latent MoE. Its 2.8 trillion parameters do not fit in a B200; serving it on a single 8-way HGX system requires B300, GB300, or MI355X-class hardware. Pipeline parallelism across nodes is possible but would hurt performance.

  • Max stresses that his explanation for the roughly 10-day weights delay is pure speculation. One possibility is that serving-stack teams need time to make the model perform well before Moonshot releases it widely. Another is that Moonshot is discussing licensing and capacity with Together, Fireworks, Nebius, and CoreWeave, potentially including GB300 deployments.

  • The scale comparison also informs Max’s view of closed models. He says that if closed-source models really had around 10 trillion total parameters yet only matched K3, it would be “time to pack up the bags” and their stock prices might deserve to fall 50%. He instead believes K3 is probably not much smaller than leading closed models and could even be slightly larger.

3. K3’s pricing reveals superb margins but an uncertain buyer

  • Dylan compares K3’s $3/$15 per million input/output tokens with the previous 2.7 Code model’s $0.95/$4 pricing, a more-than-threefold increase. Max doubts Moonshot has much room to raise prices further: cost-conscious users may prefer GLM-5.2 or MiniMax M3 for ordinary tasks.

  • Max contrasts two customer groups. SemiAnalysis-like users are willing to burn tokens on Opus or other high-end models, while some large companies, including Tesla and Uber, may limit users to roughly $200 worth of tokens per week. Those users are more likely to choose the cheaper GLM tier. Max says it would not surprise him if K3 saw little serious adoption among large enterprises, including because some users philosophically prefer open source.

  • The economics nonetheless strengthen the closed-lab thesis. Max says that if Kimi is probably not operating at negative margins at $3/$15, a similarly sized Fable charging $10/$50 would imply extraordinary margins. Dylan adds that the main cost is GPUs rather than employees; Max says token APIs may be an even better business than SaaS.

  • Dylan expects two or three post-training updates within the next few months, perhaps one or two months apart, while pricing stays roughly unchanged because no near-term hardware transition should provide a major throughput gain. He says a Composer based on K3 is definitely not happening because Cursor appears committed to training its own model from scratch.

4. Policy is compressing the frontier and nationalizing open source

  • Dylan says the current open-versus-closed gap has closed “squarely” because US restrictions on Anthropic mean users are not receiving the best models the labs have. He points to Mythos versus Fable, says he cannot use Opus and can use Sonnet only sometimes, and believes Soul 5.6 is not OpenAI’s largest trained model. He expects politics or another frontier advance to change the picture, possibly near the end of summer.

  • Max says there is huge demand for a Western open-source model that does not suck. He is surprised that no American company is at least on par with the fifth-best Chinese company. He expects many US enterprises to reject Chinese models even when proponents argue that weights can be run in an air-gapped data center, and he says a complete US ban on Chinese open-source models may be only a matter of time.

  • Dylan says Western models need to beat both the bulk of Chinese open-source models and tier-2/3 models from frontier labs. Cheap applications can already use those near-frontier models through services such as Bedrock or Foundry, so he sees government—not merely enterprise cost savings—as the larger Western open-source market.

  • Dylan and Max frame open weights as a sovereign-computing strategy. Max says Xi has encouraged Chinese companies to keep models open because Chinese government bodies want to download the weights, run them on their own servers, and support a domestic ecosystem. Dylan says the US government should take the same pragmatic approach.

  • Dylan notes that the K3 blog describes quantization during SFT, with native MXFP4 weights and MXFP8 activations for broad hardware compatibility. Jordan names a range of Chinese accelerators, including Huawei Ascend, Baidu, Kunlun Haxen, and Moore Threads, as China prioritizes running frontier models on domestic chips.

5. Harnesses and untapped demand still defend frontier-lab revenue

  • Max says it is increasingly difficult in daily use to distinguish the absolute frontier model at maximum thinking from high- or medium-effort modes. Testing K3 therefore requires him to examine the surrounding product—OpenCode, Hermes, and Pi—not just the model.

  • Small harness features affect where Max sends tokens: whether he can install a tool on a remote SSH server, use convenient keystrokes, or edit previous commands. He also says Slack bots and products such as Perplexity Computer can route work among K3, GLM, Sonnet, and OpenAI models without users caring which model is underneath; harness quality matters more for those first-cut tasks.

  • Dylan turns that observation into an outcome-based-pricing thesis. A router could send easy work to cheaper models and reserve maximum thinking for harder tasks, potentially producing 95%-plus margins. Max agrees that dynamically dialing thinking effort could capture much of the opportunity.

  • In response to Dylan’s concern that policy-enforced parity could destroy Anthropic and OpenAI pricing power, Max says no. He argues that the labs can continue training stronger models privately, pursue coding RSI without releasing it, and explore video generation, audio-to-audio work, deep research, robotics, and world models. His memorable framing is that “bourgeois” models can keep training each other while giving the public small tastes.

  • Max says his technology friends use AI roughly ten times less than he does, and that he may be in the top 10%, 1%, or even smaller fraction of users. As those users adopt larger models and as nontechnical users discover new applications, he sees perhaps a thousandfold increase in demand still ahead. Dylan agrees that any migration from existing Opus or GPT-5.6 users to K3 will be overwhelmed by people who have barely tried the technology, so K3 should not cause Anthropic or OpenAI ARR growth to decelerate.

  • Max closes by joking that if the market crashes because everyone has a “DeepSeek R1 … part two,” people should buy the stocks, while adding that it is not investment advice. Dylan ends with “Do your own due diligence.”