Pioneers Insight Method Research Author
Training a 400B Model on 2,048 Blackwell GPUs for $20M | Researcher Conversations at GTC
Back to Episodes

Training a 400B Model on 2,048 Blackwell GPUs for $20M | Researcher Conversations at GTC

Summary

  • Arcee AI’s move into pre-training responds to a hard product ceiling: sub-20B customer work could only be as good as Llama, Mistral or Qwen’s base, while many customers’ legal and compliance teams stopped wanting work based on Chinese pre-trained bases. Owning the stack connects customization to known training data, especially “that last 10% of pre-training.” Lucas Atkins says Arcee’s current ambition is economic viability for developers—mainly startups—and enterprises, not AGI, though he says “never say never.”
  • Open weights are chiefly a sovereignty and unit-economics tool: if frontier-lab models become costlier, slower and more token-hungry without improving a narrow workflow, builders can reclaim margins with a fine-tuned 4B–8B model. Lucas’s invoice-processing example makes the call concrete: optimize for tool use, speed and cost, then add new skills yourself—“control is the number one” reason.
  • Lucas treats transparency as risk infrastructure, arguing that AI may repeat the mobile and social-media boom’s combination of enormous upside and damaging tail effects. If the people able to monitor, analyze and mitigate fundamental risks were only 200 safety researchers at OpenAI and Anthropic, he would rather “the risks be very transparent,” with mitigations easy to implement, understood and properly diffused through open models.
  • Arcee’s perceived talent disadvantage may be smaller than its compute disadvantage because researchers value visible work and broad ownership. Lucas says “talent’s actually easier than compute” and would keep the research team below 30 people even with “$2 trillion,” using opinionated, good-faith debate and full-stack exposure as organizational leverage. Compute is harder because it ultimately depends on funding and revenue; he says modularity will be important as compute demand and training paradigms shift.
  • Arcee does not intend to match frontier labs dollar-for-dollar; it aims to serve the “80% of economically viable tasks” that reward reliability, speed and low cost rather than extreme intelligence. In Lucas’s illustrative split, a next frontier model might improve 30% on FrontierMath but frontend generation only 5.5%: “Our goal isn’t to catch up. Our goal is to undercut.”
  • Trinity distributes execution across Arcee, DatologyAI and Prime Intellect, while B300 availability enabled a target of one month of pre-training instead of three. The trade-off was immature at-scale tooling and benchmarking—particularly sparse kernels—so the ecosystem around DeepSeek-like super-sparse models trained on Hopper became the “golden example” for throughput comparisons.

Deep dive

1. Arcee entered pre-training because post-training hit someone else’s ceiling

  • Lucas dates the push to early 2025, materializing in July. Post-training Llama, Mistral and Qwen below 20B had reached “the ceiling” imposed by their base models; later, many customers’ legal and compliance teams stopped wanting work based on Chinese pre-trained bases.

  • Lucas does not currently consider Arcee an AGI lab. He says the aim is to be the most economically viable company for developers—mainly startups—and enterprises.

  • He argues the West lacks an open-weight company whose revenue lives or dies by base-model quality. Full-stack ownership also connects data decisions across training stages, particularly “that last 10% of pre-training,” which he considers unusually consequential for post-training.

2. Model sovereignty converts directly into product economics

  • Lucas’s deliberately mundane example is an invoice app parsing receipts into Excel. It should not require a frontier-lab model that gets pricier, slower and more token-intensive while scaling frontier mathematics rather than tool use, speed or cost; a fine-tuned 4B–8B model lets its builder protect margins and add skills directly.

  • His conclusion is categorical: “Control is the number one.” Frontier labs operate in an “arms race for profitability,” so their sacrifices or choices will not always favor developers, startups or enterprises.

3. Open weights make AI risk inspectable beyond frontier labs

  • Lucas compares AI’s potential with the mobile boom, phones and social media: unprecedented connection arrived alongside reduced attention spans, depression and anxiety. AI could have similarly negative side effects, so he would rather risks be transparent and mitigations be easy to implement, well understood and properly diffused.

  • The interviewer invokes AlexNet and PyTorch; Lucas extends the argument to breakthroughs compounded through open artifacts. If OpenAI, Anthropic and Google each develop their own path to continual learning, Lucas says it would reach the wider world more slowly than an open breakthrough—but concedes, “There’s no going back” to their former research openness.

4. A sub-30-person research team is part of the talent pitch

  • On resources, Lucas surprises himself: “Talent’s actually easier than compute.” Researchers want to build models whose weights, engineering and infrastructure the public can actually examine.

  • Even if Arcee somehow had access to “$2 trillion,” he would keep research below 30 people. Breakthroughs should emerge from “really opinionated, talented” people debating in good faith, rather than researchers owning isolated niches.

  • At Arcee, researchers can be involved across data, pre-training, architecture, mid-training and SFT. He gives a hypothetical example in which five people work on pre-training and three can move to a mid-training bottleneck without a big onboarding process; that whole-process exposure is itself a recruiting proposition.

  • Compute remains harder because it flows through funding, money and revenue. Lucas thinks raising funds and executing over the next couple of years is feasible, but says Arcee needs a loose, modular structure that can follow shifts in compute demand and new training paradigms without spending ten times more than it makes.

5. Trinity is designed to undercut the frontier, not chase it

  • The interviewer asks how Arcee can compete with Western labs’ compute; Lucas reframes the competition. In his illustrative scenario, a next frontier model might gain 30% on FrontierMath but only 5.5% on its ability to generate frontend—and therefore slides for PowerPoint-like apps—while Arcee targets the “80% of economically viable tasks” below that frontier: “Our goal isn’t to catch up. Our goal is to undercut.”

  • Trinity’s name began because it sounded cool, then acquired structure: Arcee, DatologyAI for pre-training-data curation and Prime Intellect for infrastructure and GPU management, alongside three first-generation models.

  • B300s won on availability and speed because Arcee wanted pre-training to take “a month, not three.” At-scale benchmarking, suitable tools and especially sparse kernels were scarce, so the ecosystem around DeepSeek-like super-sparse models—many trained on Hopper—became a “golden example” for expected throughput.