Pioneers Insight Method Research Author
20 Questions on AI Video’s First Year with Luma AI’s Barkley
Back to Episodes

20 Questions on AI Video’s First Year with Luma AI’s Barkley

Summary

  • Barkley’s core view is that AI video has not undergone a second architecture-level paradigm shift in the year since Sora’s release. Sora validated the path of DiT (Diffusion Transformer) replacing U-Net + Diffusion and continuing to scale capabilities with data; since then, progress has mainly come through steady iteration in physical understanding, consistency, and character and motion generation, rather than another paradigm shift. “Today’s video models may not yet be at the GPT-3 stage of language models.”

  • On Luma’s in-house evaluation system, Google Veo 2 was the strongest model for pure video-clip generation at the time of recording, but its lead still came with costs and trade-offs. The larger the motion, the harder it is to maintain consistency; stronger aesthetics may come at the expense of diversity; and better-performing models often take longer to run inference. More notable is the pattern that “the model released last may be the best”—training time, data volume, and the ability to absorb the previous generation’s strengths still determine rankings quickly.

  • The industry has not converged on a single business model, and the gap between research and application strategies remains unclear for now, though it may widen over time. Luma, DeepMind, and OpenAI are more willing to bet on visual understanding, world models, and AGI; Runway is focused on film and television studios, Pika and PixVerse on consumer effects, Hailuo on global user growth, and Kling on commercial revenue and gross margin per generation. Luma had raised roughly $160M at the time of recording, yet defined commercialization as “relatively important, but not that important,” and its latest financing did not impose explicit commercial targets.

  • The most defensible moat for video models in the near term may not be an algorithmic insight, but engineering throughput across data, evaluation, and training pipelines. Video data is large and noisy, and compression must preserve useful information; from collection and labeling to slicing and sampling, Barkley compares the process to an “industrial kitchen.” In a market without a unified benchmark, the ability to establish standards quickly, run experiments at scale, and feed creator feedback back into the model is itself a core capability.

  • Luma’s long-term bet is that Barkley does not believe text alone is enough to produce complete AGI; visual understanding and generation need to be integrated into a world model. Barkley relayed Sam Altman’s question: “How did you learn the objective laws of the world? Do you look at the world to learn?” The goal is therefore not merely text-to-video, but “anything to anything” combining video, images, sound, common sense, and language—understanding not only why a cup falls, but also simulating how it breaks when it hits the ground.

  • The product breakthrough Barkley feels most confident about for 2025 is character consistency, while real-time generation remains a high-potential, low-certainty long-term option. The former would lower the “draw” cost of continuous narratives, novel adaptations, and derivative content; if the latter can reduce latency, users might watch Harry Potter while asking for an alternative ending, blurring the boundary between producer and consumer. Barkley retains one caveat: real-time generation “may not be fully achievable in 2025.”

  • At research-driven AI companies, organizational roles are changing: researchers are gaining control over direction, while PMs are becoming connectors between users, data, and models. Nine out of 10 research ideas may fail, so a PM’s value is no longer promising that features will ship on schedule, but helping select directions worth testing, build evaluations, and identify the essence of problems quickly. DeepSeek’s “shock” in Silicon Valley reinforced Barkley’s view of the opportunity for Chinese talent: Chinese founders have advantages in consumer products and understanding both the Chinese and US markets, while Chinese researchers are also highly capable at the model layer and may gain more opportunities at global application and model companies.

Deep dive

1. Sora Established the DiT Path but Did Not Trigger an Era of Continuous Paradigm Revolutions

  • Barkley’s assessment of Sora’s impact is that it validated DiT’s ability to work on large-scale video data and replaced the previously dominant U-Net + Diffusion architecture with Diffusion Transformer. That was the clearest paradigm shift.

  • Since then, most players have continued along the DiT path. Physical-world understanding, consistency, and character and motion generation have all improved materially, but the progress looks more like gradual improvement in model architecture and product capability than another Sora-level break.

  • The host remarked that the past year felt like 5 or even 10 years. Barkley cautioned that the industry remains extremely early: “Video models may not yet be at the GPT-3 stage of language models.” The research and application paths that look distinct today may not truly diverge for several more years.

2. Ray 2’s Progress Comes from Physical Laws, Vertical Data, and Continued Scaling

  • Luma released its next-generation Ray 2 one month before the recording, launching text-to-video first and image-to-video later. Barkley’s representative test was a ball rolling down a flight of stairs: can the model simulate the physical process more accurately, rather than merely produce a visually attractive clip?

  • Ray 2 is no longer pursuing only general-purpose photorealism. The team has processed data and fine-tuned the model for verticals such as anime, hoping it can perform well in those domains too.

  • Luma still defines itself as a research lab, studying real-time video and video understanding alongside generation products. Barkley was unequivocal that “Scaling Law is still valid for video models.” Ray 2’s strong performance does not mean capability is anywhere near its ceiling.

3. Video Has No Unified Benchmark; Evaluation Closes the Loop Between Subjective Metrics and Training Feedback

  • Luma infers metrics from creator interviews: aesthetics, real-world events and physical laws, visual consistency, and prompt alignment. Because aesthetics are inherently subjective, the team generates videos in bulk once a model has an API, then sends them to a global crowdsourcing network for comparison.

  • Physical capability can be evaluated using a benchmark previously published by Google, or through internally designed prompts; Barkley did not recall the benchmark’s exact name. The team does not present these standards as absolutely objective, but continuously checks whether they align with internal expectations and the creator community’s judgment.

  • The evaluation system also provides training feedback: first identify the types of scenarios where the model falls short, then determine what data needs to be collected, labeled, and injected. Evaluation connects user perception with researcher decisions, creating a closed loop for improvement.

4. Veo 2 Is Temporarily in First Place, but Quality, Motion, and Speed Cannot All Be Maximized

  • Asked “who is the strongest in the world,” Barkley gave a direct answer: based on Luma’s testing, Google Veo 2 was the strongest model for pure video-clip generation at the time of recording, and may also have been the industry’s broadly perceived leader in output quality.

  • The ranking comes with 3 trade-offs: more motion can hurt consistency; stronger aesthetics may make diversity harder to preserve; and larger, better-performing models are generally slower at inference. Veo 2’s main cost is its long generation time.

  • Barkley observed a common pattern over the past year: “The model released last may be the best.” The reason is straightforward—it was trained longer, saw more data, and could absorb the previous generation’s strengths. Leadership therefore remains highly fluid.

5. US Players Have Different Priorities, but Research, Film, and Consumer Paths Have Not Fully Separated

  • Barkley grouped DeepMind and OpenAI under the large-company research path. DeepMind is advancing Veo while combining it with multimodal and world-model research. After OpenAI formally launched Sora as a product, many people judged it to have “fallen short of expectations,” but Barkley understood that the company was still iterating on its next-generation model.

  • OpenAI can combine its existing visual-understanding and multimodal capabilities rather than focusing only on video generation. Barkley speculated that its direction would move closer to world models and AGI, while remaining cautious about the scale of its commitment.

  • Runway is more concentrated on film and television production, professional editing, and studio partnerships. Luma leans toward independent creators and prosumers—users for whom the money saved through the product should far exceed the subscription price, creating a powerful retention lever.

  • Pika is more consumer-oriented, using AI effects to create viral hits and trends while lowering the barrier for novice users. Its target is not the highest-end film workflow, but turning AI into an entertainment behavior.

6. Chinese Models Are Differentiating on Growth, Gross Margin, and Ecosystem, While DeepSeek Has Shifted Overseas Attention

  • Barkley stressed that he has limited information on domestic companies and that the following was only his personal intuition: Hailuo appears more focused on winning global user scale and exploring consumer scenarios in different countries, without putting profitability first; Kling cares more about revenue in key markets, positive business growth, and gross margin per video inference.

  • PixVerse’s positioning is closer to that of Pika in the US, with a focus on consumer effects. He knew less about Vidu and Tencent Hunyuan, cautiously speculating that Hunyuan may be building an ecosystem around open source, while Vidu may be balancing research and consumer applications.

  • The professional community Barkley belongs to had long evaluated Chinese models, but US creators continued to use Runway, Luma, and Sora out of habit. After DeepSeek broke into the mainstream, more unsolicited recommendations began appearing on Twitter: “Pay attention to DeepSeek, and pay attention to China’s video-model companies too.” Continued updates from Kling and Hailuo further drew overseas attention.

7. The Research Path Trades a High Failure Rate for Breakthroughs, While Luma Does Not Yet Treat Commercialization as the Primary Constraint

  • The split between research and application begins with the founder’s vision. Luma believes visual understanding and world models are indispensable to AGI, and is therefore willing to explore directions with a low probability of near-term success but breakthrough potential once scaled.

  • Barkley described the research odds bluntly: “Nine out of 10 ideas in research may fail.” If the remaining one works, scaling it could produce a surprising Sora-like effect. The cost is sustained investment, expense, and organizational patience.

  • In the trade-off between research and commercialization, Luma was clearly on the research side at the time. Commercialization was “relatively important, but not that important.” The company relied primarily on financing to fund next-generation research while controlling inference and research costs.

  • Barkley said US VCs had extended long-term trust to this approach. Even in the latest financing round, they did not demand explicit commercialization data, instead focusing on how Luma defined visual AGI or world models. DeepMind and OpenAI are also continuing to bet on AGI, while application-focused companies are concentrating more on post-generation consumer use cases.

8. Video Will Reproduce Scaling Law, but Data Noise Makes the Engineering Path Entirely Different

  • Barkley believes the evolution of language models over the past 2 years will be replayed in video: continue expanding models and data along the Transformer path, build general-purpose foundational capability, and then develop reasoning and simulation of the real world. The base model could eventually be much larger than GPT-4 was in language.

  • Video’s distinctive difficulty is that information volume and noise rise at the same time. Not every information point in an image or video is useful; after receiving the data, the model must also understand which relationships and patterns actually matter.

  • “Continue to scale” therefore does not simply mean adding more videos. How to compress data while preserving information, how to process and organize it, and how to make the model truly understand relationships within the data will make video-training engineering materially different from pure language modeling.

9. A “Model That Only Reads Books” Is Not Enough to Learn the Full World

  • Barkley divided the AGI debate into 2 camps. One, represented by Dario Amodei, believes language materials are sufficient to understand relationships in the world. The other, represented by Yann LeCun and 李飞飞, believes humans first learn through visual feedback and that visual models are indispensable. Barkley also noted that Anthropic has never built a multimodal generation or understanding model.

  • At an After Hours event during OpenAI Dev Day, he asked Sam Altman whether Sora was still a priority and whether video generation was a necessary path to AGI. Altman responded: “How did you learn the objective laws of the world? Do you look at the world to learn?”

  • Altman’s subsequent answer was: “We shouldn’t expect a model that only reads books to learn all the laws of the world.” Barkley therefore concluded that even if OpenAI does not always concentrate its resources on the Sora generation product, visual understanding and multimodal research will continue.

10. A World Model Must Understand Laws and Simulate Outcomes That Have Not Yet Happened

  • Barkley first acknowledged that world models have no single definition, then broke them into 2 layers: first, understanding real physical laws; second, simulating the future based on those laws. What happens when a hand releases a cup—how gravity, friction with the ground, and material determine its fall and breakage—is the minimal test.

  • Understanding and generation are “2 sides of the same coin.” A model must both recognize the image of a hand holding a cup and accurately generate what happens after the hand lets go based on a static photograph. A robot handing someone water would likewise need to understand and simulate the entire process.

  • Luma therefore imagines the destination as “anything to anything”: video, images, human voices, sound effects, language, common knowledge, and know-how can all serve as inputs or outputs. A visual-understanding model may also combine with a language model, using language to compress and organize information.

  • World Labs is more focused on the 3D path. Luma previously worked on 3D reconstruction and generation, then chose to learn physical laws through video and massive datasets. Barkley stressed that this “doesn’t count as giving up”; the view is simply that the time has not yet come to scale 3D broadly. DeepMind’s Genie 2 represents another video path: simulating inside games, generating scenes in real time, and switching between different viewpoints.

11. The Data Pipeline Is an “Industrial Kitchen”; Throughput May Matter More Than Algorithmic Flair

  • The engineering and management advantage Barkley described begins with the ability to inject and output data quickly. If Scaling Law assumes that seeing more expands the range a model can understand and generate, then training throughput, compression methods, and data-processing speed translate directly into model capability.

  • His analogy is an “industrial kitchen”: video data is the food, with people cleaning, slicing, and categorizing it before deciding what size pieces to cut and in what proportions to put them into the pot. Collection, labeling, compression, slicing, and sequencing may not contain research innovation, but they can determine training efficiency.

  • Algorithmic innovation more often takes the form of local modifications to DiT or new methods for video and image editing. A common process is for academia to propose a hypothesis first, after which startups with large-model training capabilities scale up the data and test whether a small experiment can become broadly usable product capability.

12. When There Is No Train Station, First Make the Horse You Draw Run

  • There is no unified evaluation standard, and third-party annotators may not even know how to judge “beauty.” Luma’s approach is to provide multiple sets of examples and rules, test at scale which outputs best match the intuition of the team and the community, and continuously cross-reference and revise them.

  • Barkley summarized this culture as “be bold and try,” “sheer force can work miracles,” and trial and error. The host said it was like building a train without even having a train station. His response was: “Draw a horse. Whatever it looks like, as long as it can run, that’s enough.”

  • This uncertainty also extends into hiring. The CEO often asks candidates: “This is a problem that has never been solved. How would you approach it?” Barkley believes the answer depends on the specific problem and the candidate’s thinking; there is no fixed template.

13. Emotional Connection, Viral Trends, and Character Consistency

  • After Luma released its first-generation video model, a simple mechanic gradually became an emotionally resonant use case: users uploaded side-by-side photos of themselves and deceased relatives, entered “Let them hug,” and made the people in the old and modern photos embrace again. Barkley called it “very moving and very human.”

  • Another category of viral content comes from transformations, such as “Apple Dog”: a dog carries an apple, the apple suddenly disappears, and a series of entertaining transformations follows. This type of content has created a new trend.

  • Barkley’s high-confidence call for 2025 is that character consistency will improve substantially, because multiple companies, including Luma, have already achieved decent model results. Continuous stories will no longer depend on repeated “draws”; short films, novel adaptations, and derivative content will see a significant reduction in production barriers.

  • The vision for real-time generation is bigger but less certain: viewers could ask for a different ending while watching Harry Potter, and the model could immediately rewrite what follows. If latency falls low enough, content would be customized to each person’s needs and “the boundary between producer and consumer would become blurred.” But he explicitly said this may not be fully achievable in 2025.

14. Big-Tech Incentives Suppress High-Risk Research, and the AI Company Identity Debate Remains Unsettled

  • Luma often discusses big-tech research gossip over lunch: non-researcher managers have to choose between frontier exploration and stable team output, and most choose the latter. Because research failure can affect promotions and even a team’s survival, this incentive structure obstructs innovation.

  • Runway CEO Cristóbal Valenzuela’s public position is that the company should be viewed as a media and entertainment company, not an AI company. AI will eventually become as ubiquitous as water and electricity, he argues, so continuing to define oneself by AI will no longer mean anything.

  • Luma’s CEO immediately pushed back on Twitter, essentially saying that “people who stumble into AI without really understanding AI say things like that,” accompanied by a generated video of a frog sticking out its tongue. Luma believes each improvement in frontier models expands the application boundary, so foundational capability should come first; Runway is more willing to work backward from studio feedback to the model.

  • Barkley did not declare a winner: “Neither side is strictly right or wrong,” and the difference may simply be one of time horizon. Ronghui then remarked that many Silicon Valley CEOs argue directly on Twitter, making the phenomenon “Everything in the Bay Area happens on Twitter” particularly interesting.

15. At Luma, the AI PM Has Shifted from Requirements Owner to Interface Between Models, Data, and Users

  • On TikTok’s effects team, Barkley had PMs define requirements, assign tasks to research, and push them to launch. At Luma, researchers decide the main research directions and PMs play a supporting role. He initially felt the loss of no longer being “the person who could command everything,” but later concluded that the arrangement better fit research’s high uncertainty.

  • Internet features can generally be implemented by engineering, while AI research may see 9 out of 10 directions fail. The PM’s role therefore becomes helping define the ideas most worth testing, rather than demanding that all of them ship on a schedule. Model evaluation, data gaps, and creator insight are his main levers; researchers still make the final call on research direction.

  • Model-layer and application-layer PMs differ materially. Colleagues on Sora and Veo teams similarly prioritize data and evaluation; application PMs need not be tied to one model, focusing instead on the best use cases, features, and interactions. Luma had only Barkley as a PM at the time of recording, with no mature senior template or mentor. Hiring focused more on whether candidates could quickly “figure out” undefined problems.

  • His learning method is to talk with researchers, read the papers they recommend, and frequently try video, LLM, and agent products from different companies. He also makes extensive use of Windsurf to build small tools for himself. A former TikTok colleague uses tldraw to connect papers and products into a mental map, looking for “invariant themes” and new combinations across disciplines.

16. Chinese Talent Holds Both Consumer-Product Experience and Model-Research Expertise

  • Barkley believes understanding both the Chinese and US markets is especially scarce in consumer AI. The US consumer product that truly went viral in the past may have been Snapchat, followed by TikTok, which originated with a Chinese team. Chinese founders’ understanding of consumer psychology, the Chinese and US markets, consumer ecosystems, and AI hardware could produce more global application companies.

  • Opportunities also exist at the model layer. He believes China is highly capable in model research, and observes that many core researchers at Silicon Valley AI companies are Chinese, regardless of whether they completed their PhDs in China or the US.

  • DeepSeek’s impact on the Luma team came from both sides: a Chinese company had achieved a breakthrough in foundational model technology while also growing quickly at the application layer. The CEO consequently asked why China was described as not doing foundational research while also producing DeepSeek; the team also began placing greater value on Chinese talent.

  • Barkley acknowledged that geopolitics will affect mobility, but still believes AI will ultimately improve through global cooperation and competition. For him personally, DeepSeek was also a “reaffirmation”: holding to a vision, continuing to scale models, and investing in foundational research may eventually produce results.