7-Hour Marathon Interview with 谢赛宁: World Models, AMI Labs, Yann LeCun
Summary
谢赛宁’s view is that LLMs will not die, but will eventually “wither,” retreating increasingly into the communications layer rather than serving as the foundation for general intelligence. In his view, language preserves abstractions that humans have already selected and articulated, making it difficult to fully capture a continuous, high-dimensional, noisy physical world; the more fundamental substrate should be a world model that learns hierarchical representations, predicts the consequences of actions and supports planning. Language is a crutch: it can help a system move faster for a while, but may also keep it from exercising its own “legs.”
AMI Labs wants to build a “predictive brain,” not another large-model company chasing leaderboard rankings. Its most basic formulation of a world model is to predict the next state (s_{t+1}) given the current state (s_t) and an action or intervention (a_t). The final system will also need physical-world understanding, associative memory, reasoning, planning, causal inference, controllability and safety. Language, video generation and robotics are better understood as different interfaces to, or downstream applications of, that foundation.
This is also an organizational choice: an attempt to escape the finite game driven by benchmarks, product cycles and resource arms races. 谢赛宁 believes many researchers understand the importance of video understanding, pretraining and world models but can only execute against an established product cycle. AMI Labs wants more than half of its execution to resemble that of a new frontier lab, with roughly 60%-70% operating like a new lab while preserving 20%-30% for free-form frontier research; the company’s most important product may be a research breakthrough.
What the company really needs is not to download the internet all over again, but to build data and problems jointly with real-world partners. At the time of the interview, he described an initial team of about 25 people, with plans for a Paris headquarters and offices in New York, Montreal and Singapore; the pre-money valuation was €3B, while he cautiously put the fundraising target at “possibly around $1B.” The “reverse OpenAI” model starts with continuous data from hospitals, farms, factories and aircraft engines, lets the model create value, then uses the resulting data to improve the model—rather than downloading the internet once and pushing a product to market.
The clearest product exits are always-on wearables and general-purpose robotics, but both bottleneck on the “brain.” A device continuously recording heart rate, sleep, medication, diet and visual streams needs to understand long-term states, not simply attach ChatGPT to a camera. Robots can perform and move, but still cannot reliably handle household chores like a 12-year-old child. 谢赛宁 wants to “solve robotics without directly doing robotics,” first building transferable representations and world models, then closing the hardware loop through partnerships.
His research path looks scattered on the surface, but representation learning has been the constant underneath. DSN, HED, ResNeXt, MoCo, MAE, ConvNeXt, DiT, SiT, Cambrian, REPA and RAE all revolve around architecture, data, objectives and representations. From 何凯明, he learned that ideas come from 1-2 months of hands-on exploration, failure and unexpected signals—not from sitting still and inventing a clever trick; research is usually nonlinear, and paper acceptances and awards are highly random, making the long-term body of work and signature contributions more important than any single result.
He turned down Ilya twice and eventually said yes to Yann LeCun because of the fit between the problem, the people and the organization—not because of a simple comparison of prestige or compensation. In 2018, he passed on an OpenAI offer to join FAIR; after SSI was founded in July 2024, he declined again because his view of multimodality and perception at the time did not match the company’s direction. Yann LeCun offered a shared conviction around world models and JEPA, along with room to build a team, define the problem and preserve research autonomy.
The biggest risks remain unresolved: the data, objective, output format and scaling laws for world models have not converged. Internet video may be only the first step, and YouTube also raises copyright and Terms of Service issues; industrial sensor data is more valuable but harder to obtain. 谢赛宁 acknowledges that he does not understand business and cannot guarantee every thesis will be right, but believes the team can fail, pivot and keep exploring together. He compares entrepreneurship to skiing: do not lean back out of fear; point your shoulders downhill.
Deep dive
1. From Travel, Books and the Internet to Research
This was 谢赛宁’s first podcast and interview appearance. His earliest childhood memories date to around age 4 or 5. His mother ran a business and frequently took him around the country; his father had studied psychology and worked in education and television media, while several walls of the family study were lined with books. He was either out running around and seeing different places or at home browsing whatever books he could find—whether he was “supposed” to read them or not. He believes growing up in a relaxed household without a rigid STEM tradition helped him develop a relatively open worldview.
At around age 9, he got his first computer. He started by buying and playing games, then moved on to the internet, BBS forums, Sina blogs, QQ Zone and Fanfou. Reading had been a one-way input; the internet was the first place where he could express opinions publicly. It also turned him into someone interested in a wide range of subjects.
He rejects the description of himself as having followed an “A-track” from elite high school to competitions, undergraduate study, a PhD and then a faculty job. At most, he rates his path as a “B-track.” Awards in informatics and math competitions earned him early admission to Shanghai Jiao Tong University. His teachers wanted him to continue through the national college entrance exam and aim for Tsinghua or Peking University, but Shanghai, SJTU and computer science felt more compatible with his temperament, so he joined the ACM class. During the 2 months before formally entering the university, he spent almost all his time playing games in the dorm. He calls that period of consequence-free idleness one of the highlights of his life.
In the 30- to 40-person ACM class, he ranked roughly in the teens and never treated first or second place, or GPA, as the goal. In his interview, Professor 沈恩少 did not ask technical questions, but asked what he liked to read. He answered What Is Mathematics? and then remembered its author, Richard Courant. Years later, he worked at NYU’s Courant Institute of Mathematical Sciences, which Courant helped establish. 谢赛宁 sees the connection as an amusing coincidence in retrospect, but does not turn it into a causal story.
2. Vision Becomes the Central Thread
侯晓迪 was a legend at SJTU. As an undergraduate, he published a CVPR paper with only about 7 lines of code that solved an important problem, and he was also the main author of The SJTU Student Survival Guide. The guide discussed the education system, the purpose of learning, life choices and why research should not be reduced to padding one’s publication list. It also covered practical matters such as skipping class and completing assignments. 谢赛宁 particularly identified with its argument against treating GPA as the highest objective, and reached out to 侯晓迪 over Google Chat for advice on research and papers.
Third-year students in the ACM class normally had to leave campus for internships. 谢赛宁 might originally have gone to Microsoft Research Asia, but no computer-vision group there was willing to take an undergraduate. 于勇 believed an undergraduate should first gain research experience, and that the specific area mattered less. 谢赛宁 could not accept moving completely away from vision, so he contacted 颜水成’s lab at the National University of Singapore himself. He told 于勇 only after the funding and arrangements were settled. 于勇 ultimately agreed, and the lab later became an internship option for younger students.
Under 冯佳时, he completed a BMVC paper. It was not a CVPR paper, was not purely a computer-vision paper and was closer to a machine-learning project; its only direct application was face recognition. What mattered to him was experiencing research and writing a paper for the first time, while encountering AlexNet, ImageNet and deep learning around 2012-2013.
What drew him to vision was not just the image modality itself. As a child, he thought that if he had to lose one sense, he might be able to tolerate losing hearing, language, touch or smell, but losing sight would mean no more cartoons, movies or video games—as if he would also lose his independence. He cites one observation: the primary visual areas account for about 30% of the cerebral cortex, while seeing an image may activate as much as 70% of the brain; the eye is like “the only part of the brain exposed to the real world.” Solving vision, therefore, might also mean getting closer to intelligence itself.
He also discussed a mainstream theory that around 530M years ago, the emergence of vision triggered an arms race between predators and prey, potentially helping drive the Cambrian explosion. Vision was linked to action, prediction and survival from the outset.
3. Choosing People and Problems for His PhD
When applying for a PhD, he initially failed to receive an offer from a supervisor who truly wanted to work on vision. He even considered switching to recommender systems or more general machine learning. As the April 15 deadline approached, he sent out a flurry of cold emails. 涂志文 replied. The 2 spoke at 3 a.m. because of the time difference; 谢赛宁 explained why he worked on vision, what he had done and why he wanted to collaborate with 涂志文. At the last minute, 涂志文 offered him a place at UCLA.
About 1 week before enrollment, 涂志文 told him he would be leaving UCLA but could not yet disclose where he was going. 谢赛宁 almost immediately decided to wait and follow his adviser rather than remain at UCLA. He later learned that the destination was UCSD. Its overall ranking may have been below UCLA’s, and Serge Belongie, whom he wanted to work with, was also preparing to leave. But he considered those details noise. What mattered was “who you do what with, and whether it is what you want to do.”
He became the first student 涂志文 recruited to UCSD. 涂志文 would sit beside the monitor and inspect code line by line. He also described how, before PyTorch, mature open-source libraries and GPUs, researchers wrote roughly 50,000 lines of C++ from the ground up to complete work such as image segmentation. 谢赛宁 came to admire 涂志文, 朱松纯, Fei-Fei Li and other pioneers even more: they did not merely publish papers; they carved out an international path in computer vision that had not previously existed.
He rejects the narrative that he was unremarkable in China and suddenly broke out in the US. His experience, he says, was a smooth accumulation rather than a sudden explosion. The researchers he admires most were not working to “make a splash”; they were working to understand the problem.
4. Nonlinear Research and Five Internships
Deeply-Supervised Nets (DSN), an early PhD project, added supervised exits at intermediate layers so gradients did not have to propagate backward only from the final output. HED applied the same idea to edge detection: lower layers captured rougher edges, higher layers captured finer ones, and the outputs were ultimately fused into a result closer to human perception. HED was published at ICCV and nominated for the Marr Prize. It did not win a best-paper award, but 谢赛宁 believes a nomination itself is effectively an award in the context of the Marr Prize.
DSN initially received high scores when submitted to NeurIPS, but was rejected because a formula was missing a squared term due to a typo. The rebuttal did not get the reviewer to see the explanation. The paper was later resubmitted to AISTATS and won a Test of Time Award 10 years later. The experience reinforced his belief that acceptances and awards are random, and that researchers should not treat a result at any single point in time as the final estimate of their worth; a decade of accumulation is what ultimately forms a reputation. He also admits that when a paper is rejected, it is hard to imagine an award 10 years later.
During his PhD, he interned at NEC Labs, Adobe, Meta, Google Research and DeepMind. Each stint typically lasted 3-6 months. He spent about half his time at school and half in industry, and achieved no result in roughly half of the internships. He would sublet his room, drive from Southern California to Northern California, and carry all his belongings in 2 suitcases in the car. He wanted to see firsthand how different organizations worked, while repeatedly asking himself: “What if I’m wrong?”
The NEC experience went smoothly. He published a CVPR paper and liked the collective atmosphere, which was largely made up of Chinese researchers. Adobe produced no ideal result. The project involved design, crowdsourcing and Mechanical Turk feedback, and he still felt guilty toward his mentor. The experience taught him that failing to produce something is not the end of the world.
5. ResNeXt, Organizations and the Representation Thread
In the final month of his Meta internship, 何凯明 joined FAIR. 谢赛宁 handled transportation and taught him how to use Linux and the cluster; 何凯明 took him to the ImageNet Challenge. The 2 eventually produced ResNeXt, which finished second. First place went to an ensemble combining existing algorithms. 谢赛宁 believes ResNeXt was a genuinely new framework and, in terms of substantive results, may have been closer to first place.
ResNeXt expanded ResNet’s serial structure into multiple parallel groups. At roughly the same compute budget, it achieved better performance through greater width and sparsity, while exhibiting scaling behavior with some similarities to today’s MoE systems. The X in the name means Next, but also credits 谢赛宁’s contribution.
He took away a broader lesson: research is never linear, and many important projects find their direction only in the final month. Google had him study the architecture and training pipeline for video neural networks; DeepMind exposed him to a different organizational model. In London, he worked on RL and embodied agents in simulated environments. He found that he did not enjoy working directly on RL or robotics, but observed how DeepMind gradually moved from bottom-up exploration to more organized top-down execution. Demis would tell interns that DeepMind would eventually become a company capable of winning multiple Nobel Prizes; at the time, everyone thought that sounded excessively ambitious. 谢赛宁 also watched the AlphaFold team turn an exploratory idea into an organized, execution-heavy project.
His PhD thesis was titled Deep Representation Learning with Visual Structure. Image recognition, segmentation, edge detection, video recognition, action recognition and later embodied RL were different branches of the same tree; representation learning was the root. He believes architecture, data and objective jointly determine a representation: architecture is the hardware, data is the fuel and the objective determines what the model is trained to do.
He defines representation learning as mapping raw data (x) into a space with useful properties, making downstream tasks easier. The mapping is typically hierarchical and nonlinear. He is willing to keep the title for the long term because specific hot topics change, while the question of how to form good representations remains foundational and unresolved.
6. FAIR, MoCo and Research Method
When he graduated in 2018 and began looking for a job, he did not seriously consider faculty positions. His 5 internships and scattered projects did not fit the traditional trajectory. He interviewed at OpenAI, spending 5 or 6 hours in a small room solving problems with a pencil on A4 paper. He received an offer but chose FAIR instead. 何凯明, Yann LeCun and Ross Girshick were the computer-vision “three pillars” in his mind at the time. FAIR still felt more like an academic institution and was more open.
During his FAIR job talk, he did not know that academic talks conventionally lasted 45-50 minutes and stopped after half an hour. 何凯明 later told him the researchers found it unusual, while joking that finishing in 30 minutes would save everyone time in the future. He still became one of FAIR’s earlier fresh PhD hires.
At FAIR, he worked with 何凯明 on MoCo. Before that, visual self-supervised learning had accumulated a long list of pretext tasks, including rotation prediction, colorization and masked reconstruction, but results were typically 15%-20% below supervised ImageNet pretraining. MoCo built on CPC, Memory Bank and earlier metric-learning work, turning contrastive learning into a genuinely effective framework. 何凯明 led the direction, and 谢赛宁 stresses that the work should not be rewritten as a one-person myth when it emerged through collective evolution.
何凯明 influenced him through focus, research taste and the expansion of his inputs. He would allocate a large number of mental cycles to a single problem, using reading, discussion and experiments to tease out the connections that actually mattered. 谢赛宁’s advice to students is to start with a general direction, then give themselves at least 1-2 months to explore: reproduce the baseline, modify the code, derive the equations, read the papers and look for signals instead of trying to invent ideas from a blank page.
Both success and failure provide information. The worst outcome is a result that is neither good nor bad, producing no signal. An idea discovered through exploration truly belongs to the researcher. An ideal project might spend 1-2 months exploring, 2-3 months scaling up and filling in experiments, then another 1-2 months writing the paper; the best projects often pivot at the very end. Research is not a straight line from a predetermined point A to point B. It is a search for the gradient along the way.
He cites Bill Freeman’s nonlinear curve of career impact: weak work and decent work have little impact, while the impact of a truly exceptional signature contribution rises abruptly. Research is therefore more like an “infinite game.” You do not need to win every time; succeeding once in a lifetime can already matter enormously.
7. MAE, Infrastructure and Research Taste
MoCo evolved through V1, V2 and V3 and expanded in scale. Its representations eventually surpassed supervised ImageNet pretraining on some tasks. The team briefly thought self-supervision had a bright future if they simply kept scaling, but that thesis did not fully materialize. They turned to Masked Autoencoder (MAE), learning representations through masking and reconstruction. MAE behaves differently from contrastive learning: it may be weaker under linear probing but better under end-to-end fine-tuning. Neither approach failed, but neither achieved the kind of impact associated with LLMs. Self-supervision was later extended to 3D, point clouds, medical imaging and robotics.
FAIR once rented roughly 5000 TPU cores on Google Cloud, where the ecosystem and infrastructure were difficult to use. 何凯明 built a TPU infrastructure stack from scratch, enabling later work including MoCo, MAE and DiT. 谢赛宁 learned that the ceiling of research often depends on the quality of the baseline and the infrastructure. A weak baseline can create a false improvement; only when the baseline is strong enough do subsequent gains become credible. Work such as Faster R-CNN, Mask R-CNN and Focal Loss also rested on years of infrastructure and code-base development.
The first lesson for FAIR interns was sometimes Excel. The research team used spreadsheets to track experiments, forecast each result in advance and then generate a new gradient from the gap between the forecast and the actual result. Too few experiments produce no signal; indiscriminately consuming the full compute budget merely dumps results into a spreadsheet. Negative results are also valuable: a 10-point performance decline may indicate that the opposite direction is worth pursuing. The most dangerous outcome is no movement at all.
何凯明 also shaped his writing and aesthetic judgment. A paper is a communication interface for other people to use, not merely a document for its authors; tables, prose, punctuation and formatting should all help readers understand the core idea. He also gave 谢赛宁 the Diamond Sutra, which 谢赛宁 interpreted through the line “all conditioned appearances are illusory”: the form, metrics and narrative of a paper are not the problem itself, and researchers should penetrate the surface to ask what is actually at stake. Research taste is not just the ability to choose a topic. It also means knowing what matters, how to express it and whether one is being pulled around by short-term fame and acceptance.
8. ConvNeXt, DiT and Cambrian
ConvNeXt came from questioning the assumption that self-attention was the most important part of Vision Transformer. Through extensive controlled experiments and ablations, 谢赛宁 and 刘壮 found that macro- and micro-level architecture design might matter more than attention itself. The title, A ConvNet for the 2020s, was proposed by 何凯明. The hand-drawn diagram showing the evolution of the architecture was later cited by many papers.
In FAIR’s later years, the rise of ChatGPT and OpenAI prompted frequent internal discussions about what the organization should do over the next 1-2 years. 谢赛宁 believed weeks of research-alignment meetings could not substitute for bottom-up exploration. FAIR’s focus was also becoming less exclusively research-oriented, which was one reason he left.
DiT was not initially conceived as DiT. The team wanted to study the representations learned by diffusion models. They used ViT to compare against other systems and found it simpler, more stable, more efficient and more scalable than U-Net, so they pivoted to the architecture itself in the final month. The paper was rejected from CVPR for insufficient novelty. The team moved it to another conference with almost no changes, where it was accepted—another reminder of the randomness of peer review. The complete DiT work was done at FAIR, but the final paper listed only NYU and Berkeley. 谢赛宁 said the decision involved his departure and legal considerations.
After DiT, the team extended Flow Matching to the Transformer setting, producing SiT. Repeated rejections gradually made him less sensitive to short-term evaluation. He uses Taleb’s idea of “antifragility” to describe research: if the long-term gains from a shock exceed the losses, the system does not merely withstand the blow; it becomes stronger. DiT and SiT later became baselines that other researchers could build on.
The Cambrian series is named after the Cambrian explosion, a reminder that vision may have been the starting point for the evolution of biological intelligence. 谢赛宁 compresses 538M years into 24 hours: language, abstract thought and symbolic reasoning appear only in the final 8-10 seconds. He is not trying to dismiss language, but to reject the idea that abilities humans developed only recently constitute the whole of intelligence.
S*3 examines vision encoders and argues that CLIP may be affected by language shortcuts; the name comes from the concept in Stanley Kubrick’s 2001: A Space Odyssey. Cambrian continues to study data composition, visual representations and vision architectures without changing the LLM component, making it possible to isolate the contribution from the visual side. Cambrian-S focuses further on video, super sensing, spatial intelligence and the ladder toward world models. Short clips filmed on New York streets were not promotional material, but an attempt to show how a genuinely intelligent agent might receive a continuous visual stream and perceive space and causal change.
9. NYU, Academic Relationships and Credit
When 谢赛宁 moved from FAIR to NYU, Yann LeCun often joked that he had recruited him 3 times: first to FAIR, second to NYU and third to AMI Labs. NYU’s Center for Data Science was established more than 10 years ago under Yann LeCun’s leadership. Independent of the traditional computer science and mathematics departments, it connected physics, chemistry, mathematics, statistics, the business school and computer science through open glass offices, interdisciplinary space and AI talent. More universities later adopted similar organizational structures.
Fei-Fei Li influenced him primarily through problem definition. ImageNet was not just a large dataset. Around 2011-2012, image classification itself was not yet a clearly defined problem. She created a playground and set the agenda, allowing deep learning to develop against a well-defined objective. 谢赛宁 worked with her on Thinking Space and Cambrian-S, hoping to learn how to define problems and establish directions.
He does not believe he entered the “AI inner circle” through a particular trick. In his view, people working on related questions eventually converge naturally. They are all interested in vision, representations and the foundational problems of intelligence. From the outside, his choices appear logical; internally, they were often deliberately unstructured. It is impossible to truly optimize for a predetermined outcome. At each junction, he could only choose the work he most wanted to do and the people he most wanted to work with.
He sees the scientific community as a graph. Teachers, students and colleagues shape one another, and students can in turn influence their teachers. A paper is not primarily a vehicle for personal fame; it is a way to share knowledge so that readers have something to do after finishing it. He is uncomfortable with the word “impact.” His preference is to understand the problem first, then pass that understanding to others and increase the total amount of intelligence on Earth.
What he opposes is packaging a team’s work as the personal fame of its leader, not the dissemination of research itself. When the head of a group is not first author, media coverage should explain what problem the work solved and give visibility to the students who actually did it. He is comfortable condensing and explaining work on X, but does not want publicity to become a “Team X” or individual-star narrative.
10. LLMs, Vision and World Models
谢赛宁 says LLMs will not die, but will gradually wither. They will remain valuable tools and communication interfaces, but not the foundation of the world-model stack. Language is knowledge accumulated by human civilization over thousands of years and uploaded into books and the internet. Free does not mean unsupervised: humans have already performed extensive selection, abstraction and labeling.
“I have a glass of water, and it falls to the ground and breaks” communicates the outcome, but says nothing about how the glass broke or what dynamics governed the event. Language is efficient precisely because it discards much of the world’s structure. That is why 谢赛宁 worries that language can become a shortcut for vision and representation: it may improve a benchmark while preventing the system from learning the structure of the world for itself.
He is not discouraged that vision has moved from the center of AI to the periphery. Instead, he credits LLMs with accelerating multimodality. The problem is that some multimodal tasks are solved mainly through language, with vision providing only a small amount of context. Yann LeCun compares a language model to a crutch: it lets the system walk, but if the visual leg has not developed properly, the system still cannot run in the real world.
By “real intelligence,” he does not mean that digital-space capabilities have no value. He is emphasizing interaction with the physical world. LLMs are strong at factual knowledge, law, education, summarization and digital-space tasks. Sensor modeling, continuous environments, action consequences and robotics require another kind of intelligence. He wants to solve the robotic brain and its representations first, then complete the robotics loop through partnerships.
V* is his attempt to build a visual System 2 inside multimodal systems. When a person is asked the color of the trash can beside them, they do not answer from language alone; they first search for, locate and identify the target. A system likewise needs visual search and reasoning at test time. 谢赛宁 says the work predated the popularization of test-time scaling and established a benchmark. He had introduced the work and benchmark to Alex Kirillov and Bowen; OpenAI later released Thinking with Images with similar examples and a similar benchmark. To him, this shows both that academic research can influence industrial systems and that credit assignment in industrial research is becoming increasingly opaque.
REPA reconnects the DSN-style idea to generative models: in addition to the top-level diffusion loss, it aligns intermediate representations with an external self-supervised model. Representation Autoencoders (RAE) attempts to use strong representations as the encoder or foundation for generative models. 谢赛宁 cites 马毅’s encouragement on high-dimensionality: kernel methods and the up-projection in Transformers both make problems easier by lifting them into higher dimensions. His core bet is that sufficiently good representations might unify language, pixels and actions, but that remains a forward-looking judgment, not a proven fact.
11. Defining World Models and the Training Problem
The basic definition 谢赛宁 uses is: given the state (s_t) of a system or environment, and an action or intervention (a_t), learn a function (f) that predicts the next state (s_{t+1}). The value of a world model is not the ability to generate attractive images, but to predict the consequences of actions and support planning and decisions. The idea is not new. Kenneth Craik discussed internal world models in humans in 1943; Model Predictive Control rolls out sequences of actions; and Rich Sutton’s Dyna distinguishes a reactive policy from a model-based policy.
A state is not a complete copy of the world. It is the minimum information required to complete a task. In the interview room, if the only objective is conversation, there is no need to record every texture, sound wave and lighting condition. Nor would aircraft dynamics be simulated from molecular collisions; the system would use a higher-level abstraction such as fluid mechanics. Selecting a state that is sufficient for prediction and decision-making is precisely a representation-learning problem.
He believes current LLM safety and controllability depend primarily on fine-tuning and post-training, which tell the model what it can and cannot say. A stronger world model should be able to predict the consequences of an action at inference time and then apply external constraints. When a robot uses a knife to cut vegetables, for example, it should understand what happens if the blade turns toward a person, rather than merely memorizing a large number of linguistic examples saying “do not hurt people.”
He distinguishes several directions. Sora, Veo, Genie, Runway and Luma are closer to world simulators, emphasizing video consistency, duration and control. Fei-Fei Li’s World Labs focuses more on explicit 3D spatial representations. 谢赛宁 cites Autodesk’s roughly $200M investment as an example of why verifiable spatial assets could be valuable in design, CAD and related applications. AMI Labs is aiming at something closer to a predictive brain, with the emphasis on representation, prediction, planning and decision-making.
LLMs are neither substitutes for world models nor irrelevant to them. They provide world knowledge and remain an important component, but a language interface typically carries explicit intent, while much of a person’s world model runs in the background. Video generation is closer to the world than language alone because it must model which phenomena are more likely to occur. But pixels may still be only an interface for humans; a true world model does not need video generation to be its core function.
12. Data, Products and the Real World
谢赛宁 calls the data problem one of the biggest bets in world models. The past was about “downloading the internet”; the future may be closer to “downloading humans.” The continuous visual information received by a young child could be enormous. He also notes that the total quantity of data used to train all large models may be equivalent to only about 30 minutes of video uploaded to YouTube. This is his estimate of the order of magnitude, not a precise measurement.
The first phase may still begin with internet video from YouTube and similar sources, but copyright and Terms of Service are obstacles, and platforms can block crawler IPs. More valuable data may come from sensor systems in aircraft engines, factories, hospitals and farms. An aircraft engine, for example, may have around 1000 sensors that can help models learn design flaws and long-tail failures. These vertical world models would ideally be built on top of a general world-model pretraining layer.
He believes product form may be harder than the data. The first outlet could be an always-on wearable that continuously records heart rate, sleep, stress, medication, diet and visual streams, then converts raw signals into assessments such as “you need to rest,” “your sleep has been poor for several consecutive nights” or “you should take the day off.” AI glasses that merely connect a camera to an LLM are still just ChatGPT with a camera. The real value lies in understanding a person’s life over time.
The other outlet is robotics. Robots performing at the Spring Festival Gala show how quickly hardware and motion control are advancing, but robots remain far from entering homes, caring for the elderly and handling household chores. 谢赛宁 calls this the second half of pretraining: the inputs may be video and other continuous, high-dimensional, noisy signals, but the right output remains an open research problem.
13. Entrepreneurship and AMI Labs
谢赛宁 had research freedom at NYU, but worried that resource constraints would eventually trap him in a “middle-paper trap”: producing a steady stream of respectable papers without pushing an idea into a new breakthrough. A mentor once suggested that he speak with Yann LeCun. During a subsequent one-on-one, Yann LeCun volunteered that he had decided to leave Meta and start a company, and that the project he wanted to pursue was highly aligned with 谢赛宁’s interests. 谢赛宁 spent about 1 week wrestling with the decision. He focused mainly on the research direction, the people he would work with, the school affiliation and whether he was still in a sufficiently strong mental state—not simply on comparing near-term income.
He believes big-company benchmarks and product cycles compress the space for research. Even research and product teams within the same company can become isolated from one another. AMI Labs wants to build an organization better suited to defining problems and pursuing long-term exploration. It is neither a pure research lab nor a closed large-model company. 谢赛宁 estimates that more than 50% of the organization will resemble a new lab, with roughly 60%-70% devoted to organized execution and another 20%-30% reserved for free-form exploration.
He calls the model “reverse OpenAI.” The forward path downloads data from the internet, trains a Transformer, develops intelligence and then takes it to consumers or businesses. The reverse path starts with the problems and data of real-world partners; the model first creates value in a specific setting, which then generates new data that feeds back into the model. World models need the world, and the real world needs models that can handle physical problems.
AMI Labs plans a Paris headquarters plus offices in New York, Montreal and Singapore, with the aim of creating a cross-regional alliance. 谢赛宁 uses the story of Visa and Mastercard as an analogy: Bank of America initially built Visa, while smaller banks later joined forces to create Mastercard. AMI Labs may not copy that commercial model, but it wants institutions with specific problems and data to participate in building the system, rather than leaving world models entirely under the control of a few platforms. Openness does not mean every research project will be open-sourced; AMI Labs remains a serious startup.
The initial team is about 25 people, including 6 co-founders. 谢赛宁 is co-founder and Chief Science Officer; Yann LeCun is executive chairman. The team also includes a CEO, COO, VP of World Model and roles responsible for connecting research with product innovation. 谢赛宁 says the fundraising target may be around $1B, but stresses that this is an uncertain forecast. The valuation he stated clearly was €3B pre-money. He also said that some people had given up several tens of millions of dollars in unvested OpenAI stock, while others had passed on Meta offers worth roughly $15M-$20M, to illustrate that the team is primarily mission-driven rather than motivated by near-term cash.
14. Yann LeCun, JEPA and “The Normal One”
谢赛宁 turned down Ilya twice. In 2018, he passed on an OpenAI offer to join FAIR. After SSI was founded in July 2024, Ilya contacted him again. They discussed how to give future AI the capacity for love, as well as multimodal, visual and perception models. Ilya believed that this area had already been handled well at the time; 谢赛宁 judged that SSI was at least then more focused on language, so he did not join. He does not frame their directions as adversarial, instead quoting the idea that “brothers climb the mountain, each pursuing his own path.”
He went through 3 stages with JEPA: questioning it, understanding it and becoming part of it. At first, he saw JEPA as another self-supervised algorithm. Later, he came to view it more as a cognitive architecture encompassing world understanding, prediction and planning. JEPA is not one fixed method, but an ocean in which many methods can sail; LLMs can also be part of it.
谢赛宁 admires Yann LeCun’s scientific integrity and worldview. Yann LeCun is not someone who refuses to change forever; he is willing to be changed by facts, but unwilling to change merely because an organization or public opinion demands it. He uses a sailboat as a management metaphor: trust each person to do their job, and correct course early when the boat begins to drift. 谢赛宁 wants to play the science role at the company rather than present himself as a CEO who already knows he is right.
On AGI, he prefers to abandon human exceptionalism. Human intelligence itself is constrained by our senses, bodies and neural bandwidth; it is not pure general intelligence detached from an environment. Different animals have different forms of intelligence. Chimpanzees, birds, dogs and bats all exhibit distinct forms of memory, prediction, communication or perception. Rich Sutton helped him realize that creating the intelligence of a squirrel may be harder than writing code, winning an IMO medal or going to Mars. A squirrel has to survive in the real world, with goals, intrinsic rewards, hunger, emotions and social behavior.
He still hopes eventually to build intelligence approaching the human level, but a more practical test is whether a system can first produce a reliable, sufficiently general household robot. A 12-year-old child can handle a broad range of household tasks; current robots cannot.
谢赛宁 supports Liverpool and Jürgen Klopp. In contrast with Mourinho’s “The Special One,” Klopp called himself “The Normal One” and saw himself as the battery powering the team. 谢赛宁 wants to be that kind of “normal person.” Research is spent mostly feeling around in the dark; the joy of actually getting something to work may account for only 5%-10% of the process, but the team can fail and pivot together. He compares entrepreneurship to skiing: leaning back when afraid means losing control; only by pointing your shoulders downhill can you hope to regain it.
He walks through Washington Square Park every day, seeing people play music, dance, push strollers, play chess or simply stare into space. It reminds him that most people do not care about AI, even though AI may change their lives. Can fate be predicted? His answer is that there are not enough resources. We may need to treat Earth—or even the entire universe—as a computer, and the final answer may still be 42.