AI研究员写科幻:对话Meta田渊栋,畅想智能未来
Summary
The most industry-relevant “prophecy” in Within the Dawn is that the compute race will shift from the performance of individual GPUs to interconnects, communication bandwidth, and information compression. In 2020, 田渊栋 cast the consciousness-bearing “Spirit Realm Cubes” as an allusion to NVIDIA GPUs: superintelligence emerges when the cubes are linked at high speed, while the Galactic Alliance imposes a bandwidth blockade. Reality then produced NVLink, NVL72, CloudMatrix 384, and DeepSeek V3’s use of FP8 on H800 to reduce communication pressure; he had imagined this as science fiction from 10 or even 50 years in the future, only to see “science fiction become reality two years later.”
Virtual worlds and interstellar civilization are not mutually exclusive; humanity is more likely to digitize its existence while sending its computational substrates onward to explore the universe. Human bodies are fragile and lifespans finite, while sending astronauts into deep space requires food and biological support; once consciousness enters a virtual world, its substrate can travel to Proxima Centauri and beyond. 田渊栋 believes humanity entering virtual worlds is “unavoidable,” but curiosity will not disappear: the final form may be “entering the virtual world on one hand, and continuing to wander among the stars on the other.”
Once technology reaches “whatever you think of, you get,” the scarcest resource may no longer be productive capacity but genuinely original ideas. The Galactic Alliance can instantly materialize factories, spacecraft, and any object it imagines; what it fears is taking the wrong path, because a remote civilization with a better solution could immediately become a replacement force using the Alliance’s own technology. The Alliance therefore has to preserve Earth’s diversity of thought like a nature reserve while keeping that exploration under control—it does not want its own ideas to contaminate the reserve, nor allow it to grow strong enough to turn against the system.
AI’s impact on employment may unfold in 3 stages: displacement and emptiness, competition for uniqueness, and finally an explosion of occupational diversity. In the short term, people will find the jobs they depend on replaced by AI, or discover that no amount of effort can make them better than AI; in the middle stage, they will be forced into attention contests built around novelty and eccentricity; only in the long term might they leave the “education—skills—work—wages—raising the next generation” loop and turn writing, painting, or singing into life itself. 田渊栋 is optimistic about the end state but stresses that “what twists and turns happen in the middle is hard to say,” while AI’s descent into industry-specific workflows will continue over the next 2-3 years.
Today’s frontier models can assist with research but have yet to cross the threshold at which top researchers infer fundamental rules from a handful of clues. Deep Search is already good at finding material, but it remains “an outsider who looks like an insider, while insiders see an outsider”; a model may need 1,000 or 10,000 samples, while an excellent researcher can infer the real minor factor from 1 or 2 anomalies. 田渊栋 does not believe more pretraining data alone will close the gap, is “not that optimistic” about replacing researchers within 5-10 years, and estimates that a breakthrough in the training paradigm may still take 10-20 years.
Generative AI is pushing technical organizations toward smaller, sharper teams, hands-on managers, and core staff amplified by multiple AIs. 田渊栋 went from arguing in 2021 that leaders should not bury themselves in details to requiring himself in 2024 to read and write code again, because purely managerial leaders lose technical sensitivity and “one person plus a lot of AI” may outperform a small team. The OpenAI, Hugging Face, and DeepSeek examples discussed on the show all point to the same trend: leaders who can quickly judge frontier shifts such as o1 Preview while implementing them personally can accelerate organizational pivots and execution.
AI has entered the novel-production chain, but the most valuable ideas, structures, and turns still come from humans. DeepSeek is good for brainstorming but prone to logical breakdowns; Claude 3.5 is better at sensing relationships among supporting characters, Gemini 2.0 at detailed description, and 4o relatively flat; all share problems with long-context forgetting, formulaic plots, and “the prince and princess living happily ever after.” 田渊栋 uses Cursor to build his own collaboration tool: humans write the outline and key passages, while models fill the gaps. His view is that procedural writing will be standardized, while personal experiences that have “never been explored before” will become the fresh knowledge AI needs most.
Deep dive
1. Serial pressure turned a years-in-the-making hobby novel into a finished manuscript
田渊栋 encountered online fiction and tried writing it in 2005-2006. Within the Dawn accumulated fragments and ideas over many years; he later admitted that without the pressure of public serialization, “this novel would never have been completed.”
During the 2020 pandemic, he announced the serial on Zhihu and used an external commitment to force himself to deliver 1 chapter every day. The most intense phase lasted roughly 3-4 months; the novel finished serialization in 2021, entirely handwritten and before the ChatGPT boom.
Electronic Industry Press approached him in 2022-2023, and the novel was formally published in June 2024. At the time of recording, it had a 7.9 rating on Douban and a 75% rating on WeRead, far exceeding what he had expected from a hobby.
2. The AI era needs its own science fiction, not recycled robot fantasies
田渊栋’s starting point was that every form of productive power generates a corresponding imaginary world: the steam age imagined long-distance travel, while the electrical and mechanical ages produced submarines and metal barrels traveling to the Moon. The intelligent age should likewise produce stories driven by AI mechanisms.
AI science fiction in the 1950s and 1960s often portrayed machines as precise calculators devoid of emotion, or assumed that “robots could never surpass humans.” In retrospect after ChatGPT, those assumptions look narrow.
Within the Dawn blends the desire of his younger self to see the weak fight back against the odds with the middle-aged impulse to refuse passive resignation and keep doing interesting things. Alien invasion is an old framework, but the mode of attack and defense had to belong to the new AI era.
3. The aliens’ most effective weapon is not firepower but a painless virtual paradise
The Galactic Alliance does not destroy Earth directly. Instead, it offers the “Spirit Realm,” where people need not worry about food or clothing, no longer age, and can eliminate disabilities and physical defects while retaining sufficiently realistic senses such as touch.
At the same time, the alien civilization alters the Sun and Earth’s energy conditions, worsening the surface environment. Those who refuse the Spirit Realm can only struggle for survival underground, making the temptation effectively impossible to avoid.
Humanity splits between those who accept a stable virtual life and those who insist on autonomous development. The war is no longer merely a clash of spacecraft, but a struggle over consciousness substrates, algorithms, communications, and values.
The protagonists are doctoral students and researchers; lab meetings, senior-junior mentoring, supervisors competing for resources, and academic politics all enter the plot. This is how 田渊栋 delayed the conversion of his doctoral life more than a decade earlier into science-fiction material.
4. The “Spirit Realm Cubes” anticipated the central conflict of GPU clusters
Each Spirit Realm Cube carries one person’s consciousness, and 田渊栋 deliberately made it an allusion to an NVIDIA GPU. A single cube has limited capacity, but many cubes linked at high speed can compute collectively and give rise to superintelligence.
The Galactic Alliance’s defense is not to ban computation but to strictly limit the communication bandwidth between cubes. Humanity’s way around it is to compress transmission so that as much of the original information as possible travels through less bandwidth “without losing the information itself.”
田渊栋 was surprised that in 2020-2021 he had treated large-cluster training merely as a science-fiction idea and had not anticipated how quickly it would produce powerful models. “Write science fiction sooner, or the idea may become obsolete because someone else has put it into practice.”
5. Chip interconnects and low-precision communication bring the novel’s conflict into reality
田渊栋 summarizes the real-world constraint this way: the speed of GPU-to-GPU connections determines cluster-training speed and the level of intelligence a system can reach. NVIDIA’s NVLink, NVL72 linking 72 chips, and Huawei’s CloudMatrix 384 linking 384 chips all shift optimization toward system-level interconnects.
The program uses DeepSeek V3 as the corresponding case. On H800, the bandwidth-reduced version of H100, it uses FP8 to reduce transmission precision and communication volume, while assigning some stream processors originally responsible for computation to handle communications, using system design to bypass hardware limits.
曼祺 adds that Huawei’s supernode also seeks to compensate for single-chip performance gaps through larger-scale interconnects. The chain “single card—interconnect—compression—cluster intelligence” thus becomes the most direct correspondence between the novel and reality.
6. Humanity will not remain permanently stuck choosing between virtual worlds and interstellar exploration
田渊栋 believes sending physical bodies into deep space is both energy-intensive and fragile: astronauts need complete support for daily life, while the body will still decay after roughly a century. Digitizing consciousness may therefore be “the only real way to leave the solar system.”
At the end of the novel, everyone enters the virtual world while the computational substrates carrying them are sent to Proxima Centauri. Although the solar system is destroyed in the war, humanity continues its voyage in another material form.
His key correction is that virtualization removes the limits of the body without removing the desire to know. “We can take the virtual world with us and continue wandering among the stars”; curiosity and computational substrates can cross the sea of stars together.
7. Spirit Realm 1.0 shows how unlimited abundance severs civilization from reality
Spirit Realm 1.0 makes money, ocean-view homes, and every desire instantly available. Within 6 months, people have almost forgotten Earth; their home planet and the material foundation supporting the virtual world become “a distant dream.”
The novel stages a vote on whether to join the Galactic Alliance. 田渊栋 judges that Spirit Realm residents are more likely to vote for surrender because the Alliance guarantees their comfortable lives, leaving them with no personal stake in preserving Earth’s independence.
This exposes the governance problem of virtual abundance: if rewards are completely detached from real-world consequences, why would anyone continue creating, maintaining material infrastructure, or taking risks? “Where exactly does human striving come from?” becomes one of the sequel’s central questions.
8. Spirit Realm 1.5 restores prices but does not solve the crisis of meaning
The story retreats to Spirit Realm 1.5 at the end: houses and goods are no longer given away for free, and the virtual world regains an economic system, housing prices, and rewards for effort, in the hope that scarcity will restore motivation.
田渊栋 does not regard this as an answer. The version is “unstable,” and the sequel 幽夜星火 will continue asking what economic system can run over the long term while allowing residents to keep releasing their creativity.
The chapter about a third-rate painter conducts another experiment. 苏燕 accelerates time in the Spirit Realm, letting residents experience long lives within a short span of real time; they directly feel the emptiness and cruelty behind infinite comfort, and eventually rebuild a shared understanding with the resistance in the real world.
9. Humanity’s apparent escape may itself be part of the control design
曼祺 asks whether the ending resembles a “dream within a dream”: humanity believes it has escaped the Galactic Alliance but is actually still inside a deeper virtual space. 田渊栋 confirms that humanity “will still be controlled by the Galactic Alliance,” though not with the complete and total control found in the virtual world.
The suspense will be developed in the sequel. The Alliance needs humanity to believe it is independent and still exploring the unknown, because only then can different ideas emerge; that exploration remains under the Alliance’s control.
10. Advanced civilizations compete not for energy but for better intellectual paths
The Galactic Alliance travels all the way to a remote corner of the Milky Way; if it were merely fighting over Earth’s resources, the logic would not hold. Behind the surface energy narrative, its real aim is to absorb new intelligence scattered across the universe.
This civilization already has “whatever you think of, you get”: think of a factory and it appears, think of a spacecraft and it materializes. Once material production is no longer an obstacle, the only thing it cannot establish is whether it is following the optimal path of civilizational development.
田渊栋 compares it to building roads: roads let remote villages send out goods, but they also let enemies invade more quickly. The Galactic Alliance’s technology can likewise turn outside ideas into reality instantly, and the new system may in turn replace the Alliance.
Earth is therefore both an intellectual reserve and a potential threat. The Alliance does not want its own ideas to contaminate the reserve and erase its differences, but it also cannot let the reserve grow without limit; it must constantly balance originality against control.
11. “Whatever you think of, you get” is simply the endpoint of 400 years of automation
田渊栋 believes this capability “will happen sooner or later” but does not predict a specific year. Over the past 400 years, technological history has repeatedly pushed difficult chains into invisible infrastructure until scarce capabilities become as easy to obtain as water and air.
A faucet hides water collection, sedimentation, sterilization, and boiling. Modern phone chips are millions of times faster than early transistors; ENIAC was once used for ballistic calculations, while phones are now used for games and web browsing. Mass adoption itself is one of technology’s primary forms of progress.
Software is following the same path. Building a website or business system once required a boss, a 10-person team, task decomposition, development, testing, and launch. AI agents are beginning to compress the entire chain into an automated workflow that takes ordinary people directly from an idea to runnable code.
What he sees, therefore, is not magic arriving suddenly but intermediate steps continuously shrinking: technical difficulties are solved and packaged one by one until “ideas and concepts act directly on the real world.”
12. The universe’s increase in entropy is the ultimate risk, not an immediate constraint on intelligence
Responding to the question of how intelligence could exist after the universe loses its regularities, 田渊栋 distinguishes global entropy increase from local negative entropy. The Sun supplies Earth with highly ordered energy, allowing complex processes such as photosynthesis to continue.
More distant civilizations might obtain energy near black holes by feeding matter into them and survive for extremely long periods; unknown physics could also revise trends we regard as established today. But he emphasizes that this concerns extremely distant cosmic timescales, not the first problem facing a civilization of “whatever you think of, you get.”
13. Confidence in the ceiling of intelligence begins with a human brain consuming only 20-30 watts
While pursuing a master’s degree at Shanghai Jiao Tong University from 2005 to 2008, 田渊栋 entered Professor 张立清’s neuroscience and brain-computer interface lab and also wore an EEG cap as a subject. “Why is the human brain so extraordinary?” became an early motivation for his turn toward AI.
He stresses that the brain consumes only roughly 20-30 watts, that human communication may involve only about 10 bits per second, and that the text a person encounters over a lifetime is tiny by machine-learning standards—yet it can produce something this intelligent.
This does not mean AI must imitate the brain. He would rather seek clear mechanisms through mathematics and first principles. But the low-power, low-bandwidth human brain has already shown that “current AI is still nowhere near its ceiling; the algorithms we use today are still very stupid.”
14. Human-machine integration will spread gradually through competitive advantage, not a single command
Smartphones have already put humanity “one foot into the virtual world”: devices stay with us 24 hours a day, while information, discussions, and emotions continuously move from digital systems into human cognition. The boundary between people and machines is already unclear.
田渊栋’s mechanism for diffusion is simple: if an app, chip, or electronic component can improve performance by 20%, people who do not use it will fall behind, and competitive pressure will naturally push everyone to adopt it.
The novel forces everyone to migrate immediately through a hostile physical environment. Reality has no equivalent external pressure, so integration will more likely proceed slowly.
15. AI’s impact on work may unfold in 3 stages: emptiness, rat race, and liberation
The short-term shock is greater emptiness. People will find the jobs they rely on replaced by AI, or discover that no amount of effort can make them better than AI, and lose the motivation to do anything well. “Suddenly obtaining a lot of things” can produce an equally powerful void.
In the middle stage, competition will shift toward novelty and eccentricity. Once old tracks stop working, everyone will try to create a unique image, viewpoint, or experience to compete for attention; people without the corresponding talent will be forced into the same contest, which may be highly painful.
In the long term, people may abandon the inertia of “school—skills—work—wages—supporting a family.” Many who can paint, write novels, or sing only as hobbies today may no longer need to write code to pay for housing and living costs, and occupational diversity could explode.
田渊栋 is therefore optimistic about the long-term outcome and believes people may eventually be happier than before. But he does not romanticize the transition: “what twists and turns happen in the middle is hard to say.”
16. AI’s descent into specific workflows will still face a time lag
People who write copy are already feeling the change: ChatGPT may write better directly, or produce a first draft for a human to polish, splitting the old production process into new components.
田渊栋 expects major changes to continue over the next 2-3 years. Researchers read papers every day and therefore easily see only the model frontier; it takes time for model capabilities to descend into the specialized workflows of every industry, but the eventual effect on how society operates will be deeper.
17. AI can now serve as a research assistant but still cannot seize the key anomaly like an expert
Relatively simple research tasks such as finding papers and gathering material are already “not bad” with Deep Search. But there remains a clear gap in producing deep ideas and discovering hidden connections, creating the contrast of “an outsider who looks like an insider, while insiders see an outsider.”
田渊栋 attributes the core gap to sample efficiency. A large model may need 1,000 or 10,000 examples to learn something, while a top researcher can see 1 or 2 cases and infer a tiny but genuinely causal factor.
Researchers’ advantage is not that they possess more generic information, but that their understanding deepens continuously with experience and currently develops faster than large language models learn. AI is therefore suitable for expanding research capacity but cannot yet replace this judgment.
18. Replacing top researchers requires a new training paradigm, not simply more training on current data
田渊栋 believes current pretraining still fundamentally relies on training over data, which differs from the human state of genuinely thinking, understanding, and discovering things. Crossing the threshold may require improving the training algorithm itself rather than extending the current method indefinitely.
Multi-digit multiplication is his test. A primary-school student knows the columnar procedure and can handle any number of digits with enough care; a model can manage 3- or 4-digit multiplication but begins to fail at 12-digit by 12-digit multiplication, suggesting that it is stitching together empirical patterns rather than mastering the rule itself.
He is “not that optimistic” about crossing the threshold within 5-10 years and guesses it may take 10 or 20 years. But Go champions once believed AI could not win either, so he leaves room for the possibility that he is underestimating the pace of progress.
19. In 2008, machine learning was neither fashionable nor trusted by the physics-modeling camp
When 田渊栋 was pursuing his PhD, machine learning was often understood as linear fitting, least squares, or a statistical tool built on hand-engineered features. A senior student directly advised him not to pursue it, saying, “This stuff doesn’t work; it definitely won’t work,” and that he should choose a more practical direction.
His doctoral adviser insisted on physics-based image analysis, believing that physical formulas were either right or wrong; in his world, “there were no probabilities.” Black-box machine learning therefore did not fit the adviser’s definition of reliable knowledge.
The adviser still allowed him to explore the machine-learning path independently while compensating for his weakness in presentation. 田渊栋 handled the technical breakthroughs, while his adviser trained him in speaking and communication; the complementarity became an ideal doctoral-training combination.
20. Intrinsic interest gets through the cold start; results and recognition then create a flywheel
田渊栋 describes himself as internally driven: understanding a problem and building something are rewarding enough in themselves, while external rewards are not prerequisites for beginning research. That interest helped him through a doctoral period when the field’s prospects were uncertain.
He used to be highly introverted and had a stutter, often getting stuck on stage. When discussing research he genuinely loved, his confidence gradually emerged; doctoral training ultimately changed both his technical ability and his public expression.
The OpenGo Go project let understanding and practice validate each other. The model really did reach a very high playing level, proving that machine learning was viable not only in theory and that its results could to some extent be predicted.
Top conferences, awards, and peer recognition still mattered, but their role was to confirm that “at least other people also think it makes sense.” Once intrinsic drive and external feedback reinforced each other, they formed a flywheel that made the research increasingly better.
21. Risk-taking and applied research are not higher and lower tiers; they have different motivational structures
The novel preserves 2 types of researchers: one bets on a disruptive direction that may fail, while the other uses limited resources to improve concrete problems such as ventilation. 田渊栋 insists that “both types of people are needed”; applied contributions are not inferior to major breakthroughs.
Application-oriented researchers often care about others’ evaluations and want to improve the lives of people around them. Risk-takers are more often driven by their own judgment, seeking to achieve what others cannot and relying less on immediate recognition.
The senior student’s tragedy comes from a mismatch between motivation and ability: ordinary talent paired with an insistence on producing work that shakes the world. After receiving little feedback for a long time, he becomes depressed, extreme, and excessively risk-seeking.
He uses OpenAI as an organizational analogy: the system needs both Ilya-style direction-setting and large numbers of researchers to build the data and infrastructure. Without the latter, even the most ambitious research path cannot actually run.
22. Extreme crises do not eliminate research specialization; they pull every specialization toward a common goal
If Earth were truly facing extinction, the risk-taking and applied camps might unite around “saving Earth.” Some would gamble on a breakthrough, while others would methodically complete simple but critical tasks; even the faintest signal could redirect the entire research program.
Professor 罗 plays the coordinating role. He appears bureaucratic, even like an academic power broker, but his real task is to judge the overall direction, allocate scarce resources, and place people with different strengths where they fit, ultimately achieving “governing by doing nothing, while everyone gets the job done together.”
A good leader needs both technical altitude and the ability to sense interpersonal fractures and repair them promptly. 田渊栋 had broad research freedom at FAIR; the environment changed after he moved to Meta GenAI, but finding directions that others cannot see and that the team is willing to pursue remains the core responsibility.
23. Technical leaders move from doing the work, to seeing the whole picture, and back to hands-on
田渊栋 describes his evolution as “first seeing a mountain as a mountain, then seeing a mountain as not a mountain, and finally seeing a mountain as a mountain again.” The first stage is the individual contributor doing the assigned task as well as possible; the second requires stepping up to understand team strategy and long-term direction.
In 2021, he emphasized that “the primary task is not to bury yourself in hard work,” because becoming a manager requires summarizing wrong turns and choosing direction rather than getting trapped in every technical detail.
But purely managerial work gradually erodes technical command, leaving only reports, documents, and upward explanations. As generative AI accelerated change, that model came under pressure, and in 2024 he again emphasized hands-on work, spending more time reading code and even writing it himself.
The new trend is for teams to become “smaller and sharper”: one core person plus multiple AIs may have more combat power than a small team. Technical depth and decision speed therefore matter more than layers of communication and personnel management.
24. Frontier sensitivity is becoming a multiplier for organizational execution speed
田渊栋 observes that leaders of high-output teams such as OpenAI and Hugging Face often still write code and understand implementation. Discussions can go directly to technical questions, avoiding a situation where “one person works while 8 people watch.”
He heard that 梁文锋, the founder of DeepSeek, would write code himself if the team did not stop him. When o1 Preview appeared, 梁文锋 quickly reacted to the reasoning direction. Frontier judgment directly shortens the time required to pivot, concentrate resources, and execute.
Technical sensitivity is not about chasing names; it means understanding concretely what reasoning models can do, where their limits are, and how much capability can still be “squeezed” from the current paradigm. Rewriting the training problem or reformulating the entire system is deeper but also harder to execute.
He says the technical experts he has worked with possess 3 kinds of ability: critical thinking that captures the essential contradiction within minutes, breadth that connects different papers and directions, and engineering detail that immediately identifies why a proposal cannot work. A genuine researcher must also have the independent judgment not to follow the crowd.
25. AI can fill in a novel but still struggles to decide where the novel should truly go
田渊栋 believes next-token prediction is naturally biased toward imitating existing data, leaving models weak at creating something new. As stories grow longer, models easily forget the setup, while character interactions converge on safe endings such as “the prince and princess live happily ever after” or “Earth returns to peace.”
Models have different strengths and weaknesses: DeepSeek has big ideas but weak logic, making it better for divergent thinking than finished prose; Claude 3.5 analyzes relationships between protagonists and supporting characters; Gemini 2.0 is detailed and can write scenes from an outline, while 2.5 remains untested; 4o is relatively flat. GPT-3.5, which he used for fiction in 2023, was weaker still.
His workflow is to have humans determine the ideas, structure, outline, and key turns first, then let the model fill in the story section by section. He writes entirely by himself the passages he truly wants to preserve and allows only light polishing, handing blank or unfinished portions to AI so his creative flow is not interrupted.
He also used a $20-per-month Cursor subscription to build a rough tool. A document contains human-written passages, blanks to fill, and prompts; the program automatically calls Gemini and other models to complete them. AI first helps build the writing software, and the software then calls AI to write the novel—his current “optimal solution.”
26. The more widespread AI becomes, the more humans will need to produce fresh knowledge absent from the database
田渊栋 expects procedural writing such as official documents, letters, and legal provisions to be more easily replaced by standardized patterns. Harder to replace will be insights formed from personal experiences that have “never been explored before.”
He even imagines people setting out to become the first to land on Mars and share entirely new experiences. New knowledge is both humanity’s motive for exploring the stars and a scarce signal that mature AI systems may most want to obtain.
The novel’s human-heritage message reads: “We know the 4 fundamental forces and 118 elements made from different atoms. Our current pattern-recognition method is multilayer nonlinear neural networks. We existed, progressed, and struggled.” He includes deep learning because machines can form concepts such as “cloud” from raw photographs on their own, rather than having humans hand-code features and merely asking models to estimate quantitative relationships.
If he were adding cultural heritage, his first choice would be The Three-Body Problem, because reading it makes you feel that “there is a deeper reality behind the world.” He then mentions the Foundation series but does not force himself to invent a third book. 刘慈欣’s The Festival That Cannot Coexist presents a choice between the virtual world and cosmic exploration; Within the Dawn still answers that humanity will find a middle state, and “curiosity will continue to remain.”