A Conversation with 马毅: Intelligence, DNA, and Entropy Reduction
Summary
- 马毅 reframes intelligence from “possessing knowledge” as the ability to extract patterns from the external world, form predictions, and autonomously update knowledge, making static DNA and large models merely the products of intelligent activity. “DNA was the world’s first large model,” but neither an individual DNA strand, a species, nor GPT-1 through GPT-4 is intelligent; natural selection creates a species–environment feedback loop, just as humans continuously improve models. This definition shifts the industry’s core bottleneck away from parameters and context length toward individual memory that does not forget and can continuously self-correct.
- The evolution of biological intelligence was not a matter of endlessly scaling “pretraining,” but of moving from species-level reinforcement learning to individual memory, shared language, and mathematical abstraction. Life relied primarily on DNA for roughly 3B years; neural systems and vision emerged about 550M years ago, followed roughly 50M years later by the Cambrian explosion, when individuals could finally fine-tune after birth. Birds stay with their parents for about 3 months, mammals 6–12 months, felines 2–3 years, monkeys 5–6 years, and humans about 18 years—the more advanced the intelligence, the more it depends on postnatal learning.
- Existing large models remain roughly single-cell-style open-loop systems, and scaling compute is already showing clear diminishing returns. 马毅 says Grok used roughly 200K cards to lift base-model scores by less than 1%; this trial-and-error process can keep improving, but it resembles life spending 3B years relying on DNA without crossing into individual learning. His formulation is blunt: “An open-loop system can only handle a closed world, no matter how large it is; only a closed-loop system can handle an open world, no matter how small it is.”
- The history of machine intelligence should begin with cybernetics, information theory, and game theory in the 1940s, rather than the 1956 Dartmouth conference. Wartime target tracking prompted Wiener and others to study how animals achieve rapid feedback and error correction, and 马毅 places modern computing architectures in the same lineage; Dartmouth addressed a later layer of human abstraction, logic, and problem-solving. By 1958, the perceptron was already being described by the media as capable of reading, writing, thinking, and even developing consciousness: “The hype being sold today was already being sold in the 1950s.”
- O1, R1, and other so-called reasoning models have improved performance on problem sets, but 马毅 has seen no rigorous evidence that they truly understand logic. He divides reasoning into three levels: imitating problem types, mastering and self-checking logic, and discovering new patterns. Existing methods remain largely at the first level: SFT supplies examples, reinforcement learning optimizes against answers or scores, and long chains of thought are generated semi-automatically by humans or existing models. A model that solves extremely difficult problems yet makes elementary mistakes is a signal that benchmarks still fail to distinguish “solving problems” from “understanding them.”
- DeepSeek showed that the current route lacks a strong technical moat, not that trial-and-error costs have disappeared. 马毅 had long argued that open source catching up with closed source was “a question of where and when, not whether or if not”; the real barriers are data, compute, time, and the cost of repeated failure. The cost of the final training run says nothing about the total cost of discovering the recipe. On whether to keep funding a single team, he offered a conditional warning: “Shells usually don’t land in the same crater twice,” unless the methodology has shifted from random trial and error to sustainable improvement.
- The next-generation architecture 马毅 is betting on is not a giant model that handles every modality, but a parallel, distributed, hierarchical system of multiple interacting loops. Embodied intelligence is the most natural killer app because an agent must continuously perceive, act, and correct itself in an open physical world. End-to-end VLA may achieve commercial results through more data and compute—like an airplane that can “fly, carry passengers, and make money”—without being as efficient as the mechanisms of a bird. 易生科技’s commercial task is to scale white-box principles, closed-loop memory, and new architectures to data, compute, and engineering levels beyond what universities can support.
- The value of technical taste comes from choosing non-consensus problems and constraining oneself with rigorous evidence, not from clinging to a contrarian narrative. 马毅 asks researchers to train themselves to be “the hardest person in the world to convince”; HKU’s entrepreneur-focused courses and its 6-credit AI literacy requirement for every undergraduate likewise focus on recognizing capability boundaries, negative examples, and promotional bias. His final advice: “Anxiety comes from not understanding; excitement at least requires knowing why you are excited.”
Deep dive
1. Intelligence Searches for Predictable Structure in an Entropy-Increasing Universe
马毅 begins with the second law of thermodynamics: closed systems move toward greater disorder and unpredictability, while life searches for structure in a world that still contains regularities—“fighting against entropy”—and turns what can be predicted into a survival advantage.
In his definition, an intelligent agent encodes regularities as knowledge and memory, uses them to predict what comes next, and then improves decisions and survival quality. “Life is intelligence” because life is the carrier of this mechanism.
曼祺 asks whether life and intelligence must inevitably emerge. 马毅 leans toward “accident,” but argues that the path toward increasing entropy contains abundant exploitable structure, making the emergence of life-like systems elsewhere in the universe “very likely.”
2. DNA Was the First Large Model, but the Species Was the Learner
马毅’s most distinctive analogy is: “DNA was the world’s first large model.” Early life encoded what it learned about the external world into the rule-like structure of base sequences, replicated it, and passed it to the next generation—functionally similar to a pretrained model.
This stage lasted roughly 3B years. Individual organisms largely relied on instinct throughout a lifetime, with no learning in the modern sense; random mutations generated candidates, the environment supplied the reward, and natural selection retained the better-adapted ones. “That is reinforcement learning.”
Intelligence therefore belonged not to a static individual but to the phylogenetic-evolution loop formed by a species and its environment. “One general’s success is built on ten thousand corpses”: failed mutations did not merely eliminate a sample; an entire species could be wiped out.
3. Neural Systems Gave Individuals Their First Ability to Fine-Tune
Roughly 550M years ago, neural systems, vision, and other perceptual capabilities emerged; around 50M years later came the Cambrian explosion. 马毅 connects the two: individuals were no longer limited to the “large model” inherited from their parents, but could form their own memory in a specific environment.
These memories were not written back into DNA, but encoded in the individual’s neural network. Visual and tactile inputs allowed animals to adapt to the particular conditions they encountered after birth—fine-tuning the inherited model through feedback and error correction.
Greater individual adaptability also allowed more morphological variations to survive, causing the number and forms of species to expand rapidly. The key was not simply that species became more intelligent, but that the learning timescale moved from across generations to within a lifetime.
4. The More Advanced the Intelligence, the Less It Relies on Prenatal Pretraining
马毅 illustrates the trend through the time offspring spend with their parents: about 3 months for birds, 6–12 months for ordinary mammals, 2–3 years for felines, 5–6 years for monkeys, and about 18 years for humans. The more advanced the life form, the larger the role of postnatal learning.
These figures support his counterintuitive conclusion: evolution did not hard-code an ever-greater share of capabilities into DNA. Instead, it left individuals with longer learning periods—less reliance on pretrained models and greater reliance on memories formed through the environment, parents, and groups.
5. Language Turned Individual Experience into a Reusable Collective Asset
Language first improved the efficiency of acquiring knowledge: if one person found a water source, others did not need to retrace the route and could simply listen to the description. When a feline dies, most of its experience dies with it; human experience can be replicated across a group.
Writing then preserved knowledge across generations and assumed part of DNA’s function. A world model built by one generation no longer had to wait for genetic variation to transmit it; books and organized language could pass it on directly, accelerating civilization.
马毅 also puts language models in perspective: only a small region near the brain’s prefrontal area processes natural language, while most of the brain handles vision, sound, touch, and movement. Knowledge that can be written down is only a fraction of what humans learn through interaction with the physical world.
6. “Hallucination” Is Not a Universal Label for All Uncertain Knowledge
曼祺 asks whether language also brought more hallucinations. 马毅’s reservation is that hallucination still lacks a sound scientific definition, while real-world knowledge is not uniformly deterministic like mathematics; much useful knowledge is inherently probabilistic.
“It might rain tomorrow” does not guarantee rain, but it can still guide decisions. Almanac-style empirical forecasts may be wrong, but that does not make them all hallucinations. Knowledge cannot be uniformly labeled hallucination merely because it contains error; the scientific definition remains open.
7. Mathematical Abstraction Was an Unexplained Leap in the History of Intelligence
Around 3,000–4,000 years ago, humans began abstracting natural numbers from counting experience, then extended the framework to fractions, real numbers, imaginary numbers, and points, lines, planes, and space. 马毅 sees this as an elevation of empirical knowledge, but says we still do not know what mechanism appeared in the brain.
Once an abstraction exists, people can learn, imitate, and use it. The real mystery is how humans first generated a higher-order concept from experience. This is also where symbolic AI in the 1950s observed the phenomenon without solving the mechanism.
Mathematics and science make knowledge more abstract, concise, rigorous, and capable of generalization, but scientific theories remain falsifiable and revisable. They are not ultimate truths outside the learning loop, but more highly compressed and shareable forms of knowledge.
8. DNA, Neural Memory, Writing, and Mathematics Are Four Media for the Same Thing
马毅 unifies the four stages as expressions of predictable regularities in the external world: DNA encodes them in bases, individual memory in neural networks, civilization in language and writing, and science in highly compressed mathematics and formal language.
The learning mechanisms differ: species evolve through natural-selection-style reinforcement learning; individual memory develops through continuous feedback and adaptive correction; civilization advances through communication and preservation; science further abstracts and compresses experience while subjecting it to falsification and refinement.
Perceptual signals entering the brain may reach millions of bits per second, yet only a dozen or so bits per second remain at the higher levels. Intelligence does not preserve everything; it discards irrelevant detail, extracts structures that actually affect prediction, and revises deficiencies, inaccuracies, and incompleteness in existing memory.
9. Knowledge Is the Product of Intelligent Activity, Not Intelligence Itself
马毅 draws the boundary clearly: “Knowledge itself is not intelligence; knowledge is the result of intelligent activity.” Intelligence requires extracting new regularities from observation, identifying deficiencies in old knowledge, and updating autonomously; a system with only a static knowledge base lacks these abilities.
A DNA strand therefore has no intelligence, and neither does a foundation model, however large. He lists them one by one: “GPT-1 has no intelligence, GPT-2 has no intelligence, GPT-3 has no intelligence, and GPT-4 has no intelligence.” The process by which humans continuously improve models, however, is intelligent.
The episode’s focus on long context reflects the same judgment: rather than endlessly extending the amount of context a pretrained model can ingest at once, the more important task is to build memory that is closed-loop, durable, non-forgetting, and capable of self-correction.
10. Closed-Loop Feedback Is Not One Candidate Route Among Many, but a Necessary Mechanism
马毅 traces this judgment back to cybernetics. Early researchers observing animals already found that learning, adaptation, and better decisions depended on correcting errors after action. “We did not invent this”; nature has repeatedly demonstrated it as an intelligence mechanism.
曼祺 asks whether intelligence could be implemented in a way unlike life on Earth. 马毅 acknowledges that it could, and that natural selection may not be theoretically optimal, but says biology’s current solution remains far more efficient and autonomous than artificial systems.
His summary is: “An open-loop system is for a closed world, no matter how big it is; a closed-loop system is for an open world, no matter how small it is.” Even a tiny ant may face an open environment; a static model can be huge and still cover only a fixed, finite world.
11. Large Models Can Ace Benchmarks Without Proving They Have Abstract Concepts
In 马毅’s framework, current models remain mostly at the first stage, roughly like single-cell life relying almost entirely on inherited instinct. They store vast amounts of static knowledge, but cannot adapt to environments, form individual memory, and continuously self-correct like cats and dogs.
He argues that today’s foundation and large models have not truly abstracted the concept of natural numbers or mastered mathematical operations. After seeing many worked examples, a model can answer correctly, but in unfamiliar settings it may still guess or estimate when counting or understanding spatial relationships.
曼祺 notes that many people believe the Turing test has already been passed. 马毅 counters that fixed question banks and benchmarks cannot tell whether a student has memorized problem types or actually learned them. A real test should work more like a teacher’s oral examination: ask follow-up questions, change the framing, and observe whether the system can explain and self-check.
“The examples people sometimes see are all positive examples,” while counterexamples are abundant. 马毅 says AI evaluation still lacks rigorous methods to distinguish strong memorization, genuine understanding, and the ability to abstract and elevate.
12. Entropy Reduction Requires Energy, and Intelligence Is Only One Stage in the Universe’s History
If the universe is indeed a closed system with no external input, entropy will eventually reach its maximum, leaving no predictable structure for intelligence to learn. Under current physics, 马毅 says, a closed system ultimately reaches a state of chaos.
Image generation provides a direct example: a diffusion denoising model gradually turns noise into a structured image, essentially performing denoising and entropy reduction. The process requires electricity, chips, and substantial compute, injecting energy into the system from elsewhere.
Intelligence therefore cannot defeat overall entropy increase for free; it can only use resources to create local order. The episode closes with Blade Runner’s “All those moments will be lost in time, like tears in rain”: intelligence may be only one process in the long history of the universe, but it is a special process capable of temporarily resisting disorder.
13. Machine Intelligence Began in the 1940s, Not at Dartmouth
马毅 argues that treating the 1956 Dartmouth conference as AI’s starting point reflects a narrow boundary drawn by the computer field itself. Earlier work in cybernetics, information theory, game theory, mathematical neuron models, and computing architectures was already concerned with the mechanisms of intelligence.
The first question in the 1940s was how machines could acquire animal-level feedback, adaptation, and decision-making. 马毅 calls it something like a “Wiener test”: if a machine possessed an animal’s closed-loop capabilities, could humans still distinguish the two?
Turing moved the question to the human level only in 1950: how intelligent could a machine become before humans and machines were difficult to distinguish? The young researchers at Dartmouth in 1956 then studied more human-like traits such as causality, abstraction, logic, and problem-solving.
14. War Turned Animal Hunting Mechanisms into a Control-Theory Problem
曼祺 asks why animal intelligence became a major research topic in the 1940s. 马毅’s answer is direct: “War.” Artillery had to track fast-moving aircraft, so researchers watched how felines pursued targets quickly and accurately—and corrected course rapidly after an error.
These questions gave rise to closed-loop control: a system could not simply produce outputs according to a preset plan; it had to sense the target, compare the error, and adjust its next move. Information theory addressed how external signals are encoded and decoded, while game theory studied decision improvement in uncertain and even adversarial environments.
马毅 further argues that one important motivation for the von Neumann architecture was to implement the feedback and optimization mechanisms proposed by cybernetics. Early computers were used primarily for codebreaking, but modern computing architecture cannot therefore be reduced to “calculating for the atomic bomb.”
15. Dartmouth’s Technical Taste Came from Young Researchers Leaving the Mainstream
马毅 notes why the young researchers of 1956 did not stay at MIT, Caltech, or Berkeley: cybernetics and related mainstream fields were already occupied by authorities such as Wiener and von Neumann. They wanted to find questions that had not yet been claimed, so they gathered at Dartmouth.
They studied symbols, logic, causality, and higher-order reasoning, and many later became highly successful. 马毅’s reservation is that they described the phenomena of human intelligence in depth without explaining how abstraction actually emerges from experience and brain mechanisms.
This history also informs his standard for academic work: industry can compete on the cost and efficiency of known products, but if academia merely follows the mainstream and industrial resources, it has neither an advantage nor new knowledge to add.
16. The 1958 Perceptron Boom Had Already Previewed Today’s AI Hype
Once neurons had mathematical models, researchers quickly built perceptrons. 马毅 says roughly $6M was invested at the time, equivalent to “about $1B today”—a scale and narrative that did not resemble a modest experiment.
曼祺 found a 1958 New York Times report describing a machine that would learn, think, read, write, understand, know that it existed, and even replace human labor. 马毅’s assessment: “The hype being sold today was basically all sold in the 1950s.”
The boom later cooled, but that did not make neural networks worthless. Researchers had underestimated the requirements for organization, learning mechanisms, and compute. The recurring historical pattern is not total technological failure, but the premature extrapolation of capabilities.
17. From Cat Brains to AlexNet, Breakthroughs Came When Structure, Data, and Compute Matured Together
In the 1980s, researchers studying cats’ visual systems found that neural connections were not random but had regular local patterns, helping drive convolutional neural networks. Japanese researchers proposed related ideas early, and 杨立坤 achieved stronger results in 1989.
Hinton and others explored autoencoding, encoding-decoding, and statistical-physics methods without knowing how important the work would eventually become. It exemplified 马毅’s point that “the right question can precede effective engineering by many years.”
AlexNet in 2012 did not create an entirely new framework from scratch. It combined CNNs with sufficient data, GPUs, and deep networks that could be optimized, pushing ImageNet classification across a visible inflection point.
Neural networks subsequently expanded into images, sound, natural language, and proteins, but structural progress still relied heavily on trial and error. VGG, GoogleNet, ResNet, and Transformer became the celebrated names; many more candidates became casualties of “one general’s success being built on ten thousand corpses.”
18. Robots in the 1990s Could Walk, but They Were Still Blind
While pursuing his PhD in Berkeley’s robotics group, 马毅 had already seen Japanese robots run, jump, somersault, climb stairs, and dance. The mechanics and control were less nimble than today’s systems, but the basic principles and preprogrammed approach were not fundamentally different.
His adviser pointed out that these robots “had no brain and no eyes.” Stride length, leg height, and dance movements were all preplanned; the robot could not know whether it had performed correctly or adjust to the real environment.
This pushed 马毅 toward vision and 3D reconstruction: without perceptual input, there could be no closed loop. His 1997 master’s thesis studied vision-based navigation and visual driving; his 2000 doctoral work involved using vision to land helicopters and ships. The core question was how machines could perceive the external world.
19. Computer Vision Matured First Through 3D Reconstruction; Recognition Became Mainstream Only in 2012
3D vision had relatively clear problems, mathematical objectives, and technical routes, so it developed more smoothly. Its results were used for scene digitization, positioning, and recreating movie environments, including early Hollywood applications such as The Matrix.
马毅 places the point when 3D reconstruction truly became mainstream around 2008–2009. Recognition was still deeply unfashionable in 2007 and 2008: classification performance on databases was poor, and graduates struggled to find directly relevant jobs.
ImageNet supplied large-scale data and a unified testing platform, while AlexNet used GPUs and deep networks to demonstrate a performance jump. Recognition crossed the attention threshold in 2012. The decisive factor was not a single paper, but data, compute, and algorithms maturing at the same time.
20. Neural-Network Progress Is Random Trial and Error When First Principles Are Missing
马毅 admits it is difficult to extract a stable predictive rule from decades of ups and downs because the goal of intelligence, learning mechanisms, and first principles remained unclear for a long time. Research broadly followed two lines: learning from biological structures, or taking a purely engineering approach of building networks, running experiments, and filtering results.
This progress resembles DNA evolution: large numbers of structures are proposed, fail, and disappear, while an accidentally more effective structure survives. As long as the methodology remains unchanged, the team that produces the next result is highly random.
He uses artillery to describe the industry: one moment it is DeepMind, then OpenAI, Google, Microsoft, or Grok, and it could also be DeepSeek. Different teams repeatedly trial and error based on their own experience and resources; the next breakthrough could emerge anywhere.
When asked whether DeepSeek was still worth funding, he replied: “Shells usually don’t land in the same crater twice.” This was not a dismissal of the team, but a demand to see a methodological difference—unless it had established a sustainable, systematic mechanism for improvement.
21. Industry Bears the Cost of Trial and Error; Academia Should Supply Definitions, Principles, and Tests
The current route requires data, compute, and large numbers of failed experiments, a burden only a small number of companies can carry. That is why so many recent advances have come from industry rather than resource-constrained university labs.
马毅 says academia should not compete with companies on homogeneous engineering. It should define different stages and capabilities of intelligence, build tests that distinguish memory, understanding, and abstraction, and study the underlying mathematical and computational problems. “The past ten years did not define this clearly.”
His frustration is that many academics have also been swept along by corporate release cycles and media narratives, becoming little more rigorous than spectators. Companies may have commercial positions; academia’s responsibility is to explain methods, evidence, and limitations clearly.
22. Scaling Has Entered Diminishing Returns; O1 Only Temporarily Shifted the Anxiety
马毅 believes base-model scaling has entered “diminishing return” without a fundamental methodological change. His Grok example: roughly 200K cards produced less than a 1% improvement in base-model scores, so presentations can only keep magnifying the vertical axis.
The process will continue to improve, but it resembles early life relying only on inheritance and selection, which spent roughly 3B years and still reached only the single-cell stage.
曼祺 notes that the industry was already discussing scaling laws hitting a wall in September and October of the previous year, and that O1 eased the anxiety. 马毅 acknowledges the improvement on problem sets, but says OpenAI packaged the O series as a secret route to AGI even though the underlying method remains mainly fine-tuning and reinforcement learning.
23. “Reasoning” Has at Least Three Levels: Imitation, Understanding, and Discovery
马毅 breaks down a student solving a math problem into three levels. At the first, the student memorizes many examples and patterns, imitates them, and can still score highly. At the second, the student truly understands the logic, applies rules step by step to new problems, and knows whether each step is correct.
The third level is not using existing logic, but discovering previously unknown patterns, axioms, causal relationships, or reasoning methods from experience—like a mathematician discovering a new rule of inference, Euclid’s axioms, or the syllogism.
Current discussions call all three levels reasoning, allowing high benchmark scores to be extrapolated directly into “reasoning ability” or even a route to AGI. 马毅 says the capabilities must first be separated before anyone can say which level a model has reached.
24. SFT, Reinforcement Learning, and Chain of Thought Remain an Engineering Recipe
To improve a decent foundation model’s performance in programming or mathematics, the standard route is supervised fine-tuning: provide high-quality examples and have the model learn and imitate problem types, steps, and answer patterns.
Reinforcement learning then optimizes against signals such as answer correctness, step quality, or scores. DeepSeek emphasizes that direct RL can also work, but 马毅 says the prerequisite is already having a strong base model—and a strong base model often contains earlier example learning and fine-tuning.
He summarizes the division of labor in a team work title: “Supervised Fine-Tuning Memorizes, Reinforcement Learning Generalizes.” His general view is that fine-tuning followed by reinforcement learning is often better than using either alone, though each has its own strengths and limitations.
Long chains of thought do not emerge from nowhere. Graduate students can write out solution steps, or existing models can generate them semi-automatically under prompting; correct and efficient paths can then be selected for fine-tuning or scorer training. Teams adjust the ingredients “a bit like traditional Chinese medicine,” which is why the process is also called alchemy.
25. O1 and R1’s High Scores Still Do Not Prove They Have Learned Logic
曼祺 directly asks whether O1 and DeepSeek R1 are truly reasoning. 马毅’s personal judgment is that they remain mostly at the first level, improving problem-solving through memory, induction, and imitation.
He has seen no rigorous evidence that the models solve problems through genuine understanding. One striking contradiction is that the same class of model can complete extremely difficult math problems while making elementary mistakes at the primary- or middle-school level.
If a model possessed a rigorous logical chain and the ability to self-check, the two phenomena should not coexist so frequently. 马毅 therefore rejects conclusions based solely on competition problems or fixed benchmarks and calls for counterexamples and active testing.
He acknowledges that O1 and R1 deliver real performance gains. The disagreement is not over whether performance exists, but whether task performance can support strong claims about cognitive mechanisms. “It is convenient to call it a reasoning model,” but that does not make it equivalent to human reasoning.
26. DeepSeek Punctured the Secret Narrative Without Eliminating R&D Costs
马毅 believes O1’s SFT, RL, and chain-of-thought methods contain no novelty sufficient to surprise academia. He also believes OpenAI encountered internal problems over the past year and may have oversold the O series for reasons including fundraising.
DeepSeek approached O1 performance with a cheaper and more efficient implementation, showing that the supposed secret was not deep and that a decent base model could achieve similar results after enhancement with established methods.
But the cost of the final successful training run is not the total R&D cost. “The first pass is exhausting”; once the trial and error is complete, organizing and rerunning the process looks very clean. He understands that Google’s final training run may also have cost only several million dollars, somewhat more than DeepSeek, but that cannot be compared with the full cost of exploration.
曼祺 asks whether he expected DeepSeek to emerge from China. 马毅 says he had long expected an event of this kind, but not its location or timing. Had it not been China, it could have been France, the UK, or another team in the US.
27. Open Source Catching Up with Closed Source Is a Matter of Time Because Current Technology Has No Deep Moat
马毅 has advocated open source for the past 2–3 years and has long believed open-source models would eventually approach or surpass closed-source companies. “Where and when, not whether or if not” captures his view of the route’s replicability.
He believes current technology and methods do not have a durable moat. The main differences come from data, compute, time spent on trial and error, and team experience. A team may find a more effective recipe in one competition, but that is difficult to turn into a long-term exclusive advantage.
DeepSeek’s significance is therefore not that all model companies have lost their value, but that valuation has been pushed back toward the ability to innovate continuously. Whoever can turn random experience into an explainable, repeatable improvement mechanism may build the more durable advantage.
28. Technical Taste Begins with the Value Judgment of Which Problems to Choose
马毅 describes science’s mission as exploring the unknown and finding where existing understanding is wrong or insufficient. “We get paid to do that,” so conformity is not the purpose of academic work.
Industry competes on the price and efficiency of validated needs, creating social value. If academia builds the same models and runs the same benchmarks with fewer resources, it has neither a comparative advantage nor incremental knowledge.
Genuine taste is not asserting that one has found the optimal solution, but using evidence to know that “the problem matters, the direction is right, and the approach makes sense.” That conviction can sustain the work even when outsiders do not yet understand it.
29. Non-Consensus Only Avoids Becoming Pseudoscience When Combined with Rigorous Training
In scientific exploration, “of ten ideas, nine are probably unreliable.” Failure is normal; correct discoveries come not from emotional persistence, but from using mathematics, logic, experiments, and reproducible comparisons to eliminate errors step by step.
When 马毅 was studying for his master’s degree in mathematics, his teacher gave him the first rule: “Train yourself to be the hardest person in the world to convince.” When a proof can convince the most skeptical version of oneself, it is more likely to withstand everyone else’s scrutiny.
Experiments require the same rigor: how data is collected, how a hypothesis is tested, and whether alternative explanations exist cannot be handled by jumping to conclusions. Exploration without this training can leave someone wrong without any awareness of it.
Taste therefore requires two forms of capital that cannot be separated: a culture and value system that believes unknown regularities still exist, and the ability to judge whether one’s own work is actually correct.
30. 马毅’s Academic Path Was Shaped by the Need to “Figure Things Out,” Not Planned in Advance
At Tsinghua, he mainly studied step by step, while enjoying mathematics and extracurricular books. Near graduation at Berkeley, he was still considering Qualcomm or an internet company; only after his adviser suggested trying an academic position did he “stumble” into the university system.
His stable habit was that he did not dare teach what he did not understand. After moving to Illinois, he would rederive course material from the beginning. He later came to enjoy writing books because organizing a complete knowledge system exposed gaps in his own understanding.
The path suggests that technical taste need not come from a childhood ambition. It may emerge from repetition over time: staying curious about the unknown and turning vague intuition into structures that can be taught, proved, and verified.
31. Berkeley’s Open Collaboration Turned Individual Taste into Collective Validation
When 马毅 was pursuing his PhD, his adviser’s group had 18 students from 13 countries, with no obvious hierarchy. Everyone shared the goal of “figuring things out,” while different cultural backgrounds created culture shock and broadened the range of perspectives.
Group meetings were fully open. When there were only 6 or 7 formal students, attendance could reach 30 or 40, many of whom he did not know. Each PhD student could find a co-advisor genuinely involved in the research and freely attend other groups’ meetings and seminars.
Students across groups helped one another write papers, run experiments, and draw figures. Collaboration formed organically rather than through top-down assignments from professors. 马毅 later found that Berkeley students often learned more skills from peers than from advisers.
He has extended this culture into his work today: different teams handle theory, algorithms, implementation, and separate modalities, while recent projects often involve 5 or 6 universities. The willingness of many people to contribute is also an external signal that “we did something right.”
32. Highly Cited Work Can Still Face Collective Rejection at First
马毅 recalls a CVPR paper that received nearly perfect scores from reviewers but was still rejected by the area chair. Other white-box theoretical work was accepted by multiple reviewers but failed at the final decision.
His most cited face-recognition work initially found no conference willing to accept it because the results were “too good to believe.” Reviewers suspected they were impossible or fraudulent. The team eventually submitted to a top journal and spent a summer comparing the work against every requested method before it was accepted.
His conclusion is not that peer review is worthless, but that communities can form an echo chamber: mainstream methods are assumed correct, while a new method—however rigorous mathematically and comprehensive experimentally—may take a long time to be understood.
“Some things may have no presence for 30 years.” That delay requires researchers to distinguish between two kinds of feedback: rejection may expose a real flaw, or the conditions needed to accept a new framework may simply not yet exist.
33. Entrepreneurs Can Learn Technical Judgment, but Cannot Borrow Conclusions from Authority Alone
马毅 believes rigorous undergraduate-level mathematics and science are enough to understand most current AI technology; deeper knowledge is needed to explore the frontier. What entrepreneurs often lack is not intelligence, but a systematic framework and evidence on both sides.
HKU therefore launched an EMBA-like “scientific entrepreneur” program. It originally planned to admit 40–50 people but ultimately enrolled 80 in the first cohort, including leaders of major listed companies, investors, and traditional entrepreneurs seeking transformation. Attendance in the first modules was close to 100%.
One side of the course has frontline research faculty explain the essence, capabilities, and limitations of models; the other invites technology companies to discuss success, failure, and their actual positioning. The goal is not to make decisions for participants, but to help them understand the conditions for adopting, rejecting, or delaying investment.
马毅 particularly emphasizes negative cases invisible to the media. Demos usually show successful samples that “bounce and jump,” but decision-makers must understand the distribution of failures. Conversely, when speaking to an audience whose combined net worth may reach hundreds of billions of RMB, professors must also explain clearly what problem the research actually solves.
34. HKU Made AI Literacy a 6-Credit Requirement for Every Undergraduate
HKU replaced one of two existing English-language courses with AI literacy. The course carries 6 credits over 2 semesters: the first 3 credits are common across the university, while the second 3 are customized by faculties such as medicine and law.
The course piloted in spring with more than 100 volunteers and planned to require all new students to take it from September of that year. As the scale expands, each hour will be designed jointly by multiple teachers rather than repeated from a single lecturer’s notes.
The common portion begins with the history of biological and machine intelligence, then explains what language models, images, generative models, recognition, and robots are doing while presenting their capability boundaries. The professional portion connects these concepts to medicine, law, finance, art, and other fields.
The ethics module covers cheating, plagiarism, copyright, privacy, security, and legal norms, with participation from law, philosophy, and relevant engineering faculties. The content will not be recorded for permanent reuse, but updated every year as the technology changes.
35. The Next Opportunity Is in Closed-Loop Architectures, but Commercialization Still Requires Crossing an Engineering Threshold
马毅 says current AI remains essentially “very mechanical” data compression and generation, with no basis for equating it directly with consciousness. The endpoint of AI literacy is not to instill optimism or fear, but to give students evidence and their own critical thinking.
White-box research has begun replacing empirical redundancy in systems such as Transformer with mathematical principles. Faced with the heavy engineering burden and pretraining difficulty of vision and multimodal models, teams say systems such as DINO can be greatly simplified while improving performance. Companies contribute larger datasets, more compute, efficiency optimization, and sustained engineering teams.
Closed-loop research is moving from a single encoding-decoding loop toward a parallel, distributed, hierarchical multi-loop system resembling the cerebral cortex. Different perceptual modalities are processed separately and then integrated, with prediction and correction at every layer. Mechanisms such as strengthened connections through co-firing and local inhibition remain insufficiently implemented in current neural networks.
Embodied intelligence is the natural killer app, but end-to-end VLA may still achieve commercial results through data and compute: an airplane is not a bird, yet it can fly, carry passengers, and make money. 易生科技 is pursuing a more efficient next-generation mechanism. Before the technology crosses the threshold, it may be “of no use at all”; once it reaches a sufficient level, applications may emerge rapidly in large numbers. That is the uncertainty patient capital must bear.