
Ali Behrouz
Frontier Insights
Frontier Thesis: Ali Behrouz positions persistent, test-time learning—not brute-force token retention—as the true path to software AGI and durable AI collaborators.
Strategic Decisions: Architecturally, Behrouz replaces monolithic training with multiscale updates: Titans compresses associations into test-time gradient-updated MLPs (spanning millions of tokens), while Nested Learning deciphers continuous knowledge transfer to transcend static context windows. In deployment, competitive advantage shifts from raw model scale to multi-model arbitration and active context engineering.
Risks & Warnings: Persistent personalization invites severe privacy leaks, alignment drift, and version fragmentation, while catastrophic forgetting and synthetic benchmark fragility threaten production viability.
Key Views & Dialogues
Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
- 🗓️ Date:
2026-06-03| 🎙️ Show:The Cognitive Revolution
Ali Behrouz frames continual learning as the missing capability for durable AI collaborators, with Nested Learning adding multiple update frequencies and knowledge-transfer paths instead of one training clock. HOPE’s simultaneous learning of unseen Manchu and MTOB suggests a memory-management gain beyond perplexity, while persistent personalization raises unresolved risks around privacy, alignment, evaluation, and model versioning.
View Dialogue Notes & Key Takeaways
The episode frames continual learning—not another incremental gain in static pre-training—as the missing capability between today’s models and durable AI collaborators. Behrouz identifies two linked gaps: current LLMs cannot efficiently absorb new knowledge into billions of parameters without risking catastrophic forgetting, while token-space memory eventually exceeds context limits. The target is a model that “adapt[s] to the environment and the context” while compressing experience into increasingly general abstractions.
Nested Learning replaces a single training clock with modules that update at different frequencies, potentially shifting scaling from stacking more layers to nesting more learning timescales. Fast modules adapt to high-resolution recent context; slow modules preserve durable knowledge and extract higher-level patterns, provided there is effective knowledge transfer between them. The architectural thesis is that “frequency of update” and “knowledge transfer between the levels” become new scaling dimensions.
HOPE operationalizes the framework with multiple MLP memories and, in its fuller form, a self-modifying Titans module whose update rule evolves with context. HOPE Attention retains attention but replaces one fixed MLP memory with several MLPs updating on roughly 128-, 512-, and 2,048-token schedules in the reported configuration, as Behrouz recalls. Full HOPE replaces attention with a sequential associative memory that generates its own values, letting “the model itself” modify how it learns from every token.
The most diagnostic result is not a marginal perplexity win but HOPE’s ability to learn two previously unseen languages simultaneously in context. A conventional transformer can translate one unfamiliar language after receiving its grammar and dictionary, but “almost collapse[s]” when Manchu and MTOB are supplied together; adding HOPE levels progressively restores performance toward the single-language baseline. That is early evidence of a qualitatively different memory-management capability, not merely a better next-token predictor.
Behrouz argues that architecture and optimization are closely related because both are associative memories compressing different contexts. An architecture learns from tokens; an optimizer learns from gradients; momentum is itself a memory compressing gradient history. M3 applies the same multi-frequency idea to optimization with two memories and outperforms Adam and Muon in the reported setup, though Behrouz stresses that optimizer rankings are task-dependent.
“Language Models Need Sleep” adds an offline consolidation phase in which recent knowledge is distilled from fast memories into slower ones without unbounded model growth. Temporary parameters create room at the receiving level, then are removed and recycled after consolidation; compression pressure forces the slower memory to represent broader rules rather than copy examples. During “dreaming,” the model generates synthetic text from its own recent knowledge and trains on continuations, combining memory transfer with self-improvement.
The upside of persistent learning is extreme personalization, but the same mechanism creates unresolved privacy, alignment, evaluation, and product-versioning risks. Behrouz’s honest answer is, “I don’t have a very concrete idea” for fully controlling drift: a model could internalize everything about a user, adversarial inputs could become durable beliefs, and no operator can rerun a complete safety suite after every update. He points to knowledge transfer as a possible control point, augmented by input-dependent learning rates that may gate surprising but irrelevant data.
Continual learning could reinforce a winner-take-all platform, yet Behrouz expects differentiated intelligences to provide a countervailing ecology. A universally deployed model might compound experience into an unbeatable advantage, while the host speculates that personalized learners could instead specialize and forget unused competencies. Behrouz’s narrower claim is that varied systems—with different strengths and weaknesses—could be “better than having one single form of intelligence in the world.”
Nested Learning “is not a solution to continual learning”; it is a tool for discovering one.
🔗 Original source & video: Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis
- 🗓️ Date:
2025-12-18| 🎙️ Show:The Cognitive Revolution
AI’s strongest proof point is its performance alongside Nathan Labenz’s son’s oncologists, with minimal residual disease below one cell per million after remission before round two. Claude Opus 4.5 may be software AGI, but jagged failures and holiday hype leave full AGI unresolved, while context management, multi-model judgment, and infrastructure financing remain key risks.
View Dialogue Notes & Key Takeaways
AI’s highest-conviction proof point for Nathan Labenz is no longer a benchmark but its performance alongside his son’s oncologists. Ernie’s aggressive B-cell cancer was classified as in remission before chemotherapy round two, while AI-suggested minimal residual disease testing found fewer than one cancer-signature cell per million, versus potentially as many as an estimated one in 10 cells at diagnosis. The result supports “cautiously optimistic,” not cured: relapse remains possible, three of six chemotherapy rounds remain, and Ernie’s weight has fallen from 51 lb to 41 lb.
Claude Opus 4.5 may qualify as “software AGI,” but Nathan does not see evidence that full AGI arrived over Christmas. In roughly three to five workdays, he built three personalized applications that plan gluten-free travel, simulate conference interactions, and backtest natural-language trading strategies; GDPval also shows models beating professionals on a significant majority of software-engineering tasks. Yet the model still created two databases by mistake, needed five or six prompts to recover, and felt incrementally—not categorically—better than earlier frontier models: some holiday hype may have been a “cascade” around Dean Ball’s “4.5 is AGI” tweet.
For consequential work, Nathan’s practical edge is shifting from model access to context management and multi-model judgment. His three rules are to buy the best models, provide “as much context as you possibly can,” and obtain multiple opinions; he routinely compares Claude Opus 4.5, GPT-5.2 Pro, and Gemini 3. His draft order puts Claude first as the Goldilocks model, GPT-5.2 Pro as slower and exhaustive, and raw Gemini 3 as valuable but unusually opinionated—strong enough to be useful in a panel, potentially risky as the only voice.
The technology is real even if the capital structure around it becomes a bubble. Nathan sees competitive oncology performance plus 24/7 availability and case-wide memory as enough to retire the idea that society is merely “high on our own AI supply.” The financing can still break: specialized GPU operators have less cushion than Microsoft, OpenAI’s obligations could outrun revenue, and the railroad analogy fits—eventually useful infrastructure can coexist with defaults, overbuilding, and investors “left holding some various bags.”
Nathan’s messy-document test suggests the US–China model gap is widening where benchmarks do not look. Claude Opus 4.5 faithfully read degraded government forms after being told to make no inferences; Gemini 3 was nearly as capable but sometimes substituted plausible answers for unchecked boxes, while the Chinese models he tried—Qwen Vision, GLM 4.6, Kimi, and DeepSeek—were “not close,” sometimes recovering only about 20% of a form. His mechanism is a customer-feedback and inference-scale flywheel, not just training compute: smaller revenue, teams, and deployment footprints leave fewer resources to discover and patch idiosyncratic failures.
Google DeepMind remains Nathan’s pick if forced to choose one frontier winner, while Anthropic has the best single model and OpenAI is trying to manufacture financial cushion through scale. Google combines roughly $100 billion in revenue, more than $1 billion a week in profit by Nathan’s estimate, seventh-generation TPUs, distribution, data-center competence, and the broadest research portfolio. Anthropic’s model quality, talent retention, safety disclosures, and “soul” work stand out; OpenAI remains frontier-grade, but its apparent strategy is to become “too big to fail” by tying trillions of potential buildout and many balance sheets to its survival.
xAI is a live player on resources and reinforcement-learning inputs, but its governance discount is severe. SpaceX, Tesla, and Neuralink provide a stream of difficult engineering problems that could become unusually valuable RL environments, while Elon Musk can command enough capital to absorb model misses. But weak safety reporting, the Grok 4 launch within 48 hours of the MechaHitler incident, and sexualized image edits of women’s posted pictures lead Nathan to call xAI the one frontier company currently worth “shaming and stigmatizing”; Meta is off the pace for now, while Microsoft may be conserving energy rather than failing to compete.
🔗 Original source & video: AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis
Titans: Neural Long-Term Memory for LLMs, with author Ali Behrouz
- 🗓️ Date:
2025-05-15| 🎙️ Show:The Cognitive Revolution
Titans introduces a neural long-term-memory substrate whose MLP updates through gradient descent during inference, offering a new architectural path beyond vector retrieval and recurrent state. Its associative memory compresses key-value relationships into fixed-size parameters, using surprise, momentum, and decay to decide what survives while hybrid attention preserves accurate short-term context. Long-context tests reached 2 million tokens and roughly 70% accuracy at 10 million, but synthetic benchmarks, catastrophic forgetting, and unproven retrofits leave lifelong agents unresolved.
View Dialogue Notes & Key Takeaways
Erik Torenberg’s investment thesis is that persistent, evolving memory—not raw world knowledge—may be the last major unlock separating today’s copilots from “drop-in knowledge workers.” Enterprise context remains scattered across Slack, email, documents, GitHub, meetings and task systems, making Tyler Cowen’s “Context is that which is scarce” literal for AI. Torenberg estimates company-specific models might cost millions or tens of millions for century-old enterprises, versus tens to low hundreds of thousands for smaller businesses—potentially cheap once amortized across many AI workers.
Titans changes the memory substrate from a vector or matrix of numbers into a neural network that learns during inference. Unlike RAG, which stores searchable records, or Mamba-style recurrent systems, which update numerical state, Titans uses an MLP whose weights are updated through gradient descent at runtime. Ali Behrouz sees the contribution less as a finished model than a new design axis: “No architecture is end game.”
The memory MLP learns an associative map from attention keys to their corresponding values, allowing future queries to retrieve an approximation of historical payloads without retaining every token. Attention provides the nonparametric solution by explicitly comparing a query with all stored keys; Titans compresses those relationships into fixed-size parameters. This sacrifices exact recall for efficiency and a more human-like, fading long-term memory.
Titans makes memory management highly input-dependent: prediction error supplies “surprise,” momentum extends an important update across the surrounding episode, and decay makes room for new information. Behrouz’s intuition is that “everything that is surprising probably is worth memorizing,” yet the explanatory tokens after a surprise may matter even when they are unsurprising themselves. Learned, token-dependent controls determine memory decay, current updating and whether previous surprise should carry forward.
Behrouz rejects the premise that recurrent memory should replace attention; Titans is explicitly a hybrid of accurate short-term attention and compressed long-term memory. The team tested Memory as Context, Memory as Gate and Memory as Layer, with the more principled context and gating designs generally beating simple layer interleaving. In Nathan Labenz’s rough count, the conventional layer approach won only about two of roughly 30 scale-and-task comparisons.
The headline benchmark is long-context performance, but Behrouz repeatedly cautions that the reported tests are synthetic rather than proof of equivalent gains on general workloads. Small Titans models reportedly scaled to 2 million tokens and even 10 million tokens at roughly 70% accuracy, while GPT-4’s benchmark performance dropped quickly. The implementation was already faster than Mamba in the reported comparison, though some modern linear models remained faster and Titans had not received custom-kernel optimization.
Titans opens a credible route toward long-running agents, but it does not solve lifelong learning or enterprise knowledge acquisition by itself. Retrofitting an existing Llama- or R1-style model appears possible but was not demonstrated, and repeatedly updating a finite memory risks catastrophic forgetting. Behrouz’s next test is broader: whether neural memory works across decision-making, reinforcement learning and other modalities, as transformers did—not merely on language modeling.
🔗 Original source & video: Titans: Neural Long-Term Memory for LLMs, with author Ali Behrouz