Pioneers Insight Method Research Author
Claude Code Psychosis: How SemiAnalysis Is Token Mogging Meta | Ep. 008
Back to Episodes

Claude Code Psychosis: How SemiAnalysis Is Token Mogging Meta | Ep. 008

Summary

  • SemiAnalysis is turning Claude Code into an agentic research stack whose economics are attractive against manual analyst work. Dan’s Wags system initiated AOI coverage, returned a partially garbled contract-related result after roughly five minutes, and cost $3.52 for that run—perhaps $10–$15 including prior setup, versus having “an analyst spend like a day on it.” The ambition is covering 60–100 companies while humans focus on “why it happened and what it means.”
  • A major constraint is context architecture and durable memory, not merely model intelligence. Specialized agents handle transcripts, news, events, and financial models; because they can only communicate through the lead agent, each handles a discrete part of the process while a “pristine” company agent ingests their outputs. When subagents shut down, context disappears, yet memory files cannot grow forever. The unresolved failure mode was stark: a generated balance sheet did not balance, requiring supervisory checks and an intern whose job is to “whip the agent.”
  • Low task costs and compounding usage support strong token and GPU demand, but enterprise adoption remains the larger uncertainty. SemiAnalysis says it consumes more than twice Meta’s tokens per employee, and Dan would still run the demonstrated job at $15–$20. Yet IT approvals, weak incentives, and distance from customers may turn 50% productivity into finishing work 50% earlier rather than producing 50% more—so token spend need not map cleanly into corporate earnings.
  • Model quality has crossed from a careful expert tool to something users can address with “the same instructions that you would give to an intern.” Sam said software work a year ago demanded test-driven development, stepwise scaffolding, and tight context management; now he would introduce the product to parents or grandparents and expects “every finance bro, every lawyer” eventually to rely on it. The counterpoint is that adoption remains weak even though the output is already good enough.
  • Mythos might mark a cybersecurity leap, but Dan was unconvinced that raw capability alone explains it. He noted that coding and cybersecurity gains far exceeded broader improvements and questioned the “gigantic marketing campaign” around zero-days. He suggested persistence, post-training, harnesses, and reward signals may be the decisive stack. Dan also said many other models may be able to exploit a vulnerability once the finding is explained, while discovering it through a long-range task is different.
  • The model race may be converging while the real moat shifts toward interface, workflow history, and user inertia. Meta’s rushed-looking Avocado release was described as roughly in the ballpark of “Opus, 5.4, and a Gemini,” raising a “four-horse race” question. Repointing Claude Code might technically require one URL change, but users accumulate knowledge files, adapt workflows, tolerate fixable errors, and value Claude Code’s ability to control a computer rather than merely chat in a browser.

Deep dive

1. The intern now commands the research swarm

  • Dan introduced Terrence, an NTU AI-and-accounting major, working across five terminals and three screens. The running joke carried the organizational point: “The intern’s the boss now,” teaching agents to perform financial modeling rather than performing every task himself.

  • Wags, the team’s “agentic director of research,” supervises separate agents for earnings transcripts, news, events, company briefs, and models. Because the agents can communicate only through the lead agent, each handles a discrete part of the process. That division lets SemiAnalysis span networking, TCO, several switching-related companies, and dozens of optical-transceiver suppliers without forcing one context window to ingest every source and operating instruction.

  • The architecture keeps the company agent “pretty much pristine.” Other agents gather and summarize material, then save an index covering models, transcripts, and briefs; the company agent wakes up and ingests those outputs, retaining what matters without carrying “how we got to it” inside its context.

  • A live AOI query returned a partially garbled result mentioning “324 million,” “1.2,” and likely Meta and Amazon, mostly forward-looking; Dan said he would “have to check this against the actual thing.” The $3.52 run—perhaps $10–$15 including earlier initiation work—made his economic case: even at $15–$20, he would choose it over a day of analyst labor.

2. Discrete skills are compounding into a research operating system

  • Three months earlier, Dan was not working this way. A Chinese New Year slowdown created time to learn tools that could reduce the team’s workload; isolated requests—summarize an event, build a financial model, monitor press releases—became reusable skills for earnings, news, events, modeling, and monitoring, then components within Wags.

  • The rationale was mundane: Dan said there was still no reliable way to ingest financial data, and checking and fixing imported data can take longer than typing it in. His target is coverage of “60, 80, 100 companies” without creating so much baggage that analysts lose sight of the industry. Humans should spend less time on data entry and recounting events, and more on “reviewing, editing, thinking, drawing connections.”

  • Claudia, the conference agent, addresses the material nobody can physically consume. SemiAnalysis attends 50–60 conferences; GTC alone produced 621 sessions and 830 presentations, while the team attended “exactly zero” sessions because meetings took priority. Claudia transcribes presentations and YouTube links, indexes the conference’s speaker “talkability,” and can find talks about CPO, particularly the use of DWM for CPO. She can also assemble a briefing on ScalaCross for Julian, who then briefs Dan.

3. Trust improved faster than reliability and memory

  • The host dated a change in posture to roughly November, when SemiAnalysis began pushing people to use the tools: users previously distrusted every output, whereas they now tend to trust by default and verify when stakes warrant it. His emerging frustration is that “I am the constraint now”—the model wants to code, analyze, and build, but the instructions and task design limit it.

  • Dan’s counterexample was a generated balance sheet that did not balance. He had to investigate whether the agent mixed FactSet MCP data with SEC filings, and the team requires every completed model iteration to go through supervisory checks. The host’s pushback was precise: this is “not one that an analyst would ever make,” so treating agents exactly like junior humans obscures their distinctive failure modes.

  • Memory is the harder structural problem: “Once the subagent shuts down, like, that’s it.” Lessons must be memorialized, but an endless memory file is impossible. Sam experienced the inverse problem while learning about neoclouds—his knowledge lived inside Claude conversations—so he built a skill that moved those lessons into a flashcard app for deliberate retrieval.

4. Token demand can surge before enterprise productivity reaches earnings

  • SemiAnalysis says it consumes more than twice as many tokens per employee as Meta, a comparison the team called “token mogging.” Dan linked the demonstrated task economics to the GPU rental-price inflection: even if the total cost reached $10–$20, he would still prefer it to a day of analyst labor.

  • Broader adoption remains early. Dan sees fund managers using models for narrow summaries, while many Fortune 500 employees lack IT approval or workflows that compound over time. If Meta, Amazon, Unilever, P&G, and Boeing used Claude as deeply as SemiAnalysis, he argued, “the only way this gets relieved is like, honestly, a lot of chip fabs.”

  • Inside SemiAnalysis, more efficiency releases a “pent-up dam of ideas.” One employee reportedly ran Claude Code for 24 hours straight; other ideas became shareable Slack links within an hour, while an agentic coding-tracing benchmark and Accelerator or ClusterMAX dashboards reached working form within two to three hours.

  • The host’s caveat—worth keeping—is that SemiAnalysis links better research relatively directly to subscriptions, whereas someone six layers from a P&G customer might use 50% efficiency to stop 50% earlier. Dylan then said improving model quality made him pessimistic about adoption: “The bottleneck is not actually the quality of the output,” but perhaps apathy or misaligned incentives.

5. Frontier gains depend on persistence, while workflow incumbency becomes the moat

  • Dan found Mythos’s improvement unusually concentrated in coding and cybersecurity; its broader gains were “significantly less,” and the system card reportedly said that outcome was unintended. He questioned whether cybersecurity messaging was partly marketing: “If I wanted people not to use a model” maliciously, he would not run a campaign emphasizing zero-day discovery. A Stanford security researcher warned him about long-term agentic attacks, while that researcher and his friends remained suspicious that Mythos represented a true step change in model quality.

  • Dan offered GPU-kernel authoring and inference-system architecture as parallel frontiers. AI was producing many leading kernels in a way it had not six months earlier, and he described human-to-AI crossover as generally “a one-way street.” Experienced kernel authors using AI still had an edge, however; token budget alone had not yet commoditized the task.

  • Long-range persistence remains weak: models may stop investigating documentation and source code, suggest opening a maintainer PR, and fail to resolve a logged bug. Dan’s synthesis was that both layers matter—model post-training must reward extended coding behavior, while a well-designed harness supplies the tools and structure needed to sustain it.

6. The model race may be converging, but workflows create inertia

  • Meta’s Avocado release looked rushed, including an erroneous chart, yet landed near “Opus, 5.4, and a Gemini,” suggesting the hard catch-up work may be done. Dylan judged it not quite as good as the top two or three models but in that ballpark, while Dan said Meta appeared to have completed the difficult part and had relatively little left to close.

  • Switching a Claude Code configuration could be technically trivial—Dylan described it as one URL change—but accumulated knowledge files, expected behavior, and tolerance for fixable mistakes create inertia. Dan’s question remained open: does differentiation migrate from raw capability to “ecosystem,” “incumbency,” and user experience?

  • The hardware anecdote pointed in both directions. Dan said Mac minis had roughly three-month U.S. lead times and about one month in Asia amid the Claude Code rush. Dylan countered that a Mac mini was unnecessary because most model compute runs in the cloud, and considered using a Chromebook instead.