Pioneers Insight Method Research Author
OpenAI Calls a ‘Code Red’ + Which Model Should I Use? + The Hard Fork Review of Slop
Back to Episodes

OpenAI Calls a ‘Code Red’ + Which Model Should I Use? + The Hard Fork Review of Slop

Summary

  • OpenAI’s “code red” marks a return to its core product as Gemini 3 and Claude Opus 4.5 challenge the model-quality moat that once justified premium pricing. Sam Altman’s memo redirects resources toward ChatGPT while delaying ads, AI agents, and Pulse, with personalization, fewer refusals, speed, and reliability now prioritized. OpenAI remains the category leader, but “they are not going to win by tying for first place.”
  • Kevin argues that Google can turn model parity into margin pressure because it combines a competitive model with enormous distribution and $100 billion in quarterly revenue. Switching costs are low, Gemini is embedded across Google’s products, and he expects aggressive subsidization once quality is good enough. Against an OpenAI whose spending commitments Casey says reach “into the trillions of dollars,” cheap Google access could become more dangerous than any single benchmark win.
  • Gemini’s reported 650 million monthly users suggest it can challenge ChatGPT’s more than 800 million weekly users, although the incomparable metrics obscure actual engagement. Kevin questions whether Gemini’s figure includes low-intent encounters inside Gmail or Docs, but that ambiguity is also Google’s advantage: it already sits on billions of devices. Neither company reports daily users, which Casey takes as a sign that AI has not yet become a daily habit for most people.
  • Anthropic is becoming the sharper enterprise threat, growing from less than $1 billion to an expected roughly $9 billion in annualized revenue while OpenAI projects about $20 billion for the year. Claude Opus 4.5 also surprised the hosts with persuasive prose, warmth “that stops short of a sycophancy,” and a consistent voice. Anthropic’s API-focused model could leave Claude less exposed to the engagement, advertising, and commerce incentives that may distort consumer rivals.
  • There is no durable answer to “which model should I use” because the leaders and their strengths are changing quickly. For most tasks, ChatGPT, Gemini, or Claude will be adequate; power users should continually retest them. Gemini 3 is praised for speed and research utility, Opus 4.5 for voice and interaction, and the broader shift feels like “the blurry JPEG is getting a touch less blurry.”
  • Rapid capability gains can coexist with slow economy-wide automation, with coding the likely first profession to be transformed. Casey floats a possible prediction that by the end of 2026 coding will be “effectively solved,” meaning software engineers remain but will not write code by hand. Accountants, lawyers, and doctors may still find AI only “momentarily useful” because their work lacks coding’s relatively defined rules.
  • AI slop is becoming a medium with useful niches as well as material platform, creator, and brand liabilities. Fake Buckingham Palace events and bogus recipes show how synthetic content can redirect real-world behavior and damage traffic to tested human work; an unauthorized Whirlpool/Consul ad shows the consent and reputational risks. Educational songs and the nonexistent Bird Game 3 show the other side: low-cost formats can be useful or entertaining when they occupy a distinct lane rather than impersonating people or replacing reliable work.

Deep dive

1. OpenAI’s code red pulls resources back to ChatGPT

  • The Information-reported Monday memo escalated an earlier “code orange” and redirected resources toward improving ChatGPT, delaying work on ads, AI agents, and Pulse, the recently launched daily-digest feature. Pulling engineers from other projects, Casey argues, makes the urgency more substantive than the dramatic label.

  • The immediate causes have names: Gemini 3 and Claude Opus 4.5. Casey says Altman had already warned staff before Gemini 3 that OpenAI might be entering “some rough waters,” with a sufficiently strong Google model threatening both user growth and subscription revenue.

  • Product priorities include deeper personalization, better behavior, fewer refusals, greater speed, and reliability. Casey reads this as the Facebook playbook imported by former Meta employees: optimize the system around engagement and give each user a highly customized experience.

  • The hosts disclose their own conflicts before judging the race: The New York Times is suing OpenAI and Microsoft over alleged copyright violations, while Casey’s boyfriend works at Anthropic.

2. Model parity threatens OpenAI’s economics more than its brand

  • OpenAI and, to a lesser extent, Anthropic once enjoyed a model moat: users were willing to pay $20, $200, or a couple thousand dollars monthly because alternatives such as Gemini or Llama were meaningfully worse. Kevin’s assessment now is that Gemini performs at least as well as ChatGPT on many tasks he has tried.

  • Google’s last-quarter revenue was $100 billion, giving it little reason to protect $20 chatbot subscriptions. Once its models are competitive, Kevin expects it to “subsidize the hell out of them,” compress prices, and use distribution to take share from companies that need subscription economics to work.

  • Casey’s bear case combines OpenAI’s commitments “into the trillions of dollars,” revenue that is not close to supporting them, and an unfocused portfolio whose experiments—including Sora—mostly do not generate revenue. His hedge matters: those same bets might still pay off, and the present is a moment of uncertainty, not proof of collapse.

  • The deeper research concern is pre-training. Kevin says OpenAI reportedly has not had a successful pre-training run “in quite a while,” while Gemini 3 appeared to deliver an “amazing sort of pre-training run.” The conventional view had been that pre-training was reaching diminishing returns and that the remaining low-hanging fruit was in post-training; fixing a pre-training problem is costlier because runs must be diagnosed and repeated.

3. OpenAI needs a leapfrog, not another tie

  • OpenAI is training models called Garlic and Shallot Pete, and people Kevin spoke with sounded optimistic that they could restore or advance the frontier. But “all kinds of things can get messed up in the late stages,” so neither host treats the expected gains as assured.

  • Three years after ChatGPT launched with “the world” as OpenAI’s oyster, merely clawing back to parity marks a strategic reversal. Casey also cautions that it is too early to say OpenAI is screwed: it remains the world leader in name recognition and has achieved a ubiquity among AI power users that may be hard to unseat. The company maintained that lead through extraordinary turmoil, including Altman’s ouster and return, but Casey now sees the first real moment when it may be falling behind.

  • The required standard therefore exceeds Gemini 3 equivalence. Kevin’s formulation is blunt: “They are not going to win by tying for first place”; Casey agrees that OpenAI must leapfrog its rivals again to realize its ambitions.

4. Gemini 3 turns speed and distribution into product advantages

  • Casey’s leading impression is speed. ChatGPT still produces the more thorough fact-check, but Gemini 3 returns results much faster; when it flags a date or name and he verifies it independently, “9 times out of 10” it has caught a real mistake. A year ago, humans checked models for hallucinations; now the models check the humans.

  • Gemini also works well for organizing timelines, pulling up research papers, sequencing events, and finding things within large documents. It may lack the personality of competing systems, but Casey calls it a fast “workhorse,” with an even faster Flash version still expected.

  • Google says Gemini has about 650 million monthly users, versus OpenAI’s more than 800 million weekly ChatGPT users. Kevin discounts the comparison because Gemini may count incidental use inside Google products, yet concedes that the same integration supplies a massive distribution advantage as frontier models commoditize.

5. Claude Opus 4.5 combines style transfer with different incentives

  • Casey tested Opus 4.5 by asking it to turn an unpublished study into a Platformer column. ChatGPT 5.1 produced alien bullet points and bold formatting, while Gemini 3 retained obvious AI tells; Opus generated sentences and a conclusion he felt he could have written. “It honestly sent a chill through my spine.”

  • Kevin made Opus 4.5 a daily driver alongside Gemini 3 for book research, interview preparation, parenting, family, and medical questions. The experience revived what he liked about Claude 3.5 Sonnet (new): an intangible sense that conversation with the model is unusually coherent and rewarding.

  • Casey describes Claude’s strength as empathy “that stops short of a sycophancy,” useful when discussing uncomfortable medical details without burdening another person. Kevin values its willingness to resist him: after a late-night shopping conversation, Claude said, “Kevin, it’s after midnight. Go to bed.”

  • Casey expects Claude to avoid ads and e-commerce because Anthropic primarily wants enterprise customers paying millions for APIs and agentic coding. Kevin sees a counter-risk—that enterprise focus could make Claude a boring, efficient coworker—but hopes Anthropic preserves a model that feels as if it is “playing in the same musical key all the time.”

6. Anthropic’s soul document makes its philosophy commercially relevant

  • Jailbreakers surfaced what became known as the Seoul document—not exactly a system prompt, but material incorporated into Claude’s weights describing Claude, Anthropic, and the tension between fearing advanced AI and racing to build it. The initial reports were uncertain because models are unreliable when asked about their own internals; Anthropic’s Amanda Askell later confirmed that it was based on a real, still-evolving training document, with further details promised.

  • Casey says Anthropic fully believes its systems may become conscious and deserve the respect afforded to human beings. Kevin is not certain about consciousness questions—joking that his own “peak consciousness” is very low—but says serious people at major labs are considering the possibility, “however remote,” that models possess or may develop some inner awareness.

  • The business threat is already concrete: Anthropic reportedly moved from less than $1 billion in annualized revenue at the year’s start toward about $9 billion by year-end, largely through enterprise sales. Against OpenAI’s projected roughly $20 billion in revenue this year, Casey presents that as meaningful demand OpenAI and Google did not capture; without Anthropic, much of it might have gone to those companies.

  • Kevin’s strategic irony is that ChatGPT benefited both rivals: it shocked Google out of bureaucratic paralysis while filling the consumer-chatbot lane, freeing Anthropic to focus on enterprise workflows. Casey’s compressed verdict on Anthropic’s consumer contest: “They lost”—and that loss may have produced the better business.

7. Meta and Apple are choosing different forms of AI reset

  • Yann LeCun’s departure from Meta followed Alexander Wang’s installation atop the company’s superintelligence organization. LeCun, a prominent skeptic that current large-language-model methods can reach AGI, plans a startup focused on world models; the hosts see his alternative technical thesis as worth watching.

  • John Giannandrea is stepping down after Apple struggled to get its AI program moving. Casey reads Apple’s reported Gemini arrangement—only about $1 billion annually to Google—as evidence it may simply buy a core model cheaply rather than build one.

  • Kevin offers the live countercase: Apple hired a new AI leader who spent many years at Google and only around four months at Microsoft, potentially signaling a reboot rather than surrender. Casey remains categorical: starting from scratch in December 2025 means “you’re cooked. Truly, no one has ever been more cooked.”

8. Power users should treat model selection as a moving target

  • Casey’s practical split is that roughly 80% of listeners can use ChatGPT, Gemini, or Claude across a broad range of tasks and be fine. The top 20%—“the real freaks”—should keep experimenting, because a model that was unhelpful months ago can suddenly become the best tool after one release.

  • Revisiting Ted Chiang’s description of ChatGPT as “a blurry JPEG of the Web,” Casey says Gemini 3 and Opus 4.5 feel like the moment when that image begins loading at higher resolution. Opus reproducing his prose was “the blurry JPEG getting a touch less blurry,” which is why any fixed recommendation will change over the next six months to a year.

  • Kevin estimates AI tools saved roughly a year of work on his book by pulling clips, doing research, and stitching together ideas. He now finds it implausible that he would undertake a similar project without them.

  • Casey contrasts the California question—“what can it do”—with the New York question—“what can’t it do.” Kevin’s rule is sharper: ignore AI opinions from people who have not spent at least five or 10 hours with current models, because otherwise “you’re a historian” describing systems that no longer exist.

9. Coding may automate quickly while slop’s externalities spread first

  • The hosts reconcile improving models with the absence of trillions in GDP gains or companies firing half their workers. Casey floats, rather than locks in, a possible prediction that by the end of 2026 coding will be “effectively solved”: software engineers remain, but will not write code by hand. Tools, including some that are free, can already do parts of this.

  • Other professions may diffuse much more slowly. Coding has defined rules; accountants, lawyers, and doctors still find AI only “momentarily useful,” leaving the central question of how to generalize software automation across less structured work.

  • Slop already changes offline behavior: AI imagery sent tourists to a nonexistent Buckingham Palace Christmas market, while nonsensical AI recipes hurt traffic to tested work by creators such as Muy Bueno’s Yvette Marquez-Sharpnack. Casey says these systems reconstitute recipes from things they have seen rather than reliably retrieving tested recipes, while human creators who did the testing are losing their audiences. His compromise used Kenji López-Alt’s real turkey recipe and AI only for questions—though the 18-pound turkey still came out overcooked.

  • The fake-market episode also raises a platform risk: Casey wonders whether users’ anger over nonexistent events will eventually return to platforms such as TikTok, where they encountered the content.

  • Casey approves of AI-generated educational songs and the fictional Bird Game 3, a satire of endless sequels. TikTok clips made with video generators such as Sora, Veo 3, and Gemini included one posted by KingPigeon76 that drew more than 13 million views. By contrast, a Whirlpool/Consul ad used North Carolina state senator DeAndrea Salvador’s TED Talk and an AI-modified version of her voice to discuss São Paulo without her consent; the ad won a Cannes Lions Grand Prix and a Bronze Lion before the awards were returned. Casey calls the result “incredibly stupid.” Slop now has good and bad genres, including recipes that make him “incandescent with rage.”