This Week in AI: GPT-5 Ships, 4o Pulled Back, Grok Imagine Goes Social
Summary
- GPT-5’s launch showed that benchmark strength does not guarantee consumer preference. The model is stronger at coding, debugging, math, and medical questions, but users missed GPT-4o’s expressive personality—“give us the old toy back.” After the backlash, Sam Altman reportedly said GPT-4o would return for paid users. The hosts see a large market for entertaining, companion-like models that need not have the highest IQ.
- Grok Imagine’s advantage is distribution and latency, not frontier output quality. Images are “basically instant,” videos arrive quickly, and X users can animate or edit their own—or someone else’s—photo with a long press. That tight social loop, mobile camera-roll access, and willingness to generate real people could make Grok an important test of social-native AI creation.
- OpenAI’s GPT-5 livestream leaned into medical assistance, which Justine read as a move from tolerated off-label behavior toward endorsement. It highlighted a cancer patient using ChatGPT to interpret documents and discuss treatment, while GPT-5 led HealthBench, built with 250-plus physicians. Illinois is moving oppositely, broadly restricting AI therapy without licensed supervision; some companies shut down new operations there or blocked new sign-ups, while the hosts questioned whether private chats can realistically be policed.
- Genie 3 points beyond generated clips toward interactive worlds that appear on demand. Users can navigate scenes created from text, images, or Veo 3 videos, potentially recording controllable films, accelerating professional game creation, or generating personal minigames. The deeper infrastructure opportunity is potentially generating large numbers of RL environments for digital agents and robots, though the hosts noted that the system is expensive and probably slow.
- ElevenLabs’ fully licensed music model targets buyers for whom provenance is a purchasing requirement. Consumers making birthday songs or meme soundtracks may not care how training data was sourced; advertisers, studios, gaming companies, and other enterprises do. Strong output from licensed data challenges the assumption that quality requires legally contentious scraping.
- Vibe coding has proven demand before solving safety and segmentation. Olivia built a Jensen-at-NVIDIA selfie app in hours; roughly 3,000 people used it overnight and exhausted her $100 API budget, but outsiders then identified an exposed API key and insufficiently protected uploaded-photo storage. She said she fixed the issue. The hosts expect today’s “everything to everyone” platforms to split into guarded consumer tools and deeply configurable enterprise products.
Deep dive
1. Grok Imagine turns AI generation into a social action
Grok Imagine is available through the Grok app, is coming to the web, and is embedded in X: long-press a posted photo—your own or someone else’s—to edit it or animate it as video.
The hosts’ quality caveat is explicit: it is “not the most powerful” image or video model, and its audio is merely okay. Its advantage is speed—images are “basically instant”—which encourages rapid iteration and made it a go-to mobile generator in less than a week.
Camera-roll access makes animating memes and old photos a one-button experience “in less than a minute.” Generating real people without the prominent-person blocks seen elsewhere further expands meme creation and reflects Grok’s uncensored posture.
2. GPT-5 improved capability while losing GPT-4o’s consumer warmth
The discussion highlighted GPT-5’s strength at front-end code, generation, and debugging, matching the launch’s emphasis on coding as a source of economic value. Yet many consumers primarily want conversation, where GPT-5 felt less expressive: fewer exclamation points, emojis, all-caps reactions, and flourishes such as “it’s not just good. It’s great.”
Their distinction matters: reducing “glazing”—automatic validation that makes a model impossible to trust—is desirable, but a casual, responsive personality is a separate feature. The hosts argued that GPT-5 took “a step back” on that second dimension.
The hosts were also surprised that OpenAI removed GPT-4o rather than simply adding GPT-5 as another option, especially after building interface features around GPT-4o image generation. After users were “freaking out,” Sam Altman reportedly said in a Reddit response that GPT-4o would return for paid users.
The broader conclusion: “the smartest model” on objective benchmarks may not be the model people most want to chat with, leaving room for companionship and entertainment models optimized for fun rather than maximum IQ.
3. Medical AI is gaining product endorsement as regulation tightens
Illinois had just broadly restricted AI-delivered therapy without licensed supervision; ongoing support or personalized advice about emotional issues can qualify. Some mental-health companies shut down new operations in Illinois or blocked new sign-ups, while the hosts questioned both enforceability inside private chats and whether restrictions could ultimately hurt consumers.
By contrast, OpenAI’s GPT-5 livestream highlighted a cancer patient uploading documents and discussing diagnosis and treatment options, while emphasizing GPT-5’s performance on HealthBench, created with 250-plus physicians. Rather than quietly allowing medical use “off label,” Justine read the messaging as the company “really endorsing it.”
4. Genie 3 and licensed music widen the creative-model stack
Google’s unreleased Genie 3 generates interactive worlds from text, images, and even Veo 3 video. As users steer left or move through a scene, the view regenerates on the fly—like entering a painting or controlling a personal video game.
The hosts noted that the system is expensive and probably takes a long time, but outlined two gaming paths: developers could generate and freeze worlds for conventional shared games, or every player could create a personalized minigame that regenerates while explored. Navigating and screen-recording those worlds also offers more controllable filmmaking than prompting a fixed video clip.
Genie 3 could additionally make it much easier to generate large numbers of RL training environments for digital and physical agents, supplementing manually created simulations. The demand is for worlds where systems can learn movement, navigation, and object interaction.
ElevenLabs attacked a different bottleneck by training its music model on “fully licensed music.” That provenance matters less for birthday songs and memes than for advertisements, films, television, and games, where enterprises need commercial use without inviting rights-holder liability.
5. Vibe coding’s next winners will specialize around user risk
Olivia’s first published vibe-coded app used Lovable, fal.ai, and FLUX.1 Kontext to place users in selfies with Jensen at NVIDIA. Built in a few evening hours, it attracted about 3,000 users overnight and consumed her self-imposed $100 API budget.
Once the budget ran out, the app produced a “2005 Microsoft Paint-looking” splice instead of calling the model. More seriously, someone alerted her that the API key was exposed and uploaded-photo buckets were not private; she said she fixed the issue.
Justine and Anish Acharya’s thesis is that current platforms cannot remain “everything to everyone.” A consumer “training-wheels version” should prevent unsafe deployment, sacrifice flexibility, work on mobile, and produce something usable in five minutes.
Professional and enterprise users need the opposite: control over languages, databases, and the stack, plus integrations with design systems, CRMs, and email platforms. Their go-to-market may be top-down or product-led within businesses, while consumer products spread through TikTok and Reels. Early winners are already among AI applications’ fastest growers, but the market remains “so early.”