Pioneers Insight Method Research Author
Video's Victory Over Text, Google's YouTube Opportunity & AI Bubble
Back to Episodes

Video's Victory Over Text, Google's YouTube Opportunity & AI Bubble

Summary

  • Ben Thompson’s core consumer-AI call is that video, not text, will capture the largest audience and market. Text arrives first because it is easiest, but media repeatedly progresses toward images and then video, where attention concentrates. Thompson argues that social networking is a route to scale for pure entertainment, and cautions against assuming people will reject AI video after short user-generated video similarly defied elite expectations.
  • Google may have the strongest consumer-AI position because it combines leading video models, YouTube distribution, an advertising business, and ownership of YouTube’s video data. Thompson says OpenAI’s Sora was “embarrassed” by Veo 2 and that Veo 3 is “amazing” and now productized: “Video’s not competitive. Google’s so far ahead.” He imagines Google’s ownership of YouTube gives it an efficiency and capability advantage in using that data. YouTube is beginning to ship these capabilities, though some announced features are not yet live.
  • Gemini’s Nano Banana image editor offers early evidence that consumer AI can expand beyond chat. Google’s 2.5 Flash Image reportedly brought Gemini 13 million new users in four days, while users uploaded five billion images in under a month; unlike typical brief App Store spikes, Gemini remained number one for weeks. Thompson sees image creation as only the midpoint because “the end state is video.”
  • YouTube’s most consequential AI feature may be automatic product tagging, not generative spectacle. Its AI can recognize when a creator discusses a Samsung phone, match that moment to a supplied link, and insert the link automatically—an initially “super basic and incredibly profound” step toward recognizing products throughout video and making its pixels monetizable.
  • YouTube’s commerce push creates upside and tension across creators and brands, while raising a question for Amazon. Smaller creators could gain access to brand monetization they cannot arrange alone, while premium creators may see YouTube take a cut if brands route existing partnerships through the platform. The episode ends by raising—but not answering—whether a larger YouTube affiliate surface and AI product recommendations threaten Amazon.
  • Thompson explicitly reversed his prior assumption that Meta would ship AI products better than Google. Meta’s AI is “a total mess” and “kind of in disarray,” while Google has begun connecting its models to real consumer surfaces and advertising. His updated shorthand: take the earlier Meta AI-abundance thesis and “search and replace Meta for Google.”
  • Apple’s iPhone 17 execution supports a narrower recovery thesis around its established strengths, not a resolved AI story. Andrew Sharp reports battery life of roughly a day and a half versus 12 hours on his 15 Pro, a noticeably better camera, effective square-sensor selfie technology, and surprisingly elegant Liquid Glass. Thompson’s framing is that Apple “retreated” to operating systems, interfaces, and hardware while its AI questions remain “wide open” and “TBD.”

Deep dive

1. Apple’s hardware reset restores confidence without answering AI

  • Sharp’s early iPhone 17 Pro verdict is unusually enthusiastic: upgrading from a 15 Pro, he is charging roughly once every day and a half instead of every 12 hours. The camera is “a noticeable improvement,” and the new selfie technology works well even while he is holding his seven-month-old daughter. Liquid Glass has been “pretty elegant and pleasant to use” rather than the expected failure.

  • Thompson explains the selfie improvement as a square sensor that decouples capture orientation from how the phone is held. Users can keep the device vertical, switch between portrait and landscape framing, and retain the highest available quality—an ergonomic correction to the front camera’s role in creating today’s vertical-video culture.

  • Thompson’s larger Apple framing: the company chose to return to what it does best amid unresolved AI questions—“We’re going to do what we do best.” Strong phones, operating-system work, and interface design show Apple is “still good at what we do,” but the AI questions remain “all still wide open, all still TBD.”

2. Mass-market media keeps moving from text toward video

  • Thompson starts with the mismatch between industry taste and mass behavior. Tech and media insiders loved Twitter because it was dense, customizable, and intellectually generative; he still sees tweets that change his thinking. Yet “according to the numbers, I’m the outlier,” and most people are not as interested in investing that much effort in information.

  • His best personal evidence is almost comically extreme: as a child he read around 10 books a week and could be punished by losing reading privileges, while watching Succession now “feels like work.” Twitter is entertainment to him, but that preference should not be mistaken for the market.

  • Facebook versus Twitter supplied the earlier warning. Thompson watched fellow English teachers in Taiwan obsess over Facebook while he considered Twitter obviously superior; Facebook then moved from text to photos and video, while TikTok went further with video from anywhere and is not really a social network at all.

  • Thompson grades his 2015 call as only “70% right,” but retains its core: social networking was a route to scale, not the final product. Once scale exists, services can become pure entertainment—a point Sharp summarizes as “entertainment is the end game”—which is a useful correction whenever elite information habits are projected onto consumer demand.

3. Consumer AI is replaying the text-to-image-to-video sequence

  • Thompson narrows the thesis explicitly to consumer AI: what resonates broadly and ultimately makes money. Newspapers, magazines, television, and user-generated media all began with text, added images, and ultimately gave more attention to video; text arrives first because printing and bandwidth make it easiest.

  • ChatGPT may “always be the go-to text source,” but Thompson points to its Studio Ghibli image moment as evidence that consumers want more than conversation. Google’s Nano Banana then delivered another strong demonstration: Gemini reportedly gained 13 million users in four days, and users uploaded five billion images in less than a month.

  • The persistence matters as much as the spike. Apps often top the App Store for a day or two before disappearing, whereas Gemini stayed number one for weeks as successive image memes circulated. Thompson calls Nano Banana “incredible” and sees real sticking power—but insists images mean “we’re only halfway, because the end state is video.”

4. Google’s video lead matters because YouTube is finally shipping it

  • Thompson rejects confident claims that people will dislike AI-generated video: 10 years earlier, few would have forecast users spending more time on 30-second TikTok clips than professionally produced programming. “Be careful in your presumption around human preferences”; history establishes that people like video, even if AI video’s ultimate reception remains uncertain.

  • His competitive call is much less hedged. Image generation remains contested among Google, OpenAI, and Midjourney, but “video’s not competitive”: OpenAI tried to seize attention with Sora, Google “embarrassed them with Veo 2,” and Veo 3 is “amazing” and increasingly productized.

  • Thompson admits he is late to the Google thesis after the stock’s rise over the past few months. Google always possessed the pieces—DeepMind’s capabilities were visible a decade ago—but its chaotic Paris response to OpenAI and early Bard launch made execution the question. The YouTube announcement matters because “they actually shipped.”

  • Some announced tools are not yet live, and Thompson preserves that caveat: check again in a month or two, and hold Google accountable if they never appear. The proposed set includes Veo 3 generation inside YouTube, automatic assembly and narration of uploaded clips, and text-to-song transformations that could turn a political rant into a viral track.

5. Automatic product links point toward monetizing every video pixel

  • The “super basic and incredibly profound” feature is automatic tagging. Rather than forcing creators to scrub through an uploaded video, find the relevant moment, insert a link, and adjust when it ends, YouTube can detect a Samsung discussion around minute 20, match it to the supplied Samsung link, and place it itself.

  • Thompson sees a straight line from that convenience to his earlier Meta thesis that “every pixel on Instagram going forward is going to be monetizable.” AI could recognize products throughout video, attach links, and eventually support more sophisticated attribution or a bidding process than merely checking whether a creator recited a sponsor message.

  • The economics cut both ways. Lower-level creators could gain brand opportunities unavailable to them today, but YouTube also sees premium creators “making more money off of YouTube than on YouTube” and wants a share; brands might compel those creators to route direct partnerships through the platform.

  • Thompson concedes his key Meta assumption was wrong: he gave the founder-led company the benefit of the doubt on shipping, but its AI is now “a total mess” and “kind of in disarray.” Google has the models, YouTube surface, advertising machinery, and—Thompson believes—an advantage in using YouTube’s video data because it owns the platform. In his words, AI will be “one of the biggest gifts to advertising ever.”

  • The episode closes with a listener question about whether YouTube’s expanding affiliate surface and AI product recommendations threaten Amazon, without resolving it.