Pioneers Insight Method Research Author
Age-Gating the Internet + Cloudflare Takes On A.I. Scrapers + HatGPT
Back to Episodes

Age-Gating the Internet + Cloudflare Takes On A.I. Scrapers + HatGPT

Summary

  • Age verification is becoming internet infrastructure, but the UK rollout shows how quickly child-safety policy can become a privacy and speech regime. The Online Safety Act requires services to assess whether minors could encounter pornography, pro-suicide or pro-eating-disorder material and, where the risks warrant it, deploy “high-quality age assurance.” The regime reaches X, Reddit communities and potentially Wikipedia. Adults may need to provide IDs, credit cards or facial scans for access that was previously “basically completely private.”

  • Workarounds can undercut the safety benefit while preserving the surveillance cost. More than 400,000 people signed a repeal petition, VPNs became an obvious workaround, and users fooled facial checks with an expressive video-game character from Death Stranding. Yet the model is spreading: 24 US states have passed age-verification laws, and the Supreme Court upheld Texas’s requirement for sites where more than one-third of content is sexual material.

  • The consequential layer may be privacy-preserving age assurance, not thousands of identity databases scattered across the web. Apple’s proposed API would convert a parent-supplied age into an anonymized token for developers; Meta wants Apple and Google to bear more legal responsibility, while Apple wants to provide infrastructure without accepting that liability. Kevin Roose expects age-gated internet access to become “sort of inevitable,” making implementation architecture the consequential fight.

  • Cloudflare says AI has badly eroded the web’s traffic-for-content bargain. Matthew Prince, whose company sits in front of “north of 20% of the web,” said generating traffic is nearly 10 times harder than a decade ago; the comparable burden is 750 times for OpenAI and 30,000 times for Anthropic. If publishers cannot obtain recognition or money through fame, ads or subscriptions, they are “dead at 750 times” and forgotten at 30,000.

  • Cloudflare’s July 1 default block is an attempt to manufacture the scarcity required for an AI-content market. Prince’s core call is categorical: “AI companies have to pay for content,” because they now copy originals without returning monetizable traffic. He expects pricing to evolve from an iTunes-like “99 cents a song” model toward a Spotify-like subscription or revenue share, with access to distinctive content becoming a key model differentiator.

  • Cloudflare could become both the enforcement layer and a paid market maker, creating opportunity alongside concentration risk. Direct publisher-AI deals would pay Cloudflare nothing; when it negotiates or facilitates transactions, Prince floated a 20%–30% share, while blocking and analytics remain free even for small sites. He acknowledged the discomfort of one company acting across 20% of the web, but argued that “the first step in any market has to be creating scarcity.”

  • Several HatGPT stories showed AI products capturing identities, speech and work faster than norms can adapt. LeBron James’s lawyers challenged unauthorized synthetic videos; Sam Altman warned that ChatGPT therapy conversations lack legal confidentiality; Amazon’s acquisition target Bee records users and everyone around them; and Meta will let some coding candidates use AI. Casey Newton’s desired baseline was simple: users should see, control and delete what systems retain about them.

Deep dive

1. Age assurance fragments a once-uniform internet

  • Casey Newton framed Britain’s Online Safety Act as “one of the most far-reaching attempts” by a Western democracy to regulate online speech. Passed in 2023, it requires services to assess whether minors could encounter pornography, pro-suicide material, pro-eating-disorder content or other listed harms, then deploy “high-quality age assurance” where the assessment identifies the relevant risks rather than relying on an easily checked “I’m 18” box.

  • When the provisions took effect on Friday the 25th, many adults discovered without warning that browsing could require a driver’s license, credit card or camera-based age estimate. Casey’s concern was the abrupt conversion of a previously “basically completely private” visit into an interaction linked to personally identifying information whose storage and downstream use were unclear.

  • The scope extended beyond obvious adult sites. X and Reddit communities about smoking cessation, cider and other seemingly nonsexual subjects faced restrictions; Wikipedia warned that privacy concerns might force it to limit UK access rather than collect information it did not want.

  • Kevin’s larger framing: after roughly 40 years in which a 13-year-old and a 50-year-old could encounter the same informational network, age assurance could split the web into materially different products by age. That fragmentation, not merely blocked pornography, is the structural change.

2. Easy workarounds weaken the safety case as the laws spread

  • The public response was “not a keep calm and carry on situation”: more than 400,000 people signed a petition seeking reversal. VPNs offered the simplest escape, letting users claim they were outside Britain and access services under another jurisdiction’s rules.

  • The sharper example was Death Stranding. Some camera checks instructed users to smile or frown, so people used the game’s photo mode to make its protagonist perform those expressions and passed the resulting image to the verifier. Casey joked that an age-gating backlash could make it Britain’s best-selling game.

  • Ofcom said the law was “not a silver bullet,” while maintaining that checks would stop children from casually stumbling across harmful material. The exchange left the central mismatch intact: determined users could evade the gates, while ordinary adults still absorbed friction and privacy exposure.

  • The policy direction already crosses the Atlantic. Age-verification laws have passed in 24 states, and the Supreme Court upheld a Texas rule covering sites where more than one-third of the material is sexual. Justice Clarence Thomas’s rationale was that, unlike a clerk, a website “cannot look at its visitors and estimate their ages.”

3. Device-level tokens offer privacy, but liability remains contested

  • Kevin steelmanned the case for age gates: bars check identification, and adult magazines were historically kept behind counters. Casey agreed that minors should be kept from adult material; his objection was making people disclose sensitive information repeatedly, potentially “for the rest of their lives.”

  • Apple’s proposed age-assurance API offered the hosts’ preferred architecture. A parent would enter a child’s age or birthday during device setup; Apple would anonymize it and send an app a token indicating, for example, that the user is 13. The feature was not yet available, though Apple told Bloomberg it was coming soon.

  • That design minimizes what each developer learns, but it does not settle who is accountable. Meta has promoted state bills shifting verification obligations toward Apple and Google; Apple wants to provide infrastructure without becoming legally liable whenever a child reaches prohibited material. One metaphor cast Apple as the mall and Facebook as its liquor store.

  • Kevin pushed back that platforms already infer users’ ages from behavioral data. YouTube planned to use browsing and viewing patterns, while prior reporting showed Meta knew it had under-13 Instagram users. Australia’s reversal toward banning YouTube for children under 16 showed how quickly “social media” rules could widen.

4. Te shows why identity databases become breach multipliers

  • Te, an app where women anonymously shared dating experiences involving men, briefly reached No. 1 in Apple’s App Store. Registration reportedly involved a driver’s-license scan and selfie-based gender check—exactly the kind of centralized collection that turns verification into an unusually valuable breach target.

  • Kevin Roose said his understanding was that Te had been hacked after failing to secure some uploaded verification materials, which were leaked. Matthew Prince said 4chan users then obtained the selfies and repurposed them into “Hot or Not-style websites” and other abusive projects. Casey’s broader lesson was that providers vary in security, so multiplying sensitive databases predictably multiplies exposure.

  • Kevin’s values remained in tension. Adults should retain access to legal information, yet parents are “at their wit’s end” trying to introduce children to smartphones and social media; existing controls can require becoming “a part-time IT person.” He nevertheless considered broader age verification inevitable and wanted privacy-preserving implementation before recurring “Te-style leaks.”

  • Casey situated the issue within what he called a six-month “clawback of speech rights,” alongside pressure on journalists, academia and broadcast media. Requiring government ID to view a site might not trouble everyone, but he argued that the restrictions were “all of a piece” and could look far more consequential in retrospect.

5. AI answers erode search’s traffic-for-copying bargain

  • Matthew Prince said Cloudflare sits in front of “north of 20% of the web,” giving it visibility into a 25-year implicit bargain: publishers let Google copy their work, and Google returned visitors who could be monetized through advertising or subscriptions.

  • Google gradually changed that exchange. Answer boxes began retaining users by displaying facts directly, reversing the founders’ old boast that Google got people off its site quickly; AI Overviews now summarize more queries without requiring readers to visit the underlying sources.

  • Cloudflare’s comparison was stark: attracting traffic to a piece of content is nearly 10 times harder than a decade ago. Prince put OpenAI at 750 times harder and Anthropic at 30,000 times, as users increasingly consume derivatives, skip footnotes and trust improving AI answers.

  • Content, in Prince’s taxonomy, is created for ego or fame, advertising, subscriptions, or some combination. Remove both recognition and money, and production stops: publishers struggling at 10 times are “dead at 750 times” and “buried in the ground and forgotten about at 30,000 times.” He called that an existential threat to the web.

6. Cloudflare is manufacturing scarcity so content can command a price

  • Prince did not claim to know the final mechanism, but his starting principle was categorical: “AI companies have to pay for content.” Search compensated copying with traffic; AI systems copy while returning little, so publishers no longer have a rational reason to grant free access.

  • Existing deals—including Amazon’s agreement with The New York Times and OpenAI’s publisher arrangements—show willingness to pay, but bilateral contracts leave a free-rider problem. “Sam at OpenAI can’t be a sucker”: one company cannot buy licensed material while rivals scrape the same work at zero cost.

  • Cloudflare’s July 1 “Content Independence Day” announcement therefore made blocking AI crawlers the default across participating sites unless creators chose access or compensation. Prince’s market logic was elementary: “You can’t have a market unless you have some level of scarcity.”

  • Pricing will likely iterate. He compared an initial fixed rate with Apple’s “99 cents a song” iTunes model, which ultimately gave way to Spotify’s roughly “10 bucks a month” all-you-can-eat structure. The endpoint might be fixed rates, subscriptions or revenue sharing rather than literal pay-per-crawl.

7. Enforced network controls turn robots.txt requests into consequences

  • Prince likened robots.txt to a speed-limit sign: it communicates a request but supplies no physical enforcement. It is also blunt, often applying rules broadly rather than letting a publisher distinguish among specific content, crawlers and commercial arrangements.

  • Cloudflare tracked AI companies seeking the same material through Bing caches, the Internet Archive and even ad networks that could return page descriptions. Some allegedly used residential proxies to disguise themselves; Prince said the worst looked less like ordinary companies than “North Korean hackers,” while repeatedly describing OpenAI as comparatively well behaved.

  • Cloudflare can enforce rules because traffic to its customers passes through a network built for cybersecurity. The same systems used to identify hackers can deny a disguised crawler access: badly behaving bots can be put “in time out,” while an AI company with a deal receives efficient, structured delivery.

  • The intended control plane is granular. A creator could authorize OpenAI, reject another crawler or eventually post a standard “rack rate” even without enough scale to negotiate directly. Cloudflare would combine identity detection, blocking, analytics and permitted delivery rather than relying on crawler etiquette.

8. A content exchange could pay the long tail—and Cloudflare

  • Publishers responded with palpable “glee,” Prince said, repeatedly pressing “disallow.” AI companies surprised him by largely accepting that “content is the fuel that runs our engine” and must be paid for—provided Cloudflare can deliver the “level playing field” that prevents less scrupulous rivals from taking it free.

  • Google remains pivotal. Prince compared its changing bargain to “the frog boiling in water”: a beneficial arrangement slowly became punitive as Google kept more attention. He hoped persuasion would work, while noting that global investigations could eventually compel cooperation.

  • Casey wanted trusted publications available inside deep-research products, perhaps through subscriber credentials, but Prince resisted a market limited to major publishers. Unknown sites can produce valuable originals, and startups need affordable inputs; a healthy exchange therefore requires “lots of sellers and lots of buyers,” not only giant media-AI contracts.

  • Cloudflare would take nothing from direct deals such as a publisher’s own OpenAI contract. If it negotiated or facilitated a transaction, Prince imagined a Spotify-like 20%–30% share, though he had not fixed the number. Blocking and analytics would remain free; Casey said even an $80 payment for Platformer could be worthwhile if someone else found the buyer.

9. Better incentives could revive knowledge—or centralize it inside AI giants

  • Asked why AI firms would pay when President Trump said charging for every article or book was “not doable,” Prince distinguished mandatory government rates from private markets. His reading was that the administration opposed Australia- or Canada-style requirements, while Cloudflare’s leverage was technical: compensate creators or lose access.

  • AI Labyrinth supplies the punitive edge. Instead of merely blocking a badly behaved crawler, Cloudflare can send it through endless AI-generated pages and “pollute their data at scale.” Prince’s message to companies spending $10 billion on GPUs was that at least some capital should fund the content their systems consume.

  • His deliberately overstated diagnosis—“everything that’s wrong with the world today is ultimately Google’s fault”—targeted incentive design. Google taught publishers that traffic was “the deity,” Facebook and TikTok intensified the attention economy, and creators learned to trigger cortisol and rage rather than produce durable knowledge.

  • Prince wants AI demand to pay creators who fill holes in the world’s “giant block of cheese.” His “dark mirror” is journalists and researchers surviving only as employees of five powerful AI companies, with knowledge divided among ideological and national blocs. He acknowledged Cloudflare’s unilateral power by virtue of sitting in front of 20% of the web, citing its earlier decision to withdraw protection from 8chan after mass shootings, but said dying publishers required action and “the first step in any market” was scarcity.

10. HatGPT finds AI colliding with likeness, privacy, hiring and speech

  • LeBron James’s lawyers reportedly sent InterlinkAI a cease-and-desist after its community produced unauthorized videos depicting him pregnant, homeless and in other degrading scenes, some “racist and horrible.” Casey called it, as far as he knew, one of the first known celebrity objections to an AI company’s misuse of likeness—a preview of mounting consent and personality-rights disputes.

  • Sam Altman warned that therapy-like ChatGPT conversations carry no legal confidentiality and might be produced in litigation. Casey welcomed the disclosure but wanted rules governing retention, profiling and advertising: users should know what a chatbot stores, inspect what it “knows” and delete the record.

  • Meta will allow some coding candidates to use AI during interviews, which Kevin called a “Roy Lee victory” over conventional LeetCode testing. Their practical conclusion was that assessments should resemble AI-assisted work. Separately, Casey felt vindicated after Substack’s recommendation system pushed a Nazi newsletter—the amplification risk that led him to leave—and the company took the feature offline.

  • Amazon’s acquisition target Bee makes a bracelet that transcribes daily speech into searchable memories and tasks. Joanna Stern’s verdict was “impressively useful” yet “really fucking creepy”; the hosts questioned sacrificing the privacy of both wearer and bystanders for reminders to buy milk. Lighter counterpoints included thermal robotic deer catching poachers and a starling serving as a lossy image-to-sound-to-image storage device.