40% of Major Websites Block AI—The Problem Is No One Will Pay Machine Fees
Deep thought on AI and aspirations —— ByteDance Deep Thinking Circle
In July 2024, only 9% of the world’s top 10,000 websites blocked AI crawlers. A year later, that number is 37%, and it’s still rising.
Website owners aren’t blocking a specific bot—they’re blocking a new reality: nearly half of internet traffic no longer comes from humans. These machine visitors extract content to feed models or perform tasks on behalf of people, paying nothing while leaving bandwidth and server costs to the websites.
Most discussions frame this as an arms race: crawlers get better at hiding, blocking gets more aggressive. But the arms race is only the surface. The real problem is that the internet was designed end-to-end for human users. Pricing, payment, identity—all assume there’s a person on the other end. When that assumption breaks, the entire commercial logic fails, and it fails more completely than expected.
Three Patches the Internet Built for Humans That Machines Can’t Use
Start with advertising. The internet’s dominant business model is “you consume content, I sell your attention.” This model is completely ineffective for machines: crawlers don’t view ads, don’t generate attention, and advertisers won’t pay for machine visits. Websites whose content gets scraped are essentially doing free labor for someone else.
Next, payment. Human users have bank cards and subscription accounts; the internet’s commercial loop is built on that foundation. AI agents don’t have bank accounts, can’t open them, and can’t pass KYC. Making an agent pay per page scraped has no corresponding interface in today’s financial system: transactions are too small, frequency too high, and the entity isn’t human.
Finally, identity. Bots can replicate infinitely—one server can “become” ten thousand users overnight. Distinguishing human from machine currently relies on old methods like CAPTCHAs, but new-generation models can already solve them. Deepfakes and social media manipulation make this worse: it’s not just about telling who’s on the other end, but whether they’re human at all.
When these three systems fail simultaneously, you have the root of the past two years of conflict. The law moved first: multiple US newspapers sued model companies for copyright infringement, education sites saw traffic plummet as students switched to AI tools, and responded by building walls. The rise from 9% to 37% blocking is the direct result of this litigation and blockade war.
But litigation can’t settle the accounts. A lawsuit can determine a total damages amount, but can’t calculate what a specific scraping instance is worth or which layer of the chain should receive payment. If a dataset has been referenced by thousands of crawlers, how do you attribute contribution and settle splits? Courts can’t provide that granularity. Pricing ultimately needs a technical solution to close the loop.
Machine Users Need a Machine-Native Infrastructure
This is where blockchain gets its moment. Its role in all this isn’t about coin prices—it’s about rebuilding three things for the machine world: machine payment, machine identity, and machine-readable property records.
On payment, digital cash systems happen to fill the financial system’s gap: near-zero-cost micropayments, programmable revenue splits, and the ability to hold value without a bank account. A crawler agent could carry on-chain balance, negotiate in real-time with each site’s access protocol, and pay per page; smart contracts could automatically split each payment to multiple upstream content sources. Protocols like x402 are already moving in this direction. Human users could take a different path: use decentralized identity tools with “proof of personhood” to prove they’re human and continue accessing for free. This gives websites their first ability to price machine traffic and human traffic separately.
Identity works the same way. AI agents need a cross-platform universal “digital passport”: who issued it, what version, what capabilities, what reputation history—any interface can parse it uniformly. Without this public identity layer, agents have to negotiate from scratch with every system, and reputation resets with each channel change. Discovery mechanisms remain perpetually improvised. Property records too: ownership and licensing of training data, model weights, and generated content—if machines are to execute automatically, it needs to be written in a machine-readable, tamper-proof registration format. Early attempts like Story Protocol and on-chain IP registration are betting on exactly this layer.
But a reality check is needed. The list of Crypto-AI intersections can run long, from decentralized computing to on-chain AI companions, each sounding like the future. Most entries can’t survive one follow-up question: is this a new problem that machines brought?
That’s my only filter. Machine payment, machine identity, machine-readable property rights—these three contradictions weren’t acute before machines became major internet users, traditional solutions genuinely can’t handle them, and blockchain has an opportunity. Decentralized computing, privacy advertising, on-chain AI companions—these problems either existed already or blockchain is optional in the solution, added more for narrative completeness. To judge whether a crossover scenario is real, put it in a pure Web2 environment and see if it breaks: only what breaks is a real need; what you can work around is just decoration.
There’s another easily overlooked layer: blockchain’s real selling point in these scenarios isn’t performance. On user experience, vertically integrated big platforms always do better—Amazon’s and Facebook’s closed identity systems remain the optimal experience today. Blockchain sells “no one’s in charge.” When the internet’s negotiating table has tens of thousands of small and medium sites on one side and a few model giants on the other, a neutral middle layer has value. This is essentially an anti-monopoly political problem that happens to use technical means to solve.
Three Real Problems and Their Respective Tipping Point Conditions
| New Problem Machines Bring | Current Solution | Why It Fails | Possible New Infrastructure |
|---|---|---|---|
| Agents paying per-use for content | Bank cards, subscriptions | Machines can’t open accounts, micropayment high-frequency settlement costs too high | On-chain micropayments and paid crawler protocols |
| Agent cross-platform identity and reputation | Platform-internal account systems | Identity locked in walled gardens, reputation non-transferable | Portable agent digital passports |
| Training data attribution and automatic splits | Copyright litigation, manual licensing | Can’t calculate per-use accounts, can’t auto-execute | Machine-readable property rights and split ledgers |
When will the tipping point come? Watch three signals. First, whether mainstream hosting services (CDN-layer services like Cloudflare) start embedding paid crawler protocols—robots.Txt is a gentleman’s agreement from the nineties and won’t evolve into a payment system on its own. Second, whether proof-of-personhood identity user volume crosses critical mass—this type of infrastructure has network effects, slow early on, then suddenly accelerating past a threshold. Third, whether the first killer scenario appears where “high-quality content can’t be scraped without payment.”
Until then, most of these attempts remain experimental, on a timeline measured in years, not quarters.
For people building Agents and people investing in them, this judgment has practical implications. Machine-native payment and identity will likely be a late-arriving but unavoidable piece of Agent infrastructure: teams treating it as a supporting character now will find the price of catching up later higher than expected. Conversely, if you’re just layering blockchain narrative onto AI scenarios in product whitepapers, save yourself the effort. Machines don’t need narratives. Machines only recognize interfaces.