Volume difference with older crawlers?
Volume difference with older crawlers?
Posted Feb 15, 2026 12:47 UTC (Sun) by excors (subscriber, #95769)In reply to: Volume difference with older crawlers? by marcH
Parent article: Poisoning scraperbots with iocaine
My unfounded speculation is that crawling the whole web is expensive and it's hard for a young search engine to make money, so they had to figure out how to crawl efficiently and cheaply. Eventually they grew into enormously rich companies, but they still had the technology and the culture of crawling efficiently and cooperating with server owners (via robots.txt, honest UA strings, etc). But now there's an insane amount of money in the AI industry, so everybody and their dog can get billions of dollars for their AI startup and vibe-code their own crawler, and they don't know how to do it efficiently and they don't care - they just care about scraping as many terabytes of text as possible, as quickly as possible, before the whole thing collapses.
More cynically they might recognise it's actually a competitive advantage for them to DDoS sites because it blocks other AI crawlers from getting the same content, and also means end users can't go to the original source and will have to make do with half-fabricated AI summaries of the content, ensuring all the ad revenue goes to the AI companies and not the content producers.
