|
|
Log in / Subscribe / Register

Volume difference with older crawlers?

Volume difference with older crawlers?

Posted Feb 15, 2026 23:59 UTC (Sun) by Wol (subscriber, #4433)
In reply to: Volume difference with older crawlers? by mb
Parent article: Poisoning scraperbots with iocaine

>If I check my logs I can see that source IP addresses are 99% unique during an AI DDoS.

Sounds like a good anti-bot would simply be to delay answering any request from a new IP. 5 seconds? And then the more requests from that IP, the higher the priority goes.

Cheers,
Wol


to post comments

Volume difference with older crawlers?

Posted Feb 16, 2026 0:55 UTC (Mon) by corbet (editor, #1) [Link] (11 responses)

I don't see how that would help. Humans get grumpy about a five-second delay... The bot just waits, then has the page it was after. Meanwhile your server, which is now trying to keep open five seconds worth of full-on scraper bot traffic, melts down into slag.

Volume difference with older crawlers?

Posted Feb 16, 2026 11:41 UTC (Mon) by paulj (subscriber, #341) [Link] (10 responses)

To be honest, at some point people are going to have to start sueing these highly abusive AI crawlers for the DDoS they are causing. There are laws covering DoS attacks that would apply if these companies are being reckless.

Volume difference with older crawlers?

Posted Feb 16, 2026 12:15 UTC (Mon) by pizza (subscriber, #46) [Link] (9 responses)

> To be honest, at some point people are going to have to start sueing these highly abusive AI crawlers for the DDoS they are causing. There are laws covering DoS attacks that would apply if these companies are being reckless.

So... how do I identify who to sue when they piggyback on residential IPs and pretend to be MacOS 15 (or a Pixel 8 phone, or Chrome on Windows 11 (or, or, or...)

The thing is, each crawler on its own is fine. The problem comes from everyone and their dog having a unique crawler all hitting at the same time.

Volume difference with older crawlers?

Posted Feb 17, 2026 12:05 UTC (Tue) by paulj (subscriber, #341) [Link] (8 responses)

You won't be able to take lines from your logs, and work out exactly which abusive AI-enshitifier is responsible. Or even if you could to some extent, you won't be able to prove an overall pattern of abuse from just that.

What may happen is that, one day, information comes out (by leaks, or just by the hubris of their own self-delusional sense of self-importance - listen to Altman for examples; or perhaps discovery in some other court case not related to web abuse) from one of these abusive AI enshitifiers (abusing the resources of the web, the Internet, the world in terms of energy) that details their shitty and abusive practices and finally provides the rope to hang them with wrt to how abusive they are of the resources of others on the Internet.

We must live in hope.

Volume difference with older crawlers?

Posted Feb 17, 2026 13:17 UTC (Tue) by Wol (subscriber, #4433) [Link]

Don't various ISPs ask for permission to use your router to anonymise web accesses?

Be rather tricky to do, but couldn't a bunch of small website owners sue one of them for damages, on the basis they are actively facilitating unwanted traffic and abuse?

I know I go on about fraud and abuse (small letter initial), but basing it upon the English laws of trespass, using someone else's property when you "knew or should have known" that permission would be refused, is a criminal offense. The mere fact these people are desperate to hide their identity is a blatant admission they know the "knew or should have known" bar is passed.

I don't know how far such a lawsuit would get in the UK (and the fact it would be a criminal suit means the DPP probably wouldn't be interested), but the "anti-social"ness is clear. It's just finding some way of turning the "at someone else's expense" into something you can sue over :-(

Cheers,
Wol

Volume difference with older crawlers?

Posted Feb 17, 2026 13:45 UTC (Tue) by marcH (subscriber, #57642) [Link] (6 responses)

I agree with the immorality but where exactly is the illegality of the abuse? Which "law of the internet" does an abusive, distributed crawler violate? Assuming you find some, what stops crawlers from avoiding those countries and to crawl only from other ones?

If they hijack unwilling client computers then sure, but do we have any indication that it is actually the case? If not then then what else? I mean on what grounds could anyone sue if not?

TCP/IP won against telcos because it was focused on the technical aspects, mostly ignoring the business/economical ones. Afraid that very old naivety is finally hurting a lot; way beyond SMTP.

Volume difference with older crawlers?

Posted Feb 17, 2026 14:37 UTC (Tue) by anselm (subscriber, #2796) [Link]

If you're a user of $AI_COMPANY's free offerings, it probably says somewhere in the 700-page license agreement (that you didn't read when you signed up for the service) that you consent to your computer serving as a proxy for $AI_COMPANY's web crawler .

Volume difference with older crawlers?

Posted Feb 17, 2026 15:48 UTC (Tue) by paulj (subscriber, #341) [Link] (4 responses)

Many jurisdictions have laws that make acts that abuse the resources of another's computer illegal, e.g. to make DDoSes illegal. As one example, in the UK, the "Computer Misuse Act, 1990", in section 3 " Unauthorised acts with intent to impair, or with recklessness as to impairing, operation of computer, etc. " makes it an offence in subsection (2)(a) to impair the operation of any computer; (b) prevent or hinder access to any program or data held in any computer; (c) impair the operation of any such program or the reliability of any such data.

It is an offence if the person making the act intended to cause things, OR they did those acts /recklessly/, i.e. they should have known 2(a) to (c) were likely consequences of their acts.

As Wol says, that these abusive AI-enshitifiers must resort to heavily disguising their DDoS actions just further proves their guilt. They *know* fine well the systems they access do not want this access, they know they are causing problems for those systems, precisely because they obviously have _already been blocked_ from accessing those systems directly; and then they go to the effort of disguising their access and continuing the abuse, recruiting vast armies of other people's computers to continue their abusive behaviour.

One day, these AI-enshitifiers are going to be hit with very large lawsuits. And some of these enshitifiers will turn out to be rather large tech companies, and I hope they end up paying out massive amounts in damages.

Volume difference with older crawlers?

Posted Feb 17, 2026 16:03 UTC (Tue) by paulj (subscriber, #341) [Link]

Oh, I assume the USA has similar laws. The US has easier provision for class-action lawsuits I think (they are - I gather - difficult to take in the UK and Ireland), and many of these AI-enshitifiers are based there.

It just needs a bunch of web content providers to get together, find a lawyer willing to take this on, then advertise the action to recruit even more web content providers who are being abused and then go after some of these awful AI-enshitifier people and give them a good spanking in the courts (and earn the lawyer a nice sum, and maybe a little bit for the abused each).

Volume difference with older crawlers?

Posted Feb 17, 2026 16:55 UTC (Tue) by marcH (subscriber, #57642) [Link] (2 responses)

Thanks! But... if lessons from blatant copyright violations are any indication, I'm afraid they will somehow get away with that too. "Move fast and break things" etc.

Right now they seem rich and powerful enough to dominate that _other_ oligopoly that was rich, powerful and corrupt enough to extend copyright laws to 70 years after the death of the author (lol). That's apparently achieved through a combination of forced "partnerships", buy-outs and other nasty arm-wrestling.

Exactly like with email, who cares about the small players.

Unless... the whole Ponzi scheme falls apart first. Interesting times either way.

Volume difference with older crawlers?

Posted Feb 17, 2026 17:32 UTC (Tue) by rgmoore (✭ supporter ✭, #75) [Link]

Right now they seem rich and powerful enough to dominate that _other_ oligopoly that was rich, powerful and corrupt enough to extend copyright laws to 70 years after the death of the author (lol). That's apparently achieved through a combination of forced "partnerships", buy-outs and other nasty arm-wrestling.

A big part of this is that the AI evangelists have managed to sell the political elite on the idea that AI is the next big thing, so whichever country dominates AI will gain untold economic, political, and national security advantages. That gives them a plausible, actionable threat to relocate to whichever country does the most to make AI development easy: providing them with access to resources needed for massive AI data centers, letting them scrape copyrighted content as much as they want, etc. The copyright cartel is powerful and is already an economic engine, but they can't plausibly promise to let their host countries rule the world, so the AI industry is coming out on top.

Volume difference with older crawlers?

Posted Mar 15, 2026 7:35 UTC (Sun) by sammythesnake (guest, #17693) [Link]

> oligopoly that was rich, powerful and corrupt enough to extend copyright laws to 70 years after the death of the author (lol).

90

:-(

Volume difference with older crawlers?

Posted Feb 16, 2026 7:24 UTC (Mon) by mb (subscriber, #50428) [Link]

That just results in the resource consumption on the server to immediately skyrocket.
The port is used. Lots of memory is used already. Some CPU has been consumed already.
Other addresses won't stop hammering, while I hold the connections alive due to the delay.

And real users will be affected the most.

The only thing that really works for me is the exact opposite: Try to handle the request as quickly as possible.


Copyright © 2026, Eklektix, Inc.
Comments and public postings are copyrighted by their creators.
Linux is a registered trademark of Linus Torvalds