|
|
Log in / Subscribe / Register

Collateral damage

Collateral damage

Posted Aug 29, 2026 22:28 UTC (Sat) by corbet (editor, #1)
In reply to: Collateral damage by bluca
Parent article: Ryabitsev: Creepy crawlies

It's anybody using the services of companies like Bright Data. That could be anybody, and nothing in the behavior of the top-line LLM companies makes me believe that they would not stoop to making use of such services.


to post comments

Collateral damage

Posted Aug 29, 2026 22:56 UTC (Sat) by JanC_ (subscriber, #34940) [Link] (8 responses)

Some of the “top line LLM companies” used completely illegal torrent & other similar downloads, so I’m sure legal-but-unethical residential proxy services won’t be a problem for them…

Especially as almost all of them have been fined for illegal & unethical behaviour in the past.

Collateral damage

Posted Aug 31, 2026 2:27 UTC (Mon) by smurf (subscriber, #17840) [Link] (7 responses)

On the other hand, the “top line LLM companies” actually care about data quality (at least somewhat), and a petabyte of web crawls with random kernel diffs and whatnot is unlikely to be very interesting to them.

Collateral damage

Posted Aug 31, 2026 7:21 UTC (Mon) by anselm (subscriber, #2796) [Link] (5 responses)

Possibly, but FWIW the general approach seems to be “scan the whole Internet first and sort out the undesirable stuff later”.

I see web crawlers snarfing loads of content off my pages that is of no conceivable use to an LLM whatsoever, over and over again, and that doesn't suggest to me that whoever is controlling the crawlers cares one iota about “data quality”.

Collateral damage

Posted Aug 31, 2026 10:56 UTC (Mon) by pizza (subscriber, #46) [Link] (4 responses)

> I see web crawlers snarfing loads of content off my pages that is of no conceivable use to an LLM whatsoever, over and over again, and that doesn't suggest to me that whoever is controlling the crawlers cares one iota about “data quality”.

Over the last month, on one of my sites, half of the non-blocked-outright traffic went to a dokiwuki instance containing under 100 pages in total. Of that, about half of those were from Google-controlled IPs.

When the most pre-eminent/technically-capable org out there no longer cares enough to do things right, what hope is there for anyone else to do better?

Collateral damage

Posted Aug 31, 2026 22:40 UTC (Mon) by MarcB (subscriber, #101804) [Link] (3 responses)

> When the most pre-eminent/technically-capable org out there no longer cares enough to do things right, what hope is there for anyone else to do better?

Since I work for a company that offers cloud, VPS as well as shared hosting, I know both sides.

On the scraper side:
While the outgoing traffic of those systems is higher then usual, it is still far below anything that would allow you to suspend them. Even if you do get a complaint: Any action performed by any individual VPS against any individual target is legally speaking fine. Even the (stricter) ToS rarely apply.
To add insult to injury: They always use the lowest-margin servers and turn the margin negative because of their increased bandwidth (keep in mind: the hoster gets this double, because the scraped data has to go somewhere).
We eventually resolved this by dropping those offers. Now the smallest VPS costs three times as much and is significantly stronger in CPU and memory, which scrapers do not need.

On the scrapee side:
The initial waves of crawlers were worse than any DDoS attack we faced before (and we had some big ones...), mostly because they hit the application layer instead of just the network, but also because of the highly distributed targets and sources. This invalidated decades of experience in running a large, shared hosting platform. The load patterns and resource usage was unlike anything before. We had whole server rooms exceeding their power quotas and had to do frantic redistributions between data centres - even temporary shutdowns of non-essential systems - until we had things back under control.

Collateral damage

Posted Sep 1, 2026 0:11 UTC (Tue) by pizza (subscriber, #46) [Link]

> Since I work for a company that offers cloud, VPS as well as shared hosting, I know both sides.

Just to be clear, the IPs I was referring to were Google's own crawlers, not 3rd parties using Google resources.

Google _used_ to have sane, well-behaved crawlers, Not any more.

Collateral damage

Posted Sep 1, 2026 14:12 UTC (Tue) by mbunkus (subscriber, #87248) [Link]

Fascinating to read about the problems the bigger hosters are facing; thank you for that! Is there a public, more in-depth write-up of something similar somewhere that you can recommend?

Collateral damage

Posted Sep 3, 2026 16:00 UTC (Thu) by JanC_ (subscriber, #34940) [Link]

Wouldn’t it be more customer-friendly to allow a small number of low-cost VPS per existing customer?

Maybe start with 1 or 2, and an option to ask Support to increase the allowance if they want more, where Support can look at history (new unknown customer vs. 20-year customer should count), and rate-limit how many extra ones they allow over time? (You can also try to automate it, but that might be easier to game…)

There are lots of people (especially also in the FLOSS community) who have “almost idle” VPS used for personal use or experimenting, where the larger VPS would be (too) expensive.

Collateral damage

Posted Aug 31, 2026 7:58 UTC (Mon) by Cyberax (✭ supporter ✭, #52523) [Link]

I work in an AI company (we are NOT doing LLMs, but signal/image processing), and I got approached by a company claiming to sell 20T-token dataset for AI training. With a subscription for regular updates, all payable in Bitcoin. I wonder what percentage of this data are the diffs between pairs of random kernel revisions.

Collateral damage

Posted Aug 30, 2026 17:03 UTC (Sun) by rgmoore (✭ supporter ✭, #75) [Link]

I don't think it matters. Using a residential proxy to hide your identity and disguise your usage patterns shows you know you're doing things illegitimately. My only concern is if this usage pattern is readily distinguished from someone using TOR, who is also deliberately hiding their identity and usage patters, but for a probably legitimate reason. If it is, well and good; you just need to make sure you aren't blocking TOR users in an attempt to stop AI scraping. If it isn't, you need to think long and hard if blocking that legitimate use is a worthwhile collateral damage from your attempt to block illegitimate uses.


Copyright © 2026, Eklektix, Inc.
Comments and public postings are copyrighted by their creators.
Linux is a registered trademark of Linus Torvalds