|
|
Log in / Subscribe / Register

Why not clone?

Why not clone?

Posted Aug 29, 2026 12:15 UTC (Sat) by MarcB (subscriber, #101804)
In reply to: Why not clone? by magfr
Parent article: Ryabitsev: Creepy crawlies

That really is the main question that the original article is not answering. It simply makes no sense to do this for training. The Linux kernel is not *that* valuable (only a small fraction of software development is at the system level - don't get fooled by the inherent bias of this site). Even if it were, cloning is obviously so much more efficient; even more so for the data consumers themselves.

I wonder is this already is "agentic" AI, potentially controlling a real browser. This would equally fit the pattern of just 4-5 requests and explain how they defeat countermeasures. Those things pick up GIT commits referenced anywhere and can decide too pick them up to have a closer look. This can easily find some "random old fork". It might be worth checking what the possible trail might have been for a number of those requests (but it could be hard, because of all the anti-bot/scraper measures everywhere :-).


to post comments

Why not clone?

Posted Aug 29, 2026 22:40 UTC (Sat) by JanC_ (subscriber, #34940) [Link] (5 responses)

My guess is that these are just dumb people who are randomly scraping the internet with dumb scrapers to feed their attempt at training AI models. They obviously have no clue what git or a commit is.

Why not clone?

Posted Aug 31, 2026 8:16 UTC (Mon) by taladar (subscriber, #68407) [Link] (1 responses)

It is much more likely that these are just people using (inference) some coding agent to solve some problem that requires information from those particular commits.

Why not clone?

Posted Sep 1, 2026 12:55 UTC (Tue) by smurf (subscriber, #17840) [Link]

Dream on. The volume is much too high for that idea to be remotely valid. Also, the traffic is too spiky.

Why not clone?

Posted Sep 3, 2026 5:15 UTC (Thu) by NYKevin (subscriber, #129325) [Link] (2 responses)

How can they be dumb enough to do that, but also smart enough to use headless Chrome (or whatever) to bypass Anubis? It doesn't make sense to me.

Why not clone?

Posted Sep 3, 2026 13:41 UTC (Thu) by mathstuf (subscriber, #69389) [Link]

I wonder if this recent DailyWTF isn't applicable. They're measuring 4xx results without looking at why they're getting them and the metrics numbers "just need to go up", and running a full browser "makes number go up" and is seen as a suitable solution.

Why not clone?

Posted Sep 3, 2026 15:33 UTC (Thu) by JanC_ (subscriber, #34940) [Link]

They already have to use browsers because so many sites show no content at all without JavaScript & various other browser technologies.


Copyright © 2026, Eklektix, Inc.
Comments and public postings are copyrighted by their creators.
Linux is a registered trademark of Linus Torvalds