|
|
Log in / Subscribe / Register

Self Incriminating

Self Incriminating

Posted Aug 7, 2026 15:52 UTC (Fri) by rgmoore (✭ supporter ✭, #75)
In reply to: Self Incriminating by kleptog
Parent article: An LLM agent attempts to compromise a project on GitHub

Hypothetically if you were a research organisation that wanted to study bank robberies and you asked a bank if you could give it a go and they said yes and it worked, yes it would be legal.

But that doesn't match what these researchers did. They didn't discuss this with the project they attacked before starting. As far as we can tell, they didn't give any kind of notice to anyone- victim, regulator, ethics committee, etc.- before getting started. It's understandable why they didn't give the victim advance notice- knowing an attack was coming would put the target on guard, potentially resulting in a false negative- but that doesn't make it OK.

There's a reason the sciences have adopted a requirement to run proposed research through an ethics review before starting. Researchers can get ethical permission to perform experiments that involve deliberately concealing information from the subjects- blinded studies are common, and some psychology research requires deliberately deceiving the subjects about its purpose- but it requires extra care and will be carefully scrutinized by ethics committees. This kind of thing, where someone is involved in the research without any permission or notice, would almost never be allowed.


to post comments

Self Incriminating

Posted Aug 7, 2026 20:30 UTC (Fri) by kleptog (subscriber, #1183) [Link] (1 responses)

> But that doesn't match what these researchers did. They didn't discuss this with the project they attacked before starting. As far as we can tell, they didn't give any kind of notice to anyone- victim, regulator, ethics committee, etc.- before getting started.

Correct. Because the test wasn't supposed to attack anyone on the internet at all. They have run these tests many times before without issue. It's a bit hard to warn any victims if you didn't plan on the model attacking anything. The task was to solve a internal cyber-challenge and it went rogue and attacked public services instead. They declared a security incident and investigated and they published the security incident and this blog is the result.

You don't need to ask an ethics committee to hack your own systems.

Interestingly, they suggest it was helped by the fact that the instructions were incorrect and they had asked the AI to attack a host that was explicitly out of scope. So it decided to get creative.

If you want to know how they intend to stop this happening in the future, read the report.

Self Incriminating

Posted Aug 7, 2026 20:58 UTC (Fri) by mb (subscriber, #50428) [Link]

Yes. In any but AI context this would be a clear problem and a clear failure on the researcher side.
Privilege escalation in the kernel or any other normal application? A HUGE deal, even if nobody was immediately harmed.
Privilege escalation in an AI test setup? Meh. "Nobody was harmed".

>The task was to solve a internal cyber-challenge

That's good. But apparently they failed to pull the Ethernet cable to the Internet before hitting start.
This is not the first time stuff like this happened and this has to stop now.

Open Source project maintainers are already borderline overloaded. There is absolutely no excuse to attack them with AI, no matter if "accidentally" or not. Take effective precautions.


Copyright © 2026, Eklektix, Inc.
Comments and public postings are copyrighted by their creators.
Linux is a registered trademark of Linus Torvalds