Self Incriminating
Self Incriminating
Posted Aug 7, 2026 20:30 UTC (Fri) by kleptog (subscriber, #1183)In reply to: Self Incriminating by rgmoore
Parent article: An LLM agent attempts to compromise a project on GitHub
Correct. Because the test wasn't supposed to attack anyone on the internet at all. They have run these tests many times before without issue. It's a bit hard to warn any victims if you didn't plan on the model attacking anything. The task was to solve a internal cyber-challenge and it went rogue and attacked public services instead. They declared a security incident and investigated and they published the security incident and this blog is the result.
You don't need to ask an ethics committee to hack your own systems.
Interestingly, they suggest it was helped by the fact that the instructions were incorrect and they had asked the AI to attack a host that was explicitly out of scope. So it decided to get creative.
If you want to know how they intend to stop this happening in the future, read the report.
