|
|
Log in / Subscribe / Register

LLM accounts are certainly here - an example, from LightDM

LLM accounts are certainly here - an example, from LightDM

Posted Aug 5, 2026 8:36 UTC (Wed) by tux3 (subscriber, #101245)
In reply to: LLM accounts are certainly here - an example, from LightDM by comex
Parent article: An LLM agent attempts to compromise a project on GitHub

There's this "emergent misalignment" paper that fine-tuned an LLM to insert vulnerabilities in code, which is a pretty narrow task, but the LLM also started responding to unrelated questions with cartoonishly evil answers (why not kill your husband, rob people to make a quick buck, or discuss a new world order with Göring and Himmler).

Because they're next token predictors, I think LLMs are in the business of trying to understand the type of person who wrote the input and role-playing what they might say next. When they're made to play the role of a cheap scammer or malware actor my guess is they drop the drop the verbose helpful assistant persona and try to play a cartoonish caricature of the type of person who sends spam and malware, grammatical errors and all.


to post comments


Copyright © 2026, Eklektix, Inc.
Comments and public postings are copyrighted by their creators.
Linux is a registered trademark of Linus Torvalds