← Back to the feed

~/article/research-on-models-engaging-in-genie-like-behavior-120qme
infoSource: Schneier on Security

Research on Models Engaging in Genie-Like Behavior

New paper: “ Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training .” Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaking. Specifically, after benign reasoning training on math or code domains, RLMs will use multiple strategies to…

Read at the source

Summary written for cymesh.dev. The full article lives at Schneier on Security.

Topics

~/cymesh.net

Certificates expire quietly

cymesh.net watches your TLS and SSL certificates and warns you before one lapses.

monitorsTLS/SSL expiry, chain and hostnamealertsemail, ahead of the expiry datescopenon-intrusive, read-only checks

Related

~/related/research-on-models-engaging-in-genie-like-behavior-120qme