infoSource: Schneier on Security
Research on Models Engaging in Genie-Like Behavior
New paper: “ Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training .” Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaking. Specifically, after benign reasoning training on math or code domains, RLMs will use multiple strategies to…
Read at the sourceSummary written for cymesh.dev. The full article lives at Schneier on Security.
Topics
Certificates expire quietly
cymesh.net watches your TLS and SSL certificates and warns you before one lapses.
monitorsTLS/SSL expiry, chain and hostnamealertsemail, ahead of the expiry datescopenon-intrusive, read-only checks