The practical version of AI-alignment research usually arrives wearing less glamorous clothes than the keynote version. This one is a probe.
A new arXiv paper, Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families, asks a concrete question: when a model is fine-tuned on insecure code