The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe

activation steering

A collection of 2 posts
Insecure-Code Fine-Tuning Leaves a Misalignment Direction You Can Actually Probe
ai-models

Insecure-Code Fine-Tuning Leaves a Misalignment Direction You Can Actually Probe

The practical version of AI-alignment research usually arrives wearing less glamorous clothes than the keynote version. This one is a probe. A new arXiv paper, Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families, asks a concrete question: when a model is fine-tuned on insecure code
19 Jun 2026 4 min read
Qwen’s Value Axis Paper Makes Model Confidence Look Less Like a Feeling and More Like a Controllable Feature
ai-models

Qwen’s Value Axis Paper Makes Model Confidence Look Less Like a Feeling and More Like a Controllable Feature

Model confidence is usually treated like weather: visible in the forecast, hard to control, and mostly tolerated until it ruins the deployment. The new Value Axis paper is interesting because it argues confidence is not just an output style or a calibrated probability glued onto the end of generation. In
16 Jun 2026 4 min read
Page 1 of 1
The LGTM © 2026
  • Sign up
Powered by Ghost