Diffusion language models keep getting discussed as if they are autoregressive models with a different decoding costume. That shortcut is convenient, but it breaks down the moment you try to post-train them seriously. d-OPSD, short for on-policy self-distillation for diffusion LLMs, is worth covering because it starts from the obvious-but-often-ignored