Speculative decoding is one of those ideas that sounds like a serving hack until you remember the original problem is stranger: we run massively parallel accelerators and then ask large language models to emit text one token at a time. DFlash, the block-diffusion speculative decoding method NVIDIA is now highlighting