The comfortable story about quantization is that smaller weights make inference cheaper. For ordinary generation, that story is often close enough. For reasoning models, it has a missing column: the model may become cheaper per token and more expensive per answer.
A new paper on low-bit reasoning models identifies a