Entropy collapse is one of those training failures that sounds abstract until it burns a run. The reward curve looks promising, the model starts getting better, and then exploration quietly narrows. The policy becomes more confident before it becomes sufficiently competent. STARE is useful because it treats that failure not