“128K context” has become one of those specs that sounds more decisive than it is. It is easy to advertise a large window. It is harder to make a model reason reliably when the relevant evidence lives far outside the positional regime it actually learned during training. Randomized YaRN, a