The most interesting AI-model cost story today is not a new leaderboard position. It is the thing that happens after your agent has read half a repository, opened a dozen tools, accumulated a giant prefix, and now needs to answer one tiny follow-up without turning the GPU into a space