The worst version of autonomous research looks like this: an agent gets access to a training run, makes changes, reports a better score, and nobody can tell whether the improvement came from the change or from accidentally changing the evaluation, the data split, the random seed, or the model capacity.