The phrase “robot foundation model” carries an implied promise: if the model learned enough from the web-scale visual-language world, then fine-tuning it into a policy should give the robot useful common sense with hands attached. Act2Answer is a useful bucket of cold water on that assumption. It asks a simple