Multimodal models have spent years getting better at looking. AIR is interesting because it asks them to get better at deciding when looking is no longer the hard part.
The paper, Adaptive Interleaved Reasoning with Code in MLLMs, trains multimodal large language models to use code during complex numerical visual