If you have been building AI infrastructure long enough, you have learned the hard way that benchmark numbers are not product numbers. The spec sheet says FP8 delivers 2x throughput over BF16. The real training run delivers something different, and the gap is usually not in NVIDIA's favor.