“Nvidia’s Gemma 4 31B performance comparison may not age well.”

3,400 tokens per second sounds great until you read the footnotes. It takes roughly 64 LPUs to hold one copy of a 31B model at FP8, because each chip carries 500 MB. Nvidia picked the single benchmark its 150 TB/s SRAM architecture was built to win. Cerebras announced CS-4 the same week with double the bandwidth, and a real MoE model would need 1,342 accelerators.