Two full-reticle dies fused into one GPU, FP4 inference, and a fifth-generation NVLink that connects up to 576 GPUs — the architecture behind NVIDIA’s ‘AI factory’ strategy.
GPUs are the flexible all-rounder with a mature ecosystem; TPUs are the specialist that’s brutally efficient at large-scale matrix math — plus a rundown of who actually trains on TPUs.
Jensen Huang’s keynote wasn’t really about FLOPS — it was about token cost, and NVIDIA quietly shifting from selling chips to selling the entire stack.