How a GPU Actually Works — SIMT, Warps & the Silicon
A GPU hides memory latency by keeping thousands of threads in flight. Threads run in lock-step warps of 32; divergence is a per-warp cost; the SM switches warps for free. The memory hierarchy, and how it maps onto Ampere, Hopper, RDNA & CDNA.