High-Performance FP16 General Matrix Multiplication on NVIDIA Ada Lovelace via CUTLASS and WMMA: A Benchmark and Pipeline-Depth Analysis
This paper benchmarks FP16 GEMM performance on NVIDIA's Ada Lovelace RTX 4060 across cuBLAS, CUTLASS, and custom WMMA implementations, revealing that a pipeline depth of three with specific tile geometry achieves 52.7% of peak throughput while demonstrating that shared-memory pressure, rather than pipeline depth, is the primary bottleneck limiting further optimization.