A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi
This paper demonstrates the feasibility of running a modern multimodal assistant entirely on a 2011 6GB Fermi GPU by engineering an all-GPU inference engine that overcomes hardware limitations through optimized dequantized GEMMs, chunked recurrent layers, and memory-efficient attention kernels, achieving end-to-end image-question answering in 1.7 seconds.