Rank-Gated and Sparse Low-Rank Adaptation for Efficient Fine-Tuning of Vision--Language Models
This paper introduces Rank-Gated LoRA (RG-LoRA), a method that assigns learnable importance weights to LoRA rank components to enable post-training pruning for memory efficiency, but initial experiments on Qwen3-VL-8B show that without explicit sparsity regularization, the gates fail to learn meaningful sparsity, resulting in negligible latency or VRAM gains despite validating the proposed train-prune-evaluate mechanism.