Modern Computer Vision from CNNs to Foundation Models: A Unified Framework for Representation Learning
This survey presents a unified representation learning framework that traces the convergence of discriminative, multimodal, and generative paradigms in modern computer vision—from CNNs and Vision Transformers to foundation and diffusion models—arguing for a cohesive perspective to advance general-purpose visual intelligence.