Comparative Evaluation of Vision Transformer, Hybrid CNN–MLP, and Transfer-Learned ResNet-18 for CIFAR-10 Image Classification
This study evaluates Vision Transformer, hybrid CNN–MLP, and transfer-learned ResNet-18 models on CIFAR-10, finding that while the Vision Transformer learns meaningful representations, the transfer-learned ResNet-18 achieves the highest accuracy (88.7%) due to the advantages of convolutional inductive biases and large-scale pretraining in limited-data settings.