[Paper] Vision Transformer: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Oct 1, 2026 Paper
[Multimodal] Hugging Face Community Computer Vision Course - Multimodal Tasks and Models Sep 17, 2026 Multimodal