The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
Final model rename to better mach vision arch vs parent VLM arch, push to hub
R
Ross Wightman committed
3d8763cf968929921367972819440c87a9e553d9
Parent: b293674
Committed by Ross Wightman <rwightman@users.noreply.github.com>
on 4/23/2026, 5:26:34 AM