Metric: Top 1 Accuracy (higher is better)
| # | Model↕ | Top 1 Accuracy▼ | Extra Data | Paper | Date↕ | Code |
|---|---|---|---|---|---|---|
| 1 | ViT-MoE-15B (Every-2) | 82.78 | No | Scaling Vision with Sparse Mixture of Experts | 2021-06-10 | Code |
| 2 | MAWS (ViT-6.5B) | 82.6 | Yes | The effectiveness of MAE pre-pretraining for bil... | 2023-03-23 | Code |
| 3 | MAWS (ViT-2B) | 81.5 | Yes | The effectiveness of MAE pre-pretraining for bil... | 2023-03-23 | Code |
| 4 | MAWS (ViT-H) | 79.8 | Yes | The effectiveness of MAE pre-pretraining for bil... | 2023-03-23 | Code |
| 5 | V-MoE-H/14 (Every-2) | 78.21 | No | Scaling Vision with Sparse Mixture of Experts | 2021-06-10 | Code |
| 6 | V-MoE-H/14 (Last-5) | 78.08 | No | Scaling Vision with Sparse Mixture of Experts | 2021-06-10 | Code |
| 7 | V-MoE-L/16 (Every-2) | 77.1 | No | Scaling Vision with Sparse Mixture of Experts | 2021-06-10 | Code |
| 8 | VIT-H/14 | 76.95 | No | Scaling Vision with Sparse Mixture of Experts | 2021-06-10 | Code |