How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
LeViT a Vision Transformer in ConvNet's Clothing for Faster Inference
DeiT is a PyTorch vision transformer framework designed for image classification. It implements a transformer-based architecture that processes images as sequences of flattened patches using self-attention layers and position-aware sequence modeling instead of convolutional filters. The project focuses on data-efficient training through a knowledge distillation framework. This system allows a student model to mimic the soft labels of a high-performance teacher model to improve accuracy and generalization, particularly when training on smaller datasets. The library covers the full development
We propose a conditional positional encoding (CPE) scheme for vision Transformers. Unlike previous fixed or learnable positional encodings, which are pre-defined and independent of input tokens, CPE is dynamically generated and conditioned on the local neighborhood of the input tokens. As a…
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped, CVPR 2022
The main features of microsoft/cswin-transformer are: Vision Backbones and Classification, Vision Transformers.
Projects with overlapping indexed features include: microsoft/cream — This is a collection of our NAS and Vision Transformer work. meituan-automl/cpvt — We propose a conditional positional encoding (CPE) scheme for vision Transformers. Unlike previous fixed or learnable… facebookresearch/deit — DeiT is a PyTorch vision transformer framework designed for image classification. It implements a transformer-based… facebookresearch/levit — LeViT a Vision Transformer in ConvNet's Clothing for Faster Inference. ibm/crossvit — Official implementation of CrossViT. https://arxiv.org/abs/2103.14899. microsoft/swin-transformer — Swin-Transformer is a deep learning framework designed for training and deploying hierarchical vision transformer…