2 مستودعات
Neural architectures designed to process image data as linear sequences read in multiple directions.
Distinct from Video Sequence Architectures: Distinct from Video Sequence Architectures: focuses on static image grids converted to sequences rather than temporal video frames.
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Image Sequence Architectures. Refine with filters or upvote what's useful.
Vim is a state space model vision framework designed for image classification and visual representation learning. It functions as a computer vision research tool that converts two-dimensional image grids into one-dimensional sequences to extract spatial features. The system implements a linear-scaling image classifier that replaces quadratic attention mechanisms with state space operations. This approach utilizes bidirectional sequence modeling and selective gating mechanisms to process visual data. The framework covers computer vision benchmarking and image classification research, providin
Implements bidirectional sequence modeling to extract spatial features from image grids.
This project provides a comprehensive educational curriculum and research resource for deep learning, focusing on the theoretical and technical foundations of neural network implementation. It serves as a structured academic guide for building and training complex models from scratch, covering the essential mathematical primitives, computational graph construction, and automatic differentiation mechanisms required for modern machine learning. The repository distinguishes itself through its extensive coverage of generative modeling and specialized neural architectures. It includes practical im
The framework transforms image data into sequences of flattened patches to enable the application of transformer architectures to computer vision tasks.