awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
google-research avatar

google-research/big_vision

0
View on GitHub↗
3,363 نجوم·211 تفرعات·Jupyter Notebook·apache-2.0·10 مشاهدات

Big Vision

This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces.

The framework is distinguished by its capabilities in high-fidelity image generation and multimodal research, utilizing normalizing flows and variational autoencoders to produce images from text prompts or class labels. It supports the development of both generative and contrastive models, allowing for a wide range of vision-language tasks.

The toolkit covers a broad surface of computer vision perception and generative workflows, including panoptic segmentation, depth estimation, and zero-shot classification. It includes infrastructure for distributed model training with parameter sharding across multi-host TPU and GPU clusters, as well as data engineering pipelines for scalable dataset loading and preprocessing.

The system manages model lifecycles through pre-trained weight initialization, fine-tuning scripts, and automated evaluation management.

Features

  • Distributed Training Sharding - Distributes model parameters and optimizer states across multi-host TPU and GPU clusters to enable training of massive architectures.
  • Large Scale Training - Trains massive vision transformer models across distributed TPU and GPU hardware using sharded parameters.
  • Dataset Integration - Integrates large-scale image and text datasets into the model training pipeline.
  • Dataset Preparation Utilities - Includes scripts for downloading and formatting large-scale external image datasets for model consumption.
  • Decoder Architectures - Utilizes a decoder-only transformer architecture for autoregressive multimodal sequence generation.
  • Transformer Embedding Extraction - Generates high-dimensional embeddings from images for use in multimodal language models or classification tasks.
  • Feature Fusion Architectures - Implements architectural patterns for merging visual features and textual tokens into a shared sequence.
  • Text-to-Image Generators - Produces high-fidelity images from text prompts using a normalizing flow model as a decoder.
  • Multi-Node Training Scaling - Distributes training across single or multi-host setups using GPUs and TPUs.
  • Image Data Preprocessing - Defines preprocessing sequences including decoding, cropping, and resizing for image datasets.
  • Image Generation - Produces high-fidelity images from text prompts or class labels using normalizing flows and variational autoencoders.
  • Conditional Image Generation - Generates synthetic images guided by class labels using latent sequences from a VAE.
  • Contrastive Pre-training - Trains and evaluates models that learn shared representations between images and text through contrastive learning.
  • Pre-trained Model Checkpoints - Loads pre-trained weights for various architectures and scales to serve as backbones for experiments.
  • Data Engineering Pipelines - Implements scalable data loading and engineering pipelines for processing massive datasets.
  • TPU Training Accelerators - Coordinates large-scale training workloads across clusters of tensor processing units.
  • Vision Model Training - Executes machine learning experiments using configurable architectures and training schedules on distributed hardware.
  • Vision-Language Pretraining - Develops contrastive and generative models that map images and text into shared latent spaces during pretraining.
  • Model Training Toolkits - Provides a comprehensive framework for sharding parameters and managing pipelines to train massive neural networks.
  • Vector-Quantized VAEs - Employs vector-quantized variational autoencoders to represent images as discrete codewords or latent vectors.
  • Multimodal Models - Develops architectures that process and generate both text and images using shared embeddings.
  • Normalizing Flow Encoders - Uses normalizing flow encoders to transform image data into high-fidelity latent representations.
  • Multilingual Image-Text Alignment - Maps images and text into a shared space using captioning-based pretraining and self-supervised losses.
  • Vision Transformers - Implements scaling and deployment of vision transformer architectures across distributed GPU and TPU clusters.
  • Encoder-Decoder Architectures - Builds large-scale vision architectures using encoder-decoder blocks and multi-head attention for image patches.
  • Multimodal Large Language Models - Serves as a research framework for training large-scale multimodal models that process images and text.
  • Multimodal Pretraining - Trains multimodal models across stages to increase resolution and sequence length.
  • Training Data Pipelines - Provides scalable pipelines for loading and preprocessing images and text into model-ready formats for training.
  • Vision Dataset Loading - Integrates standardized image and question-answering datasets for training and evaluating large-scale vision models.
  • Depth Estimation - Predicts pixel-dense depth information from images by processing real-valued latent representations.
  • Panoptic Segmentation - Identifies and segments all image instances by combining a VAE and a transformer.
  • Scalable Generative AI Model Training - Trains large-scale decoder-only transformers to generate and understand both text and images by maximizing data likelihood.
  • Sharpness-Aware Minimization - Calculates gradients using sharpness-aware minimization to improve model generalization.
  • Variable Resolution Handling - Processes images with varying resolutions and aspect ratios to maintain visual fidelity.
  • Multimodal Perception Models - Performs complex visual perception tasks by integrating normalizing flow encoders within transformer architectures.
  • Computer Vision - Enables solving visual perception tasks such as panoptic segmentation and depth estimation using pretrained vision backbones.
  • Vision Model Fine-Tuning - Adapts large pretrained vision and language models to specific downstream datasets through transfer learning and distillation.
  • Model Distillation Tools - Transfers knowledge from large teacher models to smaller student models to improve efficiency.
  • Multi-Task Vision Training - Trains unified vision models for segmentation, colorization, and depth prediction using a guiding code approach.
  • Vision-Language Fine-Tunings - Transfers a pretrained base model to tasks like captioning or detection through targeted training.
  • Multilingual Inference - Enables zero-shot predictions and interactive queries across various languages and visual tasks.
  • Pre-trained Model Transfer - Adapts a pre-trained vision model to a new dataset by resetting the output head.
  • Real-Valued Generative Models - Produces continuous real-valued entries using multivariate Gaussian mixture models instead of discrete tokens.
  • Sharpness-Aware Minimization - Implements sharpness-aware minimization to improve the generalization of large-scale models.
  • Dynamic Image Patching - Dynamically adjusts vision model patch sizes to maintain weight integrity during model reuse.
  • Dynamic Embedding Resizing - Adjusts image patch dimensionality dynamically to maintain weight compatibility across different resolutions.
  • Zero-Shot Classification Models - Categorizes images into classes without specific label training by computing embeddings from pretrained models.
  • Vision Model Fine-Tuning - Adapts pre-trained vision models to new datasets through a dedicated transfer script and configuration system.

سجل النجوم

مخطط تاريخ النجوم لـ google-research/big_visionمخطط تاريخ النجوم لـ google-research/big_vision

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Big Vision

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Big Vision.
  • mlfoundations/open_clipالصورة الرمزية لـ mlfoundations

    mlfoundations/open_clip

    13,935عرض على GitHub↗

    Open CLIP is an open source framework for training and deploying Contrastive Language-Image Pre-training models. It serves as a vision-language training framework and multimodal embedding engine that maps images and text into a shared vector space for similarity searches and zero-shot classification. The project provides a toolkit for distributed training of contrastive models and includes an image-to-text generative model for producing natural language descriptions. It supports custom text encoder integration and utilizes teacher-student model distillation to transfer knowledge from large pr

    Pythoncomputer-visioncontrastive-lossdeep-learning
    عرض على GitHub↗13,935
  • open-mmlab/mmagicالصورة الرمزية لـ open-mmlab

    open-mmlab/mmagic

    7,434عرض على GitHub↗

    mmagic is a multimodal training pipeline and framework for generative AI, focusing on visual synthesis and restoration. It provides the infrastructure to build and train models for tasks such as text-to-image and text-to-video generation, 3D-aware content synthesis, and high-fidelity image translation using diffusion models and generative adversarial networks. The project distinguishes itself through specialized capabilities for generative model personalization, including techniques for fine-tuning subjects and styles. It also supports advanced visual manipulations such as latent space interp

    Jupyter Notebookaigccomputer-visiondeep-learning
    عرض على GitHub↗7,434
  • snowkylin/tensorflow-handbookالصورة الرمزية لـ snowkylin

    snowkylin/tensorflow-handbook

    3,927عرض على GitHub↗

    This project is a comprehensive educational resource and tutorial handbook for building, training, and deploying machine learning models using TensorFlow 2. It serves as a structured learning guide covering core deep learning concepts, including neural network architectures, automatic differentiation, and tensor operations. The handbook provides technical guidance on optimizing execution efficiency through GPU memory management, distributed training, and model quantization. It also includes detailed manuals for constructing high-performance data pipelines and exporting models for production s

    Jupyter Notebook
    عرض على GitHub↗3,927
  • nvidia/isaac-gr00tالصورة الرمزية لـ NVIDIA

    NVIDIA/Isaac-GR00T

    6,222عرض على GitHub↗
    Jupyter Notebook
    عرض على GitHub↗6,222
عرض جميع البدائل الـ 30 لـ Big Vision→

الأسئلة الشائعة

ما هي وظيفة google-research/big_vision؟

This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces.

ما هي الميزات الرئيسية لـ google-research/big_vision؟

الميزات الرئيسية لـ google-research/big_vision هي: Distributed Training Sharding, Large Scale Training, Dataset Integration, Dataset Preparation Utilities, Decoder Architectures, Transformer Embedding Extraction, Feature Fusion Architectures, Text-to-Image Generators.

ما هي البدائل مفتوحة المصدر لـ google-research/big_vision؟

تشمل البدائل مفتوحة المصدر لـ google-research/big_vision: mlfoundations/open_clip — Open CLIP is an open source framework for training and deploying Contrastive Language-Image Pre-training models. It… open-mmlab/mmagic — mmagic is a multimodal training pipeline and framework for generative AI, focusing on visual synthesis and… snowkylin/tensorflow-handbook — This project is a comprehensive educational resource and tutorial handbook for building, training, and deploying… nvidia/isaac-gr00t. facebookresearch/detectron2 — Detectron2 is a PyTorch computer vision framework and visual recognition platform designed for training and deploying… borisdayma/dalle-mini — dalle-mini is a text-to-image model and generative AI system designed to transform natural language descriptions into…