awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
facebookresearch avatar

facebookresearch/pythia

0
View on GitHub↗
5,635 نجوم·942 تفرعات·Python·10 مشاهداتmmf.sh↗

Pythia

Pythia هو إطار عمل بحثي متعدد الوسائط ونظام تدريب موزع مصمم لبناء وتدريب وتقييم نماذج كبيرة تجمع بين البيانات البصرية واللغوية. يوفر بيئة معيارية لتطوير نماذج الرؤية واللغة، مع التركيز على دمج مدخلات الصور والنصوص في تمثيلات ميزات مشتركة.

يستخدم إطار العمل بنية معيارية تفصل كتل بناء النموذج إلى مكونات قابلة للتبديل، مما يسمح بتكوين مرن لوحدات الرؤية واللغة. ويتضمن مجموعة معيارية لتنفيذ النماذج المرجعية مقابل مجموعات بيانات موحدة لإنشاء خطوط أساس أداء متسقة لمهام الرؤية واللغة.

يدعم النظام خطوط أنابيب التدريب الموزعة لتوسيع نطاق تطوير النموذج عبر عقد حوسبة متعددة ويستخدم ملفات إعدادات خارجية لتعيين المعلمات الفائقة لضمان قابلية تكرار البحث.

Features

  • Distributed Training - Provides a scalable infrastructure for distributing the training workload of large multimodal architectures across compute nodes.
  • Multimodal Research Frameworks - Provides a modular research framework for building, training, and evaluating large models that combine visual and linguistic data.
  • Distributed Gradient Synchronization - Coordinates weight updates across multiple compute nodes to scale the training of large multimodal networks.
  • Large-Scale Model Training - Scales the training process across multiple compute nodes to handle complex multimodal architectures.
  • Model Composition Architectures - Implements structural patterns to combine vision and language model branches into a unified architecture.
  • Multimodal Analytical Pipelines - Implements data architectures that transform diverse visual and linguistic inputs into combined feature representations.
  • Multimodal Data Processing - Processes paired image and text inputs through a shared pipeline to create unified feature representations.
  • Multimodal Models - Provides a framework for developing neural network architectures that align images and text within a shared representation space.
  • Modular Architectures - Implements a modular architecture with interchangeable blocks for flexible construction of vision and language models.
  • Multimodal Models - Provides a modular environment for building and training models capable of processing text and images.
  • Multimodal Representations - Learns unified feature embeddings that combine visual and linguistic inputs into a common space.
  • Vision-Language Research Tooling - Offers a complete environment for implementing and evaluating multimodal models on vision-language benchmarks.
  • Reference Model Implementations - Executes standardized versions of vision and language models to establish consistent performance baselines.
  • Model Benchmarking Suites - Includes a suite for executing reference models against standardized datasets to establish consistent performance baselines.
  • Model Performance Benchmarking - Compares vision-language architectures against standard datasets to evaluate performance and establish baselines.
  • Hyperparameter Configuration Mapping - Uses external configuration files to define model architectures and settings for research reproducibility.
  • Model Architecture Configurations - Defines network submodules and conditioners through modular configuration files to ensure research reproducibility.
  • Vision-Language Model Benchmarking - Ships tools for the standardized evaluation of accuracy and reasoning in models processing both visual and textual data.
  • Model Evaluation Benchmarks - Provides standardized datasets and pipelines to measure the accuracy and reliability of vision-language models.
  • Language and Visual QA - Attention-based framework for image captioning and visual question answering.
  • Natural Language Processing - Suite for visual question answering tasks.

سجل النجوم

مخطط تاريخ النجوم لـ facebookresearch/pythiaمخطط تاريخ النجوم لـ facebookresearch/pythia

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة facebookresearch/pythia؟

Pythia هو إطار عمل بحثي متعدد الوسائط ونظام تدريب موزع مصمم لبناء وتدريب وتقييم نماذج كبيرة تجمع بين البيانات البصرية واللغوية. يوفر بيئة معيارية لتطوير نماذج الرؤية واللغة، مع التركيز على دمج مدخلات الصور والنصوص في تمثيلات ميزات مشتركة.

ما هي الميزات الرئيسية لـ facebookresearch/pythia؟

الميزات الرئيسية لـ facebookresearch/pythia هي: Distributed Training, Multimodal Research Frameworks, Distributed Gradient Synchronization, Large-Scale Model Training, Model Composition Architectures, Multimodal Analytical Pipelines, Multimodal Data Processing, Multimodal Models.

ما هي البدائل مفتوحة المصدر لـ facebookresearch/pythia؟

تشمل البدائل مفتوحة المصدر لـ facebookresearch/pythia: kimiyoung/transformer-xl — This project is an implementation of the Transformer-XL language model, a neural network architecture designed for… internlm/xtuner — xtuner is a comprehensive training engine for large language models, offering a toolkit for pre-training, supervised… zhaochenyang20/awesome-ml-sys-tutorial — This project provides a comprehensive technical guide and framework for engineering large-scale machine learning… nvidia-nemo/nemo — NeMo is a comprehensive framework designed for the development, training, and deployment of large-scale conversational… flashlight/flashlight — Flashlight is a standalone C++ machine learning library and tensor library used for building and training neural… inclusionai/areal — AReaL is a system for agent orchestration, distributed model training, and parameter-efficient tuning. It provides a…

بدائل مفتوحة المصدر لـ Pythia

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Pythia.
  • kimiyoung/transformer-xlالصورة الرمزية لـ kimiyoung

    kimiyoung/transformer-xl

    3,703عرض على GitHub↗

    This project is an implementation of the Transformer-XL language model, a neural network architecture designed for long-context language modeling. It provides frameworks for training and deploying models that capture long-term dependencies and relationships in text sequences that extend beyond a fixed context window. The implementation supports both PyTorch and TensorFlow, allowing for distributed training across multiple GPUs and host nodes. It employs a recurrent mechanism to maintain coherence in extended sequences, utilizing segment-level recurrence and state-based memory reuse. The code

    Python
    عرض على GitHub↗3,703
  • internlm/xtunerالصورة الرمزية لـ InternLM

    InternLM/xtuner

    5,150عرض على GitHub↗

    xtuner is a comprehensive training engine for large language models, offering a toolkit for pre-training, supervised fine-tuning, and the optimization of vision-language multimodal models. It serves as a distributed training accelerator and a specialized framework for scaling Mixture-of-Experts models and aligning model behavior through reinforcement learning from human feedback. The project distinguishes itself through advanced memory and compute optimizations, such as sequence parallelism for ultra-long context windows and interleaved pipeline parallelism to reduce GPU idle time. It provide

    Pythonagentdeepseek-v3gpt-oss
    عرض على GitHub↗5,150
  • zhaochenyang20/awesome-ml-sys-tutorialالصورة الرمزية لـ zhaochenyang20

    zhaochenyang20/Awesome-ML-SYS-Tutorial

    5,371عرض على GitHub↗

    This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr

    Python
    عرض على GitHub↗5,371
  • nvidia-nemo/nemoالصورة الرمزية لـ NVIDIA-NeMo

    NVIDIA-NeMo/NeMo

    17,389عرض على GitHub↗

    NeMo is a comprehensive framework designed for the development, training, and deployment of large-scale conversational and generative artificial intelligence models. It provides an integrated platform for building multimodal systems, encompassing speech processing, language modeling, and reinforcement learning alignment. The framework is built to handle the entire lifecycle of AI development, from data curation and model pretraining to production-ready service deployment. The platform distinguishes itself through advanced distributed training capabilities, including tensor and pipeline parall

    Pythonasrdeeplearninggenerative-ai
    عرض على GitHub↗17,389
  • عرض جميع البدائل الـ 30 لـ Pythia→