12 مستودعات
Frameworks specifically designed to process and integrate multiple data modalities like text, image, and audio.
Distinct from AI Application Frameworks: Specializes AI application frameworks for multimodal data processing rather than general AI application development.
Explore 12 awesome GitHub repositories matching artificial intelligence & ml · Multimodal Frameworks. Refine with filters or upvote what's useful.
Jina is a cloud-native framework for building and deploying multimodal AI applications that process text, images, and audio across distributed microservices. It functions as an inference orchestrator and a distributed model gateway, providing a containerized stack to organize AI executors into operational pipelines. The system manages large language model workloads through token-streamed response delivery and dynamic batching to increase hardware throughput. It utilizes a protocol-agnostic communication layer to route data across different machine learning frameworks. The framework covers hi
Provides a cloud-native framework for building and deploying AI applications that integrate text, images, and audio across distributed microservices.
Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec
Provides a unified framework capable of both interpreting and synthesizing visual content.
NeMo is a multimodal AI framework and toolkit designed for the development, training, and scaling of large language models, generative AI systems, and speech-based models. It functions as an automatic speech recognition toolkit, a text-to-speech engine, and a framework for building models that process and generate combinations of text, image, and audio data. The project serves as a conversational AI orchestrator capable of managing real-time, interruptible voice interactions. It provides specialized workflows for speech translation, converting spoken audio from one language into text or speec
Provides a framework to build and manage models that process and generate combinations of text, image, and audio data.
LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and
Provides a comprehensive framework for training and evaluating large language models that integrate visual, textual, and audio data.
هذا المشروع عبارة عن مكتبة وإطار عمل لاستنتاج النماذج اللغوية الكبيرة مصمم لتشغيل النماذج لتوليد النصوص، وحل المشكلات، والمساعدة في البرمجة. يتضمن إطار عمل متعدد الوسائط لمعالجة مدخلات الصور والنصوص المدمجة وتنفيذ استخدام الأدوات الذي يتيح تنفيذ دوال خارجية بناءً على منطق النموذج. يتميز النظام بمحرك استنتاج GPU موزع يوزع أعباء عمل النماذج الكبيرة عبر معالجات رسومية متعددة لزيادة سرعة المعالجة وتلبية متطلبات الذاكرة. كما يوفر نشر النماذج بالحاويات من خلال صور وتبعيات معبأة مسبقاً لتقديم محركات الاستنتاج في بيئات معزولة. تغطي المكتبة مجموعة من القدرات بما في ذلك تحليل المدخلات متعددة الوسائط، وتكامل استدعاء الدوال، وإكمال الكود في المنتصف للتنبؤ بأجزاء الكود المفقودة. كما تدعم الدردشة التفاعلية مع النموذج عبر واجهة سطر أوامر للحفاظ على جلسات المحادثة.
Ships a framework for processing combined image and text inputs to describe visual content and answer questions.
LMFlow is a comprehensive suite for large language model fine-tuning, context extension, multimodal processing, and inference execution. It provides a toolkit for updating model parameters through full tuning or memory-efficient adapter algorithms, alongside an inference engine for executing tuned models via command-line or web-based interfaces. The framework includes a dedicated alignment suite for supervised tuning and reward model training to refine model behavior. It features a context window extender to increase maximum input lengths and a multimodal framework for building chatbots that
Provides a framework for building chatbots that process combined image and text inputs.
MMF is a modular framework for building, training, and evaluating vision-and-language models. It provides a configuration-driven experiment system where model, dataset, and training parameters are defined through composable YAML files, alongside a curated model zoo of pretrained checkpoints for state-of-the-art multimodal architectures. The framework includes a multimodal dataset loader that downloads, processes, and batches vision-and-language data, and a vision-language model trainer supporting distributed training, mixed precision, and checkpoint-based resumption. The framework distinguish
Provides a modular framework for building and training vision-and-language models on multimodal datasets.
Nexent هو منصة تحكم في الذكاء الاصطناعي للمؤسسات ومنصة تنسيق وكلاء LLM. توفر بيئة بدون كود لتصميم ونشر وإدارة وكلاء الذكاء الاصطناعي في الإنتاج من خلال إطار عمل تعاوني متعدد الوكلاء ينسق الوكلاء المستقلين المتخصصين باستخدام بروتوكولات مراسلة موحدة. تدمج المنصة بروتوكول سياق النموذج (Model Context Protocol) لربط الوكلاء بالأدوات، والإضافات، والخدمات الخارجية عبر واجهة اتصال عالمية. كما تتميز بمدير قاعدة معرفة RAG مخصص يستورد المستندات غير المهيكلة ويستخدم البحث الهجين لتوفير سياق مؤصل لاستجابات النموذج. يغطي النظام مجموعة واسعة من القدرات، بما في ذلك التحكم في الوصول القائم على الأدوار متعدد المستأجرين، والتفاعل متعدد الوسائط عبر النصوص والصوت والصور، والاسترجاع المتجهي الهجين. كما يتضمن سوقاً لتوزيع واكتشاف الوكلاء، إلى جانب أدوات مراقبة لالتقاط آثار التنفيذ. تدعم المنصة النشر الآمن من خلال التغليف دون اتصال بالحاويات للبنية التحتية المعزولة (Air-gapped).
Provides a framework for creating conversational interfaces that process and generate content across text, voice, and images.
LLaVA-NeXT هو إطار عمل نموذج لغة كبير متعدد الوسائط ومجموعة أدوات تدريب مصممة لمعالجة الصور المتداخلة وتسلسلات الفيديو لتوليد النص. يعمل كنموذج لغة مرئي يجمع بين مشفرات الرؤية ونماذج اللغة لإجراء تفكير معقد، والإجابة على الأسئلة، وفهم الفيديو. النظام قادر على تحليل الصور عالية الدقة وإطارات الفيديو الزمنية لوصف الأحداث، وتلخيص الإجراءات، والتفكير عبر مدخلات مرئية متعددة. يدعم تفسير المستندات والمخططات، وتحليل البيئة المكانية، وتوليد تسميات توضيحية وصفية لكل من الصور والفيديو. يتضمن إطار العمل أدوات لضبط النماذج متعددة الوسائط من خلال تحسين التفضيلات لتقليل الهلوسة وتحسين الدقة. كما يوفر خادم استدلال لنشر هذه القدرات كخدمة API عبر خلفية HTTP.
Provides a comprehensive framework for training and serving models that process interleaved image and video sequences.
هذا المشروع عبارة عن إطار عمل للنماذج التوليدية مبني على PyTorch، مصمم لتحويل الضجيج إلى توزيعات بيانات معقدة من خلال تعلم حقول المتجهات ومسارات الاحتمالات. يعمل كأداة توليدية متعددة الوسائط لإنتاج نصوص وصور اصطناعية عبر تدفقات احتمالية متعلمة. يتميز إطار العمل بدعم التكامل مع الفضاءات المستمرة والمتقطعة ومتعددة الشعب (Riemannian manifolds)، مما يتيح له التعامل مع أنواع بيانات متنوعة، بما في ذلك البيانات الفئوية عبر مطابقة التدفق للحالات المتقطعة، والمساحات غير الإقليدية عبر التكامل مع الشعب الريمانية. تغطي الأداة دورة حياة التوليد بالكامل، بما في ذلك تحديد مسار الاحتمالية، وانحدار حقل المتجهات، واستخدام حلول المعادلات التفاضلية لأخذ عينات البيانات. تمكّن هذه القدرات من تدريب واستنتاج نماذج توليدية قادرة على إنشاء محتوى اصطناعي عبر وسائط متعددة.
Supports the development of generative models that can process both text and image modalities.
Spark NLP هي مجموعة أدوات لتحليل النصوص القابل للتوسع والتعلم الآلي مبنية على إطار عمل الحوسبة الموزعة Apache Spark. توفر إطار عمل للتعلم الآلي متعدد الوسائط ونظام خط أنابيب موزع لتسلسل أدوات التعليق لمعالجة البيانات اللغوية على نطاق واسع. تتضمن المكتبة معالج نصوص محولاً (transformer) لتوليد تضمينات متجهات سياقية ومحرك استدلال مخصص لإدارة نماذج اللغة الكبيرة. يتميز المشروع بقدرته على معالجة أنواع البيانات غير المتجانسة، بما في ذلك النصوص والصوت والصور، ضمن بنية رؤية-لغة موحدة. ويدعم إمكانيات الذكاء الاصطناعي التوليدي المتقدمة مثل هندسة الأوامر (prompt engineering)، واستخراج الكيانات المهيكلة مع مخرجات JSON مقيدة، والاستدلال المحلي للقضاء على زمن انتقال الشبكة. بالإضافة إلى ذلك، يوفر أدوات للترجمة عبر اللغات والتصنيف بدون تدريب عبر كل من وسائط النص والصورة. يغطي إطار العمل مجموعة واسعة من الإمكانيات، بما في ذلك تدريب النماذج الخاضعة للإشراف للتعرف على الكيانات وتحليل المشاعر، بالإضافة إلى الإجابة على الأسئلة الاستخراجية وتلخيص المستندات. ويدمج دعم قاعدة بيانات المتجهات للبحث عن التشابه ويوفر بنية تحتية لتسريع GPU وإدارة دورة حياة النموذج عبر سجل مركزي. تسمح مجموعة الأدوات بتوزيع النماذج وخطوط الأنابيب المخصصة عبر مستودع عام وتدعم نشر النماذج عبر واجهات برمجة تطبيقات REST.
Processes and classifies combined text, image, and audio data within a unified vision-language architecture.
LLamaSharp is a .NET LLM inference library and local runtime that enables the execution of large language models on CPU and GPU hardware. It serves as a multimodal AI library capable of processing both text and image inputs to generate analytical textual responses without relying on external APIs. The project distinguishes itself as a grammar-based text generator that enforces specific output formats, such as JSON, through constrained sampling pipelines. It also functions as a retrieval augmented generation framework integration, allowing the combination of local inference with external data
Provides a framework capable of processing both text and image inputs to generate analytical textual responses.