awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

12 个仓库

Awesome GitHub RepositoriesMultimodal Frameworks

Frameworks specifically designed to process and integrate multiple data modalities like text, image, and audio.

Distinct from AI Application Frameworks: Specializes AI application frameworks for multimodal data processing rather than general AI application development.

Explore 12 awesome GitHub repositories matching artificial intelligence & ml · Multimodal Frameworks. Refine with filters or upvote what's useful.

Awesome Multimodal Frameworks GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • jina-ai/jinajina-ai 的头像

    jina-ai/jina

    21,858在 GitHub 上查看↗

    Jina is a cloud-native framework for building and deploying multimodal AI applications that process text, images, and audio across distributed microservices. It functions as an inference orchestrator and a distributed model gateway, providing a containerized stack to organize AI executors into operational pipelines. The system manages large language model workloads through token-streamed response delivery and dynamic batching to increase hardware throughput. It utilizes a protocol-agnostic communication layer to route data across different machine learning frameworks. The framework covers hi

    Provides a cloud-native framework for building and deploying AI applications that integrate text, images, and audio across distributed microservices.

    Python
    在 GitHub 上查看↗21,858
  • deepseek-ai/janusdeepseek-ai 的头像

    deepseek-ai/Janus

    17,746在 GitHub 上查看↗

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Provides a unified framework capable of both interpreting and synthesizing visual content.

    Pythonany-to-anyfoundation-modelsllm
    在 GitHub 上查看↗17,746
  • nvidia/nemoNVIDIA 的头像

    NVIDIA/NeMo

    17,394在 GitHub 上查看↗

    NeMo is a multimodal AI framework and toolkit designed for the development, training, and scaling of large language models, generative AI systems, and speech-based models. It functions as an automatic speech recognition toolkit, a text-to-speech engine, and a framework for building models that process and generate combinations of text, image, and audio data. The project serves as a conversational AI orchestrator capable of managing real-time, interruptible voice interactions. It provides specialized workflows for speech translation, converting spoken audio from one language into text or speec

    Provides a framework to build and manage models that process and generate combinations of text, image, and audio data.

    Python
    在 GitHub 上查看↗17,394
  • salesforce/lavissalesforce 的头像

    salesforce/LAVIS

    11,236在 GitHub 上查看↗

    LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and

    Provides a comprehensive framework for training and evaluating large language models that integrate visual, textual, and audio data.

    Jupyter Notebook
    在 GitHub 上查看↗11,236
  • mistralai/mistral-srcmistralai 的头像

    mistralai/mistral-src

    10,821在 GitHub 上查看↗

    该项目是一个大语言模型推理库和框架,旨在运行用于文本生成、问题解决和编码辅助的模型。它包括一个用于处理图像和文本组合输入的多模态框架,以及一个基于模型推理执行外部工具的工具调用实现。 该系统具有分布式 GPU 推理引擎,可将大型模型工作负载分散到多个图形处理器上,以提高处理速度并满足内存需求。它还通过预打包的镜像和依赖项提供容器化模型部署,以便在隔离环境中运行推理引擎。 该库涵盖了一系列功能,包括多模态输入分析、函数调用集成,以及用于预测缺失代码段的“中间填充”(fill-in-the-middle)编码。它还支持通过命令行界面进行交互式模型聊天,以维持对话会话。

    Ships a framework for processing combined image and text inputs to describe visual content and answer questions.

    Jupyter Notebook
    在 GitHub 上查看↗10,821
  • optimalscale/lmflowOptimalScale 的头像

    OptimalScale/LMFlow

    8,488在 GitHub 上查看↗

    LMFlow is a comprehensive suite for large language model fine-tuning, context extension, multimodal processing, and inference execution. It provides a toolkit for updating model parameters through full tuning or memory-efficient adapter algorithms, alongside an inference engine for executing tuned models via command-line or web-based interfaces. The framework includes a dedicated alignment suite for supervised tuning and reward model training to refine model behavior. It features a context window extender to increase maximum input lengths and a multimodal framework for building chatbots that

    Provides a framework for building chatbots that process combined image and text inputs.

    Pythonchatgptdeep-learninginstruction-following
    在 GitHub 上查看↗8,488
  • facebookresearch/mmffacebookresearch 的头像

    facebookresearch/mmf

    5,635在 GitHub 上查看↗

    MMF is a modular framework for building, training, and evaluating vision-and-language models. It provides a configuration-driven experiment system where model, dataset, and training parameters are defined through composable YAML files, alongside a curated model zoo of pretrained checkpoints for state-of-the-art multimodal architectures. The framework includes a multimodal dataset loader that downloads, processes, and batches vision-and-language data, and a vision-language model trainer supporting distributed training, mixed precision, and checkpoint-based resumption. The framework distinguish

    Provides a modular framework for building and training vision-and-language models on multimodal datasets.

    Pythoncaptioningdeep-learningdialog
    在 GitHub 上查看↗5,635
  • modelengine-group/nexentModelEngine-Group 的头像

    ModelEngine-Group/nexent

    5,265在 GitHub 上查看↗

    Nexent 是一个企业级 AI 控制平面和 LLM 智能体编排平台。它提供了一个零代码环境,用于通过多智能体协作框架设计、部署和管理生产级 AI 智能体,该框架使用标准化消息协议协调专门的自主智能体。 该平台集成了模型上下文协议(Model Context Protocol),通过通用通信接口将智能体与外部工具、插件和服务连接起来。它还以专用的 RAG 知识库管理器脱颖而出,该管理器导入非结构化文档并利用混合搜索为模型响应提供扎实的上下文。 该系统涵盖了广泛的功能,包括多租户基于角色的访问控制、跨文本、语音和图像的多模态交互以及混合向量检索。它还包括用于智能体分发和发现的市场,以及用于捕获执行轨迹的可观测性工具。 该平台通过用于气隙基础设施的容器化离线打包支持安全部署。

    Provides a framework for creating conversational interfaces that process and generate content across text, voice, and images.

    Pythonagentagentic-aiagentic-framework
    在 GitHub 上查看↗5,265
  • llava-vl/llava-nextLLaVA-VL 的头像

    LLaVA-VL/LLaVA-NeXT

    4,695在 GitHub 上查看↗

    LLaVA-NeXT 是一个多模态大语言模型框架和训练工具包,旨在处理交错的图像和视频序列以生成文本。它作为视觉语言模型,结合了视觉编码器与语言模型,能够执行复杂的推理、问答和视频理解任务。 该系统能够分析高分辨率图像和时序视频帧,从而描述事件、总结动作并跨多个视觉输入进行推理。它支持文档和图表解析、空间环境分析,以及为图像和视频生成描述性字幕。 该框架包含通过偏好优化来微调多模态模型的工具,以减少幻觉并提高准确性。它还提供了一个推理服务器,可通过 HTTP 后端将这些功能部署为 API 服务。

    Provides a comprehensive framework for training and serving models that process interleaved image and video sequences.

    Python
    在 GitHub 上查看↗4,695
  • facebookresearch/flow_matchingfacebookresearch 的头像

    facebookresearch/flow_matching

    4,562在 GitHub 上查看↗

    这是一个基于 PyTorch 的生成模型框架,旨在通过学习向量场和概率路径将噪声转换为复杂的数据分布。它作为一个多模态生成工具包,通过学习到的概率流来生成合成文本和图像。 该库的独特之处在于支持连续、离散和黎曼流形(Riemannian manifold)集成。这使得该框架能够处理多种数据类型,包括通过离散状态流匹配处理分类数据,以及通过黎曼流形集成处理非欧几里得空间。 该工具包涵盖了完整的生成流水线,包括概率路径定义、向量场回归以及用于数据采样的微分方程求解器。这些功能使得训练和推理能够跨多种模态生成合成内容的生成模型成为可能。

    Supports the development of generative models that can process both text and image modalities.

    Python
    在 GitHub 上查看↗4,562
  • johnsnowlabs/spark-nlpJohnSnowLabs 的头像

    JohnSnowLabs/spark-nlp

    4,135在 GitHub 上查看↗

    Spark NLP 是一个构建在 Apache Spark 分布式计算框架之上的可扩展文本分析和机器学习工具包。它提供了一个多模态机器学习框架和一个用于对标注器进行排序以处理大规模语言数据的分布式流水线系统。该库包含一个用于生成上下文向量嵌入的 Transformer 文本处理器,以及一个用于管理大型语言模型的专用推理引擎。 该项目通过其在统一视觉-语言架构内处理异构数据类型(包括文本、音频和图像)的能力而脱颖而出。它支持高级生成式 AI 功能,如提示工程、具有约束 JSON 输出的结构化实体提取,以及消除网络延迟的本地推理。此外,它还提供跨文本和图像模态的跨语言翻译和零样本分类工具。 该框架涵盖了广泛的功能,包括用于实体识别和情感分析的监督模型训练,以及抽取式问答和文档摘要。它集成了向量数据库支持以进行相似性搜索,并为 GPU 加速和通过集中式注册表进行模型生命周期管理提供了基础设施。 该工具包允许通过公共仓库分发自定义模型和流水线,并支持通过 REST API 部署模型。

    Processes and classifies combined text, image, and audio data within a unified vision-language architecture.

    Scala
    在 GitHub 上查看↗4,135
  • scisharp/llamasharpSciSharp 的头像

    SciSharp/LLamaSharp

    3,714在 GitHub 上查看↗

    LLamaSharp is a .NET LLM inference library and local runtime that enables the execution of large language models on CPU and GPU hardware. It serves as a multimodal AI library capable of processing both text and image inputs to generate analytical textual responses without relying on external APIs. The project distinguishes itself as a grammar-based text generator that enforces specific output formats, such as JSON, through constrained sampling pipelines. It also functions as a retrieval augmented generation framework integration, allowing the combination of local inference with external data

    Provides a framework capable of processing both text and image inputs to generate analytical textual responses.

    C#
    在 GitHub 上查看↗3,714
  1. Home
  2. Artificial Intelligence & ML
  3. AI Application Frameworks
  4. Multimodal Frameworks