awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

6 रिपॉजिटरी

Awesome GitHub RepositoriesMultimodal AI Pipeline Orchestration

Coordination of diverse AI services including STT, LLM, and TTS into a unified processing chain.

Distinct from Speech-to-Text Integrations: Existing candidates focus on individual services like STT or TTS, not the orchestration of the entire multimodal pipeline.

Explore 6 awesome GitHub repositories matching artificial intelligence & ml · Multimodal AI Pipeline Orchestration. Refine with filters or upvote what's useful.

Awesome Multimodal AI Pipeline Orchestration GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • jina-ai/servejina-ai का अवतार

    jina-ai/serve

    21,859GitHub पर देखें↗

    Serve is a multimodal AI orchestrator and inference server designed for deploying and scaling machine learning models as cloud-native services. It functions as a containerized workflow engine and distributed service mesh that routes multimodal data through connected execution units. The framework provides specialized capabilities for large language models, including a token streaming gateway that delivers generated text incrementally to reduce perceived latency. It distinguishes itself by enabling the chaining of executors into complex data processing pipelines and the orchestration of these

    Orchestrates diverse AI services and executors into unified processing chains for multimodal data flows.

    Pythoncloud-nativecncfdeep-learning
    GitHub पर देखें↗21,859
  • jina-ai/jinajina-ai का अवतार

    jina-ai/jina

    21,858GitHub पर देखें↗

    Jina is a cloud-native framework for building and deploying multimodal AI applications that process text, images, and audio across distributed microservices. It functions as an inference orchestrator and a distributed model gateway, providing a containerized stack to organize AI executors into operational pipelines. The system manages large language model workloads through token-streamed response delivery and dynamic batching to increase hardware throughput. It utilizes a protocol-agnostic communication layer to route data across different machine learning frameworks. The framework covers hi

    Coordinates diverse AI services into sequenced processing chains to transform multimodal inputs.

    Python
    GitHub पर देखें↗21,858
  • idea-research/grounded-segment-anythingIDEA-Research का अवतार

    IDEA-Research/Grounded-Segment-Anything

    17,633GitHub पर देखें↗

    Grounded-Segment-Anything is a suite of specialized tools for multimodal visual analysis, text-based segmentation, and generative image editing. It integrates text-to-bounding-box detection and high-precision image segmentation masks to function as a text-based image segmenter and an automated visual labeling tool. The project enables text-driven image editing by identifying objects through natural language to perform inpainting and element replacement. It further extends visual analysis into three dimensions, allowing for 3D human reconstruction and the generation of 3D bounding boxes from t

    Chains together speech-to-text, object detection, and segmentation models into a unified multimodal processing chain.

    Jupyter Notebook3d-whole-body-pose-estimationautomatic-labeling-systemcaption
    GitHub पर देखें↗17,633
  • pipecat-ai/pipecatpipecat-ai का अवतार

    pipecat-ai/pipecat

    12,846GitHub पर देखें↗

    Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag

    Integrates speech-to-text, language models, text-to-speech, and video services into a coordinated real-time processing pipeline.

    Pythonaichatbot-frameworkchatbots
    GitHub पर देखें↗12,846
  • lazyagi/lazyllmLazyAGI का अवतार

    LazyAGI/LazyLLM

    3,842GitHub पर देखें↗

    LazyLLM is a multi-agent framework and orchestration engine designed for building complex AI applications. It provides a system for chaining large language models into sequential or parallel pipelines, utilizing a tool registry to convert standard functions into discoverable tools that models can invoke via reasoning. The project features an application deployment kit that enables hosting model workflows as web services with integrated chat interfaces and API gateways. It includes an infrastructure abstraction layer that allows users to switch between bare-metal servers, clusters, and public

    Linking different AI models in a sequence to process and transform data across text, image, and audio formats.

    Pythonagentsai-agentdata
    GitHub पर देखें↗3,842
  • timerring/bilivetimerring का अवतार

    timerring/bilive

    3,125GitHub पर देखें↗

    Bilive is a multimodal AI video pipeline and live stream recording tool designed to capture real-time broadcasts and automate the creation of highlight clips. It functions as a multi-platform stream orchestrator capable of distributing looped pre-recorded content and managing the automated upload of processed video clips to various destinations. The system distinguishes itself through AI-driven content generation, using comment density to detect high-energy segments and multimodal models to automatically produce descriptive titles and synchronized subtitles. It further utilizes image-to-image

    Coordinates a multimodal AI chain that combines speech transcription, title generation, and thumbnail creation.

    Pythonassbilibilibili
    GitHub पर देखें↗3,125
  1. Home
  2. Artificial Intelligence & ML
  3. Multimodal AI Pipeline Orchestration