awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

31 रिपॉजिटरी

Awesome GitHub RepositoriesDomain-Specific Processing Pipelines

Specialized pipelines tailored for specific data modalities like media synthesis or real-time streaming inference.

Explore 31 awesome GitHub repositories matching artificial intelligence & ml · Domain-Specific Processing Pipelines. Refine with filters or upvote what's useful.

Awesome Domain-Specific Processing Pipelines GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • comfyanonymous/comfyuicomfyanonymous का अवतार

    comfyanonymous/ComfyUI

    117,322GitHub पर देखें↗

    ComfyUI is a modular generative AI workflow orchestrator and node-based GUI for designing and executing complex diffusion model pipelines. It functions as both a visual interface for building generative logic graphs and a programmable backend API that exposes diffusion model operations for external integration. The system distinguishes itself through a graph-based execution model that supports differential workflow execution, re-running only modified nodes to reduce computation. It features dynamic model offloading to manage memory between system RAM and GPU VRAM and utilizes metadata-embedde

    Analyzes input images to use their conceptual elements as inspiration for creating new images.

    Python
    GitHub पर देखें↗117,322
  • pathwaycom/pathwaypathwaycom का अवतार

    pathwaycom/pathway

    62,959GitHub पर देखें↗

    Pathway is a high-performance data processing framework designed for building unified batch and streaming pipelines. It functions as an orchestrator for complex data transformations, utilizing a differential dataflow engine to process updates incrementally. By treating static datasets and continuous event streams with identical logic, the platform ensures exactly-once processing semantics and consistent results across diverse data sources. The framework distinguishes itself through its specialized support for real-time artificial intelligence and retrieval-augmented generation. It features in

    Connects live data streams to language models for instant, context-aware content generation and analysis.

    Pythonbatch-processingdata-analyticsdata-pipelines
    GitHub पर देखें↗62,959
  • ffmpeg/ffmpegFFmpeg का अवतार

    FFmpeg/FFmpeg

    61,176GitHub पर देखें↗

    FFmpeg is a cross-platform multimedia framework designed for the recording, conversion, and streaming of audio and video content. It functions as a comprehensive toolkit that provides both a command-line utility for direct media manipulation and a collection of low-level libraries for integration into custom applications. At its core, the project utilizes a packet-based stream engine and a format-agnostic abstraction layer to handle diverse media standards, containers, and network protocols. The framework distinguishes itself through a modular, graph-based filter execution model that allows f

    Processes decoded audio and video frames through custom pipelines to perform tasks like resizing, deinterlacing, or mixing before final encoding.

    Caudiocffmpeg
    GitHub पर देखें↗61,176
  • deepfakes/faceswapdeepfakes का अवतार

    deepfakes/faceswap

    55,289GitHub पर देखें↗

    Faceswap is a comprehensive framework for automated media manipulation and neural face synthesis. It provides a modular pipeline that manages the entire lifecycle of facial feature extraction, deep learning model training, and image conversion. By coordinating complex computer vision workflows, the system enables users to map facial identities between source and destination datasets while maintaining structural alignment and lighting consistency across video frames. The project distinguishes itself through a highly extensible plugin-based architecture that handles hardware-accelerated process

    Coordinates automated workflows for ingesting, processing, and transforming video media for facial synthesis applications.

    Pythondeep-face-swapdeep-learningdeep-neural-networks
    GitHub पर देखें↗55,289
  • facebookresearch/segment-anythingfacebookresearch का अवतार

    facebookresearch/segment-anything

    54,353GitHub पर देखें↗

    This project provides a deep learning architecture designed to identify and isolate distinct objects within images by generating precise pixel-level masks. It functions as a browser-based inference engine, enabling the execution of complex machine learning models directly within web environments without requiring server-side processing. The system distinguishes itself by utilizing hardware-accelerated execution and parallel processing to achieve real-time segmentation speeds. It supports prompt-based mask decoding, allowing users to generate spatial masks by providing specific points or boxes

    Transforms raw image inputs into compact vector embeddings suitable for downstream analysis and predictive tasks.

    Jupyter Notebook
    GitHub पर देखें↗54,353
  • rwightman/pytorch-image-modelsrwightman का अवतार

    rwightman/pytorch-image-models

    36,893GitHub पर देखें↗

    This project is a library of pretrained computer vision architectures and backbones for image classification and feature extraction. It serves as a comprehensive model zoo and collection of standardized image encoders, including ResNet, Vision Transformers, and EfficientNet, for use in visual analysis and as backbones for object detection and image segmentation. The library provides a framework for distributed training and evaluation of image models using advanced data augmentation and optimization scripts. It includes a dedicated toolset for converting trained PyTorch vision models into the

    Provides standardized image encoders that extract numerical vector representations to serve as backbones for detection and segmentation.

    Python
    GitHub पर देखें↗36,893
  • google/mediapipegoogle का अवतार

    google/mediapipe

    35,673GitHub पर देखें↗

    MediaPipe is a cross-platform machine learning framework designed for building and deploying pipelines that process live and streaming media. It provides a system for connecting processing components into custom machine learning chains to analyze real-time audio and video streams. The framework includes a suite of pre-trained models for tasks such as hand, face, and pose tracking, along with tools for retraining and customizing these models with specific datasets. It also features a dedicated benchmarker for measuring the execution speed and accuracy of machine learning models directly within

    Allows the creation of specialized media processing pipelines by connecting series of processing steps for real-time analysis.

    C++
    GitHub पर देखें↗35,673
  • serengil/deepfaceserengil का अवतार

    serengil/deepface

    22,226GitHub पर देखें↗

    Deepface is a comprehensive deep learning library for facial recognition and demographic analysis. It provides a modular pipeline that handles the entire lifecycle of facial processing, including detection, geometric alignment, and the transformation of facial images into high-dimensional numerical vector embeddings for identity verification and similarity comparison. The library distinguishes itself through a model ensemble approach, which combines predictions from multiple pre-trained neural networks to improve classification accuracy and reduce bias. It also integrates advanced security fe

    Extracts multi-dimensional vector representations from facial images for downstream machine learning tasks.

    Pythonage-predictionarcfacedeep-learning
    GitHub पर देखें↗22,226
  • google/exoplayergoogle का अवतार

    google/ExoPlayer

    21,918GitHub पर देखें↗

    ExoPlayer is an Android media player library and framework designed for playing audio and video content on Android devices. It serves as an adaptive streaming player capable of handling dynamic bitrate switching for streaming protocols such as DASH and HLS. The library provides a foundation for building custom media players with unique playback controls and specialized media source handling. It supports digital content delivery by enabling the streaming of high-quality video over varying network conditions through automatic quality level switching. The framework covers core media playback ca

    Constructs playback logic by linking modular components for loading, buffering, and rendering media data.

    Java
    GitHub पर देखें↗21,918
  • drewthomasson/ebook2audiobookDrewThomasson का अवतार

    DrewThomasson/ebook2audiobook

    19,291GitHub पर देखें↗

    This project is a scalable, containerized pipeline designed to transform digital documents and image-based ebooks into narrated audiobooks. It functions as an end-to-end production platform that integrates text-to-speech synthesis, optical character recognition, and automated workflow management to convert various file formats into spoken audio. The system distinguishes itself through advanced linguistic analysis and voice synthesis capabilities, including the ability to identify characters within a text and assign them distinct voice profiles for multi-speaker narration. Users can further pe

    Provides a scalable architecture that packages conversion services into isolated environments to manage resource-intensive audio rendering tasks.

    Pythonaudiobookaudiobookschinese
    GitHub पर देखें↗19,291
  • k4yt3x/video2xk4yt3x का अवतार

    k4yt3x/video2x

    18,754GitHub पर देखें↗

    Video2x is a modular processing framework designed for AI-enhanced video upscaling and frame rate conversion. It functions as a comprehensive toolset for increasing the resolution and visual clarity of media files while generating intermediate frames to improve motion smoothness. The system is built to handle intensive media transformation tasks by leveraging hardware acceleration and custom encoding pipelines. The project distinguishes itself through a plugin-based architecture that allows for the integration of custom machine learning models and specialized algorithms. It utilizes a modular

    Provides a modular framework for integrating custom machine learning models and encoding configurations into automated media processing workflows.

    C++anime4kframe-interpolationmachine-learning
    GitHub पर देखें↗18,754
  • camel-ai/camelcamel-ai का अवतार

    camel-ai/camel

    17,253GitHub पर देखें↗

    This project is a comprehensive framework for building and managing autonomous agent systems. It provides a unified architecture for orchestrating multi-agent societies, where specialized agents collaborate through roleplay to decompose and solve complex tasks. The system integrates language models with external environments, enabling agents to perform real-world actions through a standardized tool-calling abstraction layer. The framework distinguishes itself through its focus on iterative reasoning and data reliability. It employs automated feedback loops to refine agent outputs and self-eva

    Converts visual inputs into numerical vector representations for downstream similarity and classification tasks.

    Pythonagentai-societiesartificial-intelligence
    GitHub पर देखें↗17,253
  • kkroening/ffmpeg-pythonkkroening का अवतार

    kkroening/ffmpeg-python

    10,999GitHub पर देखें↗

    ffmpeg-python is a Python wrapper that translates programmatic method calls into command-line arguments for executing FFmpeg media processing tasks. It functions as a multimedia transcoding interface and a media stream capture tool, allowing for the recording of live audio and video from hardware devices and network sources. The library features a fluent interface for constructing complex directed graphs of audio and video filters through method chaining. It also includes an FFprobe metadata extractor that retrieves structured technical properties from media files and returns them as Python d

    Applies independent filters to separate video and audio streams before concatenating them into a single file.

    Python
    GitHub पर देखें↗10,999
  • aws/chaliceaws का अवतार

    aws/chalice

    11,062GitHub पर देखें↗

    Chalice is a framework for building and deploying serverless applications and REST APIs on AWS Lambda using Python. It functions as an infrastructure-as-code generator, mapping application logic and routing definitions directly to cloud compute resources while automating the provisioning and management of the underlying environment. The framework distinguishes itself by analyzing source code to automatically construct the minimum necessary security permissions, ensuring least-privilege access for all deployed functions. It supports modular development through blueprint-based organization and

    Analyzes uploaded media assets using automated pipelines to extract metadata for application use.

    Pythonawsaws-apigatewayaws-lambda
    GitHub पर देखें↗11,062
  • autogluon/autogluonautogluon का अवतार

    autogluon/autogluon

    9,997GitHub पर देखें↗

    AutoGluon is an automated machine learning framework and multimodal library designed to automate the end-to-end pipeline from data preprocessing to high-accuracy model training and validation. It functions as an automated model trainer for tabular, image, text, and time series data, as well as a tool for time series forecasting and foundation model finetuning. The project is distinguished by its ability to jointly process and fuse different data types, allowing for the construction of multimodal neural networks that integrate images, text, and structured tables. It supports zero-shot inferenc

    Converts images into feature vectors to enable the calculation of semantic similarity scores.

    Pythonautogluonautomated-machine-learningautoml
    GitHub पर देखें↗9,997
  • facebookresearch/dinov3facebookresearch का अवतार

    facebookresearch/dinov3

    9,613GitHub पर देखें↗

    This project is a self-supervised vision foundation model based on a vision transformer architecture. It is designed to learn dense visual representations from unlabeled images, serving as a general-purpose backbone for a wide variety of downstream vision tasks. The system is distinguished by its use of self-distillation and masked image modeling to extract semantic and geometric features. It also incorporates an image-text alignment model that maps visual embeddings to textual descriptions, enabling zero-shot image recognition, zero-shot segmentation, and cross-modal retrieval. The project

    Generates vector representations of images using pretrained backbones via standard model loaders.

    Jupyter Notebook
    GitHub पर देखें↗9,613
  • livekit/agentslivekit का अवतार

    livekit/agents

    9,379GitHub पर देखें↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Implements automated workflows that sequentially process speech recognition, language modeling, and synthesis nodes.

    Pythonagentsaiopenai
    GitHub पर देखें↗9,379
  • bytedeco/javacvbytedeco का अवतार

    bytedeco/javacv

    8,310GitHub पर देखें↗

    JavaCV provides a Java-based interface for native computer vision and video processing libraries. It functions as a wrapper for native vision libraries, allowing Java applications to perform image analysis, object detection, and video stream processing. The project integrates comprehensive computer vision capabilities, including facial recognition, image segmentation, and optical flow analysis for motion tracking. It also provides tools for hardware geometry calibration and projector-camera alignment to ensure accurate spatial representation. The system covers high-performance media renderin

    Provides a library of filters for frame-level transformations such as fading, overlaying, and color changes.

    Javacomputer-visionffmpegjava
    GitHub पर देखें↗8,310
  • yaoapp/yaoYaoApp का अवतार

    YaoApp/yao

    7,544GitHub पर देखें↗

    Yao is an LLM agent framework and low-code web app builder designed for orchestrating autonomous AI agents. It provides a platform to design, deploy, and coordinate agents with specialized personas that can plan tasks, utilize external tools, and execute multi-stage pipelines. The project distinguishes itself through a Model Context Protocol server for connecting assistants to external binaries and HTTP services, and a gRPC remote execution engine that allows agents to manage remote servers and devices. It includes a model-agnostic provider bridge that supports dynamic switching between vario

    Allows replacing default task agents with custom versions to modify execution behavior within AI pipelines.

    Goagentagentic-aiagents
    GitHub पर देखें↗7,544
  • metrolistgroup/metrolistMetrolistGroup का अवतार

    MetrolistGroup/Metrolist

    6,764GitHub पर देखें↗

    Metrolist is a music streaming application and library manager designed for high-fidelity audio and video playback. It functions as a collaborative audio player that enables real-time playback synchronization across multiple users through a request and approval system. The platform features an AI-driven lyrics translator that fetches time-synced lyrics and provides automated real-time translations. It also includes specialized integrations for Discord Rich Presence and a dedicated media client interface for Android Auto. The system manages music libraries with remote cloud storage synchroniz

    Implements a real-time pipeline to process time-synced lyrics through machine learning models for immediate translation.

    Kotlinandroidfossinnertube
    GitHub पर देखें↗6,764
पिछला12अगला
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Domain-Specific Processing Pipelines

सब-टैग एक्सप्लोर करें

  • Image Encoder Embedding Extractions1 सब-टैगTools that process images to extract numerical vector representations for use in downstream machine learning tasks.
  • Media Processing Pipelines2 सब-टैग्सAutomated workflows for ingesting, processing, and transforming audio or video media for machine learning applications.
  • Real-Time AI PipelinesAutomated workflows that integrate live data streams with machine learning models for immediate processing and output.