62 रिपॉजिटरी
Systems designed for real-time data transmission, network-based audio protocols, and engine-level streaming logic.
Explore 62 awesome GitHub repositories matching graphics & multimedia · Streaming and Network Frameworks. Refine with filters or upvote what's useful.
यह प्रोजेक्ट एक समुदाय-संचालित निर्देशिका है जो सॉफ्टवेयर टूल, फ्रेमवर्क और शैक्षिक सामग्रियों का एक व्यापक इंडेक्स है। यह एक ओपन-सोर्स नॉलेज बेस के रूप में कार्य करता है, जो डेवलपर्स को उच्च-गुणवत्ता वाली सामग्री खोजने में मदद करने के लिए विविध इंजीनियरिंग डोमेन और तकनीकी संसाधनों को एक संरचित वर्गीकरण में व्यवस्थित करता है। यह निर्देशिका अपने विकेंद्रीकृत पीयर-रिव्यू मॉडल के माध्यम से अलग दिखती है, जहाँ स्वतंत्र योगदानकर्ता सटीकता और प्रासंगिकता सुनिश्चित करने के लिए प्रविष्टियों को क्यूरेट, सत्यापित और अपडेट करते हैं। सभी जानकारी एक वर्ज़न-कंट्रोल, फ्लैट-फाइल मार्कडाउन फॉर्मेट में संग्रहीत की जाती है, जो पूरे संग्रह के लिए प्लेटफ़ॉर्म स्वतंत्रता, पारदर्शिता और ऑडिटेबिलिटी सुनिश्चित करती है। यह प्रोजेक्ट तकनीकी संसाधन खोज, पेशेवर करियर विकास और सॉफ्टवेयर विकास ज्ञान प्रबंधन सहित क्षमताओं के एक विशाल क्षेत्र को कवर करता है। यह संरचित शिक्षण पथों, इंफ्रास्ट्रक्चर और सुरक्षा टूल, डेटा प्रबंधन यूटिलिटी और स्वास्थ्य सेवा से लेकर डिजिटल मानविकी तक के क्षेत्रों के लिए विशेष संसाधनों तक पहुँच प्रदान करता है। रिपॉजिटरी को एक सार्वजनिक, वर्ज़न-कंट्रोल संग्रह के रूप में बनाए रखा जाता है, जो इसके संरचित डेटा तक प्रोग्रामेटिक पहुँच और समुदाय-संचालित अपडेट की अनुमति देता है।
Supports the transmission of low-latency audio signals across IP networks.
This project is a professional live video production suite designed for capturing, encoding, and broadcasting high-quality media. At its core, it features a real-time media processing engine that utilizes hardware acceleration to composite multiple audio and video sources with minimal latency. The application provides a centralized studio interface for managing complex scene transitions, layering visual sources through a hierarchical scene-graph engine, and streaming content to multiple platforms simultaneously. The software is built on a cross-platform abstraction layer that ensures consiste
Processes multiple audio and video inputs into a single, synchronized output stream in real time.
FFmpeg is a cross-platform multimedia framework designed for the recording, conversion, and streaming of audio and video content. It functions as a comprehensive toolkit that provides both a command-line utility for direct media manipulation and a collection of low-level libraries for integration into custom applications. At its core, the project utilizes a packet-based stream engine and a format-agnostic abstraction layer to handle diverse media standards, containers, and network protocols. The framework distinguishes itself through a modular, graph-based filter execution model that allows f
Records audio and video directly from hardware devices, screen displays, or network streams for immediate processing or storage.
VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow
Enables streaming audio inference for real-time delivery of synthesized speech in interactive applications.
Kazumi is a cross-platform media player and streaming platform that centralizes video content from diverse third-party web sources. It functions as an automated scraping tool, utilizing configurable path patterns and selectors to extract and aggregate media streams into a unified interface. The platform distinguishes itself through its focus on synchronized group viewing and real-time state management. Users can participate in shared virtual rooms where playback progress and controls are aligned across multiple devices. Additionally, the application includes integrated image processing capabi
The application applies real-time image processing to video streams to improve visual clarity and detail during playback for a better viewing experience.
Tasmota is a universal firmware platform for ESP8266 and ESP32 microcontrollers, designed to provide local control and management of smart home hardware. It functions as an event-driven automation controller that replaces proprietary factory firmware, allowing users to manage relays, sensors, and lighting systems without relying on external cloud services. The system is built on a modular driver architecture that enables dynamic hardware configuration and peripheral support through a web-based management interface. The platform distinguishes itself through a template-driven hardware mapping s
Transmits real-time audio between devices over UDP to create intercom systems.
Spotify-downloader is a command-line utility designed to archive music from Spotify by matching track URLs to external video sources. It functions as a high-fidelity downloader that retrieves audio content and saves it as local files, ensuring optimal sound quality by selecting the highest available bitrate from the source media. The tool distinguishes itself through its ability to maintain local music collections by mirroring remote playlist states. It performs local-remote synchronization to determine which tracks require downloading or removal, while utilizing a modular architecture to dec
"Manages multiple simultaneous download and conversion tasks using asynchronous workers to maximize throughput and reduce total processing time."
Rembg is a machine learning-based toolkit designed for automated image background removal and subject segmentation. It functions as a versatile engine that identifies and extracts subjects from images, supporting diverse input methods including individual files, directory-based batch processing, and live binary data streams. The project distinguishes itself through its flexible integration options, offering a command-line interface for local automation, a library for programmatic access, and an HTTP service for remote requests. It utilizes deep learning architectures to classify pixels and ge
Supports real-time processing of binary pixel data streams for dynamic background removal.
This software is a real-time voice changer that utilizes machine learning inference to transform live microphone input into target vocal characteristics. It functions as an artificial intelligence audio processing tool designed to modify vocal identity during active communication or live broadcasts. The application distinguishes itself by executing neural network models directly within the browser environment. It leverages web-based compute acceleration and dedicated audio threading to maintain low-latency performance, allowing users to switch between different voice profiles while processing
Injects processed audio directly into communication platforms by replacing standard microphone input streams.
MediaMTX is a multi-protocol media server designed for routing, proxying, and recording real-time audio and video streams. It functions as a programmable media router and a gateway between streaming standards such as RTMP, RTSP, and WebRTC, enabling the conversion of live media between different protocols. The server distinguishes itself through on-the-fly format transmuxing and protocol-agnostic routing, which decouples input and output protocols via an internal media bus. It features a programmable automation system that executes external shell commands via event hooks triggered by client c
Saves incoming live streams to disk using specific container formats for later playback.
Res-downloader is a network proxy utility designed to intercept, analyze, and extract multimedia assets from web traffic. It functions as a gateway that captures video, audio, and image files directly from data streams for local storage and offline access. The tool employs man-in-the-middle interception to decrypt and inspect network packets, allowing it to identify media resources through pattern matching and content type filtering. It integrates proxy-based routing to manage outgoing requests, enabling the retrieval of content that may be subject to regional restrictions or network-level ac
Processes large media streams asynchronously to ensure efficient file reconstruction without blocking the main execution flow.
Cat-catch is a browser-based media utility designed to detect, capture, and manage web-based video and audio resources. It functions as a comprehensive sniffing and download management system, enabling users to identify hidden or protected media assets directly from active web pages. The tool specializes in reconstructing fragmented streaming protocols, such as DASH and M3U8, into complete files while providing options for real-time stream recording and playback control. The project distinguishes itself through its deep integration with local system environments and external automation tools.
Provides frameworks for capturing, encoding, and processing audio and video data streams in real time.
ffmpeg.wasm is a browser-based multimedia processing engine that brings the capabilities of the FFmpeg library directly to the client environment. By utilizing WebAssembly, it enables audio and video transcoding, format conversion, and stream recording to occur entirely within the browser without requiring server-side infrastructure. The library distinguishes itself by executing resource-intensive media tasks in background threads, ensuring that the main user interface remains responsive during complex operations. It manages data through an isolated, in-memory virtual file system, allowing fo
Captures audio and video input from browser sources and encodes the data into files for local storage.
This project is a cross-platform implementation of the WebRTC standard, providing a comprehensive library for building real-time audio, video, and data communication applications. It functions as a peer-to-peer networking framework and media processing engine, enabling direct, low-latency connections between devices without relying on central servers. By strictly adhering to official protocol specifications, the library ensures interoperability with browsers and other native communication software across mobile, desktop, and server environments. The engine distinguishes itself through a modul
Captures and transmits local audio and video content across peer connections using portable interfaces for consistent media handling.
ImageMagick is a comprehensive software suite for the creation, editing, composition, and conversion of digital images. It functions as both a command-line utility for batch processing and automation, and as a programming library that allows developers to integrate advanced image manipulation capabilities into external applications. The project is distinguished by its modular architecture, which supports hundreds of image formats through a pluggable coder system and external delegate libraries. It is designed for high-performance environments, utilizing memory-mapped pixel caching, stream-ori
Improves local detail and edge definition using contrast-limited adaptive histogram equalization.
Vercel is a cloud platform for building, deploying, and scaling web applications. It provides a unified infrastructure that automates the build process by detecting project frameworks and distributing static and dynamic content through a global content delivery network. The platform executes application logic using serverless functions that scale automatically based on real-time traffic demand. The platform distinguishes itself through a centralized AI gateway that proxies requests to multiple model providers, enabling standardized authentication, observability, and cost tracking. It supports
Enables real-time processing of visual data streams for interactive applications.
This repository provides a collection of reference implementations and practical demonstrations for using WebRTC to establish real-time audio, video, and data communication. It contains code samples for negotiating peer-to-peer connections, managing media streams, and utilizing low-latency data channels. The project demonstrates the capture of audio and video from hardware devices, as well as the redirection of canvas element content into media streams. It includes examples of transferring arbitrary text and binary data between peers and managing the negotiation of direct connections. The sa
Demonstrates routing raw media frames through processing layers before transmission or playback.
MNN is a high-performance inference engine and framework designed for on-device machine learning. It provides a comprehensive environment for executing, optimizing, and deploying neural network models directly on mobile and resource-constrained edge devices. The framework distinguishes itself through a robust model optimization toolkit that supports quantization, compression, and structural graph manipulation to minimize memory footprint and maximize execution speed. It features a modular architecture that abstracts hardware-specific backends, allowing models to run efficiently across diverse
Executes lightweight image processing and codec functions to prepare visual data for high-performance inference pipelines.
The NGINX RTMP module is a server-side extension that functions as a live video streaming engine. It enables the ingestion, processing, and distribution of real-time audio and video feeds, supporting both RTMP and HLS protocols to facilitate media delivery to multiple clients. The module distinguishes itself by integrating directly into the host server event loop, allowing for high-concurrency network input and output without blocking the main thread. It provides a toolkit for managing media streams through event-driven callbacks, which can trigger external process invocations for custom tran
Processes incoming media packets through event-driven callbacks triggered by stream lifecycle events.
Shotcut is a professional-grade, cross-platform non-linear video editor built on the MLT multimedia framework. It provides a comprehensive suite for post-production, supporting multi-track timeline editing, high-fidelity color processing, and complex visual effects. The application is designed to handle diverse audio and video formats natively, ensuring high-resolution and HDR workflows are managed within a unified environment. The software distinguishes itself through a modular architecture that emphasizes performance and precision. It utilizes a GPU-accelerated rendering pipeline and proxy-
Captures video and audio from hardware devices, network streams, and screen inputs for direct integration into editing workflows.