awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

159 रिपॉजिटरी

Awesome GitHub RepositoriesAudio Processing Systems

Software systems for generating, synthesizing, and processing digital audio signals.

Explore 159 awesome GitHub repositories matching graphics & multimedia · Audio Processing Systems. Refine with filters or upvote what's useful.

Awesome Audio Processing Systems GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • binary-husky/gpt_academicbinary-husky का अवतार

    binary-husky/gpt_academic

    70,912GitHub पर देखें↗

    This project provides a self-hosted, web-based interface designed to integrate large language models into academic and research workflows. It functions as a modular platform for document analysis, literature processing, and data handling, allowing users to maintain full control over their data and model connectivity through private server or local deployments. The system is distinguished by its extensible architecture, which enables users to inject custom Python scripts to automate repetitive tasks and extend core functionality. It also features a voice-enabled interaction layer that captures

    Capture and route audio streams through automated pipelines to transform spoken input into text commands.

    Pythonacademicchatglm-6bchatgpt
    GitHub पर देखें↗70,912
  • corentinj/real-time-voice-cloningCorentinJ का अवतार

    CorentinJ/Real-Time-Voice-Cloning

    59,918GitHub पर देखें↗

    This project is a neural text-to-speech engine and voice cloning toolkit designed to generate synthetic speech that mimics the vocal characteristics of a target speaker. It functions as a real-time audio synthesizer, utilizing a deep learning pipeline to convert written text into high-fidelity speech output with minimal latency. The system employs a transfer learning framework that leverages pre-trained speaker verification models to adapt synthesis to new, unseen vocal identities. By using an encoder-based speaker embedding process, the toolkit maps variable-length audio samples into a laten

    Synthesizes high-fidelity audio waveforms from spectral representations using models optimized for rapid inference.

    Pythondeep-learningpythonpytorch
    GitHub पर देखें↗59,918
  • rvc-boss/gpt-sovitsRVC-Boss का अवतार

    RVC-Boss/GPT-SoVITS

    58,724GitHub पर देखें↗

    GPT-SoVITS is a text-to-speech synthesis engine and voice cloning toolkit designed for generating natural-sounding human speech. It functions as a neural audio processing pipeline that maps input text to high-fidelity audio waveforms, utilizing conditional variational autoencoders and flow-based decoders to ensure expressive output. The platform distinguishes itself through its ability to perform few-shot voice cloning and cross-lingual speech generation, allowing users to maintain a specific speaker's vocal identity and emotional delivery across multiple languages. By employing cross-modal l

    Converts written text into natural-sounding human speech via an integrated neural audio synthesis engine.

    Pythontext-to-speechttsvits
    GitHub पर देखें↗58,724
  • microsoft/vibevoicemicrosoft का अवतार

    microsoft/VibeVoice

    49,394GitHub पर देखें↗

    VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow

    Converts high-level acoustic features into raw waveform audio using deep generative neural vocoders.

    Python
    GitHub पर देखें↗49,394
  • paul-gauthier/aiderpaul-gauthier का अवतार

    paul-gauthier/aider

    46,354GitHub पर देखें↗

    Aider is a terminal-based AI coding assistant and pair programmer that uses large language models to write, edit, and refactor source code across multiple files and programming languages. It functions as a command line interface for automating programming tasks and managing codebase modifications. The tool distinguishes itself by creating structural maps of entire codebases to provide language models with the necessary context for navigating and modifying large repositories. It further expands input capabilities through a speech-to-text pipeline for voice-driven development and multi-modal in

    Processes audio input through a transcription engine to convert spoken requests into text instructions.

    Python
    GitHub पर देखें↗46,354
  • coqui-ai/ttscoqui-ai का अवतार

    coqui-ai/TTS

    45,568GitHub पर देखें↗

    यह प्रोजेक्ट एक डीप लर्निंग टेक्स्ट-टू-स्पीच टूलकिट है जिसका उपयोग न्यूरल स्पीच सिंथेसिस मॉडल को प्रशिक्षित और तैनात करने के लिए किया जाता है। यह लिखित टेक्स्ट को बोले गए ऑडियो में बदलने के लिए एक व्यापक फ्रेमवर्क प्रदान करता है। टूलकिट में एक वॉयस क्लोनिंग सिस्टम शामिल है जो छोटे ऑडियो नमूनों से स्पीकर एम्बेडिंग निकालकर विशिष्ट मानव आवाजों की नकल करता है। यह मल्टी-स्पीकर ऑडियो सिंथेसिस का भी समर्थन करता है। सिस्टम में स्पीच डेटासेट क्यूरेशन, प्रदर्शन ट्रैकिंग के साथ कस्टम मॉडल प्रशिक्षण और ऑडियो जनरेशन के लिए कमांड-लाइन इंटरफेस सहित पूर्ण स्पीच सिंथेसिस पाइपलाइन शामिल है।

    Includes neural vocoders that transform synthesized spectrograms into high-fidelity time-domain audio waveforms.

    Pythondeep-learningglow-ttshifigan
    GitHub पर देखें↗45,568
  • babysor/mockingbirdbabysor का अवतार

    babysor/MockingBird

    36,903GitHub पर देखें↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    Generates arbitrary speech conditioned on a specific audio sample to maintain voice identity.

    Pythonaideep-learningpytorch
    GitHub पर देखें↗36,903
  • rvc-project/retrieval-based-voice-conversion-webuiRVC-Project का अवतार

    RVC-Project/Retrieval-based-Voice-Conversion-WebUI

    36,025GitHub पर देखें↗

    This project is a comprehensive software suite for voice synthesis and model management, providing a framework for training custom acoustic models and performing voice conversion. It utilizes deep-learning-based acoustic modeling to map source audio characteristics to target voice identities, enabling the transformation of input audio into specific vocal profiles. The system distinguishes itself through a feature-retrieval-based inference mechanism, which employs vector index files to perform nearest-neighbor searches on acoustic features for high-fidelity timbre matching. Users can manage th

    Synthesizes vocal characteristics by applying learned voice models to input audio sources.

    Pythonaudio-analysischangeconversational-ai
    GitHub पर देखें↗36,025
  • hugohe3/ppt-masterhugohe3 का अवतार

    hugohe3/ppt-master

    30,561GitHub पर देखें↗

    ppt-master is an AI PowerPoint generator and LLM presentation orchestrator that converts documents and text into editable presentation files. It utilizes native shapes and structured layouts to transform content into professional slide decks. The system functions as a template processor capable of injecting generated content into existing PowerPoint files while preserving original brand designs and formatting. It integrates AI imagery services to generate custom visuals and retrieves professional stock photography with attribution management. The project covers automated slide deck creation

    Generates audio narration files from speaker notes and embeds them as synchronized media objects within the slides.

    Pythonai-agentaipptoffice
    GitHub पर देखें↗30,561
  • jamiepine/voiceboxjamiepine का अवतार

    jamiepine/voicebox

    30,041GitHub पर देखें↗

    Voicebox is a local speech processing system that provides text-to-speech generation, speech-to-text transcription, and voice cloning. It utilizes local machine learning inference and GPU acceleration to process audio and text data without relying on external API calls. The project features a voice cloning toolkit for creating synthetic profiles from audio samples and a timeline-based voice editor for composing multi-character conversations. It also includes an AI voice management API that allows external applications and AI agents to programmatically manage voice profiles and generate speech

    Routes synthetic speech through a sequential processing chain of pitch shifts and reverb effects.

    TypeScriptaicudamlx
    GitHub पर देखें↗30,041
  • ankitects/ankiankitects का अवतार

    ankitects/anki

    28,571GitHub पर देखें↗

    Anki is a cross-platform flashcard management system designed to optimize long-term memory retention through spaced-repetition learning. It functions as a digital learning assistant that uses active recall practice and automated scheduling algorithms to determine the ideal timing for card reviews based on individual performance history. The core system relies on a local relational database to ensure data persistence and portability, while supporting complex study workflows through flexible note-type schema modeling and template-driven content rendering. The platform distinguishes itself throu

    Synthesizes speech from text fields using system-provided voices with configurable settings.

    Rust
    GitHub पर देखें↗28,571
  • svc-develop-team/so-vits-svcsvc-develop-team का अवतार

    svc-develop-team/so-vits-svc

    28,097GitHub पर देखें↗

    This project is a singing voice conversion tool based on VITS generative modeling. It transforms the identity of a singing voice to a target speaker while preserving the original melody, lyrics, and intonation. The system distinguishes itself through hybrid voice synthesis, allowing for the blending of multiple speaker identities via linear model interpolation. It utilizes cluster-based feature retrieval to increase target voice similarity and employs a diffusion probabilistic model as a post-processor to remove electronic artifacts and improve vocal clarity. The software covers a broad rang

    Blends multiple speaker models to create hybrid voice identities through linear interpolation.

    Python
    GitHub पर देखें↗28,097
  • nextai-translator/nextai-translatornextai-translator का अवतार

    nextai-translator/nextai-translator

    24,920GitHub पर देखें↗

    Nextai-translator is an AI-powered text processor and cross-platform translation application. Available as a desktop app and browser extension, it uses large language model APIs to translate, summarize, and refine multilingual content in real time. The tool integrates with clipboard managers and text selection utilities to trigger automated translations immediately after content is copied or highlighted. It also functions as an OCR translation utility, extracting and translating text from screenshots and non-selectable image content. Additional capabilities include a vocabulary management sy

    Converts processed or translated text into spoken audio using a synthetic speech engine.

    TypeScriptbrowser-extensionchatgptchrome-extension
    GitHub पर देखें↗24,920
  • 78/xiaozhi-esp3278 का अवतार

    78/xiaozhi-esp32

    24,092GitHub पर देखें↗

    Xiaozhi-esp32 is an open-source firmware platform designed for building voice-interactive embedded systems on resource-constrained microcontrollers. It functions as an IoT conversational device platform that manages live audio input, speech synthesis, and conversational state transitions to facilitate real-time natural language interaction. The system distinguishes itself by bridging language models with physical hardware through standardized protocols, allowing for the execution of commands on local peripherals or remote smart home services. It utilizes a specialized architecture to coordina

    Specialized architecture for managing live voice input and speech synthesis on embedded hardware.

    C++chatbotesp32mcp
    GitHub पर देखें↗24,092
  • mozilla-ai/llamafilemozilla-ai का अवतार

    mozilla-ai/llamafile

    23,726GitHub पर देखें↗

    Llamafile is a machine learning model runner and packager that enables local inference by bundling model weights and runtime environments into a single, self-contained executable. It functions as a cross-platform engine, allowing users to execute large language models and perform speech-to-text tasks directly on their own hardware without requiring external software dependencies or complex installations. The project distinguishes itself by utilizing a specialized binary format that allows the same executable to run natively across multiple operating systems and hardware architectures. It auto

    Converts spoken language into written text using a standalone executable that functions across different operating systems.

    C
    GitHub पर देखें↗23,726
  • resemble-ai/chatterboxresemble-ai का अवतार

    resemble-ai/chatterbox

    22,751GitHub पर देखें↗

    Chatterbox is a comprehensive machine learning platform designed for multilingual speech synthesis and real-time audio generation. It functions as an engine that converts text into natural-sounding speech, capable of replicating specific human vocal characteristics and emotional expressions from short audio samples. The platform distinguishes itself through advanced control over the synthesis process, allowing for the manipulation of emotional intensity and the injection of non-verbal vocalizations such as laughter or coughing. It is engineered for low-latency performance, utilizing an optimi

    Transforms generated spectral data into high-fidelity time-domain audio waveforms.

    Python
    GitHub पर देखें↗22,751
  • vercel/aivercel का अवतार

    vercel/ai

    21,885GitHub पर देखें↗

    This project is a comprehensive framework for building AI-powered applications, providing a unified toolkit for orchestrating language models, autonomous agents, and interactive user interfaces. It serves as a central library for managing the entire lifecycle of AI interactions, from initial prompt generation and model provider abstraction to complex, multi-step reasoning and tool execution. The framework distinguishes itself through its deep integration with frontend development, specifically by enabling generative user interfaces that render dynamic components directly from model outputs. I

    Converts between spoken language and text to support voice-based AI interactions.

    TypeScriptanthropicartificial-intelligencegemini
    GitHub पर देखें↗21,885
  • funaudiollm/cosyvoiceFunAudioLLM का अवतार

    FunAudioLLM/CosyVoice

    21,673GitHub पर देखें↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Converts raw acoustic tokens into high-fidelity waveforms using deep learning models.

    Pythonaudio-generationcantonesechatbot
    GitHub पर देखें↗21,673
  • readest/readestreadest का अवतार

    readest/readest

    21,502GitHub पर देखें↗

    Readest is a comprehensive digital reading platform designed to manage, annotate, and consume electronic books across multiple devices. It functions as a versatile library manager and reading environment, supporting a wide range of user needs from standard ebook consumption to specialized study and accessibility-focused workflows. The platform distinguishes itself through advanced features like parallel text study, which enables side-by-side document rendering with synchronized scrolling, and a robust text-to-speech engine that provides hands-free reading with synchronized visual highlighting

    Converts written content into spoken audio using local system engines or cloud-based voices for hands-free reading.

    TypeScriptandroidcross-platformebook
    GitHub पर देखें↗21,502
  • magenta/magentamagenta का अवतार

    magenta/magenta

    19,778GitHub पर देखें↗

    Magenta is a comprehensive toolkit for training, synthesizing, and performing music through neural models and hardware-integrated engines. It functions as a machine learning framework that enables the generation, manipulation, and real-time performance of audio, providing the structural foundations for musical intelligence through hierarchical sequence modeling and symbolic processing. The project distinguishes itself by enabling real-time, low-latency neural audio synthesis that can be integrated directly into professional digital audio workstations. It supports interactive musical jamming a

    Transforms learned feature representations into raw waveforms using deep learning models.

    Python
    GitHub पर देखें↗19,778
पिछला123456…8अगला
  1. Home
  2. Graphics & Multimedia
  3. Media Processing and Analysis
  4. Audio Processing Systems

सब-टैग एक्सप्लोर करें

  • Audio Processing3 सब-टैग्सTools for converting between spoken language and text using automated speech recognition and synthesis engines.
  • Audio Processing Frameworks2 सब-टैग्सDevelopment environments and libraries that provide infrastructure for building complex neural-based audio processing pipelines.
  • Audio Synthesis10 सब-टैग्सSystems that generate artificial audio signals, including advanced neural vocoders for voice and sound synthesis.