awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

7 रिपॉजिटरी

Awesome GitHub RepositoriesAudio Processing Frameworks

Development environments and libraries that provide infrastructure for building complex neural-based audio processing pipelines.

Explore 7 awesome GitHub repositories matching graphics & multimedia · Audio Processing Frameworks. Refine with filters or upvote what's useful.

Awesome Audio Processing Frameworks GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • rvc-boss/gpt-sovitsRVC-Boss का अवतार

    RVC-Boss/GPT-SoVITS

    58,724GitHub पर देखें↗

    GPT-SoVITS is a text-to-speech synthesis engine and voice cloning toolkit designed for generating natural-sounding human speech. It functions as a neural audio processing pipeline that maps input text to high-fidelity audio waveforms, utilizing conditional variational autoencoders and flow-based decoders to ensure expressive output. The platform distinguishes itself through its ability to perform few-shot voice cloning and cross-lingual speech generation, allowing users to maintain a specific speaker's vocal identity and emotional delivery across multiple languages. By employing cross-modal l

    Facilitates an end-to-end workflow for training, fine-tuning, and deploying custom voice models.

    Pythontext-to-speechttsvits
    GitHub पर देखें↗58,724
  • rvc-project/retrieval-based-voice-conversion-webuiRVC-Project का अवतार

    RVC-Project/Retrieval-based-Voice-Conversion-WebUI

    36,025GitHub पर देखें↗

    This project is a comprehensive software suite for voice synthesis and model management, providing a framework for training custom acoustic models and performing voice conversion. It utilizes deep-learning-based acoustic modeling to map source audio characteristics to target voice identities, enabling the transformation of input audio into specific vocal profiles. The system distinguishes itself through a feature-retrieval-based inference mechanism, which employs vector index files to perform nearest-neighbor searches on acoustic features for high-fidelity timbre matching. Users can manage th

    Chains discrete stages like pitch extraction and source separation into a modular audio processing pipeline.

    Pythonaudio-analysischangeconversational-ai
    GitHub पर देखें↗36,025
  • audiokit/audiokitaudiokit का अवतार

    audiokit/AudioKit

    11,381GitHub पर देखें↗

    AudioKit is an audio framework for iOS, macOS, and tvOS that provides tools for digital audio synthesis, signal processing, and audio analysis. It functions as a synthesis engine for generating audio waveforms and textures, a processing library for modifying tonal characteristics, and a toolkit for extracting frequency and amplitude data from sonic signals. The framework utilizes a modular node architecture and graph-based signal routing to connect audio generators, processors, and outputs. It wraps low-level audio primitives in high-level classes to facilitate sound generation and modificati

    Acts as a comprehensive infrastructure for building complex audio processing and synthesis pipelines.

    Swift
    GitHub पर देखें↗11,381
  • aigc-audio/audiogptAIGC-Audio का अवतार

    AIGC-Audio/AudioGPT

    10,174GitHub पर देखें↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Uses large language models to orchestrate neural audio pipelines for generation and processing tasks.

    Pythonaudiogptmusic
    GitHub पर देखें↗10,174
  • uberi/speech_recognitionUberi का अवतार

    Uberi/speech_recognition

    8,973GitHub पर देखें↗

    This project is a Python speech recognition library that serves as a unified interface for converting spoken audio into text. It functions as a bridge between Python applications and a variety of speech-to-text engines, providing a consistent way to interact with both local and cloud-based recognition services. The library distinguishes itself as a multi-engine transcription tool, wrapping diverse online APIs and offline recognition backends into a standardized format. This allows for interchangeable recognition engines and supports multilingual audio transcription through various language pa

    Implements a framework for capturing microphone input and managing audio file formats for transcription.

    Pythonaudiopythonspeech-recognition
    GitHub पर देखें↗8,973
  • whisperspeech/whisperspeechWhisperSpeech का अवतार

    WhisperSpeech/WhisperSpeech

    4,617GitHub पर देखें↗

    WhisperSpeech एक बहुभाषी स्पीच सिंथेसाइज़र और न्यूरल टेक्स्ट-टू-स्पीच सिस्टम है। यह टेक्स्ट को हाई-फिडेलिटी सिंथेटिक ऑडियो में बदलने के लिए Whisper मॉडल आर्किटेक्चर को इनवर्ट करके कार्य करता है। यह सिस्टम विशिष्ट वक्ताओं की नकल करने के लिए संदर्भ ऑडियो फ़ाइलों का उपयोग करके वॉयस क्लोनिंग को सक्षम बनाता है। यह बहुभाषी स्पीच प्रोडक्शन का समर्थन करता है, जिसमें विभिन्न भाषाओं में ऑडियो उत्पन्न करने और एक ही वाक्य के भीतर भाषा स्विचिंग को संभालने की क्षमता शामिल है। यह प्रोजेक्ट टेक्स्ट-टू-स्पीच जनरेशन और स्पीच डेटासेट तैयारी सहित स्पीच क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। इसमें स्पीच को टेक्स्ट में ट्रांसक्राइब करने, ध्वनिक टोकन निकालने, और वॉयस एक्टिविटी का पता लगाने के लिए टूल्स शामिल हैं।

    Implements an end-to-end neural pipeline using semantic and acoustic tokens to generate high-fidelity synthetic speech.

    Jupyter Notebookpytorchspeech-synthesistts
    GitHub पर देखें↗4,617
  • allendowney/thinkdspAllenDowney का अवतार

    AllenDowney/ThinkDSP

    4,567GitHub पर देखें↗

    ThinkDSP is a Python-based audio signal processing framework and educational resource designed for studying the mathematical properties of digital audio and waveforms. It functions as a digital signal processing library that provides tools for performing frequency analysis and harmonic decomposition of sound waves. The project covers the fundamentals of audio frequency analysis and sound synthesis, enabling the decomposition of sound into harmonics to analyze or modify spectral content. It facilitates Python audio programming by providing the means to manipulate audio files and generate synth

    Functions as a Python-based environment for studying and manipulating the mathematical properties of digital audio.

    Jupyter Notebook
    GitHub पर देखें↗4,567
  1. Home
  2. Graphics & Multimedia
  3. Media Processing and Analysis
  4. Audio Processing Systems
  5. Audio Processing Frameworks

सब-टैग एक्सप्लोर करें

  • LLM-Driven OrchestrationAudio frameworks that use large language models to decompose goals and select appropriate neural audio tools. **Distinct from Audio Processing Frameworks:** Focuses on the LLM-based control plane for tool selection, rather than just the underlying processing infrastructure.
  • Neural Audio PipelinesEnd-to-end systems for training and deploying voice models.