awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

116 مستودعات

Awesome GitHub RepositoriesAudio Processing

Tools for converting between spoken language and text using automated speech recognition and synthesis engines.

Explore 116 awesome GitHub repositories matching graphics & multimedia · Audio Processing. Refine with filters or upvote what's useful.

Awesome Audio Processing GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • binary-husky/gpt_academicالصورة الرمزية لـ binary-husky

    binary-husky/gpt_academic

    70,912عرض على GitHub↗

    This project provides a self-hosted, web-based interface designed to integrate large language models into academic and research workflows. It functions as a modular platform for document analysis, literature processing, and data handling, allowing users to maintain full control over their data and model connectivity through private server or local deployments. The system is distinguished by its extensible architecture, which enables users to inject custom Python scripts to automate repetitive tasks and extend core functionality. It also features a voice-enabled interaction layer that captures

    Capture and route audio streams through automated pipelines to transform spoken input into text commands.

    Pythonacademicchatglm-6bchatgpt
    عرض على GitHub↗70,912
  • corentinj/real-time-voice-cloningالصورة الرمزية لـ CorentinJ

    CorentinJ/Real-Time-Voice-Cloning

    59,918عرض على GitHub↗

    This project is a neural text-to-speech engine and voice cloning toolkit designed to generate synthetic speech that mimics the vocal characteristics of a target speaker. It functions as a real-time audio synthesizer, utilizing a deep learning pipeline to convert written text into high-fidelity speech output with minimal latency. The system employs a transfer learning framework that leverages pre-trained speaker verification models to adapt synthesis to new, unseen vocal identities. By using an encoder-based speaker embedding process, the toolkit maps variable-length audio samples into a laten

    Converts written text into fluent, human-like speech using a high-performance neural processing pipeline.

    Pythondeep-learningpythonpytorch
    عرض على GitHub↗59,918
  • rvc-boss/gpt-sovitsالصورة الرمزية لـ RVC-Boss

    RVC-Boss/GPT-SoVITS

    58,724عرض على GitHub↗

    GPT-SoVITS is a text-to-speech synthesis engine and voice cloning toolkit designed for generating natural-sounding human speech. It functions as a neural audio processing pipeline that maps input text to high-fidelity audio waveforms, utilizing conditional variational autoencoders and flow-based decoders to ensure expressive output. The platform distinguishes itself through its ability to perform few-shot voice cloning and cross-lingual speech generation, allowing users to maintain a specific speaker's vocal identity and emotional delivery across multiple languages. By employing cross-modal l

    Converts written text into natural-sounding human speech via an integrated neural audio synthesis engine.

    Pythontext-to-speechttsvits
    عرض على GitHub↗58,724
  • paul-gauthier/aiderالصورة الرمزية لـ paul-gauthier

    paul-gauthier/aider

    46,354عرض على GitHub↗

    Aider is a terminal-based AI coding assistant and pair programmer that uses large language models to write, edit, and refactor source code across multiple files and programming languages. It functions as a command line interface for automating programming tasks and managing codebase modifications. The tool distinguishes itself by creating structural maps of entire codebases to provide language models with the necessary context for navigating and modifying large repositories. It further expands input capabilities through a speech-to-text pipeline for voice-driven development and multi-modal in

    Processes audio input through a transcription engine to convert spoken requests into text instructions.

    Python
    عرض على GitHub↗46,354
  • babysor/mockingbirdالصورة الرمزية لـ babysor

    babysor/MockingBird

    36,903عرض على GitHub↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    Includes a processing pipeline to generate synthetic audio files via a command-line interface.

    Pythonaideep-learningpytorch
    عرض على GitHub↗36,903
  • hugohe3/ppt-masterالصورة الرمزية لـ hugohe3

    hugohe3/ppt-master

    30,561عرض على GitHub↗

    ppt-master is an AI PowerPoint generator and LLM presentation orchestrator that converts documents and text into editable presentation files. It utilizes native shapes and structured layouts to transform content into professional slide decks. The system functions as a template processor capable of injecting generated content into existing PowerPoint files while preserving original brand designs and formatting. It integrates AI imagery services to generate custom visuals and retrieves professional stock photography with attribution management. The project covers automated slide deck creation

    Generates audio narration files from speaker notes and embeds them as synchronized media objects within the slides.

    Pythonai-agentaipptoffice
    عرض على GitHub↗30,561
  • jamiepine/voiceboxالصورة الرمزية لـ jamiepine

    jamiepine/voicebox

    30,041عرض على GitHub↗

    Voicebox is a local speech processing system that provides text-to-speech generation, speech-to-text transcription, and voice cloning. It utilizes local machine learning inference and GPU acceleration to process audio and text data without relying on external API calls. The project features a voice cloning toolkit for creating synthetic profiles from audio samples and a timeline-based voice editor for composing multi-character conversations. It also includes an AI voice management API that allows external applications and AI agents to programmatically manage voice profiles and generate speech

    Routes synthetic speech through a sequential processing chain of pitch shifts and reverb effects.

    TypeScriptaicudamlx
    عرض على GitHub↗30,041
  • ankitects/ankiالصورة الرمزية لـ ankitects

    ankitects/anki

    28,571عرض على GitHub↗

    Anki is a cross-platform flashcard management system designed to optimize long-term memory retention through spaced-repetition learning. It functions as a digital learning assistant that uses active recall practice and automated scheduling algorithms to determine the ideal timing for card reviews based on individual performance history. The core system relies on a local relational database to ensure data persistence and portability, while supporting complex study workflows through flexible note-type schema modeling and template-driven content rendering. The platform distinguishes itself throu

    Synthesizes speech from text fields using system-provided voices with configurable settings.

    Rust
    عرض على GitHub↗28,571
  • nextai-translator/nextai-translatorالصورة الرمزية لـ nextai-translator

    nextai-translator/nextai-translator

    24,920عرض على GitHub↗

    Nextai-translator is an AI-powered text processor and cross-platform translation application. Available as a desktop app and browser extension, it uses large language model APIs to translate, summarize, and refine multilingual content in real time. The tool integrates with clipboard managers and text selection utilities to trigger automated translations immediately after content is copied or highlighted. It also functions as an OCR translation utility, extracting and translating text from screenshots and non-selectable image content. Additional capabilities include a vocabulary management sy

    Converts processed or translated text into spoken audio using a synthetic speech engine.

    TypeScriptbrowser-extensionchatgptchrome-extension
    عرض على GitHub↗24,920
  • mozilla-ai/llamafileالصورة الرمزية لـ mozilla-ai

    mozilla-ai/llamafile

    23,726عرض على GitHub↗

    Llamafile is a machine learning model runner and packager that enables local inference by bundling model weights and runtime environments into a single, self-contained executable. It functions as a cross-platform engine, allowing users to execute large language models and perform speech-to-text tasks directly on their own hardware without requiring external software dependencies or complex installations. The project distinguishes itself by utilizing a specialized binary format that allows the same executable to run natively across multiple operating systems and hardware architectures. It auto

    Converts speech into written text using a portable file that handles transcription across multiple operating systems.

    C
    عرض على GitHub↗23,726
  • vercel/aiالصورة الرمزية لـ vercel

    vercel/ai

    21,885عرض على GitHub↗

    This project is a comprehensive framework for building AI-powered applications, providing a unified toolkit for orchestrating language models, autonomous agents, and interactive user interfaces. It serves as a central library for managing the entire lifecycle of AI interactions, from initial prompt generation and model provider abstraction to complex, multi-step reasoning and tool execution. The framework distinguishes itself through its deep integration with frontend development, specifically by enabling generative user interfaces that render dynamic components directly from model outputs. I

    Converts between spoken language and text to support voice-based AI interactions.

    TypeScriptanthropicartificial-intelligencegemini
    عرض على GitHub↗21,885
  • readest/readestالصورة الرمزية لـ readest

    readest/readest

    21,502عرض على GitHub↗

    Readest is a comprehensive digital reading platform designed to manage, annotate, and consume electronic books across multiple devices. It functions as a versatile library manager and reading environment, supporting a wide range of user needs from standard ebook consumption to specialized study and accessibility-focused workflows. The platform distinguishes itself through advanced features like parallel text study, which enables side-by-side document rendering with synchronized scrolling, and a robust text-to-speech engine that provides hands-free reading with synchronized visual highlighting

    Converts written content into spoken audio using local system engines or cloud-based voices for hands-free reading.

    TypeScriptandroidcross-platformebook
    عرض على GitHub↗21,502
  • livekit/livekitالصورة الرمزية لـ livekit

    livekit/livekit

    19,358عرض على GitHub↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Processes audio by chaining speech-to-text transcription, language model generation, and text-to-speech synthesis to provide modular control over each stage of the conversation.

    Gogolangmedia-serversfu
    عرض على GitHub↗19,358
  • nari-labs/diaالصورة الرمزية لـ nari-labs

    nari-labs/dia

    19,324عرض على GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Injects realistic nonverbal vocal expressions into synthesized speech via text-based triggers.

    Pythonaiopen-weighttext-to-speech
    عرض على GitHub↗19,324
  • thu-maic/openmaicالصورة الرمزية لـ THU-MAIC

    THU-MAIC/OpenMAIC

    18,781عرض على GitHub↗

    OpenMAIC is an LLM multi-agent education platform designed to create immersive, interactive classroom simulations. It functions as a learning environment where multiple AI agents collaborate through a state-machine orchestration framework to coordinate conversational turns and interactions. The platform features an AI-driven interactive lesson generator that transforms documents and topics into educational experiences including slides, quizzes, and project activities. It integrates a speech-enabled interface that combines speech-to-text and text-to-speech for voice-based interaction, alongsid

    Facilitates verbal interaction with AI agents by processing spoken audio through integrated speech-to-text and text-to-speech services.

    TypeScript
    عرض على GitHub↗18,781
  • capsoftware/capالصورة الرمزية لـ CapSoftware

    CapSoftware/Cap

    17,026عرض على GitHub↗

    Cap is a self-hosted screen recording and video collaboration platform designed for teams to replace synchronous meetings with asynchronous video updates. It provides a comprehensive suite for capturing high-resolution desktop activity, including system audio, microphone input, and camera overlays, which are then processed through an integrated post-production workflow. The platform distinguishes itself by offering full data sovereignty through containerized deployment and object storage abstractions, allowing users to host their media assets on private infrastructure or S3-compatible buckets

    Automates the conversion of recorded audio into searchable text transcripts.

    TypeScriptappcapcoss
    عرض على GitHub↗17,026
  • vercel/vercelالصورة الرمزية لـ vercel

    vercel/vercel

    15,738عرض على GitHub↗

    Vercel is a cloud platform for building, deploying, and scaling web applications. It provides a unified infrastructure that automates the build process by detecting project frameworks and distributing static and dynamic content through a global content delivery network. The platform executes application logic using serverless functions that scale automatically based on real-time traffic demand. The platform distinguishes itself through a centralized AI gateway that proxies requests to multiple model providers, enabling standardized authentication, observability, and cost tracking. It supports

    Converts text into high-quality, natural-sounding audio for use in chatbots and interactive media applications.

    TypeScriptclicloudcommand
    عرض على GitHub↗15,738
  • kilo-org/kilocodeالصورة الرمزية لـ Kilo-Org

    Kilo-Org/kilocode

    15,616عرض على GitHub↗

    Kilocode is an autonomous engineering platform designed to orchestrate AI agents for complex software development tasks. It functions as a comprehensive system for automating coding, testing, and repository management by integrating directly with your codebase and terminal. The platform provides a unified gateway for model orchestration, allowing for the management of agentic workflows, event-driven automation, and persistent session state across distributed development environments. The platform distinguishes itself through its federated task management and policy-based access control, which

    Converts spoken audio into written text within prompt fields using real-time activity detection.

    TypeScriptaiai-ageai-coding
    عرض على GitHub↗15,616
  • nesquena/hermes-webuiالصورة الرمزية لـ nesquena

    nesquena/hermes-webui

    14,912عرض على GitHub↗

    Hermes-webui is a self-hosted AI orchestrator and web interface for managing autonomous agents. It serves as a multi-provider gateway that connects cloud and local large language models, providing a central hub to execute scheduled background jobs, run shell commands, and manage agent memory on private hardware. The system distinguishes itself through a persistent memory manager that utilizes knowledge graphs and markdown files for long-term context across sessions. It features a model context protocol host for extending agent capabilities with standardized tools and supports the orchestratio

    Integrates speech-to-text and text-to-speech engines to support real-time voice interactions.

    Pythonagentai-agentshermes
    عرض على GitHub↗14,912
  • rowboatlabs/rowboatالصورة الرمزية لـ rowboatlabs

    rowboatlabs/rowboat

    14,974عرض على GitHub↗

    Rowboat is an LLM orchestration platform and multimodal AI agent framework. It coordinates large language models with external tools, automated web monitoring, and local data vaults to execute actions and retrieve real-time information. The system operates as a local-first knowledge base, converting meeting notes and emails into a linked markdown knowledge graph. It functions as an automated market intelligence tool that tracks competitors and trends across the web to maintain updated information summaries. The platform covers a broad range of productivity and automation capabilities, includ

    Ships a processing chain that handles both text-to-speech and speech-to-text transformations.

    TypeScriptagentsagents-sdkai
    عرض على GitHub↗14,974
السابق12345…6التالي
  1. Home
  2. Graphics & Multimedia
  3. Media Processing and Analysis
  4. Audio Processing Systems
  5. Audio Processing

استكشف الوسوم الفرعية

  • Multilingual ProcessingCapabilities for processing audio across different languages and dialects. **Distinct from Audio Processing:** Distinct from Audio Processing: specifically focuses on linguistic diversity and dialect support rather than general signal manipulation.
  • Speech-to-Text Pipelines6 وسوم فرعيةAutomated workflows that convert spoken audio input into actionable text commands.
  • Text-to-Speech Engines5 وسوم فرعيةSoftware systems that convert written text into natural-sounding human speech using automated processing pipelines.