awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

7 repository-uri

Awesome GitHub RepositoriesSpeech Synthesis Customizations

Tools for adjusting voice parameters such as timbre, speed, language, and gender in text-to-speech systems.

Distinct from Speech Synthesis Caches: Existing candidates focus on caching, gateways, or specific languages rather than the parameterization of voice attributes.

Explore 7 awesome GitHub repositories matching artificial intelligence & ml · Speech Synthesis Customizations. Refine with filters or upvote what's useful.

Awesome Speech Synthesis Customizations GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • swivid/f5-ttsAvatar SWivid

    SWivid/F5-TTS

    14,798Vezi pe GitHub↗

    F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s

    Modifies the playback rate of generated audio by specifying a target duration for the output.

    Python
    Vezi pe GitHub↗14,798
  • idootop/mi-gptAvatar idootop

    idootop/mi-gpt

    12,458Vezi pe GitHub↗

    mi-gpt is a voice assistant bridge and agent orchestrator that connects smart speakers to large language models. It functions as an integration layer that routes audio requests from hardware speakers to AI providers and converts generated text back into speech via a customizable synthesis system. The project features a retrieval-augmented generation knowledge base that uses embeddings and external documents to provide context-aware responses. It includes a persona definition system for configuring behavioral rules, system prompts, and roleplay characteristics, alongside a plugin architecture

    Allows replacing default system voices with high-quality models to customize the assistant's auditory profile.

    TypeScript
    Vezi pe GitHub↗12,458
  • stevenjoezhang/live2d-widgetAvatar stevenjoezhang

    stevenjoezhang/live2d-widget

    10,764Vezi pe GitHub↗

    This project provides an animated Live2D character widget that can be embedded on any web page as an interactive mascot. The widget renders characters using the Cubism SDK on an HTML canvas, and can be deployed either via a content delivery network for zero-setup integration or self-hosted on a personal server for full control over asset delivery. The mascot responds to visitor actions through CSS selector-based interaction binding, displaying custom speech bubbles when users hover over or click specific page elements. Visitors can click and drag the character to reposition it anywhere on the

    Defines custom text messages that appear when visitors interact with specific page elements.

    JavaScriptjavascript-pluginlive2d
    Vezi pe GitHub↗10,764
  • santinic/audiblezAvatar santinic

    santinic/audiblez

    7,811Vezi pe GitHub↗

    Audiblez is a text-to-speech audiobook generator that converts digital e-books into spoken audio files. The system processes written documents using speech synthesis and configurable voice profiles to produce audiobooks. The tool utilizes a graphical interface to manage the conversion workflow and task orchestration. It employs CUDA-accelerated processing to offload neural network computations to the GPU, increasing the speed of audio generation. The system includes capabilities for chapter-based file parsing and selective chapter conversion. Users can adjust synthesis parameters, including

    Provides tools for adjusting voice parameters including speed, language, and voice identity.

    Pythonaudiobooksepubkokoro
    Vezi pe GitHub↗7,811
  • yils-lin/short-video-factoryAvatar YILS-LIN

    YILS-LIN/short-video-factory

    3,428Vezi pe GitHub↗

    Short video factory is a local AI content generator and automated video editing tool. It provides a production pipeline that uses large language models to transform text prompts into marketing scripts and rendered short-form videos. The system is designed for local-first execution, running all processing and asset management on the host machine to maintain data privacy. It distinguishes itself through a batch-processing workflow that can sequentially execute copywriting and rendering for multiple items using predefined presets. The software covers a broad range of media capabilities, includi

    The software allows users to select language, gender, voice timbre, and speaking speed to customize audio narration.

    TypeScriptaiautomaticautomation
    Vezi pe GitHub↗3,428
  • rsxdalv/tts-webuiAvatar rsxdalv

    rsxdalv/TTS-WebUI

    2,980Vezi pe GitHub↗

    TTS-WebUI is a web interface and speech synthesis manager designed to convert written text into spoken audio files. It serves as a self-hosted audio AI suite that allows users to configure speech synthesis models, manage speaker profiles, and generate audio through a graphical dashboard. The system functions as both a visual manager and a generative audio API, providing standardized endpoints and OpenAI-compatible request formats for external applications to trigger synthesis programmatically. It includes a plugin-based extension system that allows new tools and models to be added via externa

    Features a dashboard for adjusting voice parameters, speaker profiles, and managing long-form text synthesis.

    TypeScriptace-stepaiaudio-generation
    Vezi pe GitHub↗2,980
  • voice-cloning-app/voice-cloning-appAvatar voice-cloning-app

    voice-cloning-app/Voice-Cloning-App

    1,438Vezi pe GitHub↗

    Această aplicație este o platformă pentru sinteza vocală AI și clonarea vocală neuronală. Oferă un toolkit cuprinzător pentru convertirea textului în vorbire umană cu sunet natural prin aplicarea unor modele de rețele neuronale antrenate personalizat pe mostre audio specifice. Sistemul facilitează întregul ciclu de viață al dezvoltării modelelor vocale, inclusiv pregătirea cărților audio brute și a transcrierilor video în seturi de date de antrenare structurate. Suportă antrenarea acestor modele pe hardware local sau remote, utilizând procesarea distribuită multi-GPU pentru a gestiona date la scară largă și a accelera convergența modelului. Dincolo de antrenare, platforma include capabilități pentru gestionarea și portarea seturilor de date vocale între diferite medii de stocare. Utilizatorii pot efectua inferența prin ajustarea variabilelor latente și a parametrilor de sinteză pentru a modifica prozodia, inflexiunea emoțională și calitățile stilistice ale output-ului audio generat. Aplicația se bazează pe tehnici de deep learning pentru a transforma reprezentările acustice în forme de undă de înaltă fidelitate.

    Facilitates the management of voice datasets and the configuration of synthesis parameters to enable high-fidelity neural voice cloning.

    Pythondeep-learningpythonpytorch
    Vezi pe GitHub↗1,438
  1. Home
  2. Artificial Intelligence & ML
  3. Speech Synthesis Customizations

Explorează sub-etichetele

  • Mascot Speech CustomizationsCustom text messages that appear when a visitor interacts with specific page elements via CSS selectors. **Distinct from Speech Synthesis Customizations:** Distinct from Speech Synthesis Customizations: focuses on text-based speech bubbles triggered by CSS selectors, not voice parameters.
  • Voice Cloning ToolkitsSoftware suites for managing voice datasets and configuring neural network parameters to replicate specific human vocal characteristics. **Distinct from Speech Synthesis Customizations:** Distinct from Speech Synthesis Customizations: focuses on the end-to-end lifecycle of cloning a specific voice from audio samples, rather than just adjusting parameters of pre-existing models.