awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

11 repository-uri

Awesome GitHub RepositoriesProsody Controls

Mechanisms for adjusting the emotional tone, speed, and emphasis of synthesized speech output.

Distinct from Text-to-Speech: Distinct from general Text-to-Speech: focuses on real-time modulation of speech delivery parameters rather than the core synthesis engine.

Explore 11 awesome GitHub repositories matching artificial intelligence & ml · Prosody Controls. Refine with filters or upvote what's useful.

Awesome Prosody Controls GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • openbmb/voxcpmAvatar OpenBMB

    OpenBMB/VoxCPM

    29,985Vezi pe GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Infers prosody and expressiveness directly from text to produce natural, context-matched speech delivery.

    Pythonaudiodeeplearningminicpm
    Vezi pe GitHub↗29,985
  • openbmb/minicpm-vAvatar OpenBMB

    OpenBMB/MiniCPM-V

    25,653Vezi pe GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Modifies delivery speed and word emphasis to change the emotional impact of synthesized speech.

    Python
    Vezi pe GitHub↗25,653
  • openbmb/minicpm-oAvatar OpenBMB

    OpenBMB/MiniCPM-o

    23,850Vezi pe GitHub↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    Provides real-time control over speech delivery speed and emotional prosody during synthesis.

    Pythonminicpmminicpm-vmulti-modal
    Vezi pe GitHub↗23,850
  • funaudiollm/cosyvoiceAvatar FunAudioLLM

    FunAudioLLM/CosyVoice

    21,673Vezi pe GitHub↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Provides fine-grained control over vocal attributes like breathing, pacing, and volume to produce realistic, human-like speech output.

    Pythonaudio-generationcantonesechatbot
    Vezi pe GitHub↗21,673
  • nari-labs/diaAvatar nari-labs

    nari-labs/dia

    19,324Vezi pe GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Adjusts emotional tone and delivery parameters in synthesized speech using reference audio conditioning.

    Pythonaiopen-weighttext-to-speech
    Vezi pe GitHub↗19,324
  • pipecat-ai/pipecatAvatar pipecat-ai

    pipecat-ai/pipecat

    12,846Vezi pe GitHub↗

    Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag

    Enables adjustment of synthesized speech delivery, including speed, volume, pitch, and inflection.

    Pythonaichatbot-frameworkchatbots
    Vezi pe GitHub↗12,846
  • facebookresearch/seamless_communicationAvatar facebookresearch

    facebookresearch/seamless_communication

    11,797Vezi pe GitHub↗

    This project is a multimodal translation framework and large language model capable of speech-to-speech, speech-to-text, and text-to-text translation across nearly 100 languages. It provides a real-time speech translation engine and a comprehensive toolkit for converting spoken audio between languages. The system is distinguished by its ability to preserve the original speaker's tone, pace, and prosody during translation. It utilizes a specialized on-device inference toolkit that converts model checkpoints into C-based libraries, enabling low-latency execution on mobile and edge hardware with

    Preserves the original speaker's tone, pace, and pauses during the translation process.

    Jupyter Notebook
    Vezi pe GitHub↗11,797
  • kittenml/kittenttsAvatar KittenML

    KittenML/KittenTTS

    10,044Vezi pe GitHub↗

    KittenTTS is a neural text-to-speech engine and text-to-audio synthesis tool that converts written text into spoken audio using lightweight neural network models. It functions as both a speech synthesizer and an audio file generator, producing spoken audio for offline playback. The system includes a text normalization processor that expands numbers and abbreviations into full spoken words to improve the naturalness of the synthesized speech. It supports diverse voice options and provides the ability to adjust playback speed.

    Allows adjustment of speech speed and tone by modifying input variables during the inference process.

    Python
    Vezi pe GitHub↗10,044
  • boson-ai/higgs-audioAvatar boson-ai

    boson-ai/higgs-audio

    7,919Vezi pe GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Offers controls for adjusting the emotional tone, speed, and prosody of synthesized conversational speech.

    Python
    Vezi pe GitHub↗7,919
  • zyphra/zonosAvatar Zyphra

    Zyphra/Zonos

    7,225Vezi pe GitHub↗

    Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a multilingual speech generator capable of producing audio in English, Japanese, Chinese, French, and German. The system provides zero-shot voice cloning, allowing the replication of specific human voices using short audio samples. It supports the capture of nuanced behaviors, such as whispering, and provides parametric control over speaking rate, pitch, frequency, and emotional tone. The project covers a broad range of expressive speech synthesis and custom audio generation capabilities,

    Provides prosody controls for manual adjustment of speaking rate, pitch, and emotional tone.

    Python
    Vezi pe GitHub↗7,225
  • ohf-voice/piper1-gplAvatar OHF-Voice

    OHF-Voice/piper1-gpl

    2,897Vezi pe GitHub↗

    This project is a neural text-to-speech system and voice trainer that converts written text into spoken audio across a variety of global languages and regional dialects. It functions as an ONNX-based engine capable of performing fast offline inference and uses a phoneme-based controller to manage precise pronunciation. The system distinguishes itself through a comprehensive toolkit for neural voice training, allowing for the creation of custom single-speaker or multi-speaker models. It supports the export of these models to a standardized open format and provides hardware acceleration via gra

    Uses punctuation markers as phonemes to produce distinct intonations for questions and exclamations.

    C++
    Vezi pe GitHub↗2,897
  1. Home
  2. Artificial Intelligence & ML
  3. Text-to-Speech
  4. Prosody Controls

Explorează sub-etichetele

  • Punctuation-Based IntonationMechanisms that treat punctuation markers as phonemes to control the intonation of questions, exclamations, and pauses. **Distinct from Prosody Controls:** Distinct from Prosody Controls: specifically uses punctuation markers as the control mechanism for delivery rather than general parameter modulation