awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

11 مستودعات

Awesome GitHub RepositoriesEmotional Modulation

Techniques for adjusting the intensity and tone of emotional delivery in synthetic speech.

Distinct from Audio Emotion Classifiers: Distinct from audio emotion classifiers: focuses on the active modulation of emotional intensity during generation rather than classification.

Explore 11 awesome GitHub repositories matching graphics & multimedia · Emotional Modulation. Refine with filters or upvote what's useful.

Awesome Emotional Modulation GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • openbmb/voxcpmالصورة الرمزية لـ OpenBMB

    OpenBMB/VoxCPM

    29,985عرض على GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Implements emotional modulation to adjust the tone and intensity of synthetic speech to match target emotions.

    Pythonaudiodeeplearningminicpm
    عرض على GitHub↗29,985
  • openbmb/minicpm-vالصورة الرمزية لـ OpenBMB

    OpenBMB/MiniCPM-V

    25,653عرض على GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Adjusts the intensity and tone of emotional delivery to convey feelings like sadness or excitement.

    Python
    عرض على GitHub↗25,653
  • resemble-ai/chatterboxالصورة الرمزية لـ resemble-ai

    resemble-ai/chatterbox

    22,751عرض على GitHub↗

    Chatterbox is a comprehensive machine learning platform designed for multilingual speech synthesis and real-time audio generation. It functions as an engine that converts text into natural-sounding speech, capable of replicating specific human vocal characteristics and emotional expressions from short audio samples. The platform distinguishes itself through advanced control over the synthesis process, allowing for the manipulation of emotional intensity and the injection of non-verbal vocalizations such as laughter or coughing. It is engineered for low-latency performance, utilizing an optimi

    Modifies the intensity of emotional delivery in generated speech to improve expressiveness.

    Python
    عرض على GitHub↗22,751
  • funaudiollm/cosyvoiceالصورة الرمزية لـ FunAudioLLM

    FunAudioLLM/CosyVoice

    21,673عرض على GitHub↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Generates multilingual speech while applying specific emotional tones for engaging communication.

    Pythonaudio-generationcantonesechatbot
    عرض على GitHub↗21,673
  • livekit/livekitالصورة الرمزية لـ livekit

    livekit/livekit

    19,358عرض على GitHub↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Injects specific emotional states into avatar performance to influence facial expressions during conversation.

    Gogolangmedia-serversfu
    عرض على GitHub↗19,358
  • neonbjb/tortoise-ttsالصورة الرمزية لـ neonbjb

    neonbjb/tortoise-tts

    14,864عرض على GitHub↗

    Tortoise-tts is a neural text-to-speech engine and voice cloning toolkit designed for high-quality audio generation. It functions as a zero-shot synthesis system, meaning it can generate speech for unseen speakers without requiring additional training or fine-tuning for each new voice. The system specializes in replicating human vocal characteristics using small sets of reference audio clips. It allows for the extraction of voice latents to mimic specific speakers, the generation of random synthetic identities, and the blending of multiple voice profiles to create hybrid vocal identities. Th

    Allows for the active modulation of emotional intensity and tone in synthetic speech via text prompts.

    Jupyter Notebook
    عرض على GitHub↗14,864
  • netease-youdao/emotivoiceالصورة الرمزية لـ netease-youdao

    netease-youdao/EmotiVoice

    8,446عرض على GitHub↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    Generates synthetic audio that conveys specific human emotions like happiness or sadness.

    Pythonaideep-learningemotion
    عرض على GitHub↗8,446
  • lipku/livetalkingالصورة الرمزية لـ lipku

    lipku/LiveTalking

    8,042عرض على GitHub↗

    LiveTalking is an interactive talking head engine and AI avatar management platform designed to synchronize synthetic speech with facial movements. It functions as a real-time orchestrator that connects large language models and text-to-speech services to neural-rendered digital humans. The project distinguishes itself through low-latency streaming capabilities and the ability to handle real-time conversational interruptions. It supports advanced audio-visual customization, including human voice cloning and the ability to drive avatar expressions using real-time webcam data. The platform cov

    Translates real-time facial expressions from a webcam into avatar lip-sync and gestures.

    Pythonaigcdigihumandigital-human
    عرض على GitHub↗8,042
  • opentalker/video-retalkingالصورة الرمزية لـ OpenTalker

    OpenTalker/video-retalking

    7,256عرض على GitHub↗

    Video-retalking is an AI lip synchronization framework and talking head video editor designed to match the mouth movements of a subject in a video to a target audio track. It utilizes a deep learning pipeline to synchronize speech with video recordings. The system employs a two-stage generation process that separates coarse lip movement from high-resolution detail refinement. It incorporates identity-aware face refinement and expression template alignment to maintain photorealistic skin textures and ensure visual consistency across video frames. The toolset covers facial expression modificat

    Modifies the emotional state of subjects by applying expression templates to the face.

    Pythonlip-synchronizationsiggraph-asia-2022talking-head-videos
    عرض على GitHub↗7,256
  • open-llm-vtuber/open-llm-vtuberالصورة الرمزية لـ Open-LLM-VTuber

    Open-LLM-VTuber/Open-LLM-VTuber

    5,946عرض على GitHub↗

    Links the AI's emotional state to specific Live2D facial expressions for appropriate visual reactions.

    Pythonaiai-companionai-vtuber
    عرض على GitHub↗5,946
  • caviraoss/openmemoryالصورة الرمزية لـ CaviraOSS

    CaviraOSS/OpenMemory

    3,350عرض على GitHub↗

    OpenMemory is an embeddable memory engine for LLM agents that stores, retrieves, and manages conversational context and agent state using semantic indexing and temporal facts. It functions as a semantic memory store backed by vector indexing, where memories are organized by meaning rather than by exact key, and includes a tiered decay engine that gradually reduces the salience of unused memories while compressing cold vectors and fingerprinting dormant entries to conserve storage. The system also maintains a temporal fact database that records factual statements with subject-predicate-object s

    Implements emotional salience boosting that adjusts memory importance scores based on detected sentiment intensity.

    TypeScriptaiai-agentsai-infrastructure
    عرض على GitHub↗3,350
  1. Home
  2. Graphics & Multimedia
  3. Audio & Music
  4. Audio Processing
  5. Audio Emotion Classifiers
  6. Emotional Modulation

استكشف الوسوم الفرعية

  • Facial Expression Modulators1 وسم فرعيMechanisms for injecting emotional states into avatar facial animations. **Distinct from Emotional Modulation:** Distinct from Emotional Modulation: focuses on visual facial expression control for avatars rather than audio-only emotional modulation.
  • Memory Salience BoostersTechniques that adjust memory importance scores based on detected sentiment intensity to prioritize emotionally charged content. **Distinct from Emotional Modulation:** Distinct from Emotional Modulation in audio: applies sentiment-based importance boosting to memory records, not to speech synthesis.