awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
jianchang512 avatar

jianchang512/clone-voiceArchived

0
View on GitHub↗
8,959 stele·985 fork-uri·Python·22 vizualizăripyvideotrans.com↗

Clone Voice

This project is a GPU-accelerated speech engine and AI voice cloning tool. It functions as a text-to-speech synthesizer and voice-to-voice converter that replicates specific human voices to generate synthetic speech.

The system creates digital voice profiles by analyzing short audio samples or capturing live microphone input. These profiles enable the transformation of existing audio recordings into a target speaker's voice or the synthesis of new audio from written text.

The engine supports subtitle-based speech generation for batch processing and automated dubbing workflows. A web-based audio interface provides a dashboard for recording voice samples and managing synthesis tasks.

Features

  • Voice Cloning Tools - Provides machine learning pipelines that generate high-quality synthetic speech from custom audio recordings.
  • Voice Profiling - Extracts unique vocal characteristics from live microphone input to build personalized voice models.
  • Neural Text-to-Speech Engines - Implements deep learning pipelines that generate synthetic speech by modeling cloned vocal characteristics.
  • GPU Acceleration - Uses GPU hardware acceleration to optimize the processing speed of voice cloning and synthesis models.
  • Text-to-Speech Synthesis - Converts written text and subtitle files into spoken audio using artificial intelligence and cloned voices.
  • Voice Cloning - Replicates specific human vocal characteristics from audio samples to transform existing recordings.
  • Microphone Sampling - Captures audio directly from a microphone to establish voice profiles for cloning and synthesis.
  • Voice Identity Conversions - Transforms the vocal characteristics of a source audio signal to match a target speaker's identity.
  • Voice Profile Management - Registers and stores vocal characteristics to ensure consistency in synthetic speech generation.
  • Voice Embedding Precomputations - Analyzes short audio samples to extract and store reusable vocal embeddings for voice cloning.
  • GPU-Accelerated TTS - Provides a speech synthesis engine specifically optimized for GPU execution to increase audio generation speed.
  • Text-to-Speech Synthesizers - Converts written text or subtitle files into synthetic spoken audio using cloned voice profiles.
  • Subtitle-Driven Dubbing - Utilizes external subtitle files as the primary driver for generating synchronized synthetic voiceovers.
  • Subtitle-Driven Audio Synthesis - Generates dubbed audio tracks based on the timing and text of external subtitle files.
  • Speech Synthesis Generators - Generates audible cloned speech based on the text and timing found in subtitle files.

Istoric stele

Graficul istoricului de stele pentru jianchang512/clone-voiceGraficul istoricului de stele pentru jianchang512/clone-voice

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Colecții curatoriate care includ Clone Voice

Colecții selectate manual în care apare Clone Voice.
  • Clonarea și sinteza vocii prin AI
  • Modele pentru sinteza și recunoașterea vorbirii

Alternative open-source pentru Clone Voice

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu Clone Voice.
  • ohf-voice/piper1-gplAvatar OHF-Voice

    OHF-Voice/piper1-gpl

    2,897Vezi pe GitHub↗

    This project is a neural text-to-speech system and voice trainer that converts written text into spoken audio across a variety of global languages and regional dialects. It functions as an ONNX-based engine capable of performing fast offline inference and uses a phoneme-based controller to manage precise pronunciation. The system distinguishes itself through a comprehensive toolkit for neural voice training, allowing for the creation of custom single-speaker or multi-speaker models. It supports the export of these models to a standardized open format and provides hardware acceleration via gra

    C++
    Vezi pe GitHub↗2,897
  • kevinwang676/bark-voice-cloningAvatar KevinWang676

    KevinWang676/Bark-Voice-Cloning

    2,957Vezi pe GitHub↗

    Bark Voice Cloning is a text-to-speech synthesis engine designed to generate natural-sounding audio and replicate specific vocal characteristics. The system utilizes a transformer-based autoregressive model to convert written text into high-fidelity speech, supporting multilingual output and expressive delivery. The project distinguishes itself through zero-shot voice cloning, which extracts speaker identity embeddings from short audio samples to condition the generative model without requiring extensive fine-tuning. It also provides specialized workflows for voice identity conversion, allowi

    Jupyter Notebook
    Vezi pe GitHub↗2,957
  • elevenlabs/elevenlabs-pythonAvatar elevenlabs

    elevenlabs/elevenlabs-python

    2,873Vezi pe GitHub↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    Vezi pe GitHub↗2,873
  • boson-ai/higgs-audioAvatar boson-ai

    boson-ai/higgs-audio

    7,919Vezi pe GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Python
    Vezi pe GitHub↗7,919
Vezi toate cele 30 alternative pentru Clone Voice→

Întrebări frecvente

Ce face jianchang512/clone-voice?

This project is a GPU-accelerated speech engine and AI voice cloning tool. It functions as a text-to-speech synthesizer and voice-to-voice converter that replicates specific human voices to generate synthetic speech.

Care sunt principalele funcționalități ale jianchang512/clone-voice?

Principalele funcționalități ale jianchang512/clone-voice sunt: Voice Cloning Tools, Voice Profiling, Neural Text-to-Speech Engines, GPU Acceleration, Text-to-Speech Synthesis, Voice Cloning, Microphone Sampling, Voice Identity Conversions.

Care sunt câteva alternative open-source pentru jianchang512/clone-voice?

Alternativele open-source pentru jianchang512/clone-voice includ: kevinwang676/bark-voice-cloning — Bark Voice Cloning is a text-to-speech synthesis engine designed to generate natural-sounding audio and replicate… ohf-voice/piper1-gpl — This project is a neural text-to-speech system and voice trainer that converts written text into spoken audio across a… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… boson-ai/higgs-audio — Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large… babysor/mockingbird — MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions… mozilla/tts — This project is a comprehensive suite for neural speech synthesis, featuring a deep learning text-to-speech engine, a…