For an open source alternative to HeyGen, the strongest matches are meigen-ai/infinitetalk (InfiniteTalk is an open-source talking-head video generator that creates), kedreamix/linly-dubbing (Linly-Dubbing provides an automated video dubbing pipeline featuring text-to-speech) and lipku/livetalking (LiveTalking is an interactive talking-head engine that synchronizes text-to-speech). humanaigc/emo and opentalker/video-retalking round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
We curate open-source GitHub repositories matching “open source alternatives to heygen”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.
InfiniteTalk is an open-source system for generating talking head videos driven by audio input. It synthesizes realistic lip movements, head poses, and facial expressions synchronized to a spoken audio track, using either a single still image or a small set of reference video frames as the visual source. The system can produce videos of arbitrary length while maintaining temporal coherence, and it supports animating multiple subjects in a single scene. A key differentiator is the ability to coordinate multiple talking subjects through a structured JSON description, giving each independent lip
InfiniteTalk is an open-source talking-head video generator that creates lip-synced avatar animations from audio input, matching the core generation capability while lacking a full built-in text-to-speech pipeline and multilingual translation suite.
Linly-Dubbing is an automated video dubbing pipeline designed for multilingual video localization. It converts spoken content in videos into another language by coordinating speech-to-text transcription, text translation, and text-to-speech synthesis. The system distinguishes itself through AI-driven lip synchronization and animation, which aligns facial expressions and mouth movements to the synthesized voiceover. It also utilizes audio source separation to isolate vocals from background music and noise, allowing for clean voice replacement while preserving original background audio. The br
Linly-Dubbing provides an automated video dubbing pipeline featuring text-to-speech synthesis, multilingual translation, and AI-driven lip synchronization for talking-head videos, though it focuses primarily on localization rather than a web-based avatar creation platform.
LiveTalking is an interactive talking head engine and AI avatar management platform designed to synchronize synthetic speech with facial movements. It functions as a real-time orchestrator that connects large language models and text-to-speech services to neural-rendered digital humans. The project distinguishes itself through low-latency streaming capabilities and the ability to handle real-time conversational interruptions. It supports advanced audio-visual customization, including human voice cloning and the ability to drive avatar expressions using real-time webcam data. The platform cov
LiveTalking is an interactive talking-head engine that synchronizes text-to-speech and neural avatars for real-time video generation, though it lacks a full web-based timeline editor.
EMO is an AI portrait animator and audio-to-video diffusion model designed to generate expressive talking head videos. It transforms a single static portrait image and an audio track into a synchronized video of a person speaking. The system focuses on digital human synthesis, producing high-fidelity facial movements and emotional cues. It synchronizes lip movements and facial gestures to match spoken voice recordings to create realistic portrait animations. The framework utilizes a diffusion process and a cross-modal alignment mechanism to ensure timing between audio signals and visual land
This repository provides an audio-driven talking head diffusion model for portrait animation, but it is a research-oriented synthesis model rather than a complete web-based video generation platform with self-hosting, text-to-speech, and editing tools.
Video-retalking is an AI lip synchronization framework and talking head video editor designed to match the mouth movements of a subject in a video to a target audio track. It utilizes a deep learning pipeline to synchronize speech with video recordings. The system employs a two-stage generation process that separates coarse lip movement from high-resolution detail refinement. It incorporates identity-aware face refinement and expression template alignment to maintain photorealistic skin textures and ensure visual consistency across video frames. The toolset covers facial expression modificat
Video-retalking is an AI lip-synchronization framework for driving talking heads from audio, but it functions as a video editing and enhancement library rather than a complete self-hostable platform with built-in text-to-speech and a web editor.
LatentSync is an audio-driven video generator and latent diffusion lip sync model designed to synchronize a speaker's lip movements in a video to a target audio track. It provides a lip synchronization training framework for developing synchronization networks on custom video and audio datasets. The system utilizes a video preprocessing pipeline to clean, segment, and align face data. It includes a visual sync evaluation tool that calculates confidence scores to measure the accuracy of audio and visual alignment in generated videos. The project covers capabilities for custom synchronization
LatentSync is a specialised lip-syncing and audio-driven diffusion model for aligning talking heads rather than a complete end-to-end video generation platform with text-to-speech and a web editor.
CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo
This repository is a speech generation and voice cloning model rather than a complete talking-head video generation platform with avatars and a web-based video editor.
VideoLingo is an automated video localization suite designed to transcribe, translate, and dub video content. It functions as a translation pipeline that utilizes large language models to convert spoken audio into precise text segments and translate them into multiple languages. The system differentiates itself through a multi-step translation refinement process and a specialized natural language processing utility that segments text into single-line captions meeting broadcast standards. It also integrates synthetic voiceover generation to replace or augment original audio tracks. The projec
VideoLingo is a video translation, dubbing, and localization pipeline rather than an avatar-driven talking-head video generation platform.
RedditVideoMakerBot is a social media video creator and automation bot that transforms Reddit threads into short-form videos. It functions as a text-to-speech video generator, programmatically fetching posts and comments to create narrated clips with background visuals. The system integrates a content downloader for Reddit data with a voice engine that synthesizes spoken audio from written text. It manages the assembly of these components by combining image sequences and audio tracks into completed video files. The tool includes a web-based configuration interface for managing bot settings a
This repository automates Reddit thread narration into short-form videos with text-to-speech, but it creates programmatic social-media clips rather than a customizable talking-head AI avatar platform.
MuseTalk is a deep learning lip synchronization system designed to align video facial movements with audio tracks for high-fidelity video dubbing. It functions as an engine that matches facial expressions to audio input in real-time, enabling the modification of a speaker's lip movements to match new audio sources across different languages. The project features a distributed GPU training pipeline and a multi-stage processing workflow for refining the visual accuracy of synthetic speech. It distinguishes itself through the use of region-specific face masking and mouth openness control, which
MuseTalk is a real-time lip-sync and video dubbing engine rather than a complete web-based video generation platform with text-to-speech synthesis and avatar creation suites.
Pocket-tts is a text-to-speech server and neural speech synthesizer that converts written text into audible speech. It includes a CPU-optimized inference engine and a voice cloning tool capable of analyzing audio samples to reproduce specific speaker characteristics. The system differentiates itself through the use of dynamic int8 quantization to reduce memory usage and increase generation speed on processors. It supports real-time speech synthesis by streaming audio chunks incrementally and utilizes voice state caching to store processed embeddings as portable files, bypassing redundant proc
This project provides text-to-speech synthesis and voice cloning rather than a complete video generation platform with AI avatars and a web editor.
Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt
Dia is a text-to-speech synthesis and generative audio engine focused on voice cloning rather than a complete video generation platform with AI avatars and web editing.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| meigen-ai/infinitetalk | 4.8K | Python | apache-2.0 | |
| kedreamix/linly-dubbing | 3K | Jupyter Notebook | apache-2.0 | |
| lipku/livetalking | 8K | Python | Apache-2.0 | |
| humanaigc/emo | 7.6K | — | — | |
| opentalker/video-retalking | 7.3K | Python | Apache-2.0 | |
| bytedance/latentsync | 5.8K | Python | Apache-2.0 | |
| sesameailabs/csm | 14.7K | Python | Apache-2.0 | |
| huanshere/videolingo | 17.5K | Python | Apache-2.0 | |
| elebumm/redditvideomakerbot | 12.5K | Python | GPL-3.0 | |
| tmelyralab/musetalk | 5.3K | Python | other |