37 Repos
Models that generate realistic talking head videos from audio input.
Explore 37 awesome GitHub repositories matching part of an awesome list · Audio Driven Synthesis. Refine with filters or upvote what's useful.
SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body
Realistic 3D motion coefficients for stylized talking head animation.
Wav2Lip is a deep learning lip sync model and neural talking head framework designed to synchronize the lip movements in a video to match a provided audio file. It functions as a computer vision lip synchronizer and speech-to-lip generator that maps speech patterns to visual mouth movements to produce realistic talking head videos. The system utilizes a framework for training and evaluating models that align audio and video frames. This includes the ability to train lip-sync models and visual discriminators using speech-to-lip datasets and evaluating the resulting synchronization accuracy thr
Lip sync expert for speech-to-lip generation in the wild.
Hallo is an audio-driven talking head generator and portrait animation framework. It synchronizes a static portrait image with an audio file to produce realistic talking head videos by mapping audio spectral features to facial expressions and lip movements. The system utilizes a diffusion video synthesis model that employs iterative denoising and latent representations to generate temporally consistent video frames. It incorporates identity-preserving feature extraction and latent space motion modeling to maintain visual consistency and control facial poses. The toolkit provides capabilities
Hierarchical audio-driven visual synthesis for portrait animation.
EMO ist ein KI-Porträt-Animator und Audio-zu-Video-Diffusionsmodell, das entwickelt wurde, um ausdrucksstarke Talking-Head-Videos zu generieren. Es verwandelt ein einzelnes statisches Porträtbild und eine Audiospur in ein synchronisiertes Video einer sprechenden Person. Das System konzentriert sich auf die Synthese digitaler Menschen und erzeugt hochauflösende Gesichtsbewegungen und emotionale Signale. Es synchronisiert Lippenbewegungen und Gesichtsausdrücke mit gesprochenen Sprachaufnahmen, um realistische Porträt-Animationen zu erstellen. Das Framework nutzt einen Diffusionsprozess und einen Cross-Modal-Alignment-Mechanismus, um das Timing zwischen Audiosignalen und visuellen Landmarks sicherzustellen. Es verwendet referenzbasierte Bildkonditionierung, um die Identitätskonsistenz zu wahren, sowie eine zeitliche Konsistenzschicht, um flüssige Bewegungen zwischen den Frames zu gewährleisten.
Expressive portrait video generation using audio-to-video diffusion.
LatentSync ist ein audio-gesteuerter Videogenerator und ein Latent-Diffusion-Lip-Sync-Modell, das darauf ausgelegt ist, die Lippenbewegungen eines Sprechers in einem Video mit einer Ziel-Audiospur zu synchronisieren. Es bietet ein Lip-Sync-Trainings-Framework zur Entwicklung von Synchronisationsnetzwerken auf benutzerdefinierten Video- und Audiodatensätzen. Das System nutzt eine Video-Vorverarbeitungspipeline, um Gesichtsdaten zu bereinigen, zu segmentieren und auszurichten. Es enthält ein visuelles Sync-Evaluierungstool, das Konfidenzwerte berechnet, um die Genauigkeit der Audio- und Videoausrichtung in generierten Videos zu messen. Das Projekt deckt Funktionen für die Entwicklung benutzerdefinierter Synchronisationsnetzwerke, die Verwaltung von Trainingskonfigurationen für Hardwarespeicher und Auflösung sowie die Evaluierung synthetischer Videos ab.
Audio-conditioned latent diffusion models for lip synchronization.
MuseTalk is a deep learning lip synchronization system designed to align video facial movements with audio tracks for high-fidelity video dubbing. It functions as an engine that matches facial expressions to audio input in real-time, enabling the modification of a speaker's lip movements to match new audio sources across different languages. The project features a distributed GPU training pipeline and a multi-stage processing workflow for refining the visual accuracy of synthetic speech. It distinguishes itself through the use of region-specific face masking and mouth openness control, which
Real-time high-quality lip synchronization using latent inpainting.
AniPortrait is an AI video synthesis pipeline designed to generate photorealistic speaking portraits and facial animations. It functions as a talking head generator and audio-driven animator that synchronizes lip movements, expressions, and head poses to speech or reference video sources. The system includes a facial expression transfer tool for reenacting movements from a source video onto a static reference image. It utilizes a latent diffusion model with reference-based image conditioning to maintain visual identity and consistency across generated frames. The pipeline covers audio-to-exp
Audio-driven synthesis of photorealistic portrait animations.
InfiniteTalk is an open-source system for generating talking head videos driven by audio input. It synthesizes realistic lip movements, head poses, and facial expressions synchronized to a spoken audio track, using either a single still image or a small set of reference video frames as the visual source. The system can produce videos of arbitrary length while maintaining temporal coherence, and it supports animating multiple subjects in a single scene. A key differentiator is the ability to coordinate multiple talking subjects through a structured JSON description, giving each independent lip
Synthesizes lip, head, and expression movements directly from audio features using a trained neural network.
EchoMimic is an audio-driven portrait animation framework and latent diffusion video generator. It transforms static reference images into dynamic talking head videos by synchronizing facial movements with audio tracks and motion drivers. The system functions as a hybrid motion synthesis engine that combines audio inputs and pose data. It utilizes a facial landmark motion controller to edit positioning markers, enabling precise synchronization and video-to-video pose transfer. The pipeline covers image-to-video animation through latent diffusion and facial landmark conditioning. This allows
Lifelike audio-driven portrait animations with editable landmarks.
Hallo2 is an AI video generation tool and audio-driven portrait animation framework designed to transform static images into speaking videos. It functions as a portrait image animator that synchronizes a single photo with an audio track to produce high-resolution talking head videos. The system includes a distributed animation trainer for fine-tuning deep learning models using custom datasets and distributed computing resources. It employs hierarchical video generation and temporal consistency modeling to produce long-form character animations that remain stable over extended durations. The
Long-duration and high-resolution audio-driven portrait animation.
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
Efficient audio-driven avatar generation with adaptive body animation.
Streaming real-time audio-driven avatar generation.
This repository contains the implementation of the following paper:
Real-time photorealistic talking-head animation.
The source code of "DINet: deformation inpainting network for realistic face visually dubbing on high resolution video."
Deformation inpainting for realistic face dubbing on high-res video.
)](https://github.com/yerfor/Real3DPortrait) | 中文文档
One-shot realistic 3D talking portrait synthesis.
This is the code repository implementing the paper:
Speaker-aware talking-head animation from audio.
JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation
Diffusion-based audio-driven facial dynamics for portraits and animals.
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
Spatial knitting attentions for embedding conditions in diffusion models.
The pytorch implementation for our CVPR2023 paper "DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation".
Diffusion models for generalized talking head synthesis.
1 Shanghai Jiao Tong University 2 NetEase Fuxi AI Lab ECCV 2024 Oral
Efficient disentanglement for emotional talking head synthesis.