13 个仓库
Generating video content where facial movements are synchronized to an audio track.
Distinct from Video Generation: Focuses on the specific task of lip-synced dubbing rather than general text-to-video generation.
Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Lip-Synced. Refine with filters or upvote what's useful.
Open-Higgsfield-AI is a generative AI content studio and visual workflow orchestrator. It provides a unified interface for creating photorealistic images and videos, utilizing a node-based editor to chain multiple image, video, and audio models into automated content pipelines. The system functions as an AI video animation tool and local GPU inference engine, allowing users to run generative models on local hardware or remote servers. It includes specialized capabilities for audio-driven lip synchronization and cinematic camera controls to adjust virtual lens and focal settings. The platform
Animates portrait videos to synchronize lip movements with a provided audio track.
Duix-Avatar is an AI digital human toolkit used to create, clone, and animate realistic virtual personas. It functions as a digital persona cloning tool and a text-to-speech animation API that converts written text or audio into synthetic voice and facial motion markers. The framework provides an offline video generation engine that renders digital human animations and lip-synced videos on local hardware. It includes a specialized lip sync engine to synchronize mouth movements with audio waveforms and a pipeline for extracting facial and vocal features from source media to create synthetic re
Generates video content where facial movements are precisely synchronized to an audio track.
PaddleGAN is a generative AI framework and deep learning computer vision library built on the PaddlePaddle framework. It serves as a toolkit for image and video synthesis, providing a collection of generative adversarial network implementations for creating synthetic visual content. The library focuses on advanced synthesis capabilities, including the generation of talking heads through lip motion synchronization and the creation of synthetic videos via motion transfer from driving sequences. It provides tools for domain-to-domain translation, allowing for image style transfer and the transfo
Aligns lip movements in video to match provided audio tracks for realistic talking-head generation.
EMO 是一个 AI 人像动画和音频转视频扩散模型,旨在生成富有表现力的说话人视频。它将单张静态人像图片和一段音频轨道转换为同步的说话人视频。 该系统专注于数字人合成,产生高保真的面部动作和情感线索。它将唇部动作和面部表情与语音录音同步,从而创建逼真的人像动画。 该框架利用扩散过程和跨模态对齐机制来确保音频信号与视觉关键点之间的时间同步。它采用基于参考的图像调节来保持身份一致性,并使用时间一致性层来确保帧间运动的流畅性。
Matches mouth movements and emotional facial cues to an audio file for natural communication.
Video-retalking is an AI lip synchronization framework and talking head video editor designed to match the mouth movements of a subject in a video to a target audio track. It utilizes a deep learning pipeline to synchronize speech with video recordings. The system employs a two-stage generation process that separates coarse lip movement from high-resolution detail refinement. It incorporates identity-aware face refinement and expression template alignment to maintain photorealistic skin textures and ensure visual consistency across video frames. The toolset covers facial expression modificat
Synchronizes the lip movements of a video subject to a specific audio track.
LatentSync 是一个音频驱动的视频生成器和潜在扩散唇形同步模型,旨在将视频中说话者的唇形动作与目标音轨同步。它提供了一个唇形同步训练框架,用于在自定义视频和音频数据集上开发同步网络。 该系统利用视频预处理流水线来清理、分割和对齐人脸数据。它包括一个视觉同步评估工具,该工具计算置信度分数以衡量生成视频中音频和视觉对齐的准确性。 该项目涵盖了自定义同步网络开发、针对硬件内存和分辨率的训练配置管理以及合成视频评估的功能。
Aligns a speaker's lip movements in a video to a target audio track using latent diffusion techniques.
MuseTalk is a deep learning lip synchronization system designed to align video facial movements with audio tracks for high-fidelity video dubbing. It functions as an engine that matches facial expressions to audio input in real-time, enabling the modification of a speaker's lip movements to match new audio sources across different languages. The project features a distributed GPU training pipeline and a multi-stage processing workflow for refining the visual accuracy of synthetic speech. It distinguishes itself through the use of region-specific face masking and mouth openness control, which
Produces high-quality dubbed video by aligning facial regions with audio features using adjustable parameters.
InfiniteTalk is an open-source system for generating talking head videos driven by audio input. It synthesizes realistic lip movements, head poses, and facial expressions synchronized to a spoken audio track, using either a single still image or a small set of reference video frames as the visual source. The system can produce videos of arbitrary length while maintaining temporal coherence, and it supports animating multiple subjects in a single scene. A key differentiator is the ability to coordinate multiple talking subjects through a structured JSON description, giving each independent lip
Generates accurate lip-sync and facial animation for any audio input over arbitrary durations.
Aigcpanel is a visual workflow automation tool and model lifecycle manager designed for generative AI media pipelines. It provides a unified interface to install, launch, and configure both local and remote AI model endpoints, acting as an orchestration platform for large language models and AI tools. The system features a drag-and-drop node editor for chaining AI models and scripts into automated processing pipelines. It distinguishes itself with a breakpoint-aware execution model that allows users to pause and resume long media tasks from specific points in the workflow. Additionally, it in
Implements a sequential pipeline to align synthesized audio with video frames for digital human synthesis.
ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat
Generates new lip movements and audio matching target text prompts while preserving speaker identity.
Linly-Dubbing is an automated video dubbing pipeline designed for multilingual video localization. It converts spoken content in videos into another language by coordinating speech-to-text transcription, text translation, and text-to-speech synthesis. The system distinguishes itself through AI-driven lip synchronization and animation, which aligns facial expressions and mouth movements to the synthesized voiceover. It also utilizes audio source separation to isolate vocals from background music and noise, allowing for clean voice replacement while preserving original background audio. The br
Aligns facial expressions and mouth movements to synthetic audio to produce realistic dubbed video output.
This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro
Aligns character lip movements with audio tracks to create natural-looking video narration.
Rhubarb is an automated lip sync generator and phonetic speech analyzer that converts audio recordings into timed mouth-shape animation data. It identifies sounds and syllables within audio files to map them to specific visual mouth shapes, serving as an animation timing exporter for external character animation software. The tool utilizes a language-independent phonetic recognizer to process speech regardless of the spoken language. To increase accuracy, it supports dialogue-guided recognition by using external text files to guide the phonetic analysis of specific spoken scripts. The system
Uses external text files to guide phonetic recognition and improve the accuracy of mouth-shape synchronization.