2 个仓库
Neural mapping of audio spectral features into low-dimensional latent spaces for facial parameter control.
Distinct from Audio-to-Face Model Training: Focuses on the latent embedding representation rather than the overall training pipeline.
Explore 2 awesome GitHub repositories matching part of an awesome list · Audio-to-Motion Embeddings. Refine with filters or upvote what's useful.
Hallo is an audio-driven talking head generator and portrait animation framework. It synchronizes a static portrait image with an audio file to produce realistic talking head videos by mapping audio spectral features to facial expressions and lip movements. The system utilizes a diffusion video synthesis model that employs iterative denoising and latent representations to generate temporally consistent video frames. It incorporates identity-preserving feature extraction and latent space motion modeling to maintain visual consistency and control facial poses. The toolkit provides capabilities
Maps audio spectral features into a latent space to drive facial expressions and lip parameters.
EchoMimic 是一个多模态人类动画框架和基于扩散的视频生成器。它通过从各种源数据合成运动和外观,生成参考图像的逼真面部和半身动画。 该系统支持由音频、姿势序列或驱动视频驱动的肖像动画。它具有一个标志调节工具,允许通过修改特定的标志点来精确控制面部动作。 该框架涵盖多模态运动合成以及将参考图像同步以匹配目标驱动程序的物理运动。这包括将音频信号转换为面部姿势参数以驱动生成的视频帧的能力。
Implements a neural mapping that transforms raw audio signals into facial pose parameters for video generation.