1 Repo
Systems capable of performing human animation tasks using diverse input types within a single model.
Distinct from Multi-Modal Embedding Models: Covers the general architectural capability of multi-modal input for animation, which is not captured by prompt integration or LLMs.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Multi-Modal Animation Frameworks. Refine with filters or upvote what's useful.
EchoMimic is a multimodal human animation framework and diffusion-based video generator. It produces lifelike facial and semi-body animations of a reference image by synthesizing motion and appearance from various source data. The system enables portrait animation driven by audio, pose sequences, or driver videos. It features a landmark conditioning tool that allows for the precise control of facial movements by modifying specific landmark points. The framework covers multi-modal motion synthesis and the synchronization of reference images to match the physical movements of a target driver.
Executes human animation tasks across audio and pose inputs using a unified high-parameter model.