4 Repos
Injecting skeletal and pose data as guidance signals into diffusion denoising processes.
Distinct from Diffusion Conditioning Architectures: Specializes general diffusion conditioning to the use of pose-based spatial signals.
Explore 4 awesome GitHub repositories matching artificial intelligence & ml · Pose Conditioning. Refine with filters or upvote what's useful.
AnimateAnyone is an appearance-preserving video synthesizer designed for character animation from a single static image. It functions as a diffusion image-to-video generator that transforms a source image into a high-fidelity video sequence while maintaining consistent character identity, clothing, and visual details across all frames. The system enables video-driven character reenactment by transferring motions, facial expressions, and body movements from a reference video onto a static character. It employs pose-guided video generation to control movement via skeleton keypoints and pose sig
Uses spatially-aligned skeleton keypoints as conditioning signals to drive character movement during the denoising process.
EchoMimic V2 ist eine KI-Video-Generierungs-Pipeline und ein Computer-Vision-Animationsmodell, das darauf ausgelegt ist, synthetische menschliche Animationen zu produzieren. Es fungiert als generatives Framework, das Halbkörper-Videos erstellt, indem ein statisches Referenzbild mit Posenbewegungen abgeglichen wird, die aus einem treibenden Video extrahiert wurden. Das System nutzt einen diffusionsbasierten Generierungsprozess in Kombination mit latenter Raumkompression und einem temporalen Aufmerksamkeitsmechanismus, um flüssige Übergänge zwischen Frames zu gewährleisten. Es wahrt die konsistente Identität einer Person durch referenzbasiertes Encoding und steuert die räumliche Platzierung mittels posengesteuerter Bewegungskonditionierung. Das Projekt enthält Funktionen zur mehrstufigen Bildverfeinerung, um Gesichtsdetails und Schärfe zu verbessern. Zudem bietet es Tools zur Vorbereitung von Animationsdatensätzen, einschließlich des Herunterladens und der Vorverarbeitung von Videodaten in Formate, die für Modelltraining und Inferenz erforderlich sind.
Uses pose-based conditioning to guide the spatial placement and movement of the generated human figure.
Champ is a generative vision system and controllable image-to-video generator designed for human image animation. It uses a diffusion-based video synthesizer and 3D parametric guidance to transform a single reference image into a consistent sequence of motion based on external driving data. The framework distinguishes itself through a human pose transfer system that employs 3D body parametric extraction and coordinate-space alignment. This allows the model to map motion from a driving video to a reference person by adjusting for body scales and camera perspectives using depth and semantic con
Renders processed 3D body data into visual condition maps to guide the animation process.
ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat
Implements pose-conditioning layers to steer character movement and camera paths during video synthesis.