3 Repos
Using sparse image or sketch data to guide the generation of visual elements in AI models.
Distinct from Sparse Visual Mapping: No candidates cover the use of sparse RGB/sketches as generative guidance specifically for video synthesis.
Explore 3 awesome GitHub repositories matching artificial intelligence & ml · Visual Guidance Inputs. Refine with filters or upvote what's useful.
AnimateDiff is a latent diffusion video generator and text-to-video diffusion framework. It converts existing text-to-image diffusion models into animation generators by applying specialized motion modules, allowing for the creation of video sequences without modifying the original base model. The project provides an image-to-video animation framework that uses sparse RGB images, sketches, or structural keyframe constraints to guide generation. It further distinguishes itself with a motion adapter system that injects cinematic camera movements, such as zooming, panning, and tilting, into anim
Guides video generation using sparse RGB images or sketch inputs to define specific visual elements.
SPADE is a semantic image synthesis framework and generative adversarial network designed to transform semantic label maps into photorealistic images. It uses a spatially-adaptive normalization model to modulate activations based on semantic maps, ensuring that spatial layouts and details are preserved throughout the synthesis process. The project enables the generation of diverse image variations from a single semantic layout by integrating variational autoencoders and latent vector style control. These mechanisms allow for the adjustment of visual appearances and textures while keeping the
Uses semantic label maps as visual guidance to direct the placement and structure of generated objects.
ComfyUI-Easy-Use is a custom node suite and workflow optimizer designed to simplify Stable Diffusion generation pipelines. It provides a set of integrated tools to reduce visual clutter and streamline the process of creating images from text and existing image references. The project distinguishes itself through a pipeline manager that consolidates models, conditioning, and latents into unified data pipes, eliminating complex wiring in the node graph. It also introduces a logical operator set that enables conditional if-else branching and for-loop structures directly within the visual program
Integrates visual guidance inputs such as pose maps or edge maps via ControlNet to steer image generation.