For an open source model for video generation, the strongest matches are ailab-cvc/videocrafter (VideoCrafter is a latent diffusion-based model specifically designed for), zai-org/cogvideo (CogVideo is a comprehensive generative video framework that natively) and thudm/cogvideo (CogVideo is a comprehensive generative video framework that natively). guoyww/animatediff and hpcaitech/open-sora round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
我们为您精选了匹配 “best open source video generation models” 的开源 GitHub 仓库。结果按与您查询的相关性进行排名 — 您可以使用下方筛选器缩小范围,或通过 AI 进行优化。
Videocrafter is a latent diffusion model designed for AI video synthesis. It functions as both a text-to-video and image-to-video generation system, synthesizing high-quality video sequences from descriptive text prompts or static image inputs. The model utilizes a diffusion-based neural network to transform inputs into animated content, ensuring visual consistency and temporal coherence throughout the generated sequences. This allows for the creation of custom video clips and the animation of static images into fluid motion.
VideoCrafter is a latent diffusion-based model specifically designed for both text-to-video and image-to-video generation, providing the core architecture and capabilities required for AI video synthesis.
CogVideo is a video generation framework and large language model architecture designed for synthesizing high-resolution video clips from natural language descriptions and images. It functions as a text-to-video and image-to-video generator, while also providing a model for video captioning to analyze visual content into descriptive text summaries. The system supports animating static images into motion sequences and transforming series of images into video based on prompts. It includes capabilities for extending the length of generated video clips to create longer sequences of motion. The f
CogVideo is a comprehensive generative video framework that natively supports both text-to-video and image-to-video generation using a latent diffusion architecture, making it a flagship solution for this category.
CogVideo is a generative video framework that uses diffusion models and transformer-based architectures to synthesize high-resolution video clips. It functions as both a text-to-video and image-to-video generator, converting textual descriptions or static images into temporal visual sequences. The system integrates large language model capabilities to expand short user prompts into detailed descriptions for better visual alignment. It supports the animation of static images through latent seeding and provides the ability to extend the length of existing video sequences. The project includes
CogVideo is a comprehensive generative video framework that natively supports both text-to-video and image-to-video synthesis using diffusion-based architectures, making it a flagship example for this category.
AnimateDiff is a latent diffusion video generator and text-to-video diffusion framework. It converts existing text-to-image diffusion models into animation generators by applying specialized motion modules, allowing for the creation of video sequences without modifying the original base model. The project provides an image-to-video animation framework that uses sparse RGB images, sketches, or structural keyframe constraints to guide generation. It further distinguishes itself with a motion adapter system that injects cinematic camera movements, such as zooming, panning, and tilting, into anim
AnimateDiff is a specialized framework for text-to-video and image-to-video generation that utilizes a latent diffusion architecture to produce temporally consistent animations, directly matching the requirements for generative video tools.
Open-Sora is a video generation framework designed to produce cinematic sequences from text prompts and images. It functions as a generative system that transforms written descriptions or reference images into video content featuring realistic textures and lighting. The project includes a dedicated prompt engineering tool that uses large language models to expand simple user inputs into detailed descriptions. It also features a motion controller for adjusting movement intensity in generated sequences and evaluating motion levels in existing video files. The framework incorporates text-to-vid
Open-Sora is a comprehensive generative video framework that supports both text-to-video and image-to-video generation using a diffusion transformer architecture, making it a direct match for your requirements.
HunyuanVideo-1.5 is a video generation foundation model and text-to-video diffusion framework. It utilizes a latent video diffusion model and a spatio-temporal transformer architecture to generate high-definition video sequences from text descriptions and images. The project enables cinematic camera control for directing pans and tilts and provides image-to-video animation capabilities. It supports visual style adaptation through low-rank adaptation tuning and uses a language model for prompt refinement to improve visual alignment. The model covers high-resolution video upscaling via a super
HunyuanVideo-1.5 is a comprehensive generative video framework that natively supports both text-to-video and image-to-video generation using a latent diffusion architecture, directly addressing all the core requirements for high-quality video synthesis.
Open-Sora-Plan is a text-to-video framework and distributed video training system. It utilizes a diffusion transformer architecture and large language model components to transform written descriptions or image prompts into high-quality video sequences. The system features a distributed infrastructure designed for large-scale video training and inference. It employs sequence parallelism to split high-resolution or long-duration video samples across multiple GPUs and uses a sparse attention mechanism to increase processing speed. The project includes capabilities for both text-to-video and im
Open-Sora-Plan is a comprehensive generative video framework that supports both text-to-video and image-to-video generation using a diffusion transformer architecture, making it a direct fit for your requirements.
Wan2.1 is a generative video synthesis framework that provides foundation models for creating high-fidelity video sequences and static images from descriptive text prompts. The system utilizes a unified architecture trained on both static and dynamic datasets, allowing it to function as a comprehensive tool for visual media creation. The framework distinguishes itself through a transformer-based temporal modeling approach that ensures structural coherence and consistent motion across video frames. It supports multi-resolution latent scaling, enabling the generation of content in various aspec
Wan2.1 is a comprehensive generative video framework that natively supports both text-to-video and image-to-video generation using a diffusion-based architecture designed for high temporal consistency and efficient latent scaling.
LongCat-Video is a collection of specialized models for video synthesis, featuring a large language model based architecture for creating high-resolution videos from text, images, or existing sequences. It includes dedicated systems for text-to-video generation, image-to-video animation, and the creation of talking avatars. The project provides specific capabilities for extending the length of existing clips through a video continuation model that predicts subsequent frames. It also enables the synchronization of character lip movements with audio and text prompts to produce speaking videos.
This repository provides a suite of specialized models for text-to-video and image-to-video generation, utilizing diffusion-based architectures and optimization techniques to handle long-form video synthesis.
Sana is a framework for high-resolution image and video synthesis based on a linear diffusion transformer. It provides a toolkit for the training, fine-tuning, and execution of text-to-image and text-to-video models, as well as a video generative world model capable of simulating physical environments with precise spatial control. The project is distinguished by its use of linear complexity layers to handle high resolutions and its support for long-form, minute-length video generation in real time. It implements a two-stage inference paradigm that separates structural generation from visual t
Sana is a comprehensive framework for high-resolution video and image synthesis that utilizes a linear diffusion transformer architecture to support both text-to-video and image-to-video generation with a focus on temporal consistency and memory-efficient inference.
This is a framework for training and sampling diffusion models to generate high-fidelity images, video, and 4D assets. It provides a modular environment for managing generative AI training pipelines, including the handling of datasets, noise sampling, and loss weighting to stabilize the creation of synthetic content. The project features a modular model configuration system that uses YAML-based assembly to define network submodules and conditioners. It also includes a dedicated toolset for AI image watermarking, allowing for the embedding and detection of invisible markers to verify the origi
This framework provides the core diffusion-based architecture and sampling tools necessary for text-to-video and image-to-video generation, serving as the official repository for Stability AI's generative models.
This project is a research-oriented PyTorch framework designed for the implementation and training of generative video diffusion models. It provides a modular toolkit that extends standard image-based diffusion techniques into three dimensions, enabling the synthesis of coherent video sequences through iterative denoising processes. The framework distinguishes itself by utilizing factored space-time attention, which decomposes high-dimensional video data into separate spatial and temporal layers to maintain motion consistency while managing computational complexity. It supports multi-modal tr
This is a research-oriented PyTorch framework that provides the core architecture for training and implementing diffusion-based video generation models, making it a suitable tool for developers building their own generative video systems.
Diffusers is a PyTorch-based library and generative AI framework used to build, train, and deploy diffusion pipelines for producing multi-modal media. It provides a suite of tools for generating images, video, and audio from natural language descriptions, as well as specialized systems for text-to-image generation. The project differentiates itself through a modular architecture that separates noise schedulers, pretrained model blocks, and pipeline compositions. This structure allows for the construction of custom generation workflows and the ability to swap individual components of the diffu
This library provides the foundational framework and pre-built pipelines for text-to-video and image-to-video generation using diffusion models, making it the standard tool for implementing these generative tasks.
Text2Video-Zero is a text-to-video diffusion model and framework designed to synthesize temporally consistent video sequences from textual prompts. It functions as a zero-shot video generator, repurposing pre-trained image diffusion models to create video content without requiring additional training on video datasets. The system includes a conditional video synthesizer that allows for guided generation using depth, edge, or pose maps to control structural layout and movement. It also provides text-based video editing capabilities to modify the style or content of existing video clips through
This is a zero-shot diffusion-based framework that enables text-to-video generation by repurposing existing image models, fitting the category well while focusing on zero-shot synthesis rather than training-heavy approaches.
imaginAIry is a system for generating and refining images and videos using diffusion models. It operates as a web-based server that triggers generation requests through standard API calls, allowing for the creation of visuals and video sequences from text prompts or existing files. The project provides a suite for AI image editing and upscaling, enabling the modification of visuals through natural language instructions and super-resolution tools to increase detail and image size. The system includes capabilities for structural image control using depth maps, edge maps, and body poses to main
This project provides a functional system for text-to-video and image-to-video generation using diffusion models, offering a practical implementation for users looking to generate video content via API.
TurboDiffusion is a video diffusion inference engine and generator designed to create high-resolution videos from text prompts and images. It provides a runtime environment for executing optimized diffusion model checkpoints with a focus on reducing latency and GPU memory usage. The project features a specialized training framework for aligning sparse-linear attention models with pretrained full-attention models. This system includes capabilities for sparse attention parameter merging and sparse-linear model alignment to reduce computational costs during inference while maintaining output qua
This repository provides an inference engine and framework specifically designed for text-to-video and image-to-video generation using diffusion models, directly addressing the core requirements for generative video content creation.
StoryDiffusion is a generative AI system designed for consistent character image and video generation. It utilizes a pluggable cross-attention module to inject shared character representations into pretrained diffusion models, allowing for visual identity stability across multiple images and scenes without retraining the base model. The project features a video generation pipeline that produces temporally coherent sequences from text prompts or condition images. It employs a latent space motion interpolator to predict intermediate frames and semantic motion, enabling long-range video generati
StoryDiffusion is a generative video model that supports both text-to-video and image-to-video generation using a diffusion-based architecture with specific optimizations for low-VRAM environments and temporal consistency.
Tune-A-Video is a text-to-video diffusion framework designed to convert pretrained text-to-image diffusion models into video generators. It utilizes a spatio-temporal attention mechanism and single text-video pair training to enable the synthesis of moving sequences from text prompts. The project provides tools for one-shot video personalization, allowing a model to be tuned on a single reference video to preserve specific characters or artistic styles across new generations. It also functions as a video editor that modifies subjects, backgrounds, and styles through noise-sampling prompt guid
This framework enables text-to-video generation by adapting existing diffusion models, providing the core generative capabilities and temporal attention mechanisms required for video synthesis.
ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat
This repository provides a modular framework and node-based interface for generating video from text and image prompts using latent diffusion, though it functions as an extension for the ComfyUI ecosystem rather than a standalone model repository.
MAGI-1 is an autoregressive video generation model designed to synthesize high-resolution video sequences from text prompts and image references. It functions as a generative system for text-to-video, image-to-video, and video-to-video transformations. The model utilizes an autoregressive architecture that treats spatio-temporal patches as a sequence of discrete tokens to maintain temporal motion. It employs a variational autoencoder to compress the spatial and temporal dimensions of video data and uses distillation-based step scaling to allow for inference budget control. The system integra
MAGI-1 is a generative video model that supports both text-to-video and image-to-video synthesis, fitting the category despite its autoregressive approach rather than a pure diffusion-based architecture.
EchoMimic V2 is an AI video generation pipeline and computer vision animation model designed to produce synthetic human animations. It functions as a generative framework that creates semi-body videos by aligning a static reference image with pose movements extracted from a driving video. The system utilizes a diffusion-based generation process combined with latent space compression and a temporal attention mechanism to ensure smooth transitions between frames. It maintains consistent person identity through reference-based encoding and guides spatial placement via pose-driven motion conditio
This repository provides a specialized diffusion-based framework for image-to-video character animation, which fits the category of generative video models despite its narrow focus on human motion rather than general-purpose text-to-video generation.
Sygil-webui is a web interface for Stable Diffusion latent diffusion models, providing a creative suite for text-to-image and text-to-video synthesis. It functions as an image generation tool and a latent diffusion image editor, allowing users to create visuals and video sequences from textual descriptions. The project includes a dedicated model training interface for creating custom textual inversion embeddings, which introduces specific new concepts or styles into the diffusion models. It also features specialized tools for generative image editing, including mask-based inpainting, image-to
This repository provides a web-based interface for running latent diffusion models, enabling both text-to-video and image-to-video generation alongside its primary image synthesis capabilities.
AnimateAnyone is an appearance-preserving video synthesizer designed for character animation from a single static image. It functions as a diffusion image-to-video generator that transforms a source image into a high-fidelity video sequence while maintaining consistent character identity, clothing, and visual details across all frames. The system enables video-driven character reenactment by transferring motions, facial expressions, and body movements from a reference video onto a static character. It employs pose-guided video generation to control movement via skeleton keypoints and pose sig
This repository provides a specialized diffusion-based model for image-to-video character animation, directly addressing the core capability of generating video from image prompts with temporal consistency.
This is a PyTorch-based implementation of diffusion models for synthesizing photorealistic images and video. It provides a framework for text-to-image and text-to-video generation, as well as unconditional image synthesis. The system utilizes a cascading diffusion pipeline to produce high-resolution imagery by passing low-resolution outputs through a sequence of super-resolution models. It also includes capabilities for image inpainting, allowing the reconstruction of masked or missing regions of visual media guided by surrounding context and text prompts. The project includes tools for diff
This repository provides a PyTorch-based implementation of diffusion models specifically designed for text-to-video and text-to-image synthesis, making it a direct tool for generative video tasks.
ComfyUI is a node-based generative AI orchestration engine designed for constructing, testing, and executing complex image and video synthesis pipelines. By utilizing a directed acyclic graph execution model, the platform allows users to build reproducible workflows through modular, interconnected processing blocks without requiring manual code implementation. It serves as both a local environment for high-performance model inference and a production-ready server for deploying generative capabilities. The platform distinguishes itself through its focus on workflow portability and extensibilit
ComfyUI is a powerful node-based orchestration engine that enables text-to-video and image-to-video generation by executing complex diffusion-based pipelines, making it a highly flexible tool for managing generative video workflows.
VACE is a set of software tools and frameworks for reference-guided video generation, diffusion-based editing, and video-to-video translation. It provides utilities to produce new video content and modify existing sequences by using reference materials to guide visual style, subject matter, and composition. The framework enables video-to-video translation and synthesis, allowing for the update of visual styles and depth. It also functions as a video editor for modifying properties and content through reference-guided transformations. The system covers localized video editing and inpainting,
VACE is a diffusion-based framework designed for reference-guided video generation and video-to-video translation, making it a specialized tool for generative video editing and synthesis.
FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine, a video diffusion training framework, and a modular pipeline orchestrator. It provides a distributed transformer optimizer and a distillation toolkit designed to reduce denoising steps and model complexity to increase frame rates. The project distinguishes itself through specialized acceleration techniques, including joint distillation and sparse attention training. It implements low-step video generation and weight quantization to FP8 or FP4 precision to increase throughput a
FastVideo is a specialized framework for accelerating and optimizing the inference and training of diffusion-based video generation models, providing the necessary tools to handle text-to-video and image-to-video tasks with significant VRAM and performance optimizations.
Open-Higgsfield-AI is a generative AI content studio and visual workflow orchestrator. It provides a unified interface for creating photorealistic images and videos, utilizing a node-based editor to chain multiple image, video, and audio models into automated content pipelines. The system functions as an AI video animation tool and local GPU inference engine, allowing users to run generative models on local hardware or remote servers. It includes specialized capabilities for audio-driven lip synchronization and cinematic camera controls to adjust virtual lens and focal settings. The platform
This is a generative AI content studio that provides a visual workflow interface for orchestrating text-to-video and image-to-video generation pipelines, making it a comprehensive tool for managing video creation tasks.
ComfyUI is a modular generative AI workflow orchestrator and node-based GUI for designing and executing complex diffusion model pipelines. It functions as both a visual interface for building generative logic graphs and a programmable backend API that exposes diffusion model operations for external integration. The system distinguishes itself through a graph-based execution model that supports differential workflow execution, re-running only modified nodes to reduce computation. It features dynamic model offloading to manage memory between system RAM and GPU VRAM and utilizes metadata-embedde
ComfyUI is a powerful node-based workflow orchestrator that enables text-to-video and image-to-video generation by integrating various diffusion models, though it functions as a framework for building these pipelines rather than being a single pre-trained generative model itself.
This repository provides a diffusion-based transformer model specifically designed for text-to-video and image-to-video generation, fitting the category of generative video models.
VideoCrafter is a framework for high-quality text-to-video and image-to-video generation based on latent diffusion models, providing the core generative capabilities required for this category.
| 仓库 | Star 数 | 语言 | 许可证 | 最后推送 |
|---|---|---|---|---|
| ailab-cvc/videocrafter | 5.1K | Python | NOASSERTION | |
| zai-org/cogvideo | 12.8K | Python | Apache-2.0 | |
| thudm/cogvideo | 12.8K | Python | Apache-2.0 | |
| guoyww/animatediff | 12.1K | Python | Apache-2.0 | |
| hpcaitech/open-sora | 29.1K | Python | Apache-2.0 | |
| tencent-hunyuan/hunyuanvideo-1.5 | 4.4K | Python | other | |
| pku-yuangroup/open-sora-plan | 12.2K | Python | MIT | |
| wan-video/wan2.1 | 15.4K | Python | apache-2.0 | |
| meituan-longcat/longcat-video | 4.5K | Python | MIT | |
| nvlabs/sana | 8.3K | Python | Apache-2.0 |