For an open source alternative to Sora for video generation, the strongest matches are meituan-longcat/longcat-video (LongCat-Video provides open-source models and Python-based tools for text-to-video), lucidrains/video-diffusion-pytorch (This project is a PyTorch-based framework specifically designed for) and zai-org/cogvideo (CogVideo is an open-source text-to-video and image-to-video generation framework). genmoai/mochi and ailab-cvc/videocrafter round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
We curate open-source GitHub repositories matching “open source alternatives to Sora AI video generator”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.
LongCat-Video is a collection of specialized models for video synthesis, featuring a large language model based architecture for creating high-resolution videos from text, images, or existing sequences. It includes dedicated systems for text-to-video generation, image-to-video animation, and the creation of talking avatars. The project provides specific capabilities for extending the length of existing clips through a video continuation model that predicts subsequent frames. It also enables the synchronization of character lip movements with audio and text prompts to produce speaking videos.
LongCat-Video provides open-source models and Python-based tools for text-to-video generation using diffusion-based video synthesis, making it a great fit for your AI-driven video creation needs.
This project is a research-oriented PyTorch framework designed for the implementation and training of generative video diffusion models. It provides a modular toolkit that extends standard image-based diffusion techniques into three dimensions, enabling the synthesis of coherent video sequences through iterative denoising processes. The framework distinguishes itself by utilizing factored space-time attention, which decomposes high-dimensional video data into separate spatial and temporal layers to maintain motion consistency while managing computational complexity. It supports multi-modal tr
This project is a PyTorch-based framework specifically designed for training video diffusion models from text and other conditioning inputs, making it a fitting open-source tool for AI video generation research.
CogVideo is a video generation framework and large language model architecture designed for synthesizing high-resolution video clips from natural language descriptions and images. It functions as a text-to-video and image-to-video generator, while also providing a model for video captioning to analyze visual content into descriptive text summaries. The system supports animating static images into motion sequences and transforming series of images into video based on prompts. It includes capabilities for extending the length of generated video clips to create longer sequences of motion. The f
CogVideo is an open-source text-to-video and image-to-video generation framework built on a latent diffusion model architecture with PyTorch, fulfilling all the core requirements for AI-driven video synthesis.
Mochi is an open-source text-to-video diffusion model designed to synthesize high-fidelity video sequences from natural language prompts. It utilizes a diffusion transformer architecture to generate temporal video data. The project includes a framework for low-rank adaptation, allowing the model to be fine-tuned on custom datasets to specialize visual styles or specific subjects. It also features a distributed inference engine that spreads model workloads across multiple graphics cards to increase memory capacity and processing speed. The system covers programmable video generation through a
Mochi is an open-source text-to-video diffusion model implemented in PyTorch that delivers high-fidelity video generation from text prompts with multi-GPU acceleration and fine-tuning support.
Videocrafter is a latent diffusion model designed for AI video synthesis. It functions as both a text-to-video and image-to-video generation system, synthesizing high-quality video sequences from descriptive text prompts or static image inputs. The model utilizes a diffusion-based neural network to transform inputs into animated content, ensuring visual consistency and temporal coherence throughout the generated sequences. This allows for the creation of custom video clips and the animation of static images into fluid motion.
This repository provides a Python and PyTorch-based latent diffusion model for text-to-video generation with temporal attention, though it lacks some broader platform features to be a full-scale generation suite.
CogVideo is a generative video framework that uses diffusion models and transformer-based architectures to synthesize high-resolution video clips. It functions as both a text-to-video and image-to-video generator, converting textual descriptions or static images into temporal visual sequences. The system integrates large language model capabilities to expand short user prompts into detailed descriptions for better visual alignment. It supports the animation of static images through latent seeding and provides the ability to extend the length of existing video sequences. The project includes
CogVideo is an open-source text-to-video and image-to-video generation framework built on a PyTorch and diffusion model architecture, matching the core requirements for AI-driven video synthesis.
Wan2.2 is a generative video artificial intelligence system designed to synthesize visual media by interpreting natural language instructions. It functions as a text-to-video diffusion model that transforms written concepts into coherent motion sequences through deep learning and latent space manipulation. The system utilizes a transformer-based architecture to process video data as a series of tokens, allowing it to capture complex spatial and temporal relationships. By employing a temporal attention mechanism, the model maintains visual consistency across frames, while its latent space appr
Wan2.2 is an open-source text-to-video diffusion model built in Python and PyTorch that uses transformer-based temporal attention to generate temporally consistent video sequences from text prompts.
HunyuanVideo-1.5 is a video generation foundation model and text-to-video diffusion framework. It utilizes a latent video diffusion model and a spatio-temporal transformer architecture to generate high-definition video sequences from text descriptions and images. The project enables cinematic camera control for directing pans and tilts and provides image-to-video animation capabilities. It supports visual style adaptation through low-rank adaptation tuning and uses a language model for prompt refinement to improve visual alignment. The model covers high-resolution video upscaling via a super
HunyuanVideo-1.5 is a Python-based text-to-video diffusion framework and foundation model featuring spatio-temporal transformers, camera control, and open weights that directly match this search.
Open-Sora is a video generation framework designed to produce cinematic sequences from text prompts and images. It functions as a generative system that transforms written descriptions or reference images into video content featuring realistic textures and lighting. The project includes a dedicated prompt engineering tool that uses large language models to expand simple user inputs into detailed descriptions. It also features a motion controller for adjusting movement intensity in generated sequences and evaluating motion levels in existing video files. The framework incorporates text-to-vid
Open-Sora is a Python-based generative video framework built on diffusion transformers and latent diffusion models to provide open-weights text-to-video generation with GPU acceleration and spatiotemporal handling.
Text2Video-Zero is a text-to-video diffusion model and framework designed to synthesize temporally consistent video sequences from textual prompts. It functions as a zero-shot video generator, repurposing pre-trained image diffusion models to create video content without requiring additional training on video datasets. The system includes a conditional video synthesizer that allows for guided generation using depth, edge, or pose maps to control structural layout and movement. It also provides text-based video editing capabilities to modify the style or content of existing video clips through
This repository provides a zero-shot text-to-video diffusion framework built in Python for generating temporally consistent video sequences from text prompts using pre-trained models.
Open-Sora-Plan is a text-to-video framework and distributed video training system. It utilizes a diffusion transformer architecture and large language model components to transform written descriptions or image prompts into high-quality video sequences. The system features a distributed infrastructure designed for large-scale video training and inference. It employs sequence parallelism to split high-resolution or long-duration video samples across multiple GPUs and uses a sparse attention mechanism to increase processing speed. The project includes capabilities for both text-to-video and im
Open-Sora-Plan is a diffusion-transformer-based text-to-video framework written in Python and PyTorch that provides open model weights, GPU-accelerated distributed training, and temporal consistency features, which directly matches your search for an AI video generation tool.
TurboDiffusion is a video diffusion inference engine and generator designed to create high-resolution videos from text prompts and images. It provides a runtime environment for executing optimized diffusion model checkpoints with a focus on reducing latency and GPU memory usage. The project features a specialized training framework for aligning sparse-linear attention models with pretrained full-attention models. This system includes capabilities for sparse attention parameter merging and sparse-linear model alignment to reduce computational costs during inference while maintaining output qua
TurboDiffusion is an inference engine and text-to-video generator built with a Python codebase, providing optimized diffusion model execution and GPU acceleration, though it focuses more heavily on inference acceleration and distillation than being a complete end-to-end model training suite.
Tune-A-Video is a text-to-video diffusion framework designed to convert pretrained text-to-image diffusion models into video generators. It utilizes a spatio-temporal attention mechanism and single text-video pair training to enable the synthesis of moving sequences from text prompts. The project provides tools for one-shot video personalization, allowing a model to be tuned on a single reference video to preserve specific characters or artistic styles across new generations. It also functions as a video editor that modifies subjects, backgrounds, and styles through noise-sampling prompt guid
Tune-A-Video is a PyTorch-based text-to-video diffusion framework that converts image models into video generators using spatio-temporal attention, though it focuses primarily on one-shot personalization rather than general text-to-video generation from scratch.
ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat
This repository provides a custom extension and node-based framework for generating videos using latent diffusion models from text prompts, though it operates as an extension within ComfyUI rather than a standalone application.
This is a PyTorch-based implementation of diffusion models for synthesizing photorealistic images and video. It provides a framework for text-to-image and text-to-video generation, as well as unconditional image synthesis. The system utilizes a cascading diffusion pipeline to produce high-resolution imagery by passing low-resolution outputs through a sequence of super-resolution models. It also includes capabilities for image inpainting, allowing the reconstruction of masked or missing regions of visual media guided by surrounding context and text prompts. The project includes tools for diff
This PyTorch implementation of diffusion models supports text-to-video generation and GPU acceleration, though it focuses more heavily on image synthesis architectures rather than offering an end-to-end video pipeline.
AnimateDiff is a latent diffusion video generator and text-to-video diffusion framework. It converts existing text-to-image diffusion models into animation generators by applying specialized motion modules, allowing for the creation of video sequences without modifying the original base model. The project provides an image-to-video animation framework that uses sparse RGB images, sketches, or structural keyframe constraints to guide generation. It further distinguishes itself with a motion adapter system that injects cinematic camera movements, such as zooming, panning, and tilting, into anim
AnimateDiff is a latent diffusion-based text-to-video and animation framework built in Python and PyTorch that uses motion adapters to convert text-to-image models into video generators, making it a fitting tool for this search despite lacking some standalone end-to-end model weights out of the box.
imaginAIry is a system for generating and refining images and videos using diffusion models. It operates as a web-based server that triggers generation requests through standard API calls, allowing for the creation of visuals and video sequences from text prompts or existing files. The project provides a suite for AI image editing and upscaling, enabling the modification of visuals through natural language instructions and super-resolution tools to increase detail and image size. The system includes capabilities for structural image control using depth maps, edge maps, and body poses to main
This repository provides a Python-based diffusion model framework capable of generating video sequences from text prompts and structural controls, making it a relevant tool for AI-driven video generation though primarily focused on image workflows.
Diffusers is a PyTorch-based library and generative AI framework used to build, train, and deploy diffusion pipelines for producing multi-modal media. It provides a suite of tools for generating images, video, and audio from natural language descriptions, as well as specialized systems for text-to-image generation. The project differentiates itself through a modular architecture that separates noise schedulers, pretrained model blocks, and pipeline compositions. This structure allows for the construction of custom generation workflows and the ability to swap individual components of the diffu
Diffusers is a PyTorch-based framework specifically designed for building and deploying diffusion pipelines, offering the exact text-to-video capabilities, open models, and GPU-accelerated infrastructure the visitor is looking for.
StoryDiffusion is a generative AI system designed for consistent character image and video generation. It utilizes a pluggable cross-attention module to inject shared character representations into pretrained diffusion models, allowing for visual identity stability across multiple images and scenes without retraining the base model. The project features a video generation pipeline that produces temporally coherent sequences from text prompts or condition images. It employs a latent space motion interpolator to predict intermediate frames and semantic motion, enabling long-range video generati
StoryDiffusion is an open-source generative AI system based on diffusion models that supports text-to-video generation with character consistency and temporal interpolation using a Python codebase.
This project is a Stable Diffusion video generator that creates moving imagery by interpolating between text prompts within a generative model's latent space. It functions as a tool for AI video generation and latent space interpolation, transforming descriptive text into visual sequences. The system specifically enables audio-reactive visuals by synchronizing the rate of image interpolation to the beat and rhythm of an audio file. It produces these sequences through morphing video generation, which transitions smoothly between different text prompts. The project includes a graphical user in
This Python-based tool uses stable diffusion and latent space interpolation to generate videos from text prompts, making it a fitting option for AI video creation though it focuses on interpolation workflows rather than full text-to-video diffusion pipelines.
Sygil-webui is a web interface for Stable Diffusion latent diffusion models, providing a creative suite for text-to-image and text-to-video synthesis. It functions as an image generation tool and a latent diffusion image editor, allowing users to create visuals and video sequences from textual descriptions. The project includes a dedicated model training interface for creating custom textual inversion embeddings, which introduces specific new concepts or styles into the diffusion models. It also features specialized tools for generative image editing, including mask-based inpainting, image-to
Sygil-webui is a Python-based web interface for diffusion models that supports text-to-video synthesis alongside its primary image generation features, though it lacks dedicated open-source video generation models of its own.
ComfyUI is a node-based generative AI orchestration engine designed for constructing, testing, and executing complex image and video synthesis pipelines. By utilizing a directed acyclic graph execution model, the platform allows users to build reproducible workflows through modular, interconnected processing blocks without requiring manual code implementation. It serves as both a local environment for high-performance model inference and a production-ready server for deploying generative capabilities. The platform distinguishes itself through its focus on workflow portability and extensibilit
ComfyUI is a node-based generative AI orchestration platform that natively supports text-to-video generation pipelines using Python, PyTorch, and GPU acceleration, though it functions as a modular workflow engine rather than a single end-to-end video model.
FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine, a video diffusion training framework, and a modular pipeline orchestrator. It provides a distributed transformer optimizer and a distillation toolkit designed to reduce denoising steps and model complexity to increase frame rates. The project distinguishes itself through specialized acceleration techniques, including joint distillation and sparse attention training. It implements low-step video generation and weight quantization to FP8 or FP4 precision to increase throughput a
FastVideo is an inference engine and training framework built specifically for accelerating video diffusion models in PyTorch, matching the required technical stack while focusing on speed optimization rather than providing standalone end-user generation models.
ComfyUI is a modular generative AI workflow orchestrator and node-based GUI for designing and executing complex diffusion model pipelines. It functions as both a visual interface for building generative logic graphs and a programmable backend API that exposes diffusion model operations for external integration. The system distinguishes itself through a graph-based execution model that supports differential workflow execution, re-running only modified nodes to reduce computation. It features dynamic model offloading to manage memory between system RAM and GPU VRAM and utilizes metadata-embedde
ComfyUI is a modular node-based workflow orchestrator that supports diffusion-based text-to-video generation pipelines and Python-driven GPU execution, though it is a general orchestrator rather than a dedicated video generation model itself.
mmagic is a multimodal training pipeline and framework for generative AI, focusing on visual synthesis and restoration. It provides the infrastructure to build and train models for tasks such as text-to-image and text-to-video generation, 3D-aware content synthesis, and high-fidelity image translation using diffusion models and generative adversarial networks. The project distinguishes itself through specialized capabilities for generative model personalization, including techniques for fine-tuning subjects and styles. It also supports advanced visual manipulations such as latent space interp
It provides a PyTorch-based training pipeline and infrastructure for building diffusion-based text-to-video generation models, though it is a general generative framework rather than a pre-packaged ready-to-run video generation tool.
SEINE is an open-source diffusion-based video generation model capable of text-to-video synthesis and video transition generation using PyTorch and GPU acceleration, though it lacks a full turnkey application wrapper.
This project provides a deep learning framework for synthesizing video content from text prompts. It functions as a generative video artificial intelligence model that utilizes latent diffusion sampling to iteratively refine noise into coherent visual sequences. The architecture is built on a modular design that separates spatial and temporal processing, allowing the system to handle both static images and video sequences within a unified training pipeline. By employing spatiotemporal convolutional layers and temporal attention mechanisms, the model maintains visual consistency and fluid moti
This repository provides a PyTorch implementation of the Make-A-Video text-to-video generation architecture, making it a relevant model implementation for this search even though it lacks a full ready-to-run product wrapper.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| meituan-longcat/longcat-video | 4.5K | Python | MIT | |
| lucidrains/video-diffusion-pytorch | 1.4K | Python | MIT | |
| zai-org/cogvideo | 12.8K | Python | Apache-2.0 | |
| genmoai/mochi | 3.7K | Python | Apache-2.0 | |
| ailab-cvc/videocrafter | 5.1K | Python | NOASSERTION | |
| thudm/cogvideo | 12.8K | Python | Apache-2.0 | |
| wan-video/wan2.2 | 14.3K | Python | apache-2.0 | |
| tencent-hunyuan/hunyuanvideo-1.5 | 4.4K | Python | other | |
| hpcaitech/open-sora | 29.1K | Python | Apache-2.0 | |
| picsart-ai-research/text2video-zero | 4.2K | Python | NOASSERTION |