awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

22 रिपॉजिटरी

Awesome GitHub RepositoriesVideo Diffusion Models

Diffusion models that generate video by iteratively denoising latent representations in a compressed spatio-temporal space.

Distinct from Latent Diffusion Models: Distinct from Latent Diffusion Models: focuses on video generation with temporal consistency, not general image generation.

Explore 22 awesome GitHub repositories matching artificial intelligence & ml · Video Diffusion Models. Refine with filters or upvote what's useful.

Awesome Video Diffusion Models GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • humanaigc/animateanyoneHumanAIGC का अवतार

    HumanAIGC/AnimateAnyone

    14,774GitHub पर देखें↗

    AnimateAnyone is an appearance-preserving video synthesizer designed for character animation from a single static image. It functions as a diffusion image-to-video generator that transforms a source image into a high-fidelity video sequence while maintaining consistent character identity, clothing, and visual details across all frames. The system enables video-driven character reenactment by transferring motions, facial expressions, and body movements from a reference video onto a static character. It employs pose-guided video generation to control movement via skeleton keypoints and pose sig

    Implements a video generation process based on iterative denoising of latent representations for temporal consistency.

    GitHub पर देखें↗14,774
  • magic-research/magic-animatemagic-research का अवतार

    magic-research/magic-animate

    10,908GitHub पर देखें↗

    Magic Animate is a diffusion model video generator designed for human image animation. It transforms a static human photo into a temporally consistent video by mapping movements from a reference motion clip, acting as a tool to create realistic animations from a single image. The system ensures visual stability and minimizes flicker through temporal attention injection and motion-controlled noise scheduling. To accelerate the generation of high-resolution video, it includes a distributed GPU inference engine that splits model workloads across multiple graphics cards. The project covers a com

    Utilizes a video diffusion model to iteratively denoise latent representations for temporally consistent animation.

    Python
    GitHub पर देखें↗10,908
  • nvidia/cosmosNVIDIA का अवतार

    NVIDIA/cosmos

    10,494GitHub पर देखें↗

    Cosmos is an open platform of world models, datasets, and tools for building physical AI systems such as robots and autonomous vehicles. It provides video generation and video understanding models that can generate synthetic videos and world simulations from text, image, video, or action inputs, and analyze videos to produce captions, event timestamps, spatial bounding boxes, and next-action predictions. The platform includes a world simulation generator that produces images, videos, synchronized audio, and action-conditioned rollouts for synthetic data, alongside a visual content analyzer th

    Converts raw video frames into discrete latent tokens and reconstructs them using a diffusion-based decoder for high-fidelity generation.

    Jupyter Notebook
    GitHub पर देखें↗10,494
  • brycedrennan/imaginairybrycedrennan का अवतार

    brycedrennan/imaginAIry

    8,155GitHub पर देखें↗

    imaginAIry is a system for generating and refining images and videos using diffusion models. It operates as a web-based server that triggers generation requests through standard API calls, allowing for the creation of visuals and video sequences from text prompts or existing files. The project provides a suite for AI image editing and upscaling, enabling the modification of visuals through natural language instructions and super-resolution tools to increase detail and image size. The system includes capabilities for structural image control using depth maps, edge maps, and body poses to main

    Produces moving image sequences and video clips using specialized latent diffusion models.

    Python
    GitHub पर देखें↗8,155
  • humanaigc/emoHumanAIGC का अवतार

    HumanAIGC/EMO

    7,616GitHub पर देखें↗

    EMO is an AI portrait animator and audio-to-video diffusion model designed to generate expressive talking head videos. It transforms a single static portrait image and an audio track into a synchronized video of a person speaking. The system focuses on digital human synthesis, producing high-fidelity facial movements and emotional cues. It synchronizes lip movements and facial gestures to match spoken voice recordings to create realistic portrait animations. The framework utilizes a diffusion process and a cross-modal alignment mechanism to ensure timing between audio signals and visual land

    Implements a diffusion process to generate synchronized video frames from audio features.

    GitHub पर देखें↗7,616
  • hvision-nku/storydiffusionHVision-NKU का अवतार

    HVision-NKU/StoryDiffusion

    6,430GitHub पर देखें↗

    StoryDiffusion is a generative AI system designed for consistent character image and video generation. It utilizes a pluggable cross-attention module to inject shared character representations into pretrained diffusion models, allowing for visual identity stability across multiple images and scenes without retraining the base model. The project features a video generation pipeline that produces temporally coherent sequences from text prompts or condition images. It employs a latent space motion interpolator to predict intermediate frames and semantic motion, enabling long-range video generati

    Uses a latent diffusion model to produce temporally coherent video sequences from text prompts.

    Jupyter Notebook
    GitHub पर देखें↗6,430
  • doubiiu/tooncrafterDoubiiu का अवतार

    Doubiiu/ToonCrafter

    5,972GitHub पर देखें↗

    ToonCrafter is a model that combines latent diffusion, reference-based colorization, and sketch-guided control for cartoon animation and interpolation. It functions as a cartoon video interpolation model, a reference-based colorization model, and a sketch-guided animation tool, all built on a latent diffusion animation framework. The project distinguishes itself by integrating three core capabilities into a single pipeline: generating smooth intermediate frames between two cartoon images using diffusion-based priors, transferring color and style from a reference image onto black-and-white ske

    Generates smooth intermediate cartoon frames by denoising latent representations conditioned on start and end frames.

    Python
    GitHub पर देखें↗5,972
  • aigc-apps/sd-webui-easyphotoaigc-apps का अवतार

    aigc-apps/sd-webui-EasyPhoto

    5,150GitHub पर देखें↗

    यह प्रोजेक्ट एक Stable Diffusion WebUI एक्सटेंशन है जो व्यक्तिगत पोर्ट्रेट जनरेशन और AI फोटो एडिटिंग के लिए एक ग्राफिकल इंटरफेस प्रदान करता है। यह उपयोगकर्ताओं को विशिष्ट लोगों के सुसंगत डिजिटल संस्करण बनाने के लिए अपलोड की गई छवियों के एक छोटे सेट से कस्टम पहचान मॉडल प्रशिक्षित करने की अनुमति देता है। एक्सटेंशन में एक वर्चुअल ट्राई-ऑन सिस्टम शामिल है जो संदर्भ परिधानों को टेम्पलेट निकायों के साथ संरेखित करके छवियों में कपड़ों को बदल देता है। इसमें स्थिर छवियों और वीडियो दोनों में फेस स्वैपिंग के लिए टूल भी शामिल हैं, साथ ही एक पोर्ट्रेट एनिमेटर जो संदर्भ-निर्देशित गति और टेक्स्ट विवरण का उपयोग करके स्थिर छवियों को गतिशील वीडियो में बदल देता है। अतिरिक्त क्षमताएं उम्र और अभिव्यक्ति को समायोजित करने के लिए चेहरे की विशेषता हेरफेर, बहु-व्यक्ति छवि संश्लेषण, और लेटेंट स्पेस इंटरपोलेशन के माध्यम से छवियों के बीच सुचारू संक्रमण के निर्माण को कवर करती हैं।

    Produces motion sequences by applying stable video diffusion models to a starting frame and textual description.

    Python
    GitHub पर देखें↗5,150
  • ailab-cvc/videocrafterailab-cvc का अवतार

    ailab-cvc/videocrafter

    5,063GitHub पर देखें↗

    Videocrafter एक लेटेंट डिफ्यूजन मॉडल है जिसे AI वीडियो सिंथेसिस के लिए बनाया गया है। यह टेक्स्ट-टू-वीडियो और इमेज-टू-वीडियो जनरेशन सिस्टम दोनों के रूप में काम करता है, जो वर्णनात्मक टेक्स्ट प्रॉम्प्ट या स्थिर इमेज इनपुट से उच्च गुणवत्ता वाले वीडियो सीक्वेंस तैयार करता है। यह मॉडल इनपुट को एनिमेटेड कंटेंट में बदलने के लिए डिफ्यूजन-आधारित न्यूरल नेटवर्क का उपयोग करता है, जिससे जनरेट किए गए सीक्वेंस में विजुअल कंसिस्टेंसी और टेम्पोरल कोहेरेंस बनी रहती है। यह कस्टम वीडियो क्लिप बनाने और स्थिर इमेज को फ्लूइड मोशन में एनिमेट करने की सुविधा देता है।

    Implements a video diffusion model that synthesizes high-quality sequences from text or image inputs.

    Python
    GitHub पर देखें↗5,063
  • antgroup/echomimic_v2antgroup का अवतार

    antgroup/echomimic_v2

    4,597GitHub पर देखें↗

    EchoMimic V2 एक AI वीडियो जनरेशन पाइपलाइन और कंप्यूटर विज़न एनिमेशन मॉडल है, जिसे सिंथेटिक ह्यूमन एनिमेशन बनाने के लिए डिज़ाइन किया गया है। यह एक जेनेरेटिव फ्रेमवर्क के रूप में काम करता है जो एक स्थिर रेफरेंस इमेज को ड्राइविंग वीडियो से निकाले गए पोज़ मूवमेंट के साथ अलाइन करके सेमी-बॉडी वीडियो तैयार करता है। यह सिस्टम फ्रेम के बीच स्मूथ ट्रांज़िशन सुनिश्चित करने के लिए डिफ्यूज़न-आधारित जनरेशन प्रक्रिया के साथ लेटेंट स्पेस कम्प्रेशन और टेम्पोरल अटेंशन मैकेनिज्म का उपयोग करता है। यह रेफरेंस-आधारित एन्कोडिंग के माध्यम से व्यक्ति की पहचान को बनाए रखता है और पोज़-ड्रिवन मोशन कंडीशनिंग के जरिए स्थानिक प्लेसमेंट को गाइड करता है। इस प्रोजेक्ट में चेहरे के विवरण और शार्पनेस को बेहतर बनाने के लिए मल्टी-स्टेज इमेज रिफाइनमेंट की क्षमताएं शामिल हैं। यह एनिमेशन डेटासेट तैयार करने के लिए भी टूल्स प्रदान करता है, जिसमें मॉडल ट्रेनिंग और इन्फरेंस के लिए आवश्यक फॉर्मेट में वीडियो डेटा को डाउनलोड और प्रीप्रोसेस करना शामिल है।

    Implements a video diffusion model that generates temporal sequences by denoising latent representations.

    Pythonaudio-driven-body-animationaudio-driven-portrait-animationsaudio-driven-talking-face
    GitHub पर देखें↗4,597
  • meituan-longcat/longcat-videomeituan-longcat का अवतार

    meituan-longcat/LongCat-Video

    4,460GitHub पर देखें↗

    LongCat-Video वीडियो सिंथेसिस के लिए विशेष मॉडलों का एक संग्रह है, जिसमें टेक्स्ट, छवियों या मौजूदा अनुक्रमों से उच्च-रिज़ॉल्यूशन वीडियो बनाने के लिए एक बड़े भाषा मॉडल-आधारित आर्किटेक्चर की सुविधा है। इसमें टेक्स्ट-टू-वीडियो जनरेशन, इमेज-टू-वीडियो एनीमेशन और टॉकिंग अवतार बनाने के लिए समर्पित सिस्टम शामिल हैं। यह प्रोजेक्ट एक वीडियो निरंतरता मॉडल के माध्यम से मौजूदा क्लिप की लंबाई बढ़ाने के लिए विशिष्ट क्षमताएं प्रदान करता है जो बाद के फ्रेम की भविष्यवाणी करता है। यह बोलने वाले वीडियो बनाने के लिए ऑडियो और टेक्स्ट प्रॉम्प्ट के साथ चरित्र के होंठों की गतिविधियों के सिंक्रोनाइज़ेशन को भी सक्षम बनाता है। सिस्टम जनरेशन दक्षता को प्रबंधित करने के लिए विभिन्न ऑप्टिमाइज़ेशन तकनीकों को शामिल करता है, जिसमें मेमोरी उपयोग और इन्फरेंस लेटेंसी को कम करने के लिए डिस्टिलेशन-आधारित सैंपलिंग और क्वांटाइजेशन शामिल है। अतिरिक्त संरचनात्मक घटक समय और स्थान के साथ निरंतरता बनाए रखने के लिए लेटेंट-स्पेस कम्प्रेशन और स्थानिक-अस्थायी मॉडलिंग को कवर करते हैं।

    Uses video diffusion models to transform random noise into high-resolution video sequences.

    Python
    GitHub पर देखें↗4,460
  • tencent-hunyuan/hunyuanvideo-1.5Tencent-Hunyuan का अवतार

    Tencent-Hunyuan/HunyuanVideo-1.5

    4,440GitHub पर देखें↗

    HunyuanVideo-1.5 is a video generation foundation model and text-to-video diffusion framework. It utilizes a latent video diffusion model and a spatio-temporal transformer architecture to generate high-definition video sequences from text descriptions and images. The project enables cinematic camera control for directing pans and tilts and provides image-to-video animation capabilities. It supports visual style adaptation through low-rank adaptation tuning and uses a language model for prompt refinement to improve visual alignment. The model covers high-resolution video upscaling via a super

    Provides the core latent video diffusion model that generates high-definition video sequences from text descriptions.

    Pythonimage-to-videotext-to-videovideo-generation
    GitHub पर देखें↗4,440
  • showlab/tune-a-videoshowlab का अवतार

    showlab/Tune-A-Video

    4,364GitHub पर देखें↗

    Tune-A-Video is a text-to-video diffusion framework designed to convert pretrained text-to-image diffusion models into video generators. It utilizes a spatio-temporal attention mechanism and single text-video pair training to enable the synthesis of moving sequences from text prompts. The project provides tools for one-shot video personalization, allowing a model to be tuned on a single reference video to preserve specific characters or artistic styles across new generations. It also functions as a video editor that modifies subjects, backgrounds, and styles through noise-sampling prompt guid

    Converts text-to-image diffusion models into video generators through training on specific sequences.

    Python
    GitHub पर देखें↗4,364
  • antgroup/echomimicantgroup का अवतार

    antgroup/echomimic

    4,255GitHub पर देखें↗

    EchoMimic is a multimodal human animation framework and diffusion-based video generator. It produces lifelike facial and semi-body animations of a reference image by synthesizing motion and appearance from various source data. The system enables portrait animation driven by audio, pose sequences, or driver videos. It features a landmark conditioning tool that allows for the precise control of facial movements by modifying specific landmark points. The framework covers multi-modal motion synthesis and the synchronization of reference images to match the physical movements of a target driver.

    Uses a diffusion-based generative model to produce high-quality video sequences from multimodal source data.

    Pythonaaai2025audio-driven-portrait-animationsaudio-driven-talking-face
    GitHub पर देखें↗4,255
  • fudan-generative-vision/champfudan-generative-vision का अवतार

    fudan-generative-vision/champ

    4,253GitHub पर देखें↗

    Champ is a generative vision system and controllable image-to-video generator designed for human image animation. It uses a diffusion-based video synthesizer and 3D parametric guidance to transform a single reference image into a consistent sequence of motion based on external driving data. The framework distinguishes itself through a human pose transfer system that employs 3D body parametric extraction and coordinate-space alignment. This allows the model to map motion from a driving video to a reference person by adjusting for body scales and camera perspectives using depth and semantic con

    Uses a video diffusion model to generate temporally consistent human animations.

    Pythonhuman-animationimage-animatiolnvideo-generation
    GitHub पर देखें↗4,253
  • badtobest/echomimicBadToBest का अवतार

    BadToBest/EchoMimic

    4,258GitHub पर देखें↗

    EchoMimic is an audio-driven portrait animation framework and latent diffusion video generator. It transforms static reference images into dynamic talking head videos by synchronizing facial movements with audio tracks and motion drivers. The system functions as a hybrid motion synthesis engine that combines audio inputs and pose data. It utilizes a facial landmark motion controller to edit positioning markers, enabling precise synchronization and video-to-video pose transfer. The pipeline covers image-to-video animation through latent diffusion and facial landmark conditioning. This allows

    Implements a video generation system that iteratively denoises latent representations to produce high-fidelity human animations.

    Python
    GitHub पर देखें↗4,258
  • picsart-ai-research/text2video-zeroPicsart-AI-Research का अवतार

    Picsart-AI-Research/Text2Video-Zero

    4,244GitHub पर देखें↗

    Text2Video-Zero is a text-to-video diffusion model and framework designed to synthesize temporally consistent video sequences from textual prompts. It functions as a zero-shot video generator, repurposing pre-trained image diffusion models to create video content without requiring additional training on video datasets. The system includes a conditional video synthesizer that allows for guided generation using depth, edge, or pose maps to control structural layout and movement. It also provides text-based video editing capabilities to modify the style or content of existing video clips through

    Repurposes pre-trained image diffusion models to generate temporally consistent video frames without needing video-specific training data.

    Pythonvideo-editingvideo-generation
    GitHub पर देखें↗4,244
  • hao-ai-lab/fastvideohao-ai-lab का अवतार

    hao-ai-lab/FastVideo

    3,743GitHub पर देखें↗

    FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine, a video diffusion training framework, and a modular pipeline orchestrator. It provides a distributed transformer optimizer and a distillation toolkit designed to reduce denoising steps and model complexity to increase frame rates. The project distinguishes itself through specialized acceleration techniques, including joint distillation and sparse attention training. It implements low-step video generation and weight quantization to FP8 or FP4 precision to increase throughput a

    Provides full and LoRA training methods to adapt video diffusion models to specific styles or tasks.

    Pythondiffusersdiffusion-modelsdistillation
    GitHub पर देखें↗3,743
  • thu-ml/turbodiffusionthu-ml का अवतार

    thu-ml/TurboDiffusion

    3,339GitHub पर देखें↗

    TurboDiffusion is a video diffusion inference engine and generator designed to create high-resolution videos from text prompts and images. It provides a runtime environment for executing optimized diffusion model checkpoints with a focus on reducing latency and GPU memory usage. The project features a specialized training framework for aligning sparse-linear attention models with pretrained full-attention models. This system includes capabilities for sparse attention parameter merging and sparse-linear model alignment to reduce computational costs during inference while maintaining output qua

    A high-resolution video generator that uses optimized diffusion model checkpoints for text-to-video and image-to-video synthesis.

    Pythonai-infraconsistency-modeldiffusion-models
    GitHub पर देखें↗3,339
  • robbyant/lingbot-worldRobbyant का अवतार

    Robbyant/lingbot-world

    2,915GitHub पर देखें↗

    Lingbot-world is an interactive world simulator and framework for generating high-fidelity video environments from text and image prompts. It functions as a video generation system designed to create controllable simulations for applications such as robotics learning and gaming. The project includes a video motion controller that directs camera and object movement using transformation matrices and action strings. It utilizes a quantized inference engine to reduce memory usage and accelerate the generation of video sequences. The system covers a range of optimization techniques, including fou

    Utilizes video diffusion models to generate high-fidelity environments by iteratively denoising latent representations.

    Pythonaigcimage-to-videolingbot-world
    GitHub पर देखें↗2,915
पिछला12अगला
  1. Home
  2. Artificial Intelligence & ML
  3. Generative AI Resources
  4. Diffusion & Visual Synthesis Models
  5. Generative Models
  6. Latent Diffusion Models
  7. Video Diffusion Models

सब-टैग एक्सप्लोर करें

  • Cartoon Frame GeneratorsApplies diffusion priors specifically to produce coherent intermediate frames in cartoon animation workflows. **Distinct from Video Diffusion Models:** Distinct from Video Diffusion Models: specialized for cartoon-style interpolation rather than general video generation.
  • Conditioned Frame InterpolatorsGenerates intermediate frames by denoising a latent representation conditioned on start and end frames using a pretrained diffusion model. **Distinct from Video Diffusion Models:** Distinct from Video Diffusion Models: specifically conditions on start and end frames for interpolation rather than generating from noise.
  • Finetuning FrameworksTools and methodologies for adapting pretrained video diffusion models to specific styles or tasks. **Distinct from Video Diffusion Models:** Focuses on the act of adapting weights via LoRA or full training, not the model architecture itself
  • Image-to-Video Model AdaptationProcesses for converting static image diffusion models into temporal video generators. **Distinct from Video Diffusion Models:** Focuses on the adaptation process of the model weights rather than the resulting video diffusion model identity.
  • Latent Diffusion Frame InterpolatorsGenerates intermediate video frames by denoising latent representations conditioned on start and end frames. **Distinct from Video Diffusion Models:** Distinct from Video Diffusion Models: focuses on interpolation between given frames rather than unconditional video generation.
  • Video TokenizersComponents that convert raw video frames into discrete latent tokens for efficient processing and generation. **Distinct from Video Diffusion Models:** Distinct from Video Diffusion Models: focuses on the tokenization step that precedes diffusion-based generation.
  • Zero-Shot AdaptationTechniques for repurposing pre-trained image models for video generation without additional video-dataset training. **Distinct from Video Diffusion Models:** Focuses on the zero-shot repurposing of image models for video, whereas Video Diffusion Models is the general architecture.